Alternative data datasets
Browse by provider, category, marketplace, access, delivery and license
A sharded Parquet archive of NSE equities and index prices from India spanning 2000 to 2026, covering more than 2,500 tickers and organized into roughly 1.5 GB files for efficient streaming.
credit_card_transactions is a small tabular dataset published by aegisheld containing anonymized customer-level credit card account and spending summaries. It is distributed as CSV with a single training split of 8,950 rows.
S&P 500 earnings-call transcripts in Parquet under MIT, tens of thousands of calls. Independently compiled and comparable to other public transcript corpora for sentiment and tone work.
A collection of earnings call transcripts covering S&P 500 firms and other U.S. large-cap companies across the years 2005 through 2025, intended for financial analysis, NLP modeling, and sentiment research.
A small text-classification dataset by ClarusC64 that labels whether borrow-rate, short-interest and price signals cohere into a real short squeeze or diverge into a false signal. It is published on a model hub with an MIT license tag and a 10-row train split.
Indian equity market price history covering NSE-listed stocks and indices from 2000 through 2026, offered in Parquet shards of roughly 1.5GB each for efficient streaming on the Hugging Face hub. The collection draws on more than 2,500 tickers and supports both minute-level and end-of-day intervals.
View source information and access options.
Quarterly 10-Q filings parsed from EDGAR into JSON, millions of records in English (MIT). The quarterly companion to the 10-K corpus for tracking sequential changes in operating performance and risk factors.
Earnings call transcripts for S&P 500 and other large U.S. companies covering 2005 through 2025, useful for financial research, natural language processing, and sentiment analysis.
A collection of 2,000 Malaysian property transaction records spanning every state offers a detailed look at the national housing market in 2025. Compiled from Brickz, a recognized source for real estate transaction data, it covers location, tenure, property type, median pricing, and transaction volumes.
10-K filings for S&P 500 companies retrieved from EDGAR, stored as Parquet with the filing text attached (MIT). A focused corpus for the large-cap annual disclosure cycle: risk factors, MD&A and segment detail for every S&P 500 name.
Earnings Call LLM Insights 📚 Read the Full Story: For a deep dive into the methodology, the wildest moments we found, and key takeaways, check out the blog post:KnowTrend.ai: Auto-Grading Ten Years of Earnings Calls for Prescience and Delusion This dataset contains LLM-generated analysis of ~70,000+ earnings call transcripts. The analysis was performed using Kimi k2-0905-preview, focusing on extra
Earnings-call transcripts for S&P 500 companies, tens of thousands of calls in English with timestamps, in Parquet (MIT). Direct input for call-tone and question-and-answer sentiment analysis across the large-cap earnings cycle.
A collection featuring the primary independent claim of United States utility patents issued between 2005 and 2025 is provided. Because the opening claim usually establishes the widest legal boundaries of an invention, it carries particular significance, and each entry includes the patent identifier alongside its claim text.
FinSight gathers plain-text 10-K and 10-Q SEC filings for twenty major US public firms spanning six sectors, yielding 97 records intended for BERT fine-tuning, retrieval-augmented generation, and multi-agent financial research.
Designed to aid Nifty50 trading research, this dataset merges financial news, sentiment indicators, and stock market data into a single analysis-ready resource, enabling pipelines that cover news cleaning, impact categorization, FinBERT-based sentiment scoring, and Temporal Fusion Transformer forecasting.
A daily record of trading activity across multiple major global stock market indices, structured for trend analysis and forecasting work in international equity markets.
Minute- and daily-resolution price histories for more than 2,500 National Stock Exchange securities and indices from 2000 through 2026 are consolidated into roughly 1.5 gigabyte Parquet shards to enable rapid streaming on Hugging Face infrastructure.
clinicaltrials.gov-summary_and_eligibility is a small Parquet mirror of selected ClinicalTrials.gov records, distributed on the Hugging Face Hub under the rjac publisher. It packages study identifiers, recruitment status, titles, brief summaries, and eligibility criteria into a single tidy table of roughly three thousand rows.
Indian stock market price history spanning 2000 to 2026 for over 2,500 NSE-listed stocks and indices, packaged in large Parquet shards around 1.5 GB each for efficient streaming.
A resource built for forecasting patent approval outcomes, introduced in research employing domain-specific fine-grained claim dependency graphs, with further documentation to be supplied later.
stock_market_dataset is a Hugging Face community dataset published by SkyWalkertT1 that aggregates short Turkish-language market commentary paired with sentiment labels. It is distributed under an MIT-licensed CSV format and is intended for text-classification and token-classification tasks rather than live trading signals.
foot_traffic is a small MIT-licensed tabular dataset on Hugging Face that records hourly pedestrian counts alongside weather conditions. It is published by user supersam7 and contains roughly 11,000 rows.
Vector embeddings of Apple patent text totalling about 4 million entries, stored as parquet shards with a 768-dimensional float64 feature per record.
Indian Domestic Airline Flights 2018-2025 (TsFile) Apache TsFile version of Gokul99400/IndianDomesticAirlineDataset. Overview Indian domestic airline schedule records covering 2018-2025 across ~100 airports: airline, flight number, route, operating days of week, scheduled departure/arrival times and the schedule validity window (validFrom..validTo). The repo ships two CSVs: Air-Clean.csv (33,734 d
View source information and access options.
The linkedin-job-postings dataset, published by xanderios on Hugging Face, contains a single CSV table of U.S. LinkedIn job listings with about 33,000 rows. It is tagged with a MIT license on the platform, though the underlying LinkedIn terms of service place restrictions on redistribution and use.
Minute-level open-high-low-close-volume bars for Indian equities, hundreds of millions of rows in Parquet under MIT. Gives high-frequency microstructure across the Indian cash market for intraday and market-on-close strategies.