Alternative data datasets
Browse by provider, category, marketplace, access, delivery and license
Parsed eligibility criteria from 13,229 ClinicalTrials.gov studies representing the candidate pool gathered by two first-stage retrievers during TREC Clinical Trials 2021–2023, formatted as typed entity-relation graphs for reranking models.
The clinical-trials-patient-graphs dataset is a small tabular collection of patient-level clinical records distributed in Parquet format by the 2001jdev publisher. It organizes entities and relations extracted from trial data into structured rows spanning 2021 through 2023.
clinical-trials-trec-qrels is a tabular relevance-judgment file distributed by the 2001jdev user on the Hugging Face Hub. It maps clinical-trial topic identifiers to NCT registry IDs with graded relevance scores.
The clinical-trials-trec-topics dataset is a small tabular collection of clinical-trial topic records distributed in Parquet format on the Hugging Face Hub. It appears to be derived from TREC clinical-trial retrieval topics, with rows partitioned by year.
The job-postings-raw dataset is a large-scale tabular collection of scraped job postings distributed in Parquet format. It aggregates records across multiple sources and locales, with US coverage indicated in the dataset metadata.
Comprehensive E-Commerce Sales & Customer Analytics Dataset
Global Maritime Chokepoints: Oil Transport & Geopolitical Risk Forecast(2000-36)
Digital habits of Gen Z and their influence on stress and mood levels.
Washington State Electric Vehicle Demographic Data
A sharded Parquet archive of NSE equities and index prices from India spanning 2000 to 2026, covering more than 2,500 tickers and organized into roughly 1.5 GB files for efficient streaming.
credit_card_transactions is a small tabular dataset published by aegisheld containing anonymized customer-level credit card account and spending summaries. It is distributed as CSV with a single training split of 8,950 rows.
Real Estate listings (2.2M+) in the US broken by State and zip code
About 1.5 million US patent claims split into training and test partitions in CSV, organized for claim-level classification and summarization work (Apache 2.0). A large, ready-split corpus for IP-claims modeling.
Historic electricity consumption in the UK (National Grid) between 2009 and 2024
The kl3m-index-edgar-filings dataset, published by ALEA Institute, is a Parquet-format index of SEC EDGAR filings distributed via the Hub datasets library. It contains roughly 20 million tabular records spanning the 10M–100M size category, with an unstated license.
The KL3M Index of EDGAR 10-K Filings is a tabular index published by the Alea Institute that catalogs SEC EDGAR filings with associated issuer metadata. It is distributed in Parquet format and is tagged for use with pandas, polars, and the mlcroissant dataset library.
The kl3m-index-edgar-filings-8-k dataset is an indexed catalog of Form 8-K filings from the SEC's EDGAR system, published by the ALEA Institute. It provides structured metadata for over 1.8 million corporate event reports distributed in Parquet format.
U.S. Electricity Prices & Sales To Consumers (2001 - Now)
Natural Gas Consumption By End Use In The USA From The Past 10+ Years
A static tabular extension of the Bose345/sp500_earnings_transcripts collection covering the same 2005 to 2025 calendar of S&P 500 earnings events, where each row represents a single company-quarter call identified by a stable episode_id and bundles the full transcript along with related SEC press materials and pre-event inputs suitable for supervised learning or reinforcement-style experimentation.
Geospatial Telemetry, PM2.5 Epidemiology & Disaster Econometrics
Exploring the Impact of product positioning on sales and consumer behavior
SEC insider transactions 2006-2026 with stock & insider metadata
EDGAR_FILINGS_DATASET_2016_2021 is a Hugging Face mirror of parsed SEC EDGAR filings spanning 2016 through 2021, distributed as a single train split of about 6 million rows in Parquet format. It is published by anonymous-md with an unstated license.
EDGAR_FILINGS_DATASET_2022_2026H1 is a Hugging Face dataset that compiles parsed SEC EDGAR filings into a single tabular corpus. It covers roughly 2022 through the first half of 2026 and is distributed as Parquet with about 1.7 million records.
5 years and 200k building permits
23 tables · 2.53M rows · 6 domains · grounded in real hospitality industry bench
ToP 250 movies on imdb in 2026 till 31st July
Indian Agriculture Crop Production, Area and Yield 1998 to 2020-21
A curated dataset of companies listed on the London Stock Exchange
Brasil real estate Dataset For Prediction
Maximize agricultural yield by recommending appropriate crops
2022 vending machine sales data from different locations in Central New Jersey
Gender, Location, and Transaction Trends
Baku apartment prices for 2023 year
Released by Groundwork Data, this v1 release is the inaugural openly available structured compilation of Abuja residential property prices, capturing both rental and sale listings across 16 districts within Nigeria's Federal Capital Territory.
The airline-otp-data dataset is a large-scale, public-style tabular repository of U.S. domestic flight records covering on-time performance, delay attribution and cancellation outcomes. Published on Hugging Face under an unstated license, it contains roughly 30 million rows in CSV format and is aimed at analysts working with airline operations data.
Global Hydropower, Wind, Solar, Biofuel & Geothermal Renewable Energy Dataset.
Fed · ECB · BoE · RBI Policy Rates · CPI Inflation · FRED + World Bank
Imbalanced dataset for the classification problem of electricity theft detection
42,302 establishments across 714 US public companies. OSHA, WHD, NLRB, EPA,etc
Daily AI startup funding intelligence and hiring signals dataset, collected from
Google Store Reviews of Tiktok NONLITE/ FULLVERSION application
Google Store Reviews of Tiktok LITE application
Extracted from SEC EDGAR DEF 14A proxy filings since 2015, the dataset contains over 500,000 structured executive compensation entries for S&P 500 and Russell 2000 issuers, capturing CEO and CFO pay, equity awards, bonuses, incentives, and totals keyed by CIK, ticker, and fiscal year.
List of CMBS deals with publicly available Schedule AL data in SEC.gov EDGAR
A collection of restaurant reviews was assembled in 2019 through Python-based web scraping focused on Dutch establishments, capturing both visit experiences and feature-related information. It is organized in the DatasetDict format with three splits: 116,693 training records, 14,587 test records, and 14,587 validation records.
A tabular extract of clinical study descriptions drawn from ClinicalTrials.gov on 5 February 2025, emphasizing eligibility, design, and objective fields for healthcare and machine learning research.
Could corporate environmental impacts be integrated into fiancial accounting?
Suitable for modelling airline performance post-COVID
6-axis IMU sensor data with 7 different drivers
DerivedFunction01's sec-filings-snippets is a public-records dataset of short text excerpts drawn from U.S. SEC filings, released on Hugging Face in Parquet format. It is sized for language-model fill-mask work and is not positioned as a commercial alternative-data product.
airline-disruption-data is a tabular dataset of U.S. domestic flight records published on the datasets hub, tagged as covering between 10 million and 100 million rows in Parquet format. The schema describes per-flight scheduling, delay, cancellation, and routing fields.
The airline-disruption-data-phase1b dataset is a tabular release hosted by publisher Dev123Hug456Face that compiles scheduled and actual flight-level operations data with delay, cancellation and diversion flags. It is distributed as a single Parquet split of roughly 12 million rows.
ClinicalTrials.gov XML for studies registered between 2018 and 2024, parsed into CSV, hundreds of thousands of records under CC0. A clean, licensed point-in-time history of US clinical-trial registrations.
A large granted-patent text collection from the USPTO, tens of millions of records in Parquet (Apache 2.0). Broad coverage of claims and descriptions for IP-intensity and technology-trend measurement.
Use for Foundation Models, GNNs, and More
This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "names": [ "shoulder_pan.pos", "shoulder_lift.pos", "elbow_flex.pos", "wrist_flex.pos", "wrist_roll.pos", "gripper.pos" ], "shape": [ 6… See the full description on the dataset page: https://huggingface.co/datasets/Edgarium/rangement_pq_20
This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "names": [ "shoulder_pan.pos", "shoulder_lift.pos", "elbow_flex.pos", "wrist_flex.pos", "wrist_roll.pos", "gripper.pos" ], "shape": [ 6… See the full description on the dataset page: https://huggingface.co/datasets/Edgarium/rangement_pq_20
This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "names": [ "shoulder_pan.pos", "shoulder_lift.pos", "elbow_flex.pos", "wrist_flex.pos", "wrist_roll.pos", "gripper.pos" ], "shape": [ 6… See the full description on the dataset page: https://huggingface.co/datasets/Edgarium/rangement_pq_20