Alternative data datasets
Browse by provider, category, marketplace, access, delivery and license
Parsed eligibility criteria from 13,229 ClinicalTrials.gov studies representing the candidate pool gathered by two first-stage retrievers during TREC Clinical Trials 2021–2023, formatted as typed entity-relation graphs for reranking models.
This dataset is a small parquet-format subset of clinical-trial eligibility criteria represented as entity and relation graphs. It is published on a community data hub under an unspecified license.
This dataset, published by 2001jdev on Hugging Face, packages eligibility-criteria information from clinical trials into structured entity and relation records. It is distributed in Parquet format and contains a single split of roughly 104,000 rows.
The clinical-trials-patient-graphs dataset is a small tabular collection of patient-level clinical records distributed in Parquet format by the 2001jdev publisher. It organizes entities and relations extracted from trial data into structured rows spanning 2021 through 2023.
clinical-trials-synth-patient-profiles2 is a synthetic patient-profile dataset distributed by publisher 2001jdev on a public dataset registry. It pairs clinical trial identifiers with short free-text profile descriptions and a coded entity taxonomy, designed for text and entity-recognition work rather than production research.
This dataset is a parsed snapshot of ClinicalTrials.gov records, distributed in Parquet format on the Hugging Face Hub under an unspecified license. It contains roughly 52,000 trial entries drawn from public registry filings.
clinical-trials-trec-qrels is a tabular relevance-judgment file distributed by the 2001jdev user on the Hugging Face Hub. It maps clinical-trial topic identifiers to NCT registry IDs with graded relevance scores.
The clinical-trials-trec-topics dataset is a small tabular collection of clinical-trial topic records distributed in Parquet format on the Hugging Face Hub. It appears to be derived from TREC clinical-trial retrieval topics, with rows partitioned by year.
The job-postings-english-clean dataset, published by 2024-mcm-everitt-ryan, is an English-language corpus of online job postings distributed as a single Parquet train split of roughly 1.76 million rows. It is tagged as US-region data and is intended for text-based machine learning workflows.
The job-postings-raw dataset is a large-scale tabular collection of scraped job postings distributed in Parquet format. It aggregates records across multiple sources and locales, with US coverage indicated in the dataset metadata.
Datamule, Teraflop AI, and Eventual jointly contributed a 590 gigabyte SEC EDGAR collection encompassing about 8 million filings and 43 billion tokens, assembled using the datamule-python library and the official datamule API.
credit_card_transactions is a small tabular dataset published by aegisheld containing anonymized customer-level credit card account and spending summaries. It is distributed as CSV with a single training split of 8,950 rows.
About 1.5 million US patent claims split into training and test partitions in CSV, organized for claim-level classification and summarization work (Apache 2.0). A large, ready-split corpus for IP-claims modeling.
A component of the ALEA Institute's KL3M Data Project supplying training material cleared of copyright concerns. The dataset card is currently a placeholder, directing readers to the GitHub repo and project paper for full documentation.
The KL3M Data Project from the ALEA Institute supplies copyright-clean training material, and the dataset listed here is one component pending further documentation on its dedicated page.
Material contracts and agreements extracted from EDGAR filings by the KL3M project: debt, M&A, employment and license agreements, millions of documents in Parquet. Contract language is a niche but real input for M&A, financing and litigation-event signals.
The kl3m-index-edgar-filings dataset, published by ALEA Institute, is a Parquet-format index of SEC EDGAR filings distributed via the Hub datasets library. It contains roughly 20 million tabular records spanning the 10M–100M size category, with an unstated license.
The KL3M Index of EDGAR 10-K Filings is a tabular index published by the Alea Institute that catalogs SEC EDGAR filings with associated issuer metadata. It is distributed in Parquet format and is tagged for use with pandas, polars, and the mlcroissant dataset library.
The kl3m-index-edgar-filings-8-k dataset is an indexed catalog of Form 8-K filings from the SEC's EDGAR system, published by the ALEA Institute. It provides structured metadata for over 1.8 million corporate event reports distributed in Parquet format.
EDGAR_FILINGS_DATASET_2016_2021 is a Hugging Face mirror of parsed SEC EDGAR filings spanning 2016 through 2021, distributed as a single train split of about 6 million rows in Parquet format. It is published by anonymous-md with an unstated license.
EDGAR_FILINGS_DATASET_2022_2026H1 is a Hugging Face dataset that compiles parsed SEC EDGAR filings into a single tabular corpus. It covers roughly 2022 through the first half of 2026 and is distributed as Parquet with about 1.7 million records.
Dataset Card for "earnings21-gold-transcripts-non-normalized" More Information needed
The TREC Clinical Trials collections released in 2021, 2022, and 2023 are available through the TREC homepage, with corresponding papers for each year. Artur Guimarães ([email protected]) serves as the dataset curator, while inquiries regarding the original authors can be directed separately. The dataset addresses the task of linking patient profiles with appropriate clinical trials, containing entries formatted as JSON objects with fields including query identifiers, disease labels, and accompanying clinical text.
The awinml/earnings_calls_transcripts dataset on the Hugging Face Hub packages a small collection of earnings-call transcript segments formatted as chat-style messages for fine-tuning language models. It is distributed as a parquet file and is tagged for text modality with libraries including datasets, pandas, mlcroissant, and polars.
The airline-otp-data dataset is a large-scale, public-style tabular repository of U.S. domestic flight records covering on-time performance, delay attribution and cancellation outcomes. Published on Hugging Face under an unstated license, it contains roughly 30 million rows in CSV format and is aimed at analysts working with airline operations data.
A large-scale release produced jointly by Datamule, Teraflop AI, and Eventual covering roughly 590 GB of SEC EDGAR content. It comprises about eight million filings totaling 43 billion tokens, gathered using the datamule-python package and the official datamule API.
This dataset is a filtered version of https://huggingface.co/datasets/vincha77/filtered_yelp_restaurant_reviews
yelp_restaurant_reviews_5k is a small text dataset of 5,079 Yelp reviews distributed by bespokelabs, distributed as a single train split. It is structured for straightforward text classification with numeric and categorical fields.
PatentMatch pairs US patents with matching reference patents for retrieval evaluation, a few thousand records in JSON (Apache 2.0). A niche evaluation set for patent-similarity and citation-link models.
A reference collection drawn from the U.S. Patent Phrase to Phrase Matching Kaggle competition, containing additional details that are still required for full documentation.
Cadenza-Labs/apollo-llama3.3-insider-trading-generations is a small text dataset on Hugging Face containing 1,660 training rows used to fine-tune a Llama 3.3 model for detecting dishonest messages. It is distributed in Parquet format and released under an unstated license.
A sample of US patent titles, abstracts and CPC classification labels built for multi-class patent classification, tens of thousands of text records in Parquet. Useful as training material for mapping innovation activity onto technology categories over time.
clinical-trials-v2 is a Hugging Face dataset published by chemNLP containing processed ClinicalTrials.gov records stored in Parquet. It exposes three columns — filename, xml, and text — drawn from the official clinical-study XML schema.
Chinese patents paired with their abstractive summaries in Mandarin, a few thousand records in JSON (Apache 2.0). A window into the pace and direction of patenting inside the Chinese technology base.
A small text-classification dataset by ClarusC64 that labels whether borrow-rate, short-interest and price signals cohere into a real short squeeze or diverge into a false signal. It is published on a model hub with an MIT license tag and a 10-row train split.
GenAI-job-postings-Dataset is a small, US-focused text corpus of generative-AI and machine-learning job postings distributed in Parquet format. The dataset is published on the Hugging Face Hub under an unstated license and comprises a single train split of roughly 120 rows.
A collection of restaurant reviews was assembled in 2019 through Python-based web scraping focused on Dutch establishments, capturing both visit experiences and feature-related information. It is organized in the DatasetDict format with three splits: 116,693 training records, 14,587 test records, and 14,587 validation records.
Records of initial public offerings launched on Indian markets between 2006 and 2025, with fields covering open and close dates, listing date, face value, issue price and size, lot size, first-day price, total shares offered and their allocation across anchor, NII, QIB and retail categories, minimum investment, and subscription figures for each investor class.
A tabular extract of clinical study descriptions drawn from ClinicalTrials.gov on 5 February 2025, emphasizing eligibility, design, and objective fields for healthcare and machine learning research.
Vector embeddings of clinical-trial records, hundreds of thousands of entries in Parquet (Apache 2.0). Enables similarity and retrieval analysis of trial design, indication crowding and competitive positioning.
Dattito's clinical-trials-data is a Parquet-format dataset of roughly 24 million clinical trial records sourced from public trial registries. It aggregates identifiers, conditions, and standardized medical terminology into a single table for large-scale analysis.
An artificially generated collection of apparel product listings and accompanying advertisements produced by prompting GPT-4 to invent one hundred clothing items with descriptions and then write promotional copy for each, output in a structured product and description format without subsequent manual verification.
stk-sec-filings is a small Hugging Face dataset by publisher deerfieldgreen that consolidates U.S. SEC filings into a single Parquet resource. The file covers fewer than one thousand rows distributed across a training split, with no stated update cadence or refresh policy.
DerivedFunction01's sec-filings-snippets is a public-records dataset of short text excerpts drawn from U.S. SEC filings, released on Hugging Face in Parquet format. It is sized for language-model fill-mask work and is not positioned as a commercial alternative-data product.
airline-disruption-data is a tabular dataset of U.S. domestic flight records published on the datasets hub, tagged as covering between 10 million and 100 million rows in Parquet format. The schema describes per-flight scheduling, delay, cancellation, and routing fields.
The airline-disruption-data-phase1b dataset is a tabular release hosted by publisher Dev123Hug456Face that compiles scheduled and actual flight-level operations data with delay, cancellation and diversion flags. It is distributed as a single Parquet split of roughly 12 million rows.
ClinicalTrials.gov XML for studies registered between 2018 and 2024, parsed into CSV, hundreds of thousands of records under CC0. A clean, licensed point-in-time history of US clinical-trial registrations.
A large granted-patent text collection from the USPTO, tens of millions of records in Parquet (Apache 2.0). Broad coverage of claims and descriptions for IP-intensity and technology-trend measurement.
Restaurant Reviews Parsing NER Aspects This dataset is for the task of identifying the aspects of the restaurants mentioned in the reviews where aspect contains information about both the entities (FOOD, AMBIENCE, ...) and the attached sentiments. The input texts are from SemEval dataset. Labels for train and val datasets are generated by prompting Llama3 while the test dataset is curatedly manual
Prepared by Electric Sheep Africa, this MDPA-sourced dataset supplies 10 entries of hotel room occupancy rates for Mauritius over 2019-2023, delivered as ML-ready Parquet files accompanied by uniform Hugging Face metadata and source provenance.
A small Parquet-format collection from Electric Sheep Africa reporting room occupancy rates for large hotels in Mauritius from 2019 to 2023, containing ten records drawn from MDPA and packaged with Hugging Face metadata for machine learning workflows.
Assembled by Electric Sheep Africa from MDPA, this dataset provides 60 rows of monthly hotel room occupancy rates for Mauritius spanning 2019-2023, formatted as ML-ready Parquet with consistent Hugging Face metadata and source attribution.
Compiled by Electric Sheep Africa, this MDPA-derived dataset offers 20 records of quarterly hotel room occupancy rates for Mauritius across the 2019-2023 period, distributed in ML-ready Parquet format with standardized Hugging Face metadata and source traceability.
Electric Sheep Africa releases 60 quarterly observations from MDPA on hotel room occupancy for large establishments in Mauritius from 2019 through 2023, formatted as Parquet with consistent metadata.
高质量中文专利摘要数据集。
Indeed Job Postings (2026) is a 1,000-row sample of US Indeed listings joined to company firmographics, published by fact-den on Apify. License terms are not stated, and the publisher directs users to a commercial source actor for fresh or filtered data.
A U.S. patent collection with a minimal placeholder card and no additional documentation provided.
A legacy archive of Iran's consumer price index and inflation series from 1936 to 2022 remains available, though an updated multisource version using current Farmaanaa methodology supersedes it.
A curated collection of earnings call recordings broken into intelligent segments and paired with high-accuracy transcripts produced by Mistral's Voxtral model, intended for ASR benchmarking, transcription quality studies, and audio-to-text alignment research.
edgar-corpus-embeddings is a text and embedding corpus derived from SEC EDGAR filings, published on Hugging Face by user gagan3012. It packages sectioned 10-K content with pre-computed vector embeddings for machine learning workflows.