Alternative data datasets
Browse by provider, category, marketplace, access, delivery and license
Parsed eligibility criteria from 13,229 ClinicalTrials.gov studies representing the candidate pool gathered by two first-stage retrievers during TREC Clinical Trials 2021–2023, formatted as typed entity-relation graphs for reranking models.
This dataset is a small parquet-format subset of clinical-trial eligibility criteria represented as entity and relation graphs. It is published on a community data hub under an unspecified license.
This dataset, published by 2001jdev on Hugging Face, packages eligibility-criteria information from clinical trials into structured entity and relation records. It is distributed in Parquet format and contains a single split of roughly 104,000 rows.
clinical-trials-synth-patient-profiles2 is a synthetic patient-profile dataset distributed by publisher 2001jdev on a public dataset registry. It pairs clinical trial identifiers with short free-text profile descriptions and a coded entity taxonomy, designed for text and entity-recognition work rather than production research.
This dataset is a parsed snapshot of ClinicalTrials.gov records, distributed in Parquet format on the Hugging Face Hub under an unspecified license. It contains roughly 52,000 trial entries drawn from public registry filings.
clinical-trials-trec-qrels is a tabular relevance-judgment file distributed by the 2001jdev user on the Hugging Face Hub. It maps clinical-trial topic identifiers to NCT registry IDs with graded relevance scores.
The clinical-trials-trec-topics dataset is a small tabular collection of clinical-trial topic records distributed in Parquet format on the Hugging Face Hub. It appears to be derived from TREC clinical-trial retrieval topics, with rows partitioned by year.
Datamule, Teraflop AI, and Eventual jointly contributed a 590 gigabyte SEC EDGAR collection encompassing about 8 million filings and 43 billion tokens, assembled using the datamule-python library and the official datamule API.
Real Estate listings (2.2M+) in the US broken by State and zip code
About 1.5 million US patent claims split into training and test partitions in CSV, organized for claim-level classification and summarization work (Apache 2.0). A large, ready-split corpus for IP-claims modeling.
A component of the ALEA Institute's KL3M Data Project supplying training material cleared of copyright concerns. The dataset card is currently a placeholder, directing readers to the GitHub repo and project paper for full documentation.
The KL3M Data Project from the ALEA Institute supplies copyright-clean training material, and the dataset listed here is one component pending further documentation on its dedicated page.
Material contracts and agreements extracted from EDGAR filings by the KL3M project: debt, M&A, employment and license agreements, millions of documents in Parquet. Contract language is a niche but real input for M&A, financing and litigation-event signals.
The kl3m-index-edgar-filings dataset, published by ALEA Institute, is a Parquet-format index of SEC EDGAR filings distributed via the Hub datasets library. It contains roughly 20 million tabular records spanning the 10M–100M size category, with an unstated license.
The KL3M Index of EDGAR 10-K Filings is a tabular index published by the Alea Institute that catalogs SEC EDGAR filings with associated issuer metadata. It is distributed in Parquet format and is tagged for use with pandas, polars, and the mlcroissant dataset library.
The kl3m-index-edgar-filings-8-k dataset is an indexed catalog of Form 8-K filings from the SEC's EDGAR system, published by the ALEA Institute. It provides structured metadata for over 1.8 million corporate event reports distributed in Parquet format.
Tariff Escalation, Trade Diversion & Sector Disruption Across 30+ Nations
US patent full text assembled by the Allen Institute from the USPTO, millions of records under ODC-BY: claims, abstracts, descriptions and metadata, in Parquet. A broad base for measuring IP intensity, technology landscapes and the direction of R&D spending.
View source information and access options.
Quarterly institutional filings from 2015 - 2017
Comprises of Electricity consumption of each sector in Indian cities
An open, layout-preserving conversion of U.S. SEC EDGAR filings into MultiMarkdown format covers approximately 3.4 million documents submitted between January 2022 and June 2025, designed for long-context language modeling, financial reasoning, and document analysis.
EDGAR_FILINGS_DATASET_2016_2021 is a Hugging Face mirror of parsed SEC EDGAR filings spanning 2016 through 2021, distributed as a single train split of about 6 million rows in Parquet format. It is published by anonymous-md with an unstated license.
EDGAR_FILINGS_DATASET_2022_2026H1 is a Hugging Face dataset that compiles parsed SEC EDGAR filings into a single tabular corpus. It covers roughly 2022 through the first half of 2026 and is distributed as Parquet with about 1.7 million records.
Dataset Card for "earnings21-gold-transcripts-non-normalized" More Information needed
5 years and 200k building permits
View source information and access options.
View source information and access options.
View source information and access options.
View source information and access options.
View source information and access options.
View source information and access options.
View source information and access options.
View source information and access options.
The TREC Clinical Trials collections released in 2021, 2022, and 2023 are available through the TREC homepage, with corresponding papers for each year. Artur Guimarães ([email protected]) serves as the dataset curator, while inquiries regarding the original authors can be directed separately. The dataset addresses the task of linking patient profiles with appropriate clinical trials, containing entries formatted as JSON objects with fields including query identifiers, disease labels, and accompanying clinical text.
View source information and access options.
Building permits and residential unit information in San Francisco
Brasil real estate Dataset For Prediction
Past 1 year Building permits issued
An aggregation of 1.3 million U.S. patent records each accompanied by a human-written abstractive summary, sorted into nine Cooperative Patent Classification groupings spanning human necessities, chemistry, textiles, and related domains.
View source information and access options.
View source information and access options.
Collection of historical Swedish patent texts from 1885 to 1972 assigned multi-label Cooperative Patent Classification (CPC) codes, intended for retrieval, prior art searching, and multi-label classification tasks.
U.S. nationwide recorder and deed documentation exceeding 593 million entries, capturing sales transactions, mortgage filings, and historical ownership changes.
Risk and compliance intelligence derived from official audit and supervisory examination records applied across regulated organizations.
The awinml/earnings_calls_transcripts dataset on the Hugging Face Hub packages a small collection of earnings-call transcript segments formatted as chat-style messages for fine-tuning language models. It is distributed as a parquet file and is tagged for text modality with libraries including datasets, pandas, mlcroissant, and polars.
DSLM Fazenda is a living economic intelligence platform focused on fiscal policy, macroeconomics, and financial management. It monitors macroeconomic indicators, analyzes tax and fiscal legislation, and supports interpretation of Central Bank, Federal Revenue Service, and other regulatory norms for organizations, sectors, and public managers.
Secure Payments IVR for Amazon Connect
Agentic financial research across licensed data from S&P Global, QUODD, FRED, and SEC. Runs multi-step search, reconciles conflicting sources, and cross-verifies findings to return accurate, cited answers with every figure traced to its primary source.
A large-scale release produced jointly by Datamule, Teraflop AI, and Eventual covering roughly 590 GB of SEC EDGAR content. It comprises about eight million filings totaling 43 billion tokens, gathered using the datamule-python package and the official datamule API.
View source information and access options.
View source information and access options.
View source information and access options.
View source information and access options.
A leak-free, chronologically ordered corpus of 8-K, 10-Q, and 10-K filings for roughly 610 present-day U.S. large-cap companies, each paired with objective forward-return labels through 2026. The release is structured as a benchmark for testing whether filing language can anticipate subsequent stock performance, with clearly separated provenance classes.
View source information and access options.
BigQuery dataset of all SEC filings
View source information and access options.
View source information and access options.
View source information and access options.