Alternative data datasets
Browse by provider, category, marketplace, access, delivery and license
10-K annual filings from EDGAR in Parquet, tens of thousands of documents. The standard annual-disclosure corpus for fundamentals, risk-factor and management-discussion analysis.
A component of the ALEA Institute's KL3M Data Project supplying training material cleared of copyright concerns. The dataset card is currently a placeholder, directing readers to the GitHub repo and project paper for full documentation.
The KL3M Index of EDGAR 10-K Filings is a tabular index published by the Alea Institute that catalogs SEC EDGAR filings with associated issuer metadata. It is distributed in Parquet format and is tagged for use with pandas, polars, and the mlcroissant dataset library.
Complete filing metadata and URLs for 82,000+ annual report filings (July-Dec)
Public access to SEC filings and related submissions/XBRL data APIs. Filing forms differ in scope, structure and reporting delay; amendments and publication timestamps matter for historical research.
8 size-normalized financial ratios for ~440 US companies from SEC EDGAR 10-Ks
Raw 10-K financials for 442 US companies extracted from SEC EDGAR
Multi-modal dataset: Bridging GAAP accounting and LLM-powered NLP for the S&P 50
An open, layout-preserving conversion of U.S. SEC EDGAR filings into MultiMarkdown format covers approximately 3.4 million documents submitted between January 2022 and June 2025, designed for long-context language modeling, financial reasoning, and document analysis.
Leverage a LLM to extract data from financial reporting from 10-k, 10-Q, news and events, etc
A leak-free, chronologically ordered corpus of 8-K, 10-Q, and 10-K filings for roughly 610 present-day U.S. large-cap companies, each paired with objective forward-return labels through 2026. The release is structured as a benchmark for testing whether filing language can anticipate subsequent stock performance, with clearly separated provenance classes.
A reconstructed EDGAR corpus of company filings in English (Apache 2.0): a large sample of annual and quarterly reports and exhibits, hundreds of thousands of records. Comparable to other public SEC text sets and a solid starting point for filings-based research.
Machine-readable XBRL filings covering 10-K and 10-Q reports, cleaned and standardized so financial figures can be extracted reliably.
Drawn from SEC EDGAR's full archive of millions of filings spanning all form types and US public companies, this sample isolates 1,000 recent 8-K material event submissions together with their filing metadata and document references.
A corpus of SEC EDGAR filings in English (Apache 2.0), roughly a few hundred thousand company documents including annual and quarterly reports and exhibits. The filings text is organized for model pretraining and retrieval, but the same documents are the raw material for filings-based and disclosure signals.
An extensive archive of SEC EDGAR submissions from company filings, organized by CIK, document type, and date, and accessible through a live REST API at api.ai-analytics.org.
Structured numeric records parsed from corporate filings such as 10-K, 10-Q, and 8-K submissions to the U.S. SEC. The figures are packaged inside a compressed DuckDB instance made available through the open-source Datapond registry for fast querying.
Financial data parsed from 10-Q, 10-Q/A, 10-K, 10-K/A SEC filings from 2010.
edgar-corpus-embeddings is a text and embedding corpus derived from SEC EDGAR filings, published on Hugging Face by user gagan3012. It packages sectioned 10-K content with pre-computed vector embeddings for machine learning workflows.
Quarterly 10-Q filings parsed from EDGAR into JSON, millions of records in English (MIT). The quarterly companion to the 10-K corpus for tracking sequential changes in operating performance and risk factors.
Structured line items from SEC 10-K and 10-Q filings between 2009 and 2023, plus company facts — fundamentals extracted straight from EDGAR.
10-K filings for S&P 500 companies retrieved from EDGAR, stored as Parquet with the filing text attached (MIT). A focused corpus for the large-cap annual disclosure cycle: risk factors, MD&A and segment detail for every S&P 500 name.
FinSight gathers plain-text 10-K and 10-Q SEC filings for twenty major US public firms spanning six sectors, yielding 97 records intended for BERT fine-tuning, retrieval-augmented generation, and multi-agent financial research.
Penumbra AI SEC 10-K Risk Factor Disclosures (S&P 500, FY2025-2026) Dataset Summary This dataset contains a sample of 2,683 individual risk factor disclosures from the most recent Form 10-K filing of 50 S&P 500 companies (fiscal year ends ranging from May 2025 to February 2026). Each record is a single risk driver — one discrete point a company disclosed under "Risk Factors" — together with the ve
Complete textual SEC 10-K annual filings from 2021, suited for studying corporate disclosures and natural language processing tasks.
A suite of core financial analytics covering every SEC EDGAR filing type, including 10-K, 10-Q, and 8-K submissions.
EDGAR M&A Deal Events (2010–2024) 4,156 U.S. public-company acquisition events, discovered directly from SEC EDGAR's own quarterly filing indexes — not scraped from a vendor list or a blog. Each row is a company that filed a merger proxy or tender-offer response between 2010 and 2024, with the earliest such filing's date as an announcement-date proxy. Built as a byproduct of an M&A target-predicti
Text of SEC filings pulled from EDGAR spanning millions of records in English (Apache 2.0): 10-K, 10-Q, 8-K, annual reports and exhibits, tagged for text-generation and classification. A practical base corpus for earnings- and filings-driven signals, from risk factors and MD&A to management commentary.
A historical, quarterly index of U.S. public company 10-K and 10-Q filings