Alternative data datasets
Browse by provider, category, marketplace, access, delivery and license
BigQuery dataset of all SEC filings
An extensive archive of SEC EDGAR submissions from company filings, organized by CIK, document type, and date, and accessible through a live REST API at api.ai-analytics.org.
View source information and access options.
sec-filings is a Hugging Face dataset published by Sicheng-Chroma that distributes pre-chunked U.S. SEC filings paired with dense vector embeddings. It is distributed in optimized Parquet format and is intended for retrieval and similarity research over regulatory disclosures.
A leak-free, chronologically ordered corpus of 8-K, 10-Q, and 10-K filings for roughly 610 present-day U.S. large-cap companies, each paired with objective forward-return labels through 2026. The release is structured as a benchmark for testing whether filing language can anticipate subsequent stock performance, with clearly separated provenance classes.
Drawn from SEC EDGAR's full archive of millions of filings spanning all form types and US public companies, this sample isolates 1,000 recent 8-K material event submissions together with their filing metadata and document references.
DerivedFunction01's sec-filings-snippets is a public-records dataset of short text excerpts drawn from U.S. SEC filings, released on Hugging Face in Parquet format. It is sized for language-model fill-mask work and is not positioned as a commercial alternative-data product.
stk-sec-filings is a small Hugging Face dataset by publisher deerfieldgreen that consolidates U.S. SEC filings into a single Parquet resource. The file covers fewer than one thousand rows distributed across a training split, with no stated update cadence or refresh policy.
A compact sample set of SEC filings used for the MemGPT agent demos, tens of thousands of text records. A small, clean entry point to filings text rather than a comprehensive corpus.
FinSight gathers plain-text 10-K and 10-Q SEC filings for twenty major US public firms spanning six sectors, yielding 97 records intended for BERT fine-tuning, retrieval-augmented generation, and multi-agent financial research.
A curated collection of SEC EDGAR filings from recent and upcoming IPO companies, containing 5,179 text segments sourced from 261 filings, formatted for large language model training and financial analysis.
The cleaned public release of the EDGARCalcQA benchmark, containing example files in both JSON and JSONL formats, a human-readable summary, optional rejected candidate records, and a schema document describing field names, JSON types, and brief field-level descriptions.
A consolidated, ready-to-use repository of 13F institutional holdings drawn from SEC EDGAR filings, capturing the most recent quarterly disclosures of more than 13,000 investment managers including hedge funds, mutual fund complexes, pension funds, banks, and family offices, with each record summarizing total assets under management and individual positions.
EDGAR M&A Deal Events (2010–2024) 4,156 U.S. public-company acquisition events, discovered directly from SEC EDGAR's own quarterly filing indexes — not scraped from a vendor list or a blog. Each row is a company that filed a merger proxy or tender-offer response between 2010 and 2024, with the earliest such filing's date as an announcement-date proxy. Built as a byproduct of an M&A target-predicti