Alternative data datasets
Browse by provider, category, marketplace, access, delivery and license
Structured Form 144 filings used to track insider transactions and inform quantitative trading approaches.
Datamule, Teraflop AI, and Eventual jointly contributed a 590 gigabyte SEC EDGAR collection encompassing about 8 million filings and 43 billion tokens, assembled using the datamule-python library and the official datamule API.
View source information and access options.
Quarterly institutional filings from 2015 - 2017
SEC insider transactions 2006-2026 with stock & insider metadata
A large-scale release produced jointly by Datamule, Teraflop AI, and Eventual covering roughly 590 GB of SEC EDGAR content. It comprises about eight million filings totaling 43 billion tokens, gathered using the datamule-python package and the official datamule API.
A leak-free, chronologically ordered corpus of 8-K, 10-Q, and 10-K filings for roughly 610 present-day U.S. large-cap companies, each paired with objective forward-return labels through 2026. The release is structured as a benchmark for testing whether filing language can anticipate subsequent stock performance, with clearly separated provenance classes.
BigQuery dataset of all SEC filings
Disclosure documents submitted by institutional investment firms that manage portfolios exceeding $100 million under Section 13(f).
Ticketing information aggregated from a digital marketplace for live music, athletic, and theatrical performances.
Drawn from SEC EDGAR's full archive of millions of filings spanning all form types and US public companies, this sample isolates 1,000 recent 8-K material event submissions together with their filing metadata and document references.
View source information and access options.
Extracted from SEC EDGAR DEF 14A proxy filings since 2015, the dataset contains over 500,000 structured executive compensation entries for S&P 500 and Russell 2000 issuers, capturing CEO and CFO pay, equity awards, bonuses, incentives, and totals keyed by CIK, ticker, and fiscal year.
Quarterly, annual and TTM financials with XBRL provenance and amendment signals
A systematic crawl of business websites classified by sector and geography, giving store counts, opening hours, technology stacks and contact details per market.
All 8-k filings for 2024 with augmented fields.
Includes all company names and cik keys from the SEC database.
DerivedFunction01's sec-filings-snippets is a public-records dataset of short text excerpts drawn from U.S. SEC filings, released on Hugging Face in Parquet format. It is sized for language-model fill-mask work and is not positioned as a commercial alternative-data product.
An extensive collection of financial filings sourced from the U.S. Securities and Exchange Commission, intended for analytical and academic use.
An extensive archive of SEC EDGAR submissions from company filings, organized by CIK, document type, and date, and accessible through a live REST API at api.ai-analytics.org.
Structured numeric records parsed from corporate filings such as 10-K, 10-Q, and 8-K submissions to the U.S. SEC. The figures are packaged inside a compressed DuckDB instance made available through the open-source Datapond registry for fast querying.
Trading prices and liquidity metrics for shares of privately held companies in their later growth stages.
Data Analytics and Data Visualisation for beginners
View source information and access options.
Structured line items from SEC 10-K and 10-Q filings between 2009 and 2023, plus company facts, fundamentals extracted straight from EDGAR.
Historical Data from SEC Company Filings
A joint release by Datamule, Teraflop AI, and Eventual provides the SEC-EDGAR dataset, comprising 590 GB of content drawn from all major filings in the SEC EDGAR database and totaling roughly 8 million samples and 43 billion tokens. The data was assembled using the datamule-python library and the official datamule API developed by John Friedman, with datamule serving as a Python package for collecting and manipulating such records.
Datamule partnered with Teraflop AI and Eventual to publish the SEC-EDGAR dataset, which covers all major filings from the SEC EDGAR database and amounts to 590 GB across about 8 million samples and 43 billion tokens. Collection relied on the datamule-python library and John Friedman's official datamule API, tools designed in Python for gathering and handling SEC filing data.
Metadata for every EDGAR filing submitted to the U.S. Securities and Exchange Commission from 1994 through December 14, 2024, derived from the quarterly master index files and including CIKs, issuer names, and form types.
sec_filings is a small text dataset of 494 rows that pairs natural-language queries about U.S. securities filings with retrieved document facts and ground-truth answers. It is distributed under an unstated license on a public dataset hub.
AlphaAI's near-real-time parse of SEC EDGAR Form 4 filings, providing 20,832 individual insider transaction tranches and 8,538 grouped economic events for US public-company officers, directors, and 10% owners.
EDGAR filings of companies
Complete filing metadata and URLs for all 8-K material event disclosures in Q4
Complete filing metadata and URLs for 82,000+ annual report filings (July-Dec)
Complete filing metadata and URLs for all insider transaction filings
High-Conviction Insider Trading & Institutional Holding Signals
View source information and access options.
Index of filings submitted to the Securities & Exchange Commission since 1993
Catalog of SEC filings dating back to 1993, providing a navigational guide to the EDGAR archive.
View source information and access options.
View source information and access options.
Penumbra AI SEC 10-K Risk Factor Disclosures (S&P 500, FY2025-2026) Dataset Summary This dataset contains a sample of 2,683 individual risk factor disclosures from the most recent Form 10-K filing of 50 S&P 500 companies (fiscal year ends ranging from May 2025 to February 2026). Each record is a single risk driver — one discrete point a company disclosed under "Risk Factors" — together with the ve
Complete textual SEC 10-K annual filings from 2021, suited for studying corporate disclosures and natural language processing tasks.
Extracted 10k fillings from SEC Edgar for the fiscal year 2021
Search filings on SEC EDGAR using ticker symbols, CIK identifiers, SIC industry codes, or free-text queries, returning a single structured row per filing that includes company classification, period dates, direct links to source documents, and optional XBRL financial data. The dataset contains 2,691 rows across 27 fields, supported by 64 collector runs, with the latest observation dated 2026-08-04, and covers 771 entities, browsable at https://reapx.dev/data/sec-edgar-scraper/.
Public access to SEC filings and related submissions/XBRL data APIs. Filing forms differ in scope, structure and reporting delay; amendments and publication timestamps matter for historical research.
A suite of core financial analytics covering every SEC EDGAR filing type, including 10-K, 10-Q, and 8-K submissions.
Aggregated daily short selling volume information from TRF and ADF for transactions in exchange-traded securities.
Clean, normalized daily prices with insider trading filings
sec-filings is a Hugging Face dataset published by Sicheng-Chroma that distributes pre-chunked U.S. SEC filings paired with dense vector embeddings. It is distributed in optimized Parquet format and is intended for retrieval and similarity research over regulatory disclosures.
View source information and access options.
View source information and access options.
Parquet and csv files associating CIK (Central Index Key) with filing entities
JSON data file for SEC conformed company name, CIK, ticker, exchange association
Text of SEC filings pulled from EDGAR spanning millions of records in English (Apache 2.0): 10-K, 10-Q, 8-K, annual reports and exhibits, tagged for text-generation and classification. A practical base corpus for earnings- and filings-driven signals, from risk factors and MD&A to management commentary.
Normalized financial statements drawn from the SEC Financial Statement Data Sets and company-facts services, both built on the EDGAR APIs and spanning reported results across filers.
A historical, quarterly index of U.S. public company 10-K and 10-Q filings
AI importance scores for 18k+ SEC filings vs next-session excess moves
Secure Payments IVR for Amazon Connect
An anonymized consumer purchase dataset from Bloomberg Second Measure that projects corporate revenue ahead of official disclosures.