Alternative data datasets
Browse by provider, category, marketplace, access, delivery and license
An open, layout-preserving conversion of U.S. SEC EDGAR filings into MultiMarkdown format covers approximately 3.4 million documents submitted between January 2022 and June 2025, designed for long-context language modeling, financial reasoning, and document analysis.
EDGAR_FILINGS_DATASET_2016_2021 is a Hugging Face mirror of parsed SEC EDGAR filings spanning 2016 through 2021, distributed as a single train split of about 6 million rows in Parquet format. It is published by anonymous-md with an unstated license.
EDGAR_FILINGS_DATASET_2022_2026H1 is a Hugging Face dataset that compiles parsed SEC EDGAR filings into a single tabular corpus. It covers roughly 2022 through the first half of 2026 and is distributed as Parquet with about 1.7 million records.
A reconstructed EDGAR corpus of company filings in English (Apache 2.0): a large sample of annual and quarterly reports and exhibits, hundreds of thousands of records. Comparable to other public SEC text sets and a solid starting point for filings-based research.
edgar_xbrl_companyfacts is a Hugging Face dataset published by DenyTranDFW that aggregates U.S. SEC EDGAR XBRL company-facts filings into a single Parquet file. The training split contains roughly 125 million rows of structured financial facts tagged by reporting period and unit.
The cleaned public release of the EDGARCalcQA benchmark, containing example files in both JSON and JSONL formats, a human-readable summary, optional rejected candidate records, and a schema document describing field names, JSON types, and brief field-level descriptions.
An adaptation of the CelebA-HQ face image set at 512 resolution that retains the original photographs while appending identity-cluster labels produced automatically through face embedding grouping.
Generated with the LeRobot framework, this dataset comprises 80 episodes demonstrating a pick-and-place routine using a single orange cube positioned at varying locations within the workspace. Each episode requires the robot to rotate toward the cube, open its gripper, close it around the object, and transport it to a specified drop zone.
This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "names": [ "shoulder_pan.pos", "shoulder_lift.pos", "elbow_flex.pos", "wrist_flex.pos", "wrist_roll.pos", "gripper.pos" ], "shape": [ 6… See the full description on the dataset page: https://huggingface.co/datasets/Edgarium/rangement_pq_20
This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so101_follower", "total_episodes": 24, "total_frames": 5358, "total_tasks": 1, "total_videos": 48, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:24" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… Se
A LeRobot-formatted dataset captured with a so101_follower robot, comprising 30 episodes, 6,788 frames, 60 video files, and a single chunk at 30 fps, with training covering the full episode range.
This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so101_follower", "total_episodes": 50, "total_frames": 13379, "total_tasks": 1, "total_videos": 100, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:50" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":…
This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so101_follower", "total_episodes": 709, "total_frames": 187905, "total_tasks": 1, "total_videos": 1418, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:709" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path
View source information and access options.
View source information and access options.
View source information and access options.
View source information and access options.
A robotics dataset generated with the LeRobot framework contains two recorded episodes totaling 897 frames at 30 frames per second, captured using an SO100 robot performing a single task, with four video files and accompanying parquet data organized in a single chunk of up to 1,000 entries and partitioned as a train split covering episodes zero through one.
Another LeRobot so100 robotics dataset containing 50 episodes, 29,874 frames, and 100 videos at 30 fps, structured identically to the related edgar-block release.
A LeRobot dataset captured on the so100 robot platform, consisting of 50 episodes, 44,746 frames, and 100 videos recorded at 30 fps and split for training.
A robotics dataset assembled with the LeRobot framework for the SO100 platform, containing 20 episodes totaling 11,940 frames captured at 30 fps, stored in parquet data files alongside 40 corresponding video chunks.
View source information and access options.
View source information and access options.
This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so101_follower", "total_episodes": 50, "total_frames": 12620, "total_tasks": 1, "total_videos": 100, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:50" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":…
View source information and access options.
This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so101_follower", "total_episodes": 500, "total_frames": 136475, "total_tasks": 1, "total_videos": 1000, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:500" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path
This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so101_follower", "total_episodes": 500, "total_frames": 137269, "total_tasks": 1, "total_videos": 1000, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:500" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path
A LeRobot-formatted robotics capture set for the so101_follower arm, comprising 51 episodes, 13,109 frames, and 102 videos at 30 fps stored across parquet chunks.
A LeRobot-generated dataset comprising 149 episodes, 27,628 frames, and 298 video files captured at 30 fps on a so101_follower robot, with all episodes assigned to the training split and stored in the chunk-based Parquet layout expected by the codebase.
Recorded with the LeRobot v2.1 schema on a 6-DOF SO-ARM101 manipulator, this dataset contains 2 episodes totaling 409 frames captured at 30 FPS in 640x480 resolution from two cameras, showing the arm retrieving a red square block and depositing it into a green square region on the right.
A corpus of SEC EDGAR filings in English (Apache 2.0), roughly a few hundred thousand company documents including annual and quarterly reports and exhibits. The filings text is organized for model pretraining and retrieval, but the same documents are the raw material for filings-based and disclosure signals.
edgar-corpus-embeddings is a text and embedding corpus derived from SEC EDGAR filings, published on Hugging Face by user gagan3012. It packages sectioned 10-K content with pre-computed vector embeddings for machine learning workflows.
Quarterly 10-Q filings parsed from EDGAR into JSON, millions of records in English (MIT). The quarterly companion to the 10-K corpus for tracking sequential changes in operating performance and risk factors.
A JSONL collection of detokenized EDGAR agreement records, where a single entry may contain multiple agreements and no source-level cleaning has been applied.
EdgarItem7 is a text dataset derived from SEC filings, packaging Item 6 (Selected Financial Data) and Item 7 (Management's Discussion and Analysis) excerpts alongside filing metadata. It is hosted by publisher lealOO and distributed in Arrow format under an unstated license.
EDGAR M&A Deal Events (2010–2024) 4,156 U.S. public-company acquisition events, discovered directly from SEC EDGAR's own quarterly filing indexes — not scraped from a vendor list or a blog. Each row is a company that filed a merger proxy or tender-offer response between 2010 and 2024, with the earliest such filing's date as an announcement-date proxy. Built as a byproduct of an M&A target-predicti
Datamule, Teraflop AI, and Eventual jointly contributed a 590 gigabyte SEC EDGAR collection encompassing about 8 million filings and 43 billion tokens, assembled using the datamule-python library and the official datamule API.
A component of the ALEA Institute's KL3M Data Project supplying training material cleared of copyright concerns. The dataset card is currently a placeholder, directing readers to the GitHub repo and project paper for full documentation.
The KL3M Data Project from the ALEA Institute supplies copyright-clean training material, and the dataset listed here is one component pending further documentation on its dedicated page.
Material contracts and agreements extracted from EDGAR filings by the KL3M project: debt, M&A, employment and license agreements, millions of documents in Parquet. Contract language is a niche but real input for M&A, financing and litigation-event signals.
The kl3m-index-edgar-filings dataset, published by ALEA Institute, is a Parquet-format index of SEC EDGAR filings distributed via the Hub datasets library. It contains roughly 20 million tabular records spanning the 10M–100M size category, with an unstated license.
The KL3M Index of EDGAR 10-K Filings is a tabular index published by the Alea Institute that catalogs SEC EDGAR filings with associated issuer metadata. It is distributed in Parquet format and is tagged for use with pandas, polars, and the mlcroissant dataset library.
The kl3m-index-edgar-filings-8-k dataset is an indexed catalog of Form 8-K filings from the SEC's EDGAR system, published by the ALEA Institute. It provides structured metadata for over 1.8 million corporate event reports distributed in Parquet format.
View source information and access options.
A large-scale release produced jointly by Datamule, Teraflop AI, and Eventual covering roughly 590 GB of SEC EDGAR content. It comprises about eight million filings totaling 43 billion tokens, gathered using the datamule-python package and the official datamule API.
titer · EDGAR officer corpus 4,206,080 attested person–company–role–date tuples from SEC Forms 3/4/5, published as pointers rather than records, alongside the frozen pre-registrations that were hash-published before any measurement ran. edgar_officers.parquet: 4.2M rows, 230,405 distinct people, 20,266 issuers, 2006q1–2026q2. Column Meaning accession SEC accession number, the pointer that reconstr
Extracted from SEC EDGAR DEF 14A proxy filings since 2015, the dataset contains over 500,000 structured executive compensation entries for S&P 500 and Russell 2000 issuers, capturing CEO and CFO pay, equity awards, bonuses, incentives, and totals keyed by CIK, ticker, and fiscal year.
List of CMBS deals with publicly available Schedule AL data in SEC.gov EDGAR
Includes all company names and cik keys from the SEC database.
Structured numeric records parsed from corporate filings such as 10-K, 10-Q, and 8-K submissions to the U.S. SEC. The figures are packaged inside a compressed DuckDB instance made available through the open-source Datapond registry for fast querying.
Data Analytics and Data Visualisation for beginners
A consolidated, ready-to-use repository of 13F institutional holdings drawn from SEC EDGAR filings, capturing the most recent quarterly disclosures of more than 13,000 investment managers including hedge funds, mutual fund complexes, pension funds, banks, and family offices, with each record summarizing total assets under management and individual positions.
Structured line items from SEC 10-K and 10-Q filings between 2009 and 2023, plus company facts — fundamentals extracted straight from EDGAR.
Historical Data from SEC Company Filings
A joint release by Datamule, Teraflop AI, and Eventual provides the SEC-EDGAR dataset, comprising 590 GB of content drawn from all major filings in the SEC EDGAR database and totaling roughly 8 million samples and 43 billion tokens. The data was assembled using the datamule-python library and the official datamule API developed by John Friedman, with datamule serving as a Python package for collecting and manipulating such records.
10-K filings for S&P 500 companies retrieved from EDGAR, stored as Parquet with the filing text attached (MIT). A focused corpus for the large-cap annual disclosure cycle: risk factors, MD&A and segment detail for every S&P 500 name.
Datamule partnered with Teraflop AI and Eventual to publish the SEC-EDGAR dataset, which covers all major filings from the SEC EDGAR database and amounts to 590 GB across about 8 million samples and 43 billion tokens. Collection relied on the datamule-python library and John Friedman's official datamule API, tools designed in Python for gathering and handling SEC filing data.
EDGAR filings of companies
Complete filing metadata and URLs for 82,000+ annual report filings (July-Dec)
View source information and access options.