Alternative data datasets
Browse by provider, category, marketplace, access, delivery and license
Includes all company names and cik keys from the SEC database.
Dattito's clinical-trials-data is a Parquet-format dataset of roughly 24 million clinical trial records sourced from public trial registries. It aggregates identifiers, conditions, and standardized medical terminology into a single table for large-scale analysis.
View source information and access options.
Later and continuously updated statistical outputs capturing developments from the late 1980s through the 1990s.
Financial sector data and macroeconomic statistical releases standardized under reference 71.
A cross-domain collection of statistical frameworks covering standards 72 through 74.
A 220-million-row consumer file forms the foundation for ZIP-code-level analytics.
View source information and access options.
stk-sec-filings is a small Hugging Face dataset by publisher deerfieldgreen that consolidates U.S. SEC filings into a single Parquet resource. The file covers fewer than one thousand rows distributed across a training split, with no stated update cadence or refresh policy.
edgar_xbrl_companyfacts is a Hugging Face dataset published by DenyTranDFW that aggregates U.S. SEC EDGAR XBRL company-facts filings into a single Parquet file. The training split contains roughly 125 million rows of structured financial facts tagged by reporting period and unit.
DerivedFunction01's sec-filings-snippets is a public-records dataset of short text excerpts drawn from U.S. SEC filings, released on Hugging Face in Parquet format. It is sized for language-model fill-mask work and is not positioned as a commercial alternative-data product.
Real estate listings and sales from 2001-2021
Useful for retrieval-augmented legal history analysis, similarity-based search, and pretraining language models specialized in legal text.
ClinicalTrials.gov XML for studies registered between 2018 and 2024, parsed into CSV, hundreds of thousands of records under CC0. A clean, licensed point-in-time history of US clinical-trial registrations.
Building cost figures and benchmarks from Dodge construction data, organized by structure type, project scale and geographic area.
View source information and access options.
View source information and access options.
View source information and access options.
View source information and access options.
MolMole_Patent300 is a curated evaluation benchmark for extracting chemical information from full patent pages, supporting end-to-end testing of molecule detection, reaction parsing, and optical chemical structure recognition.
Real Estate Listings Across Major Cities in Bangladesh with Prices in Bangladesh
A large granted-patent text collection from the USPTO, tens of millions of records in Parquet (Apache 2.0). Broad coverage of claims and descriptions for IP-intensity and technology-trend measurement.
The cleaned public release of the EDGARCalcQA benchmark, containing example files in both JSON and JSONL formats, a human-readable summary, optional rejected candidate records, and a schema document describing field names, JSON types, and brief field-level descriptions.
An adaptation of the CelebA-HQ face image set at 512 resolution that retains the original photographs while appending identity-cluster labels produced automatically through face embedding grouping.
Generated with the LeRobot framework, this dataset comprises 80 episodes demonstrating a pick-and-place routine using a single orange cube positioned at varying locations within the workspace. Each episode requires the robot to rotate toward the cube, open its gripper, close it around the object, and transport it to a specified drop zone.
This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "names": [ "shoulder_pan.pos", "shoulder_lift.pos", "elbow_flex.pos", "wrist_flex.pos", "wrist_roll.pos", "gripper.pos" ], "shape": [ 6… See the full description on the dataset page: https://huggingface.co/datasets/Edgarium/rangement_pq_20
This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so101_follower", "total_episodes": 24, "total_frames": 5358, "total_tasks": 1, "total_videos": 48, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:24" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… Se
A LeRobot-formatted dataset captured with a so101_follower robot, comprising 30 episodes, 6,788 frames, 60 video files, and a single chunk at 30 fps, with training covering the full episode range.
This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so101_follower", "total_episodes": 50, "total_frames": 13379, "total_tasks": 1, "total_videos": 100, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:50" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":…
This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so101_follower", "total_episodes": 709, "total_frames": 187905, "total_tasks": 1, "total_videos": 1418, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:709" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path
View source information and access options.
View source information and access options.
View source information and access options.
View source information and access options.
A robotics dataset generated with the LeRobot framework contains two recorded episodes totaling 897 frames at 30 frames per second, captured using an SO100 robot performing a single task, with four video files and accompanying parquet data organized in a single chunk of up to 1,000 entries and partitioned as a train split covering episodes zero through one.
Another LeRobot so100 robotics dataset containing 50 episodes, 29,874 frames, and 100 videos at 30 fps, structured identically to the related edgar-block release.
A LeRobot dataset captured on the so100 robot platform, consisting of 50 episodes, 44,746 frames, and 100 videos recorded at 30 fps and split for training.
A robotics dataset assembled with the LeRobot framework for the SO100 platform, containing 20 episodes totaling 11,940 frames captured at 30 fps, stored in parquet data files alongside 40 corresponding video chunks.
View source information and access options.
View source information and access options.
This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so101_follower", "total_episodes": 50, "total_frames": 12620, "total_tasks": 1, "total_videos": 100, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:50" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":…
View source information and access options.
This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so101_follower", "total_episodes": 500, "total_frames": 136475, "total_tasks": 1, "total_videos": 1000, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:500" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path
This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so101_follower", "total_episodes": 500, "total_frames": 137269, "total_tasks": 1, "total_videos": 1000, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:500" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path
A LeRobot-formatted robotics capture set for the so101_follower arm, comprising 51 episodes, 13,109 frames, and 102 videos at 30 fps stored across parquet chunks.
A LeRobot-generated dataset comprising 149 episodes, 27,628 frames, and 298 video files captured at 30 fps on a so101_follower robot, with all episodes assigned to the training split and stored in the chunk-based Parquet layout expected by the codebase.
Recorded with the LeRobot v2.1 schema on a 6-DOF SO-ARM101 manipulator, this dataset contains 2 episodes totaling 409 frames captured at 30 FPS in 640x480 resolution from two cameras, showing the arm retrieving a red square block and depositing it into a green square region on the right.
A compact compilation tracking the count of clinical trials across African countries, sourced from Our World in Data, formatted as parquet files, tagged under health, and re-indexed by Electric Sheep Africa with harmonized metadata and loading instructions to support discovery of African statistics.
A 1K-to-10K record CSV inventory of African cancer clinical trials contributed to the Electric Sheep Africa Hugging Face catalog for standardized discovery.
A Parquet file of annual inflation as measured by consumer prices for Africa, sourced from World Bank gender and regional statistics, indexed by Electric Sheep Africa for African data discovery and released in the 1K–10K row range.
A small-scale inventory of African cancer clinical trials, listed in the Electric Sheep Africa catalog on Hugging Face and tagged for discovery with uniform metadata and loading notes.
Compilation of worldwide development metrics spanning multiple indicators.
Standardized reference structures for United States Postal Service ZIP codes.
A corpus of SEC EDGAR filings in English (Apache 2.0), roughly a few hundred thousand company documents including annual and quarterly reports and exhibits. The filings text is organized for model pretraining and retrieval, but the same documents are the raw material for filings-based and disclosure signals.
An extensive archive of SEC EDGAR submissions from company filings, organized by CIK, document type, and date, and accessible through a live REST API at api.ai-analytics.org.
Project intelligence on energy infrastructure covering pipelines, commodity facilities, permits and construction milestones.
Over 12.9 million U.S. brands evaluated for presence in AI-driven search platforms such as ChatGPT, Claude, Perplexity, and Gemini, with the list expanding continually.
A platform that consolidates public filings, government, property and corporate records and normalizes them for analytical use.
A synthetic person-level dataset spanning the entire United States, including individual records together with linked place identifiers.
Structured numeric records parsed from corporate filings such as 10-K, 10-Q, and 8-K submissions to the U.S. SEC. The figures are packaged inside a compressed DuckDB instance made available through the open-source Datapond registry for fast querying.