Alternative data datasets
Browse by provider, category, marketplace, access, delivery and license
A collection of earnings call transcripts covering S&P 500 firms and other U.S. large-cap companies across the years 2005 through 2025, intended for financial analysis, NLP modeling, and sentiment research.
Crowd-sourced research platform pooling citizen and professional input to evaluate corporate behavior across operational, regulatory and ESG dimensions.
Public-safe distribution bundle from a local r19 counterexample resolution review of arXiv patents using a language model, containing review-related outputs, manifests, checksums, and provenance receipts, but excluding the original arXiv PDFs and any expert validation.
Country-level healthcare claims platform that traces patient pathways and results across hundreds of millions of records.
Extracted from SEC EDGAR DEF 14A proxy filings since 2015, the dataset contains over 500,000 structured executive compensation entries for S&P 500 and Russell 2000 issuers, capturing CEO and CFO pay, equity awards, bonuses, incentives, and totals keyed by CIK, ticker, and fiscal year.
Research and intelligence covering pharmaceutical products, clinical trials, patent filings and regulatory progress.
A small text-classification dataset by ClarusC64 that labels whether borrow-rate, short-interest and price signals cohere into a real short squeeze or diverge into a false signal. It is published on a model hub with an MIT license tag and a 10-row train split.
Farm-level agronomic records such as planting activity, harvest output, and field imagery sourced from connected agricultural operations.
GenAI-job-postings-Dataset is a small, US-focused text corpus of generative-AI and machine-learning job postings distributed in Parquet format. The dataset is published on the Hugging Face Hub under an unstated license and comprises a single train split of roughly 120 rows.
Records of initial public offerings launched on Indian markets between 2006 and 2025, with fields covering open and close dates, listing date, face value, issue price and size, lot size, first-day price, total shares offered and their allocation across anchor, NII, QIB and retail categories, minimum investment, and subscription figures for each investor class.
Professional services firm linking subscribers with vetted industry practitioners through structured interview engagements spanning multiple verticals.
Compensation benchmarks for executives and staff assembled from proxy disclosures filed by publicly traded companies.
Full access to the Copernicus Sentinel constellation archive, Sentinel-1 SAR, Sentinel-2 optical at 10 metres, and Sentinel-3 and Sentinel-5P atmosphere products, with OpenSearch APIs and on-demand processing.
Weekly pricing and occupancy for major cruise lines and itineraries: an advance read on cruise-operator demand.
A worldwide repository of startup financing information, tracking funding rounds, backers, and exit events.
Bitcoin and Ethereum exchange flow indicators and miner activity metrics for market positioning.
A dataset enabling audience construction by identifying consumers who have previously been to designated places or establishments.
A tabular extract of clinical study descriptions drawn from ClinicalTrials.gov on 5 February 2025, emphasizing eligibility, design, and objective fields for healthcare and machine learning research.
Vector embeddings of clinical-trial records, hundreds of thousands of entries in Parquet (Apache 2.0). Enables similarity and retrieval analysis of trial design, indication crowding and competitive positioning.
Racial Bias in inmate COMPAS reoffense risk scores for Florida (ProPublica)
DAPFAM patent documents, a curated sample of US patents in Parquet with text for retrieval and classification tasks (CC BY-NC-SA 4.0). Scopus of claim text for building patent-similarity and technology-mapping models.
A systematic crawl of business websites classified by sector and geography, giving store counts, opening hours, technology stacks and contact details per market.
Insights into the Global Supply Chain: Container Transportation Data
Consolidated site visit data rolled up across each operator's full domain portfolio to gauge overall reach and engagement momentum for businesses.
Competitive shelf intelligence, prices, promotions and assortment, continuously collected across retail websites in over 50 countries for brands and investors.
Non-traditional Chinese financial datasets include capital movement patterns and broader macroeconomic signals.
A startup intelligence source centered on Europe, documenting transactions, valuations, and regional ecosystems.
Cardholder spending information from a non-traditional panel, broken down by geography and merchant category.
Commercial healthcare dataset profiling practitioners, care sites, organizational relationships and medical procedures.
edgar_xbrl_companyfacts is a Hugging Face dataset published by DenyTranDFW that aggregates U.S. SEC EDGAR XBRL company-facts filings into a single Parquet file. The training split contains roughly 125 million rows of structured financial facts tagged by reporting period and unit.
DerivedFunction01's sec-filings-snippets is a public-records dataset of short text excerpts drawn from U.S. SEC filings, released on Hugging Face in Parquet format. It is sized for language-model fill-mask work and is not positioned as a commercial alternative-data product.
Cross-sector knowledge marketplace supplying curated expert calls and research support to corporate and investment decision-makers across more than twenty-five industries.
Continuously built knowledge graph indexing over thirty billion entities such as firms, offerings and individuals, accompanied by APIs that parse arbitrary web pages into structured JSON.
Social listening across numerous languages producing opinion and topic indicators from global online conversations.
ClinicalTrials.gov XML for studies registered between 2018 and 2024, parsed into CSV, hundreds of thousands of records under CC0. A clean, licensed point-in-time history of US clinical-trial registrations.
Building cost figures and benchmarks from Dodge construction data, organized by structure type, project scale and geographic area.
Weather and crop-management information designed to support planting decisions for farmers and commodity market participants.
Supplier payment history condensed into PAYDEX credit ratings for millions of businesses.
Granular loan data infrastructure that normalizes consumer credit performance for securitization analysis.
A large granted-patent text collection from the USPTO, tens of millions of records in Parquet (Apache 2.0). Broad coverage of claims and descriptions for IP-intensity and technology-trend measurement.
Worldwide advisory community providing pay-per-use expert conversations and panel-based research studies on a broad range of industry topics.
Use for Foundation Models, GNNs, and More
Satellite imagery paired with proprietary machine-learning models that label construction, infrastructure change and land use at scale.
This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "names": [ "shoulder_pan.pos", "shoulder_lift.pos", "elbow_flex.pos", "wrist_flex.pos", "wrist_roll.pos", "gripper.pos" ], "shape": [ 6… See the full description on the dataset page: https://huggingface.co/datasets/Edgarium/rangement_pq_20
This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "names": [ "shoulder_pan.pos", "shoulder_lift.pos", "elbow_flex.pos", "wrist_flex.pos", "wrist_roll.pos", "gripper.pos" ], "shape": [ 6… See the full description on the dataset page: https://huggingface.co/datasets/Edgarium/rangement_pq_20
This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "names": [ "shoulder_pan.pos", "shoulder_lift.pos", "elbow_flex.pos", "wrist_flex.pos", "wrist_roll.pos", "gripper.pos" ], "shape": [ 6… See the full description on the dataset page: https://huggingface.co/datasets/Edgarium/rangement_pq_20
This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so101_follower", "total_episodes": 50, "total_frames": 10333, "total_tasks": 1, "total_videos": 100, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:50" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":…
This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so101_follower", "total_episodes": 50, "total_frames": 13379, "total_tasks": 1, "total_videos": 100, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:50" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":…
This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so101_follower", "total_episodes": 709, "total_frames": 187905, "total_tasks": 1, "total_videos": 1418, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:709" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path
This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so101_follower", "total_episodes": 161, "total_frames": 42852, "total_tasks": 1, "total_videos": 322, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:161" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":
This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so101_follower", "total_episodes": 101, "total_frames": 26027, "total_tasks": 1, "total_videos": 202, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:101" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":
This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so101_follower", "total_episodes": 50, "total_frames": 12620, "total_tasks": 1, "total_videos": 100, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:50" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":…
This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so101_follower", "total_episodes": 500, "total_frames": 136475, "total_tasks": 1, "total_videos": 1000, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:500" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path
This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so101_follower", "total_episodes": 500, "total_frames": 137269, "total_tasks": 1, "total_videos": 1000, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:500" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path
Automotive datasets on vehicle prices, dealer stock and transaction volumes widely used by industry analysts.
Gaming-sector specialist providing analytics, datasets and modeling for online gambling, sports wagering and fantasy contests.
A licensed and comprehensive source covering political developments, economic conditions, and business activity within Brazil.
Fully licensed real-time international news content supplied by top-tier global wire services.
Fully licensed, real-time news feed covering business, financial, and corporate developments from major Japanese media outlets.
Prepared by Electric Sheep Africa, this MDPA-sourced dataset supplies 10 entries of hotel room occupancy rates for Mauritius over 2019-2023, delivered as ML-ready Parquet files accompanied by uniform Hugging Face metadata and source provenance.