Alternative data datasets
Browse by provider, category, marketplace, access, delivery and license
View source information and access options.
View source information and access options.
View source information and access options.
View source information and access options.
View source information and access options.
View source information and access options.
View source information and access options.
View source information and access options.
View source information and access options.
View source information and access options.
View source information and access options.
Public-sector finances — tax receipts, budget execution and fiscal indicators — across multiple jurisdictions.
PatentMatch pairs US patents with matching reference patents for retrieval evaluation, a few thousand records in JSON (Apache 2.0). A niche evaluation set for patent-similarity and citation-link models.
Governance intelligence on boards of thousands of companies: directors, backgrounds, gender and skills.
S&P 500 earnings-call transcripts in Parquet under MIT, tens of thousands of calls. Independently compiled and comparable to other public transcript corpora for sentiment and tone work.
View source information and access options.
View source information and access options.
Compiled summaries of FEMA-declared disasters along with trend breakdowns for emergency management research.
Disclosure documents submitted by institutional investment firms that manage portfolios exceeding $100 million under Section 13(f).
View source information and access options.
A reference collection drawn from the U.S. Patent Phrase to Phrase Matching Kaggle competition, containing additional details that are still required for full documentation.
A reconstructed EDGAR corpus of company filings in English (Apache 2.0): a large sample of annual and quarterly reports and exhibits, hundreds of thousands of records. Comparable to other public SEC text sets and a solid starting point for filings-based research.
Cadenza-Labs/apollo-llama3.3-insider-trading-generations is a small text dataset on Hugging Face containing 1,660 training rows used to fine-tune a Llama 3.3 model for detecting dishonest messages. It is distributed in Parquet format and released under an unstated license.
Includes all company names and cik keys from the SEC database.
titer · EDGAR officer corpus 4,206,080 attested person–company–role–date tuples from SEC Forms 3/4/5, published as pointers rather than records, alongside the frozen pre-registrations that were hash-published before any measurement ran. edgar_officers.parquet: 4.2M rows, 230,405 distinct people, 20,266 issuers, 2006q1–2026q2. Column Meaning accession SEC accession number, the pointer that reconstr
Machine-readable XBRL filings covering 10-K and 10-Q reports, cleaned and standardized so financial figures can be extracted reliably.
Drawn from SEC EDGAR's full archive of millions of filings spanning all form types and US public companies, this sample isolates 1,000 recent 8-K material event submissions together with their filing metadata and document references.
A sample of US patent titles, abstracts and CPC classification labels built for multi-class patent classification, tens of thousands of text records in Parquet. Useful as training material for mapping innovation activity onto technology categories over time.
Millions of consumer complaints about financial products — mortgages, credit cards, debt collection and loans — with product, issue, company response and ZIP-level geography, together with HMDA mortgage disclosures.
A weekly dataset tracking positions held by commercial and non-commercial participants across futures, provided in legacy, disaggregated, and financial-futures formats for agricultural, energy, metals, and financial contracts.
clinical-trials-v2 is a Hugging Face dataset published by chemNLP containing processed ClinicalTrials.gov records stored in Parquet. It exposes three columns — filename, xml, and text — drawn from the official clinical-study XML schema.
Chinese patents paired with their abstractive summaries in Mandarin, a few thousand records in JSON (Apache 2.0). A window into the pace and direction of patenting inside the Chinese technology base.
View source information and access options.
View source information and access options.
Extracted from SEC EDGAR DEF 14A proxy filings since 2015, the dataset contains over 500,000 structured executive compensation entries for S&P 500 and Russell 2000 issuers, capturing CEO and CFO pay, equity awards, bonuses, incentives, and totals keyed by CIK, ticker, and fiscal year.
List of CMBS deals with publicly available Schedule AL data in SEC.gov EDGAR
Records of initial public offerings launched on Indian markets between 2006 and 2025, with fields covering open and close dates, listing date, face value, issue price and size, lot size, first-day price, total shares offered and their allocation across anchor, NII, QIB and retail categories, minimum investment, and subscription figures for each investor class.
Compensation benchmarks for executives and staff assembled from proxy disclosures filed by publicly traded companies.
View source information and access options.
View source information and access options.
Consolidated and enhanced real estate parcel data available natively within the Snowflake environment.
View source information and access options.
View source information and access options.
View source information and access options.
A tabular extract of clinical study descriptions drawn from ClinicalTrials.gov on 5 February 2025, emphasizing eligibility, design, and objective fields for healthcare and machine learning research.
Vector embeddings of clinical-trial records, hundreds of thousands of entries in Parquet (Apache 2.0). Enables similarity and retrieval analysis of trial design, indication crowding and competitive positioning.
Demographic and consumer indicators compiled from the Data Axle Consumer Database, broken down to U.S. census tracts and block groups.
Regional consumer and demographic metrics sourced from the Data Axle Consumer Database, aggregated at the U.S. county and ZIP code scale.
Population characteristics of people who relocated to a new residence within the last 30 days.
A centralized repository offering curated datasets drawn from New Zealand's comprehensive data lake.
A starting point for analytics and information drawn from datasets covering the Australian market.
View source information and access options.
Comprehensive company-level information for entities registered in the United Kingdom.
Directory of principal contacts associated with businesses registered in the United Kingdom.
DAPFAM patent documents, a curated sample of US patents in Parquet with text for retrieval and classification tasks (CC BY-NC-SA 4.0). Scopus of claim text for building patent-similarity and technology-mapping models.
Two decades of block-level commuting patterns, worker demographics, and job attributes across every U.S. state.
All 8-k filings for 2024 with augmented fields.
View source information and access options.
View source information and access options.
View source information and access options.