Alternative data datasets
Browse by provider, category, marketplace, access, delivery and license
EDGAR M&A Deal Events (2010–2024) 4,156 U.S. public-company acquisition events, discovered directly from SEC EDGAR's own quarterly filing indexes — not scraped from a vendor list or a blog. Each row is a company that filed a merger proxy or tender-offer response between 2010 and 2024, with the earliest such filing's date as an announcement-date proxy. Built as a byproduct of an M&A target-predicti
This is a Techsalerator dataset on Real Estate & Airbnb Data for UAE
Text of SEC filings pulled from EDGAR spanning millions of records in English (Apache 2.0): 10-K, 10-Q, 8-K, annual reports and exhibits, tagged for text-generation and classification. A practical base corpus for earnings- and filings-driven signals, from risk factors and MD&A to management commentary.
View source information and access options.
View source information and access options.
Online and In-Person Building Permit Submissions in Fiscal Years
Las Vegas Building Permit Issuance Details
View source information and access options.
A relational database mapping consumer attributes and connections across various touchpoints.
8 size-normalized financial ratios for ~440 US companies from SEC EDGAR 10-Ks
Raw 10-K financials for 442 US companies extracted from SEC EDGAR
Provides highly reliable and thorough records of names, addresses, telephone numbers, and email addresses.
A collection of Parquet files providing one record per Chinese invention patent family granted between 2003 and 2019, capturing a cumulative quality assessment pipeline that includes metrics such as semantic knowledge recombination, strict historical semantic novelty, and IPC-based robustness measures across global patent publications.
Governance, beneficial-ownership and related-party transaction intelligence (Tussell) used for institutional investment research.
View source information and access options.
View source information and access options.
View source information and access options.
View source information and access options.
View source information and access options.
View source information and access options.
View source information and access options.
Collection of anchor-positive pairs built from concatenated title, summary, and inclusion criteria fields for individual clinical trials, where each combined text is paired with four related questions for embedding model fine-tuning.
Iteration 2 of a clinical-trials corpus intended for embedding-model training and fine-tuning, enabling tasks such as ranked retrieval, document comparison across two or more anchors, and anchor-versus-chunk similarity scoring, distinguished from the prior release by finer-grained chunk segmentation.
A final anchor-positive pair dataset designed for fine-tuning embedding models on clinical trials, merging consolidated title-summary-inclusion chunk pairs with five-anchor three-positive chunk pairings.
US-GAAP Financial Data: SEC Filings of Listed US Companies Since January 2009
Normalized financial statements drawn from the SEC Financial Statement Data Sets and company-facts services, both built on the EDGAR APIs and spanning reported results across filers.
A knowledge graph assembled from 575,778 ClinicalTrials.gov registrations and their associated arms, outcomes, sites, sponsors, conditions, interventions, MeSH codes, and PubMed citations, containing 7,628,735 nodes and 15,531,427 edges, delivered as 623 MB of Parquet files (compared to about 7 GB of raw JSON) and built using Samyama Graph.
Educational outcome records from the U.S. Department of Education College Scorecard, including graduate salary figures.
Complete IPEDS database records spanning 2004 through the present day.
Current and former military job titles across all branches of service, kept up to date.
Mapping that links military job classifications to their civilian equivalents in the O*NET system.
Labor market role and occupation reference information published by the U.S. Department of Labor's O*NET program.
A historical, quarterly index of U.S. public company 10-K and 10-Q filings
View source information and access options.
chem-patent-eval-assets is a small image-folder dataset published under the vietmed namespace, evidently paired with a chemistry-patent evaluation task. The license and broader provenance are not stated, and only 363 training rows are documented on its dataset card.
A snapshot of processed ClinicalTrials.gov records filtered to completed-recruitment Phase 3 and Phase 4 studies, plus trials of medical devices or behavioral interventions, totaling 156,887 trials as of March 2025, used for an agentic medical retrieval and evidence-grounding system.
A 10% sample of over 573,000 ClinicalTrials.gov studies augmented with AI-classified therapeutic areas, sponsor categorization, outcome groupings, and duration metrics, offered for portfolio and market landscape analysis.
A 10% sample of a sponsor profiling resource containing entries for over 10,200 clinical trial sponsors, covering trial counts, completion ratios, phase distribution, therapeutic areas, and pipeline activity.
View source information and access options.
This dataset includes completed or partially completed building permits.
10-K annual filings from EDGAR in Parquet, tens of thousands of documents. The standard annual-disclosure corpus for fundamentals, risk-factor and management-discussion analysis.
View source information and access options.
Enriched entity records and linking attributes that map corporations, individuals and locations together to power graph-style analysis.
View source information and access options.
patent-spec-xml is a dataset of patent specification documents in XML format, published by Yehoon and licensed under unspecified terms. It is catalogued on the alternative-data directory as a public-records resource.
Real estate listings collected in the first 6 months in 2021
The 2024 Earnings Call Transcript dataset, published by yeong-hwan on Hugging Face, aggregates quarterly earnings call dialogue for individual stock tickers. It is distributed as a JSON table with 1,904 rows in a single train split.
Over 635,000 registered clinical studies from ClinicalTrials.gov together with a supplementary pharma intelligence table of roughly 1.06 million rows that map those trials to FDA drug approvals, distributed as part of an open data initiative for academic and personal use.
Derived from USPTO and PatentsView bulk releases, this dataset catalogs 9.1 million U.S. patents from 1976 onward along with complete citation links, assignees, inventors, and CPC codes, with the full-graph edition exceeding 255 million rows.
Fifteen opt-in signals indicating ACA open enrollment intent, drawn from roughly 59.6 million U.S. consumers.
Eighteen opt-in signals reflecting aging- and retirement-related intent across 37.4 million U.S. consumers.
Twenty-three opt-in automotive purchase-intent indicators spanning roughly 24.2 million people in the United States.
Five opt-in automotive defect-intent indicators covering about 1.2 million U.S. individuals.
A compilation of six opt-in cancer-related intent indicators drawn from roughly 8.9 million individuals located in the United States.
Twenty-nine opt-in signals covering borrowing, lending, and credit-related consumer behavior across an estimated 44.8 million adults in the United States.
Twenty-five opt-in indicators reflecting general health interests drawn from roughly 16.5 million American individuals.
Seventeen opt-in Medicare-related intent signals drawn from a U.S. consumer base of approximately 33.6 million individuals.
Twenty opt-in signals related to metabolic health intent covering roughly 31.7 million U.S. consumers.
Thirty-three opt-in prescription intent indicators drawn from roughly 38 million American individuals.
Thirty-seven opt-in household purchase propensity indicators covering approximately 36.1 million individuals in the United States.