Alternative data datasets
Browse by provider, category, marketplace, access, delivery and license
A healthcare claims dataset containing detailed records on patients, care providers, and the medical services rendered.
A historical archive focused on medical and surgical practices from 1847 through 1898 during the Civil War era.
The inaugural issue of what is considered the earliest medical journal published in America, dated 1797.
A curated collection designed to function as a retrieval-augmented generation knowledge base and a resource for historical research.
Medical claims sourced from a clearinghouse network, reflecting more than one-third of healthcare interactions occurring within the United States.
View source information and access options.
View source information and access options.
View source information and access options.
View source information and access options.
View source information and access options.
View source information and access options.
View source information and access options.
Records documenting medical practices and treatments at the height of the Victorian period.
A leading Canadian medical publication covering the years 1911 through 1930.
Pre-1930 textual material on genetics and inheritance that has been cleaned and formatted for machine learning use.
Curated and audited corpus drawing from two leading obstetrics and gynecology journals, formatted for artificial intelligence workflows.
Curated historical clinical text that has been cleaned and checked for bias to support better model reasoning and factual grounding.
Well-suited for embedding-based search and retrieval-augmented generation across a unified historical text collection.
Historical psychological writing reviewed for demographic and social biases, intended to support language model training and evaluation.
Sanitized archive of past medical and scientific periodicals reviewed for bias, intended to strengthen grounded reasoning in language models.
Historical record of American medical periodical publishing during the 1800s.
This collection shows high suitability for machine learning model development and retrieval-augmented generation workflows.
View source information and access options.
View source information and access options.
View source information and access options.
View source information and access options.
View source information and access options.
View source information and access options.
View source information and access options.
View source information and access options.
View source information and access options.
View source information and access options.
View source information and access options.
View source information and access options.
View source information and access options.
A neurology-focused clinical dataset containing detailed patient information intended to support both medical treatment and scientific investigation.
Parsed eligibility criteria from 13,229 ClinicalTrials.gov studies representing the candidate pool gathered by two first-stage retrievers during TREC Clinical Trials 2021–2023, formatted as typed entity-relation graphs for reranking models.
The TREC Clinical Trials collections released in 2021, 2022, and 2023 are available through the TREC homepage, with corresponding papers for each year. Artur Guimarães ([email protected]) serves as the dataset curator, while inquiries regarding the original authors can be directed separately. The dataset addresses the task of linking patient profiles with appropriate clinical trials, containing entries formatted as JSON objects with fields including query identifiers, disease labels, and accompanying clinical text.
Track healthy organs in medical scans to improve cancer treatment
China Multiple Myeloma Clinical Trials — Open Dataset Multiple myeloma clinical trials registered in China, curated from official NMPA / CDE filings by the China Myeloma Digital Network (CMDN), an independent non-profit patient advocacy organisation. This is a mirror. The citable version of record lives at doi.org/10.5281/zenodo.22690814; the source repository is chinamyeloma/china-myeloma-clinica
A tabular extract of clinical study descriptions drawn from ClinicalTrials.gov on 5 February 2025, emphasizing eligibility, design, and objective fields for healthcare and machine learning research.
Vector embeddings of clinical-trial records, hundreds of thousands of entries in Parquet (Apache 2.0). Enables similarity and retrieval analysis of trial design, indication crowding and competitive positioning.
Over twenty million adverse event reports covering drugs and medical devices are consolidated, refreshed daily, and structured for analysis.
Structured records of FDA medical device adverse event reports, formatted to support safety monitoring and risk evaluation workflows.
Continuously refreshed, query-ready dataset covering county employment figures, medically underserved regions, and Federally Qualified Health Center locations.
Dattito's clinical-trials-data is a Parquet-format dataset of roughly 24 million clinical trial records sourced from public trial registries. It aggregates identifiers, conditions, and standardized medical terminology into a single table for large-scale analysis.
Commercial healthcare dataset profiling practitioners, care sites, organizational relationships and medical procedures.
A collection examining surgical practice during a period of technological and medical advancement.
Ready-to-use collection suited for retrieval-augmented question answering on historical medical topics.
Globally recognized peer-reviewed medical publication of the highest standing.
The oldest medical periodical still in circulation across the United States has been prepared for modern artificial intelligence use cases.
A nineteenth-century medical review publication that summarized each year's clinical literature for practicing physicians.
A dataset encompassing physician services, inpatient and outpatient care, prescriber information, medical equipment, prosthetics, and acute care categories.
Yearly and quarterly pharmaceutical revenue figures along with analyst consensus projections, broken down by product, geography, and medical condition.
Medical and health-related records assembled to assist insurance providers and organizations that assume financial risk.
IQVIA's Pharmetrics medical-claims dataset marketed on Snowflake Marketplace, offering United States healthcare billing records for pharmaceutical and biotechnology research.
Metadata for hundreds of thousands of clinical trials in English, French and Spanish, structured in Parquet with study-level fields such as phase, condition, sponsor and status (Apache 2.0). Trial starts and phase transitions lead clinical development, so this works as an early indicator of biotech pipeline activity.
An analysis-ready compilation of roughly 3,000 ClinicalTrials.gov studies registered between 2000 and 2025 that involve AI, machine learning, or digital-health tools, enriched with 30 LLM-derived variables covering use case, therapeutic area, sponsor composition, trial phase, deployment score, evidence strength, and responsible-AI keyword indicators.
A queryable DuckDB mirror of the AACT flat-file export of ClinicalTrials.gov, packaged via the clinicaltrials-database project and covering every registered trial across 48 tables totaling roughly 58 million rows.
FDA-regulated products and events exposed as structured JSON: adverse drug-event reports, device recalls and registrations, labels, the NDC directory and food enforcement — used for pharmacovigilance and regulatory-risk screens.