Alternative data datasets
Browse by provider, category, marketplace, access, delivery and license
A tabular extract of clinical study descriptions drawn from ClinicalTrials.gov on 5 February 2025, emphasizing eligibility, design, and objective fields for healthcare and machine learning research.
Metadata for hundreds of thousands of clinical trials in English, French and Spanish, structured in Parquet with study-level fields such as phase, condition, sponsor and status (Apache 2.0). Trial starts and phase transitions lead clinical development, so this works as an early indicator of biotech pipeline activity.
Parsed eligibility criteria from 13,229 ClinicalTrials.gov studies representing the candidate pool gathered by two first-stage retrievers during TREC Clinical Trials 2021–2023, formatted as typed entity-relation graphs for reranking models.
This dataset, published by 2001jdev on Hugging Face, packages eligibility-criteria information from clinical trials into structured entity and relation records. It is distributed in Parquet format and contains a single split of roughly 104,000 rows.
The clinical-trials-patient-graphs dataset is a small tabular collection of patient-level clinical records distributed in Parquet format by the 2001jdev publisher. It organizes entities and relations extracted from trial data into structured rows spanning 2021 through 2023.
This dataset is a parsed snapshot of ClinicalTrials.gov records, distributed in Parquet format on the Hugging Face Hub under an unspecified license. It contains roughly 52,000 trial entries drawn from public registry filings.
clinical-trials-trec-qrels is a tabular relevance-judgment file distributed by the 2001jdev user on the Hugging Face Hub. It maps clinical-trial topic identifiers to NCT registry IDs with graded relevance scores.
The clinical-trials-trec-topics dataset is a small tabular collection of clinical-trial topic records distributed in Parquet format on the Hugging Face Hub. It appears to be derived from TREC clinical-trial retrieval topics, with rows partitioned by year.
clinical-trials-v2 is a Hugging Face dataset published by chemNLP containing processed ClinicalTrials.gov records stored in Parquet. It exposes three columns — filename, xml, and text — drawn from the official clinical-study XML schema.
Vector embeddings of clinical-trial records, hundreds of thousands of entries in Parquet (Apache 2.0). Enables similarity and retrieval analysis of trial design, indication crowding and competitive positioning.
Dattito's clinical-trials-data is a Parquet-format dataset of roughly 24 million clinical trial records sourced from public trial registries. It aggregates identifiers, conditions, and standardized medical terminology into a single table for large-scale analysis.
ClinicalTrials.gov XML for studies registered between 2018 and 2024, parsed into CSV, hundreds of thousands of records under CC0. A clean, licensed point-in-time history of US clinical-trial registrations.
The clinical-trials-xml-2018-2024 dataset on Hugging Face, published by user jvaton, distributes a small Parquet collection of clinical-trial related XML records. It contains fewer than one thousand rows and is licensed on an unstated basis.
Minimal information is provided on this clinical-trials-embedded dataset, with additional details needed.
An analysis-ready compilation of roughly 3,000 ClinicalTrials.gov studies registered between 2000 and 2025 that involve AI, machine learning, or digital-health tools, enriched with 30 LLM-derived variables covering use case, therapeutic area, sponsor composition, trial phase, deployment score, evidence strength, and responsible-AI keyword indicators.
Clinical-trial protocol documents from Parexel, tens of thousands of studies. Protocol text, covering inclusion criteria, endpoints and timelines, is a leading indicator for enrollment pace and pipeline risk.
Clinical Trials Data Ingestion & ETL Pipeline -ClinicalTrials.gov API v2
A 10% sample of over 573,000 ClinicalTrials.gov studies augmented with AI-classified therapeutic areas, sponsor categorization, outcome groupings, and duration metrics, offered for portfolio and market landscape analysis.
A 10% sample of a sponsor profiling resource containing entries for over 10,200 clinical trial sponsors, covering trial counts, completion ratios, phase distribution, therapeutic areas, and pipeline activity.
Over 635,000 registered clinical studies from ClinicalTrials.gov together with a supplementary pharma intelligence table of roughly 1.06 million rows that map those trials to FDA drug approvals, distributed as part of an open data initiative for academic and personal use.
The TREC Clinical Trials collections released in 2021, 2022, and 2023 are available through the TREC homepage, with corresponding papers for each year. Artur Guimarães ([email protected]) serves as the dataset curator, while inquiries regarding the original authors can be directed separately. The dataset addresses the task of linking patient profiles with appropriate clinical trials, containing entries formatted as JSON objects with fields including query identifiers, disease labels, and accompanying clinical text.
China Multiple Myeloma Clinical Trials — Open Dataset Multiple myeloma clinical trials registered in China, curated from official NMPA / CDE filings by the China Myeloma Digital Network (CMDN), an independent non-profit patient advocacy organisation. This is a mirror. The citable version of record lives at doi.org/10.5281/zenodo.22690814; the source repository is chinamyeloma/china-myeloma-clinica
A compact compilation tracking the count of clinical trials across African countries, sourced from Our World in Data, formatted as parquet files, tagged under health, and re-indexed by Electric Sheep Africa with harmonized metadata and loading instructions to support discovery of African statistics.
A small-scale inventory of African cancer clinical trials, listed in the Electric Sheep Africa catalog on Hugging Face and tagged for discovery with uniform metadata and loading notes.
A CC-BY-4.0 licensed compilation from Ozari Health cataloging 22 major published GLP-1 clinical trials as of May 2026, consolidating peer-reviewed trial-level data into a single structured index.
Data for the COVID-19 clinical trials registered in the US
A registry and results database of more than 450,000 clinical trials worldwide, with protocols, sites, enrollment counts, outcomes and sponsor details, queryable through a free REST API and bulk downloads — the reference source for pharmaceutical pipeline tracking.
A queryable DuckDB mirror of the AACT flat-file export of ClinicalTrials.gov, packaged via the clinicaltrials-database project and covering every registered trial across 48 tables totaling roughly 58 million rows.
Clinical-trials records combining structured metadata and narrative text, hundreds of thousands of studies in Parquet. An early read on pharmaceutical pipeline activity and trial design trends.
A structured compilation of 29,633 completed Phase 3 clinical trials sourced from ClinicalTrials.gov, including status and design details for each study record.
Iteration 2 of a clinical-trials corpus intended for embedding-model training and fine-tuning, enabling tasks such as ranked retrieval, document comparison across two or more anchors, and anchor-versus-chunk similarity scoring, distinguished from the prior release by finer-grained chunk segmentation.
A knowledge graph assembled from 575,778 ClinicalTrials.gov registrations and their associated arms, outcomes, sites, sponsors, conditions, interventions, MeSH codes, and PubMed citations, containing 7,628,735 nodes and 15,531,427 edges, delivered as 623 MB of Parquet files (compared to about 7 GB of raw JSON) and built using Samyama Graph.