Alternative data datasets
Browse by provider, category, marketplace, access, delivery and license
This dataset is a small parquet-format subset of clinical-trial eligibility criteria represented as entity and relation graphs. It is published on a community data hub under an unspecified license.
The clinical-trials-patient-graphs dataset is a small tabular collection of patient-level clinical records distributed in Parquet format by the 2001jdev publisher. It organizes entities and relations extracted from trial data into structured rows spanning 2021 through 2023.
The clinical-trials-trec-topics dataset is a small tabular collection of clinical-trial topic records distributed in Parquet format on the Hugging Face Hub. It appears to be derived from TREC clinical-trial retrieval topics, with rows partitioned by year.
A small text-classification dataset by ClarusC64 that labels whether borrow-rate, short-interest and price signals cohere into a real short squeeze or diverge into a false signal. It is published on a model hub with an MIT license tag and a 10-row train split.
GenAI-job-postings-Dataset is a small, US-focused text corpus of generative-AI and machine-learning job postings distributed in Parquet format. The dataset is published on the Hugging Face Hub under an unstated license and comprises a single train split of roughly 120 rows.
An artificially generated collection of apparel product listings and accompanying advertisements produced by prompting GPT-4 to invent one hundred clothing items with descriptions and then write promotional copy for each, output in a structured product and description format without subsequent manual verification.
stk-sec-filings is a small Hugging Face dataset by publisher deerfieldgreen that consolidates U.S. SEC filings into a single Parquet resource. The file covers fewer than one thousand rows distributed across a training split, with no stated update cadence or refresh policy.
MolMole_Patent300 is a curated evaluation benchmark for extracting chemical information from full patent pages, supporting end-to-end testing of molecule detection, reaction parsing, and optical chemical structure recognition.
View source information and access options.
View source information and access options.
View source information and access options.
View source information and access options.
A robotics dataset generated with the LeRobot framework contains two recorded episodes totaling 897 frames at 30 frames per second, captured using an SO100 robot performing a single task, with four video files and accompanying parquet data organized in a single chunk of up to 1,000 entries and partitioned as a train split covering episodes zero through one.
View source information and access options.
View source information and access options.
View source information and access options.
Recorded with the LeRobot v2.1 schema on a 6-DOF SO-ARM101 manipulator, this dataset contains 2 episodes totaling 409 frames captured at 30 FPS in 640x480 resolution from two cameras, showing the arm retrieving a red square block and depositing it into a green square region on the right.
Prepared by Electric Sheep Africa, this MDPA-sourced dataset supplies 10 entries of hotel room occupancy rates for Mauritius over 2019-2023, delivered as ML-ready Parquet files accompanied by uniform Hugging Face metadata and source provenance.
A small Parquet-format collection from Electric Sheep Africa reporting room occupancy rates for large hotels in Mauritius from 2019 to 2023, containing ten records drawn from MDPA and packaged with Hugging Face metadata for machine learning workflows.
Assembled by Electric Sheep Africa from MDPA, this dataset provides 60 rows of monthly hotel room occupancy rates for Mauritius spanning 2019-2023, formatted as ML-ready Parquet with consistent Hugging Face metadata and source attribution.
Compiled by Electric Sheep Africa, this MDPA-derived dataset offers 20 records of quarterly hotel room occupancy rates for Mauritius across the 2019-2023 period, distributed in ML-ready Parquet format with standardized Hugging Face metadata and source traceability.
Electric Sheep Africa releases 60 quarterly observations from MDPA on hotel room occupancy for large establishments in Mauritius from 2019 through 2023, formatted as Parquet with consistent metadata.
A legacy archive of Iran's consumer price index and inflation series from 1936 to 2022 remains available, though an updated multisource version using current Farmaanaa methodology supersedes it.
A curated collection of earnings call recordings broken into intelligent segments and paired with high-accuracy transcripts produced by Mistral's Voxtral model, intended for ASR benchmarking, transcription quality studies, and audio-to-text alignment research.
The patent dataset, published by HypernetworkRG, is a small graph dataset consisting of one row in a single training split. It is distributed without a stated license or size and is tagged for text modality.
sec_filings is a small text dataset of 494 rows that pairs natural-language queries about U.S. securities filings with retrieved document facts and ground-truth answers. It is distributed under an unstated license on a public dataset hub.
A distilabel-built set of restaurant reviews packaged with a pipeline.yaml for reproducing the synthetic generation workflow via the distilabel CLI.
An API-based ClinicalTrials.gov scraper that retrieves over 585,000 study records including phase, sponsor, conditions, interventions, enrollment, eligibility, and sites, filterable by condition, sponsor, and status.
patent-desc-test is a small text dataset on Hugging Face Datasets that contains short identifier-style strings derived from patent document numbers. The dataset is published by user matthias-ehrlich and is broadly tagged as public records–adjacent text data.
A benchmark collection designed to test vision-language models on the visual conventions found in U.S. design patent illustrations. It pulls 3.6 million figures from 2007 to 2022 in the IMPACT archive and pairs them with text drawn from PatentsView, stored in yearly Parquet files along with an 800-sample evaluation set.
FinSight gathers plain-text 10-K and 10-Q SEC filings for twenty major US public firms spanning six sectors, yielding 97 records intended for BERT fine-tuning, retrieval-augmented generation, and multi-agent financial research.
The patent_classification dataset is a small text corpus hosted on the Hugging Face Hub by publisher pgurazada1, comprising paired patent text snippets and their cooperative classification labels. It is distributed as a CSV file under an unstated license.
A compilation of 150 restaurant reviews for casual and fine dining venues, drawn from the publicly available DineScope project under a CC0 1.0 license, with a production-tier usage designation and a 2026 sync date.
A structured compilation of 29,633 completed Phase 3 clinical trials sourced from ClinicalTrials.gov, including status and design details for each study record.
jobpostingsamples is a small sample dataset of US job postings published by talanAI. It is distributed in CSV format and is tagged for use with the Hugging Face datasets, pandas, polars, and mlcroissant libraries.
Techsalerator's offering merges anonymized mobility signals from various providers to map population movement and location visits in urban cores, business districts, transit corridors, and broader regions.
chem-patent-eval-assets is a small image-folder dataset published under the vietmed namespace, evidently paired with a chemistry-patent evaluation task. The license and broader provenance are not stated, and only 363 training rows are documented on its dataset card.
View source information and access options.
The tech-job-postings-salaries-sample dataset is a 500-row public sample of a larger scraping-based feed of technology job postings drawn from public applicant tracking system endpoints such as Greenhouse, Lever, Ashby, Workable and SmartRecruiter. It is published by DataForge (Zalize) for evaluation and preview use.