Alternative data datasets
Browse by provider, category, marketplace, access, delivery and license
Assembled by Electric Sheep Africa from MDPA, this dataset provides 60 rows of monthly hotel room occupancy rates for Mauritius spanning 2019-2023, formatted as ML-ready Parquet with consistent Hugging Face metadata and source attribution.
Compiled by Electric Sheep Africa, this MDPA-derived dataset offers 20 records of quarterly hotel room occupancy rates for Mauritius across the 2019-2023 period, distributed in ML-ready Parquet format with standardized Hugging Face metadata and source traceability.
Electric Sheep Africa releases 60 quarterly observations from MDPA on hotel room occupancy for large establishments in Mauritius from 2019 through 2023, formatted as Parquet with consistent metadata.
A U.S. patent collection with a minimal placeholder card and no additional documentation provided.
A curated collection of earnings call recordings broken into intelligent segments and paired with high-accuracy transcripts produced by Mistral's Voxtral model, intended for ASR benchmarking, transcription quality studies, and audio-to-text alignment research.
edgar-corpus-embeddings is a text and embedding corpus derived from SEC EDGAR filings, published on Hugging Face by user gagan3012. It packages sectioned 10-K content with pre-computed vector embeddings for machine learning workflows.
A cleaned collection of French patent documents published from 1981 through 2026, each row in the Parquet file representing one complete A1 publication parsed from the original XML. It was assembled independently by a contributor with direct access to the public patent source and is optimized for distributed loading.
Raw French patent publications from 2020 through 2026, pulled from original A1 XML records by an independent API/FTP extraction and delivered one-document-per-row in streaming-ready parquet.
A dataset card entry for restaurant reviews that flags the need for additional information beyond what is currently provided.
10-K filings for S&P 500 companies retrieved from EDGAR, stored as Parquet with the filing text attached (MIT). A focused corpus for the large-cap annual disclosure cycle: risk factors, MD&A and segment detail for every S&P 500 name.
The mintic_linkedin-job-postings dataset aggregates LinkedIn job posting text into a single string column with roughly 124,000 rows. It is published on Hugging Face by user jmparejaz under an unstated license.
Metadata for every EDGAR filing submitted to the U.S. Securities and Exchange Commission from 1994 through December 14, 2024, derived from the quarterly master index files and including CIKs, issuer names, and form types.
A distilabel-built set of restaurant reviews packaged with a pipeline.yaml for reproducing the synthetic generation workflow via the distilabel CLI.
The patents-publications-dataset is a large-scale tabular collection of patent publication records distributed in Parquet format. Hosted by publisher labofsahil on the Hugging Face Hub under an unstated license, it is sized in the 100M–1B row range.
An API-based ClinicalTrials.gov scraper that retrieves over 585,000 study records including phase, sponsor, conditions, interventions, enrollment, eligibility, and sites, filterable by condition, sponsor, and status.
Minimal information is provided on this clinical-trials-embedded dataset, with additional details needed.
A corpus of around 1.3 million U.S. patent filings paired with human-authored abstractive summaries, organized into nine Cooperative Patent Classification categories ranging from human necessities to textiles and paper.
The clinical_trials dataset distributes ClinicalTrials.gov registry records in Parquet format for offline analysis. It contains a single training split of roughly 150,000 rows sourced from the public clinical-trial registry.
US Patent Descriptions This dataset contains the descriptions of granted US utility patents, filtered and deduplicated.The original data comes from all granted patents in 2025 up to May 20, available from PatentsView. Splits train: 10,000 rows for model training validation: 2,500 rows for validation test: 2,500 rows for evaluation Columns patent_id: Identifier for the patent; useful for reconcilin
A collection featuring the primary independent claim of United States utility patents issued between 2005 and 2025 is provided. Because the opening claim usually establishes the widest legal boundaries of an invention, it carries particular significance, and each entry includes the patent identifier alongside its claim text.
View source information and access options.
A subset of the BIGPATENT corpus adapted for the MTEB clustering benchmark, tens of thousands of English patents and titles under CC BY 4.0. Primarily a research and evaluation set for similarity and clustering workloads.
The restaurant_review_sentiment dataset, published on the Hugging Face MTEB hub, supplies a small Arabic-language corpus of restaurant reviews annotated for sentiment. It is intended for evaluating embedding models rather than for production analytics.
google-patents-data-preview is a community-uploaded preview of Google Patents bibliographic records, distributed in Parquet format through the Hugging Face Hub. The single training split contains roughly 340,000 rows covering patent identifiers, classifications, and localized text fields.
Global Job Postings – Multi-ATS is a single dated snapshot of 112,816 job listings aggregated from 23 applicant tracking systems and parsed into a 47-column schema by NextGig-Rocks. It is distributed as Parquet via the Hugging Face datasets library and targets labor-market research, not production job boards.
Over 1.3 million granted US patents from the BIGPATENT benchmark, each pairing its claims with an abstractive summary and full description text in English (CC BY 4.0, Parquet). Use it to track the technology areas where applicants are filing and to study how claim language has evolved over time.
phase1-metadata is a tabular dataset published by open-edgar-sec that consolidates SEC filing-level and entity-level attributes for publicly registered companies. It is distributed as parquet with a single training split of 37,547 rows and targets analysts building reproducible pipelines over public filings.
TUDelft-Electricity-Consumption-1.0 is a high-resolution, open-source time series dataset of household electricity load published on Hugging Face by OpenSynth. It draws on multi-country trials covering thousands of households under different tariff regimes and is distributed in Parquet format.
Clinical-trials records combining structured metadata and narrative text, hundreds of thousands of studies in Parquet. An early read on pharmaceutical pipeline activity and trial design trends.
Clinical-trial protocol documents from Parexel, tens of thousands of studies. Protocol text, covering inclusion criteria, endpoints and timelines, is a leading indicator for enrollment pace and pipeline risk.
ClinicalTrialSummary is a Hugging Face dataset published under the pat-jj account that pairs long-form clinical trial article text with shorter plain-language summaries. It is distributed in Parquet format with predefined train, validation, and test splits totaling 77,516 rows.
The ClinicalTrialSummary_Full resource lacks further documentation, with no additional details supplied.
Patent_FR_US_Merge_Radix_65536 is a Hugging Face dataset by publisher PhysiQuanty that merges French and US patent documents into a single text corpus. It is distributed in Parquet format and is tagged with an Apache-2.0 license, though the README provides no further description.
Dataset Card for "zephyr_pi0_gen_57k_for_offline_dpo_ipo" More Information needed
clinicaltrials.gov-summary_and_eligibility is a small Parquet mirror of selected ClinicalTrials.gov records, distributed on the Hugging Face Hub under the rjac publisher. It packages study identifiers, recruitment status, titles, brief summaries, and eligibility criteria into a single tidy table of roughly three thousand rows.
An English-language NLP corpus of cleaned earnings call transcripts gathered from public investor-relations pages, sized between 10,000 and 100,000 documents for financial analysis and LLM work.
A resource built for forecasting patent approval outcomes, introduced in research employing domain-specific fine-grained claim dependency graphs, with further documentation to be supplied later.
sec-filings is a Hugging Face dataset published by Sicheng-Chroma that distributes pre-chunked U.S. SEC filings paired with dense vector embeddings. It is distributed in optimized Parquet format and is intended for retrieval and similarity research over regulatory disclosures.
Vector embeddings of Apple patent text totalling about 4 million entries, stored as parquet shards with a 768-dimensional float64 feature per record.
The patent-phrase-similarity dataset is a text corpus that pairs patent terminology for similarity scoring. It is distributed by tasksource and contains roughly 48,548 rows across train, validation and test splits.
A collection of Parquet files providing one record per Chinese invention patent family granted between 2003 and 2019, capturing a cumulative quality assessment pipeline that includes metrics such as semantic knowledge recombination, strict historical semantic novelty, and IPC-based robustness measures across global patent publications.
View source information and access options.
Collection of anchor-positive pairs built from concatenated title, summary, and inclusion criteria fields for individual clinical trials, where each combined text is paired with four related questions for embedding model fine-tuning.
Iteration 2 of a clinical-trials corpus intended for embedding-model training and fine-tuning, enabling tasks such as ranked retrieval, document comparison across two or more anchors, and anchor-versus-chunk similarity scoring, distinguished from the prior release by finer-grained chunk segmentation.
A final anchor-positive pair dataset designed for fine-tuning embedding models on clinical trials, merging consolidated title-summary-inclusion chunk pairs with five-anchor three-positive chunk pairings.
The restaurant_review_sentiment dataset is a small text corpus pairing short restaurant reviews with binary sentiment labels and restaurant and user identifiers. It is hosted on the Hugging Face Hub under an unstated license and has been downloaded only a handful of times.
Dataset card entry for "filtered_yelp_restaurant_reviews" with additional details required.
A 10% sample of over 573,000 ClinicalTrials.gov studies augmented with AI-classified therapeutic areas, sponsor categorization, outcome groupings, and duration metrics, offered for portfolio and market landscape analysis.
A 10% sample of a sponsor profiling resource containing entries for over 10,200 clinical trial sponsors, covering trial counts, completion ratios, phase distribution, therapeutic areas, and pipeline activity.
10-K annual filings from EDGAR in Parquet, tens of thousands of documents. The standard annual-disclosure corpus for fundamentals, risk-factor and management-discussion analysis.
The restaurant_reviews dataset, published on the Hugging Face Hub by user yav1327, is a tabular and text collection of restaurant listings paired with user reviews. It is distributed in Parquet format with a single train split of 13,144 rows and an unstated license.
View source information and access options.
Over 635,000 registered clinical studies from ClinicalTrials.gov together with a supplementary pharma intelligence table of roughly 1.06 million rows that map those trials to FDA drug approvals, distributed as part of an open data initiative for academic and personal use.
The Tech Job Postings & Salary Dataset from Zalize Data contains over 449,000 monthly job postings scraped from public applicant tracking systems at US and EU technology companies. It is distributed as parquet files under a CC BY-NC 4.0 license and is part of the DataForge Open Data program.