Alternative data datasets
Browse by provider, category, marketplace, access, delivery and license
An aggregation of 1.3 million U.S. patent records each accompanied by a human-written abstractive summary, sorted into nine Cooperative Patent Classification groupings spanning human necessities, chemistry, textiles, and related domains.
The cleaned public release of the EDGARCalcQA benchmark, containing example files in both JSON and JSONL formats, a human-readable summary, optional rejected candidate records, and a schema document describing field names, JSON types, and brief field-level descriptions.
Prepared by Electric Sheep Africa, this MDPA-sourced dataset supplies 10 entries of hotel room occupancy rates for Mauritius over 2019-2023, delivered as ML-ready Parquet files accompanied by uniform Hugging Face metadata and source provenance.
A small Parquet-format collection from Electric Sheep Africa reporting room occupancy rates for large hotels in Mauritius from 2019 to 2023, containing ten records drawn from MDPA and packaged with Hugging Face metadata for machine learning workflows.
Assembled by Electric Sheep Africa from MDPA, this dataset provides 60 rows of monthly hotel room occupancy rates for Mauritius spanning 2019-2023, formatted as ML-ready Parquet with consistent Hugging Face metadata and source attribution.
Compiled by Electric Sheep Africa, this MDPA-derived dataset offers 20 records of quarterly hotel room occupancy rates for Mauritius across the 2019-2023 period, distributed in ML-ready Parquet format with standardized Hugging Face metadata and source traceability.
Electric Sheep Africa releases 60 quarterly observations from MDPA on hotel room occupancy for large establishments in Mauritius from 2019 through 2023, formatted as Parquet with consistent metadata.
Official Consumer Price Index and inflation series for Iran drawn from two providers: the Statistical Centre of Iran (monthly, from 2011/1390 onward) and the Central Bank of Iran (annual, long historical run from 1936/1315 onward), each with its own series and a combined view covering national-level index and percentage data.
A legacy archive of Iran's consumer price index and inflation series from 1936 to 2022 remains available, though an updated multisource version using current Farmaanaa methodology supersedes it.
Text of IPO filings, including S-1 and related prospectuses, in JSON: hundreds of thousands of records in English (CC BY 4.0). Pre-IPO disclosure text is one of the few information sources available ahead of a listing.
An API-based ClinicalTrials.gov scraper that retrieves over 585,000 study records including phase, sponsor, conditions, interventions, enrollment, eligibility, and sites, filterable by condition, sponsor, and status.
A corpus of around 1.3 million U.S. patent filings paired with human-authored abstractive summaries, organized into nine Cooperative Patent Classification categories ranging from human necessities to textiles and paper.
A benchmark collection designed to test vision-language models on the visual conventions found in U.S. design patent illustrations. It pulls 3.6 million figures from 2007 to 2022 in the IMPACT archive and pairs them with text drawn from PatentsView, stored in yearly Parquet files along with an 800-sample evaluation set.
AlphaAI's near-real-time parse of SEC EDGAR Form 4 filings, providing 20,832 individual insider transaction tranches and 8,538 grouped economic events for US public-company officers, directors, and 10% owners.
german-job-postings is a publicly hosted corpus of roughly 70,584 normalized German-language vacancies drawn from the Bundesagentur für Arbeit Jobbörse API, each row tagged with the KldB-2010 occupation code and machine-assigned ESCO occupation and skill labels. The dataset is distributed under CC-BY-4.0 and is intended as an open alternative to the predominantly English/US job-posting resources available to machine-learning practitioners.
A repository of roughly 2.7 million publicly available U.S. patent application records split into yearly source files spanning 2021 to 2026, covering application metadata, applicants and inventors, classifications, prosecution history, continuity and priority data, assignments, publications, and grant information where present, sourced from the USPTO.
A subset of the BIGPATENT corpus adapted for the MTEB clustering benchmark, tens of thousands of English patents and titles under CC BY 4.0. Primarily a research and evaluation set for similarity and clustering workloads.
Global Job Postings – Multi-ATS is a single dated snapshot of 112,816 job listings aggregated from 23 applicant tracking systems and parsed into a 47-column schema by NextGig-Rocks. It is distributed as Parquet via the Hugging Face datasets library and targets labor-market research, not production job boards.
Over 1.3 million granted US patents from the BIGPATENT benchmark, each pairing its claims with an abstractive summary and full description text in English (CC BY 4.0, Parquet). Use it to track the technology areas where applicants are filing and to study how claim language has evolved over time.
Penumbra AI SEC 10-K Risk Factor Disclosures (S&P 500, FY2025-2026) Dataset Summary This dataset contains a sample of 2,683 individual risk factor disclosures from the most recent Form 10-K filing of 50 S&P 500 companies (fiscal year ends ranging from May 2025 to February 2026). Each record is a single risk driver — one discrete point a company disclosed under "Risk Factors" — together with the ve
The ja_patent mirror of llm-jp-corpus-v4, assembled by the LLM-jp Corpus Building Working Group at NII, reproduces the Japanese patent sub-corpus from the larger LLM-jp Corpus v4 release. It is delivered as 621 jsonl.gz files totaling 58.2 GB in compressed form, with each line holding a JSON object that includes a text field and a meta field containing the document identifier, URL, and other provenance information.
A structured compilation of 29,633 completed Phase 3 clinical trials sourced from ClinicalTrials.gov, including status and design details for each study record.
The patent-phrase-similarity dataset is a text corpus that pairs patent terminology for similarity scoring. It is distributed by tasksource and contains roughly 48,548 rows across train, validation and test splits.
Reduced sample of the big_patent collection, balanced across text lengths to support shorter fine-tuning runs, with a roughly even spread up to one million characters suitable for training on sequences of up to 250,000 tokens.
A 10% sample of over 573,000 ClinicalTrials.gov studies augmented with AI-classified therapeutic areas, sponsor categorization, outcome groupings, and duration metrics, offered for portfolio and market landscape analysis.
A 10% sample of a sponsor profiling resource containing entries for over 10,200 clinical trial sponsors, covering trial counts, completion ratios, phase distribution, therapeutic areas, and pipeline activity.