Alternative data datasets
Browse by provider, category, marketplace, access, delivery and license
PatentMatch pairs US patents with matching reference patents for retrieval evaluation, a few thousand records in JSON (Apache 2.0). A niche evaluation set for patent-similarity and citation-link models.
Chinese patents paired with their abstractive summaries in Mandarin, a few thousand records in JSON (Apache 2.0). A window into the pace and direction of patenting inside the Chinese technology base.
An artificially generated collection of apparel product listings and accompanying advertisements produced by prompting GPT-4 to invent one hundred clothing items with descriptions and then write promotional copy for each, output in a structured product and description format without subsequent manual verification.
高质量中文专利摘要数据集。
A legacy archive of Iran's consumer price index and inflation series from 1936 to 2022 remains available, though an updated multisource version using current Farmaanaa methodology supersedes it.
Text of IPO filings, including S-1 and related prospectuses, in JSON: hundreds of thousands of records in English (CC BY 4.0). Pre-IPO disclosure text is one of the few information sources available ahead of a listing.
Quarterly 10-Q filings parsed from EDGAR into JSON, millions of records in English (MIT). The quarterly companion to the 10-K corpus for tracking sequential changes in operating performance and risk factors.
Penumbra AI SEC 10-K Risk Factor Disclosures (S&P 500, FY2025-2026) Dataset Summary This dataset contains a sample of 2,683 individual risk factor disclosures from the most recent Form 10-K filing of 50 S&P 500 companies (fiscal year ends ranging from May 2025 to February 2026). Each record is a single risk driver — one discrete point a company disclosed under "Risk Factors" — together with the ve
The ja_patent mirror of llm-jp-corpus-v4, assembled by the LLM-jp Corpus Building Working Group at NII, reproduces the Japanese patent sub-corpus from the larger LLM-jp Corpus v4 release. It is delivered as 621 jsonl.gz files totaling 58.2 GB in compressed form, with each line holding a JSON object that includes a text field and a meta field containing the document identifier, URL, and other provenance information.
A structured compilation of 29,633 completed Phase 3 clinical trials sourced from ClinicalTrials.gov, including status and design details for each study record.
tech-job-postings-labeled is a small public dataset of labeled technology job postings published on Hugging Face by SuhaibAtef. It comprises roughly 8,000 records tagged with source labels and is distributed in JSON.
View source information and access options.
The 2024 Earnings Call Transcript dataset, published by yeong-hwan on Hugging Face, aggregates quarterly earnings call dialogue for individual stock tickers. It is distributed as a JSON table with 1,904 rows in a single train split.