Alternative data datasets
Browse by provider, category, marketplace, access, delivery and license
Parsed eligibility criteria from 13,229 ClinicalTrials.gov studies representing the candidate pool gathered by two first-stage retrievers during TREC Clinical Trials 2021–2023, formatted as typed entity-relation graphs for reranking models.
Datamule, Teraflop AI, and Eventual jointly contributed a 590 gigabyte SEC EDGAR collection encompassing about 8 million filings and 43 billion tokens, assembled using the datamule-python library and the official datamule API.
The TREC Clinical Trials collections released in 2021, 2022, and 2023 are available through the TREC homepage, with corresponding papers for each year. Artur Guimarães ([email protected]) serves as the dataset curator, while inquiries regarding the original authors can be directed separately. The dataset addresses the task of linking patient profiles with appropriate clinical trials, containing entries formatted as JSON objects with fields including query identifiers, disease labels, and accompanying clinical text.
An aggregation of 1.3 million U.S. patent records each accompanied by a human-written abstractive summary, sorted into nine Cooperative Patent Classification groupings spanning human necessities, chemistry, textiles, and related domains.
Released by Groundwork Data, this v1 release is the inaugural openly available structured compilation of Abuja residential property prices, capturing both rental and sale listings across 16 districts within Nigeria's Federal Capital Territory.
A large-scale release produced jointly by Datamule, Teraflop AI, and Eventual covering roughly 590 GB of SEC EDGAR content. It comprises about eight million filings totaling 43 billion tokens, gathered using the datamule-python package and the official datamule API.
S&P 500 earnings-call transcripts in Parquet under MIT, tens of thousands of calls. Independently compiled and comparable to other public transcript corpora for sentiment and tone work.
titer · EDGAR officer corpus 4,206,080 attested person–company–role–date tuples from SEC Forms 3/4/5, published as pointers rather than records, alongside the frozen pre-registrations that were hash-published before any measurement ran. edgar_officers.parquet: 4.2M rows, 230,405 distinct people, 20,266 issuers, 2006q1–2026q2. Column Meaning accession SEC accession number, the pointer that reconstr
Drawn from SEC EDGAR's full archive of millions of filings spanning all form types and US public companies, this sample isolates 1,000 recent 8-K material event submissions together with their filing metadata and document references.
A sample of US patent titles, abstracts and CPC classification labels built for multi-class patent classification, tens of thousands of text records in Parquet. Useful as training material for mapping innovation activity onto technology categories over time.
A collection of earnings call transcripts covering S&P 500 firms and other U.S. large-cap companies across the years 2005 through 2025, intended for financial analysis, NLP modeling, and sentiment research.
Public-safe distribution bundle from a local r19 counterexample resolution review of arXiv patents using a language model, containing review-related outputs, manifests, checksums, and provenance receipts, but excluding the original arXiv PDFs and any expert validation.
A small text-classification dataset by ClarusC64 that labels whether borrow-rate, short-interest and price signals cohere into a real short squeeze or diverge into a false signal. It is published on a model hub with an MIT license tag and a 10-row train split.
Records of initial public offerings launched on Indian markets between 2006 and 2025, with fields covering open and close dates, listing date, face value, issue price and size, lot size, first-day price, total shares offered and their allocation across anchor, NII, QIB and retail categories, minimum investment, and subscription figures for each investor class.
A tabular extract of clinical study descriptions drawn from ClinicalTrials.gov on 5 February 2025, emphasizing eligibility, design, and objective fields for healthcare and machine learning research.
Vector embeddings of clinical-trial records, hundreds of thousands of entries in Parquet (Apache 2.0). Enables similarity and retrieval analysis of trial design, indication crowding and competitive positioning.
DAPFAM patent documents, a curated sample of US patents in Parquet with text for retrieval and classification tasks (CC BY-NC-SA 4.0). Scopus of claim text for building patent-similarity and technology-mapping models.
An artificially generated collection of apparel product listings and accompanying advertisements produced by prompting GPT-4 to invent one hundred clothing items with descriptions and then write promotional copy for each, output in a structured product and description format without subsequent manual verification.
edgar_xbrl_companyfacts is a Hugging Face dataset published by DenyTranDFW that aggregates U.S. SEC EDGAR XBRL company-facts filings into a single Parquet file. The training split contains roughly 125 million rows of structured financial facts tagged by reporting period and unit.
DerivedFunction01's sec-filings-snippets is a public-records dataset of short text excerpts drawn from U.S. SEC filings, released on Hugging Face in Parquet format. It is sized for language-model fill-mask work and is not positioned as a commercial alternative-data product.
Restaurant Reviews Parsing NER Aspects This dataset is for the task of identifying the aspects of the restaurants mentioned in the reviews where aspect contains information about both the entities (FOOD, AMBIENCE, ...) and the attached sentiments. The input texts are from SemEval dataset. Labels for train and val datasets are generated by prompting Llama3 while the test dataset is curatedly manual
The cleaned public release of the EDGARCalcQA benchmark, containing example files in both JSON and JSONL formats, a human-readable summary, optional rejected candidate records, and a schema document describing field names, JSON types, and brief field-level descriptions.
Prepared by Electric Sheep Africa, this MDPA-sourced dataset supplies 10 entries of hotel room occupancy rates for Mauritius over 2019-2023, delivered as ML-ready Parquet files accompanied by uniform Hugging Face metadata and source provenance.
A small Parquet-format collection from Electric Sheep Africa reporting room occupancy rates for large hotels in Mauritius from 2019 to 2023, containing ten records drawn from MDPA and packaged with Hugging Face metadata for machine learning workflows.
Assembled by Electric Sheep Africa from MDPA, this dataset provides 60 rows of monthly hotel room occupancy rates for Mauritius spanning 2019-2023, formatted as ML-ready Parquet with consistent Hugging Face metadata and source attribution.
Compiled by Electric Sheep Africa, this MDPA-derived dataset offers 20 records of quarterly hotel room occupancy rates for Mauritius across the 2019-2023 period, distributed in ML-ready Parquet format with standardized Hugging Face metadata and source traceability.
Electric Sheep Africa releases 60 quarterly observations from MDPA on hotel room occupancy for large establishments in Mauritius from 2019 through 2023, formatted as Parquet with consistent metadata.
Structured numeric records parsed from corporate filings such as 10-K, 10-Q, and 8-K submissions to the U.S. SEC. The figures are packaged inside a compressed DuckDB instance made available through the open-source Datapond registry for fast querying.
A collection of 49,023 patent filings focused on distributed ledger technology is available, comprising roughly 1.296 billion tokens. The corpus is a component of the broader DLT-Corpus initiative and is intended to facilitate natural language processing research, innovation analysis, and patent examination within the distributed ledger field.
Indeed Job Postings (2026) is a 1,000-row sample of US Indeed listings joined to company firmographics, published by fact-den on Apify. License terms are not stated, and the publisher directs users to a commercial source actor for fresh or filtered data.
A sampled extraction of IPO-related tables from SEC filings between 1994 and 2026, delivering raw table HTML with provenance fields and targeting roughly 100 tables per year, with edge years possibly containing fewer valid extractions.
Text of IPO filings, including S-1 and related prospectuses, in JSON: hundreds of thousands of records in English (CC BY 4.0). Pre-IPO disclosure text is one of the few information sources available ahead of a listing.
Earnings call transcripts for S&P 500 and other large U.S. companies covering 2005 through 2025, useful for financial research, natural language processing, and sentiment analysis.
A consolidated, ready-to-use repository of 13F institutional holdings drawn from SEC EDGAR filings, capturing the most recent quarterly disclosures of more than 13,000 investment managers including hedge funds, mutual fund complexes, pension funds, banks, and family offices, with each record summarizing total assets under management and individual positions.
A joint release by Datamule, Teraflop AI, and Eventual provides the SEC-EDGAR dataset, comprising 590 GB of content drawn from all major filings in the SEC EDGAR database and totaling roughly 8 million samples and 43 billion tokens. The data was assembled using the datamule-python library and the official datamule API developed by John Friedman, with datamule serving as a Python package for collecting and manipulating such records.
A collection of 2,000 Malaysian property transaction records spanning every state offers a detailed look at the national housing market in 2025. Compiled from Brickz, a recognized source for real estate transaction data, it covers location, tenure, property type, median pricing, and transaction volumes.
Earnings-call transcripts from an IBM research project, tens of thousands of records with finance tags, in CC0. A filings-adjacent corpus for call-tone and question-and-answer analysis.
Datamule partnered with Teraflop AI and Eventual to publish the SEC-EDGAR dataset, which covers all major filings from the SEC EDGAR database and amounts to 590 GB across about 8 million samples and 43 billion tokens. Collection relied on the datamule-python library and John Friedman's official datamule API, tools designed in Python for gathering and handling SEC filing data.
Metadata for every EDGAR filing submitted to the U.S. Securities and Exchange Commission from 1994 through December 14, 2024, derived from the quarterly master index files and including CIKs, issuer names, and form types.
Earnings-call transcripts for S&P 500 companies, tens of thousands of calls in English with timestamps, in Parquet (MIT). Direct input for call-tone and question-and-answer sentiment analysis across the large-cap earnings cycle.
Dataset Card for 中華民國專利技術名詞中英對照詞庫 中華民國專利技術名詞中英對照詞庫(Chinese-English Technical Patent Glossary)收錄逾 324 萬筆台灣專利技術名詞之中英對照資料,涵蓋國際專利分類(IPC)A 至 H 全部八大類,時間跨度自 2011 年至 2023 年。本資料集適用於專利翻譯、技術術語標準化、以及繁體中文語言模型在專業領域之詞彙增強。 Dataset Details Dataset Description 本資料集整理自中華民國經濟部智慧財產局(TIPO)公開之專利技術名詞中英對照詞庫。每筆資料包含一組繁體中文與英文之技術術語對照,並標註其對應的國際專利分類(IPC)代碼與資料來源編號。 資料涵蓋 IPC 八大類別: A — 人類生活需要(Human Necessities) B — 作業、運輸(Performin
Metadata for hundreds of thousands of clinical trials in English, French and Spanish, structured in Parquet with study-level fields such as phase, condition, sponsor and status (Apache 2.0). Trial starts and phase transitions lead clinical development, so this works as an early indicator of biotech pipeline activity.
A corpus of around 1.3 million U.S. patent filings paired with human-authored abstractive summaries, organized into nine Cooperative Patent Classification categories ranging from human necessities to textiles and paper.
A collection featuring the primary independent claim of United States utility patents issued between 2005 and 2025 is provided. Because the opening claim usually establishes the widest legal boundaries of an invention, it carries particular significance, and each entry includes the patent identifier alongside its claim text.
A benchmark collection designed to test vision-language models on the visual conventions found in U.S. design patent illustrations. It pulls 3.6 million figures from 2007 to 2022 in the IMPACT archive and pairs them with text drawn from PatentsView, stored in yearly Parquet files along with an 800-sample evaluation set.
Mindweave's Job Postings & Applications is a synthetic applicant-tracking and job-board dataset covering a simulated multi-industry hiring market. It is published under a CC BY-NC 4.0 license and distributed as a CSV.
A subset of the BIGPATENT corpus adapted for the MTEB clustering benchmark, tens of thousands of English patents and titles under CC BY 4.0. Primarily a research and evaluation set for similarity and clustering workloads.
FinSight gathers plain-text 10-K and 10-Q SEC filings for twenty major US public firms spanning six sectors, yielding 97 records intended for BERT fine-tuning, retrieval-augmented generation, and multi-agent financial research.
Global Job Postings – Multi-ATS is a single dated snapshot of 112,816 job listings aggregated from 23 applicant tracking systems and parsed into a 47-column schema by NextGig-Rocks. It is distributed as Parquet via the Hugging Face datasets library and targets labor-market research, not production job boards.
Two Parquet files hold global patent publication records paired with quality indicators for Chinese invention patent families from 2003 to 2019, with each row representing a focal family and carrying cumulative measures such as semantic knowledge recombination, strict historical semantic novelty, and IPC-based robustness.
Over 1.3 million granted US patents from the BIGPATENT benchmark, each pairing its claims with an abstractive summary and full description text in English (CC BY 4.0, Parquet). Use it to track the technology areas where applicants are filing and to study how claim language has evolved over time.
Stocks Weekly Earnings Surprise Weekly earnings surprise probabilities and outcomes for publicly traded companies. 2,204,032 rows over 6,067 symbols, 8 columns, covering 2016-12-30 to 2026-07-03. Refreshed monthly. Why It Matters This dataset supplies high-frequency earnings-surprise context for equity strategies by: Pre-event positioning: Surprise probabilities guide sizing and hedging ahead of e
Penumbra AI SEC 10-K Risk Factor Disclosures (S&P 500, FY2025-2026) Dataset Summary This dataset contains a sample of 2,683 individual risk factor disclosures from the most recent Form 10-K filing of 50 S&P 500 companies (fiscal year ends ranging from May 2025 to February 2026). Each record is a single risk driver — one discrete point a company disclosed under "Risk Factors" — together with the ve
The patent_classification dataset is a small text corpus hosted on the Hugging Face Hub by publisher pgurazada1, comprising paired patent text snippets and their cooperative classification labels. It is distributed as a CSV file under an unstated license.
ClinicalTrials.gov scraping tool that extracts study records based on criteria like condition, sponsor, phase, status, or geography, yielding one row per NCT identifier containing fields such as sponsors, phases, enrollment, outcomes, and locations, with no API key needed and approximately 11,191 rows and 63 fields collected over 65 runs.
Search filings on SEC EDGAR using ticker symbols, CIK identifiers, SIC industry codes, or free-text queries, returning a single structured row per filing that includes company classification, period dates, direct links to source documents, and optional XBRL financial data. The dataset contains 2,691 rows across 27 fields, supported by 64 collector runs, with the latest observation dated 2026-08-04, and covers 771 entities, browsable at https://reapx.dev/data/sec-edgar-scraper/.
A structured compilation of 29,633 completed Phase 3 clinical trials sourced from ClinicalTrials.gov, including status and design details for each study record.
Original TIFF drawings and grant full-text XML for 165,917 U.S. design patents omitted from the AI4Patents/IMPACT collection, filling gaps including 161,093 grants from 2023 to 2026 and 4,824 earlier patents missing from IMPACT.
A community-compiled collection of stock-market-related tweets in English, stored as CSV, hundreds of thousands of posts (CC BY 4.0). Retail chatter feeds sentiment-signal research and market-tone studies.
Text of SEC filings pulled from EDGAR spanning millions of records in English (Apache 2.0): 10-K, 10-Q, 8-K, annual reports and exhibits, tagged for text-generation and classification. A practical base corpus for earnings- and filings-driven signals, from risk factors and MD&A to management commentary.