Alternative data datasets
Browse by provider, category, marketplace, access, delivery and license
Datamule, Teraflop AI, and Eventual jointly contributed a 590 gigabyte SEC EDGAR collection encompassing about 8 million filings and 43 billion tokens, assembled using the datamule-python library and the official datamule API.
A static tabular extension of the Bose345/sp500_earnings_transcripts collection covering the same 2005 to 2025 calendar of S&P 500 earnings events, where each row represents a single company-quarter call identified by a stable episode_id and bundles the full transcript along with related SEC press materials and pre-event inputs suitable for supervised learning or reinforcement-style experimentation.
The TREC Clinical Trials collections released in 2021, 2022, and 2023 are available through the TREC homepage, with corresponding papers for each year. Artur Guimarães ([email protected]) serves as the dataset curator, while inquiries regarding the original authors can be directed separately. The dataset addresses the task of linking patient profiles with appropriate clinical trials, containing entries formatted as JSON objects with fields including query identifiers, disease labels, and accompanying clinical text.
Collection of historical Swedish patent texts from 1885 to 1972 assigned multi-label Cooperative Patent Classification (CPC) codes, intended for retrieval, prior art searching, and multi-label classification tasks.
A large-scale release produced jointly by Datamule, Teraflop AI, and Eventual covering roughly 590 GB of SEC EDGAR content. It comprises about eight million filings totaling 43 billion tokens, gathered using the datamule-python package and the official datamule API.
S&P 500 earnings-call transcripts in Parquet under MIT, tens of thousands of calls. Independently compiled and comparable to other public transcript corpora for sentiment and tone work.
Drawn from SEC EDGAR's full archive of millions of filings spanning all form types and US public companies, this sample isolates 1,000 recent 8-K material event submissions together with their filing metadata and document references.
A sample of US patent titles, abstracts and CPC classification labels built for multi-class patent classification, tens of thousands of text records in Parquet. Useful as training material for mapping innovation activity onto technology categories over time.
A collection of earnings call transcripts covering S&P 500 firms and other U.S. large-cap companies across the years 2005 through 2025, intended for financial analysis, NLP modeling, and sentiment research.
A small text-classification dataset by ClarusC64 that labels whether borrow-rate, short-interest and price signals cohere into a real short squeeze or diverge into a false signal. It is published on a model hub with an MIT license tag and a 10-row train split.
Vector embeddings of clinical-trial records, hundreds of thousands of entries in Parquet (Apache 2.0). Enables similarity and retrieval analysis of trial design, indication crowding and competitive positioning.
Indeed Job Postings (2026) is a 1,000-row sample of US Indeed listings joined to company firmographics, published by fact-den on Apify. License terms are not stated, and the publisher directs users to a commercial source actor for fresh or filtered data.
S&P 500 earnings-call transcripts preprocessed into optimized Parquet, tens of thousands of calls (MIT). Ready for sentiment, QA and summarization research on the large-cap earnings season.
A sampled extraction of IPO-related tables from SEC filings between 1994 and 2026, delivering raw table HTML with provenance fields and targeting roughly 100 tables per year, with edge years possibly containing fewer valid extractions.
Text of IPO filings, including S-1 and related prospectuses, in JSON: hundreds of thousands of records in English (CC BY 4.0). Pre-IPO disclosure text is one of the few information sources available ahead of a listing.
Earnings call transcripts for S&P 500 and other large U.S. companies covering 2005 through 2025, useful for financial research, natural language processing, and sentiment analysis.
A joint release by Datamule, Teraflop AI, and Eventual provides the SEC-EDGAR dataset, comprising 590 GB of content drawn from all major filings in the SEC EDGAR database and totaling roughly 8 million samples and 43 billion tokens. The data was assembled using the datamule-python library and the official datamule API developed by John Friedman, with datamule serving as a Python package for collecting and manipulating such records.
Earnings-call transcripts from an IBM research project, tens of thousands of records with finance tags, in CC0. A filings-adjacent corpus for call-tone and question-and-answer analysis.
Datamule partnered with Teraflop AI and Eventual to publish the SEC-EDGAR dataset, which covers all major filings from the SEC EDGAR database and amounts to 590 GB across about 8 million samples and 43 billion tokens. Collection relied on the datamule-python library and John Friedman's official datamule API, tools designed in Python for gathering and handling SEC filing data.
Earnings-call transcripts for S&P 500 companies, tens of thousands of calls in English with timestamps, in Parquet (MIT). Direct input for call-tone and question-and-answer sentiment analysis across the large-cap earnings cycle.
Metadata for hundreds of thousands of clinical trials in English, French and Spanish, structured in Parquet with study-level fields such as phase, condition, sponsor and status (Apache 2.0). Trial starts and phase transitions lead clinical development, so this works as an early indicator of biotech pipeline activity.
german-job-postings is a publicly hosted corpus of roughly 70,584 normalized German-language vacancies drawn from the Bundesagentur für Arbeit Jobbörse API, each row tagged with the KldB-2010 occupation code and machine-assigned ESCO occupation and skill labels. The dataset is distributed under CC-BY-4.0 and is intended as an open alternative to the predominantly English/US job-posting resources available to machine-learning practitioners.
A subset of the BIGPATENT corpus adapted for the MTEB clustering benchmark, tens of thousands of English patents and titles under CC BY 4.0. Primarily a research and evaluation set for similarity and clustering workloads.
FinSight gathers plain-text 10-K and 10-Q SEC filings for twenty major US public firms spanning six sectors, yielding 97 records intended for BERT fine-tuning, retrieval-augmented generation, and multi-agent financial research.
Penumbra AI SEC 10-K Risk Factor Disclosures (S&P 500, FY2025-2026) Dataset Summary This dataset contains a sample of 2,683 individual risk factor disclosures from the most recent Form 10-K filing of 50 S&P 500 companies (fiscal year ends ranging from May 2025 to February 2026). Each record is a single risk driver — one discrete point a company disclosed under "Risk Factors" — together with the ve
A daily record of trading activity across multiple major global stock market indices, structured for trend analysis and forecasting work in international equity markets.
S&P 500 earnings call episodes from 2005 to 2025 are released as static tabular rows, each identified by a company-quarter episode identifier and containing full transcripts along with SEC press materials, suited to supervised or reinforcement-style experiments.
stock_market_dataset is a Hugging Face community dataset published by SkyWalkertT1 that aggregates short Turkish-language market commentary paired with sentiment labels. It is distributed under an MIT-licensed CSV format and is intended for text-classification and token-classification tasks rather than live trading signals.
A community-compiled collection of stock-market-related tweets in English, stored as CSV, hundreds of thousands of posts (CC BY 4.0). Retail chatter feeds sentiment-signal research and market-tone studies.
Text of SEC filings pulled from EDGAR spanning millions of records in English (Apache 2.0): 10-K, 10-Q, 8-K, annual reports and exhibits, tagged for text-generation and classification. A practical base corpus for earnings- and filings-driven signals, from risk factors and MD&A to management commentary.
A collection of 943,672 tweets gathered between April 9 and July 16, 2020, harvested via the #SPX500 tag, references to the top 25 S&P 500 firms, and the #stocks tag, originally published by Bruno Taborda on IEEE.
Over 635,000 registered clinical studies from ClinicalTrials.gov together with a supplementary pharma intelligence table of roughly 1.06 million rows that map those trials to FDA drug approvals, distributed as part of an open data initiative for academic and personal use.
The tech-job-postings-salaries-sample dataset is a 500-row public sample of a larger scraping-based feed of technology job postings drawn from public applicant tracking system endpoints such as Greenhouse, Lever, Ashby, Workable and SmartRecruiter. It is published by DataForge (Zalize) for evaluation and preview use.
The Tech Job Postings & Salary Dataset from Zalize Data contains over 449,000 monthly job postings scraped from public applicant tracking systems at US and EU technology companies. It is distributed as parquet files under a CC BY-NC 4.0 license and is part of the DataForge Open Data program.
Derived from USPTO and PatentsView bulk releases, this dataset catalogs 9.1 million U.S. patents from 1976 onward along with complete citation links, assignees, inventors, and CPC codes, with the full-graph edition exceeding 255 million rows.