Alternative data datasets

Browse by provider, category, marketplace, access, delivery and license

Showing 1–35 of35 results for “task_categories:text-classification”
adityaag2k
SEC-EDGAR

Datamule, Teraflop AI, and Eventual jointly contributed a 590 gigabyte SEC EDGAR collection encompassing about 8 million filings and 43 billion tokens, assembled using the datamule-python library and the official datamule API.

Public Records & Filings Free View
Amaanaush
earnings-call-data

A static tabular extension of the Bose345/sp500_earnings_transcripts collection covering the same 2005 to 2025 calendar of S&P 500 earnings events, where each row represents a single company-quarter call identified by a stable episode_id and bundles the full transcript along with related SEC press materials and pre-event inputs suitable for supervised learning or reinforcement-style experimentation.

Estimates, Events & Fund Flows Free View
araag2
TREC_Clinical-Trials

The TREC Clinical Trials collections released in 2021, 2022, and 2023 are available through the TREC homepage, with corresponding papers for each year. Artur Guimarães ([email protected]) serves as the dataset curator, while inquiries regarding the original authors can be directed separately. The dataset addresses the task of linking patient profiles with appropriate clinical trials, containing entries formatted as JSON objects with fields including query identifiers, disease labels, and accompanying clinical text.

Public Records & Filings Free View
atheer2104
swedish-patent-cpc-subclass-new

Collection of historical Swedish patent texts from 1885 to 1972 assigned multi-label Cooperative Patent Classification (CPC) codes, intended for retrieval, prior art searching, and multi-label classification tasks.

Public Records & Filings Free View
baridhi
SEC-EDGAR

A large-scale release produced jointly by Datamule, Teraflop AI, and Eventual covering roughly 590 GB of SEC EDGAR content. It comprises about eight million filings totaling 43 billion tokens, gathered using the datamule-python package and the official datamule API.

Public Records & Filings Free View
Bose345
sp500_earnings_transcripts

S&P 500 earnings-call transcripts in Parquet under MIT, tens of thousands of calls. Independently compiled and comparable to other public transcript corpora for sentiment and tone work.

Public Records & Filings Free View
carrierone
sec-filings-sample

Drawn from SEC EDGAR's full archive of millions of filings spanning all form types and US public companies, this sample isolates 1,000 recent 8-K material event submissions together with their filing metadata and document references.

Public Records & Filings Free View
ccdv
patent-classification

A sample of US patent titles, abstracts and CPC classification labels built for multi-class patent classification, tens of thousands of text records in Parquet. Useful as training material for mapping innovation activity onto technology categories over time.

Public Records & Filings Free View
churchill1254
sp500_earnings_transcripts

A collection of earnings call transcripts covering S&P 500 firms and other U.S. large-cap companies across the years 2005 through 2025, intended for financial analysis, NLP modeling, and sentiment research.

News & Sentiment Free View
ClarusC64
market-borrow-rate-short-interest-coherence-squeeze-risk-v0.1

A small text-classification dataset by ClarusC64 that labels whether borrow-rate, short-interest and price signals cohere into a real short squeeze or diverge into a false signal. It is published on a model hub with an MIT license tag and a 10-row train split.

Estimates, Events & Fund Flows Free View
cyrilzakka
clinical-trials-embeddings

Vector embeddings of clinical-trial records, hundreds of thousands of entries in Parquet (Apache 2.0). Enables similarity and retrieval analysis of trial design, indication crowding and competitive positioning.

Public Records & Filings Free View
fact-den
indeed-job-postings-2026

Indeed Job Postings (2026) is a 1,000-row sample of US Indeed listings joined to company firmographics, published by fact-den on Apify. License terms are not stated, and the publisher directs users to a commercial source actor for fresh or filtered data.

Jobs & Workforce Free View
glopardo
sp500-earnings-transcripts

S&P 500 earnings-call transcripts preprocessed into optimized Parquet, tens of thousands of calls (MIT). Ready for sentiment, QA and summarization research on the large-cap earnings season.

News & Sentiment Free View
gtfintechlab
ipo-tables

A sampled extraction of IPO-related tables from SEC filings between 1994 and 2026, delivering raw table HTML with provenance fields and targeting roughly 100 tables per year, with edge years possibly containing fewer valid extractions.

Public Records & Filings Free View
gtfintechlab
ipo-text

Text of IPO filings, including S-1 and related prospectuses, in JSON: hundreds of thousands of records in English (CC BY 4.0). Pre-IPO disclosure text is one of the few information sources available ahead of a listing.

Estimates, Events & Fund Flows Free View
idleengine
sp500_earnings_transcripts

Earnings call transcripts for S&P 500 and other large U.S. companies covering 2005 through 2025, useful for financial research, natural language processing, and sentiment analysis.

News & Sentiment Free View
Jeremydh911
SEC-EDGAR

A joint release by Datamule, Teraflop AI, and Eventual provides the SEC-EDGAR dataset, comprising 590 GB of content drawn from all major filings in the SEC EDGAR database and totaling roughly 8 million samples and 43 billion tokens. The data was assembled using the datamule-python library and the official datamule API developed by John Friedman, with datamule serving as a Python package for collecting and manipulating such records.

Public Records & Filings Free View
jlh-ibm
earnings_call

Earnings-call transcripts from an IBM research project, tens of thousands of records with finance tags, in CC0. A filings-adjacent corpus for call-tone and question-and-answer analysis.

Public Records & Filings Free View
kapilrao
SEC-EDGAR

Datamule partnered with Teraflop AI and Eventual to publish the SEC-EDGAR dataset, which covers all major filings from the SEC EDGAR database and amounts to 590 GB across about 8 million samples and 43 billion tokens. Collection relied on the datamule-python library and John Friedman's official datamule API, tools designed in Python for gathering and handling SEC filing data.

Public Records & Filings Free View
kurry
sp500_earnings_transcripts

Earnings-call transcripts for S&P 500 companies, tens of thousands of calls in English with timestamps, in Parquet (MIT). Direct input for call-tone and question-and-answer sentiment analysis across the large-cap earnings cycle.

Public Records & Filings Free View
louisbrulenaudet
clinical-trials

Metadata for hundreds of thousands of clinical trials in English, French and Spanish, structured in Parquet with study-level fields such as phase, condition, sponsor and status (Apache 2.0). Trial starts and phase transitions lead clinical development, so this works as an early indicator of biotech pipeline activity.

Public Records & Filings Free View
mischeiwiller
german-job-postings

german-job-postings is a publicly hosted corpus of roughly 70,584 normalized German-language vacancies drawn from the Bundesagentur für Arbeit Jobbörse API, each row tagged with the KldB-2010 occupation code and machine-assigned ESCO occupation and skill labels. The dataset is distributed under CC-BY-4.0 and is intended as an open alternative to the predominantly English/US job-posting resources available to machine-learning practitioners.

Jobs & Workforce Free View
mteb
big-patent

A subset of the BIGPATENT corpus adapted for the MTEB clustering benchmark, tens of thousands of English patents and titles under CC BY 4.0. Primarily a research and evaluation set for similarity and clustering workloads.

Public Records & Filings Free View
musk1209
finsight-sec-filings

FinSight gathers plain-text 10-K and 10-Q SEC filings for twenty major US public firms spanning six sectors, yielding 97 records intended for BERT fine-tuning, retrieval-augmented generation, and multi-agent financial research.

Public Records & Filings Free View
PenumbraAI
sec-edgar-filing-risks

Penumbra AI SEC 10-K Risk Factor Disclosures (S&P 500, FY2025-2026) Dataset Summary This dataset contains a sample of 2,683 individual risk factor disclosures from the most recent Form 10-K filing of 50 S&P 500 companies (fiscal year ends ranging from May 2025 to February 2026). Each record is a single risk driver — one discrete point a company disclosed under "Risk Factors" — together with the ve

Public Records & Filings Free View
pettah
global-top-Index-exploring-trends-in-stock-Market

A daily record of trading activity across multiple major global stock market indices, structured for trend analysis and forecasting work in international equity markets.

Public Records & Filings Free View
RudrakshNanavaty
earnings-call-data

S&P 500 earnings call episodes from 2005 to 2025 are released as static tabular rows, each identified by a company-quarter episode identifier and containing full transcripts along with SEC press materials, suited to supervised or reinforcement-style experiments.

Estimates, Events & Fund Flows Free View
SkyWalkertT1
stock_market_dataset

stock_market_dataset is a Hugging Face community dataset published by SkyWalkertT1 that aggregates short Turkish-language market commentary paired with sentiment labels. It is distributed under an MIT-licensed CSV format and is intended for text-classification and token-classification tasks rather than live trading signals.

News & Sentiment Free View
StephanAkkerman
stock-market-tweets-data

A community-compiled collection of stock-market-related tweets in English, stored as CSV, hundreds of thousands of posts (CC BY 4.0). Retail chatter feeds sentiment-signal research and market-tone studies.

News & Sentiment Free View
TeraflopAI
SEC-EDGAR

Text of SEC filings pulled from EDGAR spanning millions of records in English (Apache 2.0): 10-K, 10-Q, 8-K, annual reports and exhibits, tagged for text-generation and classification. A practical base corpus for earnings- and filings-driven signals, from risk factors and MD&A to management commentary.

Public Records & Filings Free View
VeeraThakshith
stock-market-tweets-data

A collection of 943,672 tweets gathered between April 9 and July 16, 2020, harvested via the #SPX500 tag, references to the top 25 S&P 500 firms, and the #stocks tag, originally published by Bruno Taborda on IEEE.

News & Sentiment Free View
zalizedata
clinical-trials-drug-approvals-dataset

Over 635,000 registered clinical studies from ClinicalTrials.gov together with a supplementary pharma intelligence table of roughly 1.06 million rows that map those trials to FDA drug approvals, distributed as part of an open data initiative for academic and personal use.

Public Records & Filings Free View
zalizedata
tech-job-postings-salaries-sample

The tech-job-postings-salaries-sample dataset is a 500-row public sample of a larger scraping-based feed of technology job postings drawn from public applicant tracking system endpoints such as Greenhouse, Lever, Ashby, Workable and SmartRecruiter. It is published by DataForge (Zalize) for evaluation and preview use.

Jobs & Workforce Free View
zalizedata
tech-job-postings-salary-dataset

The Tech Job Postings & Salary Dataset from Zalize Data contains over 449,000 monthly job postings scraped from public applicant tracking systems at US and EU technology companies. It is distributed as parquet files under a CC BY-NC 4.0 license and is part of the DataForge Open Data program.

Jobs & Workforce Free View
zalizedata
us-patents-citation-company-dataset

Derived from USPTO and PatentsView bulk releases, this dataset catalogs 9.1 million U.S. patents from 1976 onward along with complete citation links, assignees, inventors, and CPC codes, with the full-graph edition exceeding 255 million rows.

Public Records & Filings Free View