Alternative data datasets
Browse by provider, category, marketplace, access, delivery and license
Reduced sample of the big_patent collection, balanced across text lengths to support shorter fine-tuning runs, with a roughly even spread up to one million characters suitable for training on sequences of up to 250,000 tokens.
A knowledge graph assembled from 575,778 ClinicalTrials.gov registrations and their associated arms, outcomes, sites, sponsors, conditions, interventions, MeSH codes, and PubMed citations, containing 7,628,735 nodes and 15,531,427 edges, delivered as 623 MB of Parquet files (compared to about 7 GB of raw JSON) and built using Samyama Graph.
A collection of 943,672 tweets gathered between April 9 and July 16, 2020, harvested via the #SPX500 tag, references to the top 25 S&P 500 firms, and the #stocks tag, originally published by Bruno Taborda on IEEE.
The 2024 Earnings Call Transcript dataset, published by yeong-hwan on Hugging Face, aggregates quarterly earnings call dialogue for individual stock tickers. It is distributed as a JSON table with 1,904 rows in a single train split.
Over 635,000 registered clinical studies from ClinicalTrials.gov together with a supplementary pharma intelligence table of roughly 1.06 million rows that map those trials to FDA drug approvals, distributed as part of an open data initiative for academic and personal use.
The tech-job-postings-salaries-sample dataset is a 500-row public sample of a larger scraping-based feed of technology job postings drawn from public applicant tracking system endpoints such as Greenhouse, Lever, Ashby, Workable and SmartRecruiter. It is published by DataForge (Zalize) for evaluation and preview use.
The Tech Job Postings & Salary Dataset from Zalize Data contains over 449,000 monthly job postings scraped from public applicant tracking systems at US and EU technology companies. It is distributed as parquet files under a CC BY-NC 4.0 license and is part of the DataForge Open Data program.
Derived from USPTO and PatentsView bulk releases, this dataset catalogs 9.1 million U.S. patents from 1976 onward along with complete citation links, assignees, inventors, and CPC codes, with the full-graph edition exceeding 255 million rows.
Normalized stock-trade disclosures filed by U.S. House and Senate members under the STOCK Act, regenerated on a schedule by the repository's own pipeline rather than curated manually.
ZipLime US Insider Trading Disclosures (PIT) [!CAUTION] Use knowledge_date, not transaction_date, when backtesting. transaction_date says when a trade occurred; knowledge_date says when the filing became observable through EDGAR. Using the former as the signal date introduces look-ahead bias. visible = trades.filter(pl.col("knowledge_date") <= simulation_time) This dataset normalizes corporate-ins
A curated collection of SEC EDGAR filings from recent and upcoming IPO companies, containing 5,179 text segments sourced from 261 filings, formatted for large language model training and financial analysis.