Alternative data datasets

Browse by provider, category, marketplace, access, delivery and license

Showing 61–71 of71 results for “language:en”
Trelis
big_patent_sample

Reduced sample of the big_patent collection, balanced across text lengths to support shorter fine-tuning runs, with a roughly even spread up to one million characters suitable for training on sequences of up to 250,000 tokens.

Crypto & On-Chain Free View
VaidhyaMegha
clinicaltrials-kg

A knowledge graph assembled from 575,778 ClinicalTrials.gov registrations and their associated arms, outcomes, sites, sponsors, conditions, interventions, MeSH codes, and PubMed citations, containing 7,628,735 nodes and 15,531,427 edges, delivered as 623 MB of Parquet files (compared to about 7 GB of raw JSON) and built using Samyama Graph.

Public Records & Filings Free View
VeeraThakshith
stock-market-tweets-data

A collection of 943,672 tweets gathered between April 9 and July 16, 2020, harvested via the #SPX500 tag, references to the top 25 S&P 500 firms, and the #stocks tag, originally published by Bruno Taborda on IEEE.

News & Sentiment Free View
yeong-hwan
2024-earnings-call-transcript

The 2024 Earnings Call Transcript dataset, published by yeong-hwan on Hugging Face, aggregates quarterly earnings call dialogue for individual stock tickers. It is distributed as a JSON table with 1,904 rows in a single train split.

Public Records & Filings Free View
zalizedata
clinical-trials-drug-approvals-dataset

Over 635,000 registered clinical studies from ClinicalTrials.gov together with a supplementary pharma intelligence table of roughly 1.06 million rows that map those trials to FDA drug approvals, distributed as part of an open data initiative for academic and personal use.

Public Records & Filings Free View
zalizedata
tech-job-postings-salaries-sample

The tech-job-postings-salaries-sample dataset is a 500-row public sample of a larger scraping-based feed of technology job postings drawn from public applicant tracking system endpoints such as Greenhouse, Lever, Ashby, Workable and SmartRecruiter. It is published by DataForge (Zalize) for evaluation and preview use.

Jobs & Workforce Free View
zalizedata
tech-job-postings-salary-dataset

The Tech Job Postings & Salary Dataset from Zalize Data contains over 449,000 monthly job postings scraped from public applicant tracking systems at US and EU technology companies. It is distributed as parquet files under a CC BY-NC 4.0 license and is part of the DataForge Open Data program.

Jobs & Workforce Free View
zalizedata
us-patents-citation-company-dataset

Derived from USPTO and PatentsView bulk releases, this dataset catalogs 9.1 million U.S. patents from 1976 onward along with complete citation links, assignees, inventors, and CPC codes, with the full-graph edition exceeding 255 million rows.

Public Records & Filings Free View
ZipLime
congress-trading

Normalized stock-trade disclosures filed by U.S. House and Senate members under the STOCK Act, regenerated on a schedule by the repository's own pipeline rather than curated manually.

Public Records & Filings Free View
ZipLime
insider-trading

ZipLime US Insider Trading Disclosures (PIT) [!CAUTION] Use knowledge_date, not transaction_date, when backtesting. transaction_date says when a trade occurred; knowledge_date says when the filing became observable through EDGAR. Using the former as the signal date introduces look-ahead bias. visible = trades.filter(pl.col("knowledge_date") <= simulation_time) This dataset normalizes corporate-ins

Public Records & Filings Free View
zorynthiq
zoryntiq-sec-filings

A curated collection of SEC EDGAR filings from recent and upcoming IPO companies, containing 5,179 text segments sourced from 261 filings, formatted for large language model training and financial analysis.

Public Records & Filings Free View