Alternative data datasets
Browse by provider, category, marketplace, access, delivery and license
Over 635,000 registered clinical studies from ClinicalTrials.gov together with a supplementary pharma intelligence table of roughly 1.06 million rows that map those trials to FDA drug approvals, distributed as part of an open data initiative for academic and personal use.
The tech-job-postings-salaries-sample dataset is a 500-row public sample of a larger scraping-based feed of technology job postings drawn from public applicant tracking system endpoints such as Greenhouse, Lever, Ashby, Workable and SmartRecruiter. It is published by DataForge (Zalize) for evaluation and preview use.
The Tech Job Postings & Salary Dataset from Zalize Data contains over 449,000 monthly job postings scraped from public applicant tracking systems at US and EU technology companies. It is distributed as parquet files under a CC BY-NC 4.0 license and is part of the DataForge Open Data program.
Derived from USPTO and PatentsView bulk releases, this dataset catalogs 9.1 million U.S. patents from 1976 onward along with complete citation links, assignees, inventors, and CPC codes, with the full-graph edition exceeding 255 million rows.
Sentiment analysis in Chinese aggregates Weibo, other social platforms and mainland news sources, aimed at informing A-share equity approaches.
Normalized stock-trade disclosures filed by U.S. House and Senate members under the STOCK Act, regenerated on a schedule by the repository's own pipeline rather than curated manually.
ZipLime US Insider Trading Disclosures (PIT) [!CAUTION] Use knowledge_date, not transaction_date, when backtesting. transaction_date says when a trade occurred; knowledge_date says when the filing became observable through EDGAR. Using the former as the signal date introduces look-ahead bias. visible = trades.filter(pl.col("knowledge_date") <= simulation_time) This dataset normalizes corporate-ins
US Short Interest This dataset publishes no short-interest data, and that is the finding, not a bug. Every US source of equity short positioning was audited on 13 September 2026 before any of it was collected. Two forbid the publication pattern and the third does not exist yet. What is published here is the audit itself — which clause, on which page, read on which day — plus the adapters that will
A curated collection of SEC EDGAR filings from recent and upcoming IPO companies, containing 5,179 text segments sourced from 261 filings, formatted for large language model training and financial analysis.
Web-scraping infrastructure, the Zyte API and Smart Proxy Manager, formerly Scrapinghub, plus fully managed extraction serving millions of pages daily.