Alternative data datasets
Browse by provider, category, marketplace, access, delivery and license
A private collection of cleaned earnings-call transcript segments derived from a Motley Fool Kaggle release, split into 135,306 training rows tagged with ticker, exchange, date, and chunk text.
Public-safe distribution bundle from a local r19 counterexample resolution review of arXiv patents using a language model, containing review-related outputs, manifests, checksums, and provenance receipts, but excluding the original arXiv PDFs and any expert validation.
Extracted from SEC EDGAR DEF 14A proxy filings since 2015, the dataset contains over 500,000 structured executive compensation entries for S&P 500 and Russell 2000 issuers, capturing CEO and CFO pay, equity awards, bonuses, incentives, and totals keyed by CIK, ticker, and fiscal year.
Structured numeric records parsed from corporate filings such as 10-K, 10-Q, and 8-K submissions to the U.S. SEC. The figures are packaged inside a compressed DuckDB instance made available through the open-source Datapond registry for fast querying.
A collection of 49,023 patent filings focused on distributed ledger technology is available, comprising roughly 1.296 billion tokens. The corpus is a component of the broader DLT-Corpus initiative and is intended to facilitate natural language processing research, innovation analysis, and patent examination within the distributed ledger field.
Indeed Job Postings (2026) is a 1,000-row sample of US Indeed listings joined to company firmographics, published by fact-den on Apify. License terms are not stated, and the publisher directs users to a commercial source actor for fresh or filtered data.
The Future of Patent Corpus Management One Document, Fully Decoded - Without Anyone Reading It The signal you cannot get from a search interface The highest-value competitor signal in a patent portfolio is not the invention. It is the amount of money and urgency a company committed to it. That signal is never in an abstract, and no search interface surfaces it. It comes out of this corpus for ever
An English-language NLP corpus of cleaned earnings call transcripts gathered from public investor-relations pages, sized between 10,000 and 100,000 documents for financial analysis and LLM work.
PatentPulse PatentPulse is a provenance-preserving corpus of USPTO grants and published patent applications extracted from official weekly XML bulk releases. This immutable Parquet snapshot normalizes the project's historical append-only JSONL into one schema. Historical partial snapshot. This release was captured on 2026-08-28 while the upstream backfill was still in progress. It is a stable, cit
A knowledge graph assembled from 575,778 ClinicalTrials.gov registrations and their associated arms, outcomes, sites, sponsors, conditions, interventions, MeSH codes, and PubMed citations, containing 7,628,735 nodes and 15,531,427 edges, delivered as 623 MB of Parquet files (compared to about 7 GB of raw JSON) and built using Samyama Graph.
The tech-job-postings-salaries-sample dataset is a 500-row public sample of a larger scraping-based feed of technology job postings drawn from public applicant tracking system endpoints such as Greenhouse, Lever, Ashby, Workable and SmartRecruiter. It is published by DataForge (Zalize) for evaluation and preview use.