Alternative data datasets
Browse by provider, category, marketplace, access, delivery and license
What companies are doing, buying, complying with and investing in, from the topics they hire for. 9.4k topics across 14 keyword types rolled up per company: one row per company and topic with confidence and first/last-seen dates, so you can reach out while the need is still open.
Every job posting collected since 2021: who is hiring, for what, where and at what salary. One row per posting with title, full description, location, salary and company; closed postings ship as their own file with the day each one went offline. Companies file included for firmographic joins.
Which of 33k technologies each company runs, inferred from what they hire for. One row per company and technology with confidence score, number of postings mentioning it and first/last-seen dates; technologies catalog and companies file included.
View source information and access options.
An intelligent research assistant for searching across filings, transcripts, broker research and news, built to support competitive intelligence and thesis validation.
View source information and access options.
View source information and access options.
View source information and access options.
View source information and access options.
View source information and access options.
View source information and access options.
Resources for spotting, measuring, and keeping track of credit risk using shared credit assessment data.
99.90% Accuracy · 12 Categories · 100+ File Extensions · Soft-Voting Ensemble
FactSet consensus forecasts refreshed throughout the trading day, compiled from roughly eight hundred contributing analysts.
Daily point-in-time consensus forecasts aggregated from a broad network of contributing analysts.
View source information and access options.
Japan listed firms receive analyst consensus projections refreshed every business day.
Japanese public companies have their consensus forecasts refreshed on a daily schedule.
Refine audience segmentation algorithms and filter out non-legitimate ad interactions to lift marketing ROI.
View source information and access options.
Aggregated analyst projections for corporate profits and sales, combined with price targets and professional buy or sell ratings.
Wallet labeling service that identifies high-performing addresses and monitors their token activity.
Compilation of over 11 million UK-registered companies, featuring legal particulars, industry sector codes, director details, and financial metrics such as income and profit.
Roster of present-day and past candidates running for U.S. federal and state-level political offices.
Fifteen-day outlooks derived from 51 ECMWF ensemble forecasts and aligned to user-defined coordinates.
View source information and access options.
View source information and access options.
View source information and access options.
Insider Trades India
Collection of 15 Harmful Insects Found in Farming Environments
A high-volume retail transaction analytics system that supports custom panel queries for demand forecasting.
Parsed eligibility criteria from 13,229 ClinicalTrials.gov studies representing the candidate pool gathered by two first-stage retrievers during TREC Clinical Trials 2021–2023, formatted as typed entity-relation graphs for reranking models.
This dataset is a small parquet-format subset of clinical-trial eligibility criteria represented as entity and relation graphs. It is published on a community data hub under an unspecified license.
This dataset is a parsed snapshot of ClinicalTrials.gov records, distributed in Parquet format on the Hugging Face Hub under an unspecified license. It contains roughly 52,000 trial entries drawn from public registry filings.
Programmatic access delivers app store details, chart positions and software development kit intelligence for both iOS and Google Play.
Internet usage tied to firm IP ranges exposing the subjects and suppliers staff are researching, surfacing pre-RFP purchase intent indicators.
Indicators inferred from payment processor flows that serve as proxies for merchant sales trajectories.
Quarterly conference calls are processed to detect revisions in guidance and shifts in executive language.
Block-level air pollution measurements gathered through dense networks of local environmental sensors.
Datamule, Teraflop AI, and Eventual jointly contributed a 590 gigabyte SEC EDGAR collection encompassing about 8 million filings and 43 billion tokens, assembled using the datamule-python library and the official datamule API.
Mobile measurement platform supplying install records, attribution outcomes and in-application event tracking for advertiser apps.
A sharded Parquet archive of NSE equities and index prices from India spanning 2000 to 2026, covering more than 2,500 tickers and organized into roughly 1.5 GB files for efficient streaming.
Foot-traffic panels built from mobile-device location signals, measuring visits, dwell time and cross-shopping for US retailers, malls and real-estate assets, delivered on a daily schedule to investors and real-estate desks.
credit_card_transactions is a small tabular dataset published by aegisheld containing anonymized customer-level credit card account and spending summaries. It is distributed as CSV with a single training split of 8,950 rows.
A US debit and credit expenditure dataset built from hundreds of millions of consumers and used to model same-store performance.
The largest consumer panel survey in Japan, drawing responses from over 50,000 participants.
Consumer purchase survey responses from over 53,000 individuals in Japan are used to monitor retailer sales performance.
Worldwide airport passenger volume and movement figures sourced from ACI, offering demand metrics at the individual airport level.
Movement analytics computed from anonymized cellular-network signalling that spans hundreds of millions of devices, feeding trade-area studies, commuting research and travel-demand modelling.
Numerical sentiment indicators derived from financial press coverage and social platforms to gauge mood toward publicly listed companies worldwide.
Material contracts and agreements extracted from EDGAR filings by the KL3M project: debt, M&A, employment and license agreements, millions of documents in Parquet. Contract language is a niche but real input for M&A, financing and litigation-event signals.
The kl3m-index-edgar-filings dataset, published by ALEA Institute, is a Parquet-format index of SEC EDGAR filings distributed via the Hub datasets library. It contains roughly 20 million tabular records spanning the 10M–100M size category, with an unstated license.
US patent full text assembled by the Allen Institute from the USPTO, millions of records under ODC-BY: claims, abstracts, descriptions and metadata, in Parquet. A broad base for measuring IP intensity, technology landscapes and the direction of R&D spending.
Opinion-scoring system that evaluates conventional news outlets, blog content, and social networks to produce equity-oriented sentiment signals.
A leading expert network for advisory boards and single-expert interviews across industries, companies and strategy.
Network-level DNS query data exposing traffic patterns tied to individual domains.
A private collection of cleaned earnings-call transcript segments derived from a Motley Fool Kaggle release, split into 135,306 training rows tagged with ticker, exchange, date, and chunk text.
Curated and standardized web metrics spanning online retail and digital industries, formatted for use by financial and investment analytics teams.
Text analytics that pulls sentiment signals and event markers from regulatory filings, earnings call transcripts and news articles.
Subscription research service delivering executive-level briefings, quantitative models and curated datasets authored by seasoned former leaders in technology, automotive and consumer sectors.