The Alternative Data Encyclopedia

Alternative data is data generated outside the traditional financial-reporting and market-data system, used alongside filings and prices to inform investment decisions. It is usually granular, large-scale and sourced from digital exhaust or physical sensors: card transactions, store visits, job postings, scraped prices, satellite imagery, shipping records or on-chain activity.

Research alternative data by dataset, provider, marketplace and signal category. Compare sources used by investors, from transactions and geolocation to jobs, web data, sentiment and on-chain activity, with licensing, delivery, coverage and editorial context in one open catalog.

4,477 datasets · 378 providers · 22 signal categories.

Explore by signal

All categories →

Notable datasets

All datasets →

Explore providers

Compare providers →

Guides & research

All guides →

Signal families in this catalogue

Alternative data is usually grouped by the economic signal it carries. This catalogue indexes 22 families, each with an article on how the signal is created, how it is collected and where it misleads:

How alternative data is collected

Most products are built from a handful of methods. Panels observe a sample of consumers, cards or devices and extrapolate. Web collection reads prices, catalogues, listings and reviews at scale. Public records cover filings, registries, tenders and court documents. Sensors and imagery measure physical activity, from satellites to vessel transponders. Employment and corporate signals come from job postings, technographics and people data. Each method trades coverage against precision and introduces its own bias, which is why the source of a dataset matters more than its size.

How to evaluate a dataset or provider

Six questions decide whether a product is usable: what population the sample actually represents; how the data is generated and whether the method is documented; how far back the history goes and whether it has been restated; how often it refreshes and how quickly it becomes available; what the licence permits, including redistribution and derived work; and how it is delivered, whether by file, API, warehouse or feed. Records here list the licence, delivery, coverage and source listing where the provider publishes them, so you can compare before contacting a vendor.

What this catalogue contains

4,477 dataset records from provider websites, public marketplaces and open repositories, covering 378 providers. Records are never the data itself: each one links to the original listing for files, samples, pricing and terms. Datasets are tiered as featured (first-party), premium (commercial) and open (free). Listings are informational and are not investment advice. The same catalogue is available to machines through the documented API, an MCP server and llms.txt.