Alternative data for investing

From AltData.wiki, The Alternative Data Encyclopedia | Primer | 5 min read

How funds turn alternative data into an investment edge: signal families, evaluation, backtesting, compliance and the build-vs-buy decision, mapped to the datasets and providers indexed on this encyclopedia.

Why alternative data entered the investment process

Traditional fundamental analysis relies on quarterly filings, sell-side estimates and management guidance, all of which arrive late and are produced by interested parties. Alternative data closes part of that gap by observing what companies actually do — hire, price, ship, receive cards, move foot traffic and publish — in near real time and ahead of disclosure. For a fund, the value is not the data itself but the lead time: a signal that predicts a revision in revenue, margins or guidance by days or weeks is monetizable even when the correlation is modest.

The category grew out of three changes. Commodity computing made it cheap to collect and store unstructured traces; web and mobile activity made company behavior observable; and a generation of quants proved that edge could be found outside the terminal. The result is a market of hundreds of providers across dozens of signal families, which is exactly what this encyclopedia indexes. Knowing the landscape — who has what data, how it is collected, and what it can and cannot do — is now a standard part of the research stack rather than a differentiator by itself.

From dataset to tradeable signal

A dataset becomes a signal through a pipeline with fixed steps: acquisition, ingestion, entity resolution, validation and backtesting. Acquisition is the commercial step — signing an evaluation, receiving a sample and a data dictionary, and negotiating license terms. Ingestion is the engineering step — landing raw files, normalizing schemas and handling revisions. Entity resolution is the analytical step that most often determines success: mapping a vendor's identifiers (company names, store chains, card hashes) to the tickers and securities a portfolio actually holds.

Validation is where most datasets fail. A team must confirm that history is point-in-time correct (no look-ahead bias, no restated series), that coverage is stable and representative, and that the signal is not an artifact of collection. Only then does backtesting make sense: walk-forward tests, out-of-sample periods, turnover and cost modeling, and an honest comparison against a null model. The same discipline applies whether the signal is job postings, card transactions, web traffic or satellite imagery.

The framing in this guide applies across every category in the register. Each category article — for example job postings, card transactions or foot traffic — describes how that specific signal is collected, what drives it and what its failure modes are.

The main signal families

Signal families differ by what they observe and how early they lead fundamentals. Workforce data (job postings, payroll, job changes) is a leading indicator of growth intent and cost control. Consumer data (card transactions, email receipts, foot traffic, app usage) measures realized demand directly. Corporate data (web pricing, technographics, patents, shipping) tracks strategy, adoption and supply chain. Market-structure data (options flow, short interest, fund flows, insider filings) captures positioning. And news, sentiment and event data turns unstructured text into forward-looking scores.

No single family is universally best; the edge comes from matching the family to the question. A long-short equity fund forecasting same-store sales will look first at card and foot-traffic data; a private-credit fund underwriting a borrower will look at payroll and job data; a macro fund will watch shipping, commodity flows and high-frequency price data. This encyclopedia indexes providers and datasets per family so a sourcing team can shortlist options without starting from vendor marketing.

Evaluation and onboarding

Evaluation follows a standard arc, described in more depth in How funds evaluate alternative data vendors: a free sample, a data dictionary and historical pull, a reconstruction of a known window, and a paid pilot before an enterprise license. The questions are the same everywhere — depth of history, survivorship, point-in-time correctness, entity coverage, refresh cadence and the compliance story behind collection.

Onboarding is where hidden costs live. Data arrives in inconsistent formats and schemas; entity mapping is labor-intensive; and the signal decays if the vendor changes collection methodology without notice. Funds that succeed treat data as a product with an owner, versioning and a monitoring step that alerts when coverage or freshness drifts. The sections below and the delivery formats guide cover the operational side of that commitment.

Backtesting and point-in-time integrity

Backtesting alternative data is harder than backtesting price data because the series are revised. A job-postings file downloaded today may have backfilled or corrected history, which silently leaks the future into the past. Point-in-time integrity means storing each snapshot as it was published and testing only on data that was available on the decision date. Without it, results are optimistic and the signal will disappoint in production.

The discipline extends to survivorship: a dataset of companies that exist today omits the ones that failed, biasing any study of the past. The same goes for coverage drift, when a vendor adds regions or channels over time. A rigorous test asks what the strategy would have returned using only contemporaneously available, consistently collected data — and reports the gap between that and a naive backtest.

Alternative data carries legal risk concentrated in how the data was collected. Web scraping can implicate terms of service, the Computer Fraud and Abuse Act and, in the EU, the General Data Protection Regulation; personal data and material non-public information are red lines. The web scraping and compliance guide walks through the boundaries, including the hiQ v. LinkedIn line of cases and the difference between public information and personal data.

Licensing is the other half. Contracts define permitted users, redistribution, derived-data rights and what happens at termination. Funds increasingly run vendor due diligence — the same know-your-customer discipline applied to data — before signing, and keep an audit trail of licenses so a position is never built on data a fund cannot prove it was entitled to use.

Build versus buy, and the cost question

The build-versus-buy decision is really a question of where the moat lives. Collecting raw data in-house is expensive but defensible; licensing a widely available dataset is cheaper but the edge is shared with every other subscriber. Most funds land somewhere in the middle: license a core signal, then add proprietary processing, entity mapping and feature engineering that competitors do not have.

Costs range from free open data sources through five-figure category datasets to seven-figure enterprise feeds. The register lists a price band per provider so a sourcing team can gauge the cost of the shortlist before engaging sales. The open vs commercial guide covers what is realistically available for free and where paying is unavoidable.

Frequently asked questions

What is alternative data in investing?
Non-traditional datasets — job postings, card transactions, web traffic, satellite imagery, app usage, options flow — used to forecast company fundamentals, positioning or macro conditions ahead of official disclosure.
How much lead time does alternative data provide?
It varies by family. Workforce and consumer signals often lead quarterly results by weeks to months, while market-structure signals can move intraday. The lead is only useful if the data is ingested and scored in time to act.
Which signal family should a fund start with?
The one closest to the question it is trying to answer: consumer data for same-store sales, workforce data for growth and cost, market-structure data for positioning, and geospatial or shipping data for supply-chain and macro questions.
What is point-in-time integrity and why does it matter?
It means testing a strategy only on data that was available on the decision date, not on revised or backfilled history. Without it, backtests are optimistically biased and live performance disappoints.
How much does alternative data cost?
From free (government and open datasets) to mid five figures for category datasets and seven figures for enterprise feeds. The provider register lists price bands to help shortlist before engaging sales.
Is alternative data legal?
Mostly, but collection matters: scraping can raise terms-of-service and privacy issues, and material non-public information is off limits. Funds run vendor due diligence and keep a license audit trail.

Further reading