State of Alternative Data 2026

A live snapshot of the AltData.wiki catalog · updated September 2, 2026

This report summarizes the structure of the alternative-data catalog itself rather than estimating industry revenue or vendor spend. The counts below are generated from the current AltData.wiki register, so they describe what is indexed across provider websites and public data marketplaces at the time of publication.

4,465

canonical dataset records

378

providers in the register

22

signal categories

37.2%

datasets meeting the SEO quality gate

Catalog quality and indexability

1,661 of 4,465 dataset records currently meet the encyclopedia's search-index quality threshold. The remaining records stay browsable and available through the API, but are marked noindex until their category and core metadata are strong enough to stand on their own.

750 records remain unclassified, and 765 datasets resolve confidently to a provider in the curated provider register. These gaps are exposed rather than silently filled. See the methodology and editorial policy for the scoring and inference rules.

Where dataset listings come from

SourceDataset listingsShare of catalog
Databricks Marketplace2,18148.8%
Snowflake Marketplace1,02122.9%
Kaggle60013.4%
Provider websites3237.2%
Hugging Face2615.8%
AWS Data Exchange741.7%
Datarade230.5%
TheirStack30.1%

Cloud marketplaces now account for a large share of discoverable alternative-data inventory. That breadth is useful for discovery, but marketplace listings vary sharply in metadata depth; a listing appearing in the raw catalog does not automatically qualify for search indexing.

Largest signal categories

CategoryDatasets
Foot Traffic & Mobility448
Public Records & Filings426
Technographics & Tech Stacks414
Free & Open Data302
Healthcare & Clinical291
Web & Pricing Data193
News & Sentiment178
Jobs & Workforce175
Weather & Agriculture170
Estimates, Events & Fund Flows159
Crypto & On-Chain158
Satellite & Geospatial140
Supply Chain & Shipping139
ESG & Climate135
Card & Transactions133
Credit & Lending62
App Usage & Downloads50
Sensors & IoT38
Private Markets Signals28
Auto & Vehicle Data27
Expert Networks & Surveys26
Email Receipts23

The category distribution measures listing availability, not investor demand. A category with many marketplace products can be easier to discover while still being harder to diligence if history, methodology or point-in-time behavior is poorly documented.

Provider geography

Primary geographyProviders
US207
Global84
UK26
China10
Canada8
France5
Israel4
Norway3
Ireland3
India3
Japan3
Germany2
Australia2
Lithuania2
South Korea2

How providers deliver data

Delivery methodProviders
API195
Dashboard101
Feed42
Bulk41
S314
Calls9
Report7
Survey4
Web4
Terminal1
Snowflake1
Parquet1

Provider delivery labels are normalized before counting, so “Dashboard, API” and “API, Dashboard” contribute to the same API and Dashboard totals. A provider can appear in more than one delivery row.

Provider price bands

Price bandProviders
$150k+107
$50–150k47
Free42
< $50k9
Custom / enterprise7
Subscription4
Free / Subscription2
Licensed1

Price bands are directional rather than quotes. Enterprise contracts depend on history depth, permitted use, entity coverage, redistribution rights and service levels, so the register uses bands to support comparison rather than pretending that vendor pricing is static.

What the 2026 catalog suggests

Three patterns stand out. First, discovery has become marketplace-heavy: cloud catalogs expose thousands of products that previously required direct vendor outreach. Second, metadata quality has not kept pace with inventory growth, which makes normalization, provider identity and source transparency more valuable than raw listing count. Third, useful provider comparison increasingly depends on operational fields — geography, delivery, history, refresh and licensing — rather than category labels alone.

For buyers, the practical implication is to treat catalog breadth as the beginning of diligence, not the end. Use the provider comparison to narrow candidates, open the underlying dataset records, and then follow the vendor-evaluation workflow before treating any feed as production-ready.