State of Alternative Data 2026
A live snapshot of the AltData.wiki catalog · updated September 2, 2026
This report summarizes the structure of the alternative-data catalog itself rather than estimating industry revenue or vendor spend. The counts below are generated from the current AltData.wiki register, so they describe what is indexed across provider websites and public data marketplaces at the time of publication.
4,465
canonical dataset records
378
providers in the register
22
signal categories
37.2%
datasets meeting the SEO quality gate
Catalog quality and indexability
1,661 of 4,465 dataset records currently meet the encyclopedia's search-index quality threshold. The remaining records stay browsable and available through the API, but are marked noindex until their category and core metadata are strong enough to stand on their own.
750 records remain unclassified, and 765 datasets resolve confidently to a provider in the curated provider register. These gaps are exposed rather than silently filled. See the methodology and editorial policy for the scoring and inference rules.
Where dataset listings come from
| Source | Dataset listings | Share of catalog |
|---|---|---|
| Databricks Marketplace | 2,181 | 48.8% |
| Snowflake Marketplace | 1,021 | 22.9% |
| Kaggle | 600 | 13.4% |
| Provider websites | 323 | 7.2% |
| Hugging Face | 261 | 5.8% |
| AWS Data Exchange | 74 | 1.7% |
| Datarade | 23 | 0.5% |
| TheirStack | 3 | 0.1% |
Cloud marketplaces now account for a large share of discoverable alternative-data inventory. That breadth is useful for discovery, but marketplace listings vary sharply in metadata depth; a listing appearing in the raw catalog does not automatically qualify for search indexing.
Largest signal categories
The category distribution measures listing availability, not investor demand. A category with many marketplace products can be easier to discover while still being harder to diligence if history, methodology or point-in-time behavior is poorly documented.
Provider geography
| Primary geography | Providers |
|---|---|
| US | 207 |
| Global | 84 |
| UK | 26 |
| China | 10 |
| Canada | 8 |
| France | 5 |
| Israel | 4 |
| Norway | 3 |
| Ireland | 3 |
| India | 3 |
| Japan | 3 |
| Germany | 2 |
| Australia | 2 |
| Lithuania | 2 |
| South Korea | 2 |
How providers deliver data
| Delivery method | Providers |
|---|---|
| API | 195 |
| Dashboard | 101 |
| Feed | 42 |
| Bulk | 41 |
| S3 | 14 |
| Calls | 9 |
| Report | 7 |
| Survey | 4 |
| Web | 4 |
| Terminal | 1 |
| Snowflake | 1 |
| Parquet | 1 |
Provider delivery labels are normalized before counting, so “Dashboard, API” and “API, Dashboard” contribute to the same API and Dashboard totals. A provider can appear in more than one delivery row.
Provider price bands
| Price band | Providers |
|---|---|
| $150k+ | 107 |
| $50–150k | 47 |
| Free | 42 |
| < $50k | 9 |
| Custom / enterprise | 7 |
| Subscription | 4 |
| Free / Subscription | 2 |
| Licensed | 1 |
Price bands are directional rather than quotes. Enterprise contracts depend on history depth, permitted use, entity coverage, redistribution rights and service levels, so the register uses bands to support comparison rather than pretending that vendor pricing is static.
What the 2026 catalog suggests
Three patterns stand out. First, discovery has become marketplace-heavy: cloud catalogs expose thousands of products that previously required direct vendor outreach. Second, metadata quality has not kept pace with inventory growth, which makes normalization, provider identity and source transparency more valuable than raw listing count. Third, useful provider comparison increasingly depends on operational fields — geography, delivery, history, refresh and licensing — rather than category labels alone.
For buyers, the practical implication is to treat catalog breadth as the beginning of diligence, not the end. Use the provider comparison to narrow candidates, open the underlying dataset records, and then follow the vendor-evaluation workflow before treating any feed as production-ready.