Web & Pricing Data
From AltData.wiki, The Alternative Data Encyclopedia
Web & Pricing Data is one of the categories of alternative data covered by The Alternative Data Encyclopedia: Structured collection of product-level web data: prices, availability, catalogs, reviews and promotions captured from retailer and marketplace sites on a fixed schedule. The signal answers questions official statistics cannot, such as what a specific SKU costs at a specific merchant today and how online assortment is shifting week by week.
Price scraping industrialized alongside the e-commerce boom of the 2010s, initially serving retailers and brands with competitive-pricing dashboards. The category became macro-relevant during the 2021-22 inflation wave, when academic and central-bank research demonstrated that high-frequency web prices could nowcast CPI components. Today the market spans pricing-intelligence platforms for commerce teams and data feeds sold into quantitative funds.
The signal
Structured collection of product-level web data: prices, availability, catalogs, reviews and promotions captured from retailer and marketplace sites on a fixed schedule. The signal answers questions official statistics cannot, such as what a specific SKU costs at a specific merchant today and how online assortment is shifting week by week.
Raw output consists of SKU-level snapshots recording price, list versus selling price, availability, ratings and seller identity, keyed to product identifiers and timestamps. Derived layers add price-change events, discount depth, out-of-stock rates and category price indexes constructed from matched products. Vendor platforms extend this with share-of-search, content quality scores and assortment-gap comparisons across competitors.
Why investors pay for it
Prices and availability lead reported revenue and margin: promotional intensity flags demand weakness before sales are published, and stock-out waves anticipate supply constraints. During the 2021-22 inflation surge, several central banks and research groups built web-scraped price nowcasts that moved ahead of official CPI releases. Equity analysts use the same streams to benchmark competitive pricing power and track market-share shifts among retailers and brands.
Vendors operate distributed crawlers behind residential and datacenter proxy fleets, with dedicated parsers per retailer template and scheduled recrawl frequencies ranging from daily to near-real-time for high-value categories. Matching SKUs across merchants combines exact identifiers (GTIN/EAN/UPC), title normalization and embedding-based similarity; leading providers report matching accuracy above ninety-nine percent backed by human-in-the-loop verification. Point-in-time discipline requires storing each crawl as an immutable snapshot so historical index construction can be reproduced.
Who uses it
Consumer-retail and e-commerce equity analysts track price gaps, promotion cycles and assortment share; inflation researchers and macro teams consume category indexes as nowcast inputs. CPG and retail pricing teams buy the same data commercially for competitive response.
Questions to ask vendors in this category
How broad and frequent is coverage per retailer and category? What is SKU-match accuracy across merchants, and how do you handle rematching when listings change? What is your robots.txt and terms-of-service compliance posture? How far back does clean history go, and is it point-in-time? How are out-of-stocks treated in index construction?
Complementary signals
This signal pairs naturally with adjacent categories of the encyclopedia:
Caveats and limitations
Coverage and crawl frequency create survivorship-like bias because delisted products drop out of panels. A single daily snapshot can miss intraday repricing on dynamically priced marketplaces. Seller churn and geographic price variation inject noise, and category mix drives index results as much as underlying prices.
Compliance and legal considerations
Scraping sits against site terms of service and database rights, and relevant case law remains jurisdiction-dependent after disputes such as hiQ versus LinkedIn. Reputable vendors document robots.txt policies, exclude personal data from review content, and treat product imagery and copy under copyright constraints.
Further reading
Providers in this category
The register lists 33 companies for this signal family: