nyse

From AltData.wiki, The Alternative Data Encyclopedia · updated 2017-02-22

S&P 500 companies historical prices with fundamental data

dgawlik/nyse is a Web & Pricing Data data product published on Kaggle and indexed by The Alternative Data Encyclopedia.

Dgawlik publishes nyse as a Kaggle offering in the Web & Pricing Data signal family. The listing has drawn 110,062 downloads and 1402 likes on Kaggle.

The data

S&P 500 companies historical prices with fundamental data

It is one of 58 listings in the Web & Pricing Data family; between them the practical differences come down to coverage, history depth and how the raw signal is cleaned and delivered.

The source tags it with business, finance and investing.

Structure, access and licensing

License: CC0: Public Domain. Size: 32148316. Last updated: 2017-02-22.

Access is through a dataset that is downloaded from Kaggle for use in a notebook. Its most recent recorded snapshot is from 2017-02-22.

The signal

Structured collection of product-level web data: prices, availability, catalogs, reviews and promotions captured from retailer and marketplace sites on a fixed schedule. The signal answers questions official statistics cannot, such as what a specific SKU costs at a specific merchant today and how online assortment is shifting week by week.

Prices and availability lead reported revenue and margin: promotional intensity flags demand weakness before sales are published, and stock-out waves anticipate supply constraints. During the 2021-22 inflation surge, several central banks and research groups built web-scraped price nowcasts that moved ahead of official CPI releases. Equity analysts use the same streams to benchmark competitive pricing power and track market-share shifts among retailers and brands. Raw output consists of SKU-level snapshots recording price, list versus selling price, availability, ratings and seller identity, keyed to product identifiers and timestamps. Derived layers add price-change events, discount depth, out-of-stock rates and category price indexes constructed from matched products. Vendor platforms extend this with share-of-search, content quality scores and assortment-gap comparisons across competitors.

Vendors operate distributed crawlers behind residential and datacenter proxy fleets, with dedicated parsers per retailer template and scheduled recrawl frequencies ranging from daily to near-real-time for high-value categories. Matching SKUs across merchants combines exact identifiers (GTIN/EAN/UPC), title normalization and embedding-based similarity; leading providers report matching accuracy above ninety-nine percent backed by human-in-the-loop verification. Point-in-time discipline requires storing each crawl as an immutable snapshot so historical index construction can be reproduced.

Caveats and compliance

Coverage and crawl frequency create survivorship-like bias because delisted products drop out of panels. A single daily snapshot can miss intraday repricing on dynamically priced marketplaces. Seller churn and geographic price variation inject noise, and category mix drives index results as much as underlying prices.

Scraping sits against site terms of service and database rights, and relevant case law remains jurisdiction-dependent after disputes such as hiQ versus LinkedIn. Reputable vendors document robots.txt policies, exclude personal data from review content, and treat product imagery and copy under copyright constraints.

Who uses this signal

Consumer-retail and e-commerce equity analysts track price gaps, promotion cycles and assortment share; inflation researchers and macro teams consume category indexes as nowcast inputs. CPG and retail pricing teams buy the same data commercially for competitive response.

Complementary signals

This kind of signal pairs naturally with adjacent categories of the encyclopedia:

Further reading

Discussion

Anchored on 𝕏 with the commit-style tag #… — tweet with it and the thread picks it up.

Discuss on 𝕏

No comments yet — start the thread on 𝕏.

More in Web & Pricing Data