Web scraping and compliance in alternative data

From AltData.wiki, The Alternative Data Encyclopedia · Primer · 8 min read

How scraped datasets are actually collected, where terms of service and computer-muse laws draw lines, what hiQ v. LinkedIn settled and did not, and the diligence checklist for legally fragile data.

Scraping built half this industry

A large share of alternative data originates as web scraping: prices collected from retailer sites, postings from career pages, reviews from platforms, menus, flight fares, and app metadata. The practice is neither inherently legal nor inherently illegal; it sits on a spectrum defined by what is scraped, how, from whom, and under which terms of service. Understanding that spectrum is a core diligence requirement for any buyer of scraped-origin data rather than an afterthought. Many dataset families surveyed in introductions to alternative data, including job postings and web-priced goods, trace their raw inputs to collection pipelines of this kind.

The economic logic is straightforward. The web is the largest continuously updated record of commercial behavior ever assembled, and much of it is publicly readable at zero marginal cost. Firms that convert those pages into structured, timestamped, entity-resolved tables create products that no single source publishes natively. This explains why scraping remains widespread even among large, well-funded vendors who could license comparable feeds: in many categories no official feed exists at all, and the scraped corpus is the primary source from which any commercial product would have to be built.

A short history

Automated collection predates the alternative-data industry itself. Search-engine crawlers normalized the idea of machine access to public pages in the 1990s, alongside conventions such as robots.txt that signaled which paths collectors should avoid. Through the 2000s, price comparison services, academic archiving projects, and open efforts such as Common Crawl demonstrated that useful structured corpora could be assembled from the open web at scale. Quantitative funds began experimenting with scraped retail prices and listings during the early 2010s, initially building collection capability in-house.

Professionalization followed quickly. Collection moved from bespoke scripts toward commercial infrastructure, with specialized providers selling proxy networks, browser automation, and managed extraction, some of which are documented in directories of tools such as Bright Data and similar platforms. At the same time, legal departments, compliance review, and formal licensing replaced the informal posture of earlier years. The result is a layered market today: a few vendors own end-to-end pipelines for flagship datasets, while many smaller firms resell, clean, or repackage collections originally gathered elsewhere.

United States: the central statutory question under the Computer Fraud and Abuse Act — whether accessing publicly accessible pages without authorization exceeds authorized access — was substantially narrowed by hiQ v. LinkedIn and subsequent case law, though login-walled scraping remains exposed and later procedural developments left the underlying contract questions unsettled. Contracts add an independent layer: terms of service can prohibit scraping regardless of the CFAA, and breach-of-contract claims do not require any criminal statute. Copyright can attach to creative content displayed on pages, while facts and data themselves enjoy thinner protection, although compilation rights vary.

Europe: the GDPR governs personal data appearing within scraped content — names, photographs, profiles, contact details — and requires a lawful basis for processing that indiscriminate scraping rarely supplies. This is why EU-linked scraped personal data is the category most frequently abandoned or geofenced by US-focused vendors, even when identical collection would raise no US privacy issue. Separately, the sui generis database right protects substantial European compilations against extraction, adding a property-style layer on top of privacy law. Nothing here constitutes legal advice; outcomes depend on jurisdiction, data type, and purpose.

Elsewhere the picture is more fragmented. Some jurisdictions rely on general contract and unfair-competition doctrines rather than dedicated computer-misuse statutes; others protect databases explicitly. Enforcement priorities also differ: consumer privacy regulators, competition authorities, and platform operators may each police the same collection activity under different theories. Because these regimes interact unpredictably — a dataset lawful in one country can be problematic in another — vendors operating internationally typically adopt the strictest applicable standard as their default posture, and sophisticated buyers treat cross-border provenance as a diligence item in its own right.

Technical courtesy and its evidentiary weight

Respectful collectors honor robots.txt where present, rate-limit politely, identify themselves honestly in user-agent headers, avoid login walls, and cache aggressively to minimize server load. These practices do not legalize otherwise prohibited scraping, but they matter in three distinct ways. They reduce the harm arguments available to a complaining site in any dispute; they demonstrate good faith to regulators and courts assessing reasonableness; and they correlate strongly with vendors whose collection survives platform crackdowns over long periods. Courtesy is therefore best read as evidence of operational maturity rather than as a safe harbor.

Buyers can test these practices directly. Ask vendors about their robots.txt policy and how exceptions are handled; request sample identification headers; ask how quickly a takedown request from a source site is honored and whether honoring one has ever materially changed coverage. Vendors that answer precisely and document their answers tend to manage collection professionally, while evasive answers about basics often foreshadow weaker controls on matters that are harder to verify, such as source concentration or the handling of personal data encountered incidentally during collection.

Platform-specific risk

Large platforms enforce their terms both technically and legally, and enforcement intensity shifts with corporate mood, leadership changes, and litigation posture. Datasets sourced wholly from a single platform concentrate that risk: one enforcement action, one API closure, or one redesign can end supply abruptly and without appeal. Diversified collectors, drawing each signal from many independent sites, degrade gracefully when individual sources disappear. Source diversification is thus not only a legal hedge but also an operational resilience property that determines whether a vendor's history looks continuous or gappy.

When diligencing a scraped dataset, ask for the source-site distribution over time, not just at present. A top-five-sites concentration above half the corpus deserves scrutiny, as does any recent cliff in the contribution curve that suggests a lost source being quietly backfilled. Coverage maps should be compared against the marketing description: a product sold as economy-wide but sourced predominantly from a handful of marketplaces behaves very differently from its pitch. Historical reconstruction quality also varies, since backfilled histories may mix collection eras with inconsistent schemas.

How practitioners use scraped-origin data

In practice, scraped data reaches investment workflows through several standard routes. Price series scraped from retailer sites support nowcasts of inflation components and promotional intensity; job listings feed hiring and expansion indicators; reviews and ratings proxy product quality shifts; availability signals track supply shortages before they appear in earnings. Most funds do not consume raw pages at all. Instead they buy indicator-level products built on scraped foundations, delivered through the channels described in guides to alt data delivery formats, and evaluate them exactly as they would any other third-party input.

A second, quieter use is competitive and market intelligence outside investing: monitoring distributor pricing, catalog churn, and channel compliance. These consumers are typically less sensitive to statistical rigor and more sensitive to freshness and completeness, which pushes vendors to maintain parallel quality tiers for the same underlying collection. The dual demand base helps sustain scraping-focused businesses through periods when investor interest cools, and it explains why some datasets remain commercially available despite limited demonstrated value in financial applications.

Common failure modes

Scraped datasets fail in characteristic ways. Silent collector breakage occurs when a source site redesigns and fields begin populating incorrectly or not at all, sometimes for weeks before anyone notices. Schema drift produces subtle discontinuities that contaminate backtests. Coverage survivorship arises when delisted items vanish from historical views because the archive mirrors the live site rather than a true panel. Personal data contamination happens when names, emails, or identifiers ride along in fields sold as purely commercial, creating privacy obligations the buyer never priced.

Licensing-era mismatches are another recurring problem: terms of service change mid-contract, and a collection that was defensible at license inception becomes contested before renewal. Buyers sometimes discover that the vendor's own supply chain involves subcontracted collectors whose practices nobody audited. None of these failures announce themselves; they surface through reconciliation breaks against ground truth, unexplained level shifts, or a regulator's inquiry. The mitigation is procedural — monitoring, source audits, contractual warranties — rather than heroic single checks performed at onboarding.

A buyer's checklist

Before licensing scraped-origin data: confirm the collection methodology is described accurately in writing, including sites, frequency, and exclusion rules; verify that no personal data is included beyond lawful bases; check the vendor's terms-of-service posture and any litigation history; understand source concentration and fallback sites; confirm that takedown and correction processes exist and have been exercised; and ensure your own use case, including redistribution and client reporting, falls within the license scope. Structured vendor-evaluation frameworks such as those covering web and pricing data apply directly here.

Treat the checklist as ongoing rather than one-time. Re-verify methodology descriptions annually, monitor reconciliation against your own observations, and maintain an exit plan for datasets whose legal posture could shift. Data that cannot survive these questions can still be explored opportunistically, but it should never anchor a strategy, because the cost of losing an anchoring dataset mid-process extends well beyond the subscription fee: models, research narratives, and client commitments built atop the signal all need rebuilding.

Limitations and criticism

Critics raise substantive objections from several directions. Publishers argue that wholesale appropriation of listings, reviews, and prices free-rides on investments made by the sites that created the content, and several high-profile disputes have tested that argument. Data-quality researchers note that scraped corpora inherit every bias of their sources — what a site chooses to list, display, and rank — while presenting themselves as neutral measurements. And some observers contend that the legal ambiguity itself imposes costs, favoring well-capitalized firms able to litigate over smaller participants.

These criticisms have not halted the practice, but they shape its evolution. Vendors increasingly emphasize licensed sources, partnerships, and first-party relationships where feasible; buyers increasingly demand provenance documentation; and the industry's center of gravity continues shifting from raw page capture toward derived indicators where the underlying facts carry thinner protection than their creative presentation. The likely trajectory is continued growth of scraped-derived analytics alongside a narrowing band of collection practices considered professionally acceptable.

Frequently asked questions

Is web scraping illegal?
Not categorically. Publicly accessible, non-personal data collected respectfully is generally defensible in the US; login-walled content, personal data in Europe, and ToS-prohibited collection each add distinct legal exposure.
Did hiQ v. LinkedIn make scraping legal?
It narrowed CFAA exposure for public pages but decided neither copyright, contract, nor GDPR questions — and the parties later settled with LinkedIn's terms prevailing contractually.
Should funds buy scraped data?
Most do, knowingly. The diligence goal is understanding exactly what was collected, from where, under which terms — and matching that to your risk appetite.

Further reading