Zyte

From AltData.wiki, The Alternative Data Encyclopedia

Zyte is a web data extraction company that provides a scraping API, managed data services, and cloud hosting for crawlers. It is best known as the maintainer of Scrapy, the widely used open-source Python web-crawling framework, which it has maintained since 2011 when it operated under the name Scrapinghub[2].

Zyte's products are used by e-commerce, market research, and financial data teams to collect product, pricing, news, job posting, real estate, and search engine results data at scale, making it an upstream infrastructure provider for several alternative data workflows[1].

What It Does

The company builds and operates the infrastructure required to extract structured data from websites: request routing with ban handling, headless browser rendering, automatic parsing and extraction, and delivery of cleaned data feeds. Its flagship Zyte API bundles unblocking, rendering, and AI-based extraction in a single service, and the company also offers a fully managed data service in which Zyte builds and maintains client data feeds[1].

Products

Product lines include Zyte API (unblocking, headless browser, AI extraction, SERP results, and an enterprise tier), Scrapy Cloud for running and monitoring Scrapy spiders, and Zyte Managed Data for outsourced feed development[1]. The company publishes data type catalogs for products and e-commerce, news and articles, job postings, real estate, business locations, social media, and search results[1].

Open Source

Zyte maintains Scrapy, a free and open-source web-crawling framework written in Python and released under the BSD license.[2][3] Scrapy originated at London-based e-commerce company Mydeco with contributions from Insophia, and Scrapinghub (renamed Zyte in 2020) became its official maintainer in 2011[2]. The framework is built around self-contained crawlers called spiders and is used by companies including Lyst and Parse.ly[2].

Buyers and Use Cases

Customers cited by the company include ZoomInfo, Yahoo, Sage, Red Bull, and McDonald's[1]. In alternative data contexts, Zyte's services are used to build price tracking, product availability, job postings, and news monitoring datasets, either by funds collecting their own data or by other data vendors outsourcing collection[1].

Compliance

Zyte positions legal compliance as a core differentiator, publishing web data compliance guidelines, maintaining a trust center, and holding ISO 27001 certification[1]. The company participates in the Ethical Web Data Collection Initiative[1].

History

The business traces its origins to 2010, when it began operating as Scrapinghub, and it took the name Zyte in 2020[1][2]. The company states it now handles billions of monthly requests across more than 100 countries and runs the Extract Summit conference and a large web-scraping developer community[1].

Datasets from Zyte (1)