Diffbot
From AltData.wiki, The Alternative Data Encyclopedia
Diffbot is an American technology company that applies machine learning and computer vision to extract structured data from web pages and to maintain a large structured database of public web content known as the Diffbot Knowledge Graph. The company is headquartered in Menlo Park, California, and was founded by Mike Tung, who serves as chief executive officer.[1][2][3].
Diffbot's products include page extraction and crawling APIs, a natural language processing API, a queryable knowledge graph of entities such as organizations, people, products, and articles, and a self-hostable web search index.[1] The extracted data is used by developers and enterprises, including hedge funds and due-diligence teams, as a source of web-derived alternative data.
What It Does
Diffbot converts unstructured web content into structured data. Its Extract API classifies any submitted URL into a page type using computer-vision models, renders the page, and returns clean structured fields in roughly 300 milliseconds without hand-written rules or per-site scrapers.[1] Its Crawl product walks entire websites from a seed URL and consolidates thousands of pages into a single structured dataset that can be queried or exported.[1]
Data And Methodology
The company's core methodology is the visual parsing of web pages: rather than relying on site-specific markup rules, Diffbot's models visually identify the important elements of a page and return them in a structured format, ignoring elements not central to the primary content.[2] Text sent to the Natural Language API is analyzed for entities, relationships, facts, and sentiment, with each entity resolved against the Diffbot Knowledge Graph.[1]
The Knowledge Graph, first announced in 2015 and released in 2019, is built by crawling the web and applying automatic extraction. According to the company and press coverage, it has grown to include more than two billion entities and ten trillion facts covering corporations, people, articles, products, and discussions.[2]
Products
Diffbot's current product line comprises Extract, Crawl, the Natural Language API, the Knowledge Graph, and Web Search, a packaged web index of roughly four terabytes of news, documentation, homepages, and other content that can be queried via API or self-hosted; the company reports P90 query latency under 300 milliseconds for the hosted service.[1] The company also operates LeadGraph, a separate offering built on its data.[1]
History
Diffbot released its Page Classifier API in August 2012, which automatically categorizes web pages into page types; as part of that launch the company analyzed roughly 750,000 links shared on Twitter and found photos, articles, and videos to be the predominant media types.[2] In May 2012 the company raised $2 million from investors including Andy Bechtolsheim and Sky Dayton.[2]
In June 2015 Diffbot announced work on an automated knowledge graph built by crawling the web,[2] and in 2019 it released the Diffbot Knowledge Graph.[2] In September 2020 the company launched a natural language processing API for building knowledge graphs from text.[2]
Buyers And Use Cases
Diffbot sells API access to developers and enterprises. Customers named in press coverage include Adobe, AOL, Cisco, DuckDuckGo, eBay, Instapaper, Microsoft, Onswipe, and Springpad.[2] On its website the company highlights use cases such as sales research briefs, financial due diligence, open-source intelligence network mapping, and grounding AI agents on structured web facts.[1]