Alternative data datasets
Browse by provider, category, marketplace, access, delivery and license
A collection of Parquet files providing one record per Chinese invention patent family granted between 2003 and 2019, capturing a cumulative quality assessment pipeline that includes metrics such as semantic knowledge recombination, strict historical semantic novelty, and IPC-based robustness measures across global patent publications.
Disambiguated US patent data, inventors, assignees, locations, claims and citation networks, from PatentsView and Patent Examination Research, reaching back to 1976 and including pre-grant publications, all queryable for free.
US patent full text assembled by the Allen Institute from the USPTO, millions of records under ODC-BY: claims, abstracts, descriptions and metadata, in Parquet. A broad base for measuring IP intensity, technology landscapes and the direction of R&D spending.
A large granted-patent text collection from the USPTO, tens of millions of records in Parquet (Apache 2.0). Broad coverage of claims and descriptions for IP-intensity and technology-trend measurement.
The Future of Patent Corpus Management One Document, Fully Decoded - Without Anyone Reading It The signal you cannot get from a search interface The highest-value competitor signal in a patent portfolio is not the invention. It is the amount of money and urgency a company committed to it. That signal is never in an abstract, and no search interface surfaces it. It comes out of this corpus for ever
A repository of roughly 2.7 million publicly available U.S. patent application records split into yearly source files spanning 2021 to 2026, covering application metadata, applicants and inventors, classifications, prosecution history, continuity and priority data, assignments, publications, and grant information where present, sourced from the USPTO.
patent-ate termhood on HUPD This dataset is a ranked phrase table: each row is a multiword (or surface) key, a nested-frequency C-value, and how many filings contain that key (df). patent-ate corpus then score built it from the Harvard USPTO Patent Dataset (HUPD). It is not HUPD itself. Start with the slice (3,393 rows, tens of kilobytes). The full table is 88,607,764 keys from 4,518,254 filings a
Original TIFF drawings and grant full-text XML for 165,917 U.S. design patents omitted from the AI4Patents/IMPACT collection, filling gaps including 161,093 grants from 2023 to 2026 and 4,824 earlier patents missing from IMPACT.
PatentPulse PatentPulse is a provenance-preserving corpus of USPTO grants and published patent applications extracted from official weekly XML bulk releases. This immutable Parquet snapshot normalizes the project's historical append-only JSONL into one schema. Historical partial snapshot. This release was captured on 2026-08-28 while the upstream backfill was still in progress. It is a stable, cit
Derived from USPTO and PatentsView bulk releases, this dataset catalogs 9.1 million U.S. patents from 1976 onward along with complete citation links, assignees, inventors, and CPC codes, with the full-graph edition exceeding 255 million rows.