Alternative data datasets
Browse by provider, category, marketplace, access, delivery and license
A large granted-patent text collection from the USPTO, tens of millions of records in Parquet (Apache 2.0). Broad coverage of claims and descriptions for IP-intensity and technology-trend measurement.
About 1.5 million US patent claims split into training and test partitions in CSV, organized for claim-level classification and summarization work (Apache 2.0). A large, ready-split corpus for IP-claims modeling.
The patents-publications-dataset is a large-scale tabular collection of patent publication records distributed in Parquet format. Hosted by publisher labofsahil on the Hugging Face Hub under an unstated license, it is sized in the 100M–1B row range.
US patent full text assembled by the Allen Institute from the USPTO, millions of records under ODC-BY: claims, abstracts, descriptions and metadata, in Parquet. A broad base for measuring IP intensity, technology landscapes and the direction of R&D spending.
Monitor the intellectual property filings that influence pharmaceutical innovation and competition.
A collection of 49,023 patent filings focused on distributed ledger technology is available, comprising roughly 1.296 billion tokens. The corpus is a component of the broader DLT-Corpus initiative and is intended to facilitate natural language processing research, innovation analysis, and patent examination within the distributed ledger field.
Raw French patent publications from 2020 through 2026, pulled from original A1 XML records by an independent API/FTP extraction and delivered one-document-per-row in streaming-ready parquet.
google-patents-data-preview is a community-uploaded preview of Google Patents bibliographic records, distributed in Parquet format through the Hugging Face Hub. The single training split contains roughly 340,000 rows covering patent identifiers, classifications, and localized text fields.
Two Parquet files hold global patent publication records paired with quality indicators for Chinese invention patent families from 2003 to 2019, with each row representing a focal family and carrying cumulative measures such as semantic knowledge recombination, strict historical semantic novelty, and IPC-based robustness.
A read-only archival repository serving as the official worldwide record for 'River of Cognition' technology patents and compliance authorizations, used for public filing and prior-art preservation of complete invention patent documents.
Original TIFF drawings and grant full-text XML for 165,917 U.S. design patents omitted from the AI4Patents/IMPACT collection, filling gaps including 161,093 grants from 2023 to 2026 and 4,824 earlier patents missing from IMPACT.
Vector embeddings of Apple patent text totalling about 4 million entries, stored as parquet shards with a 768-dimensional float64 feature per record.
Disambiguated US patent data, inventors, assignees, locations, claims and citation networks, from PatentsView and Patent Examination Research, reaching back to 1976 and including pre-grant publications, all queryable for free.
Derived from USPTO and PatentsView bulk releases, this dataset catalogs 9.1 million U.S. patents from 1976 onward along with complete citation links, assignees, inventors, and CPC codes, with the full-graph edition exceeding 255 million rows.
Chinese patents paired with their abstractive summaries in Mandarin, a few thousand records in JSON (Apache 2.0). A window into the pace and direction of patenting inside the Chinese technology base.
Public-safe distribution bundle from a local r19 counterexample resolution review of arXiv patents using a language model, containing review-related outputs, manifests, checksums, and provenance receipts, but excluding the original arXiv PDFs and any expert validation.
Research and intelligence covering pharmaceutical products, clinical trials, patent filings and regulatory progress.
DAPFAM patent documents, a curated sample of US patents in Parquet with text for retrieval and classification tasks (CC BY-NC-SA 4.0). Scopus of claim text for building patent-similarity and technology-mapping models.
Intellectual-property datasets that capture patent claims, ownership chains and dispute records to support IP-driven strategies.
Patent analytics offering portfolio insights on innovation trends, assignee behaviour and citation patterns (IPQwery).
The Future of Patent Corpus Management One Document, Fully Decoded - Without Anyone Reading It The signal you cannot get from a search interface The highest-value competitor signal in a patent portfolio is not the invention. It is the amount of money and urgency a company committed to it. That signal is never in an abstract, and no search interface surfaces it. It comes out of this corpus for ever
Machine-readable extraction of financial and operating figures drawn from publicly filed corporate documents.
Patent All Claims EP granted claim sets (indepdendent and dependent claims). One claim per line. Claims from 20210915-20250806 Splits Train: 96% Validation: 2% Test: 2% Usage from datasets import load_dataset ds = load_dataset("mhurhangee/ep-patent-all-claims")
US Patent Descriptions This dataset contains the descriptions of granted US utility patents, filtered and deduplicated.The original data comes from all granted patents in 2025 up to May 20, available from PatentsView. Splits train: 10,000 rows for model training validation: 2,500 rows for validation test: 2,500 rows for evaluation Columns patent_id: Identifier for the patent; useful for reconcilin
A collection featuring the primary independent claim of United States utility patents issued between 2005 and 2025 is provided. Because the opening claim usually establishes the widest legal boundaries of an invention, it carries particular significance, and each entry includes the patent identifier alongside its claim text.
Over 1.3 million granted US patents from the BIGPATENT benchmark, each pairing its claims with an abstractive summary and full description text in English (CC BY 4.0, Parquet). Use it to track the technology areas where applicants are filing and to study how claim language has evolved over time.
Patent-linked intelligence that connects intellectual property activity to research strategy and competitive standing.
patent-ate termhood on HUPD This dataset is a ranked phrase table: each row is a multiword (or surface) key, a nested-frequency C-value, and how many filings contain that key (df). patent-ate corpus then score built it from the Harvard USPTO Patent Dataset (HUPD). It is not HUPD itself. Start with the slice (3,393 rows, tens of kilobytes). The full table is 88,607,764 keys from 4,518,254 filings a
PatentPulse PatentPulse is a provenance-preserving corpus of USPTO grants and published patent applications extracted from official weekly XML bulk releases. This immutable Parquet snapshot normalizes the project's historical append-only JSONL into one schema. Historical partial snapshot. This release was captured on 2026-08-28 while the upstream backfill was still in progress. It is a stable, cit
A collection of Parquet files providing one record per Chinese invention patent family granted between 2003 and 2019, capturing a cumulative quality assessment pipeline that includes metrics such as semantic knowledge recombination, strict historical semantic novelty, and IPC-based robustness measures across global patent publications.