Alternative data datasets
Browse by provider, category, marketplace, access, delivery and license
高质量中文专利摘要数据集。
View source information and access options.
View source information and access options.
View source information and access options.
A U.S. patent collection with a minimal placeholder card and no additional documentation provided.
Continually expanding collection of publicly available statistical and reference datasets sourced globally.
Official Consumer Price Index and inflation series for Iran drawn from two providers: the Statistical Centre of Iran (monthly, from 2011/1390 onward) and the Central Bank of Iran (annual, long historical run from 1936/1315 onward), each with its own series and a combined view covering national-level index and percentage data.
Mobile App Permission Analysis and Classification
Zillow Home Listings in the U.S: Buyable Homes Dataset. Explore U.S. Real Estate
Financial data parsed from 10-Q, 10-Q/A, 10-K, 10-K/A SEC filings from 2010.
Daily updates covering the entire historical record of more than 250,000 Federal Reserve Economic Data series such as interest rates, inflation, economic growth, employment, and credit metrics.
edgar-corpus-embeddings is a text and embedding corpus derived from SEC EDGAR filings, published on Hugging Face by user gagan3012. It packages sectioned 10-K content with pre-computed vector embeddings for machine learning workflows.
Structured datasets on United States federal expenditures, contracting activity and agency disbursements intended for public-sector demand analysis.
View source information and access options.
A sampled extraction of IPO-related tables from SEC filings between 1994 and 2026, delivering raw table HTML with provenance fields and targeting roughly 100 tables per year, with edge years possibly containing fewer valid extractions.
Data Analytics and Data Visualisation for beginners
View source information and access options.
View source information and access options.
Quarterly 10-Q filings parsed from EDGAR into JSON, millions of records in English (MIT). The quarterly companion to the 10-K corpus for tracking sequential changes in operating performance and risk factors.
Single-source schema containing roughly four billion mobile advertising identifiers spanning thirteen Asia-Pacific markets plus the United States, classified under IAB taxonomy version 1.1.
Nationwide United States coverage of approximately two billion mobile advertising identifiers alongside around 910 million household exposure markers, organized into 295 IAB-aligned branded audience categories.
The patent dataset, published by HypernetworkRG, is a small graph dataset consisting of one row in a single training split. It is distributed without a stated license or size and is tagged for text modality.
Intellectual-property datasets that capture patent claims, ownership chains and dispute records to support IP-driven strategies.
View source information and access options.
View source information and access options.
View source information and access options.
View source information and access options.
View source information and access options.
Corporate filing and deadline intelligence (IN-Filings) that monitors reporting calendars, submissions and compliance timelines.
View source information and access options.
View source information and access options.
A cleaned collection of French patent documents published from 1981 through 2026, each row in the Parquet file representing one complete A1 publication parsed from the original XML. It was assembled independently by a contributor with direct access to the public patent source and is optimized for distributed loading.
Raw French patent publications from 2020 through 2026, pulled from original A1 XML records by an independent API/FTP extraction and delivered one-document-per-row in streaming-ready parquet.
Worldwide business-entity and firmographic records consolidating corporate registries, ownership hierarchies and contact details.
Patent analytics offering portfolio insights on innovation trends, assignee behaviour and citation patterns (IPQwery).
A consolidated, ready-to-use repository of 13F institutional holdings drawn from SEC EDGAR filings, capturing the most recent quarterly disclosures of more than 13,000 investment managers including hedge funds, mutual fund complexes, pension funds, banks, and family offices, with each record summarizing total assets under management and individual positions.
Structured line items from SEC 10-K and 10-Q filings between 2009 and 2023, plus company facts — fundamentals extracted straight from EDGAR.
Historical Data from SEC Company Filings
A joint release by Datamule, Teraflop AI, and Eventual provides the SEC-EDGAR dataset, comprising 590 GB of content drawn from all major filings in the SEC EDGAR database and totaling roughly 8 million samples and 43 billion tokens. The data was assembled using the datamule-python library and the official datamule API developed by John Friedman, with datamule serving as a Python package for collecting and manipulating such records.
A collection of 2,000 Malaysian property transaction records spanning every state offers a detailed look at the national housing market in 2025. Compiled from Brickz, a recognized source for real estate transaction data, it covers location, tenure, property type, median pricing, and transaction volumes.
Earnings-call transcripts from an IBM research project, tens of thousands of records with finance tags, in CC0. A filings-adjacent corpus for call-tone and question-and-answer analysis.
10-K filings for S&P 500 companies retrieved from EDGAR, stored as Parquet with the filing text attached (MIT). A focused corpus for the large-cap annual disclosure cycle: risk factors, MD&A and segment detail for every S&P 500 name.
View source information and access options.
View source information and access options.
View source information and access options.
Schema information created by PatentHubLayoutV2 describing JusticeDAO's BM25 indexing resources for patent and legal text. It points to the CFR Title 37 2024 edition as the underlying source and uses the patent-legal-v2.0.0 layout tag.
Layout metadata produced by PatentHubLayoutV2 for a JusticeDAO repository holding patent and legal material. It references the CFR Title 37 annual edition for 2024 and carries the tag patent-legal-v2.0.0, with publication handled separately as an operator decision.
The Patent Legal IR GraphRAG release follows the Publicus retrieval format, providing dense corpus shards, BM25 documents with CID-keyed length and entry metadata, and BM25 postings arranged for inverted-index use.
Repository layout details generated by PatentHubLayoutV2 for JusticeDAO's patent and legal knowledge graph. The metadata references the CFR Title 37 2024 annual edition under a U.S. public-domain license and is tagged patent-legal-v2.0.0.
A vector and layout configuration for patent and legal texts produced via PatentHubLayoutV2, derived from Title 37 of the Code of Federal Regulations annual edition 2024 and marked as U.S. government public domain.
Datamule partnered with Teraflop AI and Eventual to publish the SEC-EDGAR dataset, which covers all major filings from the SEC EDGAR database and amounts to 590 GB across about 8 million samples and 43 billion tokens. Collection relied on the datamule-python library and John Friedman's official datamule API, tools designed in Python for gathering and handling SEC filing data.
Metadata for every EDGAR filing submitted to the U.S. Securities and Exchange Commission from 1994 through December 14, 2024, derived from the quarterly master index files and including CIKs, issuer names, and form types.
sec_filings is a small text dataset of 494 rows that pairs natural-language queries about U.S. securities filings with retrieved document facts and ground-truth answers. It is distributed under an unstated license on a public dataset hub.
Machine-readable extraction of financial and operating figures drawn from publicly filed corporate documents.
Earnings-call transcripts for S&P 500 companies, tens of thousands of calls in English with timestamps, in Parquet (MIT). Direct input for call-tone and question-and-answer sentiment analysis across the large-cap earnings cycle.
The patents-publications-dataset is a large-scale tabular collection of patent publication records distributed in Parquet format. Hosted by publisher labofsahil on the Hugging Face Hub under an unstated license, it is sized in the 100M–1B row range.
Privacy-protected statistics on Florida holders of dormant property interests, broken out by holding size, loan-to-value ratio, portfolio scale, risk profile, and tier.
Information on roughly 95,000 Florida property investors, showing a median equity of $1.6 million, an 11% LTV, and no transactions in the past year.
Segmentation of Florida's 1.96 million real estate investors according to portfolio size, trading activity, equity position, investment approach, and risk exposure.
Approximately 20.8 million US real estate investors can be profiled by state, portfolio size, equity held, and recent transaction activity.