Alternative data datasets
Browse by provider, category, marketplace, access, delivery and license
Long-running series of news, social, and ESG sentiment measures distributed through LSEG channels.
Spatial analytics covering construction sites, mining operations and port logistics, produced from optical and synthetic aperture radar imagery and packaged as organized data feeds.
An opt-in email receipt panel that turns online purchase confirmations into estimated merchant-level sales figures.
Media intelligence solution that processes several million articles each day to track brand exposure and share of voice.
Crypto asset and protocol performance datasets paired with curated research.
Standardized financial datasets that clean and structure company financial statements from filings for downstream analytics.
Patent All Claims EP granted claim sets (indepdendent and dependent claims). One claim per line. Claims from 20210915-20250806 Splits Train: 96% Validation: 2% Test: 2% Usage from datasets import load_dataset ds = load_dataset("mhurhangee/ep-patent-all-claims")
A collection featuring the primary independent claim of United States utility patents issued between 2005 and 2025 is provided. Because the opening claim usually establishes the widest legal boundaries of an invention, it carries particular significance, and each entry includes the patent identifier alongside its claim text.
Entertainment industry datasets, music-streaming subscriber counts, video, gaming and creator-economy metrics, assembled from disclosures and web measurement.
AlphaAI's near-real-time parse of SEC EDGAR Form 4 filings, providing 20,832 individual insider transaction tranches and 8,538 grouped economic events for US public-company officers, directors, and 10% owners.
german-job-postings is a publicly hosted corpus of roughly 70,584 normalized German-language vacancies drawn from the Bundesagentur für Arbeit Jobbörse API, each row tagged with the KldB-2010 occupation code and machine-assigned ESCO occupation and skill labels. The dataset is distributed under CC-BY-4.0 and is intended as an open alternative to the predominantly English/US job-posting resources available to machine-learning practitioners.
Sales intelligence for the advertising technology sector, monitoring ad creative distribution, estimated expenditure and publisher connections.
Specialized measurements track corporate media presence, flagging atypical surges in coverage and quantifying alignment across sources.
Visitor presence and behaviour captured through venue WiFi and beacons, measuring foot counts and dwell inside malls and big-box retail spaces.
Credit analytics for structured products and residential mortgages are shared via Moody's partnership channels.
A worldwide banking directory covering financial institutions across all nations, available as a complimentary sample and through a complete paid license.
A repository of roughly 2.7 million publicly available U.S. patent application records split into yearly source files spanning 2021 to 2026, covering application metadata, applicants and inventors, classifications, prosecution history, continuity and priority data, assignments, publications, and grant information where present, sourced from the USPTO.
View source information and access options.
Scalable web harvesting service that delivers scheduled, structured data feeds ready for direct incorporation into business intelligence environments.
FinSight gathers plain-text 10-K and 10-Q SEC filings for twenty major US public firms spanning six sectors, yielding 97 records intended for BERT fine-tuning, retrieval-augmented generation, and multi-agent financial research.
Designed to aid Nifty50 trading research, this dataset merges financial news, sentiment indicators, and stock market data into a single analysis-ready resource, enabling pipelines that cover news cleaning, impact categorization, FinBERT-based sentiment scoring, and Temporal Fusion Transformer forecasting.
One marketplace where dozens of vetted location datasets can be evaluated under standard schemas with no-minimum commitments: a fast route to comparing mobility feeds.
A distribution platform, previously known as Quandl, delivers both alternative and traditional financial datasets.
A queryable DuckDB mirror of the AACT flat-file export of ClinicalTrials.gov, packaged via the clinicaltrials-database project and covering every registered trial across 48 tables totaling roughly 58 million rows.
Research marketplace for locating subject authorities and commissioning first-hand studies from practitioners.
Targeted research and datasets focused on niche sectors, private firms and specialised investment topics.
Global Job Postings – Multi-ATS is a single dated snapshot of 112,816 job listings aggregated from 23 applicant tracking systems and parsed into a 47-column schema by NextGig-Rocks. It is distributed as Parquet via the Hugging Face datasets library and targets labor-market research, not production job boards.
Japanese consumer purchase data sourced via Nikkei for retail consumption modeling.
Business registry entries covering Cyprus, Guernsey, Liechtenstein, Luxembourg, Malta, and Monaco.
Over 1.3 million granted US patents from the BIGPATENT benchmark, each pairing its claims with an abstractive summary and full description text in English (CC BY 4.0, Parquet). Use it to track the technology areas where applicants are filing and to study how claim language has evolved over time.
International credit records are converted into formats usable by lenders assessing overseas applicants.
A household panel that records purchase receipts and consumer shopping patterns for measuring retail performance.
Crowd-sourced mobile speed and connectivity test results reflecting user experience across networks.
TUDelft-Electricity-Consumption-1.0 is a high-resolution, open-source time series dataset of household electricity load published on Hugging Face by OpenSynth. It draws on multi-country trials covering thousands of households under different tariff regimes and is distributed in Parquet format.
Automated satellite-image analysis that quantifies construction, industrial-site activity, mining and energy infrastructure.
Commercial scraping interfaces for online retail, general websites, and search engine results, returning organized public-web records backed by worldwide proxy networks.
A CC-BY-4.0 licensed compilation from Ozari Health cataloging 22 major published GLP-1 clinical trials as of May 2026, consolidating peer-reviewed trial-level data into a single structured index.
The restaurant_reviews dataset is a small text corpus of roughly one thousand customer review snippets paired with binary sentiment labels, published on the Hugging Face Hub under the user account pachequinho. It is distributed in CSV format with an Apache-2.0 license declaration.
ClinicalTrialSummary is a Hugging Face dataset published under the pat-jj account that pairs long-form clinical trial article text with shorter plain-language summaries. It is distributed in Parquet format with predefined train, validation, and test splits totaling 77,516 rows.
Intelligence on digital advertising expenditure, including creative delivery and approximate channel-level budget allocations.
Patent-linked intelligence that connects intellectual property activity to research strategy and competitive standing.
Payments-focused datasets and analysis tracking transaction volumes, processing networks and overall market composition.
Penumbra AI SEC 10-K Risk Factor Disclosures (S&P 500, FY2025-2026) Dataset Summary This dataset contains a sample of 2,683 individual risk factor disclosures from the most recent Form 10-K filing of 50 S&P 500 companies (fiscal year ends ranging from May 2025 to February 2026). Each record is a single risk driver — one discrete point a company disclosed under "Risk Factors" — together with the ve
A developer interface supplying individual and organizational profiles, including work history, for talent sourcing and company data enrichment.
A daily record of trading activity across multiple major global stock market indices, structured for trend analysis and forecasting work in international equity markets.
Patent_FR_US_Merge_Radix_65536 is a Hugging Face dataset by publisher PhysiQuanty that merges French and US patent documents into a single text corpus. It is distributed in Parquet format and is tagged with an Apache-2.0 license, though the README provides no further description.
A data engineering platform that consolidates commerce, shipping and supplier information into standardized pipelines for supply-chain use.
Private equity and venture capital dataset with information on deals, funds, and valuations across private-market transactions.
Analytics product that links weather conditions to changes in category-level and regional sales demand.
Sub-metre optical images collected daily by the largest constellation of small imaging satellites, with analytics ready for detecting land change.
The ja_patent mirror of llm-jp-corpus-v4, assembled by the LLM-jp Corpus Building Working Group at NII, reproduces the Japanese patent sub-corpus from the larger LLM-jp Corpus v4 release. It is delivered as 621 jsonl.gz files totaling 58.2 GB in compressed form, with each line holding a JSON object that includes a text field and a meta field containing the document identifier, URL, and other provenance information.
Real-world observations gathered by a distributed workforce running structured micro-surveys: price checks, shelf-availability readings and local market conditions.
Logistics monitoring service that follows billions of freight movements, offering predicted arrival times, delay alerts, and carrier performance metrics across transport modes.
Consultancy arranging conversations between seasoned practitioners and clients in technology, finance and consumer industries.
Curated data and analytics products covering financial services, payments ecosystems and technology sectors.
Airline and online travel agency fares collected to track travel demand and ticket pricing trends.
Analytics derived from space assets that merge image and signal intelligence to track defense, infrastructure and commercial activity.
patent-ate termhood on HUPD This dataset is a ranked phrase table: each row is a multiword (or surface) key, a nested-frequency C-value, and how many filings contain that key (df). patent-ate corpus then score built it from the Harvard USPTO Patent Dataset (HUPD). It is not HUPD itself. Start with the slice (3,393 rows, tens of kilobytes). The full table is 88,607,764 keys from 4,518,254 filings a
Consumer panels in China quantifying monthly active users, usage duration and behavior across multiple applications.
Provider-submitted clinical performance measures consolidated into outcomes-focused datasets.