Alternative data datasets
Browse by provider, category, marketplace, access, delivery and license
Vector embeddings of Apple patent text totalling about 4 million entries, stored as parquet shards with a 768-dimensional float64 feature per record.
An interview-based network used to recruit industry professionals and conduct bespoke expert research engagements.
Alternative data sourced from China tracking online retail pricing and transaction levels, mobile app leaderboards, and distribution-channel observations to support consumer-sector investment views.
Hyperspectral and multispectral Earth observation analytics applied to farming and environmental assessment.
Financial behavior signals from emerging-market lending activity in developing countries.
Research on employment trends, detailing workforce demand, required competencies and pay levels by region.
Market data for the recruitment sector, monitoring the volume of job advertisements and the skills employers seek.
Countrywide mobile telemetry from China quantifying device penetration and in-app activity levels.
EDGAR M&A Deal Events (2010–2024) 4,156 U.S. public-company acquisition events, discovered directly from SEC EDGAR's own quarterly filing indexes — not scraped from a vendor list or a blog. Each row is a company that filed a merger proxy or tender-offer response between 2010 and 2024, with the earliest such filing's date as an announcement-date proxy. Built as a byproduct of an M&A target-predicti
An expert research platform, now part of LSEG, offering on-demand interviews and institutional research.
Text of SEC filings pulled from EDGAR spanning millions of records in English (Apache 2.0): 10-K, 10-Q, 8-K, annual reports and exhibits, tagged for text-generation and classification. A practical base corpus for earnings- and filings-driven signals, from risk factors and MD&A to management commentary.
Street-level signal quality data for wireless carriers collected from crowdsourced mobile device measurements.
Get a data sample: Consensus-verified big-box retail POIs including IKEA, Carrefour, Tesco, Walmart, and Target. 100% NAICS coverage at the first 3 levels. Weekly incremental updates.
Streaming market intelligence drawn from digital-asset news coverage and online sentiment channels.
PatentPulse PatentPulse is a provenance-preserving corpus of USPTO grants and published patent applications extracted from official weekly XML bulk releases. This immutable Parquet snapshot normalizes the project's historical append-only JSONL into one schema. Historical partial snapshot. This release was captured on 2026-08-28 while the upstream backfill was still in progress. It is a stable, cit
A machine learning service evaluates online material to assign sentiment and emotional tone scores.
A service that connects clients with practitioners worldwide, offering recorded interview archives and studies that draw on multiple specialists.
Indian Domestic Airline Flights 2018-2025 (TsFile) Apache TsFile version of Gokul99400/IndianDomesticAirlineDataset. Overview Indian domestic airline schedule records covering 2018-2025 across ~100 airports: airline, flight number, route, operating days of week, scheduled departure/arrival times and the schedule validity window (validFrom..validTo). The repo ships two CSVs: Air-Clean.csv (33,734 d
Short-range forecasting service offering hyperlocal predictions, real-time nowcasts, and a proprietary method of combining multiple sensor inputs.
Venture financing intelligence covering funding rounds, industry sectors, and developing regional markets.
Vehicle passage records from tolling and traffic management infrastructure on paid roadways.
An international dataset cataloging employment announcements alongside facility openings, expansions and shutdowns.
A worldwide research resource built on de-identified records from more than 130 million patients to enable cohort identification.
Research on private technology firms including valuation estimates and operational performance indicators.
Rankings of Chinese applications and user engagement benchmarks compiled from extensive device samples.
A platform pooling electronic health record information from U.S. healthcare providers to support real-world evidence studies.
Governance, beneficial-ownership and related-party transaction intelligence (Tussell) used for institutional investment research.
Broadcast television and radio are monitored using logo and object detection to tally minutes of brand visibility.
Curated datasets produced by UBS Evidence Lab through its in-house research across numerous alternative-data subject areas.
Machine-learning and imaging analytics applied to Earth observation data for environmental and infrastructure surveillance.
Mobility scores describing movement around venues and neighbourhoods, used in real-estate and retail analysis and including migration flows and travel-distance statistics.
Surface weather station network measuring rainfall and wind conditions at neighborhood resolution.
Provider of commercial synthetic aperture radar imagery and derived geospatial analytics aimed at surveillance and situational awareness use cases.
The full register of federal contracts, grants and loans, recipient names, NAICS codes, place of performance and dollar amounts, downloadable or API-accessible, mapping government revenue exposure down to the district.
Application programming interfaces that standardize customer account and consumption information from utility providers.
Survey responses from 1,037 technology and operations executives managing software budgets between $500,000 and $10 million.
A knowledge graph assembled from 575,778 ClinicalTrials.gov registrations and their associated arms, outcomes, sites, sponsors, conditions, interventions, MeSH codes, and PubMed citations, containing 7,628,735 nodes and 15,531,427 edges, delivered as 623 MB of Parquet files (compared to about 7 GB of raw JSON) and built using Samyama Graph.
Suite of orbital sensing offerings spanning spectral and radar modalities along with analytics for agriculture, infrastructure security and defense-oriented monitoring.
The restaurant_review_sentiment dataset is a small text corpus pairing short restaurant reviews with binary sentiment labels and restaurant and user identifiers. It is hosted on the Hugging Face Hub under an unstated license and has been downloaded only a handful of times.
Veeva CRM data capturing how life-sciences sales teams interact with healthcare professionals in the field.
A clean, scalable feed of raw device pings for teams that want to build their own visitation and mobility metrics from the ground up.
Worldwide archive of automatic identification system transmissions containing past ship locations, port visits, and fleet attributes for millions of vessels.
VesselsValue database supplying ship valuations, age records, and operating metrics across the worldwide fleet to support asset and lending decisions.
Vortexa dataset recording energy commodity shipments and tanker voyages, including port operations and oil product flows.
Trading calendars populated by detected corporate events including investor relations conferences and public filing dates.
A 10% sample of over 573,000 ClinicalTrials.gov studies augmented with AI-classified therapeutic areas, sponsor categorization, outcome groupings, and duration metrics, offered for portfolio and market landscape analysis.
A 10% sample of a sponsor profiling resource containing entries for over 10,200 clinical trial sponsors, covering trial counts, completion ratios, phase distribution, therapeutic areas, and pipeline activity.
Production and delivery statistics tracking passenger and commercial vehicle output from global automakers.
End-to-end data pipelines delivering normalized JSON streams drawn from news outlets, weblogs, discussion boards, consumer review platforms, and darknet marketplaces.
View source information and access options.
Windward platform delivering maritime risk assessments, sanctions screening, and trade monitoring tools designed to identify illicit shipping networks.
Analytical insights on the shift to cleaner power, covering renewable energy, electricity markets and commodity research.
Historical records of dividend distributions listed on stock exchanges around the world.
The linkedin-job-postings dataset, published by xanderios on Hugging Face, contains a single CSV table of U.S. LinkedIn job listings with about 33,000 rows. It is tagged with a MIT license on the platform, though the underlying LinkedIn terms of service place restrictions on redistribution and use.
Enriched entity records and linking attributes that map corporations, individuals and locations together to power graph-style analysis.
Xeneta rate index reporting negotiated and market-based container shipping prices across international trade corridors.
Minute-level open-high-low-close-volume bars for Indian equities, hundreds of millions of rows in Parquet under MIT. Gives high-frequency microstructure across the Indian cash market for intraday and market-on-close strategies.
patent-spec-xml is a dataset of patent specification documents in XML format, published by Yehoon and licensed under unspecified terms. It is catalogued on the alternative-data directory as a public-records resource.
Investment-ready consumer research assembled from scraped email receipts, transaction panels and app-level data across internet-driven consumer sectors.
A worldwide consumer panel of more than 2.5 million adults used to generate survey, brand-tracking and market-research outputs.