Alternative data datasets
Browse by provider, category, marketplace, access, delivery and license
A Europe-based SAR satellite network capturing Earth imagery at 2–5 meter resolution with once-per-day worldwide coverage, supplied either as raw scenes or as processed analytic layers.
Intellectual-property datasets that capture patent claims, ownership chains and dispute records to support IP-driven strategies.
Managed pipelines that convert public web pages into clean structured feeds at scale, with quality control suited to pricing and catalogue intelligence.
Trade records derived from bills of lading that expose importing and exporting parties, their vendors and specific cargo shipments.
U.S. import shipment information drawn from customs filings that traces the connections between overseas sellers, shippers and U.S. receivers.
Analytics service that converts global news streams into event signals and assigns impact scores to corporate and macroeconomic developments.
Corporate filing and deadline intelligence (IN-Filings) that monitors reporting calendars, submissions and compliance timelines.
View source information and access options.
View source information and access options.
A cleaned collection of French patent documents published from 1981 through 2026, each row in the Parquet file representing one complete A1 publication parsed from the original XML. It was assembled independently by a contributor with direct access to the public patent source and is optimized for distributed loading.
Daily refreshed worldwide monitoring of store-level prices, product mixes and promotional activity across thousands of brands and retail chains.
Worldwide business-entity and firmographic records consolidating corporate registries, ownership hierarchies and contact details.
Independent fixed-income price evaluations and credit benchmark indices published by Intercontinental Exchange.
Global energy balances from the IEA: production, trade, sector consumption, CO2 emissions, EV registrations and clean-tech deployment for 150-plus countries.
Financial content hub offering market schedules, price data, and real-time news updates.
Patent analytics offering portfolio insights on innovation trends, assignee behaviour and citation patterns (IPQwery).
Television and streaming advertising measured from automatic-content-recognition panels: impressions, creative libraries and estimated spend, demand proxies for media companies.
Supply-chain and demand assessments based on on-the-ground retail and distribution checks within Chinese markets.
A consolidated, ready-to-use repository of 13F institutional holdings drawn from SEC EDGAR filings, capturing the most recent quarterly disclosures of more than 13,000 investment managers including hedge funds, mutual fund complexes, pension funds, banks, and family offices, with each record summarizing total assets under management and individual positions.
A joint release by Datamule, Teraflop AI, and Eventual provides the SEC-EDGAR dataset, comprising 590 GB of content drawn from all major filings in the SEC EDGAR database and totaling roughly 8 million samples and 43 billion tokens. The data was assembled using the datamule-python library and the official datamule API developed by John Friedman, with datamule serving as a Python package for collecting and manipulating such records.
Aviation analytics dataset tracing private jet itineraries back to publicly traded corporations, offering visibility into executive travel for event-driven investment analysis.
A collection of 2,000 Malaysian property transaction records spanning every state offers a detailed look at the national housing market in 2025. Compiled from Brickz, a recognized source for real estate transaction data, it covers location, tenure, property type, median pricing, and transaction volumes.
Earnings-call transcripts from an IBM research project, tens of thousands of records with finance tags, in CC0. A filings-adjacent corpus for call-tone and question-and-answer analysis.
10-K filings for S&P 500 companies retrieved from EDGAR, stored as Parquet with the filing text attached (MIT). A focused corpus for the large-cap annual disclosure cycle: risk factors, MD&A and segment detail for every S&P 500 name.
The mintic_linkedin-job-postings dataset aggregates LinkedIn job posting text into a single string column with roughly 124,000 rows. It is published on Hugging Face by user jmparejaz under an unstated license.
S&P 500 Corporate Filings (SEC EDGAR 10-K & 8-K Archive) Overview This dataset contains clean, structured public data exported directly from production runs of Apify actors. It serves as a benchmark and sample for lead qualification, market intelligence, research, and machine learning pipelines. Source Actor: captainhandsome/sec-edgar-filings-search Dataset Page: Public sample and schema Preconfig
Modeling of physical climate threats to physical assets under projected future warming conditions.
Schema information created by PatentHubLayoutV2 describing JusticeDAO's BM25 indexing resources for patent and legal text. It points to the CFR Title 37 2024 edition as the underlying source and uses the patent-legal-v2.0.0 layout tag.
Layout metadata produced by PatentHubLayoutV2 for a JusticeDAO repository holding patent and legal material. It references the CFR Title 37 annual edition for 2024 and carries the tag patent-legal-v2.0.0, with publication handled separately as an operator decision.
The Patent Legal IR GraphRAG release follows the Publicus retrieval format, providing dense corpus shards, BM25 documents with CID-keyed length and entry metadata, and BM25 postings arranged for inverted-index use.
Repository layout details generated by PatentHubLayoutV2 for JusticeDAO's patent and legal knowledge graph. The metadata references the CFR Title 37 2024 annual edition under a U.S. public-domain license and is tagged patent-legal-v2.0.0.
The clinical-trials-xml-2018-2024 dataset on Hugging Face, published by user jvaton, distributes a small Parquet collection of clinical-trial related XML records. It contains fewer than one thousand rows and is licensed on an unstated basis.
Aggregated trade, order book and derivatives feeds from more than one hundred crypto venues for institutional users.
Datamule partnered with Teraflop AI and Eventual to publish the SEC-EDGAR dataset, which covers all major filings from the SEC EDGAR database and amounts to 590 GB across about 8 million samples and 43 billion tokens. Collection relied on the datamule-python library and John Friedman's official datamule API, tools designed in Python for gathering and handling SEC filing data.
Restaurant_Reviews.tsv is a small text corpus of customer restaurant reviews paired with a binary sentiment label. It is published on the Hugging Face Hub by user KarthikaRajagopal under an unstated license.
A satellite-based carbon intelligence platform that gauges industrial production and greenhouse gas output for more than 100 million facilities, supporting ESG analysis and macroeconomic supply surveillance.
Expenditure intelligence on video games and digital entertainment assembled from consumer panels and distribution channel reporting.
sec_filings is a small text dataset of 494 rows that pairs natural-language queries about U.S. securities filings with retrieved document facts and ground-truth answers. It is distributed under an unstated license on a public dataset hub.
**Subtitle:** Hotel Booking & Revenue Analysis Dataset, Insights into Bookings,
The Future of Patent Corpus Management One Document, Fully Decoded - Without Anyone Reading It The signal you cannot get from a search interface The highest-value competitor signal in a patent portfolio is not the invention. It is the amount of money and urgency a company committed to it. That signal is never in an abstract, and no search interface surfaces it. It comes out of this corpus for ever
Earnings Call LLM Insights 📚 Read the Full Story: For a deep dive into the methodology, the wildest moments we found, and key takeaways, check out the blog post:KnowTrend.ai: Auto-Grading Ten Years of Earnings Calls for Prescience and Delusion This dataset contains LLM-generated analysis of ~70,000+ earnings call transcripts. The analysis was performed using Kimi k2-0905-preview, focusing on extra
De-identified patient-level analytics combining insurance claims, laboratory results and clinical encounters into connected care maps.
Kpler energy intelligence: oil, gas and coal trade flows, vessel tracking and infrastructure for commodity trading.
Machine-readable extraction of financial and operating figures drawn from publicly filed corporate documents.
Earnings-call transcripts for S&P 500 companies, tens of thousands of calls in English with timestamps, in Parquet (MIT). Direct input for call-tone and question-and-answer sentiment analysis across the large-cap earnings cycle.
The patents-publications-dataset is a large-scale tabular collection of patent publication records distributed in Parquet format. Hosted by publisher labofsahil on the Hugging Face Hub under an unstated license, it is sized in the 100M–1B row range.
EdgarItem7 is a text dataset derived from SEC filings, packaging Item 6 (Selected Financial Data) and Item 7 (Management's Discussion and Analysis) excerpts alongside filing metadata. It is hosted by publisher lealOO and distributed in Arrow format under an unstated license.
Self-hosted text analysis toolkit that detects sentiment, extracts entities, and categorizes themes within a customer's own infrastructure.
Enterprise-grade datasets spanning lawsuits, regulatory actions, real-estate records, corporate information and identity risk for due-diligence workflows.
Brand- and ticker-level consumer mood indicators extracted from publicly available social media posts.
A cleaned, structured repository of job listings from multiple boards, with duplicate removal and records of when posts open and close.
Survey of technology deployment at more than five thousand colleges and universities globally, tracking learning management, student information and enterprise resource planning platforms.
Signals on corporate workforce shifts, including new hiring, employee departures and internal reorganizations.
Automated news feed with sentiment ratings, delivered in machine-readable form via the LSEG (formerly Refinitiv) distribution network.
Visual quantitative research workspace that blends sentiment, social, and web-traffic feeds into customizable backtesting strategies.
Near-real-time consumer expenditure analytics sourced from several payment processor collaborations.
Wholesale pricing index and transaction records from Manheim's used-car auctions that monitor dealer-level inflation.
A corpus of around 1.3 million U.S. patent filings paired with human-authored abstractive summaries, organized into nine Cooperative Patent Classification categories ranging from human necessities to textiles and paper.
Near-real-time AIS-based maritime tracking covering global vessel locations, port visits and overall fleet movements.
A live feed of vehicle inventory from tens of thousands of US and Canadian dealers, with VINs, trims, days on lot and pricing, near-real-time visibility into automotive supply and demand.