Alternative data datasets
Browse by provider, category, marketplace, access, delivery and license
View source information and access options.
Curated training corpus with an AI Readiness Score of 84 out of 100, positioned for regulated industry model development.
Twenty-nine opt-in signals covering borrowing, lending, and credit-related consumer behavior across an estimated 44.8 million adults in the United States.
Data drawn from decentralized exchanges, automated market makers, and lending protocols.
Agentic Ai Lex Chatbot AI-powered Financial Chatbot integrated with an AWS Agent System to streamline credit card customer support.
Agentic financial research across licensed data from S&P Global, QUODD, FRED, and SEC. Runs multi-step search, reconciles conflicting sources, and cross-verifies findings to return accurate, cited answers with every figure traced to its primary source.
View source information and access options.
Macroeconomic indicators from the EU, including GDP measurements, residential property prices, public sector accounts, monetary rates, balance of payments, purchasing power parities, and the harmonized consumer price index.
Financial sector data and macroeconomic statistical releases standardized under reference 71.
Researcher-driven sustainability ratings spanning more than 90% of the leveraged finance markets in Europe and the United States.
Mobile Advertising IDs are classified by behavioral attributes and location visits relevant to business and financial services.
View source information and access options.
Information sourced from decentralized exchanges, lending platforms, derivatives venues, and broader on-chain blockchain records.
A compilation of 150 restaurant reviews for casual and fine dining venues, drawn from the publicly available DineScope project under a CC0 1.0 license, with a production-tier usage designation and a 2026 sync date.
View source information and access options.
View source information and access options.
Datamule, Teraflop AI, and Eventual jointly contributed a 590 gigabyte SEC EDGAR collection encompassing about 8 million filings and 43 billion tokens, assembled using the datamule-python library and the official datamule API.
A sharded Parquet archive of NSE equities and index prices from India spanning 2000 to 2026, covering more than 2,500 tickers and organized into roughly 1.5 GB files for efficient streaming.
Derived Financial Ratios from 503 Public Company Trial Balances in Turkey
A free forward catalyst calendar for clinical-stage biotech assets
15 years of net shareholder capital flow adjusted for executive stock dilution
S&P 500 Institutional Derivatives: Vectorized Black-Scholes & Z-Score Analysis
Dataset for Counterfeit Detection and Sales Pattern Analysis
World top economic indicators
7,079 Trading Days | 6 Splits | IPO to $4.49 Trillion
A static tabular extension of the Bose345/sp500_earnings_transcripts collection covering the same 2005 to 2025 calendar of S&P 500 earnings events, where each row represents a single company-quarter call identified by a stable episode_id and bundles the full transcript along with related SEC press materials and pre-event inputs suitable for supervised learning or reinforcement-style experimentation.
Quarterly institutional filings from 2015 - 2017
SEC insider transactions 2006-2026 with stock & insider metadata
An open, layout-preserving conversion of U.S. SEC EDGAR filings into MultiMarkdown format covers approximately 3.4 million documents submitted between January 2022 and June 2025, designed for long-context language modeling, financial reasoning, and document analysis.
Unifi Value Frameworks PDF Lifting Competition
Predict gold price from basics to bleeding edge techniques
A curated dataset of companies listed on the London Stock Exchange
Quarterly Earnings Call Transcripts for 10 NASDAQ companies from 2016-2020
This solution identifies finance related tables and extracts information from financial statements.
Gender, Location, and Transaction Trends
A large-scale release produced jointly by Datamule, Teraflop AI, and Eventual covering roughly 590 GB of SEC EDGAR content. It comprises about eight million filings totaling 43 billion tokens, gathered using the datamule-python package and the official datamule API.
A leak-free, chronologically ordered corpus of 8-K, 10-Q, and 10-K filings for roughly 610 present-day U.S. large-cap companies, each paired with objective forward-return labels through 2026. The release is structured as a benchmark for testing whether filing language can anticipate subsequent stock performance, with clearly separated provenance classes.
A comprehensive dataset tracking macroeconomic shifts, venture funding cycles.
42,302 establishments across 714 US public companies. OSHA, WHD, NLRB, EPA,etc
BigQuery dataset of all SEC filings
World Bank collection of global development indicators (BigQuery)
Public-sector finances — tax receipts, budget execution and fiscal indicators — across multiple jurisdictions.
S&P 500 earnings-call transcripts in Parquet under MIT, tens of thousands of calls. Independently compiled and comparable to other public transcript corpora for sentiment and tone work.
Daily liquidations, ETF flows by fund, funding, OI, options, COT, premium
A reconstructed EDGAR corpus of company filings in English (Apache 2.0): a large sample of annual and quarterly reports and exhibits, hundreds of thousands of records. Comparable to other public SEC text sets and a solid starting point for filings-based research.
Includes all company names and cik keys from the SEC database.
Drawn from SEC EDGAR's full archive of millions of filings spanning all form types and US public companies, this sample isolates 1,000 recent 8-K material event submissions together with their filing metadata and document references.
Millions of consumer complaints about financial products — mortgages, credit cards, debt collection and loans — with product, issue, company response and ZIP-level geography, together with HMDA mortgage disclosures.
A collection of earnings call transcripts covering S&P 500 firms and other U.S. large-cap companies across the years 2005 through 2025, intended for financial analysis, NLP modeling, and sentiment research.
Extracted from SEC EDGAR DEF 14A proxy filings since 2015, the dataset contains over 500,000 structured executive compensation entries for S&P 500 and Russell 2000 issuers, capturing CEO and CFO pay, equity awards, bonuses, incentives, and totals keyed by CIK, ticker, and fiscal year.
List of CMBS deals with publicly available Schedule AL data in SEC.gov EDGAR
Records of initial public offerings launched on Indian markets between 2006 and 2025, with fields covering open and close dates, listing date, face value, issue price and size, lot size, first-day price, total shares offered and their allocation across anchor, NII, QIB and retail categories, minimum investment, and subscription figures for each investor class.
Could corporate environmental impacts be integrated into fiancial accounting?
Suitable for modelling airline performance post-COVID
The searchable index behind US open data, listing roughly 300,000 federal, state and local datasets across agriculture, climate, finance, health and safety, each pointing to its source agency via metadata APIs.
All 8-k filings for 2024 with augmented fields.
Insider Trading of Yahoo
DerivedFunction01's sec-filings-snippets is a public-records dataset of short text excerpts drawn from U.S. SEC filings, released on Hugging Face in Parquet format. It is sized for language-model fill-mask work and is not positioned as a commercial alternative-data product.
S&P 500 companies historical prices with fundamental data
Location and Transaction Trends