Alternative data datasets
Browse by provider, category, marketplace, access, delivery and license
The ClinicalTrialSummary_Full resource lacks further documentation, with no additional details supplied.
View source information and access options.
View source information and access options.
Penumbra AI SEC 10-K Risk Factor Disclosures (S&P 500, FY2025-2026) Dataset Summary This dataset contains a sample of 2,683 individual risk factor disclosures from the most recent Form 10-K filing of 50 S&P 500 companies (fiscal year ends ranging from May 2025 to February 2026). Each record is a single risk driver — one discrete point a company disclosed under "Risk Factors" — together with the ve
A daily record of trading activity across multiple major global stock market indices, structured for trend analysis and forecasting work in international equity markets.
The patent_classification dataset is a small text corpus hosted on the Hugging Face Hub by publisher pgurazada1, comprising paired patent text snippets and their cooperative classification labels. It is distributed as a CSV file under an unstated license.
Patent_FR_US_Merge_Radix_65536 is a Hugging Face dataset by publisher PhysiQuanty that merges French and US patent documents into a single text corpus. It is distributed in Parquet format and is tagged with an Apache-2.0 license, though the README provides no further description.
The ja_patent mirror of llm-jp-corpus-v4, assembled by the LLM-jp Corpus Building Working Group at NII, reproduces the Japanese patent sub-corpus from the larger LLM-jp Corpus v4 release. It is delivered as 621 jsonl.gz files totaling 58.2 GB in compressed form, with each line holding a JSON object that includes a text field and a meta field containing the document identifier, URL, and other provenance information.
Demographic profiles covering more than 300 million individuals and households.
Complete textual SEC 10-K annual filings from 2021, suited for studying corporate disclosures and natural language processing tasks.
Extracted 10k fillings from SEC Edgar for the fiscal year 2021
View source information and access options.
View source information and access options.
A read-only archival repository serving as the official worldwide record for 'River of Cognition' technology patents and compliance authorizations, used for public filing and prior-art preservation of complete invention patent documents.
Dataset Card for "zephyr_pi0_gen_57k_for_offline_dpo_ipo" More Information needed
ClinicalTrials.gov scraping tool that extracts study records based on criteria like condition, sponsor, phase, status, or geography, yielding one row per NCT identifier containing fields such as sponsors, phases, enrollment, outcomes, and locations, with no API key needed and approximately 11,191 rows and 63 fields collected over 65 runs.
Search filings on SEC EDGAR using ticker symbols, CIK identifiers, SIC industry codes, or free-text queries, returning a single structured row per filing that includes company classification, period dates, direct links to source documents, and optional XBRL financial data. The dataset contains 2,691 rows across 27 fields, supported by 64 collector runs, with the latest observation dated 2026-08-04, and covers 771 entities, browsable at https://reapx.dev/data/sec-edgar-scraper/.
View source information and access options.
View source information and access options.
View source information and access options.
View source information and access options.
clinicaltrials.gov-summary_and_eligibility is a small Parquet mirror of selected ClinicalTrials.gov records, distributed on the Hugging Face Hub under the rjac publisher. It packages study identifiers, recruitment status, titles, brief summaries, and eligibility criteria into a single tidy table of roughly three thousand rows.
The datasets contain real estate data by Properati Data
**First end-to-end UAVs-security corpus
An English-language NLP corpus of cleaned earnings call transcripts gathered from public investor-relations pages, sized between 10,000 and 100,000 documents for financial analysis and LLM work.
A structured compilation of 29,633 completed Phase 3 clinical trials sourced from ClinicalTrials.gov, including status and design details for each study record.
Data from SEC form 4, USA equities, Insider Trades
View source information and access options.
Procurement and spend records aggregated from public sources across countries, for supplier analysis and due-diligence.
Directory of more than 136,000 K-12 institutions across the United States, both public and private, including enrollment breakdowns and institutional attributes.
Records covering over 136,000 American public and private K-12 schools, featuring demographic details alongside SchoolDigger performance rankings.
Compilation of 136,000+ U.S. public and private K-12 school entries with street addresses, geographic coordinates, and descriptive attributes.
A directory of U.S. schools covering a five-year period.
A nationwide U.S. school directory available for a single year.
A directory of U.S. schools covering a ten-year period.
Clean, normalized daily prices with insider trading filings
Demographic data on population, income, education, households, and consumer behavior at the block-group level, delivered through Snowflake.
Enrichment attributes covering population, earnings, household composition, consumer characteristics, and geographic context at the ZIP code level for use within Snowflake analytics.
A resource built for forecasting patent approval outcomes, introduced in research employing domain-specific fine-grained claim dependency graphs, with further documentation to be supplied later.
Keyword frequencies, LDA topics & store counts from SEC EDGAR (FY1996-2025)
View source information and access options.
View source information and access options.
Real Estate Market Trends & Pricing Analysis (India 2025)
sec-filings is a Hugging Face dataset published by Sicheng-Chroma that distributes pre-chunked U.S. SEC filings paired with dense vector embeddings. It is distributed in optimized Parquet format and is intended for retrieval and similarity research over regulatory disclosures.
Original TIFF drawings and grant full-text XML for 165,917 U.S. design patents omitted from the AI4Patents/IMPACT collection, filling gaps including 161,093 grants from 2023 to 2026 and 4,824 earlier patents missing from IMPACT.
Complete records from the Bureau of Labor Statistics National Compensation Survey along with all previously issued releases.
Yearly IRS Form 990 submissions for American nonprofit entities, beginning with the 2019 tax year.
Earnings Call Transcript Lite is a small public-records dataset of corporate earnings call transcripts paired with short reference summaries. It is distributed as a CSV and is published on the Hugging Face Hub under an unstated license.
View source information and access options.
Space-sector database (Spacelist) tracking launch activity, satellite fleets, operators and the structure of the commercial space market.
Dataset for examining recent movers and acquiring new customers.
Residential profile data providing fresh analytical perspectives.
View source information and access options.
View source information and access options.
View source information and access options.
Vector embeddings of Apple patent text totalling about 4 million entries, stored as parquet shards with a 768-dimensional float64 feature per record.
Parquet and csv files associating CIK (Central Index Key) with filing entities
JSON data file for SEC conformed company name, CIK, ticker, exchange association
Uncover zero-day exploits, lateral moves, and AI-driven attack blueprints
The patent-phrase-similarity dataset is a text corpus that pairs patent terminology for similarity scoring. It is distributed by tasksource and contains roughly 48,548 rows across train, validation and test splits.