Alternative data datasets
Browse by provider, category, marketplace, access, delivery and license
EdgarItem7 is a text dataset derived from SEC filings, packaging Item 6 (Selected Financial Data) and Item 7 (Management's Discussion and Analysis) excerpts alongside filing metadata. It is hosted by publisher lealOO and distributed in Arrow format under an unstated license.
View source information and access options.
View source information and access options.
View source information and access options.
View source information and access options.
View source information and access options.
Enterprise-grade datasets spanning lawsuits, regulatory actions, real-estate records, corporate information and identity risk for due-diligence workflows.
Dataset Card for 中華民國專利技術名詞中英對照詞庫 中華民國專利技術名詞中英對照詞庫(Chinese-English Technical Patent Glossary)收錄逾 324 萬筆台灣專利技術名詞之中英對照資料,涵蓋國際專利分類(IPC)A 至 H 全部八大類,時間跨度自 2011 年至 2023 年。本資料集適用於專利翻譯、技術術語標準化、以及繁體中文語言模型在專業領域之詞彙增強。 Dataset Details Dataset Description 本資料集整理自中華民國經濟部智慧財產局(TIPO)公開之專利技術名詞中英對照詞庫。每筆資料包含一組繁體中文與英文之技術術語對照,並標註其對應的國際專利分類(IPC)代碼與資料來源編號。 資料涵蓋 IPC 八大類別: A — 人類生活需要(Human Necessities) B — 作業、運輸(Performin
A consolidated JSONL file of 1,406 Traditional Chinese and English intellectual property term pairs sourced from Taiwan's TIPO website, covering patents, trademarks, trade secrets, designs, and copyrights, suitable for IP-related translation or bilingual pretraining, with original TIPO translations preserved.
An API-based ClinicalTrials.gov scraper that retrieves over 585,000 study records including phase, sponsor, conditions, interventions, enrollment, eligibility, and sites, filterable by condition, sponsor, and status.
Metadata for hundreds of thousands of clinical trials in English, French and Spanish, structured in Parquet with study-level fields such as phase, condition, sponsor and status (Apache 2.0). Trial starts and phase transitions lead clinical development, so this works as an early indicator of biotech pipeline activity.
Minimal information is provided on this clinical-trials-embedded dataset, with additional details needed.
A Comprehensive Dataset on Real Estate Asking Prices and Property Features
Unlocking Real Estate Insights: Analyzing, Visualizing, and Predict with ML
A corpus of around 1.3 million U.S. patent filings paired with human-authored abstractive summaries, organized into nine Cooperative Patent Classification categories ranging from human necessities to textiles and paper.
The clinical_trials dataset distributes ClinicalTrials.gov registry records in Parquet format for offline analysis. It contains a single training split of roughly 150,000 rows sourced from the public clinical-trial registry.
patent-desc-test is a small text dataset on Hugging Face Datasets that contains short identifier-style strings derived from patent document numbers. The dataset is published by user matthias-ehrlich and is broadly tagged as public records–adjacent text data.
An analytical project examining the Paris Housing Dataset to uncover which property attributes most strongly affect asking prices. It involves cleaning, visualizing, and statistically modeling the records to answer targeted research questions.
A compact sample set of SEC filings used for the MemGPT agent demos, tens of thousands of text records. A small, clean entry point to filings text rather than a comprehensive corpus.
2.8K+ venture funding rounds across 35 countries & 20 sectors with investor data
Standardized financial datasets that clean and structure company financial statements from filings for downstream analytics.
Number and valuation of new housing units authorized by building permits
US Patent Descriptions This dataset contains the descriptions of granted US utility patents, filtered and deduplicated.The original data comes from all granted patents in 2025 up to May 20, available from PatentsView. Splits train: 10,000 rows for model training validation: 2,500 rows for validation test: 2,500 rows for evaluation Columns patent_id: Identifier for the patent; useful for reconcilin
A collection featuring the primary independent claim of United States utility patents issued between 2005 and 2025 is provided. Because the opening claim usually establishes the widest legal boundaries of an invention, it carries particular significance, and each entry includes the patent identifier alongside its claim text.
A benchmark collection designed to test vision-language models on the visual conventions found in U.S. design patent illustrations. It pulls 3.6 million figures from 2007 to 2022 in the IMPACT archive and pairs them with text drawn from PatentsView, stored in yearly Parquet files along with an 800-sample evaluation set.
Financials data for 3000+ companies parsed directly from the SEC since 2006
AlphaAI's near-real-time parse of SEC EDGAR Form 4 filings, providing 20,832 individual insider transaction tranches and 8,538 grouped economic events for US public-company officers, directors, and 10% owners.
View source information and access options.
Real estate listings in Madrid crawled from popular internet portals
Clean and structured apartment prices data from Greater Cairo real estate market
EDGAR filings of companies
An analysis-ready compilation of roughly 3,000 ClinicalTrials.gov studies registered between 2000 and 2025 that involve AI, machine learning, or digital-health tools, enriched with 30 LLM-derived variables covering use case, therapeutic area, sponsor composition, trial phase, deployment score, evidence strength, and responsible-AI keyword indicators.
View source information and access options.
A subset of the BIGPATENT corpus adapted for the MTEB clustering benchmark, tens of thousands of English patents and titles under CC BY 4.0. Primarily a research and evaluation set for similarity and clustering workloads.
FinSight gathers plain-text 10-K and 10-Q SEC filings for twenty major US public firms spanning six sectors, yielding 97 records intended for BERT fine-tuning, retrieval-augmented generation, and multi-agent financial research.
A queryable DuckDB mirror of the AACT flat-file export of ClinicalTrials.gov, packaged via the clinicaltrials-database project and covering every registered trial across 48 tables totaling roughly 58 million rows.
Complete filing metadata & URLs for institutional holdings, insider trades, etc
Complete filing metadata and URLs for 8 core SEC forms across Q1-Q4 2024
Complete filing metadata and URLs for all 8-K material event disclosures in Q4
Complete filing metadata and URLs for 82,000+ annual report filings (July-Dec)
High-Conviction Insider Trading & Institutional Holding Signals
Targeted research and datasets focused on niche sectors, private firms and specialised investment topics.
Two Parquet files hold global patent publication records paired with quality indicators for Chinese invention patent families from 2003 to 2019, with each row representing a focal family and carrying cumulative measures such as semantic knowledge recombination, strict historical semantic novelty, and IPC-based robustness.
View source information and access options.
5,000-row sample: Canadian building permits, licences, planning & inspections
View source information and access options.
Over 1.3 million granted US patents from the BIGPATENT benchmark, each pairing its claims with an abstractive summary and full description text in English (CC BY 4.0, Parquet). Use it to track the technology areas where applicants are filing and to study how claim language has evolved over time.
View source information and access options.
Index of filings submitted to the Securities & Exchange Commission since 1993
Catalog of SEC filings dating back to 1993, providing a navigational guide to the EDGAR archive.
phase1-metadata is a tabular dataset published by open-edgar-sec that consolidates SEC filing-level and entity-level attributes for publicly registered companies. It is distributed as parquet with a single training split of 37,547 rows and targets analysts building reproducible pipelines over public filings.
A broad registry of business identifiers and name variants maintained for organizations operating in over 140 countries.
Comprehensive officer records spanning more than 140 jurisdictions to support corporate governance research and due diligence checks.
Corporate linkage and beneficial ownership intelligence spanning more than 140 regulatory jurisdictions worldwide.
Roster of present-day and past candidates running for U.S. federal and state-level political offices.
The us-patents dataset, published by osanseviero, is a CSV text corpus of roughly 36,000 U.S. patent records relating terms, phrases, and classification codes with similarity scores. It is distributed under an unstated license and is intended primarily for natural language processing experimentation rather than as a source of structured patent metadata.
A CC-BY-4.0 licensed compilation from Ozari Health cataloging 22 major published GLP-1 clinical trials as of May 2026, consolidating peer-reviewed trial-level data into a single structured index.
Clinical-trials records combining structured metadata and narrative text, hundreds of thousands of studies in Parquet. An early read on pharmaceutical pipeline activity and trial design trends.
Clinical-trial protocol documents from Parexel, tens of thousands of studies. Protocol text, covering inclusion criteria, endpoints and timelines, is a leading indicator for enrollment pace and pipeline risk.
ClinicalTrialSummary is a Hugging Face dataset published under the pat-jj account that pairs long-form clinical trial article text with shorter plain-language summaries. It is distributed in Parquet format with predefined train, validation, and test splits totaling 77,516 rows.