Alternative data datasets
Browse by provider, category, marketplace, access, delivery and license
This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so101_follower", "total_episodes": 24, "total_frames": 5358, "total_tasks": 1, "total_videos": 48, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:24" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… Se
This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so101_follower", "total_episodes": 50, "total_frames": 13379, "total_tasks": 1, "total_videos": 100, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:50" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":…
This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so101_follower", "total_episodes": 709, "total_frames": 187905, "total_tasks": 1, "total_videos": 1418, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:709" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path
This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so101_follower", "total_episodes": 50, "total_frames": 12620, "total_tasks": 1, "total_videos": 100, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:50" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":…
This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so101_follower", "total_episodes": 500, "total_frames": 136475, "total_tasks": 1, "total_videos": 1000, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:500" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path
This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so101_follower", "total_episodes": 500, "total_frames": 137269, "total_tasks": 1, "total_videos": 1000, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:500" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path
Prepared by Electric Sheep Africa, this MDPA-sourced dataset supplies 10 entries of hotel room occupancy rates for Mauritius over 2019-2023, delivered as ML-ready Parquet files accompanied by uniform Hugging Face metadata and source provenance.
An Electric Sheep Africa compilation drawn from MDPA provides 10 rows tracking hotel room occupancy rates for large hotels in Mauritius across 2019 to 2023, distributed as Parquet with Hugging Face metadata and source lineage.
A small Parquet-format collection from Electric Sheep Africa reporting room occupancy rates for large hotels in Mauritius from 2019 to 2023, containing ten records drawn from MDPA and packaged with Hugging Face metadata for machine learning workflows.
A Parquet dataset of 60 records providing monthly hotel room occupancy rates across one African country for 2019 to 2023, adapted from MDPA by Electric Sheep Africa and distributed with metadata and source provenance.
Assembled by Electric Sheep Africa from MDPA, this dataset provides 60 rows of monthly hotel room occupancy rates for Mauritius spanning 2019-2023, formatted as ML-ready Parquet with consistent Hugging Face metadata and source attribution.
Compiled by Electric Sheep Africa, this MDPA-derived dataset offers 20 records of quarterly hotel room occupancy rates for Mauritius across the 2019-2023 period, distributed in ML-ready Parquet format with standardized Hugging Face metadata and source traceability.
Electric Sheep Africa releases 60 quarterly observations from MDPA on hotel room occupancy for large establishments in Mauritius from 2019 through 2023, formatted as Parquet with consistent metadata.
A compact compilation tracking the count of clinical trials across African countries, sourced from Our World in Data, formatted as parquet files, tagged under health, and re-indexed by Electric Sheep Africa with harmonized metadata and loading instructions to support discovery of African statistics.
A Parquet file of annual inflation as measured by consumer prices for Africa, sourced from World Bank gender and regional statistics, indexed by Electric Sheep Africa for African data discovery and released in the 1K–10K row range.
A small-scale inventory of African cancer clinical trials, listed in the Electric Sheep Africa catalog on Hugging Face and tagged for discovery with uniform metadata and loading notes.
Structured numeric records parsed from corporate filings such as 10-K, 10-Q, and 8-K submissions to the U.S. SEC. The figures are packaged inside a compressed DuckDB instance made available through the open-source Datapond registry for fast querying.
View source information and access options.
Indeed Job Postings (2026) is a 1,000-row sample of US Indeed listings joined to company firmographics, published by fact-den on Apify. License terms are not stated, and the publisher directs users to a commercial source actor for fresh or filtered data.
Official Consumer Price Index and inflation series for Iran drawn from two providers: the Statistical Centre of Iran (monthly, from 2011/1390 onward) and the Central Bank of Iran (annual, long historical run from 1936/1315 onward), each with its own series and a combined view covering national-level index and percentage data.
Mobile Application Usage: a Statistical Analysis of User Behaviour
Construction Work Permits in Buenos Aires, public info from BA gov agency
20 Years of Quarterly Earning Call Transcripts of Apple (2005-2025)
Indian equity market price history covering NSE-listed stocks and indices from 2000 through 2026, offered in Parquet shards of roughly 1.5GB each for efficient streaming on the Hugging Face hub. The collection draws on more than 2,500 tickers and supports both minute-level and end-of-day intervals.
Factors affecting crop yield
72 Financial Indicators of all companies from S&P500
S&P500 companies Insider Trading with large wallets such as CEO, Director etc..
Tesla insider trading with large wallets such as CEO, Director, Chief etc.
A cleaned collection of French patent documents published from 1981 through 2026, each row in the Parquet file representing one complete A1 publication parsed from the original XML. It was assembled independently by a contributor with direct access to the public patent source and is optimized for distributed loading.
A consolidated, ready-to-use repository of 13F institutional holdings drawn from SEC EDGAR filings, capturing the most recent quarterly disclosures of more than 13,000 investment managers including hedge funds, mutual fund complexes, pension funds, banks, and family offices, with each record summarizing total assets under management and individual positions.
Historical Data from SEC Company Filings
country and city wise listed hotels
A collection of 2,000 Malaysian property transaction records spanning every state offers a detailed look at the national housing market in 2025. Compiled from Brickz, a recognized source for real estate transaction data, it covers location, tenure, property type, median pricing, and transaction volumes.
Earnings-call transcripts from an IBM research project, tens of thousands of records with finance tags, in CC0. A filings-adjacent corpus for call-tone and question-and-answer analysis.
10-K filings for S&P 500 companies retrieved from EDGAR, stored as Parquet with the filing text attached (MIT). A focused corpus for the large-cap annual disclosure cycle: risk factors, MD&A and segment detail for every S&P 500 name.
Metadata for every EDGAR filing submitted to the U.S. Securities and Exchange Commission from 1994 through December 14, 2024, derived from the quarterly master index files and including CIKs, issuer names, and form types.
Extensive dataset features for personalized ecommerce recommendations using cont
Recommend people products based on what similar users purchased
Earnings Call LLM Insights 📚 Read the Full Story: For a deep dive into the methodology, the wildest moments we found, and key takeaways, check out the blog post:KnowTrend.ai: Auto-Grading Ten Years of Earnings Calls for Prescience and Delusion This dataset contains LLM-generated analysis of ~70,000+ earnings call transcripts. The analysis was performed using Kimi k2-0905-preview, focusing on extra
77 AI data centers, 12 GW of IT power, and the training runs they feed.
The patents-publications-dataset is a large-scale tabular collection of patent publication records distributed in Parquet format. Hosted by publisher labofsahil on the Hugging Face Hub under an unstated license, it is sized in the 100M–1B row range.
Minimal information is provided on this clinical-trials-embedded dataset, with additional details needed.
A Comprehensive Dataset on Real Estate Asking Prices and Property Features
2000 Property Listings Across Malaysia with Detailed Attributes and Pricing
20 year stock dataset of military arms corporation
Unlocking Real Estate Insights: Analyzing, Visualizing, and Predict with ML
25K Indian loan applicants with realistic CIBIL-based risk
An analytical project examining the Paris Housing Dataset to uncover which property attributes most strongly affect asking prices. It involves cleaning, visualizing, and statistically modeling the records to answer targeted research questions.
A Comprehensive Collection of Movie Information, Ratings and Genres
US Patent Descriptions This dataset contains the descriptions of granted US utility patents, filtered and deduplicated.The original data comes from all granted patents in 2025 up to May 20, available from PatentsView. Splits train: 10,000 rows for model training validation: 2,500 rows for validation test: 2,500 rows for evaluation Columns patent_id: Identifier for the patent; useful for reconcilin
co2 emission of cars
Clean and structured apartment prices data from Greater Cairo real estate market
An analysis-ready compilation of roughly 3,000 ClinicalTrials.gov studies registered between 2000 and 2025 that involve AI, machine learning, or digital-health tools, enriched with 30 LLM-derived variables covering use case, therapeutic area, sponsor composition, trial phase, deployment score, evidence strength, and responsible-AI keyword indicators.
Comprehensive Housing Listings with Property Details
A queryable DuckDB mirror of the AACT flat-file export of ClinicalTrials.gov, packaged via the clinicaltrials-database project and covering every registered trial across 48 tables totaling roughly 58 million rows.
google-patents-data-preview is a community-uploaded preview of Google Patents bibliographic records, distributed in Parquet format through the Hugging Face Hub. The single training split contains roughly 340,000 rows covering patent identifiers, classifications, and localized text fields.
Global Job Postings – Multi-ATS is a single dated snapshot of 112,816 job listings aggregated from 23 applicant tracking systems and parsed into a 47-column schema by NextGig-Rocks. It is distributed as Parquet via the Hugging Face datasets library and targets labor-market research, not production job boards.
21.7M processed spatial-temporal AIS & weather records for advanced maritime AI
Two Parquet files hold global patent publication records paired with quality indicators for Chinese invention patent families from 2003 to 2019, with each row representing a focal family and carrying cumulative measures such as semantic knowledge recombination, strict historical semantic novelty, and IPC-based robustness.
5,000-row sample: Canadian building permits, licences, planning & inspections