Alternative data datasets
Browse by provider, category, marketplace, access, delivery and license
credit_card_transactions is a small tabular dataset published by aegisheld containing anonymized customer-level credit card account and spending summaries. It is distributed as CSV with a single training split of 8,950 rows.
About 1.5 million US patent claims split into training and test partitions in CSV, organized for claim-level classification and summarization work (Apache 2.0). A large, ready-split corpus for IP-claims modeling.
The airline-otp-data dataset is a large-scale, public-style tabular repository of U.S. domestic flight records covering on-time performance, delay attribution and cancellation outcomes. Published on Hugging Face under an unstated license, it contains roughly 30 million rows in CSV format and is aimed at analysts working with airline operations data.
A small text-classification dataset by ClarusC64 that labels whether borrow-rate, short-interest and price signals cohere into a real short squeeze or diverge into a false signal. It is published on a model hub with an MIT license tag and a 10-row train split.
Records of initial public offerings launched on Indian markets between 2006 and 2025, with fields covering open and close dates, listing date, face value, issue price and size, lot size, first-day price, total shares offered and their allocation across anchor, NII, QIB and retail categories, minimum investment, and subscription figures for each investor class.
View source information and access options.
ClinicalTrials.gov XML for studies registered between 2018 and 2024, parsed into CSV, hundreds of thousands of records under CC0. A clean, licensed point-in-time history of US clinical-trial registrations.
Indeed Job Postings (2026) is a 1,000-row sample of US Indeed listings joined to company firmographics, published by fact-den on Apify. License terms are not stated, and the publisher directs users to a commercial source actor for fresh or filtered data.
A collection of 2,000 Malaysian property transaction records spanning every state offers a detailed look at the national housing market in 2025. Compiled from Brickz, a recognized source for real estate transaction data, it covers location, tenure, property type, median pricing, and transaction volumes.
Restaurant_Reviews.tsv is a small text corpus of customer restaurant reviews paired with a binary sentiment label. It is published on the Hugging Face Hub by user KarthikaRajagopal under an unstated license.
sec_filings is a small text dataset of 494 rows that pairs natural-language queries about U.S. securities filings with retrieved document facts and ground-truth answers. It is distributed under an unstated license on a public dataset hub.
An analytical project examining the Paris Housing Dataset to uncover which property attributes most strongly affect asking prices. It involves cleaning, visualizing, and statistically modeling the records to answer targeted research questions.
The restaurant_reviews dataset is a small text corpus of roughly one thousand customer review snippets paired with binary sentiment labels, published on the Hugging Face Hub under the user account pachequinho. It is distributed in CSV format with an Apache-2.0 license declaration.
A daily record of trading activity across multiple major global stock market indices, structured for trend analysis and forecasting work in international equity markets.
A compilation of 150 restaurant reviews for casual and fine dining venues, drawn from the publicly available DineScope project under a CC0 1.0 license, with a production-tier usage designation and a 2026 sync date.
Uzbek-restaurant-domain-sentiment-reviews is a small text dataset of Uzbek-language restaurant reviews paired with numeric ratings, published by Sanatbek. The dataset is distributed in CSV form and is small enough (1K–10K rows) for quick prototyping of sentiment classifiers.
Original TIFF drawings and grant full-text XML for 165,917 U.S. design patents omitted from the AI4Patents/IMPACT collection, filling gaps including 161,093 grants from 2023 to 2026 and 4,824 earlier patents missing from IMPACT.
Earnings Call Transcript Lite is a small public-records dataset of corporate earnings call transcripts paired with short reference summaries. It is distributed as a CSV and is published on the Hugging Face Hub under an unstated license.
foot_traffic is a small MIT-licensed tabular dataset on Hugging Face that records hourly pedestrian counts alongside weather conditions. It is published by user supersam7 and contains roughly 11,000 rows.
jobpostingsamples is a small sample dataset of US job postings published by talanAI. It is distributed in CSV format and is tagged for use with the Hugging Face datasets, pandas, polars, and mlcroissant libraries.
Techsalerator's offering merges anonymized mobility signals from various providers to map population movement and location visits in urban cores, business districts, transit corridors, and broader regions.
Reduced sample of the big_patent collection, balanced across text lengths to support shorter fine-tuning runs, with a roughly even spread up to one million characters suitable for training on sequences of up to 250,000 tokens.
The linkedin-job-postings dataset, published by xanderios on Hugging Face, contains a single CSV table of U.S. LinkedIn job listings with about 33,000 rows. It is tagged with a MIT license on the platform, though the underlying LinkedIn terms of service place restrictions on redistribution and use.