Alternative data datasets
Browse by provider, category, marketplace, access, delivery and license
A final anchor-positive pair dataset designed for fine-tuning embedding models on clinical trials, merging consolidated title-summary-inclusion chunk pairs with five-anchor three-positive chunk pairings.
The restaurant_review_sentiment dataset is a small text corpus pairing short restaurant reviews with binary sentiment labels and restaurant and user identifiers. It is hosted on the Hugging Face Hub under an unstated license and has been downloaded only a handful of times.
Dataset card entry for "filtered_yelp_restaurant_reviews" with additional details required.
10-K annual filings from EDGAR in Parquet, tens of thousands of documents. The standard annual-disclosure corpus for fundamentals, risk-factor and management-discussion analysis.
View source information and access options.
The linkedin-job-postings dataset, published by xanderios on Hugging Face, contains a single CSV table of U.S. LinkedIn job listings with about 33,000 rows. It is tagged with a MIT license on the platform, though the underlying LinkedIn terms of service place restrictions on redistribution and use.
The restaurant_reviews dataset, published on the Hugging Face Hub by user yav1327, is a tabular and text collection of restaurant listings paired with user reviews. It is distributed in Parquet format with a single train split of 13,144 rows and an unstated license.
The 2024 Earnings Call Transcript dataset, published by yeong-hwan on Hugging Face, aggregates quarterly earnings call dialogue for individual stock tickers. It is distributed as a JSON table with 1,904 rows in a single train split.
View source information and access options.
The Tech Job Postings & Salary Dataset from Zalize Data contains over 449,000 monthly job postings scraped from public applicant tracking systems at US and EU technology companies. It is distributed as parquet files under a CC BY-NC 4.0 license and is part of the DataForge Open Data program.
Normalized stock-trade disclosures filed by U.S. House and Senate members under the STOCK Act, regenerated on a schedule by the repository's own pipeline rather than curated manually.