Alternative data datasets

Browse by provider, category, marketplace, access, delivery and license

Showing 121–143 of143 results for “library:datasets”
Sicheng-Chroma
sec-filings

sec-filings is a Hugging Face dataset published by Sicheng-Chroma that distributes pre-chunked U.S. SEC filings paired with dense vector embeddings. It is distributed in optimized Parquet format and is intended for retrieval and similarity research over regulatory disclosures.

Public Records & Filings Free View
SkyWalkertT1
stock_market_dataset

stock_market_dataset is a Hugging Face community dataset published by SkyWalkertT1 that aggregates short Turkish-language market commentary paired with sentiment labels. It is distributed under an MIT-licensed CSV format and is intended for text-classification and token-classification tasks rather than live trading signals.

News & Sentiment Free View
soumakchak
earnings_call_transcript_lite

Earnings Call Transcript Lite is a small public-records dataset of corporate earnings call transcripts paired with short reference summaries. It is distributed as a CSV and is published on the Hugging Face Hub under an unstated license.

Public Records & Filings Free View
SuhaibAtef
tech-job-postings-labeled

tech-job-postings-labeled is a small public dataset of labeled technology job postings published on Hugging Face by SuhaibAtef. It comprises roughly 8,000 records tagged with source labels and is distributed in JSON.

Jobs & Workforce Free View
sutro
apple-patents-embeddings

Vector embeddings of Apple patent text totalling about 4 million entries, stored as parquet shards with a 768-dimensional float64 feature per record.

Public Records & Filings Free View
talanAI
jobpostingsamples

jobpostingsamples is a small sample dataset of US job postings published by talanAI. It is distributed in CSV format and is tagged for use with the Hugging Face datasets, pandas, polars, and mlcroissant libraries.

Jobs & Workforce Free View
tasksource
patent-phrase-similarity

The patent-phrase-similarity dataset is a text corpus that pairs patent terminology for similarity scoring. It is distributed by tasksource and contains roughly 48,548 rows across train, validation and test splits.

Public Records & Filings Free View
TechsaleratorLLC
FootTrafficandMobility

Techsalerator's offering merges anonymized mobility signals from various providers to map population movement and location visits in urban cores, business districts, transit corridors, and broader regions.

Foot Traffic & Mobility Free View
trentmkelly
uspto-patent-data

A collection of Parquet files providing one record per Chinese invention patent family granted between 2003 and 2019, capturing a cumulative quality assessment pipeline that includes metrics such as semantic knowledge recombination, strict historical semantic novelty, and IPC-based robustness measures across global patent publications.

Public Records & Filings Free View
vab46
Clinical_trials_anchor-contextORpositive-ground-truth_LLM_LORA_ft

View source information and access options.

Public Records & Filings Free View
vab46
Clinical_trials_anchor-positive-pairs_EmbeddingModel-data

Collection of anchor-positive pairs built from concatenated title, summary, and inclusion criteria fields for individual clinical trials, where each combined text is paired with four related questions for embedding model fine-tuning.

Public Records & Filings Free View
vab46
Clinical_trials_anchor-positive-pairs_EmbeddingModel-data2

Iteration 2 of a clinical-trials corpus intended for embedding-model training and fine-tuning, enabling tasks such as ranked retrieval, document comparison across two or more anchors, and anchor-versus-chunk similarity scoring, distinguished from the prior release by finer-grained chunk segmentation.

Public Records & Filings Free View
vab46
Clinical_trials_anchor-positive-pairs_EmbeddingModel-data_final

A final anchor-positive pair dataset designed for fine-tuning embedding models on clinical trials, merging consolidated title-summary-inclusion chunk pairs with five-anchor three-positive chunk pairings.

Public Records & Filings Free View
vatolinalex
restaurant_review_sentiment

The restaurant_review_sentiment dataset is a small text corpus pairing short restaurant reviews with binary sentiment labels and restaurant and user identifiers. It is hosted on the Hugging Face Hub under an unstated license and has been downloaded only a handful of times.

News & Sentiment Free View
vietmed
chem-patent-eval-assets

chem-patent-eval-assets is a small image-folder dataset published under the vietmed namespace, evidently paired with a chemistry-patent evaluation task. The license and broader provenance are not stated, and only 363 training rows are documented on its dataset card.

Public Records & Filings Free View
vincha77
filtered_yelp_restaurant_reviews

Dataset card entry for "filtered_yelp_restaurant_reviews" with additional details required.

News & Sentiment Free View
winterForestStump
10-K_sec_filings

10-K annual filings from EDGAR in Parquet, tens of thousands of documents. The standard annual-disclosure corpus for fundamentals, risk-factor and management-discussion analysis.

Public Records & Filings Free View
wyx-ucl
SUM-DATASET-BASED-EDGAR-CORPUS

View source information and access options.

Public Records & Filings Free View
xanderios
linkedin-job-postings

The linkedin-job-postings dataset, published by xanderios on Hugging Face, contains a single CSV table of U.S. LinkedIn job listings with about 33,000 rows. It is tagged with a MIT license on the platform, though the underlying LinkedIn terms of service place restrictions on redistribution and use.

Jobs & Workforce Free View
yav1327
restaurant_reviews

The restaurant_reviews dataset, published on the Hugging Face Hub by user yav1327, is a tabular and text collection of restaurant listings paired with user reviews. It is distributed in Parquet format with a single train split of 13,144 rows and an unstated license.

News & Sentiment Free View
yeong-hwan
2024-earnings-call-transcript

The 2024 Earnings Call Transcript dataset, published by yeong-hwan on Hugging Face, aggregates quarterly earnings call dialogue for individual stock tickers. It is distributed as a JSON table with 1,904 rows in a single train split.

Public Records & Filings Free View
yoonholee
completions_AIME2025_qwen3-8b-hint-ipo-0.01-5e-7_Qwen3-1.7B

View source information and access options.

Estimates, Events & Fund Flows Free View
zalizedata
tech-job-postings-salary-dataset

The Tech Job Postings & Salary Dataset from Zalize Data contains over 449,000 monthly job postings scraped from public applicant tracking systems at US and EU technology companies. It is distributed as parquet files under a CC BY-NC 4.0 license and is part of the DataForge Open Data program.

Jobs & Workforce Free View