Alternative data datasets

Browse by provider, category, marketplace, access, delivery and license

Showing 61–95 of95 results for “library:polars”
open-edgar-sec
phase1-metadata

phase1-metadata is a tabular dataset published by open-edgar-sec that consolidates SEC filing-level and entity-level attributes for publicly registered companies. It is distributed as parquet with a single training split of 37,547 rows and targets analysts building reproducible pipelines over public filings.

Public Records & Filings Free View
OpenSynth
TUDelft-Electricity-Consumption-1.0

TUDelft-Electricity-Consumption-1.0 is a high-resolution, open-source time series dataset of household electricity load published on Hugging Face by OpenSynth. It draws on multi-country trials covering thousands of households under different tariff regimes and is distributed in Parquet format.

Sensors & IoT Free View
osanseviero
us-patents

The us-patents dataset, published by osanseviero, is a CSV text corpus of roughly 36,000 U.S. patent records relating terms, phrases, and classification codes with similarity scores. It is distributed under an unstated license and is intended primarily for natural language processing experimentation rather than as a source of structured patent metadata.

Public Records & Filings Free View
pachequinho
restaurant_reviews

The restaurant_reviews dataset is a small text corpus of roughly one thousand customer review snippets paired with binary sentiment labels, published on the Hugging Face Hub under the user account pachequinho. It is distributed in CSV format with an Apache-2.0 license declaration.

News & Sentiment Free View
pankajrajdeo
Clinical_Trials

Clinical-trials records combining structured metadata and narrative text, hundreds of thousands of studies in Parquet. An early read on pharmaceutical pipeline activity and trial design trends.

Public Records & Filings Free View
Parexel
clinical-trials-protocols

Clinical-trial protocol documents from Parexel, tens of thousands of studies. Protocol text, covering inclusion criteria, endpoints and timelines, is a leading indicator for enrollment pace and pipeline risk.

Public Records & Filings Free View
pat-jj
ClinicalTrialSummary

ClinicalTrialSummary is a Hugging Face dataset published under the pat-jj account that pairs long-form clinical trial article text with shorter plain-language summaries. It is distributed in Parquet format with predefined train, validation, and test splits totaling 77,516 rows.

Public Records & Filings Free View
pat-jj
ClinicalTrialSummary_Full

The ClinicalTrialSummary_Full resource lacks further documentation, with no additional details supplied.

Public Records & Filings Free View
pgurazada1
patent_classification

The patent_classification dataset is a small text corpus hosted on the Hugging Face Hub by publisher pgurazada1, comprising paired patent text snippets and their cooperative classification labels. It is distributed as a CSV file under an unstated license.

Public Records & Filings Free View
PhysiQuanty
Patent_FR_US_Merge_Radix_65536

Patent_FR_US_Merge_Radix_65536 is a Hugging Face dataset by publisher PhysiQuanty that merges French and US patent documents into a single text corpus. It is distributed in Parquet format and is tagged with an Apache-2.0 license, though the README provides no further description.

Public Records & Filings Free View
raftrsf
zephyr_pi0_gen_57k_for_offline_dpo_ipo

Dataset Card for "zephyr_pi0_gen_57k_for_offline_dpo_ipo" More Information needed

Public Records & Filings Free View
rjac
clinicaltrials.gov-summary_and_eligibility

clinicaltrials.gov-summary_and_eligibility is a small Parquet mirror of selected ClinicalTrials.gov records, distributed on the Hugging Face Hub under the rjac publisher. It packages study identifiers, recruitment status, titles, brief summaries, and eligibility criteria into a single tidy table of roughly three thousand rows.

Public Records & Filings Free View
Rogersurf
earnings-call-transcripts

An English-language NLP corpus of cleaned earnings call transcripts gathered from public investor-relations pages, sized between 10,000 and 100,000 documents for financial analysis and LLM work.

Public Records & Filings Free View
Roy229
fetch_huggingface_playwright_with_chunk_terminal_yahoo-finance_7945_m3x9qk_reviews_restaurant

A compilation of 150 restaurant reviews for casual and fine dining venues, drawn from the publicly available DineScope project under a CC0 1.0 license, with a production-tier usage designation and a 2026 sync date.

News & Sentiment Free View
shangdatalab-ucsd
PatentAP

A resource built for forecasting patent approval outcomes, introduced in research employing domain-specific fine-grained claim dependency graphs, with further documentation to be supplied later.

Public Records & Filings Free View
Sicheng-Chroma
sec-filings

sec-filings is a Hugging Face dataset published by Sicheng-Chroma that distributes pre-chunked U.S. SEC filings paired with dense vector embeddings. It is distributed in optimized Parquet format and is intended for retrieval and similarity research over regulatory disclosures.

Public Records & Filings Free View
soumakchak
earnings_call_transcript_lite

Earnings Call Transcript Lite is a small public-records dataset of corporate earnings call transcripts paired with short reference summaries. It is distributed as a CSV and is published on the Hugging Face Hub under an unstated license.

Public Records & Filings Free View
SuhaibAtef
tech-job-postings-labeled

tech-job-postings-labeled is a small public dataset of labeled technology job postings published on Hugging Face by SuhaibAtef. It comprises roughly 8,000 records tagged with source labels and is distributed in JSON.

Jobs & Workforce Free View
sutro
apple-patents-embeddings

Vector embeddings of Apple patent text totalling about 4 million entries, stored as parquet shards with a 768-dimensional float64 feature per record.

Public Records & Filings Free View
talanAI
jobpostingsamples

jobpostingsamples is a small sample dataset of US job postings published by talanAI. It is distributed in CSV format and is tagged for use with the Hugging Face datasets, pandas, polars, and mlcroissant libraries.

Jobs & Workforce Free View
tasksource
patent-phrase-similarity

The patent-phrase-similarity dataset is a text corpus that pairs patent terminology for similarity scoring. It is distributed by tasksource and contains roughly 48,548 rows across train, validation and test splits.

Public Records & Filings Free View
TechsaleratorLLC
FootTrafficandMobility

Techsalerator's offering merges anonymized mobility signals from various providers to map population movement and location visits in urban cores, business districts, transit corridors, and broader regions.

Foot Traffic & Mobility Free View
trentmkelly
uspto-patent-data

A collection of Parquet files providing one record per Chinese invention patent family granted between 2003 and 2019, capturing a cumulative quality assessment pipeline that includes metrics such as semantic knowledge recombination, strict historical semantic novelty, and IPC-based robustness measures across global patent publications.

Public Records & Filings Free View
vab46
Clinical_trials_anchor-contextORpositive-ground-truth_LLM_LORA_ft

View source information and access options.

Public Records & Filings Free View
vab46
Clinical_trials_anchor-positive-pairs_EmbeddingModel-data

Collection of anchor-positive pairs built from concatenated title, summary, and inclusion criteria fields for individual clinical trials, where each combined text is paired with four related questions for embedding model fine-tuning.

Public Records & Filings Free View
vab46
Clinical_trials_anchor-positive-pairs_EmbeddingModel-data2

Iteration 2 of a clinical-trials corpus intended for embedding-model training and fine-tuning, enabling tasks such as ranked retrieval, document comparison across two or more anchors, and anchor-versus-chunk similarity scoring, distinguished from the prior release by finer-grained chunk segmentation.

Public Records & Filings Free View
vab46
Clinical_trials_anchor-positive-pairs_EmbeddingModel-data_final

A final anchor-positive pair dataset designed for fine-tuning embedding models on clinical trials, merging consolidated title-summary-inclusion chunk pairs with five-anchor three-positive chunk pairings.

Public Records & Filings Free View
vatolinalex
restaurant_review_sentiment

The restaurant_review_sentiment dataset is a small text corpus pairing short restaurant reviews with binary sentiment labels and restaurant and user identifiers. It is hosted on the Hugging Face Hub under an unstated license and has been downloaded only a handful of times.

News & Sentiment Free View
vincha77
filtered_yelp_restaurant_reviews

Dataset card entry for "filtered_yelp_restaurant_reviews" with additional details required.

News & Sentiment Free View
winterForestStump
10-K_sec_filings

10-K annual filings from EDGAR in Parquet, tens of thousands of documents. The standard annual-disclosure corpus for fundamentals, risk-factor and management-discussion analysis.

Public Records & Filings Free View
wyx-ucl
SUM-DATASET-BASED-EDGAR-CORPUS

View source information and access options.

Public Records & Filings Free View
xanderios
linkedin-job-postings

The linkedin-job-postings dataset, published by xanderios on Hugging Face, contains a single CSV table of U.S. LinkedIn job listings with about 33,000 rows. It is tagged with a MIT license on the platform, though the underlying LinkedIn terms of service place restrictions on redistribution and use.

Jobs & Workforce Free View
yav1327
restaurant_reviews

The restaurant_reviews dataset, published on the Hugging Face Hub by user yav1327, is a tabular and text collection of restaurant listings paired with user reviews. It is distributed in Parquet format with a single train split of 13,144 rows and an unstated license.

News & Sentiment Free View
yeong-hwan
2024-earnings-call-transcript

The 2024 Earnings Call Transcript dataset, published by yeong-hwan on Hugging Face, aggregates quarterly earnings call dialogue for individual stock tickers. It is distributed as a JSON table with 1,904 rows in a single train split.

Public Records & Filings Free View
yoonholee
completions_AIME2025_qwen3-8b-hint-ipo-0.01-5e-7_Qwen3-1.7B

View source information and access options.

Estimates, Events & Fund Flows Free View