Alternative data datasets

Browse by provider, category, marketplace, access, delivery and license

Showing 1–60 of75 results for “library:pandas”
2001jdev
clinical-trials-eligibility-graphs_subset_gpt4.5mini

This dataset is a small parquet-format subset of clinical-trial eligibility criteria represented as entity and relation graphs. It is published on a community data hub under an unspecified license.

Public Records & Filings Free View
2001jdev
clinical-trials-patient-graphs

The clinical-trials-patient-graphs dataset is a small tabular collection of patient-level clinical records distributed in Parquet format by the 2001jdev publisher. It organizes entities and relations extracted from trial data into structured rows spanning 2021 through 2023.

Healthcare & Clinical Free View
2001jdev
clinical-trials-synth-patient-profiles2

clinical-trials-synth-patient-profiles2 is a synthetic patient-profile dataset distributed by publisher 2001jdev on a public dataset registry. It pairs clinical trial identifiers with short free-text profile descriptions and a coded entity taxonomy, designed for text and entity-recognition work rather than production research.

Public Records & Filings Free View
2001jdev
clinical-trials-trec-parsed

This dataset is a parsed snapshot of ClinicalTrials.gov records, distributed in Parquet format on the Hugging Face Hub under an unspecified license. It contains roughly 52,000 trial entries drawn from public registry filings.

Public Records & Filings Free View
2001jdev
clinical-trials-trec-qrels

clinical-trials-trec-qrels is a tabular relevance-judgment file distributed by the 2001jdev user on the Hugging Face Hub. It maps clinical-trial topic identifiers to NCT registry IDs with graded relevance scores.

Public Records & Filings Free View
2001jdev
clinical-trials-trec-topics

The clinical-trials-trec-topics dataset is a small tabular collection of clinical-trial topic records distributed in Parquet format on the Hugging Face Hub. It appears to be derived from TREC clinical-trial retrieval topics, with rows partitioned by year.

Public Records & Filings Free View
aegishield
credit_card_transactions

credit_card_transactions is a small tabular dataset published by aegisheld containing anonymized customer-level credit card account and spending summaries. It is distributed as CSV with a single training split of 8,950 rows.

Card & Transactions Free View
AI-Growth-Lab
patents_claims_1.5m_traim_test

About 1.5 million US patent claims split into training and test partitions in CSV, organized for claim-level classification and summarization work (Apache 2.0). A large, ready-split corpus for IP-claims modeling.

Public Records & Filings Free View
alea-institute
kl3m-index-edgar-filings-10-k

The KL3M Index of EDGAR 10-K Filings is a tabular index published by the Alea Institute that catalogs SEC EDGAR filings with associated issuer metadata. It is distributed in Parquet format and is tagged for use with pandas, polars, and the mlcroissant dataset library.

Public Records & Filings Free View
anonymoussubmissions
earnings21-gold-transcripts-non-normalized

Dataset Card for "earnings21-gold-transcripts-non-normalized" More Information needed

Public Records & Filings Free View
araag2
TREC_Clinical-Trials

The TREC Clinical Trials collections released in 2021, 2022, and 2023 are available through the TREC homepage, with corresponding papers for each year. Artur Guimarães ([email protected]) serves as the dataset curator, while inquiries regarding the original authors can be directed separately. The dataset addresses the task of linking patient profiles with appropriate clinical trials, containing entries formatted as JSON objects with fields including query identifiers, disease labels, and accompanying clinical text.

Public Records & Filings Free View
awinml
earnings_calls_transcripts

The awinml/earnings_calls_transcripts dataset on the Hugging Face Hub packages a small collection of earnings-call transcript segments formatted as chat-style messages for fine-tuning language models. It is distributed as a parquet file and is tagged for text modality with libraries including datasets, pandas, mlcroissant, and polars.

Public Records & Filings Free View
bespokelabs
yelp_restaurant_reviews

This dataset is a filtered version of https://huggingface.co/datasets/vincha77/filtered_yelp_restaurant_reviews

News & Sentiment Free View
bespokelabs
yelp_restaurant_reviews_5k

yelp_restaurant_reviews_5k is a small text dataset of 5,079 Yelp reviews distributed by bespokelabs, distributed as a single train split. It is structured for straightforward text classification with numeric and categorical fields.

News & Sentiment Free View
bstds
us_patent

A reference collection drawn from the U.S. Patent Phrase to Phrase Matching Kaggle competition, containing additional details that are still required for full documentation.

Public Records & Filings Free View
Cadenza-Labs
apollo-llama3.3-insider-trading-generations

Cadenza-Labs/apollo-llama3.3-insider-trading-generations is a small text dataset on Hugging Face containing 1,660 training rows used to fine-tune a Llama 3.3 model for detecting dishonest messages. It is distributed in Parquet format and released under an unstated license.

Public Records & Filings Free View
ccdv
patent-classification

A sample of US patent titles, abstracts and CPC classification labels built for multi-class patent classification, tens of thousands of text records in Parquet. Useful as training material for mapping innovation activity onto technology categories over time.

Public Records & Filings Free View
chenmingxuan
Chinese-Patent-Summary

Chinese patents paired with their abstractive summaries in Mandarin, a few thousand records in JSON (Apache 2.0). A window into the pace and direction of patenting inside the Chinese technology base.

Public Records & Filings Free View
ClarusC64
market-borrow-rate-short-interest-coherence-squeeze-risk-v0.1

A small text-classification dataset by ClarusC64 that labels whether borrow-rate, short-interest and price signals cohere into a real short squeeze or diverge into a false signal. It is published on a model hub with an MIT license tag and a 10-row train split.

Estimates, Events & Fund Flows Free View
cmagganas
GenAI-job-postings-Dataset

GenAI-job-postings-Dataset is a small, US-focused text corpus of generative-AI and machine-learning job postings distributed in Parquet format. The dataset is published on the Hugging Face Hub under an unstated license and comprises a single train split of roughly 120 rows.

Jobs & Workforce Free View
cmotions
NL_restaurant_reviews

A collection of restaurant reviews was assembled in 2019 through Python-based web scraping focused on Dutch establishments, capturing both visit experiences and feature-related information. It is organized in the DatasetDict format with three splits: 116,693 training records, 14,587 test records, and 14,587 validation records.

News & Sentiment Free View
Coder-Dragon
Indian-IPO-2006-2025

Records of initial public offerings launched on Indian markets between 2006 and 2025, with fields covering open and close dates, listing date, face value, issue price and size, lot size, first-day price, total shares offered and their allocation across anchor, NII, QIB and retail categories, minimum investment, and subscription figures for each investor class.

Public Records & Filings Free View
deelow
restaurant-reviews

An artificially generated collection of apparel product listings and accompanying advertisements produced by prompting GPT-4 to invent one hundred clothing items with descriptions and then write promotional copy for each, output in a structured product and description format without subsequent manual verification.

News & Sentiment Free View
deerfieldgreen
stk-sec-filings

stk-sec-filings is a small Hugging Face dataset by publisher deerfieldgreen that consolidates U.S. SEC filings into a single Parquet resource. The file covers fewer than one thousand rows distributed across a training split, with no stated update cadence or refresh policy.

Public Records & Filings Free View
DerivedFunction01
sec-filings-snippets

DerivedFunction01's sec-filings-snippets is a public-records dataset of short text excerpts drawn from U.S. SEC filings, released on Hugging Face in Parquet format. It is sized for language-model fill-mask work and is not positioned as a commercial alternative-data product.

Public Records & Filings Free View
dmariko
clinical-trials-xml-2018-2024

ClinicalTrials.gov XML for studies registered between 2018 and 2024, parsed into CSV, hundreds of thousands of records under CC0. A clean, licensed point-in-time history of US clinical-trial registrations.

Public Records & Filings Free View
dvquys
restaurant-reviews-public-sources

Restaurant Reviews Parsing NER Aspects This dataset is for the task of identifying the aspects of the restaurants mentioned in the reviews where aspect contains information about both the entities (FOOD, AMBIENCE, ...) and the attached sentiments. The input texts are from SemEval dataset. Labels for train and val datasets are generated by prompting Llama3 while the test dataset is curatedly manual

Crypto & On-Chain Free View
Edgarium
rangement_pq_20260906_232911

This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "names": [ "shoulder_pan.pos", "shoulder_lift.pos", "elbow_flex.pos", "wrist_flex.pos", "wrist_roll.pos", "gripper.pos" ], "shape": [ 6… See the full description on the dataset page: https://huggingface.co/datasets/Edgarium/rangement_pq_20

Public Records & Filings Free View
Euterpezz
Chinese-Patent-Summary

高质量中文专利摘要数据集。

Public Records & Filings Free View
fact-den
indeed-job-postings-2026

Indeed Job Postings (2026) is a 1,000-row sample of US Indeed listings joined to company firmographics, published by fact-den on Apify. License terms are not stated, and the publisher directs users to a commercial source actor for fresh or filtered data.

Jobs & Workforce Free View
Farmaanaa
iran_inflation_and_cpi_1936_2022

A legacy archive of Iran's consumer price index and inflation series from 1936 to 2022 remains available, though an updated multisource version using current Farmaanaa methodology supersedes it.

Free & Open Data Free View
finosfoundation
EarningsCallTranscript

A curated collection of earnings call recordings broken into intelligent segments and paired with high-accuracy transcripts produced by Mistral's Voxtral model, intended for ASR benchmarking, transcription quality studies, and audio-to-text alignment research.

Estimates, Events & Fund Flows Free View
Gokce
Generated_Restaurant_Reviews_GPT3.5

license: cc-by-4.0 task_categories: text-classification language: tr tags: food Generated Review size_categories: 1K<n<10K

News & Sentiment Free View
him1411
EDGAR10-Q

Quarterly 10-Q filings parsed from EDGAR into JSON, millions of records in English (MIT). The quarterly companion to the 10-K corpus for tracking sequential changes in operating performance and risk factors.

Public Records & Filings Free View
INPI-France
French-Patent-1981-2026-Clean

A cleaned collection of French patent documents published from 1981 through 2026, each row in the Parquet file representing one complete A1 publication parsed from the original XML. It was assembled independently by a contributor with direct access to the public patent source and is optimized for distributed loading.

Public Records & Filings Free View
INPI-France
French-Patents-2020-2026-Raw

Raw French patent publications from 2020 through 2026, pulled from original A1 XML records by an independent API/FTP extraction and delivered one-document-per-row in streaming-ready parquet.

Public Records & Filings Free View
jannikseus
restaurant-reviews

A dataset card entry for restaurant reviews that flags the need for additional information beyond what is currently provided.

News & Sentiment Free View
jmparejaz
mintic_linkedin-job-postings

The mintic_linkedin-job-postings dataset aggregates LinkedIn job posting text into a single string column with roughly 124,000 rows. It is published on Hugging Face by user jmparejaz under an unstated license.

Jobs & Workforce Free View
KarthikaRajagopal
Restaurant_Reviews.tsv

Restaurant_Reviews.tsv is a small text corpus of customer restaurant reviews paired with a binary sentiment label. It is published on the Hugging Face Hub by user KarthikaRajagopal under an unstated license.

News & Sentiment Free View
kellyhongg
sec_filings

sec_filings is a small text dataset of 494 rows that pairs natural-language queries about U.S. securities filings with retrieved document facts and ground-truth answers. It is distributed under an unstated license on a public dataset hub.

Public Records & Filings Free View
Khaliladib
restaurant-reviews-dataset

A distilabel-built set of restaurant reviews packaged with a pipeline.yaml for reproducing the synthetic generation workflow via the distilabel CLI.

News & Sentiment Free View
MAY199
paris-housing-prices

An analytical project examining the Paris Housing Dataset to uncover which property attributes most strongly affect asking prices. It involves cleaning, visualizing, and statistically modeling the records to answer targeted research questions.

Public Records & Filings Free View
mindweave
job-postings-applications

Mindweave's Job Postings & Applications is a synthetic applicant-tracking and job-board dataset covering a simulated multi-industry hiring market. It is published under a CC BY-NC 4.0 license and distributed as a CSV.

Jobs & Workforce Free View
mteb
restaurant_review_sentiment

The restaurant_review_sentiment dataset, published on the Hugging Face MTEB hub, supplies a small Arabic-language corpus of restaurant reviews annotated for sentiment. It is intended for evaluating embedding models rather than for production analytics.

News & Sentiment Free View
NextGig-Rocks
global-job-postings-multi-ats

Global Job Postings – Multi-ATS is a single dated snapshot of 112,816 job listings aggregated from 23 applicant tracking systems and parsed into a 47-column schema by NextGig-Rocks. It is distributed as Parquet via the Hugging Face datasets library and targets labor-market research, not production job boards.

Jobs & Workforce Free View
open-edgar-sec
phase1-metadata

phase1-metadata is a tabular dataset published by open-edgar-sec that consolidates SEC filing-level and entity-level attributes for publicly registered companies. It is distributed as parquet with a single training split of 37,547 rows and targets analysts building reproducible pipelines over public filings.

Public Records & Filings Free View
osanseviero
us-patents

The us-patents dataset, published by osanseviero, is a CSV text corpus of roughly 36,000 U.S. patent records relating terms, phrases, and classification codes with similarity scores. It is distributed under an unstated license and is intended primarily for natural language processing experimentation rather than as a source of structured patent metadata.

Public Records & Filings Free View
pachequinho
restaurant_reviews

The restaurant_reviews dataset is a small text corpus of roughly one thousand customer review snippets paired with binary sentiment labels, published on the Hugging Face Hub under the user account pachequinho. It is distributed in CSV format with an Apache-2.0 license declaration.

News & Sentiment Free View
pankajrajdeo
Clinical_Trials

Clinical-trials records combining structured metadata and narrative text, hundreds of thousands of studies in Parquet. An early read on pharmaceutical pipeline activity and trial design trends.

Public Records & Filings Free View
pat-jj
ClinicalTrialSummary

ClinicalTrialSummary is a Hugging Face dataset published under the pat-jj account that pairs long-form clinical trial article text with shorter plain-language summaries. It is distributed in Parquet format with predefined train, validation, and test splits totaling 77,516 rows.

Public Records & Filings Free View
pat-jj
ClinicalTrialSummary_Full

The ClinicalTrialSummary_Full resource lacks further documentation, with no additional details supplied.

Public Records & Filings Free View
pettah
global-top-Index-exploring-trends-in-stock-Market

A daily record of trading activity across multiple major global stock market indices, structured for trend analysis and forecasting work in international equity markets.

Public Records & Filings Free View
pgurazada1
patent_classification

The patent_classification dataset is a small text corpus hosted on the Hugging Face Hub by publisher pgurazada1, comprising paired patent text snippets and their cooperative classification labels. It is distributed as a CSV file under an unstated license.

Public Records & Filings Free View
PhysiQuanty
Patent_FR_US_Merge_Radix_65536

Patent_FR_US_Merge_Radix_65536 is a Hugging Face dataset by publisher PhysiQuanty that merges French and US patent documents into a single text corpus. It is distributed in Parquet format and is tagged with an Apache-2.0 license, though the README provides no further description.

Public Records & Filings Free View
raftrsf
zephyr_pi0_gen_57k_for_offline_dpo_ipo

Dataset Card for "zephyr_pi0_gen_57k_for_offline_dpo_ipo" More Information needed

Public Records & Filings Free View
rjac
clinicaltrials.gov-summary_and_eligibility

clinicaltrials.gov-summary_and_eligibility is a small Parquet mirror of selected ClinicalTrials.gov records, distributed on the Hugging Face Hub under the rjac publisher. It packages study identifiers, recruitment status, titles, brief summaries, and eligibility criteria into a single tidy table of roughly three thousand rows.

Public Records & Filings Free View
Rogersurf
earnings-call-transcripts

An English-language NLP corpus of cleaned earnings call transcripts gathered from public investor-relations pages, sized between 10,000 and 100,000 documents for financial analysis and LLM work.

Public Records & Filings Free View
Roy229
fetch_huggingface_playwright_with_chunk_terminal_yahoo-finance_7945_m3x9qk_reviews_restaurant

A compilation of 150 restaurant reviews for casual and fine dining venues, drawn from the publicly available DineScope project under a CC0 1.0 license, with a production-tier usage designation and a 2026 sync date.

News & Sentiment Free View
soumakchak
earnings_call_transcript_lite

Earnings Call Transcript Lite is a small public-records dataset of corporate earnings call transcripts paired with short reference summaries. It is distributed as a CSV and is published on the Hugging Face Hub under an unstated license.

Public Records & Filings Free View
SuhaibAtef
tech-job-postings-labeled

tech-job-postings-labeled is a small public dataset of labeled technology job postings published on Hugging Face by SuhaibAtef. It comprises roughly 8,000 records tagged with source labels and is distributed in JSON.

Jobs & Workforce Free View