Alternative data datasets

Browse by provider, category, marketplace, access, delivery and license

Showing 1–60 of114 results for “format:parquet”
2001jdev
clinical-trials-eligibility-graphs-rerank

Parsed eligibility criteria from 13,229 ClinicalTrials.gov studies representing the candidate pool gathered by two first-stage retrievers during TREC Clinical Trials 2021–2023, formatted as typed entity-relation graphs for reranking models.

Public Records & Filings Free View
2001jdev
clinical-trials-eligibility-graphs_subset_gpt4.5mini

This dataset is a small parquet-format subset of clinical-trial eligibility criteria represented as entity and relation graphs. It is published on a community data hub under an unspecified license.

Public Records & Filings Free View
2001jdev
clinical-trials-eligibility-graphs_v2

This dataset, published by 2001jdev on Hugging Face, packages eligibility-criteria information from clinical trials into structured entity and relation records. It is distributed in Parquet format and contains a single split of roughly 104,000 rows.

Public Records & Filings Free View
2001jdev
clinical-trials-patient-graphs

The clinical-trials-patient-graphs dataset is a small tabular collection of patient-level clinical records distributed in Parquet format by the 2001jdev publisher. It organizes entities and relations extracted from trial data into structured rows spanning 2021 through 2023.

Healthcare & Clinical Free View
2001jdev
clinical-trials-synth-patient-profiles2

clinical-trials-synth-patient-profiles2 is a synthetic patient-profile dataset distributed by publisher 2001jdev on a public dataset registry. It pairs clinical trial identifiers with short free-text profile descriptions and a coded entity taxonomy, designed for text and entity-recognition work rather than production research.

Public Records & Filings Free View
2001jdev
clinical-trials-trec-parsed

This dataset is a parsed snapshot of ClinicalTrials.gov records, distributed in Parquet format on the Hugging Face Hub under an unspecified license. It contains roughly 52,000 trial entries drawn from public registry filings.

Public Records & Filings Free View
2001jdev
clinical-trials-trec-qrels

clinical-trials-trec-qrels is a tabular relevance-judgment file distributed by the 2001jdev user on the Hugging Face Hub. It maps clinical-trial topic identifiers to NCT registry IDs with graded relevance scores.

Public Records & Filings Free View
2001jdev
clinical-trials-trec-topics

The clinical-trials-trec-topics dataset is a small tabular collection of clinical-trial topic records distributed in Parquet format on the Hugging Face Hub. It appears to be derived from TREC clinical-trial retrieval topics, with rows partitioned by year.

Public Records & Filings Free View
2024-mcm-everitt-ryan
job-postings-english-clean

The job-postings-english-clean dataset, published by 2024-mcm-everitt-ryan, is an English-language corpus of online job postings distributed as a single Parquet train split of roughly 1.76 million rows. It is tagged as US-region data and is intended for text-based machine learning workflows.

Jobs & Workforce Free View
2024-mcm-everitt-ryan
job-postings-raw

The job-postings-raw dataset is a large-scale tabular collection of scraped job postings distributed in Parquet format. It aggregates records across multiple sources and locales, with US coverage indicated in the dataset metadata.

Jobs & Workforce Free View
alea-institute
kl3m-data-edgar-10-k

A component of the ALEA Institute's KL3M Data Project supplying training material cleared of copyright concerns. The dataset card is currently a placeholder, directing readers to the GitHub repo and project paper for full documentation.

Public Records & Filings Free View
alea-institute
kl3m-data-edgar-10-q

The KL3M Data Project from the ALEA Institute supplies copyright-clean training material, and the dataset listed here is one component pending further documentation on its dedicated page.

Public Records & Filings Free View
alea-institute
kl3m-data-edgar-agreements

Material contracts and agreements extracted from EDGAR filings by the KL3M project: debt, M&A, employment and license agreements, millions of documents in Parquet. Contract language is a niche but real input for M&A, financing and litigation-event signals.

Public Records & Filings Free View
alea-institute
kl3m-index-edgar-filings

The kl3m-index-edgar-filings dataset, published by ALEA Institute, is a Parquet-format index of SEC EDGAR filings distributed via the Hub datasets library. It contains roughly 20 million tabular records spanning the 10M–100M size category, with an unstated license.

Public Records & Filings Free View
alea-institute
kl3m-index-edgar-filings-10-k

The KL3M Index of EDGAR 10-K Filings is a tabular index published by the Alea Institute that catalogs SEC EDGAR filings with associated issuer metadata. It is distributed in Parquet format and is tagged for use with pandas, polars, and the mlcroissant dataset library.

Public Records & Filings Free View
alea-institute
kl3m-index-edgar-filings-8-k

The kl3m-index-edgar-filings-8-k dataset is an indexed catalog of Form 8-K filings from the SEC's EDGAR system, published by the ALEA Institute. It provides structured metadata for over 1.8 million corporate event reports distributed in Parquet format.

Public Records & Filings Free View
anonymous-md
EDGAR_FILINGS_DATASET_2016_2021

EDGAR_FILINGS_DATASET_2016_2021 is a Hugging Face mirror of parsed SEC EDGAR filings spanning 2016 through 2021, distributed as a single train split of about 6 million rows in Parquet format. It is published by anonymous-md with an unstated license.

Public Records & Filings Free View
anonymous-md
EDGAR_FILINGS_DATASET_2022_2026H1

EDGAR_FILINGS_DATASET_2022_2026H1 is a Hugging Face dataset that compiles parsed SEC EDGAR filings into a single tabular corpus. It covers roughly 2022 through the first half of 2026 and is distributed as Parquet with about 1.7 million records.

Public Records & Filings Free View
anonymoussubmissions
earnings21-gold-transcripts-non-normalized

Dataset Card for "earnings21-gold-transcripts-non-normalized" More Information needed

Public Records & Filings Free View
araag2
TREC_Clinical-Trials

The TREC Clinical Trials collections released in 2021, 2022, and 2023 are available through the TREC homepage, with corresponding papers for each year. Artur Guimarães ([email protected]) serves as the dataset curator, while inquiries regarding the original authors can be directed separately. The dataset addresses the task of linking patient profiles with appropriate clinical trials, containing entries formatted as JSON objects with fields including query identifiers, disease labels, and accompanying clinical text.

Public Records & Filings Free View
aswanth007
big_patent

An aggregation of 1.3 million U.S. patent records each accompanied by a human-written abstractive summary, sorted into nine Cooperative Patent Classification groupings spanning human necessities, chemistry, textiles, and related domains.

Public Records & Filings Free View
atheer2104
swedish-patent-cpc-subclass-new

Collection of historical Swedish patent texts from 1885 to 1972 assigned multi-label Cooperative Patent Classification (CPC) codes, intended for retrieval, prior art searching, and multi-label classification tasks.

Public Records & Filings Free View
awinml
earnings_calls_transcripts

The awinml/earnings_calls_transcripts dataset on the Hugging Face Hub packages a small collection of earnings-call transcript segments formatted as chat-style messages for fine-tuning language models. It is distributed as a parquet file and is tagged for text modality with libraries including datasets, pandas, mlcroissant, and polars.

Public Records & Filings Free View
bespokelabs
yelp_restaurant_reviews

This dataset is a filtered version of https://huggingface.co/datasets/vincha77/filtered_yelp_restaurant_reviews

News & Sentiment Free View
bespokelabs
yelp_restaurant_reviews_5k

yelp_restaurant_reviews_5k is a small text dataset of 5,079 Yelp reviews distributed by bespokelabs, distributed as a single train split. It is structured for straightforward text classification with numeric and categorical fields.

News & Sentiment Free View
bstds
us_patent

A reference collection drawn from the U.S. Patent Phrase to Phrase Matching Kaggle competition, containing additional details that are still required for full documentation.

Public Records & Filings Free View
Cadenza-Labs
apollo-llama3.3-insider-trading-generations

Cadenza-Labs/apollo-llama3.3-insider-trading-generations is a small text dataset on Hugging Face containing 1,660 training rows used to fine-tune a Llama 3.3 model for detecting dishonest messages. It is distributed in Parquet format and released under an unstated license.

Public Records & Filings Free View
caiotheodoro
titer-edgar-officers

titer · EDGAR officer corpus 4,206,080 attested person–company–role–date tuples from SEC Forms 3/4/5, published as pointers rather than records, alongside the frozen pre-registrations that were hash-published before any measurement ran. edgar_officers.parquet: 4.2M rows, 230,405 distinct people, 20,266 issuers, 2006q1–2026q2. Column Meaning accession SEC accession number, the pointer that reconstr

Public Records & Filings Free View
ccdv
patent-classification

A sample of US patent titles, abstracts and CPC classification labels built for multi-class patent classification, tens of thousands of text records in Parquet. Useful as training material for mapping innovation activity onto technology categories over time.

Public Records & Filings Free View
chemNLP
clinical-trials-v2

clinical-trials-v2 is a Hugging Face dataset published by chemNLP containing processed ClinicalTrials.gov records stored in Parquet. It exposes three columns — filename, xml, and text — drawn from the official clinical-study XML schema.

Public Records & Filings Free View
cmagganas
GenAI-job-postings-Dataset

GenAI-job-postings-Dataset is a small, US-focused text corpus of generative-AI and machine-learning job postings distributed in Parquet format. The dataset is published on the Hugging Face Hub under an unstated license and comprises a single train split of roughly 120 rows.

Jobs & Workforce Free View
cmotions
NL_restaurant_reviews

A collection of restaurant reviews was assembled in 2019 through Python-based web scraping focused on Dutch establishments, capturing both visit experiences and feature-related information. It is organized in the DatasetDict format with three splits: 116,693 training records, 14,587 test records, and 14,587 validation records.

News & Sentiment Free View
cyrilzakka
clinical-trials

A tabular extract of clinical study descriptions drawn from ClinicalTrials.gov on 5 February 2025, emphasizing eligibility, design, and objective fields for healthcare and machine learning research.

Public Records & Filings Free View
cyrilzakka
clinical-trials-embeddings

Vector embeddings of clinical-trial records, hundreds of thousands of entries in Parquet (Apache 2.0). Enables similarity and retrieval analysis of trial design, indication crowding and competitive positioning.

Public Records & Filings Free View
Dattito
clinical-trials-data

Dattito's clinical-trials-data is a Parquet-format dataset of roughly 24 million clinical trial records sourced from public trial registries. It aggregates identifiers, conditions, and standardized medical terminology into a single table for large-scale analysis.

Public Records & Filings Free View
deerfieldgreen
stk-sec-filings

stk-sec-filings is a small Hugging Face dataset by publisher deerfieldgreen that consolidates U.S. SEC filings into a single Parquet resource. The file covers fewer than one thousand rows distributed across a training split, with no stated update cadence or refresh policy.

Public Records & Filings Free View
DenyTranDFW
edgar_xbrl_companyfacts

edgar_xbrl_companyfacts is a Hugging Face dataset published by DenyTranDFW that aggregates U.S. SEC EDGAR XBRL company-facts filings into a single Parquet file. The training split contains roughly 125 million rows of structured financial facts tagged by reporting period and unit.

Public Records & Filings Free View
DerivedFunction01
sec-filings-snippets

DerivedFunction01's sec-filings-snippets is a public-records dataset of short text excerpts drawn from U.S. SEC filings, released on Hugging Face in Parquet format. It is sized for language-model fill-mask work and is not positioned as a commercial alternative-data product.

Public Records & Filings Free View
Dev123Hug456Face
airline-disruption-data

airline-disruption-data is a tabular dataset of U.S. domestic flight records published on the datasets hub, tagged as covering between 10 million and 100 million rows in Parquet format. The schema describes per-flight scheduling, delay, cancellation, and routing fields.

Foot Traffic & Mobility Free View
Dev123Hug456Face
airline-disruption-data-phase1b

The airline-disruption-data-phase1b dataset is a tabular release hosted by publisher Dev123Hug456Face that compiles scheduled and actual flight-level operations data with delay, cancellation and diversion flags. It is distributed as a single Parquet split of roughly 12 million rows.

Foot Traffic & Mobility Free View
dvdmrs09
patents

A large granted-patent text collection from the USPTO, tens of millions of records in Parquet (Apache 2.0). Broad coverage of claims and descriptions for IP-intensity and technology-trend measurement.

Public Records & Filings Free View
dvquys
restaurant-reviews-public-sources

Restaurant Reviews Parsing NER Aspects This dataset is for the task of identifying the aspects of the restaurants mentioned in the reviews where aspect contains information about both the entities (FOOD, AMBIENCE, ...) and the attached sentiments. The input texts are from SemEval dataset. Labels for train and val datasets are generated by prompting Llama3 while the test dataset is curatedly manual

Crypto & On-Chain Free View
edgarcancinoe
celebahq_512_id_clusters

An adaptation of the CelebA-HQ face image set at 512 resolution that retains the original photographs while appending identity-cluster labels produced automatically through face embedding grouping.

Public Records & Filings Free View
edgarcancinoe
soarm101_pickplace_orange_080e_ts_closed

Generated with the LeRobot framework, this dataset comprises 80 episodes demonstrating a pick-and-place routine using a single orange cube positioned at varying locations within the workspace. Each episode requires the robot to rotate toward the cube, open its gripper, close it around the object, and transport it to a specified drop zone.

Public Records & Filings Free View
Edgarium
rangement_pq_20260906_232911

This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "names": [ "shoulder_pan.pos", "shoulder_lift.pos", "elbow_flex.pos", "wrist_flex.pos", "wrist_roll.pos", "gripper.pos" ], "shape": [ 6… See the full description on the dataset page: https://huggingface.co/datasets/Edgarium/rangement_pq_20

Public Records & Filings Free View
edgarkim
data_test_left_22

This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so101_follower", "total_episodes": 24, "total_frames": 5358, "total_tasks": 1, "total_videos": 48, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:24" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… Se

Public Records & Filings Free View
edgarkim
left_only

A LeRobot-formatted dataset captured with a so101_follower robot, comprising 30 episodes, 6,788 frames, 60 video files, and a single chunk at 30 fps, with training covering the full episode range.

Public Records & Filings Free View
edgarkim
mimic_0223

This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so101_follower", "total_episodes": 50, "total_frames": 13379, "total_tasks": 1, "total_videos": 100, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:50" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":…

Public Records & Filings Free View
edgarkim
mimic_0223_final

This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so101_follower", "total_episodes": 709, "total_frames": 187905, "total_tasks": 1, "total_videos": 1418, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:709" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path

Public Records & Filings Free View
edgarkim
so100_test

A robotics dataset generated with the LeRobot framework contains two recorded episodes totaling 897 frames at 30 frames per second, captured using an SO100 robot performing a single task, with four video files and accompanying parquet data organized in a single chunk of up to 1,000 entries and partitioned as a train split covering episodes zero through one.

Public Records & Filings Free View
edgarkim
so100_test_edgar_2

Another LeRobot so100 robotics dataset containing 50 episodes, 29,874 frames, and 100 videos at 30 fps, structured identically to the related edgar-block release.

Public Records & Filings Free View
edgarkim
so100_test_edgar_block

A LeRobot dataset captured on the so100 robot platform, consisting of 50 episodes, 44,746 frames, and 100 videos recorded at 30 fps and split for training.

Public Records & Filings Free View
edgarkim
so100_test_edgar_blue_block_1

A robotics dataset assembled with the LeRobot framework for the SO100 platform, containing 20 episodes totaling 11,940 frames captured at 30 fps, stored in parquet data files alongside 40 corresponding video chunks.

Public Records & Filings Free View
edgarkim
so101_0212_random_50

This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so101_follower", "total_episodes": 50, "total_frames": 12620, "total_tasks": 1, "total_videos": 100, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:50" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":…

Public Records & Filings Free View
edgarkim
so101_test_0113_mimic

This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so101_follower", "total_episodes": 500, "total_frames": 136475, "total_tasks": 1, "total_videos": 1000, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:500" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path

Public Records & Filings Free View
edgarkim
so101_test_0113_mimic_3

This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so101_follower", "total_episodes": 500, "total_frames": 137269, "total_tasks": 1, "total_videos": 1000, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:500" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path

Public Records & Filings Free View
edgarkim
so101_test_0130

A LeRobot-formatted robotics capture set for the so101_follower arm, comprising 51 episodes, 13,109 frames, and 102 videos at 30 fps stored across parquet chunks.

Public Records & Filings Free View
edgarkim
so101_test_0209_random

A LeRobot-generated dataset comprising 149 episodes, 27,628 frames, and 298 video files captured at 30 fps on a so101_follower robot, with all episodes assigned to the training split and stored in the chunk-based Parquet layout expected by the codebase.

Public Records & Filings Free View
electricsheepafrica
africa-mauritius-hotel-room-occupancy-rate-2019-to-2023-for-all-hotels-by-s-ef4516bd

Prepared by Electric Sheep Africa, this MDPA-sourced dataset supplies 10 entries of hotel room occupancy rates for Mauritius over 2019-2023, delivered as ML-ready Parquet files accompanied by uniform Hugging Face metadata and source provenance.

Estimates, Events & Fund Flows Free View
electricsheepafrica
africa-mauritius-hotel-room-occupancy-rate-2019-to-2023-for-large-hotels-by-bcab993e

A small Parquet-format collection from Electric Sheep Africa reporting room occupancy rates for large hotels in Mauritius from 2019 to 2023, containing ten records drawn from MDPA and packaged with Hugging Face metadata for machine learning workflows.

Estimates, Events & Fund Flows Free View