Alternative data datasets

Browse by provider, category, marketplace, access, delivery and license

Showing 1–40 of40 results for “library:dask”
2001jdev
clinical-trials-eligibility-graphs_v2

This dataset, published by 2001jdev on Hugging Face, packages eligibility-criteria information from clinical trials into structured entity and relation records. It is distributed in Parquet format and contains a single split of roughly 104,000 rows.

Public Records & Filings Free View
2024-mcm-everitt-ryan
job-postings-english-clean

The job-postings-english-clean dataset, published by 2024-mcm-everitt-ryan, is an English-language corpus of online job postings distributed as a single Parquet train split of roughly 1.76 million rows. It is tagged as US-region data and is intended for text-based machine learning workflows.

Jobs & Workforce Free View
2024-mcm-everitt-ryan
job-postings-raw

The job-postings-raw dataset is a large-scale tabular collection of scraped job postings distributed in Parquet format. It aggregates records across multiple sources and locales, with US coverage indicated in the dataset metadata.

Jobs & Workforce Free View
alea-institute
kl3m-data-edgar-10-k

A component of the ALEA Institute's KL3M Data Project supplying training material cleared of copyright concerns. The dataset card is currently a placeholder, directing readers to the GitHub repo and project paper for full documentation.

Public Records & Filings Free View
alea-institute
kl3m-data-edgar-10-q

The KL3M Data Project from the ALEA Institute supplies copyright-clean training material, and the dataset listed here is one component pending further documentation on its dedicated page.

Public Records & Filings Free View
alea-institute
kl3m-data-edgar-agreements

Material contracts and agreements extracted from EDGAR filings by the KL3M project: debt, M&A, employment and license agreements, millions of documents in Parquet. Contract language is a niche but real input for M&A, financing and litigation-event signals.

Public Records & Filings Free View
alea-institute
kl3m-index-edgar-filings

The kl3m-index-edgar-filings dataset, published by ALEA Institute, is a Parquet-format index of SEC EDGAR filings distributed via the Hub datasets library. It contains roughly 20 million tabular records spanning the 10M–100M size category, with an unstated license.

Public Records & Filings Free View
alea-institute
kl3m-index-edgar-filings-8-k

The kl3m-index-edgar-filings-8-k dataset is an indexed catalog of Form 8-K filings from the SEC's EDGAR system, published by the ALEA Institute. It provides structured metadata for over 1.8 million corporate event reports distributed in Parquet format.

Public Records & Filings Free View
anonymous-md
EDGAR_FILINGS_DATASET_2016_2021

EDGAR_FILINGS_DATASET_2016_2021 is a Hugging Face mirror of parsed SEC EDGAR filings spanning 2016 through 2021, distributed as a single train split of about 6 million rows in Parquet format. It is published by anonymous-md with an unstated license.

Public Records & Filings Free View
anonymous-md
EDGAR_FILINGS_DATASET_2022_2026H1

EDGAR_FILINGS_DATASET_2022_2026H1 is a Hugging Face dataset that compiles parsed SEC EDGAR filings into a single tabular corpus. It covers roughly 2022 through the first half of 2026 and is distributed as Parquet with about 1.7 million records.

Public Records & Filings Free View
Babbi21SA
airline-otp-data

The airline-otp-data dataset is a large-scale, public-style tabular repository of U.S. domestic flight records covering on-time performance, delay attribution and cancellation outcomes. Published on Hugging Face under an unstated license, it contains roughly 30 million rows in CSV format and is aimed at analysts working with airline operations data.

Unclassified Free View
chemNLP
clinical-trials-v2

clinical-trials-v2 is a Hugging Face dataset published by chemNLP containing processed ClinicalTrials.gov records stored in Parquet. It exposes three columns, filename, xml, and text, drawn from the official clinical-study XML schema.

Public Records & Filings Free View
cyrilzakka
clinical-trials

A tabular extract of clinical study descriptions drawn from ClinicalTrials.gov on 5 February 2025, emphasizing eligibility, design, and objective fields for healthcare and machine learning research.

Public Records & Filings Free View
Dattito
clinical-trials-data

Dattito's clinical-trials-data is a Parquet-format dataset of roughly 24 million clinical trial records sourced from public trial registries. It aggregates identifiers, conditions, and standardized medical terminology into a single table for large-scale analysis.

Public Records & Filings Free View
Dev123Hug456Face
airline-disruption-data

airline-disruption-data is a tabular dataset of U.S. domestic flight records published on the datasets hub, tagged as covering between 10 million and 100 million rows in Parquet format. The schema describes per-flight scheduling, delay, cancellation, and routing fields.

Foot Traffic & Mobility Free View
dvdmrs09
patents

A large granted-patent text collection from the USPTO, tens of millions of records in Parquet (Apache 2.0). Broad coverage of claims and descriptions for IP-intensity and technology-trend measurement.

Public Records & Filings Free View
edgarkim
data_0324_2

This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so101_follower", "total_episodes": 50, "total_frames": 10333, "total_tasks": 1, "total_videos": 100, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:50" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":…

Public Records & Filings Free View
edgarkim
mimic_0223

This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so101_follower", "total_episodes": 50, "total_frames": 13379, "total_tasks": 1, "total_videos": 100, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:50" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":…

Public Records & Filings Free View
edgarkim
mimic_0223_final

This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so101_follower", "total_episodes": 709, "total_frames": 187905, "total_tasks": 1, "total_videos": 1418, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:709" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path

Public Records & Filings Free View
edgarkim
new_random_150_0311

This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so101_follower", "total_episodes": 161, "total_frames": 42852, "total_tasks": 1, "total_videos": 322, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:161" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":

Public Records & Filings Free View
edgarkim
so101_0212_random_100

This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so101_follower", "total_episodes": 101, "total_frames": 26027, "total_tasks": 1, "total_videos": 202, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:101" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":

Public Records & Filings Free View
edgarkim
so101_0212_random_50

This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so101_follower", "total_episodes": 50, "total_frames": 12620, "total_tasks": 1, "total_videos": 100, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:50" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":…

Public Records & Filings Free View
edgarkim
so101_test_0113_mimic

This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so101_follower", "total_episodes": 500, "total_frames": 136475, "total_tasks": 1, "total_videos": 1000, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:500" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path

Public Records & Filings Free View
edgarkim
so101_test_0113_mimic_3

This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so101_follower", "total_episodes": 500, "total_frames": 137269, "total_tasks": 1, "total_videos": 1000, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:500" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path

Public Records & Filings Free View
factored
us_patent_hub

A U.S. patent collection with a minimal placeholder card and no additional documentation provided.

Public Records & Filings Free View
gagan3012
edgar-corpus-embeddings

edgar-corpus-embeddings is a text and embedding corpus derived from SEC EDGAR filings, published on Hugging Face by user gagan3012. It packages sectioned 10-K content with pre-computed vector embeddings for machine learning workflows.

Public Records & Filings Free View
jlohding
sp500-edgar-10k

10-K filings for S&P 500 companies retrieved from EDGAR, stored as Parquet with the filing text attached (MIT). A focused corpus for the large-cap annual disclosure cycle: risk factors, MD&A and segment detail for every S&P 500 name.

Public Records & Filings Free View
labofsahil
patents-publications-dataset

The patents-publications-dataset is a large-scale tabular collection of patent publication records distributed in Parquet format. Hosted by publisher labofsahil on the Hugging Face Hub under an unstated license, it is sized in the 100M–1B row range.

Public Records & Filings Free View
louisbrulenaudet
clinical-trials-embedded

Minimal information is provided on this clinical-trials-embedded dataset, with additional details needed.

Public Records & Filings Free View
marianna13
clinical_trials

The clinical_trials dataset distributes ClinicalTrials.gov registry records in Parquet format for offline analysis. It contains a single training split of roughly 150,000 rows sourced from the public clinical-trial registry.

Public Records & Filings Free View
mhurhangee
us-patent-descriptions

US Patent Descriptions This dataset contains the descriptions of granted US utility patents, filtered and deduplicated.The original data comes from all granted patents in 2025 up to May 20, available from PatentsView. Splits train: 10,000 rows for model training validation: 2,500 rows for validation test: 2,500 rows for evaluation Columns patent_id: Identifier for the patent; useful for reconcilin

Public Records & Filings Free View
mhurhangee
us_patent_claim1

A collection featuring the primary independent claim of United States utility patents issued between 2005 and 2025 is provided. Because the opening claim usually establishes the widest legal boundaries of an invention, it carries particular significance, and each entry includes the patent identifier alongside its claim text.

Crypto & On-Chain Free View
Mingjuu
pubmed_clinicaltrials

View source information and access options.

Public Records & Filings Free View
nbettencourt
google-patents-data-preview

google-patents-data-preview is a community-uploaded preview of Google Patents bibliographic records, distributed in Parquet format through the Hugging Face Hub. The single training split contains roughly 340,000 rows covering patent identifiers, classifications, and localized text fields.

Public Records & Filings Free View
OpenSynth
TUDelft-Electricity-Consumption-1.0

TUDelft-Electricity-Consumption-1.0 is a high-resolution, open-source time series dataset of household electricity load published on Hugging Face by OpenSynth. It draws on multi-country trials covering thousands of households under different tariff regimes and is distributed in Parquet format.

Sensors & IoT Free View
Parexel
clinical-trials-protocols

Clinical-trial protocol documents from Parexel, tens of thousands of studies. Protocol text, covering inclusion criteria, endpoints and timelines, is a leading indicator for enrollment pace and pipeline risk.

Public Records & Filings Free View
shangdatalab-ucsd
PatentAP

A resource built for forecasting patent approval outcomes, introduced in research employing domain-specific fine-grained claim dependency graphs, with further documentation to be supplied later.

Public Records & Filings Free View
Sicheng-Chroma
sec-filings

sec-filings is a Hugging Face dataset published by Sicheng-Chroma that distributes pre-chunked U.S. SEC filings paired with dense vector embeddings. It is distributed in optimized Parquet format and is intended for retrieval and similarity research over regulatory disclosures.

Public Records & Filings Free View
sutro
apple-patents-embeddings

Vector embeddings of Apple patent text totalling about 4 million entries, stored as parquet shards with a 768-dimensional float64 feature per record.

Public Records & Filings Free View
trentmkelly
uspto-patent-data

A collection of Parquet files providing one record per Chinese invention patent family granted between 2003 and 2019, capturing a cumulative quality assessment pipeline that includes metrics such as semantic knowledge recombination, strict historical semantic novelty, and IPC-based robustness measures across global patent publications.

Public Records & Filings Free View