Alternative data datasets

Browse by provider, category, marketplace, access, delivery and license

Showing 1–60 of65 results for “modality:tabular”
2001jdev
clinical-trials-eligibility-graphs-rerank

Parsed eligibility criteria from 13,229 ClinicalTrials.gov studies representing the candidate pool gathered by two first-stage retrievers during TREC Clinical Trials 2021–2023, formatted as typed entity-relation graphs for reranking models.

Public Records & Filings Free View
2001jdev
clinical-trials-patient-graphs

The clinical-trials-patient-graphs dataset is a small tabular collection of patient-level clinical records distributed in Parquet format by the 2001jdev publisher. It organizes entities and relations extracted from trial data into structured rows spanning 2021 through 2023.

Healthcare & Clinical Free View
2001jdev
clinical-trials-trec-qrels

clinical-trials-trec-qrels is a tabular relevance-judgment file distributed by the 2001jdev user on the Hugging Face Hub. It maps clinical-trial topic identifiers to NCT registry IDs with graded relevance scores.

Public Records & Filings Free View
2001jdev
clinical-trials-trec-topics

The clinical-trials-trec-topics dataset is a small tabular collection of clinical-trial topic records distributed in Parquet format on the Hugging Face Hub. It appears to be derived from TREC clinical-trial retrieval topics, with rows partitioned by year.

Public Records & Filings Free View
2024-mcm-everitt-ryan
job-postings-raw

The job-postings-raw dataset is a large-scale tabular collection of scraped job postings distributed in Parquet format. It aggregates records across multiple sources and locales, with US coverage indicated in the dataset metadata.

Jobs & Workforce Free View
aegishield
credit_card_transactions

credit_card_transactions is a small tabular dataset published by aegisheld containing anonymized customer-level credit card account and spending summaries. It is distributed as CSV with a single training split of 8,950 rows.

Card & Transactions Free View
AI-Growth-Lab
patents_claims_1.5m_traim_test

About 1.5 million US patent claims split into training and test partitions in CSV, organized for claim-level classification and summarization work (Apache 2.0). A large, ready-split corpus for IP-claims modeling.

Public Records & Filings Free View
alea-institute
kl3m-index-edgar-filings

The kl3m-index-edgar-filings dataset, published by ALEA Institute, is a Parquet-format index of SEC EDGAR filings distributed via the Hub datasets library. It contains roughly 20 million tabular records spanning the 10M–100M size category, with an unstated license.

Public Records & Filings Free View
alea-institute
kl3m-index-edgar-filings-10-k

The KL3M Index of EDGAR 10-K Filings is a tabular index published by the Alea Institute that catalogs SEC EDGAR filings with associated issuer metadata. It is distributed in Parquet format and is tagged for use with pandas, polars, and the mlcroissant dataset library.

Public Records & Filings Free View
alea-institute
kl3m-index-edgar-filings-8-k

The kl3m-index-edgar-filings-8-k dataset is an indexed catalog of Form 8-K filings from the SEC's EDGAR system, published by the ALEA Institute. It provides structured metadata for over 1.8 million corporate event reports distributed in Parquet format.

Public Records & Filings Free View
anonymous-md
EDGAR_FILINGS_DATASET_2016_2021

EDGAR_FILINGS_DATASET_2016_2021 is a Hugging Face mirror of parsed SEC EDGAR filings spanning 2016 through 2021, distributed as a single train split of about 6 million rows in Parquet format. It is published by anonymous-md with an unstated license.

Public Records & Filings Free View
anonymous-md
EDGAR_FILINGS_DATASET_2022_2026H1

EDGAR_FILINGS_DATASET_2022_2026H1 is a Hugging Face dataset that compiles parsed SEC EDGAR filings into a single tabular corpus. It covers roughly 2022 through the first half of 2026 and is distributed as Parquet with about 1.7 million records.

Public Records & Filings Free View
Babbi21SA
airline-otp-data

The airline-otp-data dataset is a large-scale, public-style tabular repository of U.S. domestic flight records covering on-time performance, delay attribution and cancellation outcomes. Published on Hugging Face under an unstated license, it contains roughly 30 million rows in CSV format and is aimed at analysts working with airline operations data.

Foot Traffic & Mobility Free View
cmotions
NL_restaurant_reviews

A collection of restaurant reviews was assembled in 2019 through Python-based web scraping focused on Dutch establishments, capturing both visit experiences and feature-related information. It is organized in the DatasetDict format with three splits: 116,693 training records, 14,587 test records, and 14,587 validation records.

News & Sentiment Free View
DerivedFunction01
sec-filings-snippets

DerivedFunction01's sec-filings-snippets is a public-records dataset of short text excerpts drawn from U.S. SEC filings, released on Hugging Face in Parquet format. It is sized for language-model fill-mask work and is not positioned as a commercial alternative-data product.

Public Records & Filings Free View
Dev123Hug456Face
airline-disruption-data

airline-disruption-data is a tabular dataset of U.S. domestic flight records published on the datasets hub, tagged as covering between 10 million and 100 million rows in Parquet format. The schema describes per-flight scheduling, delay, cancellation, and routing fields.

Foot Traffic & Mobility Free View
Dev123Hug456Face
airline-disruption-data-phase1b

The airline-disruption-data-phase1b dataset is a tabular release hosted by publisher Dev123Hug456Face that compiles scheduled and actual flight-level operations data with delay, cancellation and diversion flags. It is distributed as a single Parquet split of roughly 12 million rows.

Foot Traffic & Mobility Free View
dmariko
clinical-trials-xml-2018-2024

ClinicalTrials.gov XML for studies registered between 2018 and 2024, parsed into CSV, hundreds of thousands of records under CC0. A clean, licensed point-in-time history of US clinical-trial registrations.

Public Records & Filings Free View
dvdmrs09
patents

A large granted-patent text collection from the USPTO, tens of millions of records in Parquet (Apache 2.0). Broad coverage of claims and descriptions for IP-intensity and technology-trend measurement.

Public Records & Filings Free View
edgarcancinoe
soarm101_pickplace_orange_080e_ts_closed

Generated with the LeRobot framework, this dataset comprises 80 episodes demonstrating a pick-and-place routine using a single orange cube positioned at varying locations within the workspace. Each episode requires the robot to rotate toward the cube, open its gripper, close it around the object, and transport it to a specified drop zone.

Public Records & Filings Free View
Edgarium
rangement_pq_20260906_232911

This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "names": [ "shoulder_pan.pos", "shoulder_lift.pos", "elbow_flex.pos", "wrist_flex.pos", "wrist_roll.pos", "gripper.pos" ], "shape": [ 6… See the full description on the dataset page: https://huggingface.co/datasets/Edgarium/rangement_pq_20

Public Records & Filings Free View
edgarkim
data_test_left_22

This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so101_follower", "total_episodes": 24, "total_frames": 5358, "total_tasks": 1, "total_videos": 48, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:24" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… Se

Public Records & Filings Free View
edgarkim
left_only

A LeRobot-formatted dataset captured with a so101_follower robot, comprising 30 episodes, 6,788 frames, 60 video files, and a single chunk at 30 fps, with training covering the full episode range.

Public Records & Filings Free View
edgarkim
mimic_0223

This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so101_follower", "total_episodes": 50, "total_frames": 13379, "total_tasks": 1, "total_videos": 100, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:50" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":…

Public Records & Filings Free View
edgarkim
mimic_0223_final

This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so101_follower", "total_episodes": 709, "total_frames": 187905, "total_tasks": 1, "total_videos": 1418, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:709" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path

Public Records & Filings Free View
edgarkim
so100_test

A robotics dataset generated with the LeRobot framework contains two recorded episodes totaling 897 frames at 30 frames per second, captured using an SO100 robot performing a single task, with four video files and accompanying parquet data organized in a single chunk of up to 1,000 entries and partitioned as a train split covering episodes zero through one.

Public Records & Filings Free View
edgarkim
so100_test_edgar_2

Another LeRobot so100 robotics dataset containing 50 episodes, 29,874 frames, and 100 videos at 30 fps, structured identically to the related edgar-block release.

Public Records & Filings Free View
edgarkim
so100_test_edgar_block

A LeRobot dataset captured on the so100 robot platform, consisting of 50 episodes, 44,746 frames, and 100 videos recorded at 30 fps and split for training.

Public Records & Filings Free View
edgarkim
so100_test_edgar_blue_block_1

A robotics dataset assembled with the LeRobot framework for the SO100 platform, containing 20 episodes totaling 11,940 frames captured at 30 fps, stored in parquet data files alongside 40 corresponding video chunks.

Public Records & Filings Free View
edgarkim
so101_0212_random_50

This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so101_follower", "total_episodes": 50, "total_frames": 12620, "total_tasks": 1, "total_videos": 100, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:50" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":…

Public Records & Filings Free View
edgarkim
so101_test_0113_mimic

This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so101_follower", "total_episodes": 500, "total_frames": 136475, "total_tasks": 1, "total_videos": 1000, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:500" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path

Public Records & Filings Free View
edgarkim
so101_test_0113_mimic_3

This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so101_follower", "total_episodes": 500, "total_frames": 137269, "total_tasks": 1, "total_videos": 1000, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:500" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path

Public Records & Filings Free View
edgarkim
so101_test_0130

A LeRobot-formatted robotics capture set for the so101_follower arm, comprising 51 episodes, 13,109 frames, and 102 videos at 30 fps stored across parquet chunks.

Public Records & Filings Free View
edgarkim
so101_test_0209_random

A LeRobot-generated dataset comprising 149 episodes, 27,628 frames, and 298 video files captured at 30 fps on a so101_follower robot, with all episodes assigned to the training split and stored in the chunk-based Parquet layout expected by the codebase.

Public Records & Filings Free View
electricsheepafrica
africa-mauritius-hotel-room-occupancy-rate-2019-to-2023-for-all-hotels-by-s-ef4516bd

Prepared by Electric Sheep Africa, this MDPA-sourced dataset supplies 10 entries of hotel room occupancy rates for Mauritius over 2019-2023, delivered as ML-ready Parquet files accompanied by uniform Hugging Face metadata and source provenance.

Estimates, Events & Fund Flows Free View
electricsheepafrica
africa-mauritius-hotel-room-occupancy-rate-2019-to-2023-for-large-hotels-by-bcab993e

A small Parquet-format collection from Electric Sheep Africa reporting room occupancy rates for large hotels in Mauritius from 2019 to 2023, containing ten records drawn from MDPA and packaged with Hugging Face metadata for machine learning workflows.

Estimates, Events & Fund Flows Free View
electricsheepafrica
africa-mauritius-monthly-hotel-room-occupancy-rate-2019-to-2023-for-all-hot-657f1ac6

Assembled by Electric Sheep Africa from MDPA, this dataset provides 60 rows of monthly hotel room occupancy rates for Mauritius spanning 2019-2023, formatted as ML-ready Parquet with consistent Hugging Face metadata and source attribution.

Free & Open Data Free View
electricsheepafrica
africa-mauritius-quarterly-hotel-room-occupancy-rate-2019-to-2023-for-all-h-04a9bd9a

Compiled by Electric Sheep Africa, this MDPA-derived dataset offers 20 records of quarterly hotel room occupancy rates for Mauritius across the 2019-2023 period, distributed in ML-ready Parquet format with standardized Hugging Face metadata and source traceability.

Estimates, Events & Fund Flows Free View
electricsheepafrica
africa-mauritius-quarterly-hotel-room-occupancy-rate-2019-to-2023-for-large-4fa9bd56

Electric Sheep Africa releases 60 quarterly observations from MDPA on hotel room occupancy for large establishments in Mauritius from 2019 through 2023, formatted as Parquet with consistent metadata.

Estimates, Events & Fund Flows Free View
fact-den
indeed-job-postings-2026

Indeed Job Postings (2026) is a 1,000-row sample of US Indeed listings joined to company firmographics, published by fact-den on Apify. License terms are not stated, and the publisher directs users to a commercial source actor for fresh or filtered data.

Jobs & Workforce Free View
INPI-France
French-Patent-1981-2026-Clean

A cleaned collection of French patent documents published from 1981 through 2026, each row in the Parquet file representing one complete A1 publication parsed from the original XML. It was assembled independently by a contributor with direct access to the public patent source and is optimized for distributed loading.

Public Records & Filings Free View
jienweng
housing-prices-malaysia-2025

A collection of 2,000 Malaysian property transaction records spanning every state offers a detailed look at the national housing market in 2025. Compiled from Brickz, a recognized source for real estate transaction data, it covers location, tenure, property type, median pricing, and transaction volumes.

Public Records & Filings Free View
jlh-ibm
earnings_call

Earnings-call transcripts from an IBM research project, tens of thousands of records with finance tags, in CC0. A filings-adjacent corpus for call-tone and question-and-answer analysis.

Public Records & Filings Free View
jlohding
sp500-edgar-10k

10-K filings for S&P 500 companies retrieved from EDGAR, stored as Parquet with the filing text attached (MIT). A focused corpus for the large-cap annual disclosure cycle: risk factors, MD&A and segment detail for every S&P 500 name.

Public Records & Filings Free View
kapilrao
SEC_filings_1994_2024

Metadata for every EDGAR filing submitted to the U.S. Securities and Exchange Commission from 1994 through December 14, 2024, derived from the quarterly master index files and including CIKs, issuer names, and form types.

Public Records & Filings Free View
labofsahil
patents-publications-dataset

The patents-publications-dataset is a large-scale tabular collection of patent publication records distributed in Parquet format. Hosted by publisher labofsahil on the Hugging Face Hub under an unstated license, it is sized in the 100M–1B row range.

Public Records & Filings Free View
louisbrulenaudet
clinical-trials-embedded

Minimal information is provided on this clinical-trials-embedded dataset, with additional details needed.

Public Records & Filings Free View
MAY199
paris-housing-prices

An analytical project examining the Paris Housing Dataset to uncover which property attributes most strongly affect asking prices. It involves cleaning, visualizing, and statistically modeling the records to answer targeted research questions.

Public Records & Filings Free View
mhurhangee
us-patent-descriptions

US Patent Descriptions This dataset contains the descriptions of granted US utility patents, filtered and deduplicated.The original data comes from all granted patents in 2025 up to May 20, available from PatentsView. Splits train: 10,000 rows for model training validation: 2,500 rows for validation test: 2,500 rows for evaluation Columns patent_id: Identifier for the patent; useful for reconcilin

Public Records & Filings Free View
mindweave
job-postings-applications

Mindweave's Job Postings & Applications is a synthetic applicant-tracking and job-board dataset covering a simulated multi-industry hiring market. It is published under a CC BY-NC 4.0 license and distributed as a CSV.

Jobs & Workforce Free View
nbettencourt
google-patents-data-preview

google-patents-data-preview is a community-uploaded preview of Google Patents bibliographic records, distributed in Parquet format through the Hugging Face Hub. The single training split contains roughly 340,000 rows covering patent identifiers, classifications, and localized text fields.

News & Sentiment Free View
NextGig-Rocks
global-job-postings-multi-ats

Global Job Postings – Multi-ATS is a single dated snapshot of 112,816 job listings aggregated from 23 applicant tracking systems and parsed into a 47-column schema by NextGig-Rocks. It is distributed as Parquet via the Hugging Face datasets library and targets labor-market research, not production job boards.

Jobs & Workforce Free View
open-edgar-sec
phase1-metadata

phase1-metadata is a tabular dataset published by open-edgar-sec that consolidates SEC filing-level and entity-level attributes for publicly registered companies. It is distributed as parquet with a single training split of 37,547 rows and targets analysts building reproducible pipelines over public filings.

Public Records & Filings Free View
Rogersurf
earnings-call-transcripts

An English-language NLP corpus of cleaned earnings call transcripts gathered from public investor-relations pages, sized between 10,000 and 100,000 documents for financial analysis and LLM work.

Public Records & Filings Free View
shangdatalab-ucsd
PatentAP

A resource built for forecasting patent approval outcomes, introduced in research employing domain-specific fine-grained claim dependency graphs, with further documentation to be supplied later.

Public Records & Filings Free View
TechsaleratorLLC
FootTrafficandMobility

Techsalerator's offering merges anonymized mobility signals from various providers to map population movement and location visits in urban cores, business districts, transit corridors, and broader regions.

Foot Traffic & Mobility Free View
vab46
Clinical_trials_anchor-contextORpositive-ground-truth_LLM_LORA_ft

View source information and access options.

Public Records & Filings Free View
vab46
Clinical_trials_anchor-positive-pairs_EmbeddingModel-data

Collection of anchor-positive pairs built from concatenated title, summary, and inclusion criteria fields for individual clinical trials, where each combined text is paired with four related questions for embedding model fine-tuning.

Public Records & Filings Free View
vab46
Clinical_trials_anchor-positive-pairs_EmbeddingModel-data2

Iteration 2 of a clinical-trials corpus intended for embedding-model training and fine-tuning, enabling tasks such as ranked retrieval, document comparison across two or more anchors, and anchor-versus-chunk similarity scoring, distinguished from the prior release by finer-grained chunk segmentation.

Public Records & Filings Free View
vab46
Clinical_trials_anchor-positive-pairs_EmbeddingModel-data_final

A final anchor-positive pair dataset designed for fine-tuning embedding models on clinical trials, merging consolidated title-summary-inclusion chunk pairs with five-anchor three-positive chunk pairings.

Public Records & Filings Free View