Alternative data datasets

Browse by provider, category, marketplace, access, delivery and license

Showing 61–120 of143 results for “library:datasets”
edgarkim
so101_test_0113_mimic

This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so101_follower", "total_episodes": 500, "total_frames": 136475, "total_tasks": 1, "total_videos": 1000, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:500" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path

Public Records & Filings Free View
edgarkim
so101_test_0113_mimic_3

This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so101_follower", "total_episodes": 500, "total_frames": 137269, "total_tasks": 1, "total_videos": 1000, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:500" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path

Public Records & Filings Free View
edgarkim
so101_test_0130

A LeRobot-formatted robotics capture set for the so101_follower arm, comprising 51 episodes, 13,109 frames, and 102 videos at 30 fps stored across parquet chunks.

Public Records & Filings Free View
edgarkim
so101_test_0209_random

A LeRobot-generated dataset comprising 149 episodes, 27,628 frames, and 298 video files captured at 30 fps on a so101_follower robot, with all episodes assigned to the training split and stored in the chunk-based Parquet layout expected by the codebase.

Public Records & Filings Free View
edgarkim
so_arm101

Recorded with the LeRobot v2.1 schema on a 6-DOF SO-ARM101 manipulator, this dataset contains 2 episodes totaling 409 frames captured at 30 FPS in 640x480 resolution from two cameras, showing the arm retrieving a red square block and depositing it into a green square region on the right.

Public Records & Filings Free View
electricsheepafrica
africa-mauritius-hotel-room-occupancy-rate-2019-to-2023-for-all-hotels-by-s-ef4516bd

Prepared by Electric Sheep Africa, this MDPA-sourced dataset supplies 10 entries of hotel room occupancy rates for Mauritius over 2019-2023, delivered as ML-ready Parquet files accompanied by uniform Hugging Face metadata and source provenance.

Estimates, Events & Fund Flows Free View
electricsheepafrica
africa-mauritius-hotel-room-occupancy-rate-2019-to-2023-for-large-hotels-by-bcab993e

A small Parquet-format collection from Electric Sheep Africa reporting room occupancy rates for large hotels in Mauritius from 2019 to 2023, containing ten records drawn from MDPA and packaged with Hugging Face metadata for machine learning workflows.

Estimates, Events & Fund Flows Free View
electricsheepafrica
africa-mauritius-monthly-hotel-room-occupancy-rate-2019-to-2023-for-all-hot-657f1ac6

Assembled by Electric Sheep Africa from MDPA, this dataset provides 60 rows of monthly hotel room occupancy rates for Mauritius spanning 2019-2023, formatted as ML-ready Parquet with consistent Hugging Face metadata and source attribution.

Free & Open Data Free View
electricsheepafrica
africa-mauritius-quarterly-hotel-room-occupancy-rate-2019-to-2023-for-all-h-04a9bd9a

Compiled by Electric Sheep Africa, this MDPA-derived dataset offers 20 records of quarterly hotel room occupancy rates for Mauritius across the 2019-2023 period, distributed in ML-ready Parquet format with standardized Hugging Face metadata and source traceability.

Estimates, Events & Fund Flows Free View
electricsheepafrica
africa-mauritius-quarterly-hotel-room-occupancy-rate-2019-to-2023-for-large-4fa9bd56

Electric Sheep Africa releases 60 quarterly observations from MDPA on hotel room occupancy for large establishments in Mauritius from 2019 through 2023, formatted as Parquet with consistent metadata.

Estimates, Events & Fund Flows Free View
Euterpezz
Chinese-Patent-Summary

高质量中文专利摘要数据集。

Public Records & Filings Free View
fact-den
indeed-job-postings-2026

Indeed Job Postings (2026) is a 1,000-row sample of US Indeed listings joined to company firmographics, published by fact-den on Apify. License terms are not stated, and the publisher directs users to a commercial source actor for fresh or filtered data.

Jobs & Workforce Free View
factored
us_patent_hub

A U.S. patent collection with a minimal placeholder card and no additional documentation provided.

Public Records & Filings Free View
Farmaanaa
iran_inflation_and_cpi_1936_2022

A legacy archive of Iran's consumer price index and inflation series from 1936 to 2022 remains available, though an updated multisource version using current Farmaanaa methodology supersedes it.

Free & Open Data Free View
finosfoundation
EarningsCallTranscript

A curated collection of earnings call recordings broken into intelligent segments and paired with high-accuracy transcripts produced by Mistral's Voxtral model, intended for ASR benchmarking, transcription quality studies, and audio-to-text alignment research.

Estimates, Events & Fund Flows Free View
gagan3012
edgar-corpus-embeddings

edgar-corpus-embeddings is a text and embedding corpus derived from SEC EDGAR filings, published on Hugging Face by user gagan3012. It packages sectioned 10-K content with pre-computed vector embeddings for machine learning workflows.

Public Records & Filings Free View
Gokce
Generated_Restaurant_Reviews_GPT3.5

license: cc-by-4.0 task_categories: text-classification language: tr tags: food Generated Review size_categories: 1K<n<10K

News & Sentiment Free View
him1411
EDGAR10-Q

Quarterly 10-Q filings parsed from EDGAR into JSON, millions of records in English (MIT). The quarterly companion to the 10-K corpus for tracking sequential changes in operating performance and risk factors.

Public Records & Filings Free View
INPI-France
French-Patent-1981-2026-Clean

A cleaned collection of French patent documents published from 1981 through 2026, each row in the Parquet file representing one complete A1 publication parsed from the original XML. It was assembled independently by a contributor with direct access to the public patent source and is optimized for distributed loading.

Public Records & Filings Free View
INPI-France
French-Patents-2020-2026-Raw

Raw French patent publications from 2020 through 2026, pulled from original A1 XML records by an independent API/FTP extraction and delivered one-document-per-row in streaming-ready parquet.

Public Records & Filings Free View
jannikseus
restaurant-reviews

A dataset card entry for restaurant reviews that flags the need for additional information beyond what is currently provided.

News & Sentiment Free View
jienweng
housing-prices-malaysia-2025

A collection of 2,000 Malaysian property transaction records spanning every state offers a detailed look at the national housing market in 2025. Compiled from Brickz, a recognized source for real estate transaction data, it covers location, tenure, property type, median pricing, and transaction volumes.

Public Records & Filings Free View
jlh-ibm
earnings_call

Earnings-call transcripts from an IBM research project, tens of thousands of records with finance tags, in CC0. A filings-adjacent corpus for call-tone and question-and-answer analysis.

Public Records & Filings Free View
jlohding
sp500-edgar-10k

10-K filings for S&P 500 companies retrieved from EDGAR, stored as Parquet with the filing text attached (MIT). A focused corpus for the large-cap annual disclosure cycle: risk factors, MD&A and segment detail for every S&P 500 name.

Public Records & Filings Free View
jmparejaz
mintic_linkedin-job-postings

The mintic_linkedin-job-postings dataset aggregates LinkedIn job posting text into a single string column with roughly 124,000 rows. It is published on Hugging Face by user jmparejaz under an unstated license.

Jobs & Workforce Free View
kapilrao
SEC_filings_1994_2024

Metadata for every EDGAR filing submitted to the U.S. Securities and Exchange Commission from 1994 through December 14, 2024, derived from the quarterly master index files and including CIKs, issuer names, and form types.

Public Records & Filings Free View
KarthikaRajagopal
Restaurant_Reviews.tsv

Restaurant_Reviews.tsv is a small text corpus of customer restaurant reviews paired with a binary sentiment label. It is published on the Hugging Face Hub by user KarthikaRajagopal under an unstated license.

News & Sentiment Free View
kellyhongg
sec_filings

sec_filings is a small text dataset of 494 rows that pairs natural-language queries about U.S. securities filings with retrieved document facts and ground-truth answers. It is distributed under an unstated license on a public dataset hub.

Public Records & Filings Free View
Khaliladib
restaurant-reviews-dataset

A distilabel-built set of restaurant reviews packaged with a pipeline.yaml for reproducing the synthetic generation workflow via the distilabel CLI.

News & Sentiment Free View
labofsahil
patents-publications-dataset

The patents-publications-dataset is a large-scale tabular collection of patent publication records distributed in Parquet format. Hosted by publisher labofsahil on the Hugging Face Hub under an unstated license, it is sized in the 100M–1B row range.

Public Records & Filings Free View
lealOO
EdgarItem7

EdgarItem7 is a text dataset derived from SEC filings, packaging Item 6 (Selected Financial Data) and Item 7 (Management's Discussion and Analysis) excerpts alongside filing metadata. It is hosted by publisher lealOO and distributed in Arrow format under an unstated license.

Public Records & Filings Free View
louisbrulenaudet
clinical-trials-embedded

Minimal information is provided on this clinical-trials-embedded dataset, with additional details needed.

Public Records & Filings Free View
marianna13
clinical_trials

The clinical_trials dataset distributes ClinicalTrials.gov registry records in Parquet format for offline analysis. It contains a single training split of roughly 150,000 rows sourced from the public clinical-trial registry.

Public Records & Filings Free View
matthias-ehrlich
patent-desc-test

patent-desc-test is a small text dataset on Hugging Face Datasets that contains short identifier-style strings derived from patent document numbers. The dataset is published by user matthias-ehrlich and is broadly tagged as public records–adjacent text data.

Public Records & Filings Free View
MAY199
paris-housing-prices

An analytical project examining the Paris Housing Dataset to uncover which property attributes most strongly affect asking prices. It involves cleaning, visualizing, and statistically modeling the records to answer targeted research questions.

Public Records & Filings Free View
MemGPT
example-sec-filings

A compact sample set of SEC filings used for the MemGPT agent demos, tens of thousands of text records. A small, clean entry point to filings text rather than a comprehensive corpus.

Public Records & Filings Free View
mhurhangee
us-patent-descriptions

US Patent Descriptions This dataset contains the descriptions of granted US utility patents, filtered and deduplicated.The original data comes from all granted patents in 2025 up to May 20, available from PatentsView. Splits train: 10,000 rows for model training validation: 2,500 rows for validation test: 2,500 rows for evaluation Columns patent_id: Identifier for the patent; useful for reconcilin

Public Records & Filings Free View
mhurhangee
us_patent_claim1

A collection featuring the primary independent claim of United States utility patents issued between 2005 and 2025 is provided. Because the opening claim usually establishes the widest legal boundaries of an invention, it carries particular significance, and each entry includes the patent identifier alongside its claim text.

Public Records & Filings Free View
mindweave
job-postings-applications

Mindweave's Job Postings & Applications is a synthetic applicant-tracking and job-board dataset covering a simulated multi-industry hiring market. It is published under a CC BY-NC 4.0 license and distributed as a CSV.

Jobs & Workforce Free View
Mingjuu
pubmed_clinicaltrials

View source information and access options.

Public Records & Filings Free View
mteb
restaurant_review_sentiment

The restaurant_review_sentiment dataset, published on the Hugging Face MTEB hub, supplies a small Arabic-language corpus of restaurant reviews annotated for sentiment. It is intended for evaluating embedding models rather than for production analytics.

News & Sentiment Free View
nbettencourt
google-patents-data-preview

google-patents-data-preview is a community-uploaded preview of Google Patents bibliographic records, distributed in Parquet format through the Hugging Face Hub. The single training split contains roughly 340,000 rows covering patent identifiers, classifications, and localized text fields.

News & Sentiment Free View
NextGig-Rocks
global-job-postings-multi-ats

Global Job Postings – Multi-ATS is a single dated snapshot of 112,816 job listings aggregated from 23 applicant tracking systems and parsed into a 47-column schema by NextGig-Rocks. It is distributed as Parquet via the Hugging Face datasets library and targets labor-market research, not production job boards.

Jobs & Workforce Free View
open-edgar-sec
phase1-metadata

phase1-metadata is a tabular dataset published by open-edgar-sec that consolidates SEC filing-level and entity-level attributes for publicly registered companies. It is distributed as parquet with a single training split of 37,547 rows and targets analysts building reproducible pipelines over public filings.

Public Records & Filings Free View
OpenSynth
TUDelft-Electricity-Consumption-1.0

TUDelft-Electricity-Consumption-1.0 is a high-resolution, open-source time series dataset of household electricity load published on Hugging Face by OpenSynth. It draws on multi-country trials covering thousands of households under different tariff regimes and is distributed in Parquet format.

Sensors & IoT Free View
osanseviero
us-patents

The us-patents dataset, published by osanseviero, is a CSV text corpus of roughly 36,000 U.S. patent records relating terms, phrases, and classification codes with similarity scores. It is distributed under an unstated license and is intended primarily for natural language processing experimentation rather than as a source of structured patent metadata.

Public Records & Filings Free View
pachequinho
restaurant_reviews

The restaurant_reviews dataset is a small text corpus of roughly one thousand customer review snippets paired with binary sentiment labels, published on the Hugging Face Hub under the user account pachequinho. It is distributed in CSV format with an Apache-2.0 license declaration.

News & Sentiment Free View
pankajrajdeo
Clinical_Trials

Clinical-trials records combining structured metadata and narrative text, hundreds of thousands of studies in Parquet. An early read on pharmaceutical pipeline activity and trial design trends.

Public Records & Filings Free View
Parexel
clinical-trials-protocols

Clinical-trial protocol documents from Parexel, tens of thousands of studies. Protocol text, covering inclusion criteria, endpoints and timelines, is a leading indicator for enrollment pace and pipeline risk.

Public Records & Filings Free View
pat-jj
ClinicalTrialSummary

ClinicalTrialSummary is a Hugging Face dataset published under the pat-jj account that pairs long-form clinical trial article text with shorter plain-language summaries. It is distributed in Parquet format with predefined train, validation, and test splits totaling 77,516 rows.

Public Records & Filings Free View
pat-jj
ClinicalTrialSummary_Full

The ClinicalTrialSummary_Full resource lacks further documentation, with no additional details supplied.

Public Records & Filings Free View
pettah
global-top-Index-exploring-trends-in-stock-Market

A daily record of trading activity across multiple major global stock market indices, structured for trend analysis and forecasting work in international equity markets.

Public Records & Filings Free View
pgurazada1
patent_classification

The patent_classification dataset is a small text corpus hosted on the Hugging Face Hub by publisher pgurazada1, comprising paired patent text snippets and their cooperative classification labels. It is distributed as a CSV file under an unstated license.

Public Records & Filings Free View
PhysiQuanty
Patent_FR_US_Merge_Radix_65536

Patent_FR_US_Merge_Radix_65536 is a Hugging Face dataset by publisher PhysiQuanty that merges French and US patent documents into a single text corpus. It is distributed in Parquet format and is tagged with an Apache-2.0 license, though the README provides no further description.

Public Records & Filings Free View
Podtech
llm-jp-corpus-v4-ja_patent

The ja_patent mirror of llm-jp-corpus-v4, assembled by the LLM-jp Corpus Building Working Group at NII, reproduces the Japanese patent sub-corpus from the larger LLM-jp Corpus v4 release. It is delivered as 621 jsonl.gz files totaling 58.2 GB in compressed form, with each line holding a JSON object that includes a text field and a meta field containing the document identifier, URL, and other provenance information.

Public Records & Filings Free View
raftrsf
zephyr_pi0_gen_57k_for_offline_dpo_ipo

Dataset Card for "zephyr_pi0_gen_57k_for_offline_dpo_ipo" More Information needed

Public Records & Filings Free View
rjac
clinicaltrials.gov-summary_and_eligibility

clinicaltrials.gov-summary_and_eligibility is a small Parquet mirror of selected ClinicalTrials.gov records, distributed on the Hugging Face Hub under the rjac publisher. It packages study identifiers, recruitment status, titles, brief summaries, and eligibility criteria into a single tidy table of roughly three thousand rows.

Public Records & Filings Free View
Rogersurf
earnings-call-transcripts

An English-language NLP corpus of cleaned earnings call transcripts gathered from public investor-relations pages, sized between 10,000 and 100,000 documents for financial analysis and LLM work.

Public Records & Filings Free View
Roy229
fetch_huggingface_playwright_with_chunk_terminal_yahoo-finance_7945_m3x9qk_reviews_restaurant

A compilation of 150 restaurant reviews for casual and fine dining venues, drawn from the publicly available DineScope project under a CC0 1.0 license, with a production-tier usage designation and a 2026 sync date.

News & Sentiment Free View
shangdatalab-ucsd
PatentAP

A resource built for forecasting patent approval outcomes, introduced in research employing domain-specific fine-grained claim dependency graphs, with further documentation to be supplied later.

Public Records & Filings Free View