Alternative data datasets

Browse by provider, category, marketplace, access, delivery and license

Showing 1–37 of37 results for “license:apache-2.0”
adityaag2k
SEC-EDGAR

Datamule, Teraflop AI, and Eventual jointly contributed a 590 gigabyte SEC EDGAR collection encompassing about 8 million filings and 43 billion tokens, assembled using the datamule-python library and the official datamule API.

Public Records & Filings Free View
baridhi
SEC-EDGAR

A large-scale release produced jointly by Datamule, Teraflop AI, and Eventual covering roughly 590 GB of SEC EDGAR content. It comprises about eight million filings totaling 43 billion tokens, gathered using the datamule-python package and the official datamule API.

Public Records & Filings Free View
BNNT
PatentMatch

PatentMatch pairs US patents with matching reference patents for retrieval evaluation, a few thousand records in JSON (Apache 2.0). A niche evaluation set for patent-similarity and citation-link models.

Public Records & Filings Free View
chenmingxuan
Chinese-Patent-Summary

Chinese patents paired with their abstractive summaries in Mandarin, a few thousand records in JSON (Apache 2.0). A window into the pace and direction of patenting inside the Chinese technology base.

Public Records & Filings Free View
Coder-Dragon
Indian-IPO-2006-2025

Records of initial public offerings launched on Indian markets between 2006 and 2025, with fields covering open and close dates, listing date, face value, issue price and size, lot size, first-day price, total shares offered and their allocation across anchor, NII, QIB and retail categories, minimum investment, and subscription figures for each investor class.

Public Records & Filings Free View
cyrilzakka
clinical-trials

A tabular extract of clinical study descriptions drawn from ClinicalTrials.gov on 5 February 2025, emphasizing eligibility, design, and objective fields for healthcare and machine learning research.

Public Records & Filings Free View
cyrilzakka
clinical-trials-embeddings

Vector embeddings of clinical-trial records, hundreds of thousands of entries in Parquet (Apache 2.0). Enables similarity and retrieval analysis of trial design, indication crowding and competitive positioning.

Public Records & Filings Free View
DerivedFunction01
sec-filings-snippets

DerivedFunction01's sec-filings-snippets is a public-records dataset of short text excerpts drawn from U.S. SEC filings, released on Hugging Face in Parquet format. It is sized for language-model fill-mask work and is not positioned as a commercial alternative-data product.

Public Records & Filings Free View
dvdmrs09
patents

A large granted-patent text collection from the USPTO, tens of millions of records in Parquet (Apache 2.0). Broad coverage of claims and descriptions for IP-intensity and technology-trend measurement.

Public Records & Filings Free View
edgarcancinoe
soarm101_pickplace_orange_080e_ts_closed

Generated with the LeRobot framework, this dataset comprises 80 episodes demonstrating a pick-and-place routine using a single orange cube positioned at varying locations within the workspace. Each episode requires the robot to rotate toward the cube, open its gripper, close it around the object, and transport it to a specified drop zone.

Public Records & Filings Free View
Edgarium
rangement_pq_20260906_232911

This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "names": [ "shoulder_pan.pos", "shoulder_lift.pos", "elbow_flex.pos", "wrist_flex.pos", "wrist_roll.pos", "gripper.pos" ], "shape": [ 6… See the full description on the dataset page: https://huggingface.co/datasets/Edgarium/rangement_pq_20

Public Records & Filings Free View
edgarkim
data_test_left_22

This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so101_follower", "total_episodes": 24, "total_frames": 5358, "total_tasks": 1, "total_videos": 48, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:24" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… Se

Public Records & Filings Free View
edgarkim
left_only

A LeRobot-formatted dataset captured with a so101_follower robot, comprising 30 episodes, 6,788 frames, 60 video files, and a single chunk at 30 fps, with training covering the full episode range.

Public Records & Filings Free View
edgarkim
mimic_0223

This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so101_follower", "total_episodes": 50, "total_frames": 13379, "total_tasks": 1, "total_videos": 100, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:50" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":…

Public Records & Filings Free View
edgarkim
mimic_0223_final

This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so101_follower", "total_episodes": 709, "total_frames": 187905, "total_tasks": 1, "total_videos": 1418, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:709" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path

Public Records & Filings Free View
edgarkim
so100_test

A robotics dataset generated with the LeRobot framework contains two recorded episodes totaling 897 frames at 30 frames per second, captured using an SO100 robot performing a single task, with four video files and accompanying parquet data organized in a single chunk of up to 1,000 entries and partitioned as a train split covering episodes zero through one.

Public Records & Filings Free View
edgarkim
so100_test_edgar_2

Another LeRobot so100 robotics dataset containing 50 episodes, 29,874 frames, and 100 videos at 30 fps, structured identically to the related edgar-block release.

Public Records & Filings Free View
edgarkim
so100_test_edgar_block

A LeRobot dataset captured on the so100 robot platform, consisting of 50 episodes, 44,746 frames, and 100 videos recorded at 30 fps and split for training.

Public Records & Filings Free View
edgarkim
so100_test_edgar_blue_block_1

A robotics dataset assembled with the LeRobot framework for the SO100 platform, containing 20 episodes totaling 11,940 frames captured at 30 fps, stored in parquet data files alongside 40 corresponding video chunks.

Public Records & Filings Free View
edgarkim
so101_0212_random_50

This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so101_follower", "total_episodes": 50, "total_frames": 12620, "total_tasks": 1, "total_videos": 100, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:50" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":…

Public Records & Filings Free View
edgarkim
so101_test_0113_mimic

This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so101_follower", "total_episodes": 500, "total_frames": 136475, "total_tasks": 1, "total_videos": 1000, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:500" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path

Public Records & Filings Free View
edgarkim
so101_test_0113_mimic_3

This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so101_follower", "total_episodes": 500, "total_frames": 137269, "total_tasks": 1, "total_videos": 1000, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:500" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path

Public Records & Filings Free View
edgarkim
so101_test_0130

A LeRobot-formatted robotics capture set for the so101_follower arm, comprising 51 episodes, 13,109 frames, and 102 videos at 30 fps stored across parquet chunks.

Public Records & Filings Free View
edgarkim
so101_test_0209_random

A LeRobot-generated dataset comprising 149 episodes, 27,628 frames, and 298 video files captured at 30 fps on a so101_follower robot, with all episodes assigned to the training split and stored in the chunk-based Parquet layout expected by the codebase.

Public Records & Filings Free View
Euterpezz
Chinese-Patent-Summary

高质量中文专利摘要数据集。

Public Records & Filings Free View
INPI-France
FR-Patent-2000-2026-Raw

View source information and access options.

Public Records & Filings Free View
INPI-France
French-Patent-1981-2026-Clean

A cleaned collection of French patent documents published from 1981 through 2026, each row in the Parquet file representing one complete A1 publication parsed from the original XML. It was assembled independently by a contributor with direct access to the public patent source and is optimized for distributed loading.

Public Records & Filings Free View
Jeremydh911
SEC-EDGAR

A joint release by Datamule, Teraflop AI, and Eventual provides the SEC-EDGAR dataset, comprising 590 GB of content drawn from all major filings in the SEC EDGAR database and totaling roughly 8 million samples and 43 billion tokens. The data was assembled using the datamule-python library and the official datamule API developed by John Friedman, with datamule serving as a Python package for collecting and manipulating such records.

Public Records & Filings Free View
kapilrao
SEC-EDGAR

Datamule partnered with Teraflop AI and Eventual to publish the SEC-EDGAR dataset, which covers all major filings from the SEC EDGAR database and amounts to 590 GB across about 8 million samples and 43 billion tokens. Collection relied on the datamule-python library and John Friedman's official datamule API, tools designed in Python for gathering and handling SEC filing data.

Public Records & Filings Free View
KarthikaRajagopal
Restaurant_Reviews.tsv

Restaurant_Reviews.tsv is a small text corpus of customer restaurant reviews paired with a binary sentiment label. It is published on the Hugging Face Hub by user KarthikaRajagopal under an unstated license.

News & Sentiment Free View
Mouuns
Patent-search

View source information and access options.

Public Records & Filings Free View
pachequinho
restaurant_reviews

The restaurant_reviews dataset is a small text corpus of roughly one thousand customer review snippets paired with binary sentiment labels, published on the Hugging Face Hub under the user account pachequinho. It is distributed in CSV format with an Apache-2.0 license declaration.

News & Sentiment Free View
PhysiQuanty
Patent_FR_US_Merge_Radix_65536

Patent_FR_US_Merge_Radix_65536 is a Hugging Face dataset by publisher PhysiQuanty that merges French and US patent documents into a single text corpus. It is distributed in Parquet format and is tagged with an Apache-2.0 license, though the README provides no further description.

Public Records & Filings Free View
TeraflopAI
SEC-EDGAR

Text of SEC filings pulled from EDGAR spanning millions of records in English (Apache 2.0): 10-K, 10-Q, 8-K, annual reports and exhibits, tagged for text-generation and classification. A practical base corpus for earnings- and filings-driven signals, from risk factors and MD&A to management commentary.

Public Records & Filings Free View
vibrantlabsai
ragas-airline-dataset

View source information and access options.

Foot Traffic & Mobility Free View
ZipLime
insider-trading

ZipLime US Insider Trading Disclosures (PIT) [!CAUTION] Use knowledge_date, not transaction_date, when backtesting. transaction_date says when a trade occurred; knowledge_date says when the filing became observable through EDGAR. Using the former as the signal date introduces look-ahead bias. visible = trades.filter(pl.col("knowledge_date") <= simulation_time) This dataset normalizes corporate-ins

Public Records & Filings Free View
zorynthiq
zoryntiq-sec-filings

A curated collection of SEC EDGAR filings from recent and upcoming IPO companies, containing 5,179 text segments sourced from 261 filings, formatted for large language model training and financial analysis.

Public Records & Filings Free View