Alternative data datasets

Browse by provider, category, marketplace, access, delivery and license

Showing 1–13 of13 results for “task_categories:text-generation”
adityaag2k
SEC-EDGAR

Datamule, Teraflop AI, and Eventual jointly contributed a 590 gigabyte SEC EDGAR collection encompassing about 8 million filings and 43 billion tokens, assembled using the datamule-python library and the official datamule API.

Crypto & On-Chain Free View
allenai
us-patents

US patent full text assembled by the Allen Institute from the USPTO, millions of records under ODC-BY: claims, abstracts, descriptions and metadata, in Parquet. A broad base for measuring IP intensity, technology landscapes and the direction of R&D spending.

Public Records & Filings Free View
baridhi
SEC-EDGAR

A large-scale release produced jointly by Datamule, Teraflop AI, and Eventual covering roughly 590 GB of SEC EDGAR content. It comprises about eight million filings totaling 43 billion tokens, gathered using the datamule-python package and the official datamule API.

Crypto & On-Chain Free View
Bose345
sp500_earnings_transcripts

S&P 500 earnings-call transcripts in Parquet under MIT, tens of thousands of calls. Independently compiled and comparable to other public transcript corpora for sentiment and tone work.

News & Sentiment Free View
churchill1254
sp500_earnings_transcripts

A collection of earnings call transcripts covering S&P 500 firms and other U.S. large-cap companies across the years 2005 through 2025, intended for financial analysis, NLP modeling, and sentiment research.

News & Sentiment Free View
cyrilzakka
clinical-trials-embeddings

Vector embeddings of clinical-trial records, hundreds of thousands of entries in Parquet (Apache 2.0). Enables similarity and retrieval analysis of trial design, indication crowding and competitive positioning.

Public Records & Filings Free View
Jeremydh911
SEC-EDGAR

A joint release by Datamule, Teraflop AI, and Eventual provides the SEC-EDGAR dataset, comprising 590 GB of content drawn from all major filings in the SEC EDGAR database and totaling roughly 8 million samples and 43 billion tokens. The data was assembled using the datamule-python library and the official datamule API developed by John Friedman, with datamule serving as a Python package for collecting and manipulating such records.

Crypto & On-Chain Free View
kapilrao
SEC-EDGAR

Datamule partnered with Teraflop AI and Eventual to publish the SEC-EDGAR dataset, which covers all major filings from the SEC EDGAR database and amounts to 590 GB across about 8 million samples and 43 billion tokens. Collection relied on the datamule-python library and John Friedman's official datamule API, tools designed in Python for gathering and handling SEC filing data.

Crypto & On-Chain Free View
kurry
sp500_earnings_transcripts

Earnings-call transcripts for S&P 500 companies, tens of thousands of calls in English with timestamps, in Parquet (MIT). Direct input for call-tone and question-and-answer sentiment analysis across the large-cap earnings cycle.

News & Sentiment Free View
lianghsun
chinese-english-technical-patent-glossary

Dataset Card for 中華民國專利技術名詞中英對照詞庫 中華民國專利技術名詞中英對照詞庫(Chinese-English Technical Patent Glossary)收錄逾 324 萬筆台灣專利技術名詞之中英對照資料,涵蓋國際專利分類(IPC)A 至 H 全部八大類,時間跨度自 2011 年至 2023 年。本資料集適用於專利翻譯、技術術語標準化、以及繁體中文語言模型在專業領域之詞彙增強。 Dataset Details Dataset Description 本資料集整理自中華民國經濟部智慧財產局(TIPO)公開之專利技術名詞中英對照詞庫。每筆資料包含一組繁體中文與英文之技術術語對照,並標註其對應的國際專利分類(IPC)代碼與資料來源編號。 資料涵蓋 IPC 八大類別: A — 人類生活需要(Human Necessities) B — 作業、運輸(Performin

Public Records & Filings Free View
Podtech
llm-jp-corpus-v4-ja_patent

The ja_patent mirror of llm-jp-corpus-v4, assembled by the LLM-jp Corpus Building Working Group at NII, reproduces the Japanese patent sub-corpus from the larger LLM-jp Corpus v4 release. It is delivered as 621 jsonl.gz files totaling 58.2 GB in compressed form, with each line holding a JSON object that includes a text field and a meta field containing the document identifier, URL, and other provenance information.

Public Records & Filings Free View
TeraflopAI
SEC-EDGAR

Text of SEC filings pulled from EDGAR spanning millions of records in English (Apache 2.0): 10-K, 10-Q, 8-K, annual reports and exhibits, tagged for text-generation and classification. A practical base corpus for earnings- and filings-driven signals, from risk factors and MD&A to management commentary.

Crypto & On-Chain Free View
Z-Edgar
CoER-RL

CoER RL Project page · Paper · Code Executable task configurations for Stage 2 bilateral Co-PPO, with separate training and internal-validation splits. These are not generated rollouts or the official evaluation panel. Contents Split File Size train train.parquet 12,705 configurations validation validation.parquet 3,186 configurations Training and validation configurations are disjoint. Internal v

Public Records & Filings Free View