Company filings corpus (10-K, 10-Q, 8-K, 13F)
From AltData.wiki, The Alternative Data Encyclopedia
The complete corpus of US public-company filings since the mid-1990s — annual reports, quarterlies, current reports, insider forms and institutional holdings — with full-text search, structured financial statement data sets and XBRL company facts available through free APIs.
SEC EDGAR/Company filings corpus (10-K, 10-Q, 8-K, 13F) is a Free & Open Data data product published on web and indexed by The Alternative Data Encyclopedia.
The complete corpus of US public-company filings since the mid-1990s — annual reports, quarterlies, current reports, insider forms and institutional holdings — with full-text search, structured financial statement data sets and XBRL company facts available through free APIs.
Open-data mandates accelerated in the late 2000s when the United States launched its national catalog and Europe committed its flagship Earth-observation program to a free and open data policy, releasing continuous global satellite archives. Successive open-government directives expanded machine-readable release requirements across agencies, normalizing the expectation that non-sensitive public records be published in usable formats.
The signal
Free, openly licensed datasets published by governments, space agencies and international organizations: statistical series, satellite imagery, geographic layers and administrative records released without access fees. The category is the substrate on which much of the commercial alternative-data industry is built.
Because everyone can access raw open data, edge comes from processing speed, feature engineering and fusion rather than exclusivity: early reads of crop conditions from free imagery, port congestion from AIS tracks, or inflation nowcasts from scraped official price indices. Open satellite archives also serve as free validation for paid geospatial products before committing budget. Portals index hundreds of thousands of datasets spanning demographics, trade, health, energy, land cover and Earth observation, including full archives of civilian radar and optical satellite missions with open licenses. Derived value comes from combining layers — night lights with electricity access, land-cover change with agricultural supply, corporate filings with procurement records — into indicators no single source provides.
Data characteristics and access
License: Public domain. Delivery: API, CSV.
Agencies publish through catalog portals with APIs, bulk downloads and STAC-compliant image catalogs; cloud platforms host analysis-ready subsets so users process data in place. Teams build pipelines that monitor versioning and reprocessing notices, harmonize projections and classifications, and document provenance so downstream signals survive upstream methodology changes.
Caveats and compliance
Openness does not imply fitness: schemas change, releases slip during budget disruptions, and documentation quality varies widely. Free imagery carries revisit-time and resolution limits, and popular datasets attract crowded trades where any informational edge decays quickly.
Licenses range from public domain to attribution-required copyleft, and misreading terms creates redistribution risk. Personal-level administrative records remain exempt from open release under privacy law, so datasets that appear to expose individuals warrant scrutiny before use.
Who uses this signal
Quant funds prototype signals at zero data cost; GIS and climate teams build risk models on open imagery; journalists and NGOs use the same records for accountability work. Commercial vendors differentiate by cleaning, joining and servicing what governments publish for free.
Complementary signals
This kind of signal pairs naturally with adjacent categories of the encyclopedia:
Further reading
Discussion
Anchored on 𝕏 with the commit-style tag #… —
tweet with it and the thread picks it up.
No comments yet — start the thread on 𝕏.
More in Free & Open Data
Multi-asset history combining Bitcoin, gold and US macro (Fed) series — a compact base for cross-asset signal work.
Independent inventory of greenhouse-gas emissions for 350M+ assets and facilities worldwide — power plants, oil & gas fields, shipping, manufacturing — estimated from satellites and sensors, downloadable openly.
The world's largest clinical-trial registry with study protocols, sites, enrollment counts, outcomes and sponsor data for 450,000+ studies worldwide — queryable through a free REST API and bulk downloads; the standard dataset for pharma pipeline tracking.
Free Medicare and Medicaid datasets: provider utilization, inpatient/outpatient claims summaries, durable medical equipment, prescription drug events and hospital price transparency files — the deepest open view of US healthcare demand.
Free-tier API tracking price, volume, market cap, exchange and developer metrics across 10,000+ cryptoassets and hundreds of exchanges, with historical OHLCV endpoints — the most widely used open crypto reference feed.
Monthly petabyte-scale crawls of the open web published as WARC segments with an index service (CDX) for targeted retrieval — raw material for language models, market mapping and large-scale entity extraction.