sudoku

From AltData.wiki, The Alternative Data Encyclopedia · updated 2016-12-29

1 million numpy array pairs of Sudoku games and solutions

bryanpark/sudoku is a Free & Open Data data product published on Kaggle and indexed by The Alternative Data Encyclopedia.

Bryanpark publishes sudoku as a Kaggle offering in the Free & Open Data signal family. The listing has drawn 19,231 downloads and 423 likes on Kaggle.

The data

1 million numpy array pairs of Sudoku games and solutions

It is one of 96 listings in the Free & Open Data family; between them the practical differences come down to coverage, history depth and how the raw signal is cleaned and delivered.

The source tags it with games, puzzles and computer science.

Structure, access and licensing

License: CC0: Public Domain. Size: 71415479. Last updated: 2016-12-29.

Access is through a dataset that is downloaded from Kaggle for use in a notebook. Its most recent recorded snapshot is from 2016-12-29.

The signal

Free, openly licensed datasets published by governments, space agencies and international organizations: statistical series, satellite imagery, geographic layers and administrative records released without access fees. The category is the substrate on which much of the commercial alternative-data industry is built.

Because everyone can access raw open data, edge comes from processing speed, feature engineering and fusion rather than exclusivity: early reads of crop conditions from free imagery, port congestion from AIS tracks, or inflation nowcasts from scraped official price indices. Open satellite archives also serve as free validation for paid geospatial products before committing budget. Portals index hundreds of thousands of datasets spanning demographics, trade, health, energy, land cover and Earth observation, including full archives of civilian radar and optical satellite missions with open licenses. Derived value comes from combining layers — night lights with electricity access, land-cover change with agricultural supply, corporate filings with procurement records — into indicators no single source provides.

Agencies publish through catalog portals with APIs, bulk downloads and STAC-compliant image catalogs; cloud platforms host analysis-ready subsets so users process data in place. Teams build pipelines that monitor versioning and reprocessing notices, harmonize projections and classifications, and document provenance so downstream signals survive upstream methodology changes.

Caveats and compliance

Openness does not imply fitness: schemas change, releases slip during budget disruptions, and documentation quality varies widely. Free imagery carries revisit-time and resolution limits, and popular datasets attract crowded trades where any informational edge decays quickly.

Licenses range from public domain to attribution-required copyleft, and misreading terms creates redistribution risk. Personal-level administrative records remain exempt from open release under privacy law, so datasets that appear to expose individuals warrant scrutiny before use.

Who uses this signal

Quant funds prototype signals at zero data cost; GIS and climate teams build risk models on open imagery; journalists and NGOs use the same records for accountability work. Commercial vendors differentiate by cleaning, joining and servicing what governments publish for free.

Complementary signals

This kind of signal pairs naturally with adjacent categories of the encyclopedia:

Further reading

Discussion

Anchored on 𝕏 with the commit-style tag #… — tweet with it and the thread picks it up.

Discuss on 𝕏

No comments yet — start the thread on 𝕏.

More in Free & Open Data