Methodology & editorial policy
How the encyclopedia turns source listings into comparable, indexable records
AltData.wiki is an open research index of alternative-data providers, datasets, marketplaces and signal categories. The goal is not to reproduce vendor marketing pages: it is to normalize fragmented listings into records that can be compared, traced back to their source and corrected when the source changes. The resulting catalog distributions are published in State of Alternative Data 2026.
Sources and provenance
Dataset records can originate from provider websites and public marketplace catalogs including Databricks Marketplace, Snowflake Marketplace, AWS Data Exchange, Datarade, Hugging Face and Kaggle. A dataset page keeps links to the marketplace or provider listing used as its source. Provider articles link to the provider's official website and, where available, additional research references.
Source metadata and encyclopedia metadata are intentionally separate. A source's “updated” date is displayed as a source update; it is not presented as an editorial review date.
Normalization and deduplication
Names, licenses, delivery labels and geographies are normalized so equivalent values can be compared. Marketplace duplicates are consolidated when there is strong evidence that multiple listings represent the same product. Duplicate URLs remain accessible when needed, but canonical and indexability signals point search engines toward the principal record.
Provider matching uses normalized names and unambiguous corporate-name variants. The encyclopedia does not create a provider article merely because an organization appears as a marketplace publisher; unresolved publishers remain attached to the source listing until they can be matched confidently to the provider register.
Category classification
Datasets are assigned to one of the encyclopedia's signal categories from curated metadata, provider context or deterministic text signals. An inferred category is labelled as inferred on the dataset fact sheet. If the evidence is ambiguous, the record remains unclassified instead of forcing a category.
Search-index quality gate
Every dataset receives a metadata-completeness score based on category, description, provider identity, source, license, delivery, history or coverage, freshness and size. Records that do not meet the minimum threshold remain available to users and through the API, but are excluded from the XML sitemap and marked noindex until their metadata improves.
This policy is designed to avoid programmatic-SEO pages whose only value is a title and a marketplace link. Indexability is therefore a quality state, not a requirement for a record to exist in the catalog.
Automation and language models
Automated crawlers and language models can assist with extraction, normalization, classification and drafting. They are not treated as primary sources. Where generated prose is used, it is constrained by catalog facts and source-linked context; uncertain structured fields are left blank rather than filled with invented values.
Freshness
The catalog is refreshed from its upstream sources on different schedules. Marketplace timestamps, crawl time and editorial state are different concepts. When a source does not publish a history depth, refresh cadence, size or other field, the fact sheet says that it is not stated instead of guessing.
Measurement and search gaps
The site does not need a third-party behavioral tracker to identify missing catalog coverage. The backend records zero-result internal search terms without attaching IP addresses, cookies or account identity, so repeated gaps can be turned into concrete catalog and editorial work. Google Search Console verification can be enabled through a deployment environment token without storing that token in the repository.
Corrections and conflicts
Every provider, dataset and category article exposes a correction form. Factual changes require supporting URLs and enter a review queue. When two sources disagree, preference is given to the primary provider or marketplace documentation that directly supports the field, and unresolved conflicts should remain explicit rather than being silently merged.
Scope and limitations
AltData.wiki is an informational research resource, not investment advice and not a certification of a vendor's legal, compliance or commercial claims. Presence in the catalog does not imply endorsement. Buyers should validate point-in-time behavior, methodology, permitted use, licensing and compliance directly with the relevant provider before using a dataset in production.