SAMPLE Doorda UK Commercial Real Estate Data Property Data 6M Locations from 320 Data
From AltData.wiki, The Alternative Data Encyclopedia
Doorda/SAMPLE Doorda UK Commercial Real Estate Data Property Data 6M Locations from 320 Data is a Public Records & Filings data product published on databricks-marketplace and indexed by The Alternative Data Encyclopedia.
Doorda publishes SAMPLE Doorda UK Commercial Real Estate Data Property Data 6M Locations from 320 Data as a databricks-marketplace offering in the Public Records & Filings signal family.
The data
Doorda maintains it as part of its Public Records & Filings portfolio and lists it through databricks-marketplace.
Structure, access and licensing
License: other. Pricing: Free (Databricks Marketplace).
Access is through a dataset that is obtained from the databricks-marketplace listing. The listing does not state an explicit refresh schedule, which is worth confirming before backtesting.
The signal
Government-mandated disclosures: securities filings, patent and trademark registrations, court records, corporate registries, lobbying reports and procurement contracts. The category converts legal transparency obligations into structured, searchable intelligence about companies and their principals.
Filing arrival times and text changes move prices within seconds, so low-latency ingestion of public disclosures is itself a trading strategy layer. Patent-filing momentum in specific technology classes anticipates product cycles and M&A interest, while state corporate registrations reveal new subsidiaries, lending relationships and restructuring plans before press coverage. Enforcement dockets flag counterparty risk earlier than ratings. Securities systems publish company filings daily at scale — annual reports, ownership changes, material-event disclosures — while intellectual-property offices expose application histories, grants, assignments and maintenance fees. Derived products include filing-timestamp alerts ahead of market reaction, litigation-risk scores, technology-position maps from patent families, and network graphs of officers across registries.
Agencies operate free search systems, bulk-download endpoints and developer APIs over structured XML; vendors repackage them with entity resolution, classification models and alerting. Pipelines parse XBRL financials, disambiguate inventor and officer names, link subsidiary hierarchies across jurisdictions, and archive historical snapshots to capture amendments and restatements.
Caveats and compliance
Public does not mean clean: OCR errors, inconsistent identifiers and jurisdictional format differences burden integration. Filing-based signals decay as access democratizes, and patent counts measure administrative activity more than innovation quality without citation weighting.
Records are public by statute, but bulk redistribution may be licensed rather than free, and some jurisdictions restrict commercial reuse of personal data embedded in filings. Users must respect scraping policies and rate limits enforced by agency systems.
Who uses this signal
Fundamental analysts read filings directly; quant desks trade filing events; IP strategists map competitor R&D from patent data; compliance teams screen politically exposed persons and sanctions networks through registry links. Journalists use identical sources for investigations.
Complementary signals
This kind of signal pairs naturally with adjacent categories of the encyclopedia:
Further reading
Discussion
Anchored on 𝕏 with the commit-style tag #… —
tweet with it and the thread picks it up.
No comments yet — start the thread on 𝕏.
More in Public Records & Filings
Parsed eligibility criteria from 13,229 ClinicalTrials.gov studies representing the candidate pool gathered by two first-stage retrievers during TREC Clinical Trials 2021–2023, formatted as typed entity-relation graphs for reranking models.
This dataset is a small parquet-format subset of clinical-trial eligibility criteria represented as entity and relation graphs. It is published on a community data hub under an unspecified license.
This dataset, published by 2001jdev on Hugging Face, packages eligibility-criteria information from clinical trials into structured entity and relation records. It is distributed in Parquet format and contains a single split of roughly 104,000 rows.
The clinical-trials-synth-patient-profiles dataset, published under the GitHub user 2001jdev, contains a modest collection of synthetic patient profile narratives associated with clinical trial identifiers. It is distributed as Parquet files and is small enough to load directly into pandas or polars for quick inspection.
This dataset is a parsed snapshot of ClinicalTrials.gov records, distributed in Parquet format on the Hugging Face Hub under an unspecified license. It contains roughly 52,000 trial entries drawn from public registry filings.
clinical-trials-trec-qrels is a tabular relevance-judgment file distributed by the 2001jdev user on the Hugging Face Hub. It maps clinical-trial topic identifiers to NCT registry IDs with graded relevance scores.