GEONAMES / ENGINEERING NOTE 04

Decision brief / data integration

A Wikidata identifier mapping, not a metadata store

Store only GeoNames ID → Wikidata QID. The consuming API Worker fetches Wikidata/Wikipedia metadata on demand, outside the GeoNames search path.

Decision

Publish only the identifier crosswalk. Fetch metadata in the consuming API Worker.

No labels, aliases, Wikipedia titles or extracts, images, or websites are copied into the GeoNames mapping store. First run a bounded P1566 coverage/collision pilot before approving a global import. Mapping import, lookup, and consumer fetching are not implemented by this documentation change.

01 / Existing contract

What GeoNames already knows

Source/current fieldGeoNames import/serviceEnrichment gap
allCountries.txtID, names, comma-separated aliases, country/admin1/admin2 codes, feature class/code, coordinates, population, timezone.cc2, admin3/admin4, elevation, DEM, and modification date are discarded by this importer. No imported Wikipedia link; aliases here are not language-tagged links.
countryInfo.txtCountry reference attributes including country TLD, capital name and the country's GeoNames ID.No country website field. TLD does not establish a particular place's official website.
Public LocationID, names, country/admin codes and joined labels, feature class/code/label, coordinates, population, timezone, aliases.No QID, Wikidata aliases, Wikipedia article title, Commons image, or P856 official website.

Evidence inspected in the current importer import-global.ts, schema.ts, types.ts, locations.ts, plus the GeoNames dump field definition. The recorded baseline has 13,472,195 locations.

02 / Identity, not resemblance

Use the external identifier, never a fuzzy join

Wikidata P1566 GeoNames ID is an external identifier. Match its string value to the numeric canonical GeoNames ID. Do not match solely by label, spelling, country, or coordinates; a city, municipality, metro area, and administrative region can be different entities.

GeoNames 1277333

Bengaluru · Q1355

Live entity JSON returned P1566 1277333, normal rank. Label and identifier are consistent.

Inspect Q1355 entity JSON ↗

GeoNames 1261481

New Delhi · Q987

Live entity JSON returned P1566 1261481, normal rank with a GeoNames source reference.

Inspect Q987 entity JSON ↗
Observed sample only. Two expected matches; zero conflicts observed in these two P1566 claims. No global mapping, duplicate/collision scan, unmatched count, or coverage percentage was measured. Coverage is unknown.

P1566's property page reports expected completeness as "always incomplete" and stability as "sometimes changes." Resolve redirects to current QIDs and retain provenance in extraction reports outside the serving table. Exclude deprecated claims; prefer preferred-rank claims when present, otherwise normal-rank claims. Report multiple distinct active values, conflicting QIDs for one GeoNames ID, and unresolved redirects for review instead of choosing one. A redirect is identity maintenance, not another place.

03 / Mapping-only storage

Two serving columns

One approved, non-null QID per mapped GeoNames ID. Absence means unmapped. Ambiguous candidates stay in review reports, not published rows. Multiple GeoNames IDs may share a QID, so QID is not a unique key.

PROPOSED · UNIMPLEMENTED · SQLite/D1
CREATE TABLE geonames_wikidata_mapping (
  geoname_id INTEGER PRIMARY KEY,
  qid TEXT NOT NULL
);

Keep extraction revisions, retrieval times, redirects, candidate QIDs, and collision reasons in publication/review reports outside the serving table. Version and roll back the mapping artifact, not entity metadata. Choose placement only after measuring a pilot; this decision neither provisions nor requires a separate D1/R2 resource.

The consuming API Worker resolves the QID for a selected GeoNames ID and requests labels, aliases, sitelinks, and requested claims from Wikibase. For Wikipedia content, use the actual language-specific sitelink with that wiki's API, never a title inferred from the place name. The consumer owns language selection, optional caching, request limits, and upstream failure handling. Metadata caches remain separate from the GeoNames search cache.

P856 · website

The consumer fetches this optional statement when needed and verifies suitability before display. It is not stored in the mapping.

P18 · image

The consumer resolves the filename and checks the Commons file's current license, creator credit, and restrictions before display. No image or filename is stored in the mapping.

Labels / Wikipedia

Fetch language-specific labels, aliases, and sitelinks on demand. A sitelink is not a GeoNames alias; article metadata stays consumer-owned.

Primary property sources: P18, P856, P1566, and Commons reuse guidance. Commons states each file may have its own credit/license requirements and recommends verifying rights; a filename is not a blanket reuse permission.

04 / Capacity and placement

Measure the identifier mapping before choosing placement

Current remote database

5.07 GB

5,074,595,840 bytes

Paid D1 limit used here

10.00 GB

10,000,000,000 bytes

Baseline headroom

4.93 GB

4,925,404,160 bytes

The earlier 512–1,024-byte row estimate included copied entity metadata and is superseded. The two-column mapping stores an integer key and QID text; measure actual SQLite page/index overhead, coverage, concurrent gazetteer growth, and staging/rollback copies on representative rows. No mapping-size measurement or global coverage estimate exists. The remote size above is a recorded baseline, not a new remote measurement. Cloudflare documents a 10 GB paid limit; the arithmetic uses a conservative decimal planning ceiling.

Do not infer that the mapping fits or exceeds D1 from the old metadata estimate. Storage placement requires measured mapping-only bytes and approval. No separate enrichment database, object-backed metadata store, or metadata backfill is part of this decision.

05 / Controlled update path

Stage, reconcile, validate, then publish

The sequence below is an implementation proposal, not an existing command or live import. Use an approved Wikidata API or bounded P1566 extract. No full dump or broad SPARQL scan is needed for the pilot.

Bounded P1566 extract
        ↓
normalize QID redirects, claim ranks, deletions
        ↓
intersect with existing GeoNames IDs; report unmatched/collisions
        ↓
stage idempotently → quarantine ambiguities → validate and measure
        ↓
explicit approval → publish version → retain rollback version
  1. Extract P1566 statements in a bounded, repeatable batch. Keep source revision and retrieval time in extraction reports. Process all values and ranks; follow redirects and capture deleted/missing entities.
  2. Intersect exact identifier values with existing GeoNames IDs. Report mapped, unmapped, duplicates/collisions, invalid identifiers, redirects, deprecated-only claims and conflicting active claims. Never fallback to labels or coordinates.
  3. Query small explicit GeoNames-ID batches at Wikidata Query Service, not a global scan. Example: PREFIX wdt: <http://www.wikidata.org/prop/direct/> SELECT ?item ?id WHERE { VALUES ?id { "1277333" "1261481" } ?item wdt:P1566 ?id }. Audit matched QIDs with Wikibase wbgetentities using props=info|claims to inspect full P1566 ranks/values and redirects. The truthy-query projection is not complete audit evidence. This pilot is proposed, not run.
  4. Review changed/removed claims and conflicting candidates. Measure mapping-only page growth, including staging and rollback copies.
  5. Stage only geoname_id and qid with a unique GeoNames key and deterministic upsert. Reconcile removals and additions; report ambiguity rather than overwriting canonical records. Validate counts, keys, and QID resolution; keep provenance in reports.
  6. Publish only after explicit approval and validation. Retain the prior mapping version for rollback. No refresh schedule is implemented.
  7. Separately implement mapping lookup and on-demand metadata fetching in the consuming API Worker. Use requested languages and actual sitelinks; verify canonical location responses survive missing mappings, missing articles, and upstream failures.

All future extract, stage, reconcile and publish commands are proposed unimplemented. Do not run GeoNames seed commands for this feed.

07 / Pilot exit criteria

Evidence needed before any import decision

  • 01Bounded extract reports total source claims, exact GeoNames-ID matches, unmatched IDs, duplicates, conflicting claims, redirect/deletion cases and rank handling. State the observation date and do not extrapolate without evidence.
  • 02Ambiguous mappings are quarantined; there is no name/coordinate-only matching and no accidental one-to-many GeoNames row creation.
  • 03Published rows contain only GeoNames ID and QID. Representative mapping-only page growth, staging/rollback copies, and projected coverage fit an approved measured storage budget.
  • 04Idempotent staging, deletion reconciliation, versioned publication, rollback, and missing-data behavior are demonstrated before exposing any API.
  • 05The consuming API Worker fetches requested Wikidata/Wikipedia metadata on demand. Missing mappings and upstream failures preserve canonical results; any image display verifies Commons attribution/license.

08 / Primary sources

Reference links

Verified claims are limited to repository source inspection, the current stated capacity figures, and the two live entity examples checked 2026-09-30. Coverage, conflict rate beyond the two examples, actual enrichment bytes, deployment, and all proposed commands remain unmeasured or unimplemented. No Wikidata data was imported.