Documentation
Worldwide search
Search the worldwide GeoNames index from the deployed demo at https://geonames-global-demo.vimaksh.workers.dev. The local development demo runs on port 3000. Try Mumbai, Paris, Dal Lake, airport, railway station, or the typo Mumbia.
The demo also has a separate Countries catalog, independent of search queries and geographic scope. Its explicit Fetch action loads the selected worldwide country/territory list or fixed continent list; see the catalog API and RPC contract.
Geographic scope and defaults
Section titled “Geographic scope and defaults”Search supports countryCode, admin1Code, admin2Code, parentId, and useDefaultScope; the HTTP query parameters and typed RPC SearchOptions are listed in the API reference. Example: /api/search?q=Bengaluru&countryCode=IN&admin1Code=19. The request uses raw admin1 value 19, not qualified IN.19; qualification is used only for reference joins. Repository fixture and seed-verification data identify Bengaluru as GeoNames ID 1277333, admin1 19 (Karnataka), admin2 572 (Bangalore Urban). India is GeoNames ID 1269750.
admin1Code requires a country; admin2Code requires both country and admin1. Codes are GeoNames hierarchy codes, not ISO administrative codes. parentId resolves a real GeoNames parent and returns only strict descendants. Parent results do not include the parent itself. Explicit scopes never widen to a parent, country, or global result set.
The configurable default is India-first: initial unfiltered search uses countryCode=IN when enabled. If and only if the entire India-scoped query has zero matches, and global fallback is enabled, search falls back to global results. An empty later page is not a zero-match query and never triggers global fallback. Explicit geographic filters and parentId are strict; they never use fallback. useDefaultScope:false opts out of the default scope. Admins can change defaultCountryCode, disable default scope, or disable fallback in the protected admin settings API; see Admin settings.
parentId may be combined with explicit country/admin filters to narrow its descendants further; all supplied filters compose and must match. The selected parent itself is excluded even when it matches the other filters. Invalid hierarchy combinations are rejected rather than silently broadened.
Pagination
Section titled “Pagination”Use the exact same query and filters for every request; pass the previous response’s nextOffset as the next request’s offset. Stop only when nextOffset is null. Do not drop filters or independently calculate offsets. Global fallback is decided against the whole scoped query, not a page: an exhausted or empty later page cannot switch scope.
let offset = 0;for (;;) { const page = await env.GEONAMES.searchPage("Bengaluru", 25, offset, { countryCode: "IN", admin1Code: "19", useDefaultScope: false, }); consume(page.results); if (page.nextOffset === null) break; offset = page.nextOffset;}Relevance
Section titled “Relevance”Populated places (featureClass P) have absolute highest feature-class priority. Bounded fuzzy suggestions are promoted using the established candidate ranking within that class, then remaining results follow the existing exact-name/alias/prefix/source/population/ID order. Primary-name multiword fuzzy matches receive the existing scoring bonus; it is not an absolute precedence guarantee over population. Ranking changes can reorder results compared with earlier measurements; cache namespace is search:v6 so old ordering is not reused.
Names are normalized for matching, including case and accents. Place names, location aliases, and feature-code labels can contribute matches. Aliases are sourced from allCountries.txt; the target rebuild does not import the separate V2 alternate-name dataset. For example, airport can find locations whose feature is an airport.
Two- and three-character single-word queries search primary names only. Longer queries search names, indexed aliases and feature labels, with bounded fuzzy candidates. Fuzzy candidate retrieval is not always FTS: short prefixes and prefixes containing spaces use bounded normalized-name/alias word scans; longer prefixes use FTS. At most 400 populated-place-first candidates are scored, and at most five suggestions are returned. Scoring considers both primary names and aliases; a primary-name multiword match receives a bonus over the same spelling as an alias. Fuzzy candidates participate in the same stable, populated-place-first paginated sequence; they must not introduce duplicates or a separate first-page offset scheme.
Search execution decisions
Section titled “Search execution decisions”The Search execution section records decisions at their actual branches, separately from timing spans. timings.service.execution contains the normalized query, chosen path, ordered decision/D1 references, FTS expression, fuzzy counts/decision, second-FTS decision and current-page result count. It does not contain candidate rows or full result payloads.
Each d1 record retains its original operation name, duration and endpoints, and adds a caller label, reason, retrieval goal and observed outcome. The two fts_matches operations are explicitly labeled First FTS and Second FTS at their callers. Trace references use zero-based indices into the ordered d1 array; the UI shows their execution order without reconstructing branches from durations.
Cache and engine paths
Section titled “Cache and engine paths”- Every RPC search reads current persisted settings from D1 before building its cache key. A settings change applies to the next request, including requests that would otherwise hit KV.
- KV HIT: normalize the query, validate the saved plain page and return it. No search-engine, parent-resolution, FTS or fuzzy query runs. The settings read remains visible in D1 metrics; the collector is fresh, not the earlier MISS’s trace.
- KV MISS/BYPASS: resolve an explicit parent and apply geographic predicates inside candidate retrieval. Explicit scope never broadens.
- Default scope: when enabled and no explicit geography was supplied, evaluate the default-country query. Global fallback is allowed only when that whole query has no matches, not when an offset exceeds its results.
- Short primary-name query: retrieve exact and prefix matches together, with populated-place priority, name relevance, population and ID ordering. Aliases, feature labels and fuzzy lookup are excluded from this path.
- Longer query: retrieve scoped FTS and bounded fuzzy matches, deduplicate them, then paginate a stable sequence: populated places first, fuzzy suggestions promoted within their class using the existing candidate ranking, and remaining matches in established name relevance order. Hydration adds the country/admin/feature display fields.
- A lookahead row determines
nextOffset. Every next request must preserve the query, limit and options. Exhaustion returnsnextOffset: nullwithout changing geography.
The execution trace records decisions and actual D1 operations for that invocation. Do not assume a permanent operation count or branch from a particular query word. Settings, parent resolution, fallback, page offsets and data affect the work required.
Earlier operation sequences and latency measurements are retained in the dated repository-root SEARCH_PERFORMANCE_FINDINGS.html. They predate the scoping/ranking release and are not current engine or latency guarantees.
The search API returns up to 25 results by default, with a maximum of 50 per request. The demo requests 25 at a time and offers Load more when another page is available. nextOffset is the offset for the next request, or null when there is no next page. Preserve query, limit and every scope/filter option. Changing any of these starts a new search at offset zero.
Result fields
Section titled “Result fields”Each result is a location object with these fields:
| Field | Meaning |
|---|---|
id |
GeoNames location ID |
name, asciiname |
Display name and ASCII name |
countryCode, countryName |
Country code and resolved country name, if available |
admin1Code, admin1Name |
First-level administrative code and resolved name, if available |
admin2Code, admin2Name |
Second-level administrative code and resolved name, if available |
featureClass, featureCode |
GeoNames feature classification |
featureName, featureDescription |
Resolved feature label and description, if available |
latitude, longitude |
Coordinates in decimal degrees |
population |
GeoNames population value |
timezone |
Time zone identifier |
aliases |
Alternate search names stored with the location |
The HTTP API response retains durationMs, the rounded web Worker time across the Service Binding call. It includes service-owned KV reads, D1 work on a miss, and awaited cache writes, but excludes browser network and rendering time. Additional request-scoped timings are returned through RPC and HTTP and shown in the demo.
Request timings
Section titled “Request timings”The Search timings panel shows the latest response’s total browser request time prominently, followed by a Stage / Duration / Elapsed at completion table and a waterfall. D1 operations and stage timings is always visible below the waterfall, with the full-width D1 table above the stage details. Descriptions wrap; operation identifiers and timing values stay intact. Narrow screens can scroll the table without overflowing the page. Every executed D1 operation includes its duration, start, and finish relative to the browser request start. Cache status and D1 query count remain visible. Loading another page replaces the measurements; clearing the input removes the panel.
Each operation records explicit { startMs, endMs } timestamps using performance.timeOrigin + performance.now(). Duration is endMs - startMs; elapsed at completion is endMs - browserRequestStartMs, never a sum of stage durations. The HTTP body carries timings: { rpcMs, rpc, service }, where rpc is a recorded span or null. Worker and response-serialization measurements finish after the response JSON is constructed, so their endpoints travel as start and end parameters on the worker and serialize Server-Timing metrics, alongside their existing dur values. This preserves single-pass response serialization.
| Field or UI value | Start and stop |
|---|---|
Browser requestMs |
Immediately before fetch, through reading/parsing the response JSON. Excludes the 280 ms debounce, state updates, and rendering. |
Header worker / log totalWorkerMs |
Web handler entry through construction of the JSON Response. Excludes final timing-header formatting, structured logging, and network delivery. |
rpcMs |
Immediately before awaiting GEONAMES.searchPage, through its resolution or rejection. |
Header serialize / log serializationMs |
The response body’s single JSON.stringify; excludes response/header construction. |
serviceMs |
Service RPC entry through cache lookup, query work when needed, awaited KV write, and result assembly. |
kvReadMs, kvWriteMs |
Each awaited KV operation. The write span excludes preparation/serialization of the cache value. |
normalizationMs |
Sum of query-normalization calls in the cache wrapper and query engine, not normalization of fuzzy candidate aliases. |
firstFtsMs, secondFtsMs |
Sum the durations of their corresponding matchedPage passes across the whole scoped search, including repeated country probes and fallback work. The second-FTS total is zero/null when that branch does not run. |
ftsMs |
Sum of the FTS pass durations; it is not another query. |
fuzzyMs |
Sum of fuzzy work across the scoped search, including candidate D1 work and Worker-side ranking; an early return can be zero. |
fuzzyCandidateD1Ms |
Sum of executed candidate-fetch D1 operations, also present individually in d1. |
getManyD1Ms |
Sum of executed hydration batches, also present individually in d1. Includes Drizzle query execution/materialization, not subsequent location conversion. |
queryEngineMs |
Query-engine entry through the plain result page, including normalization and all executed search stages. Excludes KV. |
d1Queries, d1, d1Ms |
Executed operation count, ordered { name, label, reason, retrieves, outcome, durationMs, startMs, endMs } records, and summed active durations. Includes settings reads and the parent, scope, FTS, fuzzy and hydration operations actually run. SQL preparation alone is not another operation. |
Spans overlap; do not add service, engine, FTS, fuzzy and D1 totals. Nullable stage fields are null when a stage did not run; numeric aggregates are zero when no matching operation ran. A cache hit has a fresh collector with the settings read but no search-engine work. Timing metadata is attached after cache retrieval and never stored in KV. Namespace search:v6 isolates the scope/ranking cutover; TTL and read-cache policies are unchanged.
service.spans contains nullable service, kvRead, kvWrite, firstFts, secondFts, fuzzy, and queryEngine spans, plus a normalization array recording every normalization call. These endpoints are available directly through the Service Binding; the HTTP adapter also returns timings.rpc.
The D1 (window) and FTS (window) table rows run from the earliest start to the latest finish of their operations. Their durations include gaps and differ from the preserved d1Ms and ftsMs active-time sums in the details. The waterfall draws each actual query/pass separately, leaving gaps visible rather than presenting discontinuous work as one continuous bar.
Duration shows how long each operation took. Elapsed shows when it completed relative to the start of the request. Some operations are nested or overlap, so durations should not be added together.
Cloudflare Workers use a zero performance.timeOrigin; timestamps are performance.timeOrigin + performance.now(). Browser timestamps use the browser’s epoch-based origin. The clocks are not one shared monotonic clock, so skew can shift server positions relative to the browser request. The UI discloses this limitation and does not clamp endpoints to manufacture containment.
Empty HTTP queries, invalid requests, and RPC failures may not have service-stage measurements. rpc is null when no binding call ran and contains observed endpoints when a call resolves or rejects. Search scope and configuration participate in cache identity: query, page, explicit/default resolved scope, and settings version must not share cached results across differing scopes. Search cache namespace is search:v6.
Optional Better Stack forwarding is prepared but inactive until both its source token and ingest endpoint are configured as Worker secrets/variables. Console observations remain available; no credentials are committed. See observability setup.
Deployed Cloudflare Worker clocks advance only after I/O. CPU-only normalization or serialization can therefore report 0 ms; these elapsed-time spans are not CPU profiling or isolated database execution latency. Local timing precision can also produce zero for short spans. No performance conclusion follows from a single request.
Search KV cache
Section titled “Search KV cache”apps/service owns the read-through KV cache for searchPage; search uses the same cache path and slices the resulting page. get and getMany remain uncached D1 lookups. Search cache namespace is search:v6.
Keys include version, effective page limit, offset, and a digest of the normalized query, normalized options and current settings. Country/admin/parent options and default/fallback configurations cannot share each other’s pages. Parent resolution and default-country existence probes run only inside the cache loader, not on a HIT. New behavior is isolated from search:v5; do not manually delete the KV namespace.
Production nonempty pages have no expiration; development nonempty pages retain the 2,592,000-second (30-day) TTL. Empty pages expire after 3,600 seconds (1 hour) in both environments. Production is selected by the service’s SEARCH_CACHE_ENVIRONMENT binding; bun run dev selects development. Reads still request cacheTtl: 3600, with KV expiration taking precedence. Writes are awaited. HIT skips the search engine but still reads current D1 settings. A D1 page successfully written reports MISS. Missing KV, policy exclusions, hash/read/malformed-value failures or write failures report BYPASS; D1 failures propagate and are not cached. A hit does not refresh expiration. Existing search:v6 entries retain their previous expiration until they expire and are rewritten under this policy.
KV is eventually consistent. A write in one location may remain invisible elsewhere for up to the configured read cache TTL, including a cached miss. Repeating a request is not a global read-after-write guarantee. See Cloudflare’s KV reads and expiration precedence.
After a dataset rebuild or a change to ranking, normalization, result shape, or cache semantics that makes saved pages incompatible, bump the version in apps/service/src/search/cache.ts before deployment. This cleanly separates new reads from old entries; old development and empty entries expire naturally, while old nonempty production entries remain stored but are no longer read. Do not delete the namespace. A local reimport likewise needs a version bump or a local cache clear for immediate freshness.
Known limits
Section titled “Known limits”- Historical ranking note: earlier full-data observations placed some administrative names ahead of same-name capitals by population. The current ranking gives populated places (
featureClassP) highest feature priority; old measurements describe the prior ranking. - Single-word queries of two or three normalized characters search primary place names only: exact names, then prefixes. They do not search aliases, feature labels, or fuzzy suggestions.
muandmumcan find Mumbai; one-character queries retain full search. - Fuzzy lookup scores at most 400 candidates and returns at most five typo suggestions. Longer prefixes use FTS; short or space-spanning prefixes use bounded normalized-name/alias word scans to avoid broad FTS prefix expansion. Scoring considers primary names and aliases, with a multiword primary-name bonus. This is not an exhaustive nearest-name search.
- Hierarchy enrichment is shallow: the result exposes admin1/admin2 codes and resolved names where available, not an arbitrary-depth administrative tree. Missing reference names remain null.
- Search uses aliases from
allCountries.txt; it does not import the separate alternateNamesV2 dataset. - Local measurements are not proof that production latency targets are met. See benchmarking for the measured cohort and thresholds.
For the public response fields and endpoints, see the API reference. For source files and the five-source import, see data import.