Documentation
Search benchmarking
The search benchmark measures the demo HTTP endpoint and the service’s request-scoped timing data. It is a local comparison tool, not evidence that production latency targets are met. See search behavior and timing definitions, the API reference, and data import for the surrounding contracts and database setup.
Run the current local MISS/HIT benchmark
Section titled “Run the current local MISS/HIT benchmark”From the repository root, start the demo on the benchmark’s default loopback port:
bun run --cwd apps/web dev --host 127.0.0.1 --port 5174In another terminal, run:
bun run --cwd apps/service benchmark:search -- --samples 100 --mode local-cold --url http://127.0.0.1:5174 --out .import/benchmark-search.jsonThe fixed nine-query cohort uses limit 25, offset 0, and current persisted default-scope settings. Each round reads the settings to construct each exact search:v6 key, evicts that local key, then confirms a MISS followed by an identical HIT. The HIT must execute exactly one D1 operation, the search_settings read, and no search-engine D1 operations. Current settings are included in the key; settings changes isolate cached pages. This is KV-cold, not disk- or process-cold. The loopback-only mode disables remote bindings and does not reset SQLite, process/runtime caches or other keys, clear a namespace, or touch production KV. Keep local workloads and settings changes paused during measurement.
observe samples without eviction or mutation and groups actual returned cache states. It cannot manufacture an absent MISS or HIT cohort. The CLI defaults to the loopback URL and 100 samples per query; --help lists --url, --samples, --mode, and --out. The output path is relative to the process working directory and contains raw samples plus summaries.
Reading measurements
Section titled “Reading measurements”Client total is HTTP request through response JSON parsing, including network and Worker-boundary costs, but excluding UI rendering. Service duration includes the persisted settings read, KV reads, D1 search work on a miss, awaited KV writes, and result assembly. D1 query count is executed operations, not latency; operation durations are separately summed as active D1 time. A HIT skips search-engine queries but still reads current settings from D1. Service, D1, and client-total measurements overlap and must not be added together.
The CLI summarizes percentiles with nearest rank: sorted value at rank ceil(p*n). Small cohorts have coarse tail resolution; with 100 samples, p99 is the second-largest value. These summaries are not SLO guarantees.
Recorded local benchmark evidence
Section titled “Recorded local benchmark evidence”The tables and investigation observations in the dated sections below are historical baseline material, collected before geographic scoping and populated-place-first ranking. They are retained as recorded; they do not describe current engine behavior or current release performance.
Earlier 18-request smoke (2026-09-30)
Section titled “Earlier 18-request smoke (2026-09-30)”The saved apps/service/.import/geographic-cache-smoke.json records one exact-key local MISS and immediate HIT for each of the nine queries (18 requests total, limit 25, offset 0). This is smoke evidence only, not a benchmark cohort or latency percentile. The uncached bangalore, banglore, paris, airport, railway station, and lake observations individually exceeded the reference service (75 ms) and D1 (50 ms) targets; these single samples do not establish p95 or production SLO results. The recorded HITs each executed one search_settings D1 read and no search-engine D1 operations. Current running instructions are above; these measurements must not be presented as current benchmark results.
Recorded local baseline (pre-scoping and ranking)
Section titled “Recorded local baseline (pre-scoping and ranking)”This baseline predates geographic scoping and populated-place-first ranking. It has 100 samples per query and cache state, 1,800 requests total, limit 25, offset 0, collected with the populated local database. Times are milliseconds; D1 queries is an operation count. These are local observations, not current release measurements or production target evidence.
| Query | Cache | Total p50 | Total p95 | Total p99 | Service p95 | D1 p95 | D1 queries |
|---|---|---|---|---|---|---|---|
| bengaluru | MISS | 28.80 | 39.93 | 43.12 | 16 | 10 | 2 |
| bengaluru | HIT | 21.26 | 30.97 | 44.15 | 2 | 0 | 0 |
| bangalore | MISS | 295.14 | 330.39 | 360.92 | 307 | 251 | 5 |
| bangalore | HIT | 24.56 | 31.06 | 47.09 | 2 | 0 | 0 |
| banglore | MISS | 291.35 | 318.95 | 339.31 | 301 | 243 | 5 |
| banglore | HIT | 24.28 | 29.33 | 32.96 | 2 | 0 | 0 |
| new delhi | MISS | 29.01 | 34.80 | 49.50 | 14 | 9 | 1 |
| new delhi | HIT | 20.28 | 27.39 | 32.64 | 2 | 0 | 0 |
| new delh | MISS | 95.87 | 111.74 | 118.74 | 88 | 68 | 4 |
| new delh | HIT | 21.83 | 29.88 | 40.65 | 2 | 0 | 0 |
| paris | MISS | 25.45 | 34.94 | 47.25 | 9 | 5 | 1 |
| paris | HIT | 20.33 | 26.16 | 28.92 | 2 | 0 | 0 |
| airport | MISS | 159.98 | 180.50 | 187.76 | 161 | 154 | 2 |
| airport | HIT | 24.33 | 29.29 | 34.75 | 1 | 0 | 0 |
| railway station | MISS | 381.41 | 418.28 | 428.86 | 399 | 367 | 2 |
| railway station | HIT | 24.74 | 29.83 | 37.60 | 2 | 0 | 0 |
| lake | MISS | 25.18 | 29.84 | 34.61 | 8 | 4 | 1 |
| lake | HIT | 20.46 | 27.32 | 29.94 | 2 | 0 | 0 |
The requested local thresholds were D1 p95 below 50 ms and service p95 below 75 ms. Cold bangalore, banglore, new delh, airport, and railway station miss both; the other cold queries and all HITs meet them. No client-total or Service Binding RPC target was assigned. None of these local results establishes production target compliance.
Historical bottleneck notes
Section titled “Historical bottleneck notes”These measurements and recommendations describe the prior implementation, not current candidate retrieval. Operation percentiles are separate distributions; do not sum them.
bangaloreandbanglore:fuzzy_candidatesdominated, with operation p95 of 233 ms and 226 ms. The prior query selected broad FTS-prefix matches, joined full location rows, then sorted by population before limiting to 400. Non-D1 service work had per-sample median 53/54 ms, including alias parsing/edit-distance scoring and KV/gaps.new delh: four D1 operations; fuzzy-candidate p95 was 48 ms.airportandrailway station: FTS-operation p95 was 150 ms and 301 ms respectively, in location/feature unions, deduplication, enrichment, and relevance/population ordering. Railway station also ran fuzzy candidate retrieval (p95 68 ms).- Historical engine note: the old
parisandlakemeasurements used an exact-name shortcut for that first-page workload; they do not document current branch behavior.lakedid not measure an exhaustive lake-feature search. In the full dataset, more-populated Paris administrative names ranked before the capital; Mysore’s exact administrative name preceded the Mysuru city alias.
The prior .codex/test-logs/batch3-search-quality.log is user-run test evidence for that checkpoint, not a benchmark run. Production latency targets remain unverified by these local measurements.
For the manual UI and API checks, see the manual verification guide.