Add sao_paulo, the 10th benchmark split (#98): the second non-US point, at HIGH-confidence GT - #100
Merged
Merged
Conversation
Phase 1 of docs/adding_a_benchmark_city.md. 125 panos sampled from a 22,741-pano GSV run over 13.95 km2 of central Sao Paulo, Brazil. WHAT THE SPLIT IS FOR (phase 0, recorded here because it is the thing that gets lost) A second non-US split. budapest_district5 is currently n=1 for "does the rubric transfer outside the US", and it is held out of every recommendation because its reviewer rated their own pass low confidence -- so the one non-US data point the benchmark has is also its least trusted. Sao Paulo gives that finding a second observation. It also de-confounds it. Budapest is non-US *and* Mapillary/GoPro Max; Sao Paulo is non-US on GSV, the same imagery path as bend/paterson/gainesville. If the rubric fights back here too, that is about the infrastructure rather than the rig. Bonus, not the reason: a live Project Sidewalk deployment (sidewalk-sao-paulo.cs.washington.edu) sits inside the footprint, so an agree-rate comparison is possible later. It is Bras only and 8.2% audited, so that is thin. NAMING The split is `sao_paulo` but the footprint is four central districts -- Bras plus its Centro-side neighbours Se, Cambuci and Bom Retiro -- not the 1,521 km2 municipality. Named for the run dir and the manifest's run_name rather than the geometry; the footprint is stated wherever the split is described. DECISIONS THAT PHASE 5 WILL NEED - Not pooled. US_SPLITS is the pooled basis and this is not a US city; a HELD_OUT reason gets added when the split is registered, after scoring. - Tier is `gsv` and tier_of() needs no new branch: every pano is source `launch`, which it already maps. No SPLIT_IMAGERY_FALLBACK entry either -- camera_make and camera_model are empty, as they are on every GSV split, but source carries the tier. - Parity in phase 4 will not be bit-exact. That is expected for GSV (the production path resamples through a 4096x2048 intermediate); the gate is 0.5 R, not equality. WHAT IS IN THE BUNDLE Strata are the canonical 5 top / 95 random / 25 empty, defaults unchanged, seed 0, and the spatial pass did not shrink them. 100 of 125 panos carry >=1 operational detection (251 labels at >=0.55), and the 25 empty panos are empty by construction. The imagery is the freshest of any split: 105 of 125 panos (84%) were captured 2024-26, against 80.8% across the full run. Native sizes are 16384x8192 (122) and 13312x6656 (3). Export gated clean: 125 fetched, 0 failed, archive reconciles against records 1:1, nothing decayed, exit 0. imagery_manifest.json is written *before* the review rather than after, so it pins the exact bytes the reviewer will see rather than the bytes that happen to survive to publication. Digest 650bd9ba2498b156 over 125 panos, 2.01 GB. Run provenance: area_hash 6ace56a2522ebd54408861e08077ea4e9e3dfb4b3715c401d91b04758acd9ba7, streetlevel 0.12.10, detection_storage_floor 0.1, model rampnet-model / 08-21-2025. NOT DONE Phases 2-6. No verdicts.json, no scoring, no operating-point regeneration, and the split is deliberately absent from every registry until it has GT -- registering it now would add empty rows to tables that read as measurements. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
… confidence Phases 2-3 of docs/adding_a_benchmark_city.md. Reviewer jonf, 2026-08-01, via gt_gallery.py at model resolution. All 125 panos judged: 190 correct / 23 FP / 1 duplicate / 37 unsure detections, 91 missed marks (+43 unsure), zero panos excluded from the recall pool. Headline (unbiased, random+empty): P 0.869 (CI 0.811-0.911), R 0.626 (CI 0.563-0.684). All-panos: P 0.888 / R 0.676. The second non-US point diverges from the first: budapest_district5 is P 0.873 / R 0.503 at LOW reviewer confidence; sao_paulo lands at the same precision level but with recall inside the ~0.63-0.65 band the US GSV splits replicate (paterson 0.650, gainesville 0.647) - and at HIGH confidence on the same rubric. On this point the non-US recall story is imagery/rig, not infrastructure. Phase 3 gates: empty stratum 20/25 attested clean, 5 panos hold 11 missed ramps (annapolis-shaped negative check); 1 duplicate; abstention is the benchmark's highest at 14.7% of detections and 32.1% of missed marks - review_notes attribute it to mid-block captures forcing far-away judgments, not rubric failure. Two reviewer-flagged mechanisms recorded in review_notes: mid-block captures (>10 panos, assessment at the far-away nearest intersection) and white-painted curbs that rarely carry actual ramps despite crosswalks - a candidate split-specific FP mechanism the model visibly fell for. Suite: 498 passed with the verdicts in place. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…0, and a diagnosed parity failure Phases 4-6 of docs/adding_a_benchmark_city.md, minus the #55 tagging pass (48-item gallery generated, awaiting the human pass) and the challenger runs. PARITY: sao_paulo formally FAILS the gate's count arm (23 of 251 record detections have no cache peak at >=0.55; 9.2% against the 5% allowance). The failure is diagnosed, in docs/operating_point.md next to the gate table: displacement signature identical to every GSV split (98.2% within tol, max 0.439 R - the fourth split to land on that exact number); same-cell score jitter has the same two-sided shape as bend/paterson/gainesville but ~40% wider (sd 0.068 vs 0.043-0.058) and shifted +0.022 - the production tile pyramid scores slightly higher than the bundle's one-step bilinear from 16384 px, and Sao Paulo's out-of-domain scores hug the threshold, so the same jitter drops more peaks under 0.55. Strict positional decomposition: 39 unreproduced = 20 threshold-straddlers (at-site peaks 0.39-0.54) + 19 NMS pair-merges (the (10,10) Chebyshev corner of min_distance=10). Input-rebuild probe: box-pyramid/LANCZOS recover +0.03-0.07 and re-cross 0.55 in 1 of 3. Rows at <=0.38 - including the recommended 0.30 - are unaffected, and the split is held out of every pooled row regardless. Sweep: 0.55 -> 0.30 buys +14.6 recall points (0.651 -> 0.797), the largest threshold response of any trusted-GT split; F1-max sits at 0.31. The gain concentrates mid/far (+0.194/+0.217 vs +0.101 near) - exactly where the reviewer's mid-block capture geometry put 47% of the GT. Recall ceiling 0.861 vs deployed 0.651 = +0.210 recoverable, the gainesville mechanism (under-confident misses, not silent ones); only 2 ramps are floor-lost. Registration: CITY_SPLITS + HELD_OUT (non-US geography, not GT quality), pool_of()/--include-sao-paulo so the pooled basis stays 7 US splits; third neutral-ink dashed series + '‡ non-US, high confidence' footnotes in both figures (regenerated and visually inspected); README both tables + prose section; model_comparison coverage row with the challenger gap stated; operating_point per-split / A-rate / storage-floor / distance rows with the parity caveat inline. Extraction ran locally (RTX 3070, ~7 min, cache meta records device+model); the slurm launcher was not needed for 125 panos. Suite: 498 passed. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…plit — its 0.30 precision cost is real Second human pass done (jonf, 2026-08-01): all 48 incremental FPs in [0.25, 0.55) tagged — A 6 / B 38 / unsure 4. tagcheck resolves 48/48. The pairing is the finding: the split with the LARGEST threshold response (+14.6 R at 0.30) has the LOWEST A-rate (12.5%, vs 13-35% elsewhere). Its incremental false positives are overwhelmingly genuine, so sao_paulo's precision penalty at 0.30 is the model mis-firing on real out-of-domain distractors (the reviewer's white-painted-curb pattern), not GT incompleteness. gainesville is the mirror image (35.3% A-rate: mostly the GT's fault). Corrected at 0.30: P 0.803 -> 0.821 (band hi 0.828), R 0.797 -> 0.801, F1 0.800 -> 0.811. No A tag sits within 2 R of an already-detected ramp. Pooled US rows unchanged (split held out). Parity-gate exception ratified by Jon 2026-08-01 (recorded on PR #100); the diagnosis in docs/operating_point.md stands as the record. Docs: A-rate table + spread now 12.5-35% over nine splits; per-split corrected table row; coverage matrix #55 column; anchoring caveat; README white-paint bullet carries the B-rate confirmation. Suite: 498 passed. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…d Qwen-32B's caution fires a 4th time - without inverting Same-day challenger + null-recall runs for the 10th split (GPU legs on klone jobs 37998087-90, Geminis via Vertex ADC locally; all rows reproduce from the committed cache, export_model_cache --verify IDENTICAL, 7 files). RampNet F1 0.777 vs gemini-3.1-pro 0.454 - lead 0.323, near the top of the 0.12-0.34 range. The claim now rests on nine splits across three countries: zero-shot generality bought the challengers nothing on NBR 9050 infrastructure; every challenger degrades OOD at least as much as RampNet. Qwen-32B: the caution mechanism replicates a 4th time (0.6 boxes/pano - same as gainesville - P 0.506 / R 0.139, challenger-best FP economy 38) but for the FIRST time lands at parity with 8B (0.218 vs 0.219) instead of below it, because 8B also degraded. Sharpens the three-inversion story: caution is the invariant, the ranking flip is its side effect. Complementarity (#35): pairwise union ceiling 0.776 TIES paterson's 0.777 for benchmark-lowest, but by the OPPOSITE mechanism - 63 ramps (22.4% of GT) found by neither model, yet RampNet's own 0.05-floor ceiling is 0.861, ABOVE the union: sao_paulo's misses fire sub-threshold (gainesville mechanism), and the 0.30 operating point buys +14.6 R with no fusion. Null recall: OWLv2 0.922 recall at 78.3 boxes/pano with a 0.688 null - density, inside its 0.56-0.77 band; every sparse model sits at 0.01-0.06. Doc: sao_paulo per-split section, coverage matrix row complete, null table rows, scope-of-claim updated to nine splits / three countries. Suite: 498 passed. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…lo in the second registry, and encode the ratified parity exception Three defects found reviewing the sao_paulo split, none of which changed a published number but two of which quietly cost coverage. 1. `recall_by_distance.csv` and `tp_origin_by_bin.csv` had been regenerated with `--cities sao_paulo`, and `_write_csv` truncates: 27 and 80 rows for the other splits were deleted. `docs/operating_point.md` still quotes per-band far-field numbers (bend 0.214, clovis 0.389, annapolis 0.490, paterson 0.523, gainesville 0.420) that this table was the committed source for. Re-ran both commands over the default split list; the restored rows are byte-identical to main's and sao_paulo is appended, so nothing was recomputed, only recovered. 2. `miss_decomposition.py` keeps a SECOND split registry — its own US_SPLITS / HELD_OUT / ALL_SPLITS — and sao_paulo was never added, so eight scripts that take their CLI defaults from it silently skipped the split. The measurable symptom: `export_model_cache.py --verify` checked 61 (model, split) pairs and now checks 68, i.e. the seven sao_paulo challenger files this PR commits were not covered by the verify pass that vouches for them. They verify IDENTICAL. sao_paulo joins HELD_OUT (not TIER — held-out splits print "-" there, same as budapest, so no non-US split can leak into a pooled tier row), and a new test asserts the two registries cover the same splits so the next split cannot repeat this. 3. `low_floor_sweep.py parity` exited 1 on a clean clone and its narrative still said sao_paulo "reproduces within tolerance ... no scoring outcome changes" — text keyed on the displacement arm alone, which is false for a count-arm failure and contradicts operating_point.md. The reviewer's ratification lived only in prose. Added PARITY_EXCEPTIONS, keyed and documented like HELD_OUT: the row still prints MISMATCH (ratified) and still prints the full diagnosis, but a ratified split no longer counts as a NEW divergence, so the exit status goes back to meaning "nothing regressed". The waiver is per-split and tested. Also: aligned the four misindented pool_of continuation lines, and fixed two doc nits — an unwrapped 150-char line, and "the largest US-protocol queue" describing a Brazilian split two lines below budapest's larger one. pytest 501 passed (was 498). `parity` PASSes with the exception stated; `export_model_cache.py --verify` 68/68 identical. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #98 — the split is now fully integrated: GT, operating point, #55 correction, and the full 8-model challenger roster with null-recall, all landed 2026-08-01.
What this adds
The benchmark''s 10th split and second non-US point: four central districts of São Paulo, Brazil (Brás, Sé, Cambuci, Bom Retiro), 125 GSV panos from a 22,741-pano run, NBR 9050 design vocabulary. GT reviewed by jonf on 2026-08-01 at HIGH self-rated confidence.
Headline: the budapest recall collapse did not replicate. P 0.869 / R 0.626 unbiased (0.888 / 0.676 all-panos) — precision at budapest''s level, recall inside the US GSV band (paterson 0.650, gainesville 0.647). With this pair the non-US axis de-confounds: unfamiliar infrastructure alone did not break the model or the rubric; budapest''s doubt localizes to its named ambiguities (full-corner diagonal aprons, level seams — reviewer retrospective recorded on #74).
The ranking survives the second non-US vocabulary at nearly full width: RampNet F1 0.777, lead 0.323 over gemini-3.1-pro (0.454). Qwen-32B''s caution mechanism fires a fourth time (0.6 boxes/pano, P 0.506 / R 0.139) but for the first time lands at parity with 8B rather than inverting — caution is the invariant, the flip is its side effect. Pairwise union ceiling 0.776 ties paterson''s for lowest, but by the opposite mechanism: RampNet''s own 0.05-floor ceiling (0.861) sits above the union — misses fire sub-threshold, and 0.30 buys +14.6 R with no fusion.
Other findings, in the committed docs:
Checklist (docs/adding_a_benchmark_city.md)
Phase 1 — bundle (auto-labeler)
main.py --source gsv)export_benchmark.py --bundleexited zero;index.csvwrittenrecords.jsonlcommitted;panos/local + klone (/gscratch/makelab/jfroehli/rampnet_benchmark_panos/sao_paulo) + pinned byimagery_manifest.json(HF archive lags — Publish deployment validation ground-truth as a HuggingFace dataset (Bend GSV + Richmond Mapillary) #21, stated in README)Phase 2–3 — ground truth
gt_gallery.pyat model resolutionreview_noteswritten (HIGH confidence; mid-block captures + white-painted curbs)verdicts.jsoncommitted;score_validation.pyrun; unbiased column recorded (0.869 / 0.626)source: launchcarries the tier (tier_ofneeds no change — tested)Phase 4 — operating point
docs/operating_point.md, exception ratified by the reviewer 2026-08-01sweep/hist/gtbias/floor/distancere-run; derived tables committedcorrectedat 0.30;tagcheck48/48op_cache/sao_paulo.json+ refreshedop/*.csvcommittedPhase 5 — code
CITY_SPLITS/ALL_SPLITS(notUS_SPLITS— pooled basis unchanged: 7 US, 859 panos);HELD_OUTreason +--include-sao-pauloflagtier_of()recognises the rig;SERIESneutral ink + own dashpytest -qgreen (498)Phase 6 — docs + challengers
benchmark/README.md: both tables + prose section (incl. the GT-completeness correction: spot-check the low-confidence incremental FPs from the operating-point curve #55 B-rate confirmation)benchmark/model_detections/and--verify''d IDENTICALdocs/operating_point.md: parity diagnosis, per-split, A-rate, corrected, storage-floor, distance🤖 Generated with Claude Code (claude-fable-5)