Skip to content

Add sao_paulo, the 10th benchmark split (#98): the second non-US point, at HIGH-confidence GT - #100

Merged
jonfroehlich merged 6 commits into
mainfrom
benchmark/sao_paulo
Aug 2, 2026
Merged

Add sao_paulo, the 10th benchmark split (#98): the second non-US point, at HIGH-confidence GT#100
jonfroehlich merged 6 commits into
mainfrom
benchmark/sao_paulo

Conversation

@jonfroehlich

@jonfroehlich jonfroehlich commented Aug 1, 2026

Copy link
Copy Markdown
Member

Closes #98 — the split is now fully integrated: GT, operating point, #55 correction, and the full 8-model challenger roster with null-recall, all landed 2026-08-01.

What this adds

The benchmark''s 10th split and second non-US point: four central districts of São Paulo, Brazil (Brás, Sé, Cambuci, Bom Retiro), 125 GSV panos from a 22,741-pano run, NBR 9050 design vocabulary. GT reviewed by jonf on 2026-08-01 at HIGH self-rated confidence.

Headline: the budapest recall collapse did not replicate. P 0.869 / R 0.626 unbiased (0.888 / 0.676 all-panos) — precision at budapest''s level, recall inside the US GSV band (paterson 0.650, gainesville 0.647). With this pair the non-US axis de-confounds: unfamiliar infrastructure alone did not break the model or the rubric; budapest''s doubt localizes to its named ambiguities (full-corner diagonal aprons, level seams — reviewer retrospective recorded on #74).

The ranking survives the second non-US vocabulary at nearly full width: RampNet F1 0.777, lead 0.323 over gemini-3.1-pro (0.454). Qwen-32B''s caution mechanism fires a fourth time (0.6 boxes/pano, P 0.506 / R 0.139) but for the first time lands at parity with 8B rather than inverting — caution is the invariant, the flip is its side effect. Pairwise union ceiling 0.776 ties paterson''s for lowest, but by the opposite mechanism: RampNet''s own 0.05-floor ceiling (0.861) sits above the union — misses fire sub-threshold, and 0.30 buys +14.6 R with no fusion.

Other findings, in the committed docs:

  • Largest threshold response of any trusted-GT split (0.55 → 0.30 = +14.6 R), concentrated mid/far where the reviewer''s mid-block capture geometry put 47% of the GT.
  • GT-completeness correction: spot-check the low-confidence incremental FPs from the operating-point curve #55 A-rate 12.5% — the lowest of any split (6 A / 38 B / 4 unsure): the incremental FPs are overwhelmingly genuine, so the 0.30 precision cost is a real OOD FP mechanism (white-painted curbs), not GT incompleteness. gainesville (35.3%) is the mirror image.
  • Parity gate: count arm formally FAILED (9.2% > 5%), diagnosed, exception ratified by the reviewer — same GSV resample mechanism as every GSV split (max 0.439 R, fourth split on that number), amplified by OOD scores hugging the threshold. Rows ≤ 0.38 unaffected; split held out of all pooled rows.

Checklist (docs/adding_a_benchmark_city.md)

Phase 1 — bundle (auto-labeler)

Phase 2–3 — ground truth

  • Reviewed with gt_gallery.py at model resolution
  • review_notes written (HIGH confidence; mid-block captures + white-painted curbs)
  • verdicts.json committed; score_validation.py run; unbiased column recorded (0.869 / 0.626)
  • Camera provenance: source: launch carries the tier (tier_of needs no change — tested)

Phase 4 — operating point

  • Low-floor extraction run (locally, RTX 3070; cache meta records device/model)
  • Parity gate: count arm FAILED (9.2% > 5%), displacement arm passed — diagnosed in docs/operating_point.md, exception ratified by the reviewer 2026-08-01
  • sweep / hist / gtbias / floor / distance re-run; derived tables committed
  • GT-completeness correction: spot-check the low-confidence incremental FPs from the operating-point curve #55 gallery tagged (jonf, 2026-08-01): A 6 / B 38 / unsure 4 — A-rate 12.5%, lowest of any split; tags committed; corrected at 0.30; tagcheck 48/48
  • op_cache/sao_paulo.json + refreshed op/*.csv committed
  • Both figures regenerated and visually inspected

Phase 5 — code

  • CITY_SPLITS / ALL_SPLITS (not US_SPLITS — pooled basis unchanged: 7 US, 859 panos); HELD_OUT reason + --include-sao-paulo flag
  • tier_of() recognises the rig; SERIES neutral ink + own dash
  • pytest -q green (498)

Phase 6 — docs + challengers

🤖 Generated with Claude Code (claude-fable-5)

jonfroehlich and others added 3 commits July 31, 2026 16:51
Phase 1 of docs/adding_a_benchmark_city.md. 125 panos sampled from a 22,741-pano
GSV run over 13.95 km2 of central Sao Paulo, Brazil.

WHAT THE SPLIT IS FOR (phase 0, recorded here because it is the thing that gets lost)

A second non-US split. budapest_district5 is currently n=1 for "does the rubric
transfer outside the US", and it is held out of every recommendation because its
reviewer rated their own pass low confidence -- so the one non-US data point the
benchmark has is also its least trusted. Sao Paulo gives that finding a second
observation.

It also de-confounds it. Budapest is non-US *and* Mapillary/GoPro Max; Sao Paulo is
non-US on GSV, the same imagery path as bend/paterson/gainesville. If the rubric
fights back here too, that is about the infrastructure rather than the rig.

Bonus, not the reason: a live Project Sidewalk deployment
(sidewalk-sao-paulo.cs.washington.edu) sits inside the footprint, so an agree-rate
comparison is possible later. It is Bras only and 8.2% audited, so that is thin.

NAMING

The split is `sao_paulo` but the footprint is four central districts -- Bras plus its
Centro-side neighbours Se, Cambuci and Bom Retiro -- not the 1,521 km2 municipality.
Named for the run dir and the manifest's run_name rather than the geometry; the
footprint is stated wherever the split is described.

DECISIONS THAT PHASE 5 WILL NEED

- Not pooled. US_SPLITS is the pooled basis and this is not a US city; a HELD_OUT
  reason gets added when the split is registered, after scoring.
- Tier is `gsv` and tier_of() needs no new branch: every pano is source `launch`,
  which it already maps. No SPLIT_IMAGERY_FALLBACK entry either -- camera_make and
  camera_model are empty, as they are on every GSV split, but source carries the tier.
- Parity in phase 4 will not be bit-exact. That is expected for GSV (the production
  path resamples through a 4096x2048 intermediate); the gate is 0.5 R, not equality.

WHAT IS IN THE BUNDLE

Strata are the canonical 5 top / 95 random / 25 empty, defaults unchanged, seed 0, and
the spatial pass did not shrink them. 100 of 125 panos carry >=1 operational detection
(251 labels at >=0.55), and the 25 empty panos are empty by construction.

The imagery is the freshest of any split: 105 of 125 panos (84%) were captured 2024-26,
against 80.8% across the full run. Native sizes are 16384x8192 (122) and 13312x6656 (3).

Export gated clean: 125 fetched, 0 failed, archive reconciles against records 1:1,
nothing decayed, exit 0.

imagery_manifest.json is written *before* the review rather than after, so it pins the
exact bytes the reviewer will see rather than the bytes that happen to survive to
publication. Digest 650bd9ba2498b156 over 125 panos, 2.01 GB.

Run provenance: area_hash 6ace56a2522ebd54408861e08077ea4e9e3dfb4b3715c401d91b04758acd9ba7,
streetlevel 0.12.10, detection_storage_floor 0.1, model rampnet-model / 08-21-2025.

NOT DONE

Phases 2-6. No verdicts.json, no scoring, no operating-point regeneration, and the
split is deliberately absent from every registry until it has GT -- registering it now
would add empty rows to tables that read as measurements.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
… confidence

Phases 2-3 of docs/adding_a_benchmark_city.md. Reviewer jonf, 2026-08-01, via
gt_gallery.py at model resolution. All 125 panos judged: 190 correct / 23 FP /
1 duplicate / 37 unsure detections, 91 missed marks (+43 unsure), zero panos
excluded from the recall pool.

Headline (unbiased, random+empty): P 0.869 (CI 0.811-0.911), R 0.626
(CI 0.563-0.684). All-panos: P 0.888 / R 0.676.

The second non-US point diverges from the first: budapest_district5 is
P 0.873 / R 0.503 at LOW reviewer confidence; sao_paulo lands at the same
precision level but with recall inside the ~0.63-0.65 band the US GSV splits
replicate (paterson 0.650, gainesville 0.647) - and at HIGH confidence on the
same rubric. On this point the non-US recall story is imagery/rig, not
infrastructure.

Phase 3 gates: empty stratum 20/25 attested clean, 5 panos hold 11 missed
ramps (annapolis-shaped negative check); 1 duplicate; abstention is the
benchmark's highest at 14.7% of detections and 32.1% of missed marks -
review_notes attribute it to mid-block captures forcing far-away judgments,
not rubric failure.

Two reviewer-flagged mechanisms recorded in review_notes: mid-block captures
(>10 panos, assessment at the far-away nearest intersection) and
white-painted curbs that rarely carry actual ramps despite crosswalks - a
candidate split-specific FP mechanism the model visibly fell for.

Suite: 498 passed with the verdicts in place.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…0, and a diagnosed parity failure

Phases 4-6 of docs/adding_a_benchmark_city.md, minus the #55 tagging pass
(48-item gallery generated, awaiting the human pass) and the challenger runs.

PARITY: sao_paulo formally FAILS the gate's count arm (23 of 251 record
detections have no cache peak at >=0.55; 9.2% against the 5% allowance). The
failure is diagnosed, in docs/operating_point.md next to the gate table:
displacement signature identical to every GSV split (98.2% within tol, max
0.439 R - the fourth split to land on that exact number); same-cell score
jitter has the same two-sided shape as bend/paterson/gainesville but ~40%
wider (sd 0.068 vs 0.043-0.058) and shifted +0.022 - the production tile
pyramid scores slightly higher than the bundle's one-step bilinear from
16384 px, and Sao Paulo's out-of-domain scores hug the threshold, so the same
jitter drops more peaks under 0.55. Strict positional decomposition: 39
unreproduced = 20 threshold-straddlers (at-site peaks 0.39-0.54) + 19 NMS
pair-merges (the (10,10) Chebyshev corner of min_distance=10). Input-rebuild
probe: box-pyramid/LANCZOS recover +0.03-0.07 and re-cross 0.55 in 1 of 3.
Rows at <=0.38 - including the recommended 0.30 - are unaffected, and the
split is held out of every pooled row regardless.

Sweep: 0.55 -> 0.30 buys +14.6 recall points (0.651 -> 0.797), the largest
threshold response of any trusted-GT split; F1-max sits at 0.31. The gain
concentrates mid/far (+0.194/+0.217 vs +0.101 near) - exactly where the
reviewer's mid-block capture geometry put 47% of the GT. Recall ceiling 0.861
vs deployed 0.651 = +0.210 recoverable, the gainesville mechanism
(under-confident misses, not silent ones); only 2 ramps are floor-lost.

Registration: CITY_SPLITS + HELD_OUT (non-US geography, not GT quality),
pool_of()/--include-sao-paulo so the pooled basis stays 7 US splits; third
neutral-ink dashed series + '‡ non-US, high confidence' footnotes in both
figures (regenerated and visually inspected); README both tables + prose
section; model_comparison coverage row with the challenger gap stated;
operating_point per-split / A-rate / storage-floor / distance rows with the
parity caveat inline.

Extraction ran locally (RTX 3070, ~7 min, cache meta records device+model);
the slurm launcher was not needed for 125 panos.

Suite: 498 passed.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
jonfroehlich and others added 3 commits August 1, 2026 07:55
…plit — its 0.30 precision cost is real

Second human pass done (jonf, 2026-08-01): all 48 incremental FPs in
[0.25, 0.55) tagged — A 6 / B 38 / unsure 4. tagcheck resolves 48/48.

The pairing is the finding: the split with the LARGEST threshold response
(+14.6 R at 0.30) has the LOWEST A-rate (12.5%, vs 13-35% elsewhere). Its
incremental false positives are overwhelmingly genuine, so sao_paulo's
precision penalty at 0.30 is the model mis-firing on real out-of-domain
distractors (the reviewer's white-painted-curb pattern), not GT
incompleteness. gainesville is the mirror image (35.3% A-rate: mostly the
GT's fault). Corrected at 0.30: P 0.803 -> 0.821 (band hi 0.828),
R 0.797 -> 0.801, F1 0.800 -> 0.811. No A tag sits within 2 R of an
already-detected ramp. Pooled US rows unchanged (split held out).

Parity-gate exception ratified by Jon 2026-08-01 (recorded on PR #100);
the diagnosis in docs/operating_point.md stands as the record.

Docs: A-rate table + spread now 12.5-35% over nine splits; per-split
corrected table row; coverage matrix #55 column; anchoring caveat;
README white-paint bullet carries the B-rate confirmation.

Suite: 498 passed.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…d Qwen-32B's caution fires a 4th time - without inverting

Same-day challenger + null-recall runs for the 10th split (GPU legs on klone
jobs 37998087-90, Geminis via Vertex ADC locally; all rows reproduce from the
committed cache, export_model_cache --verify IDENTICAL, 7 files).

RampNet F1 0.777 vs gemini-3.1-pro 0.454 - lead 0.323, near the top of the
0.12-0.34 range. The claim now rests on nine splits across three countries:
zero-shot generality bought the challengers nothing on NBR 9050
infrastructure; every challenger degrades OOD at least as much as RampNet.

Qwen-32B: the caution mechanism replicates a 4th time (0.6 boxes/pano - same
as gainesville - P 0.506 / R 0.139, challenger-best FP economy 38) but for
the FIRST time lands at parity with 8B (0.218 vs 0.219) instead of below it,
because 8B also degraded. Sharpens the three-inversion story: caution is the
invariant, the ranking flip is its side effect.

Complementarity (#35): pairwise union ceiling 0.776 TIES paterson's 0.777
for benchmark-lowest, but by the OPPOSITE mechanism - 63 ramps (22.4% of GT)
found by neither model, yet RampNet's own 0.05-floor ceiling is 0.861, ABOVE
the union: sao_paulo's misses fire sub-threshold (gainesville mechanism),
and the 0.30 operating point buys +14.6 R with no fusion.

Null recall: OWLv2 0.922 recall at 78.3 boxes/pano with a 0.688 null -
density, inside its 0.56-0.77 band; every sparse model sits at 0.01-0.06.

Doc: sao_paulo per-split section, coverage matrix row complete, null table
rows, scope-of-claim updated to nine splits / three countries.

Suite: 498 passed.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…lo in the second registry, and encode the ratified parity exception

Three defects found reviewing the sao_paulo split, none of which changed a
published number but two of which quietly cost coverage.

1. `recall_by_distance.csv` and `tp_origin_by_bin.csv` had been regenerated with
   `--cities sao_paulo`, and `_write_csv` truncates: 27 and 80 rows for the other
   splits were deleted. `docs/operating_point.md` still quotes per-band far-field
   numbers (bend 0.214, clovis 0.389, annapolis 0.490, paterson 0.523,
   gainesville 0.420) that this table was the committed source for. Re-ran both
   commands over the default split list; the restored rows are byte-identical to
   main's and sao_paulo is appended, so nothing was recomputed, only recovered.

2. `miss_decomposition.py` keeps a SECOND split registry — its own US_SPLITS /
   HELD_OUT / ALL_SPLITS — and sao_paulo was never added, so eight scripts that
   take their CLI defaults from it silently skipped the split. The measurable
   symptom: `export_model_cache.py --verify` checked 61 (model, split) pairs and
   now checks 68, i.e. the seven sao_paulo challenger files this PR commits were
   not covered by the verify pass that vouches for them. They verify IDENTICAL.
   sao_paulo joins HELD_OUT (not TIER — held-out splits print "-" there, same as
   budapest, so no non-US split can leak into a pooled tier row), and a new test
   asserts the two registries cover the same splits so the next split cannot
   repeat this.

3. `low_floor_sweep.py parity` exited 1 on a clean clone and its narrative still
   said sao_paulo "reproduces within tolerance ... no scoring outcome changes" —
   text keyed on the displacement arm alone, which is false for a count-arm
   failure and contradicts operating_point.md. The reviewer's ratification lived
   only in prose. Added PARITY_EXCEPTIONS, keyed and documented like HELD_OUT:
   the row still prints MISMATCH (ratified) and still prints the full diagnosis,
   but a ratified split no longer counts as a NEW divergence, so the exit status
   goes back to meaning "nothing regressed". The waiver is per-split and tested.

Also: aligned the four misindented pool_of continuation lines, and fixed two doc
nits — an unwrapped 150-char line, and "the largest US-protocol queue" describing
a Brazilian split two lines below budapest's larger one.

pytest 501 passed (was 498). `parity` PASSes with the exception stated;
`export_model_cache.py --verify` 68/68 identical.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@jonfroehlich
jonfroehlich merged commit dc7450e into main Aug 2, 2026
2 checks passed
@jonfroehlich
jonfroehlich deleted the benchmark/sao_paulo branch August 2, 2026 14:47
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Deferred: 10th benchmark split - Sao Paulo, Brazil (second non-US point)

1 participant