Skip to content

Assess Denver, and freeze the Stage 1 inventories for replication (#96, #59) - #97

Open
jonfroehlich wants to merge 22 commits into
mainfrom
feat/location-precision-96
Open

Assess Denver, and freeze the Stage 1 inventories for replication (#96, #59)#97
jonfroehlich wants to merge 22 commits into
mainfrom
feat/location-precision-96

Conversation

@jonfroehlich

@jonfroehlich jonfroehlich commented Jul 31, 2026

Copy link
Copy Markdown
Member

First cities through the location-precision gate filed as #96, plus the frozen Stage 1 inputs the
ICCV paper should have shipped with. Analysis and tooling only — no pipeline change, no retrain, no
city sourced. Full suite 516 passed (was 304; +212 tests).

Denver is assessed and passes. Seattle is partially assessed and produced a methodological
correction instead of a verdict.
Four things previously asserted in this repo are contradicted
below; every correction is recorded in place rather than quietly patched.


1. Denver is Good — median 0.29 m, 92.3% within 1 m

All 59 chips reviewed by Jon in one sitting, against the §5e rubric. Wilson intervals throughout,
because at n≈55 with a rate near 5% the normal approximation runs below zero.

value 95% CI
Offset median 0.29 m
p75 / p90 / max 0.64 / 0.95 / 2.07 m
Within 1 m ("on the ramp") 92.3% (48/52) 81.8–97.0
Within 0.5 m 63.5% (33/52) 49.9–75.2
Phantom (readable corner, no ramp) 5.5% (3/55) 1.9–14.9
Unjudgeable 6.8% (4/59) 2.7–16.2

The threshold is stated rather than assumed, since the paper published none: Good = ≥90% within
1 m, phantom under 10%
. A ramp is 1.2–1.8 m deep, so "within 1 m of centre" is very nearly "on the
concrete". This is our threshold, not the paper's — Denver's Table 1 peers cannot be re-scored
against it without repeating the exercise.

It also settles §5d for Denver. Zero chips where the reviewer saw more ramps than Denver
publishes; 1.37 ramps seen per corner against 1.21 records per corner. Denver's low pairing ratio is
ramp design vocabulary, not missing records — the benign reading, now confirmed on imagery.

Denver passing moves the usable pool from 276,615 to 349,385 — the single largest available
addition. It does not reach 500,000, and §6 shows nothing available does.

2. ⚠️ Denver is mid-pack, not weak, on per-ramp vs per-corner — I benchmarked against one city

inventory_geometry.py. NYC calibrates it: it publishes rampid and cornerid, so 6 m
single-link clustering can be scored against the publisher's own grouping — P 0.976 / R 0.973.
That is what licenses running it on cities with no corner key, and every remaining city gets it free
before any reviewer time.

I first called Denver "weak on pairing" by comparing it to NYC alone. NYC is the outlier:

City Records Rec/corner (6 m) vs published corner key
NYC — in training 217,679 1.61 P .976 / R .973
Portland — in training 46,101 1.38
Seattle 38,364 1.34
Bend — in training 14,805 1.31
Sioux Falls 19,991 1.22
Denver 72,770 1.21
Arlington 10,342 1.12
Charlotte 40,600 1.09 P .621 / R .406 ⚠️
Minneapolis 18,453 1.09 P .924 / R .859
Boston 24,022 1.08
San Francisco 50,096 6.64 ⚠️ broken

The real result is about the corpus, not the city: all three training cities sit above every
candidate
, so any addition dilutes pairing density. And the mechanism is mostly benign — Sioux
Falls' Diagonal-only corners hold 1.004 records against Directional-only 1.365, so a low
ratio is largely design vocabulary, which is the diversity §8 asks for.

Never benchmark a candidate against one reference city. That was the error.

3. ⚠️ Seattle's "systematic shift" is not a registration error — and the statistic needed a null

Seattle's first eleven measurable chips read median 2.33 m with a mean offset vector of 2.06 m
against a mean magnitude of 2.37 m — an 87% systematic share. I reported that as a registration
error. The direction is real; the conclusion was not.

The share statistic had no null. It was documented as "~0 for noise", but under random headings
the mean vector shrinks only as 1/sqrt(n), so the expected share is ~0.9/sqrt(n):

city n share null median p
Seattle 11 87% 39% 0.0013
Denver 52 24% 16% 0.221

Seattle's lean survives its null — and Denver's 24% never meant anything, being inside its own.

A direction still doesn't say which side is wrong, so measure the other two legs of a triangle
that must close:

(ramps vs imagery)  =  (ramps vs centrelines)  +  (centrelines vs imagery)
 review sheet, n=11      n≈31,000                   n≈2,500

Leg 1 — inventory_centerline_offset.py (new). A ramp sits half a roadway from the centreline on
one side or the other, so median(east side) + median(west side) = twice the shift while the unknown
half-width cancels and returns as a sanity check.

resultant shift half-width E / N samples
Seattle 0.00 m 6.24 / 6.26 m 31,430 / 30,809
Denver (control) 0.12 m 7.35 / 7.46 m 49,594 / 49,404

Denver is the calibration and it lands: 0.12 m here against 0.10 m from its 52 reviewer clicks,
measured independently. Flat across a 10°–30° sweep of the cardinal cutoff. Seattle's coordinates
are unbiased against Seattle's own street network — tighter than Denver's.

Leg 2 — verify_chip_georeference.py --sites-from-verdicts (new). Orthorectification error is
local, so a five-neighbourhood average cannot clear the places the verdicts came from. Measured
under the review itself, all eleven chips clear at ≤0.32 m, mean east +0.02 m — including the
8.79 m click, at 0.11 m.

The triangle predicts 0.0–0.3 m and the sheet says 2.06 m. It does not close, and that is the
result.
No city-wide displacement can hide from 31,430 samples or a per-chip imagery check. The
sheet's own quartiles agree once you stop averaging: p25 is 0.47 m against a 2.33 m median, and a
uniform shift cannot produce a chip that is dead-on. What made 87% was a handful of large
per-record errors sharing a heading.

Consequently: the SHIFT framing is withdrawn; there is no constant to subtract, which was the
optimistic branch; Seattle's Poor rating is neither confirmed nor overturned, because eleven
chips on a basemap already declared inadequate for grading settle nothing; and the
62.9%-installed-after-2019 confound is now the leading candidate.

The durable fix: inventory_review_summary.py prints the null and p-value, refuses to call a shift
on its own
, and names the two instruments that can answer it.

4. ⚠️ The obvious basemap is not good enough, and it fails silently

inventory_review_sheet.py renders the §5 positional instrument. The first version used Esri World
Imagery, which over Denver is leaf-on, hazy and effectively ~1 m, and at z=21 serves "Map data not
yet available"
as a flat grey tile the fetcher pasted in as evidence. Denver's own
Aerial2016_tilecache is leaf-off and 0.057 m/px.

probe_basemap.py exists because the lesson kept recurring in new forms — King County advertises
maxLOD 23 and 404s above z20, so the deepest level must be probed, never read. Every King County
year is leaf-on at 27–41% vegetation cover against Denver's 7.2%, which is why Seattle's sheet
carries an explicit "adequate to size a large error, not to grade a Good city" caveat.

Every city needs its municipal basemap located and probed before its sheet is worth reviewer time.
Charlotte's is found and is good (leaf-off, 0.061 m/px); its sheet is built and unreviewed.

5. ⚠️ Two disqualifications and a supply correction, found by reading data — no imagery

  • San Francisco is disqualified for Stage 1: 50,096 rows but only 7,553 distinct coordinates,
    1:1 with cnn (intersection node id) ⇒ coordinates are intersection centroids, ~4.7 identical
    labels per pixel. Separately, 14,414 rows are crexist=0 — confirmed absence, wrong polarity for
    Stage 1 and right for RampNet 2.0: measurement, condition inference, and tagging — we already collect the supervision and discard it at ingest #86.
  • Charlotte's coordinates disagree with its own corner key (P .621 / R .406): of its multi-record
    published corners, 52.5% span >6 m, 14.4% >20 m, p99 308 m. Its sheet must test this.
  • Charlotte has 5,505 RP_Type=NoRamp rows ⇒ usable 35,095, not 40,601.

Unassessed pool is 180,673, not ~236,000 (−55,327). So Good (276,615) + every unassessed city,
all passing, = 457,288 < 500,000
: the route avoiding both the OK tier and the state DOTs is
closed.

Also fixed a silent truncation found the same way: SF names its OID field ObjectId, so ID-paging
stopped at 2,000 of 50,096 and wrote a normal-looking snapshot. find_oid_field now reads it from
the layer, and a short fetch refuses to write without --allow-partial.

6. ⚠️ Denver's "2022 imagery" claim does not survive contact with the data

96.2% of records are UPDATE_STATUS=NC, 74.4% carry a 2015 CREATEDATE, only 84 carry 2022, and
Modify is used zero times across 72,770 records. Either everything was re-verified against 2022
imagery or a 2015 delineation is being carried forward; zero M favours the second, which is a ~7-year
gap comparable to what disqualified DC. §5c's ✅ for Denver is now marked unconfirmed. Nothing in
the schema records removals either, so phantoms have no upper bound from this data.

Footprint passes: 85.2% strictly inside Denver County, 98.8% of the rest within 1 km, 128 records
(0.18%) beyond 2 km — the mountain parks.

7. Frozen Stage 1 inputs — data/inventories/

The three paper inventories were never in this repo; the README points at live portals serving current
data. Eleven snapshots are now frozen, each with endpoint, exact query, fetch date,
declared-vs-retained count and a sha256.

Snapshot Records vs paper Tab. 1
nyc-ny 217,679 −0.0005% — effectively frozen
portland-or 46,101 +1.7% drift
bend-or 14,805 +8.8% drift
+ Denver, Seattle, SF, Charlotte, Boston, Sioux Falls, Minneapolis, Arlington candidates
+ seattle-wa-centerlines (34,484), denver-co-centerlines (7,866) reference geometry, not supply

Two gzip header fields are pinned (mtime=0, filename="") so identical records hash identically — a
real bug the test caught, since otherwise the digest tracked where the file was written.

Centrelines are frozen on the same terms because an analysis that exonerates a city's coordinates must
not depend on a live endpoint for the reference it exonerated them against. Each city's centrelines
must come from the same publisher and CRS as its ramps
— Seattle's SND and curb ramps are both
EPSG:2926 from org ZOyb2t4B0UYuYNYH via the same server, which is what makes a shared datum fault
cancel and a ramp-layer defect show.

Two limits stated rather than implied: these are not the paper's files (Portland and Bend have
drifted, so a re-run reproduces today's dataset, not ICCV's — recovering the paper-exact files from the
supplemental stays open), and basemap imagery is not redistributable, so verdicts.json records tile
URLs and keys instead of shipping tiles.

What this does not settle

  • Recall is structurally invisible to this instrument — it samples published records, so it can
    only find records with no ramp, never ramps with no record. §5e says so explicitly.
  • One rater, no second pass. The benchmark's standing top follow-up applies here too.
  • Denver's phantom interval reaches 14.9%, which across 72,770 records is the difference between
    ~4,000 and ~10,800 false labels. Narrowing it needs a larger sample, not a better instrument.
  • Our Good/OK/Poor scale is still a scale of one. NYC is the remaining route to placing it on the
    paper's, and its imagery service is dynamic-only — the /export fetcher is not built.

Review

  • docs/curb_ramp_data_sourcing.md §5d–§5i and §9 carry the numbers and caveats.
  • data/inventories/README.md documents the snapshot format and how to add a city.
  • Everything is regenerable; the review sheet is not committed because it embeds municipal tiles.

🤖 Generated with Claude Code (claude-opus-5[1m])

#96, #59)

Starts the location-precision gate filed as #96. Denver first: it is the largest
city inventory found (72,770) and the one where the temporal and positional gates
pull in opposite directions.

Three results, one of which contradicts an earlier grading in this repo.

1. Per-ramp vs per-corner no longer needs a reviewer, and Denver is weak on it.
   NYC publishes rampid AND cornerid, so it is ground truth: 6 m single-link
   clustering reproduces NYC's own corner grouping at P 0.976 / R 0.973. Against
   that reference, Denver records 1.21 points per corner where NYC records 1.61,
   with 79.4% singleton corners against NYC's 40.2%. A link sweep rules out the
   obvious confound -- NYC plateaus at 1.53-1.64 across 4-8 m while its groups
   stay resolved, Denver never plateaus and only reaches NYC's ratio once its
   groups have already merged across corners. Geometry cannot say whether Denver
   *records* one point per corner or *has* single diagonal ramps, so the review
   sheet requires a ramps_visible count.

2. Denver's "delineated from 2022 aerial imagery" does not survive the data.
   96.2% of records are UPDATE_STATUS=NC, 74.4% carry a 2015 CREATEDATE, only 84
   carry 2022, and the Modify code is used zero times across 72,770 records. The
   existence bound is either 2022 or 2015 depending on what NC means, and 2015
   against median 2022 GSV capture is comparable to the gap that disqualified DC.
   The 5c grading of Denver as near-contemporaneous is now marked unconfirmed.

3. The obvious basemap is not good enough, and failed silently. Esri World
   Imagery renders Denver leaf-on and upsampled to ~1 m, and serves "Map data not
   yet available" as a flat grey tile that the first version pasted into the sheet
   as evidence. Denver's own Aerial2018 tilecache is leaf-off at 0.23 m/px. Every
   city now needs its municipal basemap located first; blank tiles are detected.

Also acts on the replication rule for Stage 1 inputs. The three paper inventories
were never in this repo -- the README points at live portals -- so NYC, Portland
and Bend are now frozen alongside Denver with endpoint, query, count and a sha256.
Portland has drifted +1.7% and Bend +8.8% since the paper, so these reproduce
today's dataset, not the ICCV one; recovering the paper-exact files from the
supplemental material stays open and is called out as such.

No verdict has been recorded for Denver. verdicts.json is an unfilled template and
the Good/OK/Poor question is open pending the visual pass.

Full suite 358 passed (was 304).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…ver claim (#96)

Runs the per-corner check over every frozen inventory. It overturns the
single-baseline framing in the previous commit and removes 55k ramps from the
supply arithmetic. Suite 376 passed (was 358).

**Denver is mid-pack, not weak.** Ordered by records/corner at a 6 m link: NYC
1.61, Portland 1.38, Bend 1.31 -- then Sioux Falls 1.22, Denver 1.21, Arlington
1.12, Charlotte 1.09, Minneapolis 1.09, Boston 1.08. The previous commit compared
Denver against NYC alone and called it weak on pairing; NYC is the outlier. The
real result is about the corpus: all three training cities sit above every
candidate, so any of them dilutes pairing density relative to what we train on
now. Minneapolis's own published corner key confirms 1.09 is real (P .924 /
R .859), so this is not a clustering artifact.

The mechanism comes from the two cities that publish ramp type. In Sioux Falls,
Diagonal-only corners hold 1.004 records and Directional-only hold 1.365;
Charlotte agrees from the other side, 1.059 records on corners with a diagonal
type against 1.176 without. A low ratio is largely design vocabulary, not
under-recording -- the benign reading of Denver's 1.21, though Denver publishes
no type field, which is what ramps_visible on the sheet is for.

Two anomalies, both caught by reading data rather than imagery:

- San Francisco is disqualified for Stage 1. Its 50,096 records carry 7,553
  distinct coordinates, 1:1 with the cnn intersection id -- the coordinates are
  intersection centroids, so Stage 1 would project ~4.7 identical labels onto one
  pixel. 14,414 rows are crexist=0, confirmed absence, wrong polarity anyway.
- Charlotte's coordinates disagree with Charlotte's own corner key (P .621 /
  R .406). Of its multi-record published corners, 52.5% span more than 6 m, 14.4%
  more than 20 m, p99 308 m. Large arterial corners do not explain 308 m.

Supply correction: the unassessed pool is 180,673, not ~236,000 (SF to zero,
Charlotte less its 5,505 NoRamp records). Consequence is sharper than the number
-- Good plus every unassessed city, all passing, is 457,288, still short of
500,000. The route that avoided both the OK tier and the state DOTs is closed.

Also fixes a silent truncation in the fetcher, found the same way. SF names its
object-ID field ObjectId, so ID paging stopped after one page and wrote 2,000 of
50,096 as a normal-looking snapshot. The field is now read from the layer, and a
short fetch refuses to write without --allow-partial. Minneapolis legitimately
needs it: 4 rows carry null geometry, checked individually.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@jonfroehlich

Copy link
Copy Markdown
Member Author

Correction: Denver is mid-pack, not weak — and the unassessed supply is 23% smaller than counted.

Ran the geometry gate over all nine frozen inventories (pushed to PR #97). Two things I said in the
comment above are now wrong, and one of them is my error rather than new data.

⚠️ The Denver framing was wrong

I compared Denver against NYC alone. Against the full set, NYC is the outlier:

City Records Rec/corner (6 m) Singleton vs published corner key
NYC — Good, in training 217,679 1.61 0.402 P .976 / R .973
Portland — Good, in training 46,101 1.38 0.635
Bend — Good, in training 14,805 1.31 0.726
Sioux Falls 19,991 1.22 0.785
Denver 72,770 1.21 0.794
Arlington 10,342 1.12 0.887
Charlotte 40,600 1.09 0.917 P .621 / R .406 ⚠️
Minneapolis 18,453 1.09 0.917 P .924 / R .859
Boston 24,022 1.08 0.924
San Francisco 50,096 6.64 0.038 ⚠️ see below

Denver is second-highest of the candidates. The actual result is about the corpus, not the city:
all three training cities sit above every candidate, so any addition dilutes pairing density
relative to what RampNet trains on today. Minneapolis's own published corner key confirms 1.09 is
real, so this is not a clustering artifact.

And the mechanism is mostly benign. The two cities publishing ramp type let it be decomposed:
Sioux Falls' Diagonal-only corners hold 1.004 records, Directional-only hold 1.365;
Charlotte agrees from the other side (1.059 with a diagonal type vs 1.176 without). A low ratio is
largely ramp-design vocabulary, not under-recording — which is the diversity §8 asks for, and
plausibly the thing Paterson and Gainesville punished us for lacking. The counter-risk is real
though: fewer paired examples to learn the pair separation #46 found us failing.

⚠️ San Francisco is disqualified for Stage 1 — no imagery needed

Its 50,096 records carry only 7,553 distinct coordinates, mapping 1:1 to cnn (SF's intersection
node id). Every ramp at an intersection is stamped with the intersection's point — modal 6–8 rows
per coordinate, up to 29. Stage 1 would project ~4.7 identical labels onto one pixel, none on a ramp.
Separately, 14,414 rows are crexist = 0 — confirmed absence, wrong polarity for Stage 1 (right
polarity for #86).

⚠️ Charlotte's coordinates disagree with Charlotte's own corner key

Recovery is P .621 / R .406, far below NYC and Minneapolis. Cause is spread: of its 3,839
multi-record published corners, 52.5% span >6 m, 14.4% >20 m, p99 = 308 m. Big arterial corners
explain ~30 m, not 308 m. Charlotte's review sheet must test this specifically.

Supply correction — and it closes a route

City Listed in §3 Corrected Why
San Francisco 50,096 0 intersection centroids
Charlotte 40,601 35,095 5,505 rows are RP_Type = NoRamp
others live drift only
Total ~236,000 180,673 −55,327

Good (276,615) + every unassessed city, all passing, = 457,288 — still short of 500,000. The
route that avoided both the OK tier and the state DOTs is now closed: 500k requires accepting the
tier the paper rejected, or the state-DOT tail with its Richmond/NYC clipping hazards.

Both corrections came from reading the data, not reviewing imagery — the automated checks pay for
themselves before a reviewer is booked. Also fixed a silent fetch truncation found the same way: SF
names its OID field ObjectId, so paging stopped after 2,000 of 50,096 and wrote a normal-looking
snapshot. Short fetches now refuse to write.

jonfroehlich and others added 20 commits July 31, 2026 09:22
…o-measure (#96)

Three problems, all found by Jon actually trying to review Denver with it.

**Annotations were burned into the JPEG.** They sit exactly on top of the pixels
being judged, so the reviewer has to be able to take them away to see whether a
ramp is underneath -- and baked-in marks cannot be removed without re-rendering
the whole sheet, which is not a workflow. render_chip now returns a clean image
and the rings, crosshair and scale bar are an SVG overlay, toggled by a checkbox
or the o key. It also means the overlay redraws at any display size without
resampling the imagery.

**The chips were too small to judge.** Denver's Aerial2018 cache stops at z19, so
a 40 m chip was 174 px. Its Aerial2016 cache goes to z21 -- leaf-off 3-inch
imagery at 0.057 m/px, 698 px per chip, 4x the linear detail, and detectable
warning pads are individually visible. Two years older, which does not matter for
a positional check because ramps do not move, and is in fact closer to the 2015
vintage 74% of Denver's records carry. The generalisable lesson for the remaining
cities: the deepest available level matters more than the capture year.

**There was no way to record a verdict except hand-editing JSON.** The sheet now
has an enlarged view where clicking the image computes the offset from the
crosshair exactly -- the reviewer points at the ramp and the page does the
measuring, so nobody estimates a distance by eye and the rings serve orientation
only. Ramps-visible, on-corner, unjudgeable and a note are entered in the page,
kept in localStorage so a refresh costs nothing, and exported as a verdicts.json
matching the template the script already writes. Keyboard: 0-3 sets ramps
visible, u unjudgeable, arrows move, n jumps to the next unreviewed chip.

Caught by inspection while writing it: the modal wrote bigsvg.outerHTML, which
detaches the node, so the overlay would have rendered exactly once and then gone
stale. It sets innerHTML on a persistent wrapper instead.

Suite 381 passed (was 376); the new tests guard placeholder substitution, that
META/CHIPS stay parseable, that the overlay scales to the chip's own viewBox, and
that imagery provenance reaches the page -- a verdict must never be readable
without knowing what it was made against.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…ng it (#96)

Jon asked whether the crosshair is really on the coordinate and whether the rings
are really 1/2/5/10 m. Neither is answerable by looking at the sheet, because the
error and the measurement come from the same code, so both are now checked
against something external. New script + 15 tests; evidence in
analysis_out/georef_check/.

Tile scheme: Denver's Aerial2016 cache is standard Web Mercator -- 256 px,
EPSG:3857, origin -20037508.342787, LOD resolutions matching the standard to
3e-10 relative.

Scale: points are constructed an exact ground distance apart from the WGS84 local
radii of curvature -- maths sharing nothing with the Web Mercator cos(lat) factor
it validates, so an error there cannot cancel -- then projected and measured at
eight bearings. Worst error 0.26% at every radius: 2.6 mm on the 1 m ring, 26 mm
on the 10 m ring. Constant in relative terms across radii, which identifies it as
the sphere-vs-ellipsoid residual rather than a bug. A regression test confirms the
checker reports >20% if the latitude correction is ever dropped.

Registration: Denver's own LRS centrelines are drawn into the imagery with the
same projection that places the crosshair, and the offset to the roadway's
optical centre measured on cross-sections every 4 m. 937 usable sections over
five neighbourhoods; resultant medians 0.08-0.46 m. A NAD83/WGS84 datum mismatch
applied on one side only -- the plausible failure, since Denver publishes in
EPSG:2877 and the server reprojects -- would be ~1 m and consistent in direction.
It is not there.

The measurement was wrong twice before it was right, and both corrections are
recorded because they are easy to repeat. Offsets have to be resolved into a
geographic frame: a segment normal's sign flips with the direction the segment
was digitised, so a real eastward shift cancels in the median. And each
cross-section may only be credited to the axis it crosses -- a north-south street
constrains east, not north, and pooling both in a grid city fills each median
with structural zeros and prints 0.00 m whatever the truth is.

Stated limit, not papered over: a centreline is a cartographic construct, not a
survey of the pavement midline, so this is evidence of no gross error at the
scale being measured, not a calibration certificate.

Suite 396 passed (was 381).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…ad (#96)

Ten minutes of real reviewing produced four questions the sheet could not
answer, and each would have moved the number: what is the "correct corner"
when the schema has no corner key, where on the ramp is the reference point,
how many ramps on a chip holding four corners, and do you click when it
already looks perfect.

The rules are now a RUBRIC constant rendered beside the field it governs,
opened in full with "?", and copied verbatim into the exported manifest --
"0.9 m" is uninterpretable without the rule saying what it is 0.9 m from.

The clause with the most at stake: click the centre of the concrete apron,
never the detectable-warning pad. PROWAG R305 puts the pad at the back of
curb, ~0.6-0.9 m down-slope of the ramp centre, and the pad is the most
visible thing in 0.057 m/px imagery -- so pad-clicking is the easy mistake
and would bias every record in one direction, indistinguishable from real
positional error and enough to move the bucket on its own.

Also: a readable corner with no ramp is now a verdict rather than a gap.
Such a chip was uncompletable -- nothing to click, so offset stayed null,
so done() was never true -- leaving only the wrong exit of "unjudgeable",
which asserts "I cannot see" rather than "I can see, and it is not there".
no_ramp records a phantom, which is a headline number for a schema with no
removal mechanism.

Each chip now carries how many records Denver itself publishes within 6 m
and 10 m, the same per-corner quantity from the published side, which is
what 5d deferred to imagery. Held back until the reviewer has entered their
own count, so the published figure cannot anchor the judgment it exists to
be compared against.

That comparison already found a threshold artefact: chip 66519 is a
pork-chop island whose three ramps sit at 0.0, 5.8 and 7.0 m, so single-link
at 6 m splits it and scores one as a singleton. Reviewer counted three,
Denver publishes three -- the inventory is per-ramp there and the clustering
loses it. 6 m was calibrated on NYC's tight urban corners, which is exactly
why NYC could not have revealed this.

19 new tests (33 in the file, 407 in the suite). Sheet regenerated offline
from cached tiles; city and seed unchanged, so in-progress localStorage
verdicts are preserved.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…trings (#96)

The existing tests assert the emitted HTML contains the right substrings,
which is not the same as the page working -- and the gap matters here more
than usual, because the sheet is a single 6.7 MB self-contained app whose
only user-visible failure mode is a blank screen with no error the reviewer
can act on. A stray brace in the template's {{/}} escaping would pass every
string assertion and waste an afternoon of human labour before anyone
noticed.

So load the emitted JavaScript into Node against a minimal DOM stub and
drive the verdict state machine directly: done() completes a no_ramp chip,
no_ramp and unreadable stay mutually exclusive, measuring clears both, and
the published-neighbour count stays hidden until the reviewer has entered
their own. That last one is load-bearing -- if it leaked early,
ramps_visible would stop being independent evidence and the comparison
against the published data would measure nothing.

Skipped when node is absent rather than made a hard dependency, per the
suite's CPU-only/no-network rule.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Aerial2016 is the right basemap for resolution -- 0.057 m/px against the
2018 cache's 0.23, which would render a 40 m chip as 174 px and lose
detectable-warning pads entirely -- and a positional check normally does
not care about capture year because ramps do not move.

But 74.4% of Denver's records carry a 2015 CREATEDATE, so for most of the
frame we are checking a delineation against imagery of nearly the same
date, quite possibly the imagery it was digitised from. That measures
digitising precision, which is the right quantity for "does the coordinate
land on the physical ramp" -- and is not the whole error a Stage 1 label
carries, because the label is projected into a ~2022 panorama and
everything that changed in between is invisible here.

So both headline numbers are lower bounds on their Stage 1 equivalents:
the offset distribution excludes post-2016 drift, and the phantom rate
excludes ramps that existed in 2016 and were gone by the panorama date --
which bites harder than usual given the schema records no removals. Names
5a as the instrument for that component, to be composed with this one
rather than either quoted alone.

Also corrects a stale reference to 2018 imagery in the sample-frame note,
left over from before the basemap switch.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Reviewing chip 66519 -- the channelising island -- the sheet reported "2
within 6 m, 4 within 10 m, you see MORE than are published: under-recording?"
against a correct count of 3. The alarm was spurious, and spurious for the
reason committed two commits earlier and then not encoded: that island's
three ramps span 7.0 m, so the 6 m threshold splits it. Comparing the
reviewer against the 6 m figure alone was guaranteed to raise a false
under-recording flag on exactly the corners the threshold is known to
mis-group, which is worse than not comparing at all -- it trains the
reviewer to ignore the signal.

The published per-corner count is BRACKETED by the two radii rather than
equal to either: 6 m splits large corners, 10 m reaches across a narrow
street. So only a count outside [p6, p10] is evidence. Above the 10 m
figure suggests under-recording, below the 6 m figure suggests phantoms or
duplicates, and anything between is consistent.

Same session surfaced the other half: the rings sit directly above the count
field and nothing said they do not bound it. Four ramps fall inside 66519's
10 m ring and the answer is three, because the fourth is across the slip
lane. Now stated in the inline hint and the rubric, with that chip as the
worked example.

Three new assertions in the node harness pin the bracket so it cannot
regress to a single-radius comparison.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Second false alarm in two chips, in the opposite direction to the first.
Chip 67585 is a triangular island where the sheet reported "4 within 6 m --
fewer than are published, phantom or duplicate?" against a correct count of
3. Resolving the neighbours by bearing rather than distance settles it: the
record 5.24 m ESE is across the slip-lane crossing, and a slip lane is 4-5 m
wide, so 6 m reaches straight over it. Denver publishes three on that island
and the reviewer counted three. Agreement, reported as a discrepancy.

Taken with 66519 -- where 6 m split a single island whose ramps span 7.0 m --
a radius is not a corner, and it fails in BOTH directions on exactly the
complex geometry where the comparison would matter. No widening of the
bracket fixes that; the two failures point opposite ways.

So the panel stops issuing verdicts. Each nearby published record is now
drawn on the chip as a magenta diamond, gated behind the same anti-anchoring
rule as the counts, and the reviewer decides which belong to the corner.
Three diamonds on the island and one across the crossing is visible in a
second and arguable in none.

Marker positions are computed in the chip's own Web Mercator projection --
verified against the sampled record's neighbours at 5.22 ESE / 5.40 NW /
5.95 W, matching an independent bearing calculation -- and regression-tested
for ground distance and for north being up, because a marker in the wrong
place would look like evidence rather than like a bug.

count_neighbours is now a projection of find_neighbours, so the counts and
the markers cannot disagree about what is nearby.

The general point, recorded in 5e and not about Denver: a threshold
calibrated on one city's geometry carries an unmarked assumption about that
geometry. 6 m validated against NYC's corner key at P .976 / R .973, and no
NYC corner could have exposed either failure -- it has neither large suburban
corner radii nor many channelised islands.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Two chips got reported as missing their neighbour markers when the markers
were present in the file on disk -- the page in the browser was a build
behind. A stale 6.7 MB file:// page is easy to keep and impossible to
distinguish from a fresh one by looking at it, and diagnosing it by asking
which wording is present does not scale.

So the header now carries a short content hash of the page logic plus the
rubric, and the same id goes into the exported manifest. "Am I on the
current sheet?" is one glance, and a verdict can be traced to the exact
instrument that produced it. Current build: 3c366f6b.

Markers were also genuinely small -- a 16 px magenta diamond with a 2 px
stroke, landing on both bright concrete and dark asphalt. Now larger, with
a dark halo under the magenta so it reads on either, and labelled with its
distance. The label matters for the far ones: 69169's neighbours are at
12.9 m and 15.6 m, out near the frame edge, where an unlabelled mark reads
as a stray rather than as deliberate.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Added it last commit and buried it: last item in the header's sub-line,
after five wrapped lines of grey text, in a dimmer grey than its
surroundings. Jon reasonably could not find it, which is the same mistake
as the 16 px markers -- the thing that answers "is my page stale?" placed
where you have to hunt for it.

Now a bordered monospace badge in the controls row, between the progress
counter and the buttons, with a title explaining what to do when it does
not match.

Three tests: the badge is in the controls row and not the fine print, the
hash moves when the rubric or page logic moves (a stamp that does not move
when the instrument moves would certify a stale page as current), and it
reaches the exported manifest so a verdict traces to the instrument that
produced it.

Current build: 989d90e8.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…rded (#96)

Asked when the markers appear, and the honest answer exposed a hole. The
gate was ramps_visible alone, so working in the order "count, then click"
put the diamonds on screen while the reviewer was still deciding where to
click. A click drifts toward a nearby marker, and offset_m silently stops
being "distance to the ramp" and becomes "distance to the published record".

That is the headline number of the whole assessment, and it was the one
measurement left unprotected -- I anchored the count and left the thing
that actually decides Denver's bucket exposed.

The gate is now done(v) && ramps_visible != null: the offset (or a terminal
state) AND the count. A no_ramp chip still reveals, because its count is
auto-set to 0 and the phantom check is exactly where comparing against
published records matters most.

The panel also stops hiding itself. An absent row read as a bug rather than
as a rule and cost time; it now states that it is holding, and why.

Six harness assertions cover the gate in both directions, including the two
partial states that must NOT reveal.

Build 265513c6.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The reviewer asked mid-pass whether the goal was calibrating coordinates,
checking existence, or whether it would be better to just click every ramp
in view. The third framing names a real blind spot, and it is a property of
the design rather than a defect in it.

The sample is drawn from the record list, so every chip is somewhere Denver
already pointed. That gives offset, phantom rate, and per-corner
completeness at corners Denver knows about -- and it cannot see a corner
Denver never recorded. An inventory publishing half its ramps accurately
would score perfectly here. Not hypothetical: #46 traced 72% of Paterson's
near-field misses to adjacent-pair merges, so under-recording has already
cost us a city.

The design still stands for the primary question, because every record
becomes exactly one Stage 1 label and a per-record average needs a
record-weighted sample. And the naive alternative fails on density: at 182
ramps/km a random 40 m chip holds 0.29 ramps, so ~71% of patches would be
empty pavement.

The version that works anchors on intersections rather than records, using
the centreline layer already fetched for the registration check, and yields
recall alongside the other three. Deferred, not rejected -- recall is not
needed to make the Denver call. Written down with the reason, so the
omission is not later mistaken for a withheld result.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…96)

Verdicts from the live pass, committed so the review stops living only in a
browser's localStorage. One chip left (71198) and three chips carry no ramp
count -- 66096's was toggled off, which is what re-clicking an already
selected segment does.

Headline, and it is a strong result: offset median 0.29 m, p75 0.65, p90
0.95, max 2.07, with 92.2% of records within 1 m [81.5-96.9]. Read the left
tail as floor-limited rather than as centimetres -- 0.29 m is 5 px, and the
registration check put imagery-vs-centreline medians at 0.08-0.46 m. Read
the whole distribution as a lower bound on Stage 1 error, since the 2016
imagery is near-contemporaneous with a delineation that is 74% 2015-dated.

Phantom rate 3/54 judgeable = 5.6% [1.9-15.1]. Unjudgeable 4/59 = 6.8%.

The per-corner result answers what 5d deferred to imagery: 50 consistent, 6
fewer than published, and ZERO cases of seeing more than Denver publishes.
No evidence of under-recording. 32 of 56 counted corners hold exactly one
ramp, mean 1.36 seen against 1.21 records/corner from clustering -- so
Denver's low ratio is ramp vocabulary, the benign reading, not merged pairs.

Two defects fixed on the way. Chip 98816 carried a 4.54 m click AND
unjudgeable, with the note "very hard to tell ... they look way off" -- so a
measurement the reviewer explicitly disowned was sitting at the top of the
distribution. unjudgeable now clears the offset exactly as no_ramp does, and
the summary excludes such chips regardless while still reporting that the
click existed. Excluding the two disowned clicks moves the max from 4.54 m
to 2.07 m.

New: scripts/analysis/inventory_review_summary.py, with Wilson intervals
because at n=54 and p near 5% the normal approximation runs below zero.
17 tests pin which chips land in which denominator.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
All 59 chips reviewed by Jon against the 5e rubric. The critical-path
question of 5 is answered for the largest unassessed city.

Offset median 0.29 m, p75 0.64, p90 0.95, max 2.07, with 92.3% [81.8-97.0]
of records within 1 m. Phantom rate 5.5% [1.9-14.9] of judgeable chips,
unjudgeable 6.8%. Verdict Good, against a threshold this document states
rather than assumes -- >=90% within 1 m and phantoms under 10% -- because
the paper published none. The 1 m mark is principled: a ramp is 1.2-1.8 m
deep, so within 1 m of the centre is very nearly on the concrete.

The per-corner result settles what 5d deferred to imagery. ZERO chips where
the reviewer saw more ramps than Denver publishes, so no evidence of the
pair-merge under-recording that #46 traced to 72% of Paterson's near-field
misses. 32 of 57 corners carry exactly one ramp; mean 1.37 seen against 1.21
records/corner from clustering. The two agree within the clustering's known
under-grouping, which confirms on imagery the benign reading 5d could only
infer from Sioux Falls and Charlotte: Denver's low ratio is ramp design
vocabulary, not missing records.

Caveats travel in the same section, not a separate note. It is a LOWER bound
-- 2016 imagery against a 74%-2015-dated delineation measures digitising
precision and is blind to drift before the ~2022 panorama. The left tail is
floor-limited at 5 px, under the imagery's own registration residual. The
phantom interval reaches 14.9%, which over 72,770 records is the difference
between ~4,000 and ~10,800 false-positive labels; that needs a bigger sample,
not a better instrument. Recall is structurally unmeasured. One rater.
92078 and 135499 stay unexplained and are named as such.

For 6: Denver passing moves the usable pool 276,615 -> 349,385, the largest
single addition available. It does not reach 500k and nothing does -- Good
plus every remaining unassessed city, all passing, is 457,288. That decision
is simply no longer blocked on Denver.

Also ships the completeness fix that this pass needed: "done" now means fully
recorded, not just measured. A chip with an offset but no count used to style
as finished and be unreachable from "next unreviewed" while silently shrinking
the per-corner denominator -- 66096 was exactly that, and Jon could not have
found it. Partials now show amber, are counted separately, and are routed to.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
5f produced an offset distribution, which is not a decision without a
tolerance -- and the tolerance is a property of Stage 1, not of the aerial
imagery it was measured on. Reading download_dataset.py rather than assuming:

    azimuth = geod.inv(pano, ramp) - pano_angle
    persp = equirectangular_to_perspective(equi, 90, azimuth, -30, 1024, 1024)
    persp = persp[0:1024, 341:341+341]

The published coordinate is consumed ONLY for its bearing from the panorama.
The range is computed and discarded; the crop model then localises the ramp
inside a strip cut around that bearing, so the label's position comes from
the imagery, not the coordinate. Three consequences, none of them visible in
an offset distribution:

- tolerance is ANGULAR: the strip is +/- 18.37 deg. Not 90*341/1024 = 30 --
  a pinhole projection is not linear in angle, and the naive figure
  overstates by 63%. There is a regression test for exactly that mistake.
- RADIAL ERROR IS FREE. An offset along the line of sight does not move the
  bearing at all.
- metric tolerance scales with range at 0.332x: 1.0 m at 3 m, 3.3 m at 10 m,
  6.6 m at 20 m. The same error is fatal beside the camera and irrelevant
  across the intersection.

Ranges measured over 6,238 GT ramps in all nine benchmark bundles: median
11.1 m. Monte Carlo over Denver's offsets x those ranges x uniformly random
error direction gives P(true ramp outside its own crop) = 0.21%, and 0.00%
beyond 10 m where the median ramp sits.

The reusable output is the tolerance curve, and it says 5f's threshold was
far too strict: a median offset near 1 m loses under 10% of labels. That
bears directly on 6's conclusion that 500k requires the OK tier the paper
rejected -- "OK" may cost much less than the label implies. It does not
license skipping assessment, since a city can be Poor through phantoms or a
heavy tail rather than through its median.

Geometric bound only: it does not model whether the crop model still
localises a ramp near the strip edge, which needs the round-2 checkpoint and
a GPU. Real degradation begins earlier than this says, never later. Stated
in the doc next to the number, along with the uniform-direction and
benchmark-range assumptions.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
5f had to invent a threshold, so "Denver is Good" is a scale of one and
cannot be compared to Table 1. NYC (Good, and 78% of training data) and
Seattle (the only city the paper explicitly rated Poor, so the only
available anchor for the bottom of the scale) would fix that permanently.
Basemaps probed first, because 5e's rule is that a sheet built on the wrong
imagery is worse than no sheet.

Seattle is usable: King County KingCo_Aerial_2021, EPSG:3857, 256 px,
0.1007 m/px -- coarser than Denver's 0.0573, so a ramp is ~15 px rather
than ~26 and the offset floor roughly doubles. Acceptable for this job
specifically: what we need from the Poor anchor is the size of a LARGE
error, which 0.1 m/px resolves easily. It would not do for grading a city
expected to be Good.

A declared LOD is not a built cache. King County advertises maxLOD 23
(0.0126 m/px) and 404s above z20. Denver's failure was grey placeholders
past the depth; this one is a 404. Both mean the deepest level must be
found by probing, never by reading -- so probe_basemap.py scans downward
and reports the gap between declared and served.

NYC is harder. maps.nyc.gov/xyz returns 403. The usable service, NYS ITS
wms/Latest, is dynamic-only, so imagery comes from /export on a bbox rather
than /tile/{z}/{y}/{x} and needs a fetcher the review sheet does not have.
Not built. A visual check of a Manhattan chip also shows heavy building
shadow and roof-lean, so NYC should be expected to produce a materially
higher unjudgeable rate than Denver's 6.8% -- itself a finding about where
this method works.

Seattle's inventory located but not frozen: SDOT Curb_Ramps_(Active) at
38,498 and Curb_Ramps_CDL at 46,431, both drifted from 5c's 38,468/46,386,
which is the live-drift argument for snapshotting.

14 tests. The two that matter pin the reasoning rather than the plumbing: a
cache with the right CRS but a bespoke resolution ladder must be rejected,
and the resolution grades must match the thresholds 5e argues for.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…er (#96)

Inventory frozen: SDOT Curb_Ramps_(Active), 38,364 records of 38,498
declared. The shortfall was VERIFIED as null geometry rather than assumed --
every OBJECTID queried with returnGeometry=true, all 38,498 checked, exactly
134 null -- before overriding the fetcher's truncation guard.

The geometry gate turns up the surprise: Seattle is 1.34 records/corner,
THIRD of eleven, above every other candidate and above Bend (1.31), which is
in the training set. The one city the paper explicitly rejected on location
precision has better per-corner completeness than anything else available.
Pairing density and positional precision are independent axes and Seattle is
the case that proves it.

Then the basemap. A first sheet was built on KingCo_Aerial_2025 and looked
wrong on inspection -- chips under canopy, an unreadable landscaped traffic
circle, a building leaning across a crosshair. An excess-green index over the
sampled chips quantifies it: every King County year is leaf-on at 27-41%
vegetation cover against Denver's 7.2%. Seattle's own sharper caches are
EPSG:2926, WA State Plane, so the sheet's Web Mercator tile math does not
apply to them at all -- exactly what probe_basemap.py exists to catch.

Rebuilt on KingCo_Aerial_2019, the least leafy Web Mercator option, with both
limits written into the tile-source note so they reach the manifest: expect an
unjudgeable rate far above Denver's 6.8%, and expect a selection effect, since
the corners that stay readable are the ones without street trees. Adequate to
size a large error, which is what a Poor anchor needs. Not adequate to grade a
city expected to be Good.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…tomatic (#96)

Jon looked at the Seattle chips and said several "feel so very off". He was
right, and the number I reported an hour earlier was not measuring what I
said it measured.

Resolving his clicks into a vector rather than a magnitude: Seattle's mean
offset vector is east -2.04, north +0.30 -> resultant 2.06 m against a mean
magnitude of 2.37 m. That is 87% SYSTEMATIC, with the ramp west of the
published point in 9 of 11 chips. Denver, the same test, is 0.10 m resultant
against 0.44 m magnitude -- 24%, east-positive 19 of 52, i.e. it cancels the
way random error must.

So Seattle's "median 2.33 m" is a registration error, not coordinate
imprecision, and must not be quoted as a precision figure or fed through 5g's
tolerance curve. I did both in my previous message; that was wrong.

A second, independent confound sits on top: INSTALL_DATE shows 62.9% of dated
Seattle ramps were installed in 2019 or later, AFTER the imagery. Seattle is
under active ADA remediation, so the 2019 basemap frequently shows the ramp
that was replaced.

Which side is wrong is not yet settled. Drawing SDOT's own Street Network
Database centrelines over King County imagery, they appear to track the
roadways -- which would put the fault in the coordinates rather than the
basemap -- but that is n=3 and eyeballed, and needs the quantitative
cross-section treatment verify_chip_georeference.py gives Denver.

The durable fix is that this check is now automatic and free: the reviewer's
click already records a direction, so summarise() reports the mean vector
alongside the mean magnitude and shouts when the systematic share exceeds
50%. Denver would have passed it silently; Seattle would have been caught on
chip 11 instead of after the pass. 6 tests, including that north is up in the
reported vector -- reporting a shift in the wrong direction is worse than
reporting none.

Also: unjudgeable now reports over ATTEMPTED chips as well as all chips. Over
all chips it read 11.7% mid-review when the honest figure was 36.8% of 19
attempted, which understated exactly the leaf-on problem the basemap note
warned about.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…high n (#96)

Yesterday I told Jon Seattle's offset was 87% SYSTEMATIC and therefore a
registration error, not imprecision. The direction is real; the conclusion was
not. Two instruments that can see a city-wide displacement both say there
isn't one.

First, the statistic needed a null it never had. systematic_share was
documented as "~0 for noise", which is wrong: under random headings the mean
vector shrinks only as 1/sqrt(n), so the expected share is ~0.9/sqrt(n) --
39% at Seattle's n=11, not 0%. systematic_shift_null now resamples the
observed magnitudes with random directions. Seattle 87% -> p=0.0013, so the
lean survives. Denver 24% -> p=0.221 against a null median of 16%, so that
number never meant anything either. Both readings changed.

A directional signal still doesn't say which side is wrong, so measure the
other two legs of the triangle:

  (ramps vs imagery) = (ramps vs centrelines) + (centrelines vs imagery)

inventory_centerline_offset.py does the middle leg with no imagery and no
reviewer. A ramp sits half a roadway from the centreline on one side or the
other, so median(east side) + median(west side) = 2x the shift while the
unknown half-width cancels and comes back as a sanity check. Seattle reads
0.00 m over 31,430 samples; Denver, the control, reads 0.12 m against the
0.10 m its 52 reviewer clicks measured independently. Half-widths land at
6.2-7.5 m on both axes in both cities, and it is flat across a 10-30 degree
sweep of the cardinal cutoff. Seattle's coordinates are unbiased against
Seattle's own street network -- tighter than Denver's.

That leg is blind to an error the ramp and centreline layers share, which is
why the imagery leg is required. verify_chip_georeference.py grew a city
registry and --sites-from-verdicts, because orthorectification error is LOCAL
and a five-neighbourhood average cannot clear the places the verdicts actually
came from. Measured under the review itself, all eleven chips clear at <=0.32 m
with mean east +0.02 -- including chip 1951390, the 8.79 m click, at 0.11 m.

So the triangle predicts 0.0-0.3 m and the sheet says 2.06 m. It does not
close, and that is the result: no city-wide displacement can hide from 31,430
samples or from a per-chip imagery check. The sheet's own quartiles say the
same thing once you stop averaging -- p25 is 0.47 m against a 2.33 m median,
and a uniform shift cannot produce a chip that is dead-on. What made 87% was a
handful of large PER-RECORD errors sharing a heading.

Consequences: the SHIFT framing is withdrawn; there is no constant to subtract,
which was the optimistic branch; Seattle's Poor rating is neither confirmed nor
overturned, because eleven chips on a basemap already declared inadequate for
grading settle nothing; and the 62.9%-installed-after-2019 confound is now the
leading candidate.

The durable fix is that inventory_review_summary.py now prints the null and the
p-value, refuses to call a shift on its own, and names the two instruments that
can answer it. Both are cheap and should run before any future city is scored.

Also: centrelines are frozen like everything else (fetch_inventory.py
--geometry polyline, 34,484 SND + 7,866 Denver, both complete) -- an analysis
that exonerates a city's coordinates must not depend on a live endpoint for the
reference it exonerated them against. Each city's centrelines must come from
the same publisher and CRS as its ramps, which is what makes a shared datum
fault cancel and a ramp-layer defect show.

50 new tests, suite 516. The sign convention is pinned against synthetic data
with a planted displacement, and there is a test that a real shift is not
attenuated by which street each ramp is compared against -- picking the overall
nearest street biases toward zero, so the nearest is chosen per axis.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
--sites-from-verdicts re-renders the same two visual sites as its city's own
run, so the PNGs are byte-duplicates and only the JSON carries anything new --
the per-chip medians that cleared Seattle's imagery. Left untracked they showed
up as permanent working-tree noise on every re-run.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Jon's question: can we gate a city by just running the labels and seeing if
they land, and do we even have enough to do that for the published cities,
given the original government files were never committed?

Yes, and the second half is why. Reading generate_dataset_meta.py rather than
assuming: curb_ramps_coords is a plain 35 m radius query against
all_locations.csv, and download_dataset.py copies it into each pano's JSON
untouched. No model is in that loop, so the published dataset carries the
government coordinates verbatim -- the denominator is not contaminated by the
thing being measured, and the portal files are not needed.

The other half is that the labels encode a bearing.
perspective_to_equirectangular maps column u to lon = (u/(W-1))*2pi - pi, and
that lon IS the azimuth relative to the pano heading. So per record,

    residual = wrap((x*360 - 180) - (bearing_gov - pano_azimuth))

which is registration error in the angular units 5g proved Stage 1 cares
about. Mean catches a systematic shift, spread catches imprecision, match rate
is the label yield for free. CPU, four columns over HTTP range requests.

90,006 records / 16,808 panos. The convention validates itself: median nearest
separation 3.45 deg, 98.5% inside the +/-18.37 deg strip, where a wrong
convention gives ~90 deg and ~10%.

    nyc       62,132  match 0.847  mean +0.055 (se 0.026)  |med| 3.30 deg
    portland  26,180  match 0.898  mean -0.250 (se 0.038)  |med| 3.35 deg
    bend       1,742  match 0.925  mean +0.036 (se 0.105)  |med| 2.19 deg

This is the null 5i went looking for and never had. All three Good cities sit
within |mean| <= 0.25 deg -- Portland's is real at 6.5 se and physically 4.8 cm
-- and at n=26k-62k a 0.1 deg shift is resolvable. A genuine 2.06 m tangential
shift at the 11.1 m median range would read as ~10.5 deg, some 250 se clear.
An instrument that resolves 2 cm cannot miss 2 m. Robust to the matcher: the
pairing cap over {18.37, 40, 90} moves every |median| by <=0.12 deg.

Four structural limits, all in the doc beside the numbers: censored at the
strip (so matched_frac must be read with it); greedy matching biases low, so
these are lower bounds; min_distance=40 merges make match rate partly density;
and it cannot see records that never reached a panorama, so this is
crop-model-stage yield, not end-to-end. It also cannot separate bad
coordinates from absent ramps, nor see a shift big enough to land on the
neighbouring corner -- the visual gate stays the only phantom detector.

Two corrections to 5g while here. Its +/-18.37 deg is now imported rather than
re-derived: an averaged 170.5 px half-width silently gives 18.42, because the
strip is asymmetric about the centre column and the published figure is the
conservative side. And its "round-2 checkpoint, not in the repo" is true but
was read as unavailable -- it is on klone at
/gscratch/makelab/jsomeara/RampNet/stage_one/crop_model/ps_and_manual_model/best_model.pth
(360 MB, readable), so the empirical arm 5g deferred is runnable.

Also fixed: frac_outside_crop was mislabelled. A peak further from its matched
record than the crop half-angle cannot have come from that record's own strip,
since the combined heatmap is the max over all crops -- it is a
cross-assignment rate, a floor on matcher error, not 5g's geometric quantity.
Conflating them would turn a matching artefact into a claim about coordinates.

19 tests, including recovery of a known injected shift, which is the property
Seattle would need. 535 pass overall.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant