Skip to content

NI data layer round 2: asked quantity wins selection, tide type codes, result/schedule row filter, honest counts, trend history - #499

Merged
SecureCloudGroup merged 1 commit into
mainfrom
round19/datalayer-r2
Oct 9, 2026
Merged

SecureCloudGroup merged 1 commit into
mainfrom
round19/datalayer-r2

Conversation

@SecureCloudGroup

Copy link
Copy Markdown
Owner

NI data layer round 2 after the set D read: asked quantity wins selection, tide type codes, result/schedule row filter, honest counts, trend history

Round 19 (plan read-the-code-in-fizzy-deer). The fourth sealed blind set on main read right 12 / partial 7 / wrong 4 of 23 live cards, every wrong in the data layer (blind50d_review.md). This PR fixes the four classes and carries five partials; the forms and the client are untouched. Pairs with Library PR SecureCloudGroup/SmartBrain_Library#11 (the snow and tide answers, the 5-cell cap).

The four wrong classes

  • W1 — "any snow expected in Duluth midweek" showed the hourly temperature series under the title "snow". Root cause in answer selection: the source declares a snow_forecast answer, but it tied with the generic "Tonight, hour by hour" answer on the one shared word ("snow" is in that answer's words too) and the declared order won. Among tied list / columns answers the one whose own name or label says an ask word now wins (select_answers). Measured over all 86 recorded live cards: selection changes on exactly two, both snow asks, both to the snow forecast (a day-by-day snowfall table). The worker's first version re-titled the card to the chosen answer's label and added a note; measured over the same 86 cards it would have re-titled 27 right cards to generic labels ("ISS" to "Latitude", "Bruins" to "Next game"), so that mechanism is not in this PR. The Library side adds snowfall to the hourly tonight answer and keeps its temperature, which needed the per-answer cell cap raised from 4 to 5 in both repos.
  • W2 — "next low tide in Bar Harbor" led with the next high. Library only: the tide source declares next_high and next_low answers filtered on the H/L type cell, with the words that pick them; the generic tides answer keeps the whole list. No app change.
  • W3 — "result of the latest Sharks game" showed the league scoreboard. The subject row filter now covers result and schedule asks: a named team present in the rows seals the filter to its row (the Sharks game was row 7 of 10; the card showed rows 1–2); absent everywhere, the honest nothing path; a team on both sides over different rows keeps the whole list. Measured on every recorded trend / result / schedule / count card: no other change.
  • W4 — "how many earthquakes hit Japan this week" showed a single 2.7 magnitude from Sep 25. A count ask now always cuts the rows to its window (the "all rows in the past, show the latest anyway" leniency was letting a two-week-old row through) and never demotes a one-row list to a measure; zero rows after the cut is the honest "0 this week". (The place "Japan" resolving to a Louisiana hamlet is a Library place-resolver gap, noted, not fixed here.)

Partials carried

A trend ask with a declared history list (FRED recent) chooses the history, not the tied latest value (all seven recorded trend cards now build a series). The Kp forecast cells carry their unit (Library). Not fixed, by design: the count stat beside a list (#36) and a forecast series leading with tomorrow (#8) are forms-engine work; #49 noted.

FIT

The labeled set grows to 172 records (86 live cards from the four blind sets with their by-eye verdicts, 86 cross-paired negatives). Measured natively against the local 9B: validity 162/172; precision of "no" against WRONG 0.76 (unchanged from 0.77); authority stays advisory. The eval now names why a verdict is missing: all ten no-verdict rows are replies that were invalid twice, none a transport error; re-running the eval with the transport's reply budget raised from 400 to 600 tokens gives a byte-identical result, so they are deterministic shape failures on those inputs, not truncation (production sends no explicit cap; the gateway's model default applies). No budget change ships.

Fixtures and tests

tests/fixtures/ni_forms/live_2026-10-08d/ (23 cases from the set D run; the four wrong cases and the trend case rebuilt through the fixed code) joins the oracle: 23/23, 21/21, 17/17, 23/23. New labeled tests in tests/test_ni_flow_datalayer_r2.py (tie-break, cap, W3 ×3, W4 ×3, trend ×2). Three earlier fixtures carry a subject_filter annotation for the pre-existing "Final games" filter the now-exercised path reads.

Verified (Docker dev image)

ruff clean; engine subset 501 passed; targeted NI set 1,537 passed (oracle + round 2 tests 244 after the final fixture); ni-flow-eval.py --recorded PASS; --chaos PASS; full suite 4,106 passed, 18 skipped. Selection snapshot over all 86 recorded live cards before/after the tie-break: 2 changes, both snow asks. Live build probe (Buffalo, Duluth): the snow forecast builds a 7-day snowfall table.

Found, not fixed

"Japan" resolves to a Louisiana hamlet (place resolver has no countries); tide times display in the viewer zone (the deferred card-zone step); frame_kind_from_text classifies near-identical map asks inconsistently; the count-stat-beside-list form and the forecast "leads with tomorrow" rule (forms engine).

…, result/schedule row filter, honest counts, trend history

After the fourth sealed blind set (right 12 / partial 7 / wrong 4 of 23 live,
all data layer): among tied list/columns answers the one whose own name or
label says an ask word wins ('Snow forecast' over 'Tonight, hour by hour' on
'snow'; measured on all 86 recorded live cards: only the two snow asks change);
the subject row filter reaches result and schedule asks (a named team's row,
or the honest nothing); a count ask always cuts its window and never demotes a
one-row list to a measure; a trend ask takes the source's declared history
over a tied latest value; the per-answer cell cap is 5. A title-fallback rule
built in this round was measured (it re-titled 27 right cards) and dropped.
Fixtures: live_2026-10-08d joins the oracle (23/23, 21/21, 17/17, 23/23); FIT
labeled set 172 records, precision('no') vs WRONG 0.76 (authority stays off;
the ten missing verdicts are deterministic invalid replies, not truncation).
Docker gate: ruff, recorded PASS, chaos PASS, suite 4106 passed / 18 skipped.
@SecureCloudGroup
SecureCloudGroup merged commit 1686e5c into main Oct 9, 2026
13 checks passed
@SecureCloudGroup
SecureCloudGroup deleted the round19/datalayer-r2 branch October 9, 2026 00:41
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant