Repository navigation
NI data layer round 2: asked quantity wins selection, tide type codes, result/schedule row filter, honest counts, trend history - #499
Merged
Conversation
…, result/schedule row filter, honest counts, trend history
After the fourth sealed blind set (right 12 / partial 7 / wrong 4 of 23 live,
all data layer): among tied list/columns answers the one whose own name or
label says an ask word wins ('Snow forecast' over 'Tonight, hour by hour' on
'snow'; measured on all 86 recorded live cards: only the two snow asks change);
the subject row filter reaches result and schedule asks (a named team's row,
or the honest nothing); a count ask always cuts its window and never demotes a
one-row list to a measure; a trend ask takes the source's declared history
over a tied latest value; the per-answer cell cap is 5. A title-fallback rule
built in this round was measured (it re-titled 27 right cards) and dropped.
Fixtures: live_2026-10-08d joins the oracle (23/23, 21/21, 17/17, 23/23); FIT
labeled set 172 records, precision('no') vs WRONG 0.76 (authority stays off;
the ten missing verdicts are deterministic invalid replies, not truncation).
Docker gate: ruff, recorded PASS, chaos PASS, suite 4106 passed / 18 skipped.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
NI data layer round 2 after the set D read: asked quantity wins selection, tide type codes, result/schedule row filter, honest counts, trend history
Round 19 (plan
read-the-code-in-fizzy-deer). The fourth sealed blind set on main read right 12 / partial 7 / wrong 4 of 23 live cards, every wrong in the data layer (blind50d_review.md). This PR fixes the four classes and carries five partials; the forms and the client are untouched. Pairs with Library PR SecureCloudGroup/SmartBrain_Library#11 (the snow and tide answers, the 5-cell cap).The four wrong classes
snow_forecastanswer, but it tied with the generic "Tonight, hour by hour" answer on the one shared word ("snow" is in that answer's words too) and the declared order won. Among tied list / columns answers the one whose own name or label says an ask word now wins (select_answers). Measured over all 86 recorded live cards: selection changes on exactly two, both snow asks, both to the snow forecast (a day-by-day snowfall table). The worker's first version re-titled the card to the chosen answer's label and added a note; measured over the same 86 cards it would have re-titled 27 right cards to generic labels ("ISS" to "Latitude", "Bruins" to "Next game"), so that mechanism is not in this PR. The Library side adds snowfall to the hourly tonight answer and keeps its temperature, which needed the per-answer cell cap raised from 4 to 5 in both repos.next_highandnext_lowanswers filtered on the H/L type cell, with the words that pick them; the generic tides answer keeps the whole list. No app change.resultandscheduleasks: a named team present in the rows seals the filter to its row (the Sharks game was row 7 of 10; the card showed rows 1–2); absent everywhere, the honest nothing path; a team on both sides over different rows keeps the whole list. Measured on every recorded trend / result / schedule / count card: no other change.countask now always cuts the rows to its window (the "all rows in the past, show the latest anyway" leniency was letting a two-week-old row through) and never demotes a one-row list to a measure; zero rows after the cut is the honest "0 this week". (The place "Japan" resolving to a Louisiana hamlet is a Library place-resolver gap, noted, not fixed here.)Partials carried
A
trendask with a declared history list (FREDrecent) chooses the history, not the tied latest value (all seven recorded trend cards now build a series). The Kp forecast cells carry their unit (Library). Not fixed, by design: the count stat beside a list (#36) and a forecast series leading with tomorrow (#8) are forms-engine work; #49 noted.FIT
The labeled set grows to 172 records (86 live cards from the four blind sets with their by-eye verdicts, 86 cross-paired negatives). Measured natively against the local 9B: validity 162/172; precision of "no" against WRONG 0.76 (unchanged from 0.77); authority stays advisory. The eval now names why a verdict is missing: all ten no-verdict rows are replies that were invalid twice, none a transport error; re-running the eval with the transport's reply budget raised from 400 to 600 tokens gives a byte-identical result, so they are deterministic shape failures on those inputs, not truncation (production sends no explicit cap; the gateway's model default applies). No budget change ships.
Fixtures and tests
tests/fixtures/ni_forms/live_2026-10-08d/(23 cases from the set D run; the four wrong cases and the trend case rebuilt through the fixed code) joins the oracle: 23/23, 21/21, 17/17, 23/23. New labeled tests intests/test_ni_flow_datalayer_r2.py(tie-break, cap, W3 ×3, W4 ×3, trend ×2). Three earlier fixtures carry asubject_filterannotation for the pre-existing "Final games" filter the now-exercised path reads.Verified (Docker dev image)
ruff clean; engine subset 501 passed; targeted NI set 1,537 passed (oracle + round 2 tests 244 after the final fixture);
ni-flow-eval.py --recordedPASS;--chaosPASS; full suite 4,106 passed, 18 skipped. Selection snapshot over all 86 recorded live cards before/after the tie-break: 2 changes, both snow asks. Live build probe (Buffalo, Duluth): the snow forecast builds a 7-day snowfall table.Found, not fixed
"Japan" resolves to a Louisiana hamlet (place resolver has no countries); tide times display in the viewer zone (the deferred card-zone step);
frame_kind_from_textclassifies near-identical map asks inconsistently; the count-stat-beside-list form and the forecast "leads with tomorrow" rule (forms engine).