Skip to content

Query layer (Q3a): the ask as a typed query over a declared answer, measured - #502

Open
SecureCloudGroup wants to merge 1 commit into
mainfrom
round20/query-core
Open

SecureCloudGroup wants to merge 1 commit into
mainfrom
round20/query-core

Conversation

@SecureCloudGroup

Copy link
Copy Markdown
Owner

Step Q3a of the Round 20 plan. New code only; the flow does not use it yet (Q3b wires it in and removes the regex parsers one per measured delta).

What it is. The ask becomes one closed query object over the chosen declared answer: {answer, time, where, order} from one local model call, {select, limit, agg, params} from code, every clause validated against the declared cell types, explicit time phrases decided by two recognizer libraries, and a code-rendered interpretation line ("Earthquakes · Magnitude > 4 · past 7 days · count") for the consent screen and the card footer.

  • ni_query/recognize.py: Microsoft Recognizers-Text (community fork) + puckling (the Duckling port) merged into one closed span shape, behind a word-boundary filter and an entity mask; comparators from a documented adjacency table; no other ask regex.
  • ni_query/prompt.py + plan.py: the measured prompt (literal skeleton, answer and parameter menus, three retrieved examples), the call through ni_forms.llm.chat_json with a new optional check= hook whose retry messages state the violated rule and never list a menu; coverage per clause; the rules floor when the model is unavailable.
  • ni_query/normalize.py, clauses.py, apply.py, say.py: the v1 contract, the code-owned clauses, the row filter/cut/order/limit/count, the interpretation line.
  • tools/ni-query-eval.py + tests/fixtures/ni_query/: the gold set (109 hand-labeled objects over the live cards of five blind sets), recorded model replies, 92 recorded samples, the DEV/TEST split and the few-shot pool; --replies replays deterministically, --exec scores execution agreement, --gateway runs the live model.

Numbers. Deterministic replay of the recorded 9B replies: 50/60 fully right (few-shot TEST) and 54/82 (zero-shot ALL) under the v1 contract, matching the lead's independent re-score to the row; 53/60 and 65/82 with the rules floor; execution agreement 46/59 and 55/80. Live through the local gateway with byte-identical prompts: 51/60, 53/60 with the floor; valid 68 first try / 9 after the rule-only retry / 2 never. Per clause live: answer 55, select 57, where 52, time 51, order 55, limit 58, agg 60 of 60.

Dependencies. ms-recognizers-text-suite==1.0.1 (MIT, six sub-packages) and puckling==0.5.0 (Apache-2.0) with their transitive pins (emoji, grapheme, multipledispatch, datedelta) in requirements.lock and THIRD_PARTY_LICENSES.md; both imported lazily.

Gates (Docker smartbrain_3000:dev): ruff clean; targeted 183 passed; replay parity 109/109 and 79/79; full suite 4,280 passed, 18 skipped on the identical scratch tree.

Open for Q3b: answer replaces select_answers; time replaces _window_from_text bound to the window ops with the answer's direction and the clock parameters; where replaces the subject row filter family with a param_bounds push-down (Library key); agg replaces the count rule; spec.interpretation; a messages adapter for the flow's call_model. One ruling needed: "over the last year" as calendar 2025 (the recognizers) or rolling 365 days (the gold labels).

…l call, code-owned clauses, apply and interpretation, with the in-repo eval tool and fixtures
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant