Skip to content

perf: read a request's questions as a prefix tree in one call - #22

Merged
feder-cr merged 1 commit into
mainfrom
prefix-tree
Oct 1, 2026
Merged

feder-cr merged 1 commit into
mainfrom
prefix-tree

Conversation

@feder-cr

@feder-cr feder-cr commented Oct 1, 2026

Copy link
Copy Markdown
Owner

A choice asked beside other questions lost the shared reading: the prompts' common prefix was then only the state, so every option re-read the instructions and the whole option list.

Change: each model call reads its questions as a prefix tree (src/tree.hpp). Runs that several prompts share (a choice's instructions and options, a score's scale) are read once; every token still sees only its own prompt's tokens, so each answer is the one its prompt gets alone.

Counted with a tracing build (one line per OpenVINO infer()):

Request Before After
choice 3 / 10 options 1 call, 94 / 217 tokens same
choice 10 + 2 yes/no + score 4 2 calls, 1,332 tokens 1 call, 278 tokens

Latency, 8 threads, alternated with the release build (the laptop was training on its GPU): the mixed request 2.36 s → 0.47 s; the others unchanged within noise.

Tests

  • plan-calls-test: every question's path in its call's tree spells exactly its tokens, shared runs are counted once, budgets hold, plus same/prefix/one-token questions.
  • check.py: 15 prompts (10-option choice, 2 yes/no, 4-level score) in one request vs each alone, f32 activations: max |dP| 0.000001.
  • check.py and snapshots.py take --threads, to share the machine.
  • parity.py ignores /v1/models' release_date (the day the server started): it failed against the recording from 30 September on any other day, main included.
  • golden/jev.json re-recorded; full check.py green locally.

🤖 Generated with Claude Code

A request's prompts shared only what all of them start with. With a choice beside other questions
that was the state alone, so every option re-read the instructions and the whole option list: a
10-option choice, 2 yes/no questions and a 4-level score made 2 model calls of 1,332 tokens. Each
call now reads its questions as a prefix tree (src/tree.hpp): the runs several prompts share (a
choice's instructions and options, a score's scale) are read once, and every token still sees only
its own prompt's tokens. The same request is now 1 call of 278 tokens: 2.36 s -> 0.47 s at 8 threads
(alternated with the release build); the other requests are unchanged.

plan_calls packs questions in order by what each adds to its call's tree. Tests: plan-calls-test
checks every question's path spells its tokens and shared runs are counted once; check.py asks 15
prompts in one request and each alone with f32 activations (max |dP| 0.000001). check.py and
snapshots.py take --threads, to share the machine. parity.py ignores /v1/models' release_date, the
day the server started (it failed against the recording from another day). golden/jev.json
re-recorded.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@feder-cr
feder-cr merged commit fa3f9cf into main Oct 1, 2026
6 checks passed
@feder-cr
feder-cr deleted the prefix-tree branch October 1, 2026 02:41
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant