perf: read a request's questions as a prefix tree in one call - #22
Merged
Merged
Conversation
A request's prompts shared only what all of them start with. With a choice beside other questions that was the state alone, so every option re-read the instructions and the whole option list: a 10-option choice, 2 yes/no questions and a 4-level score made 2 model calls of 1,332 tokens. Each call now reads its questions as a prefix tree (src/tree.hpp): the runs several prompts share (a choice's instructions and options, a score's scale) are read once, and every token still sees only its own prompt's tokens. The same request is now 1 call of 278 tokens: 2.36 s -> 0.47 s at 8 threads (alternated with the release build); the other requests are unchanged. plan_calls packs questions in order by what each adds to its call's tree. Tests: plan-calls-test checks every question's path spells its tokens and shared runs are counted once; check.py asks 15 prompts in one request and each alone with f32 activations (max |dP| 0.000001). check.py and snapshots.py take --threads, to share the machine. parity.py ignores /v1/models' release_date, the day the server started (it failed against the recording from another day). golden/jev.json re-recorded. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
A choice asked beside other questions lost the shared reading: the prompts' common prefix was then only the state, so every option re-read the instructions and the whole option list.
Change: each model call reads its questions as a prefix tree (
src/tree.hpp). Runs that several prompts share (a choice's instructions and options, a score's scale) are read once; every token still sees only its own prompt's tokens, so each answer is the one its prompt gets alone.Counted with a tracing build (one line per OpenVINO
infer()):Latency, 8 threads, alternated with the release build (the laptop was training on its GPU): the mixed request 2.36 s → 0.47 s; the others unchanged within noise.
Tests
plan-calls-test: every question's path in its call's tree spells exactly its tokens, shared runs are counted once, budgets hold, plus same/prefix/one-token questions.check.py: 15 prompts (10-option choice, 2 yes/no, 4-level score) in one request vs each alone, f32 activations: max |dP| 0.000001.check.pyandsnapshots.pytake--threads, to share the machine.parity.pyignores/v1/models'release_date(the day the server started): it failed against the recording from 30 September on any other day, main included.golden/jev.jsonre-recorded; full check.py green locally.🤖 Generated with Claude Code