diff --git a/.github/workflows/ci.yml b/.github/workflows/ci.yml index ecdcd3e..985bb4d 100644 --- a/.github/workflows/ci.yml +++ b/.github/workflows/ci.yml @@ -64,3 +64,27 @@ jobs: path: | ci-results.xml ci-dependencies.txt + + report-test: + name: report appendix + runs-on: ubuntu-24.04 + steps: + - uses: actions/checkout@v4 + - uses: actions/setup-python@v5 + with: + python-version: '3.13' + cache: pip + cache-dependency-path: reports/requirements.txt + - name: Install report dependencies + run: python -m pip install --disable-pip-version-check -r reports/requirements.txt pytest==9.1.1 + - name: Verify report environment + run: | + python -m pip check + python -c "import matplotlib, pypdf, pytest; print(matplotlib.__version__, pypdf.__version__, pytest.__version__)" + - name: Test report artifact and PDF builders + env: + JEVANY_REPORT_TESTS_REQUIRED: '1' + run: >- + python -m pytest -q + tests/test_build_choice_readout_results.py + tests/test_build_external_report_appendix.py diff --git a/README.md b/README.md index 4ca72da..c298c71 100644 --- a/README.md +++ b/README.md @@ -140,7 +140,7 @@ where Jev helps and when to return control to the LLM. - [๐Ÿ› ๏ธ 1.2 JevAny Training](#training) - [๐Ÿš€ 1.3 JevAny Deployment](#deployment) - [๐Ÿค— 2. Pretrained Models](#pretrained-models) - - [Training-free letter readout](#letter-readout) + - [Training-free choice-token readout](#choice-readout) - [๐Ÿ“Š 3. Benchmark Results](#evaluation) - [โฑ๏ธ 3.1 Inference efficiency](#efficiency) - [๐Ÿ•น๏ธ 4. Examples & Test Environments](#examples--test-environments) @@ -296,30 +296,43 @@ Pointer and direct-token models share the same API. Pointer supports up to See [readout choices](docs/TRAINING.md#pointer-and-direct-token-readouts) for training and accuracy tradeoffs. -### Training-free letter readout +### Training-free choice-token readout -The experimental letter readout presents up to 26 options as Aโ€“Z, then sums -the frozen language model's next-token probability mass for every vocabulary -token that decodes exactly to that uppercase letter. It can run on a base model -without training, apply a JevAny adapter before the same readout, or combine the -letter and native pointer distributions. The letter path is text-only and does -not change the checkpoint metadata. Select it with `--readout letter` in -`jevany eval`, `jevany decide`, or `jevany serve`; the checkpoint-native -readout remains the default. +Choice-token readout assigns the ordered options to the exact one-token IDs +`Aโ€“Z, aโ€“z`, scores those 52 rows at the answer position, and renormalizes over +the available options. It needs no readout training and works on a frozen base, +after a JevAny Pointer or Direct-Token adapter, or in a log-linear stack with +the checkpoint-native distribution. It can change the ranking and the selected +answer; scalar temperature calibration only changes confidence. -[Method, limits and evaluation command](docs/LETTER_READOUT.md). +```bash +jevany eval --run SimpleJev/JevAny-Qwen3.5-4B-Direct-Token-LoRA \ + --suite /path/to/suite --out runs/choice --device cuda --readout choice +``` + +[Method, commands, limits and complete tables](docs/CHOICE_READOUT.md) ยท +[Machine-readable results](results/choice-readout-v2.json) + +[![Training-free choice-token, checkpoint-native, Transfer-dev-tuned blend, and external baseline accuracy on Typed Decisions and JevJudge text](docs/choice-readout-results.svg)](docs/CHOICE_READOUT.md#results) + +- Direct-Token 4B's Transfer-dev-tuned stack reaches **79.83%** on held-out + Transfer, **67.65%** on Typed Decisions, and **59.25%** on JevJudge text: + +0.96, +0.45, and +0.83 points over its native readout. +- Pointer 27B's tuned stack reaches **89.10%** on held-out Transfer and + **73.30%** on Typed, but drops from **66.44% to 64.36%** on JevJudge text. + The target-domain result decides whether stacking is useful. +- The frozen 4B choice path scores 52.75% on Typed; applying the Direct-Token + adapter raises the same choice path to 64.80%. On JevJudge text, the same + adapter slightly lowers it, from 58.70% to 57.60%. -[![Accuracy change from native pointer for letter readout and the fixed blend](docs/letter-readout-results.svg)](docs/LETTER_READOUT.md#results) +**Practical rule** -- The fixed 50/50 blend improved 27B on both diagnostics: +0.43 accuracy - points on JevBench public and +0.67 on Transfer-v9. Neither gain is - statistically significant. -- At 4B, the same blend gained +1.73 points on JevBench public but lost 0.10 - on Transfer-v9. Letter readout is a paired target-domain ablation, not a - universal upgrade. -- Letter-only measured a 1.05x/1.07x median speedup over native at 4B/27B - on one warmed H200 panel. Blending runs both paths and cost 1.81x/1.89x the - native median latency. +- Use choice-token readout when the model can reason over the options and the + bottleneck is extracting a decision; it cannot create missing task ability. +- Blend only when paired development errors show that native and choice rescue + each other, then freeze one weight before the target evaluation. +- Use temperature for probability calibration, not accuracy: it cannot change + the argmax. > **Benchmark units:** Cygnet's official **73.70** on JevBench v1.5.4 is a > four-axis composite over 1,624 open and sealed decisions, not accuracy. Its @@ -329,6 +342,7 @@ readout remains the default. ## ๐Ÿ“Š 3. Benchmark Results +The release table below uses each checkpoint's native readout. JevAny-Qwen3.8-27B leads both benchmarks and has the lowest NLL and Brier. Among 4B releases, direct-token leads on JevBench; pointer leads on Transfer. @@ -357,7 +371,8 @@ NLL, Brier and ECE are measured on Transfer. [Machine-readable results](results/model-family-v2.json) ยท [Method and ablation report](reports/JevAny_Tech_Report.pdf) -The external comparison uses the complete Typed Decisions test split, the full +The external comparison below also uses checkpoint-native readouts. It uses the +complete Typed Decisions test split, the full 3,220-record JevJudge multimodal suite, and its 724-record text slice. The same 13 models stay in the same order; `โ€”` means unsupported native input or no matching result. diff --git a/docs/CHOICE_READOUT.md b/docs/CHOICE_READOUT.md new file mode 100644 index 0000000..7444be9 --- /dev/null +++ b/docs/CHOICE_READOUT.md @@ -0,0 +1,216 @@ +# Training-free choice-token readout + +Choice-token readout turns an ordinary causal language model into a decision +readout without adding or training a head: + +1. Preserve the candidate order and assign the exact one-character IDs + `Aโ€“Z, aโ€“z`. +2. Ask for one case-sensitive option ID with thinking disabled. +3. At the answer position, sum the mass of every vocabulary token that decodes + exactly to each available ID. +4. Renormalize over the available IDs and map the distribution back to the + original option keys. + +This operation can change the ranking and the selected answer. It is therefore +not temperature calibration. A scalar temperature changes probability values +but cannot change argmax accuracy. A native + choice stack is a third operation: +it pools two distributions and can change the answer. + +The same readout works on a frozen base model or after loading a JevAny Pointer +or Direct-Token adapter. No additional readout training is performed. + +## Use it + +The checkpoint-native readout remains the default. Select choice-token readout +with `--readout choice`: + +```bash +jevany eval \ + --run SimpleJev/JevAny-Qwen3.5-4B-Direct-Token-LoRA \ + --suite /path/to/jevbench-public-v1.4.2.2 \ + --out runs/choice-readout/qwen35-4b-direct \ + --device cuda --readout choice --choice-temperature 1.0 + +jevany decide request.json \ + --checkpoint SimpleJev/JevAny-Qwen3.5-4B-LoRA \ + --device cuda --dtype bf16 --readout choice + +jevany serve \ + --checkpoint SimpleJev/JevAny-Qwen3.5-4B-LoRA \ + --device cuda --dtype bf16 --readout choice +``` + +Use `--choice-native-weight` to pool the choice and checkpoint-native +distributions. This executes both paths. `--choice-max-tokens` controls the +choice prompt limit. The former `letter` selector and `--letter-*` flags remain +hidden compatibility aliases. + +The matrix evaluator also accepts a frozen base directly: + +```bash +python -m scripts.evaluate_choice_readout_matrix \ + --base Qwen/Qwen3.5-4B \ + --revision 851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a \ + --base-load-path /path/to/Qwen3.5-4B \ + --model-name Frozen-Qwen3.5-4B \ + --out runs/choice-readout/base4 \ + --jevbench /path/to/jevbench-public-v1.4.2.2 \ + --transfer /path/to/transfer-v9 \ + --typed /path/to/typed-decisions \ + --jevjudge /path/to/jevjudge-public +``` + +## Limits + +- Text input only; JevJudge full multimodal is unsupported and is never + converted into a text-only score. +- 1โ€“52 options per question. +- Exact one-token aliases must exist in the tokenizer for every used ID. +- Choice-only runs use one chat-formatted prefill per question. A stack also + executes the checkpoint-native prefill. +- Matrix runs use exact math SDPA below 8,192 tokens and memory-linear fused + SDPA for longer prompts. All 724 JevJudge text records complete; no record is + truncated or rejected. + +## Evaluation protocol + +The v2 matrix uses the same prompt and data for five runs: + +- frozen Qwen3.5-4B; +- JevAny Qwen3.5-4B Pointer; +- JevAny Qwen3.5-4B Direct-Token; +- frozen Qwen3.8-27B; +- JevAny Qwen3.8-27B Pointer. + +Zero-shot rows use T=1 and no evaluation-label tuning. For tuned rows, each +model selects one native weight on the 1,046 clean/knowable Transfer-v9 +development decisions; fitted NLL and distance from 0.5 break accuracy ties. +One additional temperature is then fitted on the same development rows. The +weight and temperature are frozen before Transfer test, Typed Decisions, +JevJudge text, and the JevBench public diagnostic are scored. + +The native parity gates reproduce the released README results before any +held-out panel runs: + +| Checkpoint | Transfer development | JevBench public | +|---|---:|---:| +| JevAny 4B Pointer | 823/1,046 ยท 78.68% | 185/231 ยท 80.09% | +| JevAny 4B Direct-Token | 818/1,046 ยท 78.20% | 187/231 ยท 80.95% | +| JevAny 27B Pointer | 900/1,046 ยท 86.04% | 208/231 ยท 90.04% | + +### Evaluated model sources + +| Run | Canonical public repository | Evaluated revision | +|---|---|---| +| Frozen Qwen3.5-4B | `Qwen/Qwen3.5-4B` | `851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a` | +| JevAny 4B Pointer | `SimpleJev/JevAny-Qwen3.5-4B-LoRA` | `1c7aa9bab14ac347aeb917c0bcd757838a8a78ce` | +| JevAny 4B Direct-Token | `SimpleJev/JevAny-Qwen3.5-4B-Direct-Token-LoRA` | Not verified for the evaluated local release | +| Frozen Qwen3.8-27B | `Qwen/Qwen3.8-27B` | `1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0` | +| JevAny 27B Pointer | `SimpleJev/JevAny-Qwen3.8-27B-LoRA` | `09c9e9102d5b8cc7d56558d25da1202a761b6c0d` | + +The Direct-Token run is still byte-identifiable: its evaluated adapter is +`b768d54b9e5e20204b2db9e24c18ae6b10fa310b241d1931bac6b8882f127c53` +and its native head is +`d3909cffc156d077061114627c8aa22f60c7a8f1bdcfd20c489bb3bb009fc4a6` +(SHA-256). The evaluation manifest and release metadata do not prove which +immutable public-repository commit contains those exact files, so the result +artifact records its public revision as `null` rather than guessing. + +## Results + +![Training-free choice-token, checkpoint-native, Transfer-dev-tuned blend, and external baseline accuracy on Typed Decisions and JevJudge text](choice-readout-results.svg) + +All values below are accuracy. JevBench is its 231-item public development +diagnostic. Transfer test contains 1,046 clean/knowable held-out decisions. +Typed covers 2,000 decisions in 400 test cases. JevJudge text covers all 724 +text records; it is not the 3,220-record multimodal suite. + +| Weights | Readout | JevBench public | Transfer test | Typed test | JevJudge text | +|---|---|---:|---:|---:|---:| +| Frozen Qwen3.5-4B | Choice T=1 | 79.65% | 68.45% | 52.75% | 58.70% | +| JevAny 4B Pointer | Native shipped | 80.09% | 77.92% | **63.50%** | 58.01% | +| | Choice T=1 | **81.39%** | 73.61% | 57.80% | 52.62% | +| | Tuned stack ยท w=.78 | 80.95% | **78.01%** | 62.90% | **59.53%** | +| JevAny 4B Direct-Token | Native shipped | 80.95% | 78.87% | 67.20% | 58.43% | +| | Choice T=1 | **81.39%** | 78.59% | 64.80% | 57.60% | +| | Tuned stack ยท w=.59 | 80.95% | **79.83%** | **67.65%** | **59.25%** | +| Frozen Qwen3.8-27B | Choice T=1 | 87.88% | 80.69% | 68.25% | 61.74% | +| JevAny 27B Pointer | Native shipped | **90.04%** | 87.95% | 72.80% | **66.44%** | +| | Choice T=1 | 89.61% | 86.04% | 72.60% | 57.87% | +| | Tuned stack ยท w=.52 | **90.04%** | **89.10%** | **73.30%** | 64.36% | + +The fixed, untuned 50/50 stacks remain a separate zero-shot reference: + +| Checkpoint | JevBench public | Transfer test | Typed test | JevJudge text | +|---|---:|---:|---:|---:| +| JevAny 4B Pointer | 81.39% | 77.44% | 60.70% | 59.12% | +| JevAny 4B Direct-Token | 80.95% | 79.45% | 66.75% | 58.98% | +| JevAny 27B Pointer | 90.04% | **89.20%** | 73.20% | 64.23% | + +The selected parameters differ by checkpoint: + +| Checkpoint | Native weight | Additional temperature | +|---|---:|---:| +| JevAny 4B Pointer | .78 | 1.10 | +| JevAny 4B Direct-Token | .59 | 1.12 | +| JevAny 27B Pointer | .52 | .84 | + +### External comparison + +| Model/readout | Typed test | JevJudge text | +|---|---:|---:| +| meraGPT Decider 1 ยท published | **76.80%** | โ€” | +| JevAny 27B Pointer ยท native | 72.80% | **66.44%** | +| JevAny 27B Pointer ยท tuned stack | **73.30%** | 64.36% | +| TypeSafe Jev 1.13 | 72.70% | 65.06% | +| Kev-27B | โ€” | 64.23% | + +The 27B stack beats its native readout by 12 held-out Transfer decisions and +10 Typed decisions, while losing 15 JevJudge text decisions. Direct-Token 4B's +stack improves over native on all three held-out panels. Pointer 4B improves on +Transfer and JevJudge but loses on Typed. No single blend is universally best. + +### Practical principles + +- **Selection bottleneck:** use choice-token readout when the model can reason + over the candidates but the native extraction path is the bottleneck. +- **Complementary errors:** keep a stack only when paired development results + show that each readout rescues errors made by the other. Select one weight on + development data and freeze it. +- **Capability limit:** a different readout cannot supply knowledge, perception, + planning, or instruction-following ability missing from the underlying model. +- **Calibration limit:** temperature can improve NLL, Brier, or ECE, but it + cannot change the selected option or accuracy. +- **Domain shift:** target-domain validation wins over a global rule. The same + 27B stack helps Transfer and Typed, ties JevBench, and hurts JevJudge text. + +Complete metrics, hashes, runtime versions, selected weights, and the unsupported +full-JevJudge marker are in +[`results/choice-readout-v2.json`](../results/choice-readout-v2.json). The +deterministic builder is +[`scripts/build_choice_readout_results.py`](../scripts/build_choice_readout_results.py). + +## Benchmark units + +Cygnet's official `73.70` on JevBench v1.5.4 is a four-axis composite over +1,624 open and sealed decisions, not accuracy. Its separate public-development +result is 203/231 (87.9% accuracy). Every JevBench value on this page is +accuracy over the same 231 public-development items and cannot be compared +numerically with the composite. + +## Historical v1 + +The original prompt-v1 experiment and one-panel latency diagnostic remain +unchanged in [`results/letter-readout-v1.json`](../results/letter-readout-v1.json) +and [`letter-readout-results.svg`](letter-readout-results.svg). Those runs used a +different prompt and runtime, and their matched native reruns did not reproduce +the canonical README rows. They are retained for audit and are not mixed into +the v2 tables. + +## Attribution + +The Cygnet-compatible prompt and exact option-ID aggregation semantics are +adapted from `blockbrain-ai/cygnet-recipe` commit `3cf591c`, Copyright 2026 +Nood Co and contributors, under MIT. Cygnet credits the one-token option-ID +readout to [NInfer](https://github.com/igorls/ninfer), released under Apache-2.0. +JevAny does not incorporate NInfer source code. See [`NOTICE`](../NOTICE). diff --git a/docs/EVALUATION.md b/docs/EVALUATION.md index b6e8999..a7db080 100644 --- a/docs/EVALUATION.md +++ b/docs/EVALUATION.md @@ -2,7 +2,7 @@ ## Model family v2 -This section contains the full results for the five current LoRA SFT releases. +This section contains checkpoint-native results for the five current LoRA SFT releases. Transfer also informed model development and serves as a diagnostic comparison. **Transfer** is the cross-domain evaluation built from @@ -16,7 +16,7 @@ and ECE in the main table are Transfer metrics; every run covers every item. Every JevBench value in this document is public-development accuracy (`correct / 231`). It is not the official JevBench v1.5.4 composite, which combines Intelligence, Calibration, Speed and Cost over 1,624 open and sealed -decisions. See the [letter-readout guide](LETTER_READOUT.md#benchmark-units) for +decisions. See the [choice-token readout guide](CHOICE_READOUT.md#benchmark-units) for the side-by-side definitions. | Model | Transfer โ†‘ | JevBench โ†‘ | NLL โ†“ | Brier โ†“ | ECE โ†“ | @@ -95,7 +95,7 @@ a complete local Transfer API run and JevBench's published per-tier accuracy. - **Loss:** cross-entropy, pure InfoNCE, and mixed objectives; CE gave the best Transfer accuracy in the loss sweep, while small contrastive terms mainly improved calibration. The released checkpoints use CE. -- **Readout family:** at 4B, direct-token improves JevBench (80.95% vs 80.09%), +- **Trained checkpoint-native readout family:** at 4B, direct-token improves JevBench (80.95% vs 80.09%), while pointer is slightly stronger on Transfer (78.68% vs 78.20%). ### Reproducibility @@ -140,6 +140,35 @@ excluded. Gemma uses run/checkpoint timestamps because its earlier checkpoint format did not store cumulative elapsed seconds; the other figures come from checkpoint or terminal trainer telemetry. +## Training-free choice-token readout + +Choice-token readout constrains the next token to one of 52 exact option IDs +and renormalizes their mass. It needs no additional training and can change the +selected answer. Scalar temperature fitting is separate: it changes confidence, +not accuracy. A native + choice stack can change both. + +The stack weight and additional temperature below are selected only on 1,046 +clean/knowable Transfer-v9 development decisions and then frozen. Transfer test, +Typed Decisions, and JevJudge text are held out. JevBench is a public diagnostic. +All values are accuracy percentages. + +| Weights | Readout | JevBench public | Transfer test | Typed test | JevJudge text | +|---|---|---:|---:|---:|---:| +| Frozen Qwen3.5-4B | Choice T=1 | 79.65% | 68.45% | 52.75% | 58.70% | +| JevAny 4B Pointer | Native / Choice / Tuned | 80.09 / 81.39 / 80.95 | 77.92 / 73.61 / **78.01** | **63.50** / 57.80 / 62.90 | 58.01 / 52.62 / **59.53** | +| JevAny 4B Direct-Token | Native / Choice / Tuned | 80.95 / **81.39** / 80.95 | 78.87 / 78.59 / **79.83** | 67.20 / 64.80 / **67.65** | 58.43 / 57.60 / **59.25** | +| Frozen Qwen3.8-27B | Choice T=1 | 87.88% | 80.69% | 68.25% | 61.74% | +| JevAny 27B Pointer | Native / Choice / Tuned | **90.04** / 89.61 / **90.04** | 87.95 / 86.04 / **89.10** | 72.80 / 72.60 / **73.30** | **66.44** / 57.87 / 64.36 | + +The 4B Direct-Token stack improves over its native readout on all three held-out +panels. The 27B stack improves Transfer and Typed, ties JevBench, and hurts +JevJudge text. Retain a stack only when paired development errors are +complementary and the target distribution validates it. + +[Full method, fixed 50/50 controls and protocol](CHOICE_READOUT.md) ยท +[Machine-readable results](../results/choice-readout-v2.json) ยท +[Merged 23-page technical report](../reports/JevAny_Tech_Report.pdf) + ## Earlier releases and evaluations The [JevBench and Kev comparison](EXTERNAL_EVALUATION.md) evaluates both released diff --git a/docs/EXTERNAL_EVALUATION.md b/docs/EXTERNAL_EVALUATION.md index 731db1d..004551c 100644 --- a/docs/EXTERNAL_EVALUATION.md +++ b/docs/EXTERNAL_EVALUATION.md @@ -152,9 +152,12 @@ and are not used in the figure. [`external-zero-shot-v1.json`](../results/external-zero-shot-v1.json), and the deterministic renderer is [`plot_external_zero_shot.py`](../scripts/plot_external_zero_shot.py). -- The two external-evaluation pages appended to the single merged technical - report are generated by - [`build_external_report_appendix.py`](../scripts/build_external_report_appendix.py). +- Four evaluation pages are appended to the single merged technical report by + [`build_external_report_appendix.py`](../scripts/build_external_report_appendix.py): + two external-evaluation pages in Appendix L and two training-free + choice-token pages in Appendix M. Appendix M reads + `results/choice-readout-v2.json` directly; it does not rasterize or convert + the README SVG. - The tracked artifact contains the aggregate rows needed to regenerate the figure. Raw local run locations and SHA-256 hashes are recorded for audit but are not included in the repository. JevAny full results come from @@ -190,20 +193,25 @@ build/report-repro-env/bin/python scripts/plot_external_zero_shot.py build/report-repro-env/bin/python scripts/build_external_report_appendix.py \ --data results/external-zero-shot-v1.json \ --chart docs/external-zero-shot.png \ - --output build/JevAny_Tech_Report_External_Evaluation_Appendix.pdf \ + --choice-data results/choice-readout-v2.json \ + --output build/JevAny_Tech_Report_Evaluation_Appendices.pdf \ --base-report reports/JevAny_Tech_Report.pdf \ --merged-output reports/JevAny_Tech_Report.pdf \ --base-pages 19 +cp reports/JevAny_Tech_Report.pdf site/files/reports/JevAny_Tech_Report.pdf ``` -`--base-pages 19` always takes the original report and agent-harness pages and -discards any previously appended external pages before adding the newly built -two-page appendix. The base and merged paths may therefore be identical without -growing the report on repeated runs. Both outputs are written to temporary -files in their destination directories and installed with `os.replace`; an -interrupted build cannot leave a partial tracked PDF. Appendix and merged PDF -metadata use the fixed creation and modification timestamp -`D:20261001000000Z`. With the pinned inputs and dependencies, repeated commands +`--base-pages 19` retains the report and agent-harness pages, synchronizes the +superseded 27B release headline fields to step 44,319, and discards any +previously appended evaluation pages before rebuilding Appendices L and M. The +result is exactly 23 pages: 19 retained pages plus four generated pages. The +base and merged paths may therefore be identical without growing the report on +repeated runs. Both outputs are written to temporary files in their destination +directories and installed with `os.replace`; an interrupted build cannot leave +a partial tracked PDF. `reports/` contains only one PDFโ€”the merged report; the +`site/files/` copy is its byte-identical deployment mirror. Appendix and merged +PDF metadata use the fixed creation and modification timestamp +`D:20261002000000Z`. With the pinned inputs and dependencies, repeated commands produce byte-identical appendix and merged files. ### JevBench v1.5.4 diff --git a/docs/LETTER_READOUT.md b/docs/LETTER_READOUT.md index 789f56c..136b12f 100644 --- a/docs/LETTER_READOUT.md +++ b/docs/LETTER_READOUT.md @@ -1,170 +1,7 @@ -# Training-free option-letter readout +# Moved: training-free choice-token readout -JevAny includes an experimental evaluator for a training-free, single-answer-slot -decision readout. It adapts the prompt and probability-aggregation semantics -from [Cygnet](https://github.com/blockbrain-ai/cygnet-recipe/tree/3cf591c692dec649f7c134449814610307c7bb3a): +The public method name is **choice-token readout**. See +[`CHOICE_READOUT.md`](CHOICE_READOUT.md) for the method, commands, limits, +current results, and retained historical v1 links. -1. Render the options in their original order as `A`, `B`, โ€ฆ, `Z`. -2. Ask the model for exactly one option letter with thinking disabled. -3. At the answer position, collect every vocabulary token whose decoded surface - is exactly each available uppercase letter. This emulates Cygnet's - `structured_outputs.choice` mask; leading-space, punctuated, and lowercase - forms are not admitted. -4. Sum duplicate-token mass per letter and normalize over the available options. -5. Optionally apply a temperature fitted for that model and evaluation domain. - -The model produces no explanation and the method needs no additional training. -The same readout can be applied after loading a JevAny LoRA. For pointer -checkpoints, the letter and native pointer distributions can also be combined -log-linearly. The integrated commands call this knob -`--letter-pointer-weight`; the standalone experiment script calls it -`--pointer-weight`. - -## Run the public diagnostic - -Use `--readout letter` to score a JevAny checkpoint on any regular frozen suite -or labelled JSONL. This example keeps temperature at 1 rather than copying a -value fitted for another model: - -```bash -jevany eval \ - --run SimpleJev/JevAny-Qwen3.5-4B-LoRA \ - --suite /path/to/jevbench-public-v1.4.2.2 \ - --out runs/letter-readout/qwen35-4b \ - --device cuda --readout letter --letter-temperature 1.0 -``` - -The same selector works for one local decision or an HTTP deployment. Native -checkpoint readout remains the default when `--readout` is omitted: - -```bash -jevany decide request.json \ - --checkpoint SimpleJev/JevAny-Qwen3.5-4B-LoRA \ - --device cuda --dtype bf16 --readout letter - -jevany serve \ - --checkpoint SimpleJev/JevAny-Qwen3.5-4B-LoRA \ - --device cuda --dtype bf16 --readout letter -``` - -The standalone evaluator additionally supports a frozen base via `--base`, -tier sampling, and explicit LoRA scaling: - -```bash -python scripts/evaluate_letter_readout.py \ - --checkpoint SimpleJev/JevAny-Qwen3.5-4B-LoRA \ - --suite /path/to/jevbench-public-v1.4.2.2 \ - --out runs/letter-readout/qwen35-4b \ - --device cuda --dtype bf16 --temperature 1.0 -``` - -Add `--sample-per-tier 1` for a three-record smoke test. Use -`--lora-scale 0` with a checkpoint to evaluate its exact pinned base without -the adapter. A nonzero `--pointer-weight` requires a native pointer checkpoint -and adds a second, pointer-formatted model pass per request. - -Current limits: - -- text-only model input; -- 1โ€“26 options per question; -- safetensors weights for models with an untied language-model output head; -- one letter-formatted prefill per question, plus one native prefill when - pointer blending is enabled; -- native inference-limit and CUDA-graph flags do not apply to the chat-formatted - letter path; use `--letter-max-tokens` for its prompt limit. - -Temperature changes reported probabilities but does not change the letter-only -argmax. Fit it on a separate calibration split for each model. Cygnet's `3.4` -was fitted for its own Gemma configuration and is not a default for JevAny. - -## Benchmark units - -The official score and the public diagnostic answer different questions: - -| Result | Evaluation set | Unit | -|---|---:|---| -| Cygnet 73.70 | JevBench v1.5.4, 1,624 decisions: 904 open + 720 sealed | Equal-weight harmonic mean of Intelligence, Calibration, Speed and Cost | -| Cygnet 203/231 (87.9%) | Public development set, 231 decisions | Accuracy | -| JevAny results in the current README | Same public development set, 231 decisions | Accuracy | - -Cygnet ranks first among 106 systems on the official v1.5.4 composite and is a -statistical tie with Winnow-12B Q8. The 73.70 composite is not `73.70%` and -cannot be compared numerically with public-set accuracy. See Cygnet's -[official-submission report](https://github.com/blockbrain-ai/cygnet-recipe/pull/4) -and the [JevBench v1.5.4 board](https://benchmarkheaven.com/jev-models/v1.5.4). - -## Results - -All JevAny values below are local pinned-checkpoint runs at temperature 1. The -complete metrics, suite hashes, report hashes, and paired counts are in -[`results/letter-readout-v1.json`](../results/letter-readout-v1.json). -These are separate matched reruns of the pinned released checkpoints for this -ablation; they do not replace the main model-family table, whose frozen -checkpoints and run provenance differ. - -![Accuracy-point change from native pointer for letter and fixed-blend readouts](letter-readout-results.svg) - -JevBench public is a 231-question development diagnostic, not the sealed -v1.5.4 board. `Base + letter` disables the JevAny adapter; the other columns use -the released adapter. Deltas and two-sided exact McNemar p-values compare each -adapter readout with the native pointer on the same questions. - -| Model | Base + letter | Native | Letter | Fixed 50/50 blend | -|---|---:|---:|---:|---:| -| Qwen3.5-4B | 79.65% | 80.09% | 81.39% (+1.30, p=.749) | **81.82%** (+1.73, p=.424) | -| Qwen3.8-27B | 88.31% | 89.61% | 89.61% (+0.00, p=1.000) | **90.04%** (+0.43, p=1.000) | - -Transfer-v9 evaluates all 1,264 requests without rejection or truncation. Its -accuracy headline uses the 1,046 clean knowable decisions. - -| Model | Native | Letter | Fixed 50/50 blend | -|---|---:|---:|---:| -| Qwen3.5-4B | **79.16%** | 75.72% (-3.44, p=.0028) | 79.06% (-0.10, p=1.000) | -| Qwen3.8-27B | 86.23% | 84.23% (-2.01, p=.0375) | **86.90%** (+0.67, p=.337) | - -The 27B blend gets 909/1,046 decisions right versus 902/1,046 for native, but -the seven-question gain is not statistically significant. The fixed blend is -also not a universal improvement: at 4B it gets one fewer answer right. None of -the positive gains in either table is significant at the 0.05 level. The -letter-only Transfer-v9 losses show that a public-diagnostic gain does not by -itself establish transfer. - -## Latency diagnostic - -One warmed run per mode on H200/BF16 measured the same heterogeneous 231-record -panel with 16 warmups, one measured repeat, and concurrency 1. Values are external -end-to-end median / p95 milliseconds; loading and network transport are -excluded. - -| Model | Native | Letter | Fixed 50/50 blend | -|---|---:|---:|---:| -| Qwen3.5-4B | 138.43 / 200.98 | **131.86** / 201.55 (1.05x median speedup) | 250.34 / 388.31 (1.81x median latency) | -| Qwen3.8-27B | 194.70 / 477.67 | **181.58** / 492.51 (1.07x median speedup) | 367.40 / 943.44 (1.89x median latency) | - -Letter-only has a modest median improvement on this panel despite using more -logical input tokens on average (701 versus 602); p95 does not improve. The -blend evaluates both prefills and averages 1,303 logical input tokens. This is -not saturated server throughput or a general speed claim. - -## Practical principles - -- Treat letter readout as a model-and-domain ablation, not a drop-in upgrade. - Keep it only after a paired evaluation on the target decision distribution. -- Blend only when the two readouts make complementary errors. The fixed 50/50 - pool helped 27B on both diagnostics; at 4B it helped JevBench public but not - Transfer-v9. Select the weight on a separate development split. -- Calibrate each model, readout, and domain separately. Cygnet's fitted - temperature `3.4` does not transfer to these Qwen checkpoints. -- Re-measure latency in the deployment runtime and traffic mix. The one-pass - letter path can trim median latency, while blending requires both letter and - native passes and nearly doubles median latency here. - -## Attribution - -The Cygnet-compatible prompt and aggregation semantics are adapted from -`blockbrain-ai/cygnet-recipe` commit -`3cf591c692dec649f7c134449814610307c7bb3a`, Copyright 2026 Nood Co and -contributors, under MIT. Cygnet credits the one-token option-letter readout -method to [NInfer](https://github.com/igorls/ninfer), released under Apache-2.0. -JevAny does not incorporate NInfer source code. See [NOTICE](../NOTICE) for the -full Cygnet license notice. +The former `letter` CLI/API spelling remains a compatibility alias. diff --git a/docs/choice-readout-results.svg b/docs/choice-readout-results.svg new file mode 100644 index 0000000..a32e254 --- /dev/null +++ b/docs/choice-readout-results.svg @@ -0,0 +1,928 @@ + + + + Training-free choice readout accuracy + + + + + Training-free choice readout accuracy + Typed Decisions: Frozen Qwen3.5 4B ยท Choice 52.75%, JevAny 4B Pointer ยท Native 63.50%, JevAny 4B Pointer ยท Choice 57.80%, JevAny 4B Direct-Token ยท Native 67.20%, JevAny 4B Direct-Token ยท Choice 64.80%, Frozen Qwen3.8 27B ยท Choice 68.25%, JevAny 27B Pointer ยท Native 72.80%, JevAny 27B Pointer ยท Choice 72.60%, JevAny 4B Pointer ยท Tuned blend 62.90%, JevAny 4B Direct-Token ยท Tuned blend 67.65%, JevAny 27B Pointer ยท Tuned blend 73.30%, meraGPT Decider 1 ยท published 76.80%, TypeSafe Jev 1.13 72.70% JevJudge text: Frozen Qwen3.5 4B ยท Choice 58.70%, JevAny 4B Pointer ยท Native 58.01%, JevAny 4B Pointer ยท Choice 52.62%, JevAny 4B Direct-Token ยท Native 58.43%, JevAny 4B Direct-Token ยท Choice 57.60%, Frozen Qwen3.8 27B ยท Choice 61.74%, JevAny 27B Pointer ยท Native 66.44%, JevAny 27B Pointer ยท Choice 57.87%, JevAny 4B Pointer ยท Tuned blend 59.53%, JevAny 4B Direct-Token ยท Tuned blend 59.25%, JevAny 27B Pointer ยท Tuned blend 64.36%, TypeSafe Jev 1.13 65.06%, Kev-27B ยท open 64.23% + image/svg+xml + + + Matplotlib v3.11.2, https://matplotlib.org/ + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + 0 + + + + + + + + + 20 + + + + + + + + + 40 + + + + + + + + + 60 + + + + + + + + + 80 + + + + + + + + Frozen Qwen3.5 4B ยท Choice + + + + + + JevAny 4B Pointer ยท Native + + + + + + JevAny 4B Pointer ยท Choice + + + + + + JevAny 4B Direct-Token ยท Native + + + + + + JevAny 4B Direct-Token ยท Choice + + + + + + Frozen Qwen3.8 27B ยท Choice + + + + + + JevAny 27B Pointer ยท Native + + + + + + JevAny 27B Pointer ยท Choice + + + + + + JevAny 4B Pointer ยท Tuned blend + + + + + + JevAny 4B Direct-Token ยท Tuned blend + + + + + + JevAny 27B Pointer ยท Tuned blend + + + + + + meraGPT Decider 1 ยท published + + + + + + TypeSafe Jev 1.13 + + + + + + + + + + + 52.8 + + + + + + 63.5 + + + + + + 57.8 + + + + + + 67.2 + + + + + + 64.8 + + + + + + 68.2 + + + + + + 72.8 + + + + + + 72.6 + + + + + + 62.9 + + + + + + 67.7 + + + + + + 73.3 + + + + + + 76.8 + + + + + + 72.7 + + + ZERO-SHOT + + + TRANSFER-DEV-TUNED + + + EXTERNAL REFERENCE + + + 2,000 decisions ยท accuracy (%) + + + Typed Decisions + + + + + + + + + + + + + + + + + + + + + + + 0 + + + + + + + + + 20 + + + + + + + + + 40 + + + + + + + + + 60 + + + + + + + + + 80 + + + + + + + + Frozen Qwen3.5 4B ยท Choice + + + + + + JevAny 4B Pointer ยท Native + + + + + + JevAny 4B Pointer ยท Choice + + + + + + JevAny 4B Direct-Token ยท Native + + + + + + JevAny 4B Direct-Token ยท Choice + + + + + + Frozen Qwen3.8 27B ยท Choice + + + + + + JevAny 27B Pointer ยท Native + + + + + + JevAny 27B Pointer ยท Choice + + + + + + JevAny 4B Pointer ยท Tuned blend + + + + + + JevAny 4B Direct-Token ยท Tuned blend + + + + + + JevAny 27B Pointer ยท Tuned blend + + + + + + TypeSafe Jev 1.13 + + + + + + Kev-27B ยท open + + + + + + + + + + + 58.7 + + + + + + 58.0 + + + + + + 52.6 + + + + + + 58.4 + + + + + + 57.6 + + + + + + 61.7 + + + + + + 66.4 + + + + + + 57.9 + + + + + + 59.5 + + + + + + 59.3 + + + + + + 64.4 + + + + + + 65.1 + + + + + + 64.2 + + + ZERO-SHOT + + + TRANSFER-DEV-TUNED + + + EXTERNAL REFERENCE + + + 724 records ยท accuracy (%) + + + JevJudge text + + + + Training-free choice readout + + + Zero-shot readouts and Transfer-dev-tuned blends ยท no Typed or JevJudge labels used for tuning + + + JevJudge full is unsupported for this text-only readout. Choice temperature calibration is omitted here because it cannot change accuracy. + + + + + + + Choice T=1 ยท zero-shot + + + + + + Native shipped ยท zero-shot + + + + + + Transfer-dev-tuned blend + + + + + + TypeSafe Jev + + + + + + Published / open baseline + + + + + + + + + + + + + + + + + + + + + + + + + + diff --git a/jevany/__init__.py b/jevany/__init__.py index d999df1..a3ca0e6 100644 --- a/jevany/__init__.py +++ b/jevany/__init__.py @@ -5,8 +5,12 @@ from .api import Choice, Noul, Score, SystemOneRequest from .client import JevClient from .inference import InferenceOptions +from .readout import ChoiceReadoutOptions -__all__ = ["Choice", "Noul", "Score", "SystemOneRequest", "JevClient", "JevModel", "InferenceOptions"] +__all__ = [ + "Choice", "Noul", "Score", "SystemOneRequest", "JevClient", "JevModel", + "InferenceOptions", "ChoiceReadoutOptions", +] def __getattr__(name: str): diff --git a/jevany/benchmark.py b/jevany/benchmark.py index 0a361ca..5ae3a5c 100644 --- a/jevany/benchmark.py +++ b/jevany/benchmark.py @@ -26,7 +26,7 @@ from jevany.metrics import EPSILON, grouped_metrics, metrics, unknowable_report from jevany.model import ContextLengthError from jevany.predictors import LocalPredictor, RemotePredictor -from jevany.readout import add_readout_arguments, letter_options_from_args +from jevany.readout import add_readout_arguments, choice_options_from_args, normalize_readout from jevany.suite import ENCODING, digest, load_split, read_manifest, record_digest, write_json @@ -170,7 +170,7 @@ def evaluate_records(records, predictor, directory, temperature=1.0, heldout_sou EXAMPLES = """examples: jevany eval --run runs/my-jev --data data/starter/development.jsonl --out runs/my-jev/eval jevany eval --run SimpleJev/JevAny-Qwen3.5-4B-LoRA --suite data/eval-suite --out runs/eval - jevany eval --run SimpleJev/JevAny-Qwen3.5-4B-LoRA --readout letter --suite data/eval-suite --out runs/letter + jevany eval --run SimpleJev/JevAny-Qwen3.5-4B-LoRA --readout choice --suite data/eval-suite --out runs/choice jevany eval --remote http://127.0.0.1:8008 --data my-labelled.jsonl --out runs/remote-eval --data scores your own labelled JSONL (jevany.data.load_records); --suite scores a @@ -195,12 +195,12 @@ def main(argv=None, prog=None): ap.add_argument("--date_facts", action="store_true", help="apply jevany.api.with_date_facts to every state before scoring (the opt-in serving preprocessor); reported in report.json") a = ap.parse_args(argv) try: - letter_options = letter_options_from_args(a) + choice_options = choice_options_from_args(a) except ValueError as error: ap.error(str(error)) if bool(a.run) == bool(a.remote): ap.error("give exactly one of --run or --remote") - if a.remote and letter_options is not None: - ap.error("--readout letter is a local model setting; configure it on the remote server") + if a.remote and choice_options is not None: + ap.error("--readout choice is a local model setting; configure it on the remote server") if bool(a.suite) == bool(a.data): ap.error("give exactly one of --suite or --data") if a.data: records, heldout, split, source_hash = load_records(a.data), [], "custom", digest(Path(a.data)) @@ -212,34 +212,40 @@ def main(argv=None, prog=None): records = [{**r, "state": with_date_facts(r["state"])} for r in records] if a.remote: predictor = RemotePredictor(a.remote, a.remote_model, os.environ.get("JEVANY_REMOTE_API_KEY", "local")) - elif letter_options is not None: + elif choice_options is not None: from jevany.letter_predictor import LetterReadoutPredictor predictor = LetterReadoutPredictor( checkpoint=a.run, device=a.device, options=LoadOptions.from_env(), - temperature=letter_options.temperature, - pointer_weight=letter_options.pointer_weight, - max_tokens=letter_options.max_tokens, + temperature=choice_options.temperature, + native_weight=choice_options.native_weight, + max_tokens=choice_options.max_tokens, ) else: predictor = LocalPredictor(a.run, a.device, LoadOptions.from_env()) report, _ = evaluate_records(records, predictor, a.out, heldout_sources=tuple(heldout), skip_overlong=bool(a.data)) - pointer_temperature = (getattr(predictor.pointer_model, "temperature", None) - if letter_options is not None and predictor.pointer_weight else None) + native_temperature = (getattr(predictor.native_model, "temperature", None) + if choice_options is not None and predictor.native_weight else None) calibration_applied = (None if a.remote else predictor.temperature != 1.0 - or pointer_temperature not in (None, 1.0)) + or native_temperature not in (None, 1.0)) report.update(suite_sha256=source_hash, data=a.data, date_facts=a.date_facts, run=a.run or a.remote, split=split, calibration_applied=calibration_applied, base_loading=getattr(predictor, "base_loading", None), remote={"base_url": a.remote, "requested_model": a.remote_model, "served_model": predictor.served_model} if a.remote else None) if not a.remote: - report["readout"] = a.readout - if letter_options is not None: - report["letter_readout"] = dict(predictor.provenance) - if pointer_temperature is not None: - report["letter_readout"]["pointer_temperature"] = pointer_temperature - report["calibration"]["pointer_temperature"] = pointer_temperature + report["readout"] = normalize_readout(a.readout) + if choice_options is not None: + choice_readout = dict(predictor.provenance) + if native_temperature is not None: + choice_readout["native_temperature"] = native_temperature + choice_readout["pointer_temperature"] = native_temperature + report["calibration"]["native_temperature"] = native_temperature + report["calibration"]["pointer_temperature"] = native_temperature + report["choice_readout"] = choice_readout + # Kept as an output alias for consumers of reports written before the + # public readout name changed to ``choice``. + report["letter_readout"] = dict(choice_readout) write_json(Path(a.out) / "report.json", report) print(json.dumps({"objective": report["objective"], "clean": report["clean"], "coverage": report["coverage"]}, indent=2)) diff --git a/jevany/cli.py b/jevany/cli.py index 055a321..a04a0d7 100644 --- a/jevany/cli.py +++ b/jevany/cli.py @@ -45,7 +45,7 @@ def decide_main(argv: list[str]) -> None: from dataclasses import fields from .inference import InferenceOptions, add_inference_arguments, inference_options_from_args from .placement import add_placement_arguments - from .readout import add_readout_arguments, letter_options_from_args + from .readout import add_readout_arguments, choice_options_from_args parser = argparse.ArgumentParser(prog="jevany decide") parser.add_argument("request", help="JSON request file; - reads stdin") @@ -63,7 +63,7 @@ def decide_main(argv: list[str]) -> None: add_inference_arguments(parser) args = parser.parse_args(argv) try: - letter_options = letter_options_from_args(args) + choice_options = choice_options_from_args(args) except ValueError as error: parser.error(str(error)) content = sys.stdin.read() if args.request == "-" else Path(args.request).read_text(encoding="utf-8") @@ -73,12 +73,12 @@ def decide_main(argv: list[str]) -> None: from dataclasses import replace from .checkpoint import LoadOptions, load_options_from_args from .runtime import JevModel - if letter_options is not None and ( + if choice_options is not None and ( args.cuda_graphs or args.cuda_graph_max_tokens is not None or any(getattr(args, item.name) is not None for item in fields(InferenceOptions)) ): - parser.error("CUDA graph and native inference-limit flags cannot be used with --readout letter; " - "use --letter-max-tokens") + parser.error("CUDA graph and native inference-limit flags cannot be used with --readout choice; " + "use --choice-max-tokens") options = load_options_from_args(args) if args.cuda_graphs or args.cuda_graph_max_tokens is not None: options = options or LoadOptions.from_env() @@ -87,11 +87,11 @@ def decide_main(argv: list[str]) -> None: if args.cuda_graph_max_tokens is not None else options.cuda_graph_max_tokens)) settings = ({ - "readout": "letter", - "letter_temperature": letter_options.temperature, - "letter_pointer_weight": letter_options.pointer_weight, - "letter_max_tokens": letter_options.max_tokens, - } if letter_options is not None else { + "readout": "choice", + "choice_temperature": choice_options.temperature, + "choice_native_weight": choice_options.native_weight, + "choice_max_tokens": choice_options.max_tokens, + } if choice_options is not None else { "inference_options": inference_options_from_args(args), }) client = JevModel.from_pretrained( diff --git a/jevany/demos/static/app.js b/jevany/demos/static/app.js index b7a51d1..2f43c48 100644 --- a/jevany/demos/static/app.js +++ b/jevany/demos/static/app.js @@ -217,7 +217,8 @@ async function liveStep(action) { } function renderConnection() { const media = link.media || {}, served = link.served; - const readout = served && served.readout === "letter" ? "letter" : served && served.decision_mode; + const readout = served && (served.readout === "choice" || served.readout === "letter") + ? "choice" : served && served.decision_mode; $("connect-url").disabled = $("connect-model").disabled = !link.editable; if (!$("connect-url").value) $("connect-url").value = link.base_url || link.default_base_url || ""; if (!$("connect-model").value && link.model && link.model !== "jevany-latest") $("connect-model").value = link.model; diff --git a/jevany/letter_predictor.py b/jevany/letter_predictor.py index f5c9e4b..ea9c656 100644 --- a/jevany/letter_predictor.py +++ b/jevany/letter_predictor.py @@ -1,14 +1,15 @@ -"""Training-free option-letter inference over a frozen base or JevAny adapter. +"""Training-free choice-token inference over a frozen base or JevAny adapter. -This is the in-process backend for :mod:`jevany.letter_readout`. It keeps the -existing pointer path unchanged: a JevAny checkpoint can instead be applied to -the chat prompt before the frozen vocabulary readout, and its pointer -distribution can optionally be combined with the letter distribution. +This is the in-process backend for :mod:`jevany.letter_readout`. A JevAny +checkpoint can be applied to the chat prompt before the frozen vocabulary +readout, and its Pointer or Direct-Token native distribution can optionally be +combined with the choice distribution. """ from __future__ import annotations from collections.abc import Mapping +from contextlib import nullcontext import hashlib import json import math @@ -17,13 +18,14 @@ import torch import torch.nn.functional as F +from torch.nn.attention import SDPBackend, sdpa_kernel from .api import SystemOneRequest, question_keys from .checkpoint import Checkpoint, LoadOptions from .data import api_request, materialize from .device import sync from .letter_readout import ( - LETTERS, + CHOICE_SYMBOLS, SYSTEM_PROMPT, build_prompt, letter_token_ids, @@ -61,17 +63,19 @@ def _safe_chat_text(tokenizer, text: str) -> str: return escaped -def _release_unused_output_head(decision_model) -> bool: - """Drop references to the full-vocabulary head unused by letter readout.""" +def _release_unused_output_head(decision_model, *, retain_native: bool = False) -> bool: + """Drop the full-vocabulary head unless a native LM-token pass needs it.""" adapter = getattr(decision_model, "adapter", None) adapter_head = getattr(adapter, "_output_embeddings", None) native_head = getattr(decision_model, "lm_head", None) + if retain_native: + return False if adapter_head is None and native_head is None: return False - # The exact letter rows are copied immediately after this call. Native - # lm-token inference is never invoked by this predictor, so both aliases - # can be detached even for direct-token checkpoints. + # Exact choice rows are copied from the tied embedding table or loaded + # independently from safetensors. With no LM-token native pass, both + # aliases can therefore be detached to release the full vocabulary head. if native_head is not None: decision_model.lm_head = None if adapter_head is not None: @@ -223,9 +227,10 @@ class LetterReadoutPredictor: """Benchmark predictor for an exact constrained first-token readout. Give either ``base`` for a frozen training-free model, or ``checkpoint`` - to apply a JevAny LoRA before the same readout. ``pointer_weight > 0`` - additionally pools the checkpoint's native pointer probabilities in log - space; this costs one extra pointer-formatted prefill per request. + to apply a JevAny LoRA before the same readout. ``native_weight > 0`` + additionally pools the checkpoint's native probabilities in log space; + this costs one extra native-formatted prefill per request. The legacy + ``pointer_weight`` spelling remains an alias for ``native_weight``. """ def __init__( @@ -237,7 +242,11 @@ def __init__( options: LoadOptions | None = None, revision: str | None = None, temperature: float = 1.0, - pointer_weight: float = 0.0, + pointer_weight: float | None = None, + native_weight: float | None = None, + return_components: bool = False, + exact_kernels: bool = True, + efficient_long_context_tokens: int | None = None, max_tokens: int = 16_384, ) -> None: if (base is None) == (checkpoint is None): @@ -247,21 +256,66 @@ def __init__( self.temperature = float(temperature) if not math.isfinite(self.temperature) or self.temperature <= 0: raise ValueError("temperature must be finite and positive") + if pointer_weight is not None and native_weight is not None: + raise ValueError("give at most one of native_weight or pointer_weight") + weight_source = ( + "native_weight" if native_weight is not None + else "pointer_weight" if pointer_weight is not None + else "default" + ) + selected_weight = ( + native_weight if native_weight is not None + else pointer_weight if pointer_weight is not None + else 0.0 + ) # Reuse the same validation as the actual combination path. - geometric_blend([1.0], [1.0], pointer_weight) - self.pointer_weight = float(pointer_weight) - if checkpoint is None and self.pointer_weight: - raise ValueError("pointer_weight requires a JevAny checkpoint") + geometric_blend([1.0], [1.0], selected_weight) + self.native_weight = float(selected_weight) + # Public compatibility alias for existing runtime/CLI integrations. + self.pointer_weight = self.native_weight + if checkpoint is None and self.native_weight: + raise ValueError("native_weight requires a JevAny checkpoint") + if type(return_components) is not bool: + raise TypeError("return_components must be a boolean") + if type(exact_kernels) is not bool: + raise TypeError("exact_kernels must be a boolean") + self.return_components = return_components + self.exact_kernels = exact_kernels + self.exact_cuda_kernels_applied = str(device).startswith("cuda") and exact_kernels + if ( + efficient_long_context_tokens is not None + and ( + type(efficient_long_context_tokens) is not int + or efficient_long_context_tokens < 2 + ) + ): + raise ValueError("efficient_long_context_tokens must be an integer >= 2 or None") + self.efficient_long_context_tokens = efficient_long_context_tokens if type(max_tokens) is not int or max_tokens < 2: raise ValueError("max_tokens must be an integer >= 2") self.max_tokens = max_tokens self.device = device self.options = options or LoadOptions() if self.options.temperature is not None: - raise ValueError("native-head temperature is not used by letter readout; use temperature") + raise ValueError("native-head temperature is not used by choice readout; use temperature") if self.options.cuda_graphs: raise ValueError("CUDA graph capture is available only for native readout") + if self.exact_cuda_kernels_applied: + # Match LocalPredictor: benchmark evaluation disables approximate + # TF32 and fused SDPA kernels, while exact_kernels=False retains + # the serving-style PyTorch defaults. + torch.backends.cuda.matmul.allow_tf32 = False + torch.backends.cudnn.allow_tf32 = False + torch.backends.cuda.enable_flash_sdp(False) + torch.backends.cuda.enable_mem_efficient_sdp(False) self.checkpoint: Checkpoint | None = None + self.native_model = None + self.native_decision_mode = None + self._needs_native = checkpoint is not None and bool( + self.native_weight or self.return_components + ) + # Compatibility alias retained for callers predating generalized + # native blending. self.pointer_model = None if checkpoint is not None: @@ -269,11 +323,11 @@ def __init__( raise ValueError("revision accompanies a base; pin checkpoint revisions in owner/repo@revision") loaded = self.checkpoint = Checkpoint(checkpoint) if loaded.meta.special_embeddings: - raise ValueError("letter readout is not validated for checkpoints with trained special embeddings") - if self.pointer_weight and loaded.meta.decision_mode != "pointer": - raise ValueError("pointer_weight requires a native pointer checkpoint") + raise ValueError("choice readout is not validated for checkpoints with trained special embeddings") self.preprocessor, decision_model = loaded.load(device, self.options) + self.native_model = decision_model self.pointer_model = decision_model + self.native_decision_mode = loaded.meta.decision_mode self.language_model = decision_model.lm canonical_base = loaded.meta.base canonical_revision = loaded.meta.base_revision @@ -305,7 +359,10 @@ def __init__( projection_source, projection_revision = source, source_revision checkpoint_id, adapter_scale, adapter_applied = None, 0.0, False - released_output_head = _release_unused_output_head(decision_model) + released_output_head = _release_unused_output_head( + decision_model, + retain_native=(self._needs_native and self.native_decision_mode == "lm_token"), + ) self.language_model.eval() override_config = ( Path(self.options.base_load_path) / "config.json" @@ -339,7 +396,7 @@ def __init__( hashlib.sha256(chat_template.encode()).hexdigest() if isinstance(chat_template, str) else None ) - self.alias_token_ids = letter_token_ids(self.tokenizer, LETTERS) + self.alias_token_ids = letter_token_ids(self.tokenizer, CHOICE_SYMBOLS) flat_token_ids, alias_rows = [], {} for letter, token_ids in self.alias_token_ids.items(): start = len(flat_token_ids) @@ -349,7 +406,7 @@ def __init__( if tied: embeddings = self.language_model.get_input_embeddings().weight if max(flat_token_ids) >= embeddings.shape[0]: - raise ValueError("letter choice token ID exceeds the tied vocabulary projection") + raise ValueError("choice token ID exceeds the tied vocabulary projection") index = torch.tensor(flat_token_ids, dtype=torch.long, device=embeddings.device) projection = embeddings.index_select(0, index).detach().clone() projection_bias = None @@ -368,12 +425,12 @@ def __init__( if not math.isfinite(self.output_softcap) or self.output_softcap <= 0: raise ValueError("final_logit_softcapping must be finite and positive") self.provenance = { - "method": "exact option-letter choice projection", - "constraint_emulation": "exact decoded uppercase letters", - "prompt": "Cygnet-compatible", + "method": "training-free exact choice-token projection", + "constraint_emulation": "exact decoded one-character option IDs", + "prompt": "Cygnet-compatible A-Z; extended with a-z above 26 options", "system_prompt_sha256": hashlib.sha256(SYSTEM_PROMPT.encode()).hexdigest(), "chat_template_sha256": self.chat_template_sha256, - "prompt_format_version": 1, + "prompt_format_version": 2, "canonical_base": canonical_base, "canonical_revision": canonical_revision, "checkpoint": checkpoint_id, @@ -381,12 +438,29 @@ def __init__( "adapter_scale": adapter_scale, "tied_output_embeddings": tied, "unused_full_output_head_released": released_output_head, + "choice_token_rows": len(flat_token_ids), "letter_choice_token_rows": len(flat_token_ids), "output_softcap": self.output_softcap, "temperature": self.temperature, + "native_weight": self.native_weight, + "native_weight_argument": weight_source, + "native_decision_mode": self.native_decision_mode, "pointer_weight": self.pointer_weight, + "native_temperature": ( + float(self.native_model.temperature) if self._needs_native else None + ), "pointer_temperature": ( - float(self.pointer_model.temperature) if self.pointer_weight else None + float(self.native_model.temperature) if self._needs_native else None + ), + "return_components": self.return_components, + "exact_kernels": self.exact_kernels, + "exact_cuda_kernels_applied": self.exact_cuda_kernels_applied, + "efficient_long_context_tokens": self.efficient_long_context_tokens, + "long_context_attention": ( + "fused SDPA (Flash/Efficient, no math fallback) at or above the " + "recorded threshold; exact math SDPA below it" + if self.efficient_long_context_tokens is not None + else None ), "max_tokens": self.max_tokens, "backbone_context_window": self.backbone_context_window, @@ -395,6 +469,26 @@ def __init__( "device": device, } + def _uses_efficient_attention(self, token_count: int) -> bool: + return ( + str(self.device).startswith("cuda") + and getattr(self, "efficient_long_context_tokens", None) is not None + and token_count >= self.efficient_long_context_tokens + ) + + def _attention_context(self, token_count: int): + if not self._uses_efficient_attention(token_count): + return nullcontext() + # Exact math SDPA materializes an O(sequence^2) attention matrix. A + # few public JevJudge text records exceed 30K tokens, which can require + # roughly 60 GiB for that matrix alone. Release cached short-record + # workspaces and force a memory-linear fused backend for those records. + torch.cuda.empty_cache() + return sdpa_kernel( + [SDPBackend.FLASH_ATTENTION, SDPBackend.EFFICIENT_ATTENTION], + set_priority=True, + ) + def _chat_ids(self, prompt: str) -> torch.Tensor: messages = [ {"role": "system", "content": SYSTEM_PROMPT}, @@ -410,22 +504,25 @@ def _chat_ids(self, prompt: str) -> torch.Tensor: ids = _one_row_input_ids(ids) if ids.shape[1] + 1 > self.effective_max_tokens: raise ContextLengthError( - f"letter prompt needs {ids.shape[1] + 1} tokens including the answer; " + f"choice prompt needs {ids.shape[1] + 1} tokens including the answer; " f"limit {self.effective_max_tokens}" ) return ids.to(self.device) def _letter_question(self, state, question: dict) -> tuple[list[str], list[float], list[float], int]: keys, descriptions = question_options(question) - if len(keys) > len(LETTERS): - raise ValueError(f"exact letter readout supports at most {len(LETTERS)} options") + if len(keys) > len(CHOICE_SYMBOLS): + raise ValueError( + f"exact choice-token readout supports at most {len(CHOICE_SYMBOLS)} options" + ) prompt = build_prompt(state, question.get("instructions") or "", descriptions) ids = self._chat_ids(prompt) - output = self.language_model( - input_ids=ids, - attention_mask=torch.ones_like(ids), - use_cache=False, - ) + with self._attention_context(int(ids.shape[1])): + output = self.language_model( + input_ids=ids, + attention_mask=torch.ones_like(ids), + use_cache=False, + ) hidden = output.last_hidden_state[0, -1] projection = self.letter_projection selected_logits = F.linear( @@ -434,15 +531,17 @@ def _letter_question(self, state, question: dict) -> tuple[list[str], list[float if self.output_softcap is not None: selected_logits = self.output_softcap * torch.tanh(selected_logits / self.output_softcap) selected_logits = selected_logits.cpu() - aliases = {letter: self.alias_rows[letter] for letter in LETTERS[:len(keys)]} + aliases = { + symbol: self.alias_rows[symbol] for symbol in CHOICE_SYMBOLS[:len(keys)] + } readout = read_letter_distribution(selected_logits, aliases, self.temperature) probabilities = [readout.calibrated_probabilities[letter] for letter in aliases] calibrated_logits = [readout.raw_log_masses[letter] / self.temperature for letter in aliases] return keys, probabilities, calibrated_logits, int(ids.shape[1]) - def _pointer(self, record: dict) -> tuple[list[list[float]], int]: + def _native(self, record: dict) -> tuple[list[list[float]], int]: internal = materialize(record) - encoded = self.pointer_model.encode( + encoded = self.native_model.encode( self.preprocessor, internal, max_state=self.effective_max_tokens, @@ -451,55 +550,83 @@ def _pointer(self, record: dict) -> tuple[list[list[float]], int]: ) if len(encoded["ids"]) > self.effective_max_tokens: raise ContextLengthError( - f"pointer request needs {len(encoded['ids'])} packed tokens; " + f"native request needs {len(encoded['ids'])} packed tokens; " f"limit {self.effective_max_tokens}" ) - return [F.softmax(logits, -1).float().cpu().tolist() - for logits in self.pointer_model.forward(encoded)], len(encoded["ids"]) + with self._attention_context(len(encoded["ids"])): + rows = [F.softmax(logits, -1).float().cpu().tolist() + for logits in self.native_model.forward(encoded)] + return rows, len(encoded["ids"]) + + def _pointer(self, record: dict) -> tuple[list[list[float]], int]: + """Compatibility alias for the generalized native checkpoint pass.""" + + return self._native(record) @torch.inference_mode() def __call__(self, record: dict) -> dict: request = api_request(record) SystemOneRequest.model_validate(request) if request.get("media"): - raise ValueError("letter readout is text-only and does not accept media") + raise ValueError("choice readout is text-only and does not accept media") sync(self.device) started = time.perf_counter() letter_rows = [ self._letter_question(request["state"], question) for question in request["questions"].values() ] - pointer_rows, pointer_tokens = (None, 0) - if self.pointer_weight: - pointer_rows, pointer_tokens = self._pointer(record) - if len(pointer_rows) != len(letter_rows): - raise ValueError("pointer and letter question counts differ") + native_rows, native_tokens = (None, 0) + if self._needs_native: + native_rows, native_tokens = self._native(record) + if len(native_rows) != len(letter_rows): + raise ValueError("native and choice question counts differ") probabilities, logits = {}, {} + choice_probabilities, native_probabilities = {}, {} for index, (question_id, (keys, letter_p, letter_z, _tokens)) in enumerate( zip(request["questions"], letter_rows, strict=True) ): - if pointer_rows is None: + choice_probabilities[question_id] = dict(zip(keys, letter_p, strict=True)) + if native_rows is not None: + if len(native_rows[index]) != len(keys): + raise ValueError( + f"native and choice option counts differ for question {question_id!r}" + ) + native_probabilities[question_id] = dict( + zip(keys, native_rows[index], strict=True) + ) + if native_rows is None or not self.native_weight: selected_p, selected_z = letter_p, letter_z else: selected_p, selected_z = geometric_blend( - letter_p, pointer_rows[index], self.pointer_weight + letter_p, native_rows[index], self.native_weight ) probabilities[question_id] = dict(zip(keys, selected_p, strict=True)) logits[question_id] = dict(zip(keys, selected_z, strict=True)) sync(self.device) - return { + result = { "probabilities": probabilities, "logits": logits, "inference_temperature": self.temperature, "latency_ms": (time.perf_counter() - started) * 1000, - "input_tokens": sum(row[3] for row in letter_rows) + pointer_tokens, + "input_tokens": sum(row[3] for row in letter_rows) + native_tokens, + "efficient_long_context_attention_used": ( + any(self._uses_efficient_attention(row[3]) for row in letter_rows) + or self._uses_efficient_attention(native_tokens) + ), "readout": { "method": self.provenance["method"], "adapter_applied": self.provenance["adapter_applied"], + "native_weight": self.native_weight, + "native_decision_mode": self.native_decision_mode, "pointer_weight": self.pointer_weight, }, } + if self.return_components: + result["component_probabilities"] = {"choice": choice_probabilities} + if native_rows is not None: + result["component_probabilities"]["native"] = native_probabilities + return result __all__ = ["LetterReadoutPredictor", "geometric_blend", "question_options"] diff --git a/jevany/letter_readout.py b/jevany/letter_readout.py index 0b857f7..f38ac57 100644 --- a/jevany/letter_readout.py +++ b/jevany/letter_readout.py @@ -1,4 +1,4 @@ -"""One-token option-letter readout primitives for frozen causal LMs. +"""One-token option-ID readout primitives for frozen causal LMs. The Cygnet-compatible prompt and letter aggregation semantics are adapted from the MIT-licensed ``blockbrain-ai/cygnet-recipe`` shim at commit ``3cf591c`` @@ -18,11 +18,18 @@ from typing import Any -LETTERS = "ABCDEFGHIJKLMNOPQRSTUVWXYZ" +CYGNET_LETTERS = "ABCDEFGHIJKLMNOPQRSTUVWXYZ" +# Preserve Cygnet's exact A-Z assignment for its supported range, then extend +# with distinct lowercase one-character IDs for larger decision sets. This +# covers JevJudge's 28-30-option text requests without dropping records. +CHOICE_SYMBOLS = CYGNET_LETTERS + "abcdefghijklmnopqrstuvwxyz" +# Backward-compatible implementation name. Public documentation calls these +# option IDs and the method training-free choice readout. +LETTERS = CHOICE_SYMBOLS SYSTEM_PROMPT = ( "You are a calibration engine. You never answer in prose. You are given a state, a question and " "a numbered set of options, and you choose exactly one option. You reply with that option's " - "LETTER and nothing else โ€” a single character, no words, no punctuation, no explanation." + "case-sensitive ID and nothing else โ€” a single character, no words, no punctuation, no explanation." ) _NEGATIVE_INFINITY = float("-inf") @@ -63,7 +70,7 @@ def _ordered_descriptions(options: Mapping[Any, Any] | Sequence[Any]) -> list[An def build_prompt(state: Any, instructions: Any, options: Mapping[Any, Any] | Sequence[Any]) -> str: - """Build Cygnet's measured user prompt, assigning options A through Z. + """Build the measured user prompt, assigning one-character option IDs. Mapping insertion order or sequence order is the option order. Structured state is rendered with ``indent=1`` as in Cygnet. This function only @@ -78,22 +85,22 @@ def build_prompt(state: Any, instructions: Any, options: Mapping[Any, Any] | Seq ) lines = [state_text.rstrip(), "", instruction_text.rstrip(), "", "Options:"] lines.extend(f"{LETTERS[index]}. {description}" for index, description in enumerate(descriptions)) - lines.extend(("", "Answer with the letter of exactly one option, and nothing else:")) + lines.extend(("", "Answer with the case-sensitive ID of exactly one option, and nothing else:")) return "\n".join(lines) def _validate_letters(letters: Sequence[str]) -> tuple[str, ...]: if isinstance(letters, (bytes, bytearray)): - raise TypeError("letters must be a sequence of A-Z strings") + raise TypeError("option IDs must be a sequence of one-character strings") selected = tuple(letters) if not selected: - raise ValueError("letters must not be empty") + raise ValueError("option IDs must not be empty") if len(selected) > len(LETTERS): raise ValueError(f"letters may contain at most {len(LETTERS)} entries") - if any(type(letter) is not str or len(letter) != 1 or letter not in LETTERS for letter in selected): - raise ValueError("letters must contain only single uppercase A-Z strings") + if any(type(letter) is not str or len(letter) != 1 or letter not in CHOICE_SYMBOLS for letter in selected): + raise ValueError("option IDs must be distinct single A-Z/a-z characters") if len(set(selected)) != len(selected): - raise ValueError("letters must not contain duplicates") + raise ValueError("option IDs must not contain duplicates") return selected @@ -113,12 +120,12 @@ def _decode_token(tokenizer: Any, token_id: int) -> str: def letter_token_ids(tokenizer: Any, letters: Sequence[str]) -> dict[str, tuple[int, ...]]: - """Scan the vocabulary for every token ID decoding to each exact letter. + """Scan the vocabulary for every token ID decoding to each exact option ID. Token text is intentionally never used as a dictionary key: distinct token IDs may decode to identical text, and every such ID contributes probability - mass. Tokens decoding to ``" A"``, ``"A."``, or ``"a"`` are excluded: - Cygnet's structured-output choice admits the exact uppercase string only. + mass. Tokens decoding to ``" A"`` or ``"A."`` are excluded. Case is + significant: ``A`` and ``a`` are different option IDs when both are used. ``len(tokenizer)`` must describe the full vocabulary, including added tokens. """ @@ -338,6 +345,8 @@ def read_letter_distribution( __all__ = [ + "CHOICE_SYMBOLS", + "CYGNET_LETTERS", "LETTERS", "SYSTEM_PROMPT", "LetterReadout", diff --git a/jevany/letter_runtime.py b/jevany/letter_runtime.py index 2f41022..1ea0637 100644 --- a/jevany/letter_runtime.py +++ b/jevany/letter_runtime.py @@ -1,4 +1,4 @@ -"""System One runtime adapter for the training-free option-letter predictor.""" +"""System One runtime adapter for the training-free choice-token predictor.""" from __future__ import annotations @@ -7,6 +7,7 @@ from typing import Any from .api import SystemOneRequest, output_tokens, to_answers, to_record, validate_response +from .letter_readout import CHOICE_SYMBOLS def _unlabelled_record(request: SystemOneRequest) -> dict[str, Any]: @@ -41,7 +42,7 @@ def __post_init__(self) -> None: if not isinstance(self.model_id, str) or not self.model_id.strip(): raise ValueError("model_name must be a nonempty string") if self.predictor.checkpoint is None: - raise ValueError("serving letter readout requires a JevAny checkpoint") + raise ValueError("serving choice readout requires a JevAny checkpoint") @property def checkpoint(self): @@ -52,25 +53,26 @@ def aliases(self) -> list[str]: return [] if self.model_id == "jevany-latest" else ["jevany-latest"] def clear_cache(self) -> None: - """Letter readout currently keeps no mutable prefix cache.""" + """Choice readout currently keeps no mutable prefix cache.""" def describe(self) -> dict[str, Any]: - """Describe the effective letter readout and its checkpoint overlay.""" + """Describe the effective choice readout and its checkpoint overlay.""" checkpoint = self.checkpoint - pointer_model = self.predictor.pointer_model - capabilities = pointer_model.inference_capabilities + native_model = self.predictor.native_model + capabilities = native_model.inference_capabilities context_window = capabilities.context_window effective_window = self.predictor.effective_max_tokens - acceleration = getattr(pointer_model, "inference_acceleration", { + acceleration = getattr(native_model, "inference_acceleration", { "compile_mode": None, "lora_merged": False, "approximate_bf16_merge": False, "cuda_graphs": None, }) - letter_readout = dict(self.predictor.provenance) - if self.predictor.pointer_weight: - letter_readout["pointer_temperature"] = pointer_model.temperature + choice_readout = dict(self.predictor.provenance) + if self.predictor.native_weight: + choice_readout["native_temperature"] = native_model.temperature + choice_readout["pointer_temperature"] = native_model.temperature with self.lock: return { "id": self.model_id, @@ -79,13 +81,15 @@ def describe(self) -> dict[str, Any]: "base": checkpoint.meta.base, "lora": checkpoint.meta.lora, "device": self.predictor.device, - "device_map": getattr(pointer_model, "device_map", None), - "devices": getattr(pointer_model, "devices", [self.predictor.device]), + "device_map": getattr(native_model, "device_map", None), + "devices": getattr(native_model, "devices", [self.predictor.device]), "temperature": self.predictor.temperature, "decision_mode": checkpoint.meta.decision_mode, - "readout": "letter", - "letter_readout": letter_readout, - "backbone_adapter": pointer_model.backbone_adapter, + "readout": "choice", + "choice_readout": choice_readout, + # Legacy descriptor alias retained for older clients. + "letter_readout": dict(choice_readout), + "backbone_adapter": native_model.backbone_adapter, "branch_mode": "chat", "acceleration": acceleration, "capabilities": { @@ -98,7 +102,7 @@ def describe(self) -> dict[str, Any]: "state_tokens": effective_window, "branch_tokens": effective_window, "packed_tokens": effective_window, - "choices": 26, + "choices": len(CHOICE_SYMBOLS), }, "prefix_cache": { "enabled": False, @@ -111,12 +115,12 @@ def describe(self) -> dict[str, Any]: } def answer(self, request: SystemOneRequest) -> dict[str, Any]: - """Run letter inference and return a validated System One response.""" + """Run choice-token inference and return a validated System One response.""" if request.model not in (self.model_id, *self.aliases): raise ValueError(f"unknown model {request.model!r}; this deployment serves {self.model_id!r}") if request.media: - raise ValueError("letter readout does not support media requests") + raise ValueError("choice readout does not support media requests") _, metadata = to_record(request) with self.lock: prediction = self.predictor(_unlabelled_record(request)) diff --git a/jevany/readout.py b/jevany/readout.py index 99f8b45..387a1a1 100644 --- a/jevany/readout.py +++ b/jevany/readout.py @@ -11,69 +11,159 @@ import math -READOUTS = ("native", "letter") +# ``letter`` is accepted indefinitely as the legacy spelling, but is omitted +# from generated usage/help so new integrations consistently say ``choice``. +READOUTS = ("native", "choice") +_ACCEPTED_READOUTS = (*READOUTS, "letter") + + +def normalize_readout(readout: str) -> str: + """Return the public readout name while accepting the legacy alias.""" + + if readout == "letter": + return "choice" + if readout not in READOUTS: + raise ValueError("readout must be native or choice") + return readout @dataclass(frozen=True) -class LetterReadoutOptions: - """Runtime settings for the training-free option-letter readout.""" +class ChoiceReadoutOptions: + """Runtime settings for the training-free choice-token readout.""" temperature: float = 1.0 - pointer_weight: float = 0.0 + native_weight: float = 0.0 max_tokens: int = 16_384 def __post_init__(self) -> None: - for name in ("temperature", "pointer_weight"): + for name in ("temperature", "native_weight"): value = getattr(self, name) if isinstance(value, bool) or not isinstance(value, (int, float)) or not math.isfinite(value): - raise ValueError(f"letter {name.replace('_', ' ')} must be finite and numeric") + raise ValueError(f"choice {name.replace('_', ' ')} must be finite and numeric") if self.temperature <= 0: - raise ValueError("letter temperature must be positive") - if not 0 <= self.pointer_weight <= 1: - raise ValueError("letter pointer weight must be in [0, 1]") + raise ValueError("choice temperature must be positive") + if not 0 <= self.native_weight <= 1: + raise ValueError("choice native weight must be in [0, 1]") if type(self.max_tokens) is not int or self.max_tokens < 2: - raise ValueError("letter max tokens must be an integer >= 2") + raise ValueError("choice max tokens must be an integer >= 2") + + @property + def pointer_weight(self) -> float: + """Legacy name for :attr:`native_weight`.""" + + return self.native_weight + + +@dataclass(frozen=True) +class LetterReadoutOptions: + """Legacy spelling of :class:`ChoiceReadoutOptions`.""" + + temperature: float = 1.0 + pointer_weight: float = 0.0 + max_tokens: int = 16_384 + + def __post_init__(self) -> None: + # Validate through the public representation so both APIs have exactly + # the same accepted values. + ChoiceReadoutOptions(self.temperature, self.pointer_weight, self.max_tokens) + + @property + def native_weight(self) -> float: + return self.pointer_weight + + def as_choice(self) -> ChoiceReadoutOptions: + return ChoiceReadoutOptions(self.temperature, self.pointer_weight, self.max_tokens) def add_readout_arguments(parser: argparse.ArgumentParser) -> None: - """Add the same readout selector and letter-only knobs to a CLI parser.""" + """Add shared readout arguments, with legacy flags accepted but hidden.""" parser.add_argument( - "--readout", choices=READOUTS, default="native", - help="native checkpoint head (default) or training-free option-letter readout", + "--readout", choices=_ACCEPTED_READOUTS, metavar="{native,choice}", default="native", + help="native checkpoint head (default) or training-free choice-token readout", ) parser.add_argument( - "--letter-temperature", type=float, - help="temperature for letter probability calibration (letter readout only; default 1.0)", + "--choice-temperature", type=float, + help="probability temperature for choice readout (default 1.0)", ) parser.add_argument( - "--letter-pointer-weight", type=float, - help="log-linear weight on the native pointer distribution (letter readout only; default 0)", + "--choice-native-weight", type=float, + help="log-linear weight on the checkpoint's native distribution (default 0)", ) parser.add_argument( - "--letter-max-tokens", type=int, - help="maximum chat-prompt tokens including the one-letter answer (letter readout only; default 16384)", + "--choice-max-tokens", type=int, + help="maximum chat-prompt tokens including the one-token answer (default 16384)", + ) + # Backward-compatible spellings. argparse.SUPPRESS keeps the preferred + # surface compact without removing existing scripts. + parser.add_argument("--letter-temperature", type=float, help=argparse.SUPPRESS) + parser.add_argument("--letter-pointer-weight", type=float, help=argparse.SUPPRESS) + parser.add_argument("--letter-max-tokens", type=int, help=argparse.SUPPRESS) + + +def resolve_readout_options( + readout: str = "native", *, + choice_temperature: float | None = None, + choice_native_weight: float | None = None, + choice_max_tokens: int | None = None, + letter_temperature: float | None = None, + letter_pointer_weight: float | None = None, + letter_max_tokens: int | None = None, +) -> tuple[str, ChoiceReadoutOptions | None]: + """Normalize aliases and resolve preferred/legacy setting spellings.""" + + normalized = normalize_readout(readout) + pairs = ( + ("temperature", "--choice-temperature", choice_temperature, + "--letter-temperature", letter_temperature), + ("native_weight", "--choice-native-weight", choice_native_weight, + "--letter-pointer-weight", letter_pointer_weight), + ("max_tokens", "--choice-max-tokens", choice_max_tokens, + "--letter-max-tokens", letter_max_tokens), + ) + values, used = {}, [] + for name, preferred_flag, preferred, legacy_flag, legacy in pairs: + if preferred is not None and legacy is not None: + raise ValueError(f"give at most one of {preferred_flag} and legacy {legacy_flag}") + if preferred is not None: + values[name] = preferred + used.append(preferred_flag) + elif legacy is not None: + values[name] = legacy + used.append(legacy_flag) + if normalized == "native": + if used: + raise ValueError(f"{', '.join(used)} require --readout choice") + return normalized, None + return normalized, ChoiceReadoutOptions(**values) + + +def choice_options_from_args(args: argparse.Namespace) -> ChoiceReadoutOptions | None: + """Return preferred choice settings from a shared CLI namespace.""" + + _, options = resolve_readout_options( + getattr(args, "readout", "native"), + choice_temperature=getattr(args, "choice_temperature", None), + choice_native_weight=getattr(args, "choice_native_weight", None), + choice_max_tokens=getattr(args, "choice_max_tokens", None), + letter_temperature=getattr(args, "letter_temperature", None), + letter_pointer_weight=getattr(args, "letter_pointer_weight", None), + letter_max_tokens=getattr(args, "letter_max_tokens", None), ) + return options def letter_options_from_args(args: argparse.Namespace) -> LetterReadoutOptions | None: - """Return letter settings, rejecting letter-only flags with native readout.""" - - values = { - "temperature": getattr(args, "letter_temperature", None), - "pointer_weight": getattr(args, "letter_pointer_weight", None), - "max_tokens": getattr(args, "letter_max_tokens", None), - } - if getattr(args, "readout", "native") == "native": - used = ["--letter-" + name.replace("_", "-") for name, value in values.items() if value is not None] - if used: - raise ValueError(f"{', '.join(used)} require --readout letter") - return None - return LetterReadoutOptions(**{ - name: value for name, value in values.items() if value is not None - }) + """Legacy wrapper returning the historical options type.""" + + options = choice_options_from_args(args) + return (None if options is None else LetterReadoutOptions( + options.temperature, options.native_weight, options.max_tokens, + )) __all__ = [ - "READOUTS", "LetterReadoutOptions", "add_readout_arguments", "letter_options_from_args", + "READOUTS", "ChoiceReadoutOptions", "LetterReadoutOptions", + "add_readout_arguments", "choice_options_from_args", "letter_options_from_args", + "normalize_readout", "resolve_readout_options", ] diff --git a/jevany/runtime.py b/jevany/runtime.py index 728f43b..1d1205e 100644 --- a/jevany/runtime.py +++ b/jevany/runtime.py @@ -14,6 +14,7 @@ from .device import default_device, sync from .inference import InferenceOptions from .model import DecisionModel +from .readout import resolve_readout_options DEFAULT_CHECKPOINT = "SimpleJev/JevAny-Qwen3.8-27B-LoRA" @@ -165,7 +166,9 @@ def from_pretrained( device: str | None = None, dtype: str | None = None, model_name: str | None = None, options: LoadOptions | None = None, inference_options: InferenceOptions | None = None, - readout: str = "native", letter_temperature: float | None = None, + readout: str = "native", choice_temperature: float | None = None, + choice_native_weight: float | None = None, choice_max_tokens: int | None = None, + letter_temperature: float | None = None, letter_pointer_weight: float | None = None, letter_max_tokens: int | None = None, ) -> "JevModel": """Load a local run or Hugging Face adapter ID (optionally ``owner/repo@revision``). @@ -173,20 +176,25 @@ def from_pretrained( The full backbone must fit on the selected device unless ``options.device_map`` (or JEVANY_DEVICE_MAP) splits it over the visible GPUs. ``dtype`` accepts fp32, fp16 or bf16; omission uses the checkpoint/environment settings. - ``readout='letter'`` replaces the checkpoint head with a training-free - option-letter projection; its temperature, pointer blend, and prompt - limit are deployment settings, not checkpoint metadata. Files used by - native media requests are trusted local paths. + ``readout='choice'`` replaces the checkpoint head with a training-free + choice-token projection; its temperature, native blend, and prompt + limit are deployment settings, not checkpoint metadata. ``letter`` and + ``letter_*`` remain accepted legacy aliases. Files used by native media + requests are trusted local paths. """ import torch - if readout not in ("native", "letter"): - raise ValueError("readout must be native or letter") - letter_settings = (letter_temperature, letter_pointer_weight, letter_max_tokens) - if readout == "native" and any(value is not None for value in letter_settings): - raise ValueError("letter readout options require readout='letter'") - if readout == "letter" and inference_options is not None: - raise ValueError("inference_options apply only to native readout; use letter_max_tokens") + readout, choice_options = resolve_readout_options( + readout, + choice_temperature=choice_temperature, + choice_native_weight=choice_native_weight, + choice_max_tokens=choice_max_tokens, + letter_temperature=letter_temperature, + letter_pointer_weight=letter_pointer_weight, + letter_max_tokens=letter_max_tokens, + ) + if readout == "choice" and inference_options is not None: + raise ValueError("inference_options apply only to native readout; use choice_max_tokens") device = default_device() if device is None else device if device not in ("cpu", "mps", "cuda"): @@ -203,21 +211,21 @@ def from_pretrained( options = replace(options, dtype=dtypes[dtype]) if device == "mps" and options.attn is None: options = replace(options, attn="sdpa") - if readout == "letter" and options.cuda_graphs: + if readout == "choice" and options.cuda_graphs: raise ValueError("CUDA graph capture is available only for native readout") - if readout == "letter" and options.temperature is not None: - raise ValueError("JEVANY_TEMPERATURE applies to the native head; use letter_temperature") + if readout == "choice" and options.temperature is not None: + raise ValueError("JEVANY_TEMPERATURE applies to the native head; use choice_temperature") predictor = None - if readout == "letter": + if readout == "choice": from .letter_predictor import LetterReadoutPredictor predictor = LetterReadoutPredictor( checkpoint=checkpoint, device=device, options=options, - temperature=1.0 if letter_temperature is None else letter_temperature, - pointer_weight=0.0 if letter_pointer_weight is None else letter_pointer_weight, - max_tokens=16_384 if letter_max_tokens is None else letter_max_tokens, + temperature=choice_options.temperature, + native_weight=choice_options.native_weight, + max_tokens=choice_options.max_tokens, ) loaded = predictor.checkpoint else: diff --git a/jevany/serve.py b/jevany/serve.py index a93d798..49944ff 100644 --- a/jevany/serve.py +++ b/jevany/serve.py @@ -3,7 +3,7 @@ """FastAPI server for prefill-only decisions. Run: uv run --extra serve python -m jevany.serve --run runs/rlcr --port 8008 -Use ``--readout letter`` for the training-free option-letter deployment path. +Use ``--readout choice`` for the training-free choice-token deployment path. TypeSafe-compatible: POST /v1/systemone and GET /v1/models (no auth). JEVANY_PREFIX_CACHE / JEVANY_PREFIX_MIN_TOKENS size the state-prefix cache; JEVANY_DATE_FACTS=1 enables deterministic date preprocessing. @@ -18,7 +18,13 @@ from .api import SystemOneRequest, with_date_facts from .checkpoint import LoadOptions, add_placement_arguments, load_options_from_args from .inference import InferenceOptions, add_inference_arguments, inference_options_from_args -from .readout import LetterReadoutOptions, add_readout_arguments, letter_options_from_args +from .readout import ( + ChoiceReadoutOptions, + LetterReadoutOptions, + add_readout_arguments, + choice_options_from_args, + normalize_readout, +) from .runtime import DEFAULT_CHECKPOINT, DecisionRuntime, JevModel DATE_FACTS = os.environ.get("JEVANY_DATE_FACTS", "0") == "1" MEDIA_ROOT = os.environ.get("JEVANY_MEDIA_ROOT") @@ -148,7 +154,8 @@ def create_app( model: JevModel | None = None, device: str | None = None, dtype: str | None = None, model_name: str | None = None, options: LoadOptions | None = None, inference_options: InferenceOptions | None = None, - readout: str = "native", letter_options: LetterReadoutOptions | None = None, + readout: str = "native", choice_options: ChoiceReadoutOptions | None = None, + letter_options: LetterReadoutOptions | None = None, ) -> FastAPI: """Build an isolated app, loading one checkpoint during ASGI startup. @@ -156,18 +163,21 @@ def create_app( and cache with Python callers. Loading options cannot accompany an injected model. Each worker loads its own full model; use one worker per device. """ - if model is not None and (readout != "native" or letter_options is not None or any( + readout = normalize_readout(readout) + if choice_options is not None and letter_options is not None: + raise ValueError("give at most one of choice_options and legacy letter_options") + if letter_options is not None: + choice_options = letter_options.as_choice() + if model is not None and (readout != "native" or choice_options is not None or any( value is not None for value in (checkpoint, device, dtype, model_name, options, inference_options) )): raise ValueError("pass either a loaded model or checkpoint loading options") - if readout not in ("native", "letter"): - raise ValueError("readout must be native or letter") - if readout == "native" and letter_options is not None: - raise ValueError("letter_options require readout='letter'") - if readout == "letter" and inference_options is not None: - raise ValueError("inference_options apply only to native readout; use letter_options.max_tokens") - if readout == "letter" and letter_options is None: - letter_options = LetterReadoutOptions() + if readout == "native" and choice_options is not None: + raise ValueError("choice_options require readout='choice'") + if readout == "choice" and inference_options is not None: + raise ValueError("inference_options apply only to native readout; use choice_options.max_tokens") + if readout == "choice" and choice_options is None: + choice_options = ChoiceReadoutOptions() @asynccontextmanager async def lifespan(application: FastAPI): @@ -175,11 +185,11 @@ async def lifespan(application: FastAPI): local = model else: settings = ({ - "readout": "letter", - "letter_temperature": letter_options.temperature, - "letter_pointer_weight": letter_options.pointer_weight, - "letter_max_tokens": letter_options.max_tokens, - } if letter_options is not None else { + "readout": "choice", + "choice_temperature": choice_options.temperature, + "choice_native_weight": choice_options.native_weight, + "choice_max_tokens": choice_options.max_tokens, + } if choice_options is not None else { "inference_options": inference_options, }) local = JevModel.from_pretrained( @@ -221,15 +231,15 @@ def main(argv=None): ap.add_argument("--port", type=int, default=8008) a = ap.parse_args(argv) try: - letter_options = letter_options_from_args(a) + choice_options = choice_options_from_args(a) except ValueError as error: ap.error(str(error)) - if letter_options is not None and ( + if choice_options is not None and ( a.cuda_graphs or a.cuda_graph_max_tokens is not None or any(getattr(a, item.name) is not None for item in fields(InferenceOptions)) ): - ap.error("CUDA graph and native inference-limit flags cannot be used with --readout letter; " - "use --letter-max-tokens") + ap.error("CUDA graph and native inference-limit flags cannot be used with --readout choice; " + "use --choice-max-tokens") options = load_options_from_args(a) if a.cuda_graphs or a.cuda_graph_max_tokens is not None: options = options or LoadOptions.from_env() @@ -239,8 +249,8 @@ def main(argv=None): application = create_app(a.run, device=a.device, dtype=a.dtype, model_name=a.model_name, options=options, inference_options=(inference_options_from_args(a) - if letter_options is None else None), - readout=a.readout, letter_options=letter_options) + if choice_options is None else None), + readout=normalize_readout(a.readout), choice_options=choice_options) import uvicorn uvicorn.run(application, host=a.host, port=a.port) diff --git a/reports/JevAny_Tech_Report.pdf b/reports/JevAny_Tech_Report.pdf index de0ad58..cedc9fb 100644 Binary files a/reports/JevAny_Tech_Report.pdf and b/reports/JevAny_Tech_Report.pdf differ diff --git a/results/choice-readout-v2.json b/results/choice-readout-v2.json new file mode 100644 index 0000000..05d9a87 --- /dev/null +++ b/results/choice-readout-v2.json @@ -0,0 +1,2672 @@ +{ + "schema_version": 2, + "artifact": "choice-readout-v2", + "title": "Training-free choice-token readout evaluation", + "generated_by": "scripts/build_choice_readout_results.py", + "method": { + "readout": "Constrain the next token to exact one-character option IDs and renormalize their probability mass; this can change argmax decisions.", + "training": "No additional training for the choice-token readout.", + "option_ids": "A-Z followed by a-z; at most 52 options.", + "calibration": "Scalar temperature changes probabilities but not argmax accuracy.", + "ensemble": "Native + choice configurations use log-linear/geometric pooling." + }, + "protocol_groups": { + "zero_shot": { + "meaning": "No Typed, JevJudge, JevBench, or Transfer-test labels tune the readout configuration.", + "configurations": [ + "choice_t1", + "native_shipped", + "fixed_blend_0_5" + ] + }, + "transfer_dev_tuned": { + "meaning": "Weight and additional temperature are selected only on Transfer-v9 development, then frozen for every displayed evaluation panel.", + "configurations": [ + "transfer_dev_calibrated_choice", + "transfer_dev_tuned_blend" + ], + "accuracy_note": "Temperature-only calibrated choice has the same hard accuracy as choice_t1; only a changed blend weight can change its argmax." + } + }, + "expected_counts": { + "transfer_calibration": { + "label": "Transfer-v9 development", + "role": "selection_only", + "records": 1264, + "questions": 1264, + "headline_n": 1046 + }, + "transfer_test": { + "label": "Transfer-v9 test", + "role": "held_out_evaluation", + "records": 1264, + "questions": 1264, + "headline_n": 1046 + }, + "typed_test": { + "label": "Typed Decisions test", + "role": "held_out_external_evaluation", + "records": 400, + "questions": 2000, + "headline_n": 2000 + }, + "jevbench_public": { + "label": "JevBench public development", + "role": "public_diagnostic_not_used_for_selection", + "records": 231, + "questions": 231, + "headline_n": 231 + }, + "jevjudge_text": { + "label": "JevJudge text", + "role": "held_out_external_evaluation", + "records": 724, + "questions": 724, + "headline_n": 724 + } + }, + "dataset_hashes": { + "jevbench_development_sha256": "b6d6eb34fdbbf46ec11be8657da5e513e3538990d7a00ca5582c55d8bd7941c6", + "jevbench_manifest_sha256": "e356ad5a2c1dd3917c126e74bdee58940391568de013502abcdcd902bde3c2a5", + "transfer_manifest_sha256": "3c4f0be94509a3612678bfd3a30fd99a8d0ca3c47ddfe7318075d95b2fa365e4", + "transfer_development_sha256": "53872ba045564315cf3c3dd8fb656e76441c6c514b1a7ff985ff0db3955710cf", + "transfer_test_sha256": "048abea5e12c73c3a622ce5836898c00fa43766e0bf48fb2459e77ee9873497f", + "typed_revision": "d0e2f0c42fef86cc15d1688d25a19f5ba7c85b18", + "typed_test_sha256": "4f294f218ea1da27f3efef936359389c62ea4d3973a41457732990f1d31b647c", + "jevjudge_test_sha256": "5c50f89e9ff1beb2107648b56f2a3cbc19c946e379519caf59afc191700e83cb", + "jevjudge_manifest_sha256": "ee52a15f82629742093a89910a8f72e7ca97248b197aad6ec801f984c2a15b68" + }, + "code_revision": "ae891ed97147b57b41a78a4bb4960b018db3bf42", + "source_runs": [ + { + "key": "base4", + "model": "Frozen-Qwen3.5-4B", + "family": "Qwen3.5-4B", + "weights": "frozen_base", + "native_readout": null, + "artifacts": { + "manifest": { + "path": "/lustre-storage/fsx/tianxinwei/JevAny/runs/choice-readout-v2-20261002-r1/base4/manifest.json", + "sha256": "49e57b70600869d5235a3363a3238161b753b3082f9d348cbd25e810a562418b" + }, + "ensemble": { + "path": "/lustre-storage/fsx/tianxinwei/JevAny/runs/choice-readout-v2-20261002-r1/base4/ensemble.json", + "sha256": "388cf7b00de00c5c5090f27c8436981cadb1a71986f6855f77806c58840b03d2" + } + }, + "source": { + "base": "Qwen/Qwen3.5-4B", + "checkpoint": null + }, + "public_source": { + "repository": "Qwen/Qwen3.5-4B", + "url": "https://huggingface.co/Qwen/Qwen3.5-4B", + "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", + "revision_status": "verified", + "repository_evidence": "manifest.base_loading.canonical_base", + "revision_evidence": "manifest.base_loading.canonical_revision", + "evaluated_artifact_sha256": {}, + "base_model": { + "repository": "Qwen/Qwen3.5-4B", + "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a" + }, + "limitation": null + }, + "checkpoint_artifacts": null, + "base_loading": { + "canonical_base": "Qwen/Qwen3.5-4B", + "canonical_base_is_local": false, + "canonical_revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", + "override_used": true, + "override_config_sha256": "ddc63e1c717afa86c865bb5e01313d89d72bb53b97ad4a8a03ba8510c0621670" + }, + "predictor": { + "method": "training-free exact choice-token projection", + "constraint_emulation": "exact decoded one-character option IDs", + "prompt": "Cygnet-compatible A-Z; extended with a-z above 26 options", + "system_prompt_sha256": "e2f0ac9f62eb446d7826b3abc3f7909287f4075175e9b58e4cd3ca72718671fe", + "chat_template_sha256": "a4aee8afcf2e0711942cf848899be66016f8d14a889ff9ede07bca099c28f715", + "prompt_format_version": 2, + "canonical_base": "Qwen/Qwen3.5-4B", + "canonical_revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", + "checkpoint": null, + "adapter_applied": false, + "adapter_scale": 0.0, + "tied_output_embeddings": true, + "unused_full_output_head_released": false, + "letter_choice_token_rows": 52, + "output_softcap": null, + "temperature": 1.0, + "native_weight": 0.0, + "native_weight_argument": "native_weight", + "native_decision_mode": null, + "pointer_weight": 0.0, + "native_temperature": null, + "pointer_temperature": null, + "return_components": false, + "exact_kernels": true, + "exact_cuda_kernels_applied": true, + "efficient_long_context_tokens": 8192, + "long_context_attention": "fused SDPA (Flash/Efficient, no math fallback) at or above the recorded threshold; exact math SDPA below it", + "max_tokens": 65536, + "backbone_context_window": 262144, + "effective_max_tokens": 65536, + "dtype": "torch.bfloat16", + "device": "cuda" + }, + "runtime": { + "python": "3.12.14", + "torch": "2.8.0+cu128", + "cuda": "12.8", + "transformers": "5.17.0", + "peft": "0.21.0", + "numpy": "2.5.3", + "pyarrow": "25.0.1", + "safetensors": "0.8.0", + "code_revision": "ae891ed97147b57b41a78a4bb4960b018db3bf42" + }, + "dataset_hashes": { + "jevbench_development_sha256": "b6d6eb34fdbbf46ec11be8657da5e513e3538990d7a00ca5582c55d8bd7941c6", + "jevbench_manifest_sha256": "e356ad5a2c1dd3917c126e74bdee58940391568de013502abcdcd902bde3c2a5", + "transfer_manifest_sha256": "3c4f0be94509a3612678bfd3a30fd99a8d0ca3c47ddfe7318075d95b2fa365e4", + "transfer_development_sha256": "53872ba045564315cf3c3dd8fb656e76441c6c514b1a7ff985ff0db3955710cf", + "transfer_test_sha256": "048abea5e12c73c3a622ce5836898c00fa43766e0bf48fb2459e77ee9873497f", + "typed_revision": "d0e2f0c42fef86cc15d1688d25a19f5ba7c85b18", + "typed_test_sha256": "4f294f218ea1da27f3efef936359389c62ea4d3973a41457732990f1d31b647c", + "jevjudge_test_sha256": "5c50f89e9ff1beb2107648b56f2a3cbc19c946e379519caf59afc191700e83cb", + "jevjudge_manifest_sha256": "ee52a15f82629742093a89910a8f72e7ca97248b197aad6ec801f984c2a15b68" + }, + "evaluation_protocol": { + "kernel_policy": "LocalPredictor-compatible exact CUDA policy", + "dtype": "bf16", + "attn": "sdpa", + "temperature": 1.0, + "blend_selection": "not performed here; components are saved for fitting on Transfer-v9 development only", + "efficiency_scope": "base calls execute choice only; checkpoint calls execute choice and native. Do not compare component efficiency from these end-to-end timings", + "long_context_attention": "exact math SDPA below the configured threshold; memory-linear fused SDPA without a math fallback at or above it", + "efficient_long_context_tokens": 8192, + "resume": { + "enabled": true, + "reused_datasets": [ + "jevbench_public", + "transfer_calibration", + "transfer_test", + "typed_test" + ], + "preexisting_code_revision": "256001ca806f9d7e302bea967850db214f3d8e4e", + "dataset_code_revisions": { + "jevbench_public": "256001ca806f9d7e302bea967850db214f3d8e4e", + "transfer_calibration": "256001ca806f9d7e302bea967850db214f3d8e4e", + "transfer_test": "256001ca806f9d7e302bea967850db214f3d8e4e", + "typed_test": "256001ca806f9d7e302bea967850db214f3d8e4e", + "jevjudge_text": "ae891ed97147b57b41a78a4bb4960b018db3bf42" + } + } + }, + "ensemble_protocol": { + "selection_split": "Transfer-v9 development clean/knowable accuracy cohort; never any reported test panel", + "selection_rows": 1046, + "weight_grid": "0.00 to 1.00 inclusive, step 0.01", + "pool": "log-linear/geometric", + "temperature_fit": "bounded scalar NLL minimization, T in [0.05, 20]", + "ece": "15 equal-width bins for Typed Decisions; 10 for every other dataset", + "selection_objective": "accuracy; calibrated NLL then distance from 0.5 break ties", + "transfer": "the selected weight and temperature are frozen for every held-out dataset", + "native_shipped": "the checkpoint's native probabilities, including its shipped inference temperature; no additional temperature is applied", + "fixed_blend_0_5": "equal log-linear pooling of T=1 choice probabilities and shipped native probabilities; no additional temperature is applied", + "transfer_dev_tuned_blend": "the Transfer-dev-selected native weight followed by the reported scalar additional temperature" + }, + "selected": { + "choice_temperature": 1.8514751919432642, + "native_weight": 0.0, + "blend_temperature": 1.8514751919432642, + "selection_objective": "hard-label accuracy on the Transfer-v9 development clean/knowable cohort; calibrated NLL and distance from 0.5 break ties" + } + }, + { + "key": "pointer4", + "model": "JevAny-Qwen3.5-4B-Pointer", + "family": "Qwen3.5-4B", + "weights": "jevany_sft", + "native_readout": "pointer", + "artifacts": { + "manifest": { + "path": "/lustre-storage/fsx/tianxinwei/JevAny/runs/choice-readout-v2-20261002-r1/pointer4/manifest.json", + "sha256": "11b64050dc31f834c64dbe54c975df1fb365b9eb9465a4089e7de98ba00242b1" + }, + "ensemble": { + "path": "/lustre-storage/fsx/tianxinwei/JevAny/runs/choice-readout-v2-20261002-r1/pointer4/ensemble.json", + "sha256": "0e18398586de8760719ea6a3e8956f221a69e341e75d3364e22fda7df9cb6488" + } + }, + "source": { + "base": null, + "checkpoint": "/storage/home/tianxinwei/.cache/huggingface/hub/models--SimpleJev--JevAny-Qwen3.5-4B-LoRA/snapshots/1c7aa9bab14ac347aeb917c0bcd757838a8a78ce" + }, + "public_source": { + "repository": "SimpleJev/JevAny-Qwen3.5-4B-LoRA", + "url": "https://huggingface.co/SimpleJev/JevAny-Qwen3.5-4B-LoRA", + "revision": "1c7aa9bab14ac347aeb917c0bcd757838a8a78ce", + "revision_status": "verified", + "repository_evidence": "manifest.source.checkpoint", + "revision_evidence": "manifest.source.checkpoint", + "evaluated_artifact_sha256": { + "adapter_model.safetensors": { + "sha256": "5acbaec8d554517250bf9c626b3f4b9e2ec31fc67957f3486aa72e2c91ee907f", + "bytes": 64993432 + }, + "head.pt": { + "sha256": "e5ea982efd06ecaf71b493045ebccdbfdca39f6fe128de164bf636d73152cd44", + "bytes": 5513315 + }, + "adapter_config.json": { + "sha256": "8a138b2c091d2ba463a8107e9e19182202395188750742047f4a12e8b59e1eda", + "bytes": 1265 + } + }, + "base_model": { + "repository": "Qwen/Qwen3.5-4B", + "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a" + }, + "limitation": null + }, + "checkpoint_artifacts": { + "requested": "/storage/home/tianxinwei/.cache/huggingface/hub/models--SimpleJev--JevAny-Qwen3.5-4B-LoRA/snapshots/1c7aa9bab14ac347aeb917c0bcd757838a8a78ce", + "resolved": "/storage/home/tianxinwei/.cache/huggingface/hub/models--SimpleJev--JevAny-Qwen3.5-4B-LoRA/snapshots/1c7aa9bab14ac347aeb917c0bcd757838a8a78ce", + "files": { + "adapter_model.safetensors": { + "sha256": "5acbaec8d554517250bf9c626b3f4b9e2ec31fc67957f3486aa72e2c91ee907f", + "bytes": 64993432 + }, + "head.pt": { + "sha256": "e5ea982efd06ecaf71b493045ebccdbfdca39f6fe128de164bf636d73152cd44", + "bytes": 5513315 + }, + "adapter_config.json": { + "sha256": "8a138b2c091d2ba463a8107e9e19182202395188750742047f4a12e8b59e1eda", + "bytes": 1265 + } + } + }, + "base_loading": { + "canonical_base": "Qwen/Qwen3.5-4B", + "canonical_base_is_local": false, + "canonical_revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", + "override_used": true, + "override_config_sha256": "ddc63e1c717afa86c865bb5e01313d89d72bb53b97ad4a8a03ba8510c0621670" + }, + "predictor": { + "method": "training-free exact choice-token projection", + "constraint_emulation": "exact decoded one-character option IDs", + "prompt": "Cygnet-compatible A-Z; extended with a-z above 26 options", + "system_prompt_sha256": "e2f0ac9f62eb446d7826b3abc3f7909287f4075175e9b58e4cd3ca72718671fe", + "chat_template_sha256": "a4aee8afcf2e0711942cf848899be66016f8d14a889ff9ede07bca099c28f715", + "prompt_format_version": 2, + "canonical_base": "Qwen/Qwen3.5-4B", + "canonical_revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", + "checkpoint": "/storage/home/tianxinwei/.cache/huggingface/hub/models--SimpleJev--JevAny-Qwen3.5-4B-LoRA/snapshots/1c7aa9bab14ac347aeb917c0bcd757838a8a78ce", + "adapter_applied": true, + "adapter_scale": 1.0, + "tied_output_embeddings": true, + "unused_full_output_head_released": false, + "letter_choice_token_rows": 52, + "output_softcap": null, + "temperature": 1.0, + "native_weight": 0.5, + "native_weight_argument": "native_weight", + "native_decision_mode": "pointer", + "pointer_weight": 0.5, + "native_temperature": 1.0, + "pointer_temperature": 1.0, + "return_components": true, + "exact_kernels": true, + "exact_cuda_kernels_applied": true, + "efficient_long_context_tokens": 8192, + "long_context_attention": "fused SDPA (Flash/Efficient, no math fallback) at or above the recorded threshold; exact math SDPA below it", + "max_tokens": 65536, + "backbone_context_window": 262144, + "effective_max_tokens": 65536, + "dtype": "torch.bfloat16", + "device": "cuda" + }, + "runtime": { + "python": "3.12.14", + "torch": "2.8.0+cu128", + "cuda": "12.8", + "transformers": "5.17.0", + "peft": "0.21.0", + "numpy": "2.5.3", + "pyarrow": "25.0.1", + "safetensors": "0.8.0", + "code_revision": "ae891ed97147b57b41a78a4bb4960b018db3bf42" + }, + "dataset_hashes": { + "jevbench_development_sha256": "b6d6eb34fdbbf46ec11be8657da5e513e3538990d7a00ca5582c55d8bd7941c6", + "jevbench_manifest_sha256": "e356ad5a2c1dd3917c126e74bdee58940391568de013502abcdcd902bde3c2a5", + "transfer_manifest_sha256": "3c4f0be94509a3612678bfd3a30fd99a8d0ca3c47ddfe7318075d95b2fa365e4", + "transfer_development_sha256": "53872ba045564315cf3c3dd8fb656e76441c6c514b1a7ff985ff0db3955710cf", + "transfer_test_sha256": "048abea5e12c73c3a622ce5836898c00fa43766e0bf48fb2459e77ee9873497f", + "typed_revision": "d0e2f0c42fef86cc15d1688d25a19f5ba7c85b18", + "typed_test_sha256": "4f294f218ea1da27f3efef936359389c62ea4d3973a41457732990f1d31b647c", + "jevjudge_test_sha256": "5c50f89e9ff1beb2107648b56f2a3cbc19c946e379519caf59afc191700e83cb", + "jevjudge_manifest_sha256": "ee52a15f82629742093a89910a8f72e7ca97248b197aad6ec801f984c2a15b68" + }, + "evaluation_protocol": { + "kernel_policy": "LocalPredictor-compatible exact CUDA policy", + "dtype": "bf16", + "attn": "sdpa", + "temperature": 1.0, + "blend_selection": "not performed here; components are saved for fitting on Transfer-v9 development only", + "efficiency_scope": "base calls execute choice only; checkpoint calls execute choice and native. Do not compare component efficiency from these end-to-end timings", + "long_context_attention": "exact math SDPA below the configured threshold; memory-linear fused SDPA without a math fallback at or above it", + "efficient_long_context_tokens": 8192, + "resume": { + "enabled": true, + "reused_datasets": [ + "jevbench_public", + "transfer_calibration", + "transfer_test", + "typed_test" + ], + "preexisting_code_revision": "256001ca806f9d7e302bea967850db214f3d8e4e", + "dataset_code_revisions": { + "jevbench_public": "256001ca806f9d7e302bea967850db214f3d8e4e", + "transfer_calibration": "256001ca806f9d7e302bea967850db214f3d8e4e", + "transfer_test": "256001ca806f9d7e302bea967850db214f3d8e4e", + "typed_test": "256001ca806f9d7e302bea967850db214f3d8e4e", + "jevjudge_text": "ae891ed97147b57b41a78a4bb4960b018db3bf42" + } + } + }, + "ensemble_protocol": { + "selection_split": "Transfer-v9 development clean/knowable accuracy cohort; never any reported test panel", + "selection_rows": 1046, + "weight_grid": "0.00 to 1.00 inclusive, step 0.01", + "pool": "log-linear/geometric", + "temperature_fit": "bounded scalar NLL minimization, T in [0.05, 20]", + "ece": "15 equal-width bins for Typed Decisions; 10 for every other dataset", + "selection_objective": "accuracy; calibrated NLL then distance from 0.5 break ties", + "transfer": "the selected weight and temperature are frozen for every held-out dataset", + "native_shipped": "the checkpoint's native probabilities, including its shipped inference temperature; no additional temperature is applied", + "fixed_blend_0_5": "equal log-linear pooling of T=1 choice probabilities and shipped native probabilities; no additional temperature is applied", + "transfer_dev_tuned_blend": "the Transfer-dev-selected native weight followed by the reported scalar additional temperature" + }, + "selected": { + "choice_temperature": 0.8505257300716687, + "native_weight": 0.78, + "blend_temperature": 1.0977784557618349, + "selection_objective": "hard-label accuracy on the Transfer-v9 development clean/knowable cohort; calibrated NLL and distance from 0.5 break ties" + } + }, + { + "key": "direct4", + "model": "JevAny-Qwen3.5-4B-Direct-Token", + "family": "Qwen3.5-4B", + "weights": "jevany_sft", + "native_readout": "direct-token", + "artifacts": { + "manifest": { + "path": "/lustre-storage/fsx/tianxinwei/JevAny/runs/choice-readout-v2-20261002-r1/direct4/manifest.json", + "sha256": "244ede1e22b445634be5792977e19e594bfaeb2fe783f905a6ccf2a0424078e5" + }, + "ensemble": { + "path": "/lustre-storage/fsx/tianxinwei/JevAny/runs/choice-readout-v2-20261002-r1/direct4/ensemble.json", + "sha256": "e3a2aee6ca34b98a0780807b6b49c86dcf261b55e5559da8fc61dd17b9f03db7" + } + }, + "source": { + "base": null, + "checkpoint": "/lustre-storage/fsx/tianxinwei/JevAny/releases/model-family-v2-20260929/JevAny-Qwen3.5-4B-Direct-Token-LoRA" + }, + "public_source": { + "repository": "SimpleJev/JevAny-Qwen3.5-4B-Direct-Token-LoRA", + "url": "https://huggingface.co/SimpleJev/JevAny-Qwen3.5-4B-Direct-Token-LoRA", + "revision": null, + "revision_status": "not_verified", + "repository_evidence": "results/model-family-v2.json#released_models", + "revision_evidence": null, + "evaluated_artifact_sha256": { + "adapter_model.safetensors": { + "sha256": "b768d54b9e5e20204b2db9e24c18ae6b10fa310b241d1931bac6b8882f127c53", + "bytes": 64993432 + }, + "head.pt": { + "sha256": "d3909cffc156d077061114627c8aa22f60c7a8f1bdcfd20c489bb3bb009fc4a6", + "bytes": 4699 + }, + "adapter_config.json": { + "sha256": "f84527483265b514d0a5f68d9c44e3d7952825c61d3659a858af8342e6ce9ff2", + "bytes": 1265 + } + }, + "base_model": { + "repository": "Qwen/Qwen3.5-4B", + "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a" + }, + "limitation": "The evaluated checkpoint came from a local release. Its exact files are pinned below by SHA-256, but the run manifest and release metadata do not prove an immutable public-repository revision for those bytes." + }, + "checkpoint_artifacts": { + "requested": "/lustre-storage/fsx/tianxinwei/JevAny/releases/model-family-v2-20260929/JevAny-Qwen3.5-4B-Direct-Token-LoRA", + "resolved": "/lustre-storage/fsx/tianxinwei/JevAny/releases/model-family-v2-20260929/JevAny-Qwen3.5-4B-Direct-Token-LoRA", + "files": { + "adapter_model.safetensors": { + "sha256": "b768d54b9e5e20204b2db9e24c18ae6b10fa310b241d1931bac6b8882f127c53", + "bytes": 64993432 + }, + "head.pt": { + "sha256": "d3909cffc156d077061114627c8aa22f60c7a8f1bdcfd20c489bb3bb009fc4a6", + "bytes": 4699 + }, + "adapter_config.json": { + "sha256": "f84527483265b514d0a5f68d9c44e3d7952825c61d3659a858af8342e6ce9ff2", + "bytes": 1265 + } + } + }, + "base_loading": { + "canonical_base": "Qwen/Qwen3.5-4B", + "canonical_base_is_local": false, + "canonical_revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", + "override_used": true, + "override_config_sha256": "ddc63e1c717afa86c865bb5e01313d89d72bb53b97ad4a8a03ba8510c0621670" + }, + "predictor": { + "method": "training-free exact choice-token projection", + "constraint_emulation": "exact decoded one-character option IDs", + "prompt": "Cygnet-compatible A-Z; extended with a-z above 26 options", + "system_prompt_sha256": "e2f0ac9f62eb446d7826b3abc3f7909287f4075175e9b58e4cd3ca72718671fe", + "chat_template_sha256": "a4aee8afcf2e0711942cf848899be66016f8d14a889ff9ede07bca099c28f715", + "prompt_format_version": 2, + "canonical_base": "Qwen/Qwen3.5-4B", + "canonical_revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", + "checkpoint": "/lustre-storage/fsx/tianxinwei/JevAny/releases/model-family-v2-20260929/JevAny-Qwen3.5-4B-Direct-Token-LoRA", + "adapter_applied": true, + "adapter_scale": 1.0, + "tied_output_embeddings": true, + "unused_full_output_head_released": false, + "letter_choice_token_rows": 52, + "output_softcap": null, + "temperature": 1.0, + "native_weight": 0.5, + "native_weight_argument": "native_weight", + "native_decision_mode": "lm_token", + "pointer_weight": 0.5, + "native_temperature": 1.0, + "pointer_temperature": 1.0, + "return_components": true, + "exact_kernels": true, + "exact_cuda_kernels_applied": true, + "efficient_long_context_tokens": 8192, + "long_context_attention": "fused SDPA (Flash/Efficient, no math fallback) at or above the recorded threshold; exact math SDPA below it", + "max_tokens": 65536, + "backbone_context_window": 262144, + "effective_max_tokens": 65536, + "dtype": "torch.bfloat16", + "device": "cuda" + }, + "runtime": { + "python": "3.12.14", + "torch": "2.8.0+cu128", + "cuda": "12.8", + "transformers": "5.17.0", + "peft": "0.21.0", + "numpy": "2.5.3", + "pyarrow": "25.0.1", + "safetensors": "0.8.0", + "code_revision": "ae891ed97147b57b41a78a4bb4960b018db3bf42" + }, + "dataset_hashes": { + "jevbench_development_sha256": "b6d6eb34fdbbf46ec11be8657da5e513e3538990d7a00ca5582c55d8bd7941c6", + "jevbench_manifest_sha256": "e356ad5a2c1dd3917c126e74bdee58940391568de013502abcdcd902bde3c2a5", + "transfer_manifest_sha256": "3c4f0be94509a3612678bfd3a30fd99a8d0ca3c47ddfe7318075d95b2fa365e4", + "transfer_development_sha256": "53872ba045564315cf3c3dd8fb656e76441c6c514b1a7ff985ff0db3955710cf", + "transfer_test_sha256": "048abea5e12c73c3a622ce5836898c00fa43766e0bf48fb2459e77ee9873497f", + "typed_revision": "d0e2f0c42fef86cc15d1688d25a19f5ba7c85b18", + "typed_test_sha256": "4f294f218ea1da27f3efef936359389c62ea4d3973a41457732990f1d31b647c", + "jevjudge_test_sha256": "5c50f89e9ff1beb2107648b56f2a3cbc19c946e379519caf59afc191700e83cb", + "jevjudge_manifest_sha256": "ee52a15f82629742093a89910a8f72e7ca97248b197aad6ec801f984c2a15b68" + }, + "evaluation_protocol": { + "kernel_policy": "LocalPredictor-compatible exact CUDA policy", + "dtype": "bf16", + "attn": "sdpa", + "temperature": 1.0, + "blend_selection": "not performed here; components are saved for fitting on Transfer-v9 development only", + "efficiency_scope": "base calls execute choice only; checkpoint calls execute choice and native. Do not compare component efficiency from these end-to-end timings", + "long_context_attention": "exact math SDPA below the configured threshold; memory-linear fused SDPA without a math fallback at or above it", + "efficient_long_context_tokens": 8192, + "resume": { + "enabled": true, + "reused_datasets": [ + "jevbench_public", + "transfer_calibration", + "transfer_test", + "typed_test" + ], + "preexisting_code_revision": "256001ca806f9d7e302bea967850db214f3d8e4e", + "dataset_code_revisions": { + "jevbench_public": "256001ca806f9d7e302bea967850db214f3d8e4e", + "transfer_calibration": "256001ca806f9d7e302bea967850db214f3d8e4e", + "transfer_test": "256001ca806f9d7e302bea967850db214f3d8e4e", + "typed_test": "256001ca806f9d7e302bea967850db214f3d8e4e", + "jevjudge_text": "ae891ed97147b57b41a78a4bb4960b018db3bf42" + } + } + }, + "ensemble_protocol": { + "selection_split": "Transfer-v9 development clean/knowable accuracy cohort; never any reported test panel", + "selection_rows": 1046, + "weight_grid": "0.00 to 1.00 inclusive, step 0.01", + "pool": "log-linear/geometric", + "temperature_fit": "bounded scalar NLL minimization, T in [0.05, 20]", + "ece": "15 equal-width bins for Typed Decisions; 10 for every other dataset", + "selection_objective": "accuracy; calibrated NLL then distance from 0.5 break ties", + "transfer": "the selected weight and temperature are frozen for every held-out dataset", + "native_shipped": "the checkpoint's native probabilities, including its shipped inference temperature; no additional temperature is applied", + "fixed_blend_0_5": "equal log-linear pooling of T=1 choice probabilities and shipped native probabilities; no additional temperature is applied", + "transfer_dev_tuned_blend": "the Transfer-dev-selected native weight followed by the reported scalar additional temperature" + }, + "selected": { + "choice_temperature": 1.246278988841639, + "native_weight": 0.59, + "blend_temperature": 1.1240763000234277, + "selection_objective": "hard-label accuracy on the Transfer-v9 development clean/knowable cohort; calibrated NLL and distance from 0.5 break ties" + } + }, + { + "key": "base27", + "model": "Frozen-Qwen3.8-27B", + "family": "Qwen3.8-27B", + "weights": "frozen_base", + "native_readout": null, + "artifacts": { + "manifest": { + "path": "/lustre-storage/fsx/tianxinwei/JevAny/runs/choice-readout-v2-20261002-r1/base27/manifest.json", + "sha256": "5f1d1c6a755bdb8f8ee5d37f011345213cbfa205256457622d08a88f88d81fd2" + }, + "ensemble": { + "path": "/lustre-storage/fsx/tianxinwei/JevAny/runs/choice-readout-v2-20261002-r1/base27/ensemble.json", + "sha256": "c18a65303093f9673f22cc7a9fcbd7de89a73cd46372b04311e935f01277e6fb" + } + }, + "source": { + "base": "Qwen/Qwen3.8-27B", + "checkpoint": null + }, + "public_source": { + "repository": "Qwen/Qwen3.8-27B", + "url": "https://huggingface.co/Qwen/Qwen3.8-27B", + "revision": "1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0", + "revision_status": "verified", + "repository_evidence": "manifest.base_loading.canonical_base", + "revision_evidence": "manifest.base_loading.canonical_revision", + "evaluated_artifact_sha256": {}, + "base_model": { + "repository": "Qwen/Qwen3.8-27B", + "revision": "1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0" + }, + "limitation": null + }, + "checkpoint_artifacts": null, + "base_loading": { + "canonical_base": "Qwen/Qwen3.8-27B", + "canonical_base_is_local": false, + "canonical_revision": "1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0", + "override_used": true, + "override_config_sha256": "191e0af232104ed8b65258cf3fb2b842e288008baca7633c11b82a1ac7203aab" + }, + "predictor": { + "method": "training-free exact choice-token projection", + "constraint_emulation": "exact decoded one-character option IDs", + "prompt": "Cygnet-compatible A-Z; extended with a-z above 26 options", + "system_prompt_sha256": "e2f0ac9f62eb446d7826b3abc3f7909287f4075175e9b58e4cd3ca72718671fe", + "chat_template_sha256": "c3cf9e34abf4f9e36c2d72165aa9c132d3e2a725b6c2586aaa3a8af9d7a81041", + "prompt_format_version": 2, + "canonical_base": "Qwen/Qwen3.8-27B", + "canonical_revision": "1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0", + "checkpoint": null, + "adapter_applied": false, + "adapter_scale": 0.0, + "tied_output_embeddings": false, + "unused_full_output_head_released": false, + "letter_choice_token_rows": 52, + "output_softcap": null, + "temperature": 1.0, + "native_weight": 0.0, + "native_weight_argument": "native_weight", + "native_decision_mode": null, + "pointer_weight": 0.0, + "native_temperature": null, + "pointer_temperature": null, + "return_components": false, + "exact_kernels": true, + "exact_cuda_kernels_applied": true, + "efficient_long_context_tokens": 8192, + "long_context_attention": "fused SDPA (Flash/Efficient, no math fallback) at or above the recorded threshold; exact math SDPA below it", + "max_tokens": 65536, + "backbone_context_window": 262144, + "effective_max_tokens": 65536, + "dtype": "torch.bfloat16", + "device": "cuda" + }, + "runtime": { + "python": "3.12.14", + "torch": "2.8.0+cu128", + "cuda": "12.8", + "transformers": "5.17.0", + "peft": "0.21.0", + "numpy": "2.5.3", + "pyarrow": "25.0.1", + "safetensors": "0.8.0", + "code_revision": "ae891ed97147b57b41a78a4bb4960b018db3bf42" + }, + "dataset_hashes": { + "jevbench_development_sha256": "b6d6eb34fdbbf46ec11be8657da5e513e3538990d7a00ca5582c55d8bd7941c6", + "jevbench_manifest_sha256": "e356ad5a2c1dd3917c126e74bdee58940391568de013502abcdcd902bde3c2a5", + "transfer_manifest_sha256": "3c4f0be94509a3612678bfd3a30fd99a8d0ca3c47ddfe7318075d95b2fa365e4", + "transfer_development_sha256": "53872ba045564315cf3c3dd8fb656e76441c6c514b1a7ff985ff0db3955710cf", + "transfer_test_sha256": "048abea5e12c73c3a622ce5836898c00fa43766e0bf48fb2459e77ee9873497f", + "typed_revision": "d0e2f0c42fef86cc15d1688d25a19f5ba7c85b18", + "typed_test_sha256": "4f294f218ea1da27f3efef936359389c62ea4d3973a41457732990f1d31b647c", + "jevjudge_test_sha256": "5c50f89e9ff1beb2107648b56f2a3cbc19c946e379519caf59afc191700e83cb", + "jevjudge_manifest_sha256": "ee52a15f82629742093a89910a8f72e7ca97248b197aad6ec801f984c2a15b68" + }, + "evaluation_protocol": { + "kernel_policy": "LocalPredictor-compatible exact CUDA policy", + "dtype": "bf16", + "attn": "sdpa", + "temperature": 1.0, + "blend_selection": "not performed here; components are saved for fitting on Transfer-v9 development only", + "efficiency_scope": "base calls execute choice only; checkpoint calls execute choice and native. Do not compare component efficiency from these end-to-end timings", + "long_context_attention": "exact math SDPA below the configured threshold; memory-linear fused SDPA without a math fallback at or above it", + "efficient_long_context_tokens": 8192, + "resume": { + "enabled": true, + "reused_datasets": [ + "jevbench_public", + "transfer_calibration", + "transfer_test", + "typed_test" + ], + "preexisting_code_revision": "256001ca806f9d7e302bea967850db214f3d8e4e", + "dataset_code_revisions": { + "jevbench_public": "256001ca806f9d7e302bea967850db214f3d8e4e", + "transfer_calibration": "256001ca806f9d7e302bea967850db214f3d8e4e", + "transfer_test": "256001ca806f9d7e302bea967850db214f3d8e4e", + "typed_test": "256001ca806f9d7e302bea967850db214f3d8e4e", + "jevjudge_text": "ae891ed97147b57b41a78a4bb4960b018db3bf42" + } + } + }, + "ensemble_protocol": { + "selection_split": "Transfer-v9 development clean/knowable accuracy cohort; never any reported test panel", + "selection_rows": 1046, + "weight_grid": "0.00 to 1.00 inclusive, step 0.01", + "pool": "log-linear/geometric", + "temperature_fit": "bounded scalar NLL minimization, T in [0.05, 20]", + "ece": "15 equal-width bins for Typed Decisions; 10 for every other dataset", + "selection_objective": "accuracy; calibrated NLL then distance from 0.5 break ties", + "transfer": "the selected weight and temperature are frozen for every held-out dataset", + "native_shipped": "the checkpoint's native probabilities, including its shipped inference temperature; no additional temperature is applied", + "fixed_blend_0_5": "equal log-linear pooling of T=1 choice probabilities and shipped native probabilities; no additional temperature is applied", + "transfer_dev_tuned_blend": "the Transfer-dev-selected native weight followed by the reported scalar additional temperature" + }, + "selected": { + "choice_temperature": 1.718360083457113, + "native_weight": 0.0, + "blend_temperature": 1.718360083457113, + "selection_objective": "hard-label accuracy on the Transfer-v9 development clean/knowable cohort; calibrated NLL and distance from 0.5 break ties" + } + }, + { + "key": "pointer27", + "model": "JevAny-Qwen3.8-27B-Pointer", + "family": "Qwen3.8-27B", + "weights": "jevany_sft", + "native_readout": "pointer", + "artifacts": { + "manifest": { + "path": "/lustre-storage/fsx/tianxinwei/JevAny/runs/choice-readout-v2-20261002-r1/pointer27/manifest.json", + "sha256": "c2d6dd8bb4b9f2b9e5b9c57088f673e0839af5435acb88b561559ae0a824367a" + }, + "ensemble": { + "path": "/lustre-storage/fsx/tianxinwei/JevAny/runs/choice-readout-v2-20261002-r1/pointer27/ensemble.json", + "sha256": "2270c1a3309b20fbe77bf6e905e2677ba29ecc3a3c9d4251b9932ee887e64826" + } + }, + "source": { + "base": null, + "checkpoint": "/storage/home/tianxinwei/.cache/huggingface/hub/models--SimpleJev--JevAny-Qwen3.8-27B-LoRA/snapshots/09c9e9102d5b8cc7d56558d25da1202a761b6c0d" + }, + "public_source": { + "repository": "SimpleJev/JevAny-Qwen3.8-27B-LoRA", + "url": "https://huggingface.co/SimpleJev/JevAny-Qwen3.8-27B-LoRA", + "revision": "09c9e9102d5b8cc7d56558d25da1202a761b6c0d", + "revision_status": "verified", + "repository_evidence": "manifest.source.checkpoint", + "revision_evidence": "manifest.source.checkpoint", + "evaluated_artifact_sha256": { + "adapter_model.safetensors": { + "sha256": "f4b835f4f7be78a819067dac226feaad2fc7b9eb0b968c5e6e665ad5328e6731", + "bytes": 233584648 + }, + "head.pt": { + "sha256": "f78fa283f7e979ce418c5b137d24459350fecaea73e5766894532c1cd99684c7", + "bytes": 10756195 + }, + "adapter_config.json": { + "sha256": "31d8e262a21a45b80fd409f08d47930effd6cd9bc12156c59ad61e627f50a275", + "bytes": 1266 + } + }, + "base_model": { + "repository": "Qwen/Qwen3.8-27B", + "revision": "1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0" + }, + "limitation": null + }, + "checkpoint_artifacts": { + "requested": "/storage/home/tianxinwei/.cache/huggingface/hub/models--SimpleJev--JevAny-Qwen3.8-27B-LoRA/snapshots/09c9e9102d5b8cc7d56558d25da1202a761b6c0d", + "resolved": "/storage/home/tianxinwei/.cache/huggingface/hub/models--SimpleJev--JevAny-Qwen3.8-27B-LoRA/snapshots/09c9e9102d5b8cc7d56558d25da1202a761b6c0d", + "files": { + "adapter_model.safetensors": { + "sha256": "f4b835f4f7be78a819067dac226feaad2fc7b9eb0b968c5e6e665ad5328e6731", + "bytes": 233584648 + }, + "head.pt": { + "sha256": "f78fa283f7e979ce418c5b137d24459350fecaea73e5766894532c1cd99684c7", + "bytes": 10756195 + }, + "adapter_config.json": { + "sha256": "31d8e262a21a45b80fd409f08d47930effd6cd9bc12156c59ad61e627f50a275", + "bytes": 1266 + } + } + }, + "base_loading": { + "canonical_base": "Qwen/Qwen3.8-27B", + "canonical_base_is_local": false, + "canonical_revision": "1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0", + "override_used": true, + "override_config_sha256": "191e0af232104ed8b65258cf3fb2b842e288008baca7633c11b82a1ac7203aab" + }, + "predictor": { + "method": "training-free exact choice-token projection", + "constraint_emulation": "exact decoded one-character option IDs", + "prompt": "Cygnet-compatible A-Z; extended with a-z above 26 options", + "system_prompt_sha256": "e2f0ac9f62eb446d7826b3abc3f7909287f4075175e9b58e4cd3ca72718671fe", + "chat_template_sha256": "c3cf9e34abf4f9e36c2d72165aa9c132d3e2a725b6c2586aaa3a8af9d7a81041", + "prompt_format_version": 2, + "canonical_base": "Qwen/Qwen3.8-27B", + "canonical_revision": "1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0", + "checkpoint": "/storage/home/tianxinwei/.cache/huggingface/hub/models--SimpleJev--JevAny-Qwen3.8-27B-LoRA/snapshots/09c9e9102d5b8cc7d56558d25da1202a761b6c0d", + "adapter_applied": true, + "adapter_scale": 1.0, + "tied_output_embeddings": false, + "unused_full_output_head_released": false, + "letter_choice_token_rows": 52, + "output_softcap": null, + "temperature": 1.0, + "native_weight": 0.5, + "native_weight_argument": "native_weight", + "native_decision_mode": "pointer", + "pointer_weight": 0.5, + "native_temperature": 1.0, + "pointer_temperature": 1.0, + "return_components": true, + "exact_kernels": true, + "exact_cuda_kernels_applied": true, + "efficient_long_context_tokens": 8192, + "long_context_attention": "fused SDPA (Flash/Efficient, no math fallback) at or above the recorded threshold; exact math SDPA below it", + "max_tokens": 65536, + "backbone_context_window": 262144, + "effective_max_tokens": 65536, + "dtype": "torch.bfloat16", + "device": "cuda" + }, + "runtime": { + "python": "3.12.14", + "torch": "2.8.0+cu128", + "cuda": "12.8", + "transformers": "5.17.0", + "peft": "0.21.0", + "numpy": "2.5.3", + "pyarrow": "25.0.1", + "safetensors": "0.8.0", + "code_revision": "ae891ed97147b57b41a78a4bb4960b018db3bf42" + }, + "dataset_hashes": { + "jevbench_development_sha256": "b6d6eb34fdbbf46ec11be8657da5e513e3538990d7a00ca5582c55d8bd7941c6", + "jevbench_manifest_sha256": "e356ad5a2c1dd3917c126e74bdee58940391568de013502abcdcd902bde3c2a5", + "transfer_manifest_sha256": "3c4f0be94509a3612678bfd3a30fd99a8d0ca3c47ddfe7318075d95b2fa365e4", + "transfer_development_sha256": "53872ba045564315cf3c3dd8fb656e76441c6c514b1a7ff985ff0db3955710cf", + "transfer_test_sha256": "048abea5e12c73c3a622ce5836898c00fa43766e0bf48fb2459e77ee9873497f", + "typed_revision": "d0e2f0c42fef86cc15d1688d25a19f5ba7c85b18", + "typed_test_sha256": "4f294f218ea1da27f3efef936359389c62ea4d3973a41457732990f1d31b647c", + "jevjudge_test_sha256": "5c50f89e9ff1beb2107648b56f2a3cbc19c946e379519caf59afc191700e83cb", + "jevjudge_manifest_sha256": "ee52a15f82629742093a89910a8f72e7ca97248b197aad6ec801f984c2a15b68" + }, + "evaluation_protocol": { + "kernel_policy": "LocalPredictor-compatible exact CUDA policy", + "dtype": "bf16", + "attn": "sdpa", + "temperature": 1.0, + "blend_selection": "not performed here; components are saved for fitting on Transfer-v9 development only", + "efficiency_scope": "base calls execute choice only; checkpoint calls execute choice and native. Do not compare component efficiency from these end-to-end timings", + "long_context_attention": "exact math SDPA below the configured threshold; memory-linear fused SDPA without a math fallback at or above it", + "efficient_long_context_tokens": 8192, + "resume": { + "enabled": true, + "reused_datasets": [ + "jevbench_public", + "transfer_calibration", + "transfer_test", + "typed_test" + ], + "preexisting_code_revision": "256001ca806f9d7e302bea967850db214f3d8e4e", + "dataset_code_revisions": { + "jevbench_public": "256001ca806f9d7e302bea967850db214f3d8e4e", + "transfer_calibration": "256001ca806f9d7e302bea967850db214f3d8e4e", + "transfer_test": "256001ca806f9d7e302bea967850db214f3d8e4e", + "typed_test": "256001ca806f9d7e302bea967850db214f3d8e4e", + "jevjudge_text": "ae891ed97147b57b41a78a4bb4960b018db3bf42" + } + } + }, + "ensemble_protocol": { + "selection_split": "Transfer-v9 development clean/knowable accuracy cohort; never any reported test panel", + "selection_rows": 1046, + "weight_grid": "0.00 to 1.00 inclusive, step 0.01", + "pool": "log-linear/geometric", + "temperature_fit": "bounded scalar NLL minimization, T in [0.05, 20]", + "ece": "15 equal-width bins for Typed Decisions; 10 for every other dataset", + "selection_objective": "accuracy; calibrated NLL then distance from 0.5 break ties", + "transfer": "the selected weight and temperature are frozen for every held-out dataset", + "native_shipped": "the checkpoint's native probabilities, including its shipped inference temperature; no additional temperature is applied", + "fixed_blend_0_5": "equal log-linear pooling of T=1 choice probabilities and shipped native probabilities; no additional temperature is applied", + "transfer_dev_tuned_blend": "the Transfer-dev-selected native weight followed by the reported scalar additional temperature" + }, + "selected": { + "choice_temperature": 0.5969253208700946, + "native_weight": 0.52, + "blend_temperature": 0.8352476617411964, + "selection_objective": "hard-label accuracy on the Transfer-v9 development clean/knowable cohort; calibrated NLL and distance from 0.5 break ties" + } + } + ], + "external_source": { + "path": "results/external-zero-shot-v1.json", + "sha256": "8a589a7e5d0d5ae5232fbdf30a71cdce5399cb5f5aacb1fcaf974057be48b3b9", + "artifact_version": 5, + "provenance": { + "typed_existing": { + "path": "results/external-decision-evals-20261001/typed-decisions.json", + "sha256": "a9fbcc06a5e957630e83f18c5dc01d41ff7b3e2289b0d816f2e1459ef814341e" + }, + "typed_open_summary": { + "local_run_path": "runs/external-open-baselines-20261001/summary.json", + "sha256": "4b4f06f8f02ce6cb4a610f110450469c46b917fd9aa1d571077d0c054320eab2" + }, + "jevjudge_full_context_scores": { + "local_run_path": "runs/jevjudge-open-baselines-20261001/scores-full-context", + "protocol": "native inputs, no truncation, 65,536-token shared context ceiling", + "main_cohort_sha256_by_file": { + "bongard-mini.json": "f3fa601f617ed2e70b31d98fe2c589ee8787ee47624233f3fbd8b3d02630daa9", + "jeff-gemma4-e2b.json": "78fd3e26bf01b5630e23656049f2a33e3a77e81a0381d6106c2ad3afce4d0491", + "jeff-qwen35-08b.json": "ea2dc271c6d23c11278e6fdcdf56a9f45a1938404577d98ce101ec6669f58dbd", + "jeff-qwen35-2b.json": "8d76f061b8cd67811d48b1eb5467e9274c1648e47cb3c835974c8b32c647a404", + "jevany-gemma-4b-step2771.json": "88561342bec2909101bb8e3bb90a7f9a73626a35056468fae0d27f4e0389c39a", + "jevany-muse-glimmer-30b-step3324.json": "8022d3ff2cb33c602bf6b3aca55cdc10840f1d0735bd523c53dacb0fb966ba00", + "jevany-qwen35-4b-direct-release.json": "db0b29b095e73e6fe6f0795632686a05e0057f03f7b27eae46121db6854eee02", + "jevany-qwen35-4b-pointer-step13850.json": "c7c70a9b2adda41a597bbdd2c354320f5caf2fe4ce09d24c3fb5731615ff5cc0", + "jevany-qwen38-27b-step44319.json": "fcbdda6cb097f84b8bade8b7272f588dd049d32eb0ccd1f59cb6a76850954661", + "opendecider-small.json": "438b89596d7fe7aa593d1d024ef6576082d074e0b488eebc1c4336ee65a28032" + }, + "supplemental_sha256_by_file": { + "kev-27b.json": "056a42c2fef59f5ecffb40e342690ac4467b273f374cbf5d59ec9dbbe3df0cae", + "kev-4b.json": "c6985eef36f04b962c17bc72ca3bd6d2235bcf0619904db6dda135a649291db1", + "laya.json": "abde9d49d8c100b8c6c09ec6b3aaea442f0f9b2f1350e34701d1b615e1651b90" + }, + "supplemental_protocol": "Mixed audit protocols; these hashes are not covered by the no-truncation main-cohort claim." + }, + "jevjudge_full_jevany": { + "local_run_path": "runs/jevjudge-full-current-20261001", + "summary_sha256": "fcc3025cd29ec057d869067edc5d45f5d154b8596f3281aabc4d0c83996b6609", + "final_audit_sha256": "ad0055609f1a3dd87210738acb45a8fce113c2ad0fc9b507af8ddce8d644d773", + "official_scorer_sha256": "477b39c76dd6e6200f0919fcb56e80045ffd29b66445496d5e6e70c3d1537141", + "bootstrap_scorer_sha256": "8f7413d70290bd1009acd398478bb26cbc98f89ebd19af298fc0930b7a763133", + "protocol": "native text/image/video inputs; 65,536-token ceiling; no token, option, or media truncation; 3,220/3,220 per model" + }, + "jevjudge_full_open": { + "local_run_path": "worktrees/JevAny/jevjudge-full-open-20261001/runs/jevjudge-full-open-20261001-v2", + "original_summary_sha256": "f34fc4265b6cf5847d4f1bd46a5601111cc927f57c4e6178acbf61e496477c51", + "corrected_final_summary_sha256": "3c348c94f0494815ae4236ffb2524e92d6cdd3d237812d21465e64e18ddc00b3", + "compatibility_matrix_sha256": "128cc4d00520fa0515f3f5d5172bc4d704c08ebea63ba8b60f93f225db4ac98e", + "note": "Published headline values use returned-probability NLL and the final source-stratified group bootstrap; the original summary retained a raw-logit NLL diagnostic." + }, + "jevjudge_kev_text": { + "local_run_path": "worktrees/JevAny/kev-jevjudge-20261001/runs/jevjudge-kev-full-context-20261001", + "summary_sha256": "4e1f448503983a13a14df9fc877cd876906c09915ef8173bc676778d52e06598", + "protocol": "724/724 text records; 65,536-token ceiling; 4,096-token chunked-KV prefill; no input truncation" + }, + "jevjudge_jev_113_openrouter": { + "source_commit": "2ff9d3019bd7f910486543ce76513321537a3c6f", + "protocol": "OpenRouter full-context rerun; 724/724 JevJudge text-only records", + "note": "Aggregate metrics were added directly to the public evaluation table; item-level outputs are not tracked in this artifact." + } + } + }, + "repository_catalog": { + "path": "results/model-family-v2.json", + "sha256": "a49f03bd9371f1846d78a09682f000cba9f64004d8d5a222b3868f11ccb48d8c", + "release": "model-family-v2" + }, + "datasets": { + "transfer_calibration": { + "label": "Transfer-v9 development", + "role": "selection_only", + "records": 1264, + "questions": 1264, + "headline_n": 1046, + "runs": [ + { + "run": "base4", + "model": "Frozen-Qwen3.5-4B", + "family": "Qwen3.5-4B", + "weights": "frozen_base", + "native_readout": null, + "zero_shot": { + "choice_t1": { + "n": 1046, + "correct": 725, + "accuracy": 0.6931166347992351, + "cross_entropy_from_gold": 0.8799398258908452, + "gold_entropy": 0.0, + "kl_from_gold": 0.8799398258908452, + "brier": 0.43925607559848767, + "ece": 0.12305476565507475, + "mean_confidence": 0.8154458555896062, + "cohort": "clean and knowable", + "excluded_non_headline_rows": 218, + "ece_bins": 10 + } + }, + "transfer_dev_tuned": { + "transfer_dev_calibrated_choice": { + "n": 1046, + "correct": 725, + "accuracy": 0.6931166347992351, + "cross_entropy_from_gold": 0.7684499162138351, + "gold_entropy": 0.0, + "kl_from_gold": 0.7684499162138351, + "brier": 0.40380428797912027, + "ece": 0.0442424982205075, + "mean_confidence": 0.7062354250490573, + "cohort": "clean and knowable", + "excluded_non_headline_rows": 218, + "ece_bins": 10 + } + } + }, + { + "run": "pointer4", + "model": "JevAny-Qwen3.5-4B-Pointer", + "family": "Qwen3.5-4B", + "weights": "jevany_sft", + "native_readout": "pointer", + "zero_shot": { + "choice_t1": { + "n": 1046, + "correct": 784, + "accuracy": 0.7495219885277247, + "cross_entropy_from_gold": 0.6801893631325125, + "gold_entropy": 0.0, + "kl_from_gold": 0.6801893631325125, + "brier": 0.3338033580246269, + "ece": 0.054371329145242946, + "mean_confidence": 0.6968489272720685, + "cohort": "clean and knowable", + "excluded_non_headline_rows": 218, + "ece_bins": 10 + }, + "native_shipped": { + "n": 1046, + "correct": 823, + "accuracy": 0.7868068833652008, + "cross_entropy_from_gold": 0.5868780830556992, + "gold_entropy": 0.0, + "kl_from_gold": 0.5868780830556992, + "brier": 0.2965183466659115, + "ece": 0.03479245011625781, + "mean_confidence": 0.805131298589329, + "cohort": "clean and knowable", + "excluded_non_headline_rows": 218, + "ece_bins": 10 + }, + "fixed_blend_0_5": { + "n": 1046, + "correct": 822, + "accuracy": 0.7858508604206501, + "cross_entropy_from_gold": 0.5717414381657305, + "gold_entropy": 0.0, + "kl_from_gold": 0.5717414381657305, + "brier": 0.2919156848114442, + "ece": 0.03811670878840213, + "mean_confidence": 0.7688884953818541, + "cohort": "clean and knowable", + "excluded_non_headline_rows": 218, + "ece_bins": 10 + } + }, + "transfer_dev_tuned": { + "transfer_dev_calibrated_choice": { + "n": 1046, + "correct": 784, + "accuracy": 0.7495219885277247, + "cross_entropy_from_gold": 0.675438096110365, + "gold_entropy": 0.0, + "kl_from_gold": 0.675438096110365, + "brier": 0.3287201057208544, + "ece": 0.030612288328262394, + "mean_confidence": 0.7238557515998861, + "cohort": "clean and knowable", + "excluded_non_headline_rows": 218, + "ece_bins": 10 + }, + "transfer_dev_tuned_blend": { + "n": 1046, + "correct": 828, + "accuracy": 0.7915869980879541, + "cross_entropy_from_gold": 0.569720964551106, + "gold_entropy": 0.0, + "kl_from_gold": 0.569720964551106, + "brier": 0.2929198506354004, + "ece": 0.05359065555472778, + "mean_confidence": 0.7774203199582407, + "cohort": "clean and knowable", + "excluded_non_headline_rows": 218, + "ece_bins": 10 + } + } + }, + { + "run": "direct4", + "model": "JevAny-Qwen3.5-4B-Direct-Token", + "family": "Qwen3.5-4B", + "weights": "jevany_sft", + "native_readout": "direct-token", + "zero_shot": { + "choice_t1": { + "n": 1046, + "correct": 796, + "accuracy": 0.7609942638623327, + "cross_entropy_from_gold": 0.6010867296757888, + "gold_entropy": 0.0, + "kl_from_gold": 0.6010867296757888, + "brier": 0.3064220873561066, + "ece": 0.04138187687571737, + "mean_confidence": 0.796631394719685, + "cohort": "clean and knowable", + "excluded_non_headline_rows": 218, + "ece_bins": 10 + }, + "native_shipped": { + "n": 1046, + "correct": 818, + "accuracy": 0.7820267686424475, + "cross_entropy_from_gold": 0.5642005514977974, + "gold_entropy": 0.0, + "kl_from_gold": 0.5642005514977974, + "brier": 0.2908559938309828, + "ece": 0.029163506342423436, + "mean_confidence": 0.8005854296720307, + "cohort": "clean and knowable", + "excluded_non_headline_rows": 218, + "ece_bins": 10 + }, + "fixed_blend_0_5": { + "n": 1046, + "correct": 821, + "accuracy": 0.7848948374760994, + "cross_entropy_from_gold": 0.5635679825981634, + "gold_entropy": 0.0, + "kl_from_gold": 0.5635679825981634, + "brier": 0.28919182401119004, + "ece": 0.03032435006136549, + "mean_confidence": 0.79621407300774, + "cohort": "clean and knowable", + "excluded_non_headline_rows": 218, + "ece_bins": 10 + } + }, + "transfer_dev_tuned": { + "transfer_dev_calibrated_choice": { + "n": 1046, + "correct": 796, + "accuracy": 0.7609942638623327, + "cross_entropy_from_gold": 0.5914027552210456, + "gold_entropy": 0.0, + "kl_from_gold": 0.5914027552210456, + "brier": 0.3043378646694837, + "ece": 0.03285794273637177, + "mean_confidence": 0.765631487860452, + "cohort": "clean and knowable", + "excluded_non_headline_rows": 218, + "ece_bins": 10 + }, + "transfer_dev_tuned_blend": { + "n": 1046, + "correct": 826, + "accuracy": 0.7896749521988528, + "cross_entropy_from_gold": 0.5583414368363817, + "gold_entropy": 0.0, + "kl_from_gold": 0.5583414368363817, + "brier": 0.2874056312419155, + "ece": 0.027602128315946592, + "mean_confidence": 0.7807455671293496, + "cohort": "clean and knowable", + "excluded_non_headline_rows": 218, + "ece_bins": 10 + } + } + }, + { + "run": "base27", + "model": "Frozen-Qwen3.8-27B", + "family": "Qwen3.8-27B", + "weights": "frozen_base", + "native_readout": null, + "zero_shot": { + "choice_t1": { + "n": 1046, + "correct": 820, + "accuracy": 0.7839388145315488, + "cross_entropy_from_gold": 0.6672937347936473, + "gold_entropy": 0.0, + "kl_from_gold": 0.6672937347936473, + "brier": 0.31577326969120223, + "ece": 0.08097099874334637, + "mean_confidence": 0.8596910291171252, + "cohort": "clean and knowable", + "excluded_non_headline_rows": 218, + "ece_bins": 10 + } + }, + "transfer_dev_tuned": { + "transfer_dev_calibrated_choice": { + "n": 1046, + "correct": 820, + "accuracy": 0.7839388145315488, + "cross_entropy_from_gold": 0.5895245884799086, + "gold_entropy": 0.0, + "kl_from_gold": 0.5895245884799086, + "brier": 0.3023594256806041, + "ece": 0.029495602615825022, + "mean_confidence": 0.7806446298690051, + "cohort": "clean and knowable", + "excluded_non_headline_rows": 218, + "ece_bins": 10 + } + } + }, + { + "run": "pointer27", + "model": "JevAny-Qwen3.8-27B-Pointer", + "family": "Qwen3.8-27B", + "weights": "jevany_sft", + "native_readout": "pointer", + "zero_shot": { + "choice_t1": { + "n": 1046, + "correct": 885, + "accuracy": 0.8460803059273423, + "cross_entropy_from_gold": 0.5085522464835983, + "gold_entropy": 0.0, + "kl_from_gold": 0.5085522464835983, + "brier": 0.2411187763545078, + "ece": 0.10827608077576636, + "mean_confidence": 0.7390643592114102, + "cohort": "clean and knowable", + "excluded_non_headline_rows": 218, + "ece_bins": 10 + }, + "native_shipped": { + "n": 1046, + "correct": 900, + "accuracy": 0.8604206500956023, + "cross_entropy_from_gold": 0.3878296713265552, + "gold_entropy": 0.0, + "kl_from_gold": 0.3878296713265552, + "brier": 0.19524463105141282, + "ece": 0.026304156956950205, + "mean_confidence": 0.8776426213905197, + "cohort": "clean and knowable", + "excluded_non_headline_rows": 218, + "ece_bins": 10 + }, + "fixed_blend_0_5": { + "n": 1046, + "correct": 907, + "accuracy": 0.8671128107074569, + "cross_entropy_from_gold": 0.3773854700281526, + "gold_entropy": 0.0, + "kl_from_gold": 0.3773854700281526, + "brier": 0.19080039228840473, + "ece": 0.030952348286441715, + "mean_confidence": 0.8386108687417092, + "cohort": "clean and knowable", + "excluded_non_headline_rows": 218, + "ece_bins": 10 + } + }, + "transfer_dev_tuned": { + "transfer_dev_calibrated_choice": { + "n": 1046, + "correct": 885, + "accuracy": 0.8460803059273423, + "cross_entropy_from_gold": 0.4457110992469971, + "gold_entropy": 0.0, + "kl_from_gold": 0.4457110992469971, + "brier": 0.21999786383788478, + "ece": 0.023323136487035337, + "mean_confidence": 0.8447172036358691, + "cohort": "clean and knowable", + "excluded_non_headline_rows": 218, + "ece_bins": 10 + }, + "transfer_dev_tuned_blend": { + "n": 1046, + "correct": 909, + "accuracy": 0.8690248565965584, + "cross_entropy_from_gold": 0.3698933820845643, + "gold_entropy": 0.0, + "kl_from_gold": 0.3698933820845643, + "brier": 0.18873769003914254, + "ece": 0.020429096141296538, + "mean_confidence": 0.8663359693950322, + "cohort": "clean and knowable", + "excluded_non_headline_rows": 218, + "ece_bins": 10 + } + } + } + ] + }, + "transfer_test": { + "label": "Transfer-v9 test", + "role": "held_out_evaluation", + "records": 1264, + "questions": 1264, + "headline_n": 1046, + "runs": [ + { + "run": "base4", + "model": "Frozen-Qwen3.5-4B", + "family": "Qwen3.5-4B", + "weights": "frozen_base", + "native_readout": null, + "zero_shot": { + "choice_t1": { + "n": 1046, + "correct": 716, + "accuracy": 0.6845124282982792, + "cross_entropy_from_gold": 0.8349623273118204, + "gold_entropy": 0.0, + "kl_from_gold": 0.8349623273118204, + "brier": 0.42656058074391484, + "ece": 0.1303798195616654, + "mean_confidence": 0.8139445762365763, + "cohort": "clean and knowable", + "excluded_non_headline_rows": 218, + "ece_bins": 10 + } + }, + "transfer_dev_tuned": { + "transfer_dev_calibrated_choice": { + "n": 1046, + "correct": 716, + "accuracy": 0.6845124282982792, + "cross_entropy_from_gold": 0.7410651656609946, + "gold_entropy": 0.0, + "kl_from_gold": 0.7410651656609946, + "brier": 0.3904408201413376, + "ece": 0.04134673815393105, + "mean_confidence": 0.7051530768940606, + "cohort": "clean and knowable", + "excluded_non_headline_rows": 218, + "ece_bins": 10 + } + } + }, + { + "run": "pointer4", + "model": "JevAny-Qwen3.5-4B-Pointer", + "family": "Qwen3.5-4B", + "weights": "jevany_sft", + "native_readout": "pointer", + "zero_shot": { + "choice_t1": { + "n": 1046, + "correct": 770, + "accuracy": 0.7361376673040153, + "cross_entropy_from_gold": 0.6902431349522291, + "gold_entropy": 0.0, + "kl_from_gold": 0.6902431349522291, + "brier": 0.348668135528129, + "ece": 0.04630103938423926, + "mean_confidence": 0.701833119833475, + "cohort": "clean and knowable", + "excluded_non_headline_rows": 218, + "ece_bins": 10 + }, + "native_shipped": { + "n": 1046, + "correct": 815, + "accuracy": 0.7791586998087954, + "cross_entropy_from_gold": 0.5717563825431627, + "gold_entropy": 0.0, + "kl_from_gold": 0.5717563825431627, + "brier": 0.29765821801656817, + "ece": 0.04239090732927413, + "mean_confidence": 0.8129759944331685, + "cohort": "clean and knowable", + "excluded_non_headline_rows": 218, + "ece_bins": 10 + }, + "fixed_blend_0_5": { + "n": 1046, + "correct": 810, + "accuracy": 0.7743785850860421, + "cross_entropy_from_gold": 0.5666923610469762, + "gold_entropy": 0.0, + "kl_from_gold": 0.5666923610469762, + "brier": 0.2939336464727867, + "ece": 0.04874813932892946, + "mean_confidence": 0.7776535459474749, + "cohort": "clean and knowable", + "excluded_non_headline_rows": 218, + "ece_bins": 10 + } + }, + "transfer_dev_tuned": { + "transfer_dev_calibrated_choice": { + "n": 1046, + "correct": 770, + "accuracy": 0.7361376673040153, + "cross_entropy_from_gold": 0.6888466560394513, + "gold_entropy": 0.0, + "kl_from_gold": 0.6888466560394513, + "brier": 0.34649580012970715, + "ece": 0.03269551339782706, + "mean_confidence": 0.7287518287657895, + "cohort": "clean and knowable", + "excluded_non_headline_rows": 218, + "ece_bins": 10 + }, + "transfer_dev_tuned_blend": { + "n": 1046, + "correct": 816, + "accuracy": 0.780114722753346, + "cross_entropy_from_gold": 0.5578570106435864, + "gold_entropy": 0.0, + "kl_from_gold": 0.5578570106435864, + "brier": 0.2914736076482713, + "ece": 0.04496393628178446, + "mean_confidence": 0.7852072596268804, + "cohort": "clean and knowable", + "excluded_non_headline_rows": 218, + "ece_bins": 10 + } + } + }, + { + "run": "direct4", + "model": "JevAny-Qwen3.5-4B-Direct-Token", + "family": "Qwen3.5-4B", + "weights": "jevany_sft", + "native_readout": "direct-token", + "zero_shot": { + "choice_t1": { + "n": 1046, + "correct": 822, + "accuracy": 0.7858508604206501, + "cross_entropy_from_gold": 0.5697635579703885, + "gold_entropy": 0.0, + "kl_from_gold": 0.5697635579703885, + "brier": 0.29077517389394053, + "ece": 0.0470678026592797, + "mean_confidence": 0.8092660785307848, + "cohort": "clean and knowable", + "excluded_non_headline_rows": 218, + "ece_bins": 10 + }, + "native_shipped": { + "n": 1046, + "correct": 825, + "accuracy": 0.7887189292543021, + "cross_entropy_from_gold": 0.5304680016857037, + "gold_entropy": 0.0, + "kl_from_gold": 0.5304680016857037, + "brier": 0.2727158680694733, + "ece": 0.040601108067036665, + "mean_confidence": 0.8199335179349373, + "cohort": "clean and knowable", + "excluded_non_headline_rows": 218, + "ece_bins": 10 + }, + "fixed_blend_0_5": { + "n": 1046, + "correct": 831, + "accuracy": 0.7944550669216062, + "cross_entropy_from_gold": 0.5309958800935447, + "gold_entropy": 0.0, + "kl_from_gold": 0.5309958800935447, + "brier": 0.2725904993128285, + "ece": 0.038203239429849746, + "mean_confidence": 0.8132871591572012, + "cohort": "clean and knowable", + "excluded_non_headline_rows": 218, + "ece_bins": 10 + } + }, + "transfer_dev_tuned": { + "transfer_dev_calibrated_choice": { + "n": 1046, + "correct": 822, + "accuracy": 0.7858508604206501, + "cross_entropy_from_gold": 0.5628608839489921, + "gold_entropy": 0.0, + "kl_from_gold": 0.5628608839489921, + "brier": 0.28883932397403955, + "ece": 0.030719110557561026, + "mean_confidence": 0.7766618482775124, + "cohort": "clean and knowable", + "excluded_non_headline_rows": 218, + "ece_bins": 10 + }, + "transfer_dev_tuned_blend": { + "n": 1046, + "correct": 835, + "accuracy": 0.7982791586998088, + "cross_entropy_from_gold": 0.5262143707831856, + "gold_entropy": 0.0, + "kl_from_gold": 0.5262143707831856, + "brier": 0.26999002433749075, + "ece": 0.027780702772704318, + "mean_confidence": 0.7973056353797929, + "cohort": "clean and knowable", + "excluded_non_headline_rows": 218, + "ece_bins": 10 + } + } + }, + { + "run": "base27", + "model": "Frozen-Qwen3.8-27B", + "family": "Qwen3.8-27B", + "weights": "frozen_base", + "native_readout": null, + "zero_shot": { + "choice_t1": { + "n": 1046, + "correct": 844, + "accuracy": 0.8068833652007649, + "cross_entropy_from_gold": 0.6110678829313702, + "gold_entropy": 0.0, + "kl_from_gold": 0.6110678829313702, + "brier": 0.279721544152333, + "ece": 0.06893200869919826, + "mean_confidence": 0.8617564454102284, + "cohort": "clean and knowable", + "excluded_non_headline_rows": 218, + "ece_bins": 10 + } + }, + "transfer_dev_tuned": { + "transfer_dev_calibrated_choice": { + "n": 1046, + "correct": 844, + "accuracy": 0.8068833652007649, + "cross_entropy_from_gold": 0.5494987252614483, + "gold_entropy": 0.0, + "kl_from_gold": 0.5494987252614483, + "brier": 0.2709760690574176, + "ece": 0.027612650998887305, + "mean_confidence": 0.7874674967417389, + "cohort": "clean and knowable", + "excluded_non_headline_rows": 218, + "ece_bins": 10 + } + } + }, + { + "run": "pointer27", + "model": "JevAny-Qwen3.8-27B-Pointer", + "family": "Qwen3.8-27B", + "weights": "jevany_sft", + "native_readout": "pointer", + "zero_shot": { + "choice_t1": { + "n": 1046, + "correct": 900, + "accuracy": 0.8604206500956023, + "cross_entropy_from_gold": 0.47077064689010384, + "gold_entropy": 0.0, + "kl_from_gold": 0.47077064689010384, + "brier": 0.2114754899927417, + "ece": 0.11231046843501219, + "mean_confidence": 0.7539620937273849, + "cohort": "clean and knowable", + "excluded_non_headline_rows": 218, + "ece_bins": 10 + }, + "native_shipped": { + "n": 1046, + "correct": 920, + "accuracy": 0.8795411089866156, + "cross_entropy_from_gold": 0.3191183572882807, + "gold_entropy": 0.0, + "kl_from_gold": 0.3191183572882807, + "brier": 0.15983357286839844, + "ece": 0.02547180335097676, + "mean_confidence": 0.8916975436746409, + "cohort": "clean and knowable", + "excluded_non_headline_rows": 218, + "ece_bins": 10 + }, + "fixed_blend_0_5": { + "n": 1046, + "correct": 933, + "accuracy": 0.8919694072657743, + "cross_entropy_from_gold": 0.32150313724237106, + "gold_entropy": 0.0, + "kl_from_gold": 0.32150313724237106, + "brier": 0.1550705009584002, + "ece": 0.043199084741469725, + "mean_confidence": 0.8548688707373127, + "cohort": "clean and knowable", + "excluded_non_headline_rows": 218, + "ece_bins": 10 + } + }, + "transfer_dev_tuned": { + "transfer_dev_calibrated_choice": { + "n": 1046, + "correct": 900, + "accuracy": 0.8604206500956023, + "cross_entropy_from_gold": 0.40058733865111656, + "gold_entropy": 0.0, + "kl_from_gold": 0.40058733865111656, + "brier": 0.1880387608542386, + "ece": 0.02034085988544015, + "mean_confidence": 0.8604184603403277, + "cohort": "clean and knowable", + "excluded_non_headline_rows": 218, + "ece_bins": 10 + }, + "transfer_dev_tuned_blend": { + "n": 1046, + "correct": 932, + "accuracy": 0.8910133843212237, + "cross_entropy_from_gold": 0.3094830481103647, + "gold_entropy": 0.0, + "kl_from_gold": 0.3094830481103647, + "brier": 0.15280535574202198, + "ece": 0.034827378233759226, + "mean_confidence": 0.8807493062406941, + "cohort": "clean and knowable", + "excluded_non_headline_rows": 218, + "ece_bins": 10 + } + } + } + ] + }, + "typed_test": { + "label": "Typed Decisions test", + "role": "held_out_external_evaluation", + "records": 400, + "questions": 2000, + "headline_n": 2000, + "runs": [ + { + "run": "base4", + "model": "Frozen-Qwen3.5-4B", + "family": "Qwen3.5-4B", + "weights": "frozen_base", + "native_readout": null, + "zero_shot": { + "choice_t1": { + "n": 2000, + "correct": 1055, + "accuracy": 0.5275, + "cross_entropy_from_gold": 1.485304034737827, + "gold_entropy": 0.7673585556243288, + "kl_from_gold": 0.7179454791134983, + "brier": 0.32209870787613476, + "ece": 0.21683605197556294, + "mean_confidence": 0.7434521192146655, + "cohort": "clean and knowable", + "excluded_non_headline_rows": 0, + "ece_bins": 15 + } + }, + "transfer_dev_tuned": { + "transfer_dev_calibrated_choice": { + "n": 2000, + "correct": 1055, + "accuracy": 0.5275, + "cross_entropy_from_gold": 1.160641706252667, + "gold_entropy": 0.7673585556243288, + "kl_from_gold": 0.3932831506283383, + "brier": 0.22234376212257462, + "ece": 0.0900516370077249, + "mean_confidence": 0.6134237235271172, + "cohort": "clean and knowable", + "excluded_non_headline_rows": 0, + "ece_bins": 15 + } + } + }, + { + "run": "pointer4", + "model": "JevAny-Qwen3.5-4B-Pointer", + "family": "Qwen3.5-4B", + "weights": "jevany_sft", + "native_readout": "pointer", + "zero_shot": { + "choice_t1": { + "n": 2000, + "correct": 1156, + "accuracy": 0.578, + "cross_entropy_from_gold": 1.1689837218988208, + "gold_entropy": 0.7673585556243288, + "kl_from_gold": 0.4016251662744921, + "brier": 0.21958081006971433, + "ece": 0.11804038821625913, + "mean_confidence": 0.6283756741526655, + "cohort": "clean and knowable", + "excluded_non_headline_rows": 0, + "ece_bins": 15 + }, + "native_shipped": { + "n": 2000, + "correct": 1270, + "accuracy": 0.635, + "cross_entropy_from_gold": 1.2461303849511813, + "gold_entropy": 0.7673585556243288, + "kl_from_gold": 0.4787718293268526, + "brier": 0.21769431550066473, + "ece": 0.11542953593789118, + "mean_confidence": 0.7131901462821837, + "cohort": "clean and knowable", + "excluded_non_headline_rows": 0, + "ece_bins": 15 + }, + "fixed_blend_0_5": { + "n": 2000, + "correct": 1214, + "accuracy": 0.607, + "cross_entropy_from_gold": 1.1474089185651393, + "gold_entropy": 0.7673585556243288, + "kl_from_gold": 0.3800503629408105, + "brier": 0.20385335585187295, + "ece": 0.08838893066306003, + "mean_confidence": 0.6710193141005254, + "cohort": "clean and knowable", + "excluded_non_headline_rows": 0, + "ece_bins": 15 + } + }, + "transfer_dev_tuned": { + "transfer_dev_calibrated_choice": { + "n": 2000, + "correct": 1156, + "accuracy": 0.578, + "cross_entropy_from_gold": 1.2283795069938794, + "gold_entropy": 0.7673585556243288, + "kl_from_gold": 0.4610209513695507, + "brier": 0.23872296531625706, + "ece": 0.13219447566436857, + "mean_confidence": 0.6614429063056575, + "cohort": "clean and knowable", + "excluded_non_headline_rows": 0, + "ece_bins": 15 + }, + "transfer_dev_tuned_blend": { + "n": 2000, + "correct": 1258, + "accuracy": 0.629, + "cross_entropy_from_gold": 1.1499057537598978, + "gold_entropy": 0.7673585556243288, + "kl_from_gold": 0.38254719813556903, + "brier": 0.19860375584665835, + "ece": 0.09685551065650508, + "mean_confidence": 0.6767929224204166, + "cohort": "clean and knowable", + "excluded_non_headline_rows": 0, + "ece_bins": 15 + } + } + }, + { + "run": "direct4", + "model": "JevAny-Qwen3.5-4B-Direct-Token", + "family": "Qwen3.5-4B", + "weights": "jevany_sft", + "native_readout": "direct-token", + "zero_shot": { + "choice_t1": { + "n": 2000, + "correct": 1296, + "accuracy": 0.648, + "cross_entropy_from_gold": 1.157150279237127, + "gold_entropy": 0.7673585556243288, + "kl_from_gold": 0.38979172361279835, + "brier": 0.18515075572006146, + "ece": 0.07029919279872957, + "mean_confidence": 0.6765339723262027, + "cohort": "clean and knowable", + "excluded_non_headline_rows": 0, + "ece_bins": 15 + }, + "native_shipped": { + "n": 2000, + "correct": 1344, + "accuracy": 0.672, + "cross_entropy_from_gold": 1.2022614330547874, + "gold_entropy": 0.7673585556243288, + "kl_from_gold": 0.43490287743045863, + "brier": 0.17877494146152656, + "ece": 0.08278732481632818, + "mean_confidence": 0.7379338187164135, + "cohort": "clean and knowable", + "excluded_non_headline_rows": 0, + "ece_bins": 15 + }, + "fixed_blend_0_5": { + "n": 2000, + "correct": 1335, + "accuracy": 0.6675, + "cross_entropy_from_gold": 1.1350838501828007, + "gold_entropy": 0.7673585556243288, + "kl_from_gold": 0.3677252945584719, + "brier": 0.1702376402804755, + "ece": 0.06198979447604478, + "mean_confidence": 0.7066262108749651, + "cohort": "clean and knowable", + "excluded_non_headline_rows": 0, + "ece_bins": 15 + } + }, + "transfer_dev_tuned": { + "transfer_dev_calibrated_choice": { + "n": 2000, + "correct": 1296, + "accuracy": 0.648, + "cross_entropy_from_gold": 1.083382488931363, + "gold_entropy": 0.7673585556243288, + "kl_from_gold": 0.31602393330703427, + "brier": 0.16593211527952617, + "ece": 0.07601165366691312, + "mean_confidence": 0.6328419411033379, + "cohort": "clean and knowable", + "excluded_non_headline_rows": 0, + "ece_bins": 15 + }, + "transfer_dev_tuned_blend": { + "n": 2000, + "correct": 1353, + "accuracy": 0.6765, + "cross_entropy_from_gold": 1.0910947225952166, + "gold_entropy": 0.7673585556243288, + "kl_from_gold": 0.3237361669708878, + "brier": 0.15773140989073764, + "ece": 0.06194953110709811, + "mean_confidence": 0.6893724371900853, + "cohort": "clean and knowable", + "excluded_non_headline_rows": 0, + "ece_bins": 15 + } + } + }, + { + "run": "base27", + "model": "Frozen-Qwen3.8-27B", + "family": "Qwen3.8-27B", + "weights": "frozen_base", + "native_readout": null, + "zero_shot": { + "choice_t1": { + "n": 2000, + "correct": 1365, + "accuracy": 0.6825, + "cross_entropy_from_gold": 1.4486833046492098, + "gold_entropy": 0.7673585556243288, + "kl_from_gold": 0.681324749024881, + "brier": 0.23332950942149694, + "ece": 0.12575488372079707, + "mean_confidence": 0.8052338907255023, + "cohort": "clean and knowable", + "excluded_non_headline_rows": 0, + "ece_bins": 15 + } + }, + "transfer_dev_tuned": { + "transfer_dev_calibrated_choice": { + "n": 2000, + "correct": 1365, + "accuracy": 0.6825, + "cross_entropy_from_gold": 1.0976128315315092, + "gold_entropy": 0.7673585556243288, + "kl_from_gold": 0.3302542759071805, + "brier": 0.16553442637117138, + "ece": 0.07868470788264477, + "mean_confidence": 0.7000954995537738, + "cohort": "clean and knowable", + "excluded_non_headline_rows": 0, + "ece_bins": 15 + } + } + }, + { + "run": "pointer27", + "model": "JevAny-Qwen3.8-27B-Pointer", + "family": "Qwen3.8-27B", + "weights": "jevany_sft", + "native_readout": "pointer", + "zero_shot": { + "choice_t1": { + "n": 2000, + "correct": 1452, + "accuracy": 0.726, + "cross_entropy_from_gold": 0.9201029180583041, + "gold_entropy": 0.7673585556243288, + "kl_from_gold": 0.15274436243397538, + "brier": 0.0785204257840767, + "ece": 0.10839391090387629, + "mean_confidence": 0.6176060890961234, + "cohort": "clean and knowable", + "excluded_non_headline_rows": 0, + "ece_bins": 15 + }, + "native_shipped": { + "n": 2000, + "correct": 1456, + "accuracy": 0.728, + "cross_entropy_from_gold": 1.0602046256660616, + "gold_entropy": 0.7673585556243288, + "kl_from_gold": 0.2928460700417328, + "brier": 0.13141516555965035, + "ece": 0.05304485014212312, + "mean_confidence": 0.7527105045858369, + "cohort": "clean and knowable", + "excluded_non_headline_rows": 0, + "ece_bins": 15 + }, + "fixed_blend_0_5": { + "n": 2000, + "correct": 1464, + "accuracy": 0.732, + "cross_entropy_from_gold": 0.9466132871046253, + "gold_entropy": 0.7673585556243288, + "kl_from_gold": 0.17925473148029658, + "brier": 0.09467929687142354, + "ece": 0.034120016906592165, + "mean_confidence": 0.6998051537601903, + "cohort": "clean and knowable", + "excluded_non_headline_rows": 0, + "ece_bins": 15 + } + }, + "transfer_dev_tuned": { + "transfer_dev_calibrated_choice": { + "n": 2000, + "correct": 1452, + "accuracy": 0.726, + "cross_entropy_from_gold": 1.0152256258087546, + "gold_entropy": 0.7673585556243288, + "kl_from_gold": 0.2478670701844259, + "brier": 0.11614111181065959, + "ece": 0.020460563111670046, + "mean_confidence": 0.7331494836433007, + "cohort": "clean and knowable", + "excluded_non_headline_rows": 0, + "ece_bins": 15 + }, + "transfer_dev_tuned_blend": { + "n": 2000, + "correct": 1466, + "accuracy": 0.733, + "cross_entropy_from_gold": 1.004976949424594, + "gold_entropy": 0.7673585556243288, + "kl_from_gold": 0.23761839380026517, + "brier": 0.11476036996840079, + "ece": 0.02738445525330449, + "mean_confidence": 0.740111940631487, + "cohort": "clean and knowable", + "excluded_non_headline_rows": 0, + "ece_bins": 15 + } + } + } + ], + "external_baselines": [ + { + "model": "meraGPT Decider 1", + "kind": "published_only", + "status": "dataset-card result; not locally rerun", + "accuracy": 0.768 + }, + { + "model": "Jev 1.13 (OpenRouter)", + "canonical_model": "TypeSafe Jev 1.13.0", + "kind": "external_api", + "result_source": "published_only", + "status": "dataset-card result; not locally rerun", + "accuracy": 0.727 + } + ] + }, + "jevbench_public": { + "label": "JevBench public development", + "role": "public_diagnostic_not_used_for_selection", + "records": 231, + "questions": 231, + "headline_n": 231, + "runs": [ + { + "run": "base4", + "model": "Frozen-Qwen3.5-4B", + "family": "Qwen3.5-4B", + "weights": "frozen_base", + "native_readout": null, + "zero_shot": { + "choice_t1": { + "n": 231, + "correct": 184, + "accuracy": 0.7965367965367965, + "cross_entropy_from_gold": 0.46503985941725867, + "gold_entropy": 0.0, + "kl_from_gold": 0.46503985941725867, + "brier": 0.26023857002415907, + "ece": 0.042706500279636225, + "mean_confidence": 0.8277960488256922, + "cohort": "clean and knowable", + "excluded_non_headline_rows": 0, + "ece_bins": 10 + } + }, + "transfer_dev_tuned": { + "transfer_dev_calibrated_choice": { + "n": 231, + "correct": 184, + "accuracy": 0.7965367965367965, + "cross_entropy_from_gold": 0.5034935201937853, + "gold_entropy": 0.0, + "kl_from_gold": 0.5034935201937853, + "brier": 0.2731813432864743, + "ece": 0.08846006305082621, + "mean_confidence": 0.7245256388535549, + "cohort": "clean and knowable", + "excluded_non_headline_rows": 0, + "ece_bins": 10 + } + } + }, + { + "run": "pointer4", + "model": "JevAny-Qwen3.5-4B-Pointer", + "family": "Qwen3.5-4B", + "weights": "jevany_sft", + "native_readout": "pointer", + "zero_shot": { + "choice_t1": { + "n": 231, + "correct": 188, + "accuracy": 0.8138528138528138, + "cross_entropy_from_gold": 0.47901846432733636, + "gold_entropy": 0.0, + "kl_from_gold": 0.47901846432733636, + "brier": 0.2561308970527479, + "ece": 0.056124141655517386, + "mean_confidence": 0.7692551472534949, + "cohort": "clean and knowable", + "excluded_non_headline_rows": 0, + "ece_bins": 10 + }, + "native_shipped": { + "n": 231, + "correct": 185, + "accuracy": 0.8008658008658008, + "cross_entropy_from_gold": 0.4554565575008672, + "gold_entropy": 0.0, + "kl_from_gold": 0.4554565575008672, + "brier": 0.2585838755503942, + "ece": 0.03718147002370739, + "mean_confidence": 0.8308520442419092, + "cohort": "clean and knowable", + "excluded_non_headline_rows": 0, + "ece_bins": 10 + }, + "fixed_blend_0_5": { + "n": 231, + "correct": 188, + "accuracy": 0.8138528138528138, + "cross_entropy_from_gold": 0.4142501461307885, + "gold_entropy": 0.0, + "kl_from_gold": 0.4142501461307885, + "brier": 0.23413224994049325, + "ece": 0.02272572239740828, + "mean_confidence": 0.807826264488957, + "cohort": "clean and knowable", + "excluded_non_headline_rows": 0, + "ece_bins": 10 + } + }, + "transfer_dev_tuned": { + "transfer_dev_calibrated_choice": { + "n": 231, + "correct": 188, + "accuracy": 0.8138528138528138, + "cross_entropy_from_gold": 0.46979954158584164, + "gold_entropy": 0.0, + "kl_from_gold": 0.46979954158584164, + "brier": 0.2515032873679713, + "ece": 0.03081010600115095, + "mean_confidence": 0.7956829590845715, + "cohort": "clean and knowable", + "excluded_non_headline_rows": 0, + "ece_bins": 10 + }, + "transfer_dev_tuned_blend": { + "n": 231, + "correct": 187, + "accuracy": 0.8095238095238095, + "cross_entropy_from_gold": 0.4300189742220573, + "gold_entropy": 0.0, + "kl_from_gold": 0.4300189742220573, + "brier": 0.24497083465715366, + "ece": 0.03200768487813679, + "mean_confidence": 0.8105343163775357, + "cohort": "clean and knowable", + "excluded_non_headline_rows": 0, + "ece_bins": 10 + } + } + }, + { + "run": "direct4", + "model": "JevAny-Qwen3.5-4B-Direct-Token", + "family": "Qwen3.5-4B", + "weights": "jevany_sft", + "native_readout": "direct-token", + "zero_shot": { + "choice_t1": { + "n": 231, + "correct": 188, + "accuracy": 0.8138528138528138, + "cross_entropy_from_gold": 0.4119059642596153, + "gold_entropy": 0.0, + "kl_from_gold": 0.4119059642596153, + "brier": 0.23630630313186377, + "ece": 0.06493498624236703, + "mean_confidence": 0.8264788235127298, + "cohort": "clean and knowable", + "excluded_non_headline_rows": 0, + "ece_bins": 10 + }, + "native_shipped": { + "n": 231, + "correct": 187, + "accuracy": 0.8095238095238095, + "cross_entropy_from_gold": 0.43291971264790263, + "gold_entropy": 0.0, + "kl_from_gold": 0.43291971264790263, + "brier": 0.25567839211017723, + "ece": 0.05126707211644283, + "mean_confidence": 0.8459940130356174, + "cohort": "clean and knowable", + "excluded_non_headline_rows": 0, + "ece_bins": 10 + }, + "fixed_blend_0_5": { + "n": 231, + "correct": 187, + "accuracy": 0.8095238095238095, + "cross_entropy_from_gold": 0.4013583529376732, + "gold_entropy": 0.0, + "kl_from_gold": 0.4013583529376732, + "brier": 0.23610741709317634, + "ece": 0.04227451187274651, + "mean_confidence": 0.8360951760741018, + "cohort": "clean and knowable", + "excluded_non_headline_rows": 0, + "ece_bins": 10 + } + }, + "transfer_dev_tuned": { + "transfer_dev_calibrated_choice": { + "n": 231, + "correct": 188, + "accuracy": 0.8138528138528138, + "cross_entropy_from_gold": 0.419037743510702, + "gold_entropy": 0.0, + "kl_from_gold": 0.419037743510702, + "brier": 0.2361960585013741, + "ece": 0.04983130926953957, + "mean_confidence": 0.7972245036382399, + "cohort": "clean and knowable", + "excluded_non_headline_rows": 0, + "ece_bins": 10 + }, + "transfer_dev_tuned_blend": { + "n": 231, + "correct": 187, + "accuracy": 0.8095238095238095, + "cross_entropy_from_gold": 0.40445010913159385, + "gold_entropy": 0.0, + "kl_from_gold": 0.40445010913159385, + "brier": 0.23725278784385523, + "ece": 0.04395161391787564, + "mean_confidence": 0.8242603879406017, + "cohort": "clean and knowable", + "excluded_non_headline_rows": 0, + "ece_bins": 10 + } + } + }, + { + "run": "base27", + "model": "Frozen-Qwen3.8-27B", + "family": "Qwen3.8-27B", + "weights": "frozen_base", + "native_readout": null, + "zero_shot": { + "choice_t1": { + "n": 231, + "correct": 203, + "accuracy": 0.8787878787878788, + "cross_entropy_from_gold": 0.2864628223866121, + "gold_entropy": 0.0, + "kl_from_gold": 0.2864628223866121, + "brier": 0.1604073338283613, + "ece": 0.022674652746122903, + "mean_confidence": 0.8936596611320321, + "cohort": "clean and knowable", + "excluded_non_headline_rows": 0, + "ece_bins": 10 + } + }, + "transfer_dev_tuned": { + "transfer_dev_calibrated_choice": { + "n": 231, + "correct": 203, + "accuracy": 0.8787878787878788, + "cross_entropy_from_gold": 0.32118852252593943, + "gold_entropy": 0.0, + "kl_from_gold": 0.32118852252593943, + "brier": 0.16739722641469756, + "ece": 0.09116510874430303, + "mean_confidence": 0.821229989485614, + "cohort": "clean and knowable", + "excluded_non_headline_rows": 0, + "ece_bins": 10 + } + } + }, + { + "run": "pointer27", + "model": "JevAny-Qwen3.8-27B-Pointer", + "family": "Qwen3.8-27B", + "weights": "jevany_sft", + "native_readout": "pointer", + "zero_shot": { + "choice_t1": { + "n": 231, + "correct": 207, + "accuracy": 0.8961038961038961, + "cross_entropy_from_gold": 0.34460785579024505, + "gold_entropy": 0.0, + "kl_from_gold": 0.34460785579024505, + "brier": 0.16729832168502917, + "ece": 0.12307461523799648, + "mean_confidence": 0.7818628405160418, + "cohort": "clean and knowable", + "excluded_non_headline_rows": 0, + "ece_bins": 10 + }, + "native_shipped": { + "n": 231, + "correct": 208, + "accuracy": 0.9004329004329005, + "cross_entropy_from_gold": 0.2650469579600223, + "gold_entropy": 0.0, + "kl_from_gold": 0.2650469579600223, + "brier": 0.14556199524820898, + "ece": 0.02122154456035358, + "mean_confidence": 0.884113426683873, + "cohort": "clean and knowable", + "excluded_non_headline_rows": 0, + "ece_bins": 10 + }, + "fixed_blend_0_5": { + "n": 231, + "correct": 208, + "accuracy": 0.9004329004329005, + "cross_entropy_from_gold": 0.26310190901965735, + "gold_entropy": 0.0, + "kl_from_gold": 0.26310190901965735, + "brier": 0.1406077982333171, + "ece": 0.05751149881179743, + "mean_confidence": 0.8560666856708948, + "cohort": "clean and knowable", + "excluded_non_headline_rows": 0, + "ece_bins": 10 + } + }, + "transfer_dev_tuned": { + "transfer_dev_calibrated_choice": { + "n": 231, + "correct": 207, + "accuracy": 0.8961038961038961, + "cross_entropy_from_gold": 0.26532298607460764, + "gold_entropy": 0.0, + "kl_from_gold": 0.26532298607460764, + "brier": 0.14331108028016937, + "ece": 0.042115349045515435, + "mean_confidence": 0.8714557776895534, + "cohort": "clean and knowable", + "excluded_non_headline_rows": 0, + "ece_bins": 10 + }, + "transfer_dev_tuned_blend": { + "n": 231, + "correct": 208, + "accuracy": 0.9004329004329005, + "cross_entropy_from_gold": 0.24856632495280404, + "gold_entropy": 0.0, + "kl_from_gold": 0.24856632495280404, + "brier": 0.13696196671048083, + "ece": 0.05151408138171153, + "mean_confidence": 0.8781699403131888, + "cohort": "clean and knowable", + "excluded_non_headline_rows": 0, + "ece_bins": 10 + } + } + } + ] + }, + "jevjudge_text": { + "label": "JevJudge text", + "role": "held_out_external_evaluation", + "records": 724, + "questions": 724, + "headline_n": 724, + "runs": [ + { + "run": "base4", + "model": "Frozen-Qwen3.5-4B", + "family": "Qwen3.5-4B", + "weights": "frozen_base", + "native_readout": null, + "zero_shot": { + "choice_t1": { + "n": 724, + "correct": 425, + "accuracy": 0.5870165745856354, + "cross_entropy_from_gold": 0.9990413838295111, + "gold_entropy": 0.0, + "kl_from_gold": 0.9990413838295111, + "brier": 0.5424665612791818, + "ece": 0.124790188004864, + "mean_confidence": 0.7118067625904997, + "cohort": "clean and knowable", + "excluded_non_headline_rows": 0, + "ece_bins": 10 + } + }, + "transfer_dev_tuned": { + "transfer_dev_calibrated_choice": { + "n": 724, + "correct": 425, + "accuracy": 0.5870165745856354, + "cross_entropy_from_gold": 0.9193370649594155, + "gold_entropy": 0.0, + "kl_from_gold": 0.9193370649594155, + "brier": 0.5121791926918943, + "ece": 0.030915955188617703, + "mean_confidence": 0.5929232826389871, + "cohort": "clean and knowable", + "excluded_non_headline_rows": 0, + "ece_bins": 10 + } + } + }, + { + "run": "pointer4", + "model": "JevAny-Qwen3.5-4B-Pointer", + "family": "Qwen3.5-4B", + "weights": "jevany_sft", + "native_readout": "pointer", + "zero_shot": { + "choice_t1": { + "n": 724, + "correct": 381, + "accuracy": 0.5262430939226519, + "cross_entropy_from_gold": 1.027944601965405, + "gold_entropy": 0.0, + "kl_from_gold": 1.027944601965405, + "brier": 0.5701372991529206, + "ece": 0.05964150883024191, + "mean_confidence": 0.5692400844020692, + "cohort": "clean and knowable", + "excluded_non_headline_rows": 0, + "ece_bins": 10 + }, + "native_shipped": { + "n": 724, + "correct": 420, + "accuracy": 0.580110497237569, + "cross_entropy_from_gold": 0.9966502097264682, + "gold_entropy": 0.0, + "kl_from_gold": 0.9966502097264682, + "brier": 0.5538821469971053, + "ece": 0.07444449081079427, + "mean_confidence": 0.621924240558197, + "cohort": "clean and knowable", + "excluded_non_headline_rows": 0, + "ece_bins": 10 + }, + "fixed_blend_0_5": { + "n": 724, + "correct": 428, + "accuracy": 0.5911602209944752, + "cross_entropy_from_gold": 0.9621095236741931, + "gold_entropy": 0.0, + "kl_from_gold": 0.9621095236741931, + "brier": 0.5413174973396784, + "ece": 0.06904408118946087, + "mean_confidence": 0.5874577811451718, + "cohort": "clean and knowable", + "excluded_non_headline_rows": 0, + "ece_bins": 10 + } + }, + "transfer_dev_tuned": { + "transfer_dev_calibrated_choice": { + "n": 724, + "correct": 381, + "accuracy": 0.5262430939226519, + "cross_entropy_from_gold": 1.048869475648102, + "gold_entropy": 0.0, + "kl_from_gold": 1.048869475648102, + "brier": 0.5765749498377344, + "ece": 0.0728368732066297, + "mean_confidence": 0.5930575646329146, + "cohort": "clean and knowable", + "excluded_non_headline_rows": 0, + "ece_bins": 10 + }, + "transfer_dev_tuned_blend": { + "n": 724, + "correct": 431, + "accuracy": 0.5953038674033149, + "cross_entropy_from_gold": 0.9665595951901557, + "gold_entropy": 0.0, + "kl_from_gold": 0.9665595951901557, + "brier": 0.5427219958129672, + "ece": 0.0883382849075786, + "mean_confidence": 0.5909347225687963, + "cohort": "clean and knowable", + "excluded_non_headline_rows": 0, + "ece_bins": 10 + } + } + }, + { + "run": "direct4", + "model": "JevAny-Qwen3.5-4B-Direct-Token", + "family": "Qwen3.5-4B", + "weights": "jevany_sft", + "native_readout": "direct-token", + "zero_shot": { + "choice_t1": { + "n": 724, + "correct": 417, + "accuracy": 0.5759668508287292, + "cross_entropy_from_gold": 0.9595839832215673, + "gold_entropy": 0.0, + "kl_from_gold": 0.9595839832215673, + "brier": 0.5411152452677374, + "ece": 0.10266762314870827, + "mean_confidence": 0.6575552343312288, + "cohort": "clean and knowable", + "excluded_non_headline_rows": 0, + "ece_bins": 10 + }, + "native_shipped": { + "n": 724, + "correct": 423, + "accuracy": 0.5842541436464088, + "cross_entropy_from_gold": 0.9150453634780736, + "gold_entropy": 0.0, + "kl_from_gold": 0.9150453634780736, + "brier": 0.5160204686928853, + "ece": 0.09170853327092285, + "mean_confidence": 0.6676634582809498, + "cohort": "clean and knowable", + "excluded_non_headline_rows": 0, + "ece_bins": 10 + }, + "fixed_blend_0_5": { + "n": 724, + "correct": 427, + "accuracy": 0.5897790055248618, + "cross_entropy_from_gold": 0.9182029376845215, + "gold_entropy": 0.0, + "kl_from_gold": 0.9182029376845215, + "brier": 0.5196712705778425, + "ece": 0.07987836608665412, + "mean_confidence": 0.6592116906583616, + "cohort": "clean and knowable", + "excluded_non_headline_rows": 0, + "ece_bins": 10 + } + }, + "transfer_dev_tuned": { + "transfer_dev_calibrated_choice": { + "n": 724, + "correct": 417, + "accuracy": 0.5759668508287292, + "cross_entropy_from_gold": 0.9368946832422624, + "gold_entropy": 0.0, + "kl_from_gold": 0.9368946832422624, + "brier": 0.5297156601921297, + "ece": 0.06832163591814601, + "mean_confidence": 0.6163866223639445, + "cohort": "clean and knowable", + "excluded_non_headline_rows": 0, + "ece_bins": 10 + }, + "transfer_dev_tuned_blend": { + "n": 724, + "correct": 429, + "accuracy": 0.5925414364640884, + "cross_entropy_from_gold": 0.905162432888216, + "gold_entropy": 0.0, + "kl_from_gold": 0.905162432888216, + "brier": 0.5125121544257433, + "ece": 0.06810038720073028, + "mean_confidence": 0.6374560089027138, + "cohort": "clean and knowable", + "excluded_non_headline_rows": 0, + "ece_bins": 10 + } + } + }, + { + "run": "base27", + "model": "Frozen-Qwen3.8-27B", + "family": "Qwen3.8-27B", + "weights": "frozen_base", + "native_readout": null, + "zero_shot": { + "choice_t1": { + "n": 724, + "correct": 447, + "accuracy": 0.6174033149171271, + "cross_entropy_from_gold": 1.0435471379379775, + "gold_entropy": 0.0, + "kl_from_gold": 1.0435471379379775, + "brier": 0.5477779993257101, + "ece": 0.18642668688209224, + "mean_confidence": 0.7965262996294791, + "cohort": "clean and knowable", + "excluded_non_headline_rows": 0, + "ece_bins": 10 + } + }, + "transfer_dev_tuned": { + "transfer_dev_calibrated_choice": { + "n": 724, + "correct": 447, + "accuracy": 0.6174033149171271, + "cross_entropy_from_gold": 0.8776519496796663, + "gold_entropy": 0.0, + "kl_from_gold": 0.8776519496796663, + "brier": 0.4981210685244041, + "ece": 0.08864165712545208, + "mean_confidence": 0.6911587312881367, + "cohort": "clean and knowable", + "excluded_non_headline_rows": 0, + "ece_bins": 10 + } + } + }, + { + "run": "pointer27", + "model": "JevAny-Qwen3.8-27B-Pointer", + "family": "Qwen3.8-27B", + "weights": "jevany_sft", + "native_readout": "pointer", + "zero_shot": { + "choice_t1": { + "n": 724, + "correct": 419, + "accuracy": 0.5787292817679558, + "cross_entropy_from_gold": 0.9221648550274847, + "gold_entropy": 0.0, + "kl_from_gold": 0.9221648550274847, + "brier": 0.5384293459629866, + "ece": 0.12300846658221797, + "mean_confidence": 0.654141167166119, + "cohort": "clean and knowable", + "excluded_non_headline_rows": 0, + "ece_bins": 10 + }, + "native_shipped": { + "n": 724, + "correct": 481, + "accuracy": 0.664364640883978, + "cross_entropy_from_gold": 0.8240838957036632, + "gold_entropy": 0.0, + "kl_from_gold": 0.8240838957036632, + "brier": 0.4761621807919303, + "ece": 0.10150582984409334, + "mean_confidence": 0.7093057903641973, + "cohort": "clean and knowable", + "excluded_non_headline_rows": 0, + "ece_bins": 10 + }, + "fixed_blend_0_5": { + "n": 724, + "correct": 465, + "accuracy": 0.6422651933701657, + "cross_entropy_from_gold": 0.8340868983096242, + "gold_entropy": 0.0, + "kl_from_gold": 0.8340868983096242, + "brier": 0.48649946917918646, + "ece": 0.09293531151742315, + "mean_confidence": 0.6762654332459023, + "cohort": "clean and knowable", + "excluded_non_headline_rows": 0, + "ece_bins": 10 + } + }, + "transfer_dev_tuned": { + "transfer_dev_calibrated_choice": { + "n": 724, + "correct": 419, + "accuracy": 0.5787292817679558, + "cross_entropy_from_gold": 1.0389937594613554, + "gold_entropy": 0.0, + "kl_from_gold": 1.0389937594613554, + "brier": 0.5859846798375742, + "ece": 0.18095024427280532, + "mean_confidence": 0.7592297383784663, + "cohort": "clean and knowable", + "excluded_non_headline_rows": 0, + "ece_bins": 10 + }, + "transfer_dev_tuned_blend": { + "n": 724, + "correct": 466, + "accuracy": 0.643646408839779, + "cross_entropy_from_gold": 0.8472248615476166, + "gold_entropy": 0.0, + "kl_from_gold": 0.8472248615476166, + "brier": 0.49231968730390446, + "ece": 0.10423989831022704, + "mean_confidence": 0.7135438249208921, + "cohort": "clean and knowable", + "excluded_non_headline_rows": 0, + "ece_bins": 10 + } + } + } + ], + "external_baselines": [ + { + "model": "Jev 1.13 (OpenRouter)", + "canonical_model": "TypeSafe Jev 1.13.0", + "kind": "external_api", + "result_source": "api_rerun", + "status": "OpenRouter full-context rerun", + "accuracy": 0.6505524861878453, + "nll": 0.836, + "brier": 0.461, + "ece": 0.076, + "answered": 724, + "requested": 724 + }, + { + "model": "Kev-27B", + "kind": "open_kev", + "status": "local full-context 4,096-token chunked-KV rerun", + "accuracy": 0.6422651933701657, + "nll": 0.8774529517226382, + "brier": 0.48251125757984015, + "ece": 0.07166661519467314, + "answered": 724, + "requested": 724, + "repository": "jaredpalmer/kev-27b", + "revision": "01b81998019be550f0ae858727df49bac9511195" + } + ] + }, + "jevjudge_full": { + "label": "JevJudge full", + "records": 3220, + "status": "unsupported", + "accuracy": null, + "reason": "Training-free choice-token readout is text-only; image/video records are not stripped or relabeled as a full-suite result.", + "native_checkpoint_results": "results/external-zero-shot-v1.json#jevjudge_full" + } + }, + "legacy_results": { + "status": "retained unchanged", + "artifact": "results/letter-readout-v1.json" + } +} diff --git a/scripts/build_choice_readout_results.py b/scripts/build_choice_readout_results.py new file mode 100644 index 0000000..0a37b3c --- /dev/null +++ b/scripts/build_choice_readout_results.py @@ -0,0 +1,985 @@ +#!/usr/bin/env python3 +"""Validate, summarize, and plot the five-run choice-readout matrix. + +The input directory must contain ``base4``, ``pointer4``, ``direct4``, +``base27``, and ``pointer27`` subdirectories. Each subdirectory contains the +``manifest.json`` emitted by ``evaluate_choice_readout_matrix.py`` and the +``ensemble.json`` emitted by ``fit_choice_ensemble.py``. + +The builder fails closed on incomplete panels, component/count drift, or +cross-run dataset-hash drift. It never substitutes partial coverage or a +media-stripped score for JevJudge full. +""" + +from __future__ import annotations + +import argparse +import copy +import hashlib +import json +import math +import re +from pathlib import Path +from typing import Any + + +ROOT = Path(__file__).resolve().parents[1] +DEFAULT_EXTERNAL = ROOT / "results/external-zero-shot-v1.json" +DEFAULT_MODEL_CATALOG = ROOT / "results/model-family-v2.json" +DEFAULT_JSON_OUT = ROOT / "results/choice-readout-v2.json" +DEFAULT_SVG_OUT = ROOT / "docs/choice-readout-results.svg" + +INK = "#213248" +MUTED = "#64748B" +RULE = "#E4E9EF" +CHOICE = "#278577" +NATIVE = "#8493A6" +TUNED = "#165F55" +EXTERNAL = "#A17BB7" +TYPESAFE = "#C08A42" + +RUN_SPECS = ( + { + "key": "base4", + "family": "Qwen3.5-4B", + "public_repository": "Qwen/Qwen3.5-4B", + "weights": "frozen_base", + "native_readout": None, + "native_decision_mode": None, + }, + { + "key": "pointer4", + "family": "Qwen3.5-4B", + "public_repository": "SimpleJev/JevAny-Qwen3.5-4B-LoRA", + "weights": "jevany_sft", + "native_readout": "pointer", + "native_decision_mode": "pointer", + }, + { + "key": "direct4", + "family": "Qwen3.5-4B", + "public_repository": "SimpleJev/JevAny-Qwen3.5-4B-Direct-Token-LoRA", + "weights": "jevany_sft", + "native_readout": "direct-token", + "native_decision_mode": "lm_token", + }, + { + "key": "base27", + "family": "Qwen3.8-27B", + "public_repository": "Qwen/Qwen3.8-27B", + "weights": "frozen_base", + "native_readout": None, + "native_decision_mode": None, + }, + { + "key": "pointer27", + "family": "Qwen3.8-27B", + "public_repository": "SimpleJev/JevAny-Qwen3.8-27B-LoRA", + "weights": "jevany_sft", + "native_readout": "pointer", + "native_decision_mode": "pointer", + }, +) + +DATASETS = { + "transfer_calibration": { + "label": "Transfer-v9 development", + "role": "selection_only", + "records": 1_264, + "questions": 1_264, + "headline_n": 1_046, + }, + "transfer_test": { + "label": "Transfer-v9 test", + "role": "held_out_evaluation", + "records": 1_264, + "questions": 1_264, + "headline_n": 1_046, + }, + "typed_test": { + "label": "Typed Decisions test", + "role": "held_out_external_evaluation", + "records": 400, + "questions": 2_000, + "headline_n": 2_000, + }, + "jevbench_public": { + "label": "JevBench public development", + "role": "public_diagnostic_not_used_for_selection", + "records": 231, + "questions": 231, + "headline_n": 231, + }, + "jevjudge_text": { + "label": "JevJudge text", + "role": "held_out_external_evaluation", + "records": 724, + "questions": 724, + "headline_n": 724, + }, +} + +ZERO_SHOT_CONFIGS = ( + "choice_t1", + "native_shipped", + "fixed_blend_0_5", +) +TUNED_CONFIGS = ( + "transfer_dev_calibrated_choice", + "transfer_dev_tuned_blend", +) +BASE_CONFIGS = {"choice_t1", "transfer_dev_calibrated_choice"} +CHECKPOINT_CONFIGS = BASE_CONFIGS | { + "native_shipped", + "fixed_blend_0_5", + "transfer_dev_tuned_blend", +} + + +def _read_object(path: Path) -> dict: + try: + value = json.loads(path.read_text(encoding="utf-8")) + except FileNotFoundError as error: + raise ValueError(f"missing input artifact: {path}") from error + except json.JSONDecodeError as error: + raise ValueError(f"invalid JSON artifact: {path}: {error}") from error + if not isinstance(value, dict): + raise ValueError(f"JSON artifact must be an object: {path}") + return value + + +def _sha256(path: Path) -> str: + digest = hashlib.sha256() + with path.open("rb") as stream: + for block in iter(lambda: stream.read(1024 * 1024), b""): + digest.update(block) + return digest.hexdigest() + + +def _display_path(path: Path) -> str: + resolved = path.resolve() + try: + return str(resolved.relative_to(ROOT.resolve())) + except ValueError: + return str(resolved) + + +def _finite(value: Any) -> bool: + return ( + not isinstance(value, bool) + and isinstance(value, (int, float)) + and math.isfinite(value) + ) + + +def _probability(value: Any, context: str) -> float: + if not _finite(value) or not 0 <= value <= 1: + raise ValueError(f"{context} must be a finite number in [0, 1]") + return float(value) + + +def _positive(value: Any, context: str) -> float: + if not _finite(value) or value <= 0: + raise ValueError(f"{context} must be a finite positive number") + return float(value) + + +def _metric(metric: Any, expected_n: int, context: str) -> dict: + if not isinstance(metric, dict): + raise ValueError(f"{context} must be an object") + if metric.get("n") != expected_n: + raise ValueError( + f"{context} expected n={expected_n}, got {metric.get('n')!r}" + ) + correct = metric.get("correct") + if isinstance(correct, bool) or not isinstance(correct, int) or not 0 <= correct <= expected_n: + raise ValueError(f"{context}.correct must be an integer in [0, {expected_n}]") + accuracy = _probability(metric.get("accuracy"), f"{context}.accuracy") + if not math.isclose(accuracy, correct / expected_n, rel_tol=0, abs_tol=1e-12): + raise ValueError(f"{context}.accuracy does not equal correct / n") + for key in ( + "cross_entropy_from_gold", + "gold_entropy", + "kl_from_gold", + "brier", + "ece", + "mean_confidence", + ): + if key in metric and not _finite(metric[key]): + raise ValueError(f"{context}.{key} must be finite") + return copy.deepcopy(metric) + + +def _manifest_component_accuracy(report: dict, component: str, expected_n: int, context: str) -> float: + try: + clean = report["components"][component]["clean"] + except (KeyError, TypeError) as error: + raise ValueError(f"{context}: missing {component} clean report") from error + if clean.get("n") != expected_n: + raise ValueError( + f"{context}: {component} expected clean n={expected_n}, " + f"got {clean.get('n')!r}" + ) + return _probability(clean.get("acc"), f"{context}.{component}.clean.acc") + + +def _validate_source(spec: dict, manifest: dict, context: str) -> None: + source = manifest.get("source") + predictor = manifest.get("predictor") + if not isinstance(source, dict) or not isinstance(predictor, dict): + raise ValueError(f"{context}: missing source or predictor provenance") + is_base = spec["weights"] == "frozen_base" + if is_base: + if not source.get("base") or source.get("checkpoint") is not None: + raise ValueError(f"{context}: frozen-base run must use source.base only") + if predictor.get("adapter_applied") is not False: + raise ValueError(f"{context}: frozen-base run unexpectedly applied an adapter") + if manifest.get("checkpoint_artifacts") is not None: + raise ValueError(f"{context}: frozen-base run has checkpoint artifacts") + else: + if not source.get("checkpoint") or source.get("base") is not None: + raise ValueError(f"{context}: checkpoint run must use source.checkpoint only") + if predictor.get("adapter_applied") is not True: + raise ValueError(f"{context}: checkpoint run did not apply its adapter") + artifacts = manifest.get("checkpoint_artifacts") + if not isinstance(artifacts, dict) or not isinstance(artifacts.get("files"), dict): + raise ValueError(f"{context}: checkpoint artifact hashes are missing") + if predictor.get("native_decision_mode") != spec["native_decision_mode"]: + raise ValueError( + f"{context}: expected native_decision_mode={spec['native_decision_mode']!r}, " + f"got {predictor.get('native_decision_mode')!r}" + ) + + +def _immutable_revision(value: Any) -> str | None: + if isinstance(value, str) and re.fullmatch(r"[0-9a-f]{40}", value): + return value + return None + + +def _snapshot_revision(path: Any, repository: str) -> str | None: + """Return a Hub snapshot revision only when the path names this repository.""" + + if not isinstance(path, str): + return None + encoded_repository = repository.replace("/", "--") + match = re.search( + rf"(?:^|/)models--{re.escape(encoded_repository)}/snapshots/([0-9a-f]{{40}})(?:/|$)", + path, + ) + return match.group(1) if match else None + + +def _public_source(spec: dict, manifest: dict) -> dict: + """Expose reproducible public identity separately from local load paths. + + A local release directory is not silently treated as a public Hub revision. + In that case the evaluated files remain identified by their SHA-256 values + and the missing public revision is explicit. + """ + + repository = spec["public_repository"] + base_loading = manifest.get("base_loading") + base_loading = base_loading if isinstance(base_loading, dict) else {} + base_repository = base_loading.get("canonical_base") + base_revision = _immutable_revision(base_loading.get("canonical_revision")) + if not isinstance(base_repository, str) or not base_repository: + raise ValueError(f"{spec['key']}: canonical base repository is missing") + if base_revision is None: + raise ValueError(f"{spec['key']}: immutable canonical base revision is missing") + + artifacts = manifest.get("checkpoint_artifacts") + artifact_hashes = {} + if isinstance(artifacts, dict) and isinstance(artifacts.get("files"), dict): + for filename, metadata in artifacts["files"].items(): + if not isinstance(metadata, dict) or not re.fullmatch( + r"[0-9a-f]{64}", str(metadata.get("sha256", "")) + ): + raise ValueError( + f"{spec['key']}: checkpoint artifact {filename!r} lacks SHA-256" + ) + artifact_hashes[filename] = { + "sha256": metadata["sha256"], + "bytes": metadata.get("bytes"), + } + + if spec["weights"] == "frozen_base": + if repository != base_repository: + raise ValueError( + f"{spec['key']}: public base repository does not match canonical base" + ) + revision = base_revision + revision_evidence = "manifest.base_loading.canonical_revision" + repository_evidence = "manifest.base_loading.canonical_base" + limitation = None + else: + source = manifest["source"]["checkpoint"] + revision = _snapshot_revision(source, repository) + revision_evidence = "manifest.source.checkpoint" if revision else None + repository_evidence = ( + "manifest.source.checkpoint" + if revision + else "results/model-family-v2.json#released_models" + ) + limitation = None if revision else ( + "The evaluated checkpoint came from a local release. Its exact files " + "are pinned below by SHA-256, but the run manifest and release metadata " + "do not prove an immutable public-repository revision for those bytes." + ) + + return { + "repository": repository, + "url": f"https://huggingface.co/{repository}", + "revision": revision, + "revision_status": "verified" if revision else "not_verified", + "repository_evidence": repository_evidence, + "revision_evidence": revision_evidence, + "evaluated_artifact_sha256": artifact_hashes, + "base_model": { + "repository": base_repository, + "revision": base_revision, + }, + "limitation": limitation, + } + + +def _load_run(run_root: Path, spec: dict) -> dict: + directory = run_root / spec["key"] + manifest_path = directory / "manifest.json" + ensemble_path = directory / "ensemble.json" + manifest = _read_object(manifest_path) + ensemble = _read_object(ensemble_path) + context = spec["key"] + + model = manifest.get("model") + if not isinstance(model, str) or not model.strip(): + raise ValueError(f"{context}: manifest model must be a non-empty string") + if ensemble.get("model") != model: + raise ValueError(f"{context}: manifest and ensemble model names differ") + _validate_source(spec, manifest, context) + + expected_components = {"choice"} + if spec["native_readout"] is not None: + expected_components.add("native") + reports = manifest.get("reports") + datasets = ensemble.get("datasets") + if not isinstance(reports, dict) or not isinstance(datasets, dict): + raise ValueError(f"{context}: missing report or ensemble datasets") + if set(DATASETS) - set(reports) or set(DATASETS) - set(datasets): + raise ValueError(f"{context}: one or more required datasets are missing") + + expected_configs = BASE_CONFIGS if spec["native_readout"] is None else CHECKPOINT_CONFIGS + summarized = {} + for dataset, expected in DATASETS.items(): + report = reports[dataset] + if report.get("records") != expected["records"]: + raise ValueError( + f"{context}/{dataset}: expected {expected['records']} records, " + f"got {report.get('records')!r}" + ) + if report.get("questions") != expected["questions"]: + raise ValueError( + f"{context}/{dataset}: expected {expected['questions']} questions, " + f"got {report.get('questions')!r}" + ) + components = set(report.get("components", {})) + executed = set(report.get("components_executed", components)) + if components != expected_components or executed != expected_components: + raise ValueError( + f"{context}/{dataset}: expected components {sorted(expected_components)}, " + f"got reports={sorted(components)}, executed={sorted(executed)}" + ) + component_accuracy = { + component: _manifest_component_accuracy( + report, component, expected["headline_n"], f"{context}/{dataset}" + ) + for component in sorted(expected_components) + } + + configurations = datasets[dataset] + if not isinstance(configurations, dict) or set(configurations) != expected_configs: + raise ValueError( + f"{context}/{dataset}: expected ensemble configs " + f"{sorted(expected_configs)}, got " + f"{sorted(configurations) if isinstance(configurations, dict) else configurations!r}" + ) + checked = { + name: _metric( + values, + expected["headline_n"], + f"{context}/{dataset}/{name}", + ) + for name, values in configurations.items() + } + expected_excluded = expected["questions"] - expected["headline_n"] + for name, values in checked.items(): + if values.get("excluded_non_headline_rows") != expected_excluded: + raise ValueError( + f"{context}/{dataset}/{name}: expected " + f"excluded_non_headline_rows={expected_excluded}" + ) + if not math.isclose( + checked["choice_t1"]["accuracy"], component_accuracy["choice"], + rel_tol=0, abs_tol=1e-12, + ): + raise ValueError(f"{context}/{dataset}: choice component and ensemble are not aligned") + if checked["choice_t1"]["correct"] != checked["transfer_dev_calibrated_choice"]["correct"]: + raise ValueError( + f"{context}/{dataset}: temperature-only calibration changed hard decisions" + ) + if "native" in expected_components and not math.isclose( + checked["native_shipped"]["accuracy"], component_accuracy["native"], + rel_tol=0, abs_tol=1e-12, + ): + raise ValueError(f"{context}/{dataset}: native component and ensemble are not aligned") + + summarized[dataset] = { + "zero_shot": { + key: checked[key] for key in ZERO_SHOT_CONFIGS if key in checked + }, + "transfer_dev_tuned": { + key: checked[key] for key in TUNED_CONFIGS if key in checked + }, + } + + protocol = ensemble.get("protocol") + selected = ensemble.get("selected") + if not isinstance(protocol, dict) or protocol.get("selection_rows") != 1_046: + raise ValueError(f"{context}: ensemble must select on 1,046 Transfer-dev rows") + if not isinstance(selected, dict): + raise ValueError(f"{context}: missing selected ensemble settings") + choice_temperature = _positive( + selected.get("choice_temperature"), f"{context}.choice_temperature" + ) + blend_temperature = _positive( + selected.get("blend_temperature"), f"{context}.blend_temperature" + ) + native_weight = _probability( + selected.get("native_weight"), f"{context}.native_weight" + ) + if spec["native_readout"] is None and native_weight != 0: + raise ValueError(f"{context}: frozen base selected a native weight") + + dataset_hashes = manifest.get("datasets") + runtime = manifest.get("runtime") + if not isinstance(dataset_hashes, dict) or not dataset_hashes: + raise ValueError(f"{context}: dataset hashes are missing") + if not isinstance(runtime, dict) or not runtime.get("code_revision"): + raise ValueError(f"{context}: runtime/code revision is missing") + + return { + "key": spec["key"], + "model": model, + "family": spec["family"], + "weights": spec["weights"], + "native_readout": spec["native_readout"], + "artifacts": { + "manifest": { + "path": _display_path(manifest_path), + "sha256": _sha256(manifest_path), + }, + "ensemble": { + "path": _display_path(ensemble_path), + "sha256": _sha256(ensemble_path), + }, + }, + "source": copy.deepcopy(manifest.get("source")), + "public_source": _public_source(spec, manifest), + "checkpoint_artifacts": copy.deepcopy(manifest.get("checkpoint_artifacts")), + "base_loading": copy.deepcopy(manifest.get("base_loading")), + "predictor": copy.deepcopy(manifest.get("predictor")), + "runtime": copy.deepcopy(runtime), + "dataset_hashes": copy.deepcopy(dataset_hashes), + "evaluation_protocol": copy.deepcopy(manifest.get("protocol")), + "ensemble_protocol": copy.deepcopy(protocol), + "selected": { + **copy.deepcopy(selected), + "choice_temperature": choice_temperature, + "blend_temperature": blend_temperature, + "native_weight": native_weight, + }, + "datasets": summarized, + } + + +def _find_model(rows: Any, name: str, context: str) -> dict: + if not isinstance(rows, list): + raise ValueError(f"{context}: models must be a list") + matches = [row for row in rows if isinstance(row, dict) and row.get("model") == name] + if len(matches) != 1: + raise ValueError(f"{context}: expected exactly one {name!r} row") + _probability(matches[0].get("accuracy"), f"{context}/{name}.accuracy") + return copy.deepcopy(matches[0]) + + +def _load_external(path: Path) -> dict: + external = _read_object(path) + typed = external.get("typed_decisions") + text = external.get("jevjudge_text") + if not isinstance(typed, dict) or not isinstance(text, dict): + raise ValueError("external artifact is missing Typed Decisions or JevJudge text") + + typed_rows = typed.get("models") + if not isinstance(typed_rows, list): + raise ValueError("external Typed Decisions models must be a list") + published = [ + row for row in typed_rows + if isinstance(row, dict) + and row.get("kind") == "published_only" + and _finite(row.get("accuracy")) + ] + if not published: + raise ValueError("external artifact has no published Typed Decisions baseline") + typed_top = copy.deepcopy(max(published, key=lambda row: (row["accuracy"], row["model"]))) + _probability(typed_top.get("accuracy"), "Typed published top accuracy") + typed_typesafe = _find_model( + typed_rows, "Jev 1.13 (OpenRouter)", "external/typed_decisions" + ) + + text_rows = text.get("models") + text_typesafe = _find_model( + text_rows, "Jev 1.13 (OpenRouter)", "external/jevjudge_text" + ) + kev_rows = [ + row for row in text_rows + if isinstance(row, dict) + and row.get("kind") == "open_kev" + and _finite(row.get("accuracy")) + ] if isinstance(text_rows, list) else [] + if not kev_rows: + raise ValueError("external artifact has no scored JevJudge-text Kev baseline") + text_kev = copy.deepcopy(max(kev_rows, key=lambda row: (row["accuracy"], row["model"]))) + for row in (text_typesafe, text_kev): + if (row.get("answered"), row.get("requested")) != (724, 724): + raise ValueError( + f"external/jevjudge_text/{row['model']}: expected 724/724 coverage" + ) + + return { + "artifact": { + "path": _display_path(path), + "sha256": _sha256(path), + "artifact_version": external.get("artifact_version"), + "provenance": copy.deepcopy(external.get("provenance")), + }, + "typed_decisions": [typed_top, typed_typesafe], + "jevjudge_text": [text_typesafe, text_kev], + } + + +def _load_repository_catalog(path: Path) -> dict: + catalog = _read_object(path) + released = catalog.get("released_models") + repositories = { + row.get("repository") + for row in released + if isinstance(row, dict) and isinstance(row.get("repository"), str) + } if isinstance(released, list) else set() + required = { + spec["public_repository"] + for spec in RUN_SPECS + if spec["weights"] != "frozen_base" + } + missing = required - repositories + if missing: + raise ValueError( + "model release catalog is missing canonical repositories: " + + ", ".join(sorted(missing)) + ) + return { + "path": _display_path(path), + "sha256": _sha256(path), + "release": catalog.get("release"), + } + + +def build_results( + run_root: Path, + external_path: Path = DEFAULT_EXTERNAL, + model_catalog_path: Path = DEFAULT_MODEL_CATALOG, +) -> dict: + """Build a validated, JSON-serializable result artifact without writing it.""" + + run_root = Path(run_root) + repository_catalog = _load_repository_catalog(Path(model_catalog_path)) + runs = [_load_run(run_root, spec) for spec in RUN_SPECS] + if len({run["model"] for run in runs}) != len(runs): + raise ValueError("the five run manifests must have unique model labels") + + reference_hashes = runs[0]["dataset_hashes"] + for run in runs[1:]: + if run["dataset_hashes"] != reference_hashes: + raise ValueError( + f"{run['key']}: dataset hashes differ from {runs[0]['key']}" + ) + revisions = {run["runtime"].get("code_revision") for run in runs} + if len(revisions) != 1: + raise ValueError("all five runs must use one code revision") + + external = _load_external(Path(external_path)) + result_datasets = {} + for dataset, expected in DATASETS.items(): + result_datasets[dataset] = { + **copy.deepcopy(expected), + "runs": [ + { + "run": run["key"], + "model": run["model"], + "family": run["family"], + "weights": run["weights"], + "native_readout": run["native_readout"], + **copy.deepcopy(run["datasets"][dataset]), + } + for run in runs + ], + } + result_datasets["typed_test"]["external_baselines"] = copy.deepcopy( + external["typed_decisions"] + ) + result_datasets["jevjudge_text"]["external_baselines"] = copy.deepcopy( + external["jevjudge_text"] + ) + result_datasets["jevjudge_full"] = { + "label": "JevJudge full", + "records": 3_220, + "status": "unsupported", + "accuracy": None, + "reason": ( + "Training-free choice-token readout is text-only; image/video records " + "are not stripped or relabeled as a full-suite result." + ), + "native_checkpoint_results": "results/external-zero-shot-v1.json#jevjudge_full", + } + + public_runs = [] + for run in runs: + public_runs.append({key: copy.deepcopy(value) for key, value in run.items() if key != "datasets"}) + + return { + "schema_version": 2, + "artifact": "choice-readout-v2", + "title": "Training-free choice-token readout evaluation", + "generated_by": "scripts/build_choice_readout_results.py", + "method": { + "readout": ( + "Constrain the next token to exact one-character option IDs and " + "renormalize their probability mass; this can change argmax decisions." + ), + "training": "No additional training for the choice-token readout.", + "option_ids": "A-Z followed by a-z; at most 52 options.", + "calibration": ( + "Scalar temperature changes probabilities but not argmax accuracy." + ), + "ensemble": ( + "Native + choice configurations use log-linear/geometric pooling." + ), + }, + "protocol_groups": { + "zero_shot": { + "meaning": ( + "No Typed, JevJudge, JevBench, or Transfer-test labels tune the " + "readout configuration." + ), + "configurations": list(ZERO_SHOT_CONFIGS), + }, + "transfer_dev_tuned": { + "meaning": ( + "Weight and additional temperature are selected only on Transfer-v9 " + "development, then frozen for every displayed evaluation panel." + ), + "configurations": list(TUNED_CONFIGS), + "accuracy_note": ( + "Temperature-only calibrated choice has the same hard accuracy as " + "choice_t1; only a changed blend weight can change its argmax." + ), + }, + }, + "expected_counts": copy.deepcopy(DATASETS), + "dataset_hashes": copy.deepcopy(reference_hashes), + "code_revision": next(iter(revisions)), + "source_runs": public_runs, + "external_source": external["artifact"], + "repository_catalog": repository_catalog, + "datasets": result_datasets, + "legacy_results": { + "status": "retained unchanged", + "artifact": "results/letter-readout-v1.json", + }, + } + + +def write_results(path: Path, artifact: dict) -> None: + path = Path(path) + path.parent.mkdir(parents=True, exist_ok=True) + payload = json.dumps( + artifact, indent=2, ensure_ascii=False, allow_nan=False + ) + "\n" + path.write_text(payload, encoding="utf-8") + + +def _run_result(artifact: dict, dataset: str, run_key: str) -> dict: + matches = [ + row for row in artifact["datasets"][dataset]["runs"] + if row["run"] == run_key + ] + if len(matches) != 1: + raise ValueError(f"{dataset}: expected exactly one {run_key} result") + return matches[0] + + +def _plot_rows(artifact: dict, dataset: str) -> list[dict]: + def accuracy(run: str, group: str, config: str) -> float: + return _run_result(artifact, dataset, run)[group][config]["accuracy"] + + rows = [ + { + "label": "Frozen Qwen3.5 4B ยท Choice", + "group": "zero_shot", + "kind": "choice", + "accuracy": accuracy("base4", "zero_shot", "choice_t1"), + }, + { + "label": "JevAny 4B Pointer ยท Native", + "group": "zero_shot", + "kind": "native", + "accuracy": accuracy("pointer4", "zero_shot", "native_shipped"), + }, + { + "label": "JevAny 4B Pointer ยท Choice", + "group": "zero_shot", + "kind": "choice", + "accuracy": accuracy("pointer4", "zero_shot", "choice_t1"), + }, + { + "label": "JevAny 4B Direct-Token ยท Native", + "group": "zero_shot", + "kind": "native", + "accuracy": accuracy("direct4", "zero_shot", "native_shipped"), + }, + { + "label": "JevAny 4B Direct-Token ยท Choice", + "group": "zero_shot", + "kind": "choice", + "accuracy": accuracy("direct4", "zero_shot", "choice_t1"), + }, + { + "label": "Frozen Qwen3.8 27B ยท Choice", + "group": "zero_shot", + "kind": "choice", + "accuracy": accuracy("base27", "zero_shot", "choice_t1"), + }, + { + "label": "JevAny 27B Pointer ยท Native", + "group": "zero_shot", + "kind": "native", + "accuracy": accuracy("pointer27", "zero_shot", "native_shipped"), + }, + { + "label": "JevAny 27B Pointer ยท Choice", + "group": "zero_shot", + "kind": "choice", + "accuracy": accuracy("pointer27", "zero_shot", "choice_t1"), + }, + { + "label": "JevAny 4B Pointer ยท Tuned blend", + "group": "transfer_dev_tuned", + "kind": "tuned", + "accuracy": accuracy( + "pointer4", "transfer_dev_tuned", "transfer_dev_tuned_blend" + ), + }, + { + "label": "JevAny 4B Direct-Token ยท Tuned blend", + "group": "transfer_dev_tuned", + "kind": "tuned", + "accuracy": accuracy( + "direct4", "transfer_dev_tuned", "transfer_dev_tuned_blend" + ), + }, + { + "label": "JevAny 27B Pointer ยท Tuned blend", + "group": "transfer_dev_tuned", + "kind": "tuned", + "accuracy": accuracy( + "pointer27", "transfer_dev_tuned", "transfer_dev_tuned_blend" + ), + }, + ] + baselines = artifact["datasets"][dataset]["external_baselines"] + for baseline in baselines: + is_typesafe = baseline["model"] == "Jev 1.13 (OpenRouter)" + label = ( + "TypeSafe Jev 1.13" + if is_typesafe + else f"{baseline['model']} ยท " + + ("published" if dataset == "typed_test" else "open") + ) + rows.append({ + "label": label, + "group": "external", + "kind": "typesafe" if is_typesafe else "external", + "accuracy": baseline["accuracy"], + "published": dataset == "typed_test" and not is_typesafe, + }) + return rows + + +def render_svg(artifact: dict, output: Path) -> None: + """Render the compact two-panel README figure from a validated artifact.""" + + try: + import matplotlib + except ImportError as error: + raise RuntimeError("matplotlib is required to render the SVG") from error + + matplotlib.use("Agg") + import matplotlib.pyplot as plt + from matplotlib.patches import Patch + + plt.rcParams.update({ + "font.family": ["DejaVu Sans", "sans-serif"], + "svg.fonttype": "none", + "svg.hashsalt": "jevany-choice-readout-v2", + "text.color": INK, + "hatch.linewidth": 0.7, + }) + panels = ( + ("typed_test", "Typed Decisions", "2,000 decisions ยท accuracy (%)"), + ("jevjudge_text", "JevJudge text", "724 records ยท accuracy (%)"), + ) + fig, axes = plt.subplots(1, 2, figsize=(18, 9.3), dpi=110, facecolor="white") + fig.subplots_adjust(left=0.19, right=0.988, bottom=0.12, top=0.78, wspace=0.30) + fig.text(0.035, 0.955, "Training-free choice readout", fontsize=24, weight="bold") + fig.text( + 0.035, + 0.914, + "Zero-shot readouts and Transfer-dev-tuned blends ยท no Typed or JevJudge labels used for tuning", + fontsize=12.2, + color=MUTED, + ) + fig.legend( + handles=[ + Patch(facecolor=CHOICE, label="Choice T=1 ยท zero-shot"), + Patch(facecolor=NATIVE, label="Native shipped ยท zero-shot"), + Patch(facecolor=TUNED, hatch="///", label="Transfer-dev-tuned blend"), + Patch(facecolor=TYPESAFE, label="TypeSafe Jev"), + Patch(facecolor=EXTERNAL, hatch="////", label="Published / open baseline"), + ], + loc="upper right", + bbox_to_anchor=(0.988, 0.982), + ncol=3, + frameon=False, + fontsize=10.1, + handlelength=1.2, + columnspacing=1.25, + ) + + description_parts = [] + colors = { + "choice": CHOICE, + "native": NATIVE, + "tuned": TUNED, + "external": EXTERNAL, + "typesafe": TYPESAFE, + } + positions = list(range(8)) + list(range(9, 12)) + list(range(13, 15)) + for ax, (dataset, title, scope) in zip(axes, panels, strict=True): + rows = _plot_rows(artifact, dataset) + if len(rows) != len(positions): + raise ValueError(f"{dataset}: plot requires exactly 15 rows") + values = [row["accuracy"] * 100 for row in rows] + xmax = min(100, max(70, int(math.ceil((max(values) + 4) / 10) * 10))) + ax.axhspan(-0.7, 7.65, color="#F7FAFC", zorder=0) + ax.axhspan(8.35, 11.65, color="#EDF7F4", zorder=0) + ax.axhspan(12.35, 14.65, color="#FAF7FC", zorder=0) + for y, row, value in zip(positions, rows, values, strict=True): + bars = ax.barh( + y, + value, + height=0.64, + color=colors[row["kind"]], + edgecolor="#805A2B" if row.get("published") else "none", + linewidth=0.7 if row.get("published") else 0, + zorder=3, + ) + if row["kind"] == "tuned": + bars[0].set_hatch("///") + elif row.get("published"): + bars[0].set_hatch("////") + ax.text( + min(value + xmax * 0.015, xmax * 0.985), + y, + f"{value:.1f}", + ha="right" if value > xmax * 0.91 else "left", + va="center", + fontsize=9.2, + color=INK, + weight="bold" if row["group"] != "external" else "normal", + ) + ax.text(0, -0.78, "ZERO-SHOT", fontsize=9.0, color=MUTED, weight="bold") + ax.text(0, 8.22, "TRANSFER-DEV-TUNED", fontsize=9.0, color=TUNED, weight="bold") + ax.text(0, 12.22, "EXTERNAL REFERENCE", fontsize=9.0, color=MUTED, weight="bold") + ax.set_title(title, loc="left", fontsize=17, color=INK, weight="bold", pad=28) + ax.text(0, 1.015, scope, transform=ax.transAxes, fontsize=10.2, color=MUTED) + ax.set_xlim(0, xmax) + ticks = list(range(0, xmax + 1, 20)) + ax.set_xticks(ticks) + ax.set_xticklabels([str(value) for value in ticks], fontsize=9.2, color=MUTED) + ax.set_yticks(positions) + ax.set_yticklabels([row["label"] for row in rows], fontsize=9.0, color=INK) + for tick, row in zip(ax.get_yticklabels(), rows, strict=True): + if row["group"] != "external": + tick.set_weight("bold" if row["kind"] == "tuned" else "normal") + ax.set_ylim(15.0, -1.1) + ax.tick_params(axis="x", length=0, pad=7) + ax.tick_params(axis="y", length=0, pad=7) + ax.set_axisbelow(True) + ax.grid(axis="x", color=RULE, linewidth=0.8) + for spine in ax.spines.values(): + spine.set_visible(False) + ax.axvline(0, color=RULE, linewidth=1) + description_parts.append( + f"{title}: " + ", ".join( + f"{row['label']} {row['accuracy'] * 100:.2f}%" for row in rows + ) + ) + + fig.text( + 0.035, + 0.044, + "JevJudge full is unsupported for this text-only readout. Choice temperature calibration is omitted here because it cannot change accuracy.", + fontsize=10.1, + color=MUTED, + ) + output = Path(output) + output.parent.mkdir(parents=True, exist_ok=True) + metadata = { + "Date": None, + "Title": "Training-free choice readout accuracy", + "Description": " ".join(description_parts), + } + fig.savefig(output, metadata=metadata) + plt.close(fig) + output.write_text( + "\n".join(line.rstrip() for line in output.read_text().splitlines()) + "\n", + encoding="utf-8", + ) + + +def main(argv: list[str] | None = None) -> None: + parser = argparse.ArgumentParser(description=__doc__) + parser.add_argument("--run-root", type=Path, required=True) + parser.add_argument("--external", type=Path, default=DEFAULT_EXTERNAL) + parser.add_argument("--model-catalog", type=Path, default=DEFAULT_MODEL_CATALOG) + parser.add_argument("--json-out", type=Path, default=DEFAULT_JSON_OUT) + parser.add_argument("--svg-out", type=Path, default=DEFAULT_SVG_OUT) + args = parser.parse_args(argv) + + artifact = build_results(args.run_root, args.external, args.model_catalog) + write_results(args.json_out, artifact) + render_svg(artifact, args.svg_out) + print(f"Wrote {_display_path(args.json_out)} and {_display_path(args.svg_out)}") + + +if __name__ == "__main__": + main() diff --git a/scripts/build_external_report_appendix.py b/scripts/build_external_report_appendix.py index 435746f..341bad4 100644 --- a/scripts/build_external_report_appendix.py +++ b/scripts/build_external_report_appendix.py @@ -1,10 +1,12 @@ #!/usr/bin/env python3 -"""Build and reproducibly merge the two-page external-evaluation appendix.""" +"""Build and reproducibly merge the external and choice-readout appendices.""" import argparse from contextlib import contextmanager from datetime import datetime, timezone +import hashlib import json +import math import os from pathlib import Path import tempfile @@ -17,6 +19,14 @@ from matplotlib.backends.backend_pdf import PdfPages from matplotlib.patches import FancyBboxPatch from pypdf import PdfReader, PdfWriter +from pypdf._cmap import get_encoding +from pypdf.generic import ( + ArrayObject, + ContentStream, + IndirectObject, + NumberObject, + TextStringObject, +) ROOT = Path(__file__).resolve().parents[1] @@ -25,27 +35,78 @@ RULE = "#DCE4EC" TEAL = "#278577" PALE = "#F3F7FA" -PDF_TIMESTAMP = datetime(2026, 10, 1, tzinfo=timezone.utc) -PDF_DATE_LITERAL = "D:20261001000000Z" +PDF_TIMESTAMP = datetime(2026, 10, 2, tzinfo=timezone.utc) +PDF_DATE_LITERAL = "D:20261002000000Z" +EXTERNAL_APPENDIX_PAGES = 2 +CHOICE_APPENDIX_PAGES = 2 +APPENDIX_PAGES = EXTERNAL_APPENDIX_PAGES + CHOICE_APPENDIX_PAGES +REPORT_BASE_PAGES = 19 +MERGED_REPORT_PAGES = REPORT_BASE_PAGES + APPENDIX_PAGES APPENDIX_METADATA = { - "Title": "JevAny Technical Report โ€” External Evaluation Appendix", + "Title": "JevAny Technical Report โ€” Evaluation Appendices", "Author": "SimpleJev", - "Subject": "Typed Decisions and JevJudge-Public v0.3 evaluation", + "Subject": "External evaluation and training-free choice-token readout", "Creator": "scripts/build_external_report_appendix.py", "Producer": "Matplotlib PDF backend", "CreationDate": PDF_TIMESTAMP, "ModDate": PDF_TIMESTAMP, } MERGED_METADATA = { - "/Title": "JevAny: Toward General Decision Intelligence โ€” with Agent-Harness and External-Evaluation Appendices", + "/Title": "JevAny: Toward General Decision Intelligence โ€” merged technical report", "/Author": "SimpleJev", - "/Subject": "JevAny model, bounded decision harness, and external decision-suite evaluation", + "/Subject": ( + "JevAny model, bounded decision harness, external evaluation, and " + "training-free choice-token readout" + ), "/Creator": "scripts/build_external_report_appendix.py", "/Producer": "pypdf", "/CreationDate": PDF_DATE_LITERAL, "/ModDate": PDF_DATE_LITERAL, } +# The source report predates the final 27B release checkpoint. Synchronize the +# handful of release-headline fields while merging so the one published PDF +# does not call two different checkpoints "current". The original layout, +# resources, links, and every unrelated result remain untouched. +CURRENT_RELEASE_TEXT = { + 1: (("85.76", "86.04", 2), ("90.48", "90.04", 2)), + 4: ( + ("22,160", "44,319", 1), + ("85.76", "86.04", 1), + ("90.48", "90.04", 1), + ("0.392", "0.388", 1), + ("0.200", "0.195", 1), + ("0.030", "0.026", 1), + ), + 7: (("85.76", "86.04", 1),), + 8: ( + ("22,160", "44,319", 1), + ("18.83", "39.43", 1), + ("602.7", "1,261.7", 1), + ("1,423", "2,082", 1), + ), +} + +CURRENT_COMPUTE_PROVENANCE = ( + ( + "onds;itsvalueusestherecordedrunstartandcheckpointtimestamp." + "Direct-token,Qwen27B,and", + "onds; its value uses the recorded run start and checkpoint timestamp. " + "Direct-token and Muse use", + ), + ( + "Museusecheckpoint-nativeelapsedtime;Qwen4Bpointerusesterminaltrainertelemetrybecause", + "checkpoint-native elapsed time; Qwen 4B pointer uses terminal trainer telemetry. " + "Qwen 27B uses", + ), + ( + "thereleasedcheckpointisthecompletedstep13,850run." + "Peakwithin-runparallelismwas40GPUs.", + "estimated cumulative seconds-per-record timing. " + "Peak within-run parallelism was 40 GPUs.", + ), +) + @contextmanager def atomic_destination(destination: Path): @@ -67,15 +128,24 @@ def atomic_destination(destination: Path): temporary.unlink(missing_ok=True) -def header(fig, title: str, subtitle: str, page: int) -> None: - fig.text(0.065, 0.955, "SimpleJev ยท TECHNICAL REPORT ยท APPENDIX L", color=TEAL, +def header( + fig, + title: str, + subtitle: str, + page: int, + *, + appendix: str = "L", + section: str = "External evaluation", + pages: int = EXTERNAL_APPENDIX_PAGES, +) -> None: + fig.text(0.065, 0.955, f"SimpleJev ยท TECHNICAL REPORT ยท APPENDIX {appendix}", color=TEAL, fontsize=8.5, weight="bold") fig.text(0.065, 0.905, title, color=INK, fontsize=22, weight="bold") fig.text(0.065, 0.872, subtitle, color=MUTED, fontsize=9.5) fig.lines.append(plt.Line2D([0.065, 0.935], [0.845, 0.845], transform=fig.transFigure, color=RULE, linewidth=1)) fig.text(0.065, 0.035, "JevAny ยท Toward General Decision Intelligence", color=MUTED, fontsize=7.5) - fig.text(0.935, 0.035, f"External evaluation ยท {page}/2", color=MUTED, + fig.text(0.935, 0.035, f"{section} ยท {page}/{pages}", color=MUTED, fontsize=7.5, ha="right") @@ -92,26 +162,48 @@ def overview_page(pdf: PdfPages, data: dict, chart: Path) -> None: fig = plt.figure(figsize=(8.5, 11), facecolor="white") header(fig, "External decision-suite evaluation", "Accuracy on Typed Decisions ยท JevJudge-Public v0.3 full multimodal suite ยท text-only subset", 1) - ax = fig.add_axes([0.055, 0.42, 0.89, 0.36]) - ax.imshow(mpimg.imread(chart)) - ax.axis("off") + chart_image = mpimg.imread(chart) + # The source asset is a wide, three-panel README graphic. Rendering it as + # one image makes its labels roughly four points on a portrait report page. + # Crop the panels (including their titles and axes) and use a 1 + 2 layout + # so every label is materially larger while the source pixels stay intact. + height, width = chart_image.shape[:2] + if width >= 3 and height >= 3: + y_start, y_stop = int(height * 0.15), int(height * 0.91) + panels = ( + (chart_image[y_start:y_stop, int(width * 0.02):int(width * 0.43)], + [0.065, 0.500, 0.870, 0.300]), + (chart_image[y_start:y_stop, int(width * 0.42):int(width * 0.71)], + [0.065, 0.115, 0.420, 0.335]), + (chart_image[y_start:y_stop, int(width * 0.70):int(width * 0.995)], + [0.515, 0.115, 0.420, 0.335]), + ) + for panel, bounds in panels: + ax = fig.add_axes(bounds) + ax.imshow(panel) + ax.axis("off") + else: + # Keep tiny fixture images usable in report-generation tests. + ax = fig.add_axes([0.065, 0.115, 0.870, 0.685]) + ax.imshow(chart_image) + ax.axis("off") full = {row["model"]: row for row in data["jevjudge_full"]["models"]} text = {row["model"]: row for row in data["jevjudge_text"]["models"]} typed = {row["model"]: row for row in data["typed_decisions"]["models"]} - qwen = full["JevAny-Qwen3.8-27B"] - jeff = full["Jeff-Qwen3.5-2B"] - card(fig, 0.065, "Full-suite accuracy", - f"Qwen3.8-27B: {qwen['accuracy'] * 100:.2f}%\n" - f"Jeff-2B: {jeff['accuracy'] * 100:.2f}%\n" - f"Gap: {(qwen['accuracy'] - jeff['accuracy']) * 100:.2f} points") - card(fig, 0.365, "Typed accuracy", - f"JevAny Qwen27: {typed['JevAny-Qwen3.8-27B']['accuracy'] * 100:.2f}%\n" - f"Jev 1.13: {typed['Jev 1.13 (OpenRouter)']['accuracy'] * 100:.2f}%\n" - "Typed result published ยท gap 0.10 pt") - card(fig, 0.665, "Text-only accuracy", - f"JevAny Qwen27: {text['JevAny-Qwen3.8-27B']['accuracy'] * 100:.2f}%\n" - f"Jev 1.13: {text['Jev 1.13 (OpenRouter)']['accuracy'] * 100:.2f}%\n" - f"Kev-27B: {text['Kev-27B']['accuracy'] * 100:.2f}%") + qwen_full = full["JevAny-Qwen3.8-27B"] + qwen_typed = typed["JevAny-Qwen3.8-27B"] + qwen_text = text["JevAny-Qwen3.8-27B"] + fig.text( + 0.065, 0.818, + "Same 13-model order ยท โ€” unsupported/no result ยท teal JevAny ยท gold Jev API ยท purple Kev ยท gray other open ยท hatch published", + color=MUTED, fontsize=7.8, + ) + fig.text( + 0.065, 0.795, + f"Qwen3.8-27B topline ยท Typed {qwen_typed['accuracy'] * 100:.2f}% ยท " + f"JevJudge full {qwen_full['accuracy'] * 100:.2f}% ยท text-only {qwen_text['accuracy'] * 100:.2f}%", + color=TEAL, fontsize=8.2, weight="bold", + ) pdf.savefig(fig) plt.close(fig) @@ -173,47 +265,699 @@ def results_page(pdf: PdfPages, data: dict) -> None: plt.close(fig) -def build_appendix(data_path: Path, chart_path: Path, output: Path) -> None: - data = json.loads(data_path.read_text()) +CHOICE_RUN_SPECS = { + "base4": ("Qwen3.5-4B", "frozen_base", None), + "pointer4": ("Qwen3.5-4B", "jevany_sft", "pointer"), + "direct4": ("Qwen3.5-4B", "jevany_sft", "direct-token"), + "base27": ("Qwen3.8-27B", "frozen_base", None), + "pointer27": ("Qwen3.8-27B", "jevany_sft", "pointer"), +} +CHOICE_DATASET_SPECS = ( + ("jevbench_public", "JevBench\npublic", 231, 231, 231), + ("transfer_test", "Transfer-v9\ntest", 1_264, 1_264, 1_046), + ("typed_test", "Typed\ntest", 400, 2_000, 2_000), + ("jevjudge_text", "JevJudge\ntext", 724, 724, 724), +) + + +def _read_json_object(path: Path, name: str) -> dict: + try: + value = json.loads(path.read_text(encoding="utf-8")) + except FileNotFoundError as error: + raise ValueError(f"missing {name}: {path}") from error + except json.JSONDecodeError as error: + raise ValueError(f"invalid {name}: {path}: {error}") from error + if not isinstance(value, dict): + raise ValueError(f"{name} must be a JSON object: {path}") + return value + + +def _choice_metric(row: dict, group: str, config: str, n: int, context: str) -> dict: + try: + metric = row[group][config] + except (KeyError, TypeError) as error: + raise ValueError(f"{context}: missing {group}/{config}") from error + if not isinstance(metric, dict) or metric.get("n") != n: + raise ValueError(f"{context}/{group}/{config}: expected n={n}") + correct = metric.get("correct") + accuracy = metric.get("accuracy") + if isinstance(correct, bool) or not isinstance(correct, int) or not 0 <= correct <= n: + raise ValueError(f"{context}/{group}/{config}: invalid correct count") + if ( + isinstance(accuracy, bool) + or not isinstance(accuracy, (int, float)) + or not math.isfinite(accuracy) + or not math.isclose(float(accuracy), correct / n, rel_tol=0, abs_tol=1e-12) + ): + raise ValueError(f"{context}/{group}/{config}: accuracy does not equal correct / n") + return metric + + +def load_choice_artifact(path: Path) -> dict: + """Load the validated v2 matrix consumed by Appendix M.""" + + artifact = _read_json_object(path, "choice-readout artifact") + if artifact.get("schema_version") != 2 or artifact.get("artifact") != "choice-readout-v2": + raise ValueError("choice-readout artifact must use schema_version=2") + method = artifact.get("method") + if not isinstance(method, dict) or method.get("option_ids") != ( + "A-Z followed by a-z; at most 52 options." + ): + raise ValueError("choice-readout artifact does not declare the v2 52-ID method") + + source_rows = artifact.get("source_runs") + if not isinstance(source_rows, list): + raise ValueError("choice-readout artifact has no source_runs list") + source_by_key = { + row.get("key"): row for row in source_rows if isinstance(row, dict) + } + if set(source_by_key) != set(CHOICE_RUN_SPECS) or len(source_rows) != len(source_by_key): + raise ValueError("choice-readout artifact must contain the five required source runs") + for key, (family, weights, native_readout) in CHOICE_RUN_SPECS.items(): + row = source_by_key[key] + if (row.get("family"), row.get("weights"), row.get("native_readout")) != ( + family, weights, native_readout + ): + raise ValueError(f"choice-readout source run {key!r} has inconsistent identity") + selected = row.get("selected") + if not isinstance(selected, dict): + raise ValueError(f"choice-readout source run {key!r} has no selected settings") + for field in ("native_weight", "choice_temperature", "blend_temperature"): + value = selected.get(field) + if ( + isinstance(value, bool) + or not isinstance(value, (int, float)) + or not math.isfinite(value) + ): + raise ValueError(f"choice-readout source run {key!r} has invalid {field}") + if not 0 <= selected["native_weight"] <= 1: + raise ValueError(f"choice-readout source run {key!r} has invalid native_weight") + if selected["choice_temperature"] <= 0 or selected["blend_temperature"] <= 0: + raise ValueError(f"choice-readout source run {key!r} has invalid temperature") + if weights == "frozen_base" and selected["native_weight"] != 0: + raise ValueError(f"frozen-base run {key!r} selected a native weight") + ensemble_protocol = row.get("ensemble_protocol") + if not isinstance(ensemble_protocol, dict) or ( + ensemble_protocol.get("selection_rows") != 1_046 + ): + raise ValueError(f"choice-readout source run {key!r} has invalid selection rows") + + datasets = artifact.get("datasets") + if not isinstance(datasets, dict): + raise ValueError("choice-readout artifact has no datasets object") + for dataset, _label, records, questions, n in CHOICE_DATASET_SPECS: + panel = datasets.get(dataset) + if not isinstance(panel, dict) or ( + panel.get("records"), panel.get("questions"), panel.get("headline_n") + ) != (records, questions, n): + raise ValueError( + f"choice-readout dataset {dataset!r} must declare " + f"records/questions/headline_n={records}/{questions}/{n}" + ) + rows = panel.get("runs") + if not isinstance(rows, list): + raise ValueError(f"choice-readout dataset {dataset!r} has no runs list") + by_key = {row.get("run"): row for row in rows if isinstance(row, dict)} + if set(by_key) != set(CHOICE_RUN_SPECS) or len(rows) != len(by_key): + raise ValueError(f"choice-readout dataset {dataset!r} must contain five runs") + for key, (family, weights, native_readout) in CHOICE_RUN_SPECS.items(): + row = by_key[key] + if (row.get("family"), row.get("weights"), row.get("native_readout")) != ( + family, weights, native_readout + ): + raise ValueError(f"{dataset}/{key}: inconsistent run identity") + _choice_metric(row, "zero_shot", "choice_t1", n, f"{dataset}/{key}") + if native_readout is not None: + _choice_metric(row, "zero_shot", "native_shipped", n, f"{dataset}/{key}") + _choice_metric( + row, + "transfer_dev_tuned", + "transfer_dev_tuned_blend", + n, + f"{dataset}/{key}", + ) + + full = datasets.get("jevjudge_full") + if not isinstance(full, dict) or ( + full.get("records"), full.get("status"), full.get("accuracy") + ) != (3_220, "unsupported", None): + raise ValueError("choice-token JevJudge full must be explicitly unsupported") + return artifact + + +def _rounded_box(fig, x, y, width, height, *, facecolor=PALE, edgecolor=RULE): + patch = FancyBboxPatch( + (x, y), width, height, transform=fig.transFigure, + boxstyle="round,pad=0.010,rounding_size=0.010", + facecolor=facecolor, edgecolor=edgecolor, linewidth=0.8, + ) + fig.patches.append(patch) + + +def choice_method_page(pdf: PdfPages, data: dict) -> None: + fig = plt.figure(figsize=(8.5, 11), facecolor="white") + header( + fig, + "Training-free choice-token readout", + "A constrained readout, probability calibration, and stacking are different operations", + 1, + appendix="M", + section="Choice-token readout", + pages=CHOICE_APPENDIX_PAGES, + ) + fig.text( + 0.065, 0.815, + "The frozen language model scores one exact option ID at the answer position. " + "No explanation is generated and no additional training is required for this readout.", + color=INK, fontsize=10.0, wrap=True, linespacing=1.45, + ) + + pipeline = ( + ("1", "Ordered options", "Preserve candidate order"), + ("2", "52 exact IDs", "Aโ€“Z followed by aโ€“z"), + ("3", "One-token projection", "Read next-token mass"), + ("4", "Decision distribution", "Renormalize over IDs"), + ) + for index, (number, title, body) in enumerate(pipeline): + x = 0.065 + index * 0.222 + _rounded_box(fig, x, 0.685, 0.185, 0.095, facecolor="#F5F9F8") + fig.text(x + 0.014, 0.750, number, color=TEAL, fontsize=9, weight="bold") + fig.text(x + 0.042, 0.750, title, color=INK, fontsize=9.2, weight="bold") + fig.text(x + 0.014, 0.712, body, color=MUTED, fontsize=7.8) + if index < len(pipeline) - 1: + fig.text(x + 0.198, 0.730, "โ†’", color=TEAL, fontsize=14, weight="bold") + + semantic_cards = ( + ( + "Choice-token ยท T=1", + "Training-free constrained\nprojection. This is a readout,\n" + "not temperature calibration;\nit can disagree with native.", + ), + ( + "Temperature calibration", + "Fit scalar T on separate\ndevelopment rows. T changes\n" + "probabilities; it cannot change\nargmax or accuracy.", + ), + ( + "Native + choice stack", + "Log-linear pooling can change\nrankings. Select native weight\n" + "and additional temperature on\nTransfer-v9 development only.", + ), + ) + for index, (title, body) in enumerate(semantic_cards): + x = 0.065 + index * 0.299 + _rounded_box(fig, x, 0.475, 0.270, 0.150) + fig.text(x + 0.016, 0.590, title, color=INK, fontsize=9.6, weight="bold") + fig.text( + x + 0.016, 0.558, body, color=MUTED, fontsize=7.8, + va="top", linespacing=1.35, + ) + + sources = {row["key"]: row for row in data["source_runs"]} + table = ( + ("Frozen base", "โ€”", "Choice-token", f"{sources['base4']['family']}, {sources['base27']['family']}"), + ("JevAny SFT", "Pointer", "Choice-token", f"{sources['pointer4']['family']}, {sources['pointer27']['family']}"), + ("JevAny SFT", "Direct-Token", "Choice-token", sources["direct4"]["family"]), + ) + fig.text(0.065, 0.425, "Five-run comparison", color=INK, fontsize=12, weight="bold") + columns = (0.065, 0.245, 0.420, 0.610) + for x, label in zip(columns, ("Weights", "Native reference", "Training-free path", "Backbone family")): + fig.text(x, 0.390, label, color=MUTED, fontsize=8.0, weight="bold") + fig.lines.append(plt.Line2D([0.065, 0.935], [0.377, 0.377], transform=fig.transFigure, + color=RULE, linewidth=1)) + for index, values in enumerate(table): + y = 0.346 - index * 0.052 + for x, value in zip(columns, values): + fig.text(x, y, value, color=INK, fontsize=8.4) + fig.lines.append(plt.Line2D([0.065, 0.935], [y - 0.019, y - 0.019], + transform=fig.transFigure, color="#EEF2F6", linewidth=0.7)) + + _rounded_box( + fig, 0.065, 0.123, 0.870, 0.075, + facecolor="#FFF8E8", edgecolor="#E7C46A", + ) + fig.text(0.083, 0.176, "Release sync and protocol boundary", color="#8A5A00", fontsize=9.5, weight="bold") + fig.text( + 0.083, 0.151, + "The merged report and README identify the current 27B release checkpoint: step 44,319.", + color=INK, fontsize=8.0, + ) + fig.text( + 0.083, 0.132, + "Appendix M native rows are matched prompt-v2/runtime reruns; Appendix L's earlier external-study row can differ by one decision.", + color=INK, fontsize=8.0, + ) + _rounded_box(fig, 0.065, 0.061, 0.870, 0.047, facecolor="#F7FAFC") + fig.text(0.083, 0.091, "Scope", color=INK, fontsize=8.7, weight="bold") + fig.text( + 0.137, 0.091, + "Text only ยท up to 52 IDs ยท JevJudge text 724 supported ยท full multimodal unsupported ยท agent tasks unchanged", + color=MUTED, fontsize=7.6, + ) + pdf.savefig(fig) + plt.close(fig) + + +def _dataset_run(data: dict, dataset: str, run: str) -> dict: + matches = [row for row in data["datasets"][dataset]["runs"] if row["run"] == run] + if len(matches) != 1: + raise ValueError(f"{dataset}: expected one {run} row") + return matches[0] + + +def _choice_matrix_rows(data: dict) -> list[dict]: + sources = {row["key"]: row for row in data["source_runs"]} + definitions = ( + ("Frozen Qwen3.5-4B", "Choice T=1", "base4", "zero_shot", "choice_t1", "โ€”"), + ("JevAny 4B Pointer", "Native", "pointer4", "zero_shot", "native_shipped", "shipped"), + ("", "Choice T=1", "pointer4", "zero_shot", "choice_t1", "โ€”"), + ("", "Tuned stack", "pointer4", "transfer_dev_tuned", "transfer_dev_tuned_blend", None), + ("JevAny 4B Direct-Token", "Native", "direct4", "zero_shot", "native_shipped", "shipped"), + ("", "Choice T=1", "direct4", "zero_shot", "choice_t1", "โ€”"), + ("", "Tuned stack", "direct4", "transfer_dev_tuned", "transfer_dev_tuned_blend", None), + ("Frozen Qwen3.8-27B", "Choice T=1", "base27", "zero_shot", "choice_t1", "โ€”"), + ("JevAny 27B Pointer", "Native", "pointer27", "zero_shot", "native_shipped", "shipped"), + ("", "Choice T=1", "pointer27", "zero_shot", "choice_t1", "โ€”"), + ("", "Tuned stack", "pointer27", "transfer_dev_tuned", "transfer_dev_tuned_blend", None), + ) + rows = [] + for model, readout, run, group, config, weight in definitions: + if weight is None: + weight = f"{sources[run]['selected']['native_weight']:.2f}" + metrics = [] + for dataset, _label, _records, _questions, n in CHOICE_DATASET_SPECS: + metric = _choice_metric( + _dataset_run(data, dataset, run), group, config, n, f"{dataset}/{run}" + ) + metrics.append(metric) + rows.append({ + "model": model, + "readout": readout, + "run": run, + "weight": weight, + "metrics": metrics, + "tuned": readout == "Tuned stack", + "native": readout == "Native", + }) + return rows + + +def choice_results_page(pdf: PdfPages, data: dict) -> None: + fig = plt.figure(figsize=(8.5, 11), facecolor="white") + header( + fig, + "Choice-token accuracy matrix", + "Matched runs ยท current 27B release step 44,319 ยท Transfer-dev settings frozen before evaluation", + 2, + appendix="M", + section="Choice-token readout", + pages=CHOICE_APPENDIX_PAGES, + ) + rows = _choice_matrix_rows(data) + ax = fig.add_axes([0.055, 0.285, 0.89, 0.530]) + ax.axis("off") + x_model, x_readout, x_weight = 0.00, 0.315, 0.495 + x_metrics = (0.635, 0.755, 0.875, 0.995) + ax.text(x_model, 1.035, "Model", transform=ax.transAxes, fontsize=8.0, + color=MUTED, weight="bold") + ax.text(x_readout, 1.035, "Readout", transform=ax.transAxes, fontsize=8.0, + color=MUTED, weight="bold") + ax.text(x_weight, 1.035, "w_native", transform=ax.transAxes, fontsize=8.0, + color=MUTED, weight="bold", ha="right") + for x, (_dataset, label, _records, _questions, n) in zip( + x_metrics, CHOICE_DATASET_SPECS, strict=True + ): + ax.text(x, 1.035, f"{label}\n(n={n:,})", transform=ax.transAxes, + fontsize=7.7, color=MUTED, weight="bold", ha="right", va="bottom") + ax.plot([0, 1], [0.995, 0.995], transform=ax.transAxes, color=RULE, linewidth=1) + + group_spans = ((0, 0), (1, 3), (4, 6), (7, 7), (8, 10)) + step, first_y = 0.082, 0.930 + for group_index, (start, end) in enumerate(group_spans): + top = first_y - start * step + 0.034 + bottom = first_y - end * step - 0.034 + if group_index % 2: + patch = FancyBboxPatch( + (-0.012, bottom), 1.022, top - bottom, + transform=ax.transAxes, boxstyle="round,pad=0.004,rounding_size=0.005", + facecolor="#F7FAFC", edgecolor="none", zorder=0, + ) + ax.add_patch(patch) + + maxima = [max(row["metrics"][index]["accuracy"] for row in rows) for index in range(4)] + for row_index, row in enumerate(rows): + y = first_y - row_index * step + color = TEAL if row["tuned"] else INK + ax.text(x_model, y, row["model"], transform=ax.transAxes, fontsize=8.0, + color=INK, va="center", weight="bold" if row["model"] else "normal") + ax.text(x_readout, y, row["readout"], transform=ax.transAxes, fontsize=8.0, + color=color, va="center", weight="bold" if row["tuned"] else "normal") + ax.text(x_weight, y, row["weight"], transform=ax.transAxes, fontsize=8.0, + color=MUTED, va="center", ha="right") + for metric_index, (x, metric) in enumerate(zip(x_metrics, row["metrics"], strict=True)): + best = math.isclose(metric["accuracy"], maxima[metric_index], rel_tol=0, abs_tol=1e-12) + ax.text( + x, y, f"{metric['accuracy'] * 100:.2f}", transform=ax.transAxes, + fontsize=8.1, color=TEAL if best else color, va="center", ha="right", + weight="bold" if best or row["tuned"] else "normal", + ) + ax.plot([0, 1], [y - 0.040, y - 0.040], transform=ax.transAxes, + color="#EEF2F6", linewidth=0.6, zorder=1) + + sources = {row["key"]: row for row in data["source_runs"]} + tuned = [] + for key, label in (("pointer4", "4B Pointer"), ("direct4", "4B Direct-Token"), + ("pointer27", "27B Pointer")): + selected = sources[key]["selected"] + tuned.append( + f"{label}: w={selected['native_weight']:.2f}, T={selected['blend_temperature']:.2f}" + ) + fig.text(0.065, 0.235, "Transfer-dev selection", color=INK, fontsize=10.5, weight="bold") + fig.text(0.065, 0.205, " ยท ".join(tuned), color=TEAL, fontsize=8.4, weight="bold") + fig.text( + 0.065, 0.169, + "Weights and additional temperatures use only 1,046 clean/knowable Transfer-v9 development rows.\n" + "Scalar T changes probabilities, not the accuracy values above.", + color=MUTED, fontsize=8.2, linespacing=1.35, + ) + fig.text( + 0.065, 0.125, + "Native = matched native rerun at the checkpoint's shipped inference temperature.\n" + "JevBench is a public development diagnostic; Transfer, Typed, and JevJudge text are held-out panels.", + color=MUTED, fontsize=8.2, linespacing=1.35, + ) + fig.text( + 0.065, 0.082, + "Choice-token is text-only: JevJudge text 724 is reported; full JevJudge 3,220 is unsupported.\n" + "Source and complete calibration metrics: results/choice-readout-v2.json.", + color=MUTED, fontsize=8.2, linespacing=1.35, + ) + pdf.savefig(fig) + plt.close(fig) + + +def build_appendix( + data_path: Path, + chart_path: Path, + choice_data_path: Path, + output: Path, +) -> None: + data = _read_json_object(data_path, "external-evaluation artifact") + choice_data = load_choice_artifact(choice_data_path) with atomic_destination(output) as temporary: with PdfPages(temporary, metadata=APPENDIX_METADATA) as pdf: overview_page(pdf, data, chart_path) results_page(pdf, data) + choice_method_page(pdf, choice_data) + choice_results_page(pdf, choice_data) appendix = PdfReader(temporary) - if len(appendix.pages) != 2: - raise RuntimeError(f"expected a two-page appendix, got {len(appendix.pages)} pages") + if len(appendix.pages) != APPENDIX_PAGES: + raise RuntimeError( + f"expected a {APPENDIX_PAGES}-page appendix, got {len(appendix.pages)} pages" + ) + + +def _canonical_pdf_object(value, active: set[int] | None = None): + """Resolve PDF references into a stable, object-number-independent value.""" + + if isinstance(value, IndirectObject): + return _canonical_pdf_object(value.get_object(), active) + if active is None: + active = set() + if isinstance(value, dict): + marker = id(value) + if marker in active: + return ("cycle",) + active.add(marker) + try: + items = tuple(sorted( + ( + str(key), + _canonical_pdf_object(item, active), + ) + for key, item in value.items() + if str(key) != "/Length" + )) + stream = None + if hasattr(value, "get_data"): + stream = hashlib.sha256(value.get_data()).hexdigest() + return ("dict", items, stream) + finally: + active.remove(marker) + if isinstance(value, (list, tuple)): + return tuple(_canonical_pdf_object(item, active) for item in value) + if isinstance(value, bytes): + return ("bytes", hashlib.sha256(value).hexdigest()) + if value is None or isinstance(value, (bool, int, float, str)): + return value + return str(value) def page_invariants(page) -> tuple: """Return page properties that must survive the merge unchanged.""" contents = page.get_contents() content_bytes = contents.get_data() if contents is not None else b"" + annotations = [] + for reference in page.get("/Annots", []): + annotation = reference.get_object() + # /P is a back-reference to the owning page and would make the + # otherwise stable annotation target depend on PDF object numbers. + annotations.append(_canonical_pdf_object({ + key: value for key, value in annotation.items() if str(key) != "/P" + })) return ( tuple(float(value) for value in page.mediabox), tuple(float(value) for value in page.cropbox), page.rotation, content_bytes, page.extract_text() or "", - len(page.get("/Annots", [])), + _canonical_pdf_object(page.get("/Resources")), + tuple(annotations), ) +def _font_text_maps(page) -> dict[str, tuple[dict[str, str], dict[str, str]]]: + maps = {} + fonts = page["/Resources"]["/Font"] + for name, reference in fonts.items(): + _encoding, character_map = get_encoding(reference.get_object()) + reverse = { + unicode_text: glyph + for glyph, unicode_text in character_map.items() + if isinstance(glyph, str) + and isinstance(unicode_text, str) + and len(unicode_text) == 1 + } + maps[str(name)] = (character_map, reverse) + return maps + + +def _decode_text_object(value, character_map: dict[str, str]) -> str: + return "".join(character_map.get(glyph, glyph) for glyph in str(value)) + + +def _encode_text_object( + text: str, + reverse_map: dict[str, str], + prototype, +) -> TextStringObject: + try: + glyphs = "".join(reverse_map[character] for character in text) + except KeyError as error: + raise RuntimeError(f"release-sync font cannot encode {error.args[0]!r}") from error + original_length = len(prototype.original_bytes) + glyph_count = len(str(prototype)) + if glyph_count == 0 or original_length % glyph_count: + raise RuntimeError("release-sync text object has an unsupported encoding width") + byte_width = original_length // glyph_count + if byte_width not in {1, 2}: + raise RuntimeError(f"release-sync text uses unsupported {byte_width}-byte glyphs") + raw = b"".join(ord(glyph).to_bytes(byte_width, "big") for glyph in glyphs) + result = TextStringObject(glyphs) + result._original_bytes = raw + return result + + +def _word_array(text: str, reverse_map: dict[str, str], prototype) -> ArrayObject: + words = text.split() + result = ArrayObject() + for index, word in enumerate(words): + if index: + result.append(NumberObject(-300)) + result.append(_encode_text_object(word, reverse_map, prototype)) + return result + + +def synchronize_current_release(page, page_number: int) -> bool: + """Replace the superseded 27B headline fields in the retained report core.""" + + replacements = CURRENT_RELEASE_TEXT.get(page_number) + if replacements is None: + return False + extracted = page.extract_text() or "" + legacy_counts = {old: extracted.count(old) for old, _new, _count in replacements} + if not any(legacy_counts.values()): + return False + for old, _new, expected_count in replacements: + if legacy_counts[old] != expected_count: + raise RuntimeError( + f"base report page {page_number}: expected {expected_count} occurrences " + f"of legacy value {old!r}, found {legacy_counts[old]}" + ) + + text_maps = _font_text_maps(page) + content = ContentStream(page.get_contents(), page.indirect_reference.pdf) + active_font = None + replaced_counts = {old: 0 for old, _new, _count in replacements} + provenance_replaced = 0 + for operands, operator in content.operations: + if operator == b"Tf": + active_font = str(operands[0]) + continue + if operator not in {b"Tj", b"TJ"} or active_font is None: + continue + character_map, reverse_map = text_maps[active_font] + values = operands[0] if operator == b"TJ" else ArrayObject([operands[0]]) + decoded_line = "".join( + _decode_text_object(value, character_map) + for value in values + if hasattr(value, "original_bytes") + ) + + if page_number == 8: + provenance = next( + (replacement for legacy, replacement in CURRENT_COMPUTE_PROVENANCE + if decoded_line == legacy), + None, + ) + if provenance is not None: + prototype = next( + value for value in values if hasattr(value, "original_bytes") + ) + operands[0] = _word_array(provenance, reverse_map, prototype) + provenance_replaced += 1 + continue + + for index, value in enumerate(values): + if not hasattr(value, "original_bytes"): + continue + decoded = _decode_text_object(value, character_map) + updated = decoded + for old, new, _expected_count in replacements: + count = updated.count(old) + if count: + updated = updated.replace(old, new) + replaced_counts[old] += count + if updated != decoded: + values[index] = _encode_text_object(updated, reverse_map, value) + if operator == b"Tj": + operands[0] = values[0] + + for old, _new, expected_count in replacements: + if replaced_counts[old] != expected_count: + raise RuntimeError( + f"base report page {page_number}: replaced {replaced_counts[old]} of " + f"{expected_count} expected {old!r} values" + ) + if page_number == 8 and provenance_replaced != len(CURRENT_COMPUTE_PROVENANCE): + raise RuntimeError( + "base report page 8: could not synchronize the 27B compute provenance" + ) + # In-place list edits do not invalidate ContentStream's raw-byte cache. + content.operations = content.operations + synchronized_parts = [] + active_font = None + for operands, operator in content.operations: + if operator == b"Tf": + active_font = str(operands[0]) + continue + if operator not in {b"Tj", b"TJ"} or active_font is None: + continue + character_map, _reverse_map = text_maps[active_font] + values = operands[0] if operator == b"TJ" else [operands[0]] + synchronized_parts.extend( + _decode_text_object(value, character_map) + for value in values + if hasattr(value, "original_bytes") + ) + synchronized = "".join(synchronized_parts) + for old, new, expected_count in replacements: + if old in synchronized or synchronized.count(new) < expected_count: + raise RuntimeError( + f"base report page {page_number}: failed to synchronize {old!r} to {new!r}; " + f"old={synchronized.count(old)}, new={synchronized.count(new)}" + ) + page.replace_contents(content) + return True + + +def synchronize_report_boundary(page, page_number: int) -> bool: + """Update the old appendix claim that the merged core is byte-unchanged.""" + + if page_number != 10: + return False + extracted = page.extract_text() or "" + if "The original PDF is pre-" not in extracted: + return False + replacements = { + "original ": "merged ", + "pre-": "release-", + "serv": "synch", + "ed unchanged.": "ronized.", + } + counts = {old: 0 for old in replacements} + content = ContentStream(page.get_contents(), page.indirect_reference.pdf) + for operands, operator in content.operations: + if operator not in {b"Tj", b"TJ"}: + continue + values = operands[0] if operator == b"TJ" else ArrayObject([operands[0]]) + for index, value in enumerate(values): + if not isinstance(value, TextStringObject): + continue + updated = str(value) + for old, new in replacements.items(): + count = updated.count(old) + if count: + updated = updated.replace(old, new) + counts[old] += count + if updated != str(value): + values[index] = TextStringObject(updated) + if operator == b"Tj": + operands[0] = values[0] + if any(count != 1 for count in counts.values()): + raise RuntimeError( + f"base report page 10: could not synchronize appendix boundary: {counts}" + ) + content.operations = content.operations + page.replace_contents(content) + return True + + def merge_report(base_report: Path, appendix_report: Path, merged_output: Path, base_pages: int) -> None: - if base_pages < 1: - raise ValueError("--base-pages must be positive") + if base_pages != REPORT_BASE_PAGES: + raise ValueError(f"--base-pages must be exactly {REPORT_BASE_PAGES}") base = PdfReader(base_report) appendix = PdfReader(appendix_report) if len(base.pages) < base_pages: raise ValueError( f"base report has {len(base.pages)} pages, fewer than --base-pages={base_pages}" ) - if len(appendix.pages) != 2: - raise ValueError(f"appendix must have exactly two pages, got {len(appendix.pages)}") + if len(appendix.pages) != APPENDIX_PAGES: + raise ValueError( + f"appendix must have exactly {APPENDIX_PAGES} pages, got {len(appendix.pages)}" + ) writer = PdfWriter() writer.pdf_header = "%PDF-1.7" + synchronized_pages = set() + expected_base_invariants = {} for index in range(base_pages): writer.add_page(base.pages[index]) + changed = synchronize_current_release(writer.pages[-1], index + 1) + changed = synchronize_report_boundary(writer.pages[-1], index + 1) or changed + if changed: + synchronized_pages.add(index) + expected_base_invariants[index] = page_invariants(writer.pages[-1]) for page in appendix.pages: writer.add_page(page) writer.add_metadata(MERGED_METADATA) @@ -222,18 +966,54 @@ def merge_report(base_report: Path, appendix_report: Path, merged_output: Path, with temporary.open("wb") as stream: writer.write(stream) merged = PdfReader(temporary) - expected_pages = base_pages + len(appendix.pages) + expected_pages = MERGED_REPORT_PAGES if len(merged.pages) != expected_pages: raise RuntimeError(f"expected {expected_pages} merged pages, got {len(merged.pages)}") for index in range(base_pages): - if page_invariants(base.pages[index]) != page_invariants(merged.pages[index]): + expected = ( + expected_base_invariants[index] + if index in synchronized_pages + else page_invariants(base.pages[index]) + ) + actual = page_invariants(merged.pages[index]) + if index in synchronized_pages: + # A writer-attached page temporarily extracts its glyph IDs, + # while the serialized page resolves them through ToUnicode. + # Compare every invariant except that transient text view. + expected = expected[:4] + expected[5:] + actual = actual[:4] + actual[5:] + if expected != actual: raise RuntimeError(f"base page {index + 1} changed during merge") + for page_number, replacements in CURRENT_RELEASE_TEXT.items(): + page_text = merged.pages[page_number - 1].extract_text() or "" + is_release_page = (page_number - 1) in synchronized_pages or any( + page_text.count(new) >= expected_count + for _old, new, expected_count in replacements + ) + if not is_release_page: + continue + for old, new, expected_count in replacements: + if old in page_text or page_text.count(new) < expected_count: + raise RuntimeError( + f"serialized base page {page_number} does not contain the current " + f"release value {new!r}" + ) + boundary_text = merged.pages[9].extract_text() or "" + boundary_text = boundary_text.replace("-\n", "-") + if "The original PDF is pre-" in boundary_text or ( + "The merged PDF is release-synchronized." not in boundary_text + and 9 in synchronized_pages + ): + raise RuntimeError("serialized base page 10 has a stale appendix boundary") -def main() -> None: +def main(argv: list[str] | None = None) -> None: parser = argparse.ArgumentParser(description=__doc__) parser.add_argument("--data", type=Path, default=ROOT / "results/external-zero-shot-v1.json") parser.add_argument("--chart", type=Path, default=ROOT / "docs/external-zero-shot.png") + parser.add_argument( + "--choice-data", type=Path, default=ROOT / "results/choice-readout-v2.json" + ) parser.add_argument("--output", type=Path, required=True) parser.add_argument( "--base-report", @@ -245,12 +1025,12 @@ def main() -> None: type=Path, help="merged report; may be the same path as --base-report", ) - parser.add_argument("--base-pages", type=int, default=19) - args = parser.parse_args() + parser.add_argument("--base-pages", type=int, default=REPORT_BASE_PAGES) + args = parser.parse_args(argv) if (args.base_report is None) != (args.merged_output is None): parser.error("--base-report and --merged-output must be provided together") - if args.base_pages < 1: - parser.error("--base-pages must be positive") + if args.base_pages != REPORT_BASE_PAGES: + parser.error(f"--base-pages must be exactly {REPORT_BASE_PAGES}") if args.base_report is not None: appendix_path = args.output.expanduser().absolute() base_path = args.base_report.expanduser().absolute() @@ -258,12 +1038,12 @@ def main() -> None: if appendix_path in {base_path, merged_path}: parser.error("--output must differ from --base-report and --merged-output") - build_appendix(args.data, args.chart, args.output) + build_appendix(args.data, args.chart, args.choice_data, args.output) print(f"Wrote reproducible appendix: {args.output}") if args.base_report is not None: merge_report(args.base_report, args.output, args.merged_output, args.base_pages) print( - f"Wrote reproducible {args.base_pages + 2}-page merged report: " + f"Wrote reproducible {MERGED_REPORT_PAGES}-page merged report: " f"{args.merged_output}" ) diff --git a/scripts/evaluate_choice_readout_matrix.py b/scripts/evaluate_choice_readout_matrix.py new file mode 100644 index 0000000..fba7833 --- /dev/null +++ b/scripts/evaluate_choice_readout_matrix.py @@ -0,0 +1,447 @@ +#!/usr/bin/env python3 +"""Evaluate training-free choice readout and a checkpoint's native readout. + +The script loads one model once, then scores the same frozen model on JevBench, +Transfer-v9, Typed Decisions and JevJudge's text-only subset. Checkpoint runs +save both readout components so ensemble weights can be selected offline on a +separate calibration split without repeating inference or looking at test +labels. Frozen-base runs save the training-free choice component only. +""" + +from __future__ import annotations + +import argparse +import json +import platform +import statistics +import subprocess +import time +from collections import defaultdict +from pathlib import Path +import math + +import numpy as np +import pyarrow +import pyarrow.parquet as pq +import peft +import safetensors +import torch +import transformers + +from jevany.api import question_keys +from jevany.benchmark import prediction_rows, summarize +from jevany.checkpoint import LoadOptions +from jevany.data import api_request +from jevany.letter_predictor import LetterReadoutPredictor +from jevany.suite import ( + ENCODING, + digest, + load_split, + read_json, + record_digest, + write_json, +) + + +def _typed_rows(dataset: Path, split: str) -> list[dict]: + path = dataset / "all" / f"{split}-00000-of-00001.parquet" + return pq.read_table(path).to_pylist() + + +def _typed_record(row: dict) -> dict: + state = json.loads(row["state"]) + questions = json.loads(row["questions"]) + gold = json.loads(row["gold"]) + converted = {} + for qid, question in questions.items(): + answer = gold[qid] + value = dict(question) + keys = question_keys(value["type"], value.get("criteria")) + # Typed Decisions accuracy is ordered argmax agreement with the soft + # teacher distribution. The dataset's convenience `label` differs on + # tied rows, so derive the target exactly as the benchmark scorer does. + label_key = max(keys, key=lambda key: float(answer["probabilities"][key])) + if value["type"] == "choice": + value["label"] = label_key + elif value["type"] == "noul": + value["label"] = label_key == "true" + else: + value["label"] = int(label_key) + value["target"] = answer["probabilities"] + value["src"] = f"typed_decisions/{row['workflow']}/{qid}" + converted[qid] = value + return { + "state": state, + "questions": converted, + "_meta": { + "id": row["id"], + "group_id": row["id"], + "source": "LocalLLaMA/typed-decisions", + "variant": "clean", + "split": row["split"], + "workflow": row["workflow"], + }, + } + + +def _attach_soft_gold(rows: list[dict], record: dict) -> None: + questions = record["questions"] + for row in rows: + target = questions[row["question"]].get("target") + if target is not None: + row["gold"] = [float(target[key]) for key in row["keys"]] + + +def _component_prediction(probabilities: dict) -> dict: + return {"probabilities": probabilities} + + +def _typed_soft_metrics(rows: list[dict], bins: int = 15) -> dict: + totals = [0] * bins + correct = [0] * bins + confidence = [0.0] * bins + kl, brier = 0.0, 0.0 + for row in rows: + p = [float(value) for value in row["p"]] + gold = [float(value) for value in row["gold"]] + p_total, gold_total = sum(p), sum(gold) + p = [value / p_total for value in p] + gold = [value / gold_total for value in gold] + kl += sum(g * math.log(max(g, 1e-12) / max(value, 1e-12)) + for g, value in zip(gold, p, strict=True)) + brier += sum((value - g) ** 2 for value, g in zip(p, gold, strict=True)) + predicted = max(range(len(p)), key=p.__getitem__) + conf = max(p) + bucket = min(bins - 1, int(conf * bins)) + totals[bucket] += 1 + correct[bucket] += int(predicted == row["label"]) + confidence[bucket] += conf + n = len(rows) + return { + "n": n, + "correct": sum(correct), + "accuracy": sum(correct) / n, + "kl_from_gold": kl / n, + "brier": brier / n, + "ece": sum( + totals[index] / n * abs( + correct[index] / totals[index] - confidence[index] / totals[index] + ) + for index in range(bins) if totals[index] + ), + "mean_confidence": sum(confidence) / n, + } + + +def evaluate(name: str, records: list[dict], predictor, destination: Path) -> dict: + destination.mkdir(parents=True, exist_ok=False) + component_rows: dict[str, list[dict]] = defaultdict(list) + latencies = [] + efficient_long_context_records = 0 + started = time.time() + with (destination / "predictions.jsonl").open("w", encoding=ENCODING) as stream: + for index, record in enumerate(records, 1): + result = predictor(record) + efficient_long_context_records += int( + result.get("efficient_long_context_attention_used", False) + ) + components = result.get("component_probabilities") or { + "choice": result["probabilities"] + } + saved = {} + for component, probabilities in components.items(): + if probabilities is None: + continue + rows = prediction_rows(record, _component_prediction(probabilities)) + _attach_soft_gold(rows, record) + component_rows[component].extend(rows) + saved[component] = probabilities + stream.write(json.dumps({ + "id": record["_meta"]["id"], + "request_sha256": record_digest(api_request(record)), + "component_probabilities": saved, + "latency_ms": result["latency_ms"], + "input_tokens": result["input_tokens"], + }, allow_nan=False) + "\n") + stream.flush() + latencies.append(float(result["latency_ms"])) + if index % 25 == 0: + print(f"{name}: {index}/{len(records)}", flush=True) + reports = {} + for component, rows in component_rows.items(): + write_json(destination / f"{component}-rows.json", rows) + reports[component] = summarize(rows) + if name == "typed_test": + reports[component]["typed_soft_metrics"] = _typed_soft_metrics(rows) + report = { + "dataset": name, + "records": len(records), + "questions": sum(len(record["questions"]) for record in records), + "components": reports, + "components_executed": sorted(component_rows), + "efficient_long_context_records": efficient_long_context_records, + "latency_scope": ( + "end-to-end predictor call; checkpoint calls execute both choice and native " + "components, so these values are not per-component efficiency measurements" + ), + "latency_ms": { + "median": statistics.median(latencies), + "p95": sorted(latencies)[min(len(latencies) - 1, int(.95 * len(latencies)))], + }, + "wall_seconds": time.time() - started, + } + write_json(destination / "report.json", report) + return report + + +def reuse_completed( + name: str, + records: list[dict], + predictor, + destination: Path, +) -> dict: + """Load and validate one completed dataset from an interrupted matrix run.""" + + report_path = destination / "report.json" + if not report_path.is_file(): + raise RuntimeError( + f"cannot resume incomplete dataset directory: {destination}; preserve or " + "move that directory aside before retrying" + ) + report = read_json(report_path) + expected_questions = sum(len(record["questions"]) for record in records) + expected_components = {"choice", "native"} if predictor.return_components else {"choice"} + actual_components = set(report.get("components", {})) + if ( + report.get("dataset") != name + or report.get("records") != len(records) + or report.get("questions") != expected_questions + or actual_components != expected_components + ): + raise RuntimeError(f"completed report does not match requested resume dataset: {destination}") + for component in expected_components: + rows_path = destination / f"{component}-rows.json" + if not rows_path.is_file() or len(read_json(rows_path)) != expected_questions: + raise RuntimeError(f"completed component rows are missing or incomplete: {rows_path}") + return report + + +def _correct(report: dict, component: str) -> int: + clean = report["components"][component]["clean"] + return round(clean["n"] * clean["acc"]) + + +def _code_revision() -> str | None: + root = Path(__file__).resolve().parents[1] + result = subprocess.run( + ["git", "rev-parse", "HEAD"], cwd=root, text=True, + stdout=subprocess.PIPE, stderr=subprocess.DEVNULL, check=False, + ) + return result.stdout.strip() if result.returncode == 0 else None + + +def _checkpoint_artifacts(predictor) -> dict | None: + checkpoint = predictor.checkpoint + if checkpoint is None: + return None + root = Path(checkpoint.path) + files = {} + for name in ("adapter_model.safetensors", "head.pt", "adapter_config.json"): + path = root / name + if path.is_file(): + files[name] = {"sha256": digest(path), "bytes": path.stat().st_size} + return {"requested": checkpoint.requested, "resolved": str(root), "files": files} + + +def main() -> None: + parser = argparse.ArgumentParser(description=__doc__) + source = parser.add_mutually_exclusive_group(required=True) + source.add_argument("--base") + source.add_argument("--checkpoint") + parser.add_argument("--base-load-path", required=True) + parser.add_argument("--revision") + parser.add_argument("--model-name", required=True) + parser.add_argument("--out", type=Path, required=True) + parser.add_argument("--jevbench", type=Path, required=True) + parser.add_argument("--transfer", type=Path, required=True) + parser.add_argument("--typed", type=Path, required=True) + parser.add_argument("--jevjudge", type=Path, required=True) + parser.add_argument("--expected-jevbench-native-correct", type=int) + parser.add_argument("--expected-transfer-native-correct", type=int) + parser.add_argument("--device", default="cuda") + parser.add_argument("--dtype", choices=("fp32", "bf16"), default="bf16") + parser.add_argument("--attn", choices=("eager", "sdpa"), default="sdpa") + parser.add_argument("--max-tokens", type=int, default=65_536) + parser.add_argument( + "--efficient-long-context-tokens", type=int, default=8_192, + help=( + "use memory-linear fused SDPA at or above this many tokens; short " + "parity panels retain exact math SDPA" + ), + ) + parser.add_argument( + "--resume", action="store_true", + help="reuse validated completed dataset directories in an interrupted run", + ) + parser.add_argument( + "--preexisting-code-revision", + help="code revision that produced reused reports (required when any are reused)", + ) + args = parser.parse_args() + if args.revision and args.checkpoint: + parser.error("--revision accompanies --base, not --checkpoint") + if args.base and (args.expected_jevbench_native_correct is not None + or args.expected_transfer_native_correct is not None): + parser.error("native parity assertions require --checkpoint") + + if args.resume: + if not args.out.is_dir(): + parser.error("--resume requires an existing --out directory") + else: + if args.preexisting_code_revision: + parser.error("--preexisting-code-revision requires --resume") + args.out.mkdir(parents=True, exist_ok=False) + options = LoadOptions( + dtype={"fp32": torch.float32, "bf16": torch.bfloat16}[args.dtype], + merge=False, + temperature=None, + base_load_path=args.base_load_path, + attn=args.attn, + ) + predictor = LetterReadoutPredictor( + base=args.base, + checkpoint=args.checkpoint, + revision=args.revision, + device=args.device, + options=options, + temperature=1.0, + native_weight=0.5 if args.checkpoint else 0.0, + return_components=bool(args.checkpoint), + exact_kernels=True, + efficient_long_context_tokens=args.efficient_long_context_tokens, + max_tokens=args.max_tokens, + ) + + current_code_revision = _code_revision() + dataset_code_revisions = {} + reused_datasets = [] + + def run_dataset(name: str, records: list[dict]) -> dict: + destination = args.out / name + if args.resume and destination.exists(): + report = reuse_completed(name, records, predictor, destination) + reused_datasets.append(name) + dataset_code_revisions[name] = args.preexisting_code_revision + return report + report = evaluate(name, records, predictor, destination) + dataset_code_revisions[name] = current_code_revision + return report + + parity_datasets = { + "jevbench_public": load_split(args.jevbench, "development"), + "transfer_calibration": load_split(args.transfer, "development"), + } + reports = { + name: run_dataset(name, records) + for name, records in parity_datasets.items() + } + if args.checkpoint: + checks = ( + ("jevbench_public", args.expected_jevbench_native_correct), + ("transfer_calibration", args.expected_transfer_native_correct), + ) + for dataset, expected in checks: + if expected is not None: + actual = _correct(reports[dataset], "native") + if actual != expected: + raise RuntimeError( + f"{dataset} native parity failed: expected {expected}, got {actual}" + ) + heldout_datasets = { + "transfer_test": load_split(args.transfer, "test", allow_test=True), + "typed_test": [_typed_record(row) for row in _typed_rows(args.typed, "test")], + "jevjudge_text": [ + record for record in load_split(args.jevjudge, "test", allow_test=True) + if record["_meta"]["modality"] == "text" + ], + } + reports.update({ + name: run_dataset(name, records) + for name, records in heldout_datasets.items() + }) + if reused_datasets and not args.preexisting_code_revision: + parser.error("--preexisting-code-revision is required when reports are reused") + + typed_manifest = json.loads((args.typed / ".hf-fetch.json").read_text()) + manifest = { + "model": args.model_name, + "source": {"base": args.base, "checkpoint": args.checkpoint}, + "checkpoint_artifacts": _checkpoint_artifacts(predictor), + "base_loading": predictor.base_loading, + "predictor": predictor.provenance, + "runtime": { + "python": platform.python_version(), + "torch": torch.__version__, + "cuda": torch.version.cuda, + "transformers": transformers.__version__, + "peft": peft.__version__, + "numpy": np.__version__, + "pyarrow": pyarrow.__version__, + "safetensors": safetensors.__version__, + "code_revision": current_code_revision, + }, + "protocol": { + "kernel_policy": "LocalPredictor-compatible exact CUDA policy", + "dtype": args.dtype, + "attn": args.attn, + "temperature": 1.0, + "blend_selection": ( + "not performed here; components are saved for fitting on Transfer-v9 " + "development only" + ), + "efficiency_scope": ( + "base calls execute choice only; checkpoint calls execute choice and native. " + "Do not compare component efficiency from these end-to-end timings" + ), + "long_context_attention": ( + "exact math SDPA below the configured threshold; memory-linear fused SDPA " + "without a math fallback at or above it" + ), + "efficient_long_context_tokens": args.efficient_long_context_tokens, + "resume": { + "enabled": args.resume, + "reused_datasets": reused_datasets, + "preexisting_code_revision": args.preexisting_code_revision, + "dataset_code_revisions": dataset_code_revisions, + }, + }, + "datasets": { + "jevbench_development_sha256": digest(args.jevbench / "development.jsonl"), + "jevbench_manifest_sha256": digest(args.jevbench / "manifest.json"), + "transfer_manifest_sha256": digest(args.transfer / "manifest.json"), + "transfer_development_sha256": digest(args.transfer / "development.jsonl"), + "transfer_test_sha256": digest(args.transfer / "test.jsonl"), + "typed_revision": typed_manifest["resolved_revision"], + "typed_test_sha256": digest( + args.typed / "all" / "test-00000-of-00001.parquet" + ), + "jevjudge_test_sha256": digest(args.jevjudge / "test.jsonl"), + "jevjudge_manifest_sha256": digest(args.jevjudge / "manifest.json"), + }, + "reports": reports, + } + write_json(args.out / "manifest.json", manifest) + print(json.dumps({ + "model": args.model_name, + "results": { + name: { + component: report["clean"]["acc"] + for component, report in value["components"].items() + } + for name, value in reports.items() + }, + }, indent=2), flush=True) + + +if __name__ == "__main__": + main() diff --git a/scripts/fit_choice_ensemble.py b/scripts/fit_choice_ensemble.py new file mode 100644 index 0000000..c4ae147 --- /dev/null +++ b/scripts/fit_choice_ensemble.py @@ -0,0 +1,298 @@ +#!/usr/bin/env python3 +"""Fit choice/native pooling on Transfer-v9 development and score frozen tests. + +The blend weight and temperature are selected only from the deterministic +Transfer-v9 development partition emitted by ``evaluate_choice_readout_matrix``. +Those fixed values are then applied unchanged to every held-out dataset. +""" + +from __future__ import annotations + +import argparse +import json +import math +from pathlib import Path + +import numpy as np + +from jevany.suite import read_json, write_json + + +EPS = 1e-12 + + +def load_components(directory: Path) -> tuple[list[dict], list[dict] | None]: + choice = read_json(directory / "choice-rows.json") + native_path = directory / "native-rows.json" + native = read_json(native_path) if native_path.exists() else None + if native is not None: + left = [(row["id"], row["question"], row["keys"], row["label"]) for row in choice] + right = [(row["id"], row["question"], row["keys"], row["label"]) for row in native] + if left != right: + raise ValueError(f"component rows are not aligned: {directory}") + return choice, native + + +def headline_indices(rows: list[dict]) -> list[int]: + """Indices in the benchmark's accuracy cohort. + + Transfer-v9 also contains non-clean robustness variants and unknowable + questions. The benchmark never counts either in headline accuracy, so + neither may influence weight or temperature selection. + """ + + return [ + index for index, row in enumerate(rows) + if row.get("variant", "clean") == "clean" + and row.get("source") != "unknowable" + ] + + +def select_indices(rows: list[dict], indices: list[int]) -> list[dict]: + return [rows[index] for index in indices] + + +def targets(rows: list[dict]) -> list[np.ndarray]: + values = [] + for row in rows: + if "gold" in row: + target = np.asarray(row["gold"], dtype=np.float64) + target /= target.sum() + else: + target = np.zeros(len(row["p"]), dtype=np.float64) + target[row["label"]] = 1.0 + values.append(target) + return values + + +def pooled_logits(choice: list[dict], native: list[dict] | None, weight: float) -> list[np.ndarray]: + if native is None and weight: + raise ValueError("native weight requires native component rows") + result = [] + for index, row in enumerate(choice): + left = np.log(np.maximum(np.asarray(row["p"], dtype=np.float64), EPS)) + if native is None: + result.append(left) + else: + right = np.log(np.maximum(np.asarray(native[index]["p"], dtype=np.float64), EPS)) + result.append((1 - weight) * left + weight * right) + return result + + +def probabilities(logits: list[np.ndarray], temperature: float) -> list[np.ndarray]: + result = [] + for values in logits: + scaled = values / temperature + scaled -= scaled.max() + p = np.exp(scaled) + result.append(p / p.sum()) + return result + + +def soft_nll(rows: list[dict], ps: list[np.ndarray]) -> float: + gold = targets(rows) + return float(np.mean([ + -float(np.sum(target * np.log(np.maximum(p, EPS)))) + for target, p in zip(gold, ps, strict=True) + ])) + + +def fit_temperature(rows: list[dict], logits: list[np.ndarray]) -> tuple[float, float]: + """Golden-section search in log-temperature space with fixed bounds.""" + + lo, hi = math.log(.05), math.log(20.0) + ratio = (math.sqrt(5) - 1) / 2 + x1, x2 = hi - ratio * (hi - lo), lo + ratio * (hi - lo) + + def objective(log_temperature: float) -> float: + return soft_nll(rows, probabilities(logits, math.exp(log_temperature))) + + f1, f2 = objective(x1), objective(x2) + for _ in range(80): + if f1 <= f2: + hi, x2, f2 = x2, x1, f1 + x1 = hi - ratio * (hi - lo) + f1 = objective(x1) + else: + lo, x1, f1 = x1, x2, f2 + x2 = lo + ratio * (hi - lo) + f2 = objective(x2) + point = (lo + hi) / 2 + return math.exp(point), objective(point) + + +def fit(rows: list[dict], choice: list[dict], native: list[dict] | None) -> dict: + candidates = [0.0] if native is None else [index / 100 for index in range(101)] + best = None + trace = [] + for weight in candidates: + logits = pooled_logits(choice, native, weight) + temperature, nll = fit_temperature(rows, logits) + t1 = probabilities(logits, 1.0) + correct = sum( + int(int(p.argmax()) == row["label"]) + for row, p in zip(rows, t1, strict=True) + ) + item = { + "native_weight": weight, + "temperature": temperature, + "correct": correct, + "accuracy": correct / len(rows), + "nll": nll, + } + trace.append(item) + if best is None or ( + -correct, nll, abs(weight - .5) + ) < ( + -best["correct"], best["nll"], abs(best["native_weight"] - .5) + ): + best = item + return {"selected": best, "grid": trace} + + +def ece(rows: list[dict], ps: list[np.ndarray], bins: int = 15) -> float: + totals = np.zeros(bins) + correct = np.zeros(bins) + confidence = np.zeros(bins) + for row, p in zip(rows, ps, strict=True): + conf = float(p.max()) + bucket = min(bins - 1, int(conf * bins)) + totals[bucket] += 1 + correct[bucket] += int(int(p.argmax()) == row["label"]) + confidence[bucket] += conf + n = len(rows) + return float(sum( + totals[index] / n * abs(correct[index] / totals[index] - confidence[index] / totals[index]) + for index in range(bins) if totals[index] + )) + + +def metrics(rows: list[dict], ps: list[np.ndarray], *, ece_bins: int = 10) -> dict: + gold = targets(rows) + cross_entropy = soft_nll(rows, ps) + gold_entropy = float(np.mean([ + -float(np.sum(target * np.log(np.maximum(target, EPS)))) for target in gold + ])) + return { + "n": len(rows), + "correct": sum(int(int(p.argmax()) == row["label"]) for row, p in zip(rows, ps, strict=True)), + "accuracy": float(np.mean([ + int(int(p.argmax()) == row["label"]) + for row, p in zip(rows, ps, strict=True) + ])), + "cross_entropy_from_gold": cross_entropy, + "gold_entropy": gold_entropy, + "kl_from_gold": cross_entropy - gold_entropy, + "brier": float(np.mean([ + float(np.sum((p - target) ** 2)) + for target, p in zip(gold, ps, strict=True) + ])), + "ece": ece(rows, ps, bins=ece_bins), + "mean_confidence": float(np.mean([p.max() for p in ps])), + } + + +def evaluate(directory: Path, selected: dict) -> dict: + choice, native = load_components(directory) + configurations = { + "choice_t1": (0.0, 1.0), + "transfer_dev_calibrated_choice": (0.0, selected["choice_temperature"]), + } + if native is not None: + configurations.update({ + "native_shipped": (1.0, 1.0), + "fixed_blend_0_5": (.5, 1.0), + "transfer_dev_tuned_blend": ( + selected["native_weight"], selected["blend_temperature"] + ), + }) + indices = headline_indices(choice) + ece_bins = 15 if directory.name == "typed_test" else 10 + results = {} + for name, (weight, temperature) in configurations.items(): + ps = probabilities(pooled_logits(choice, native, weight), temperature) + headline = metrics( + select_indices(choice, indices), + [ps[index] for index in indices], + ece_bins=ece_bins, + ) + headline["cohort"] = "clean and knowable" + headline["excluded_non_headline_rows"] = len(choice) - len(indices) + headline["ece_bins"] = ece_bins + results[name] = headline + return results + + +def main() -> None: + parser = argparse.ArgumentParser(description=__doc__) + parser.add_argument("--run", type=Path, required=True) + parser.add_argument("--out", type=Path, required=True) + args = parser.parse_args() + manifest = read_json(args.run / "manifest.json") + calibration_dir = args.run / "transfer_calibration" + choice, native = load_components(calibration_dir) + indices = headline_indices(choice) + calibration_choice = select_indices(choice, indices) + calibration_native = select_indices(native, indices) if native is not None else None + choice_fit = fit(calibration_choice, calibration_choice, None) + blend_fit = fit(calibration_choice, calibration_choice, calibration_native) + selected = { + "choice_temperature": choice_fit["selected"]["temperature"], + "native_weight": blend_fit["selected"]["native_weight"], + "blend_temperature": blend_fit["selected"]["temperature"], + "selection_objective": ( + "hard-label accuracy on the Transfer-v9 development clean/knowable " + "cohort; calibrated NLL and distance from 0.5 break ties" + ), + } + datasets = {} + for name in ( + "transfer_calibration", "transfer_test", "typed_test", "jevbench_public", + "jevjudge_text" + ): + datasets[name] = evaluate(args.run / name, selected) + result = { + "model": manifest["model"], + "source_manifest": str(args.run / "manifest.json"), + "protocol": { + "selection_split": ( + "Transfer-v9 development clean/knowable accuracy cohort; never any " + "reported test panel" + ), + "selection_rows": len(calibration_choice), + "weight_grid": "0.00 to 1.00 inclusive, step 0.01", + "pool": "log-linear/geometric", + "temperature_fit": "bounded scalar NLL minimization, T in [0.05, 20]", + "ece": "15 equal-width bins for Typed Decisions; 10 for every other dataset", + "selection_objective": "accuracy; calibrated NLL then distance from 0.5 break ties", + "transfer": "the selected weight and temperature are frozen for every held-out dataset", + "native_shipped": ( + "the checkpoint's native probabilities, including its shipped inference " + "temperature; no additional temperature is applied" + ), + "fixed_blend_0_5": ( + "equal log-linear pooling of T=1 choice probabilities and shipped native " + "probabilities; no additional temperature is applied" + ), + "transfer_dev_tuned_blend": ( + "the Transfer-dev-selected native weight followed by the reported scalar " + "additional temperature" + ), + }, + "selected": selected, + "selection": {"choice": choice_fit, "blend": blend_fit}, + "datasets": datasets, + } + write_json(args.out, result) + print(json.dumps({ + "model": result["model"], + "selected": selected, + "test_accuracy": { + dataset: {name: values["accuracy"] for name, values in scores.items()} + for dataset, scores in datasets.items() if dataset != "transfer_calibration" + }, + }, indent=2), flush=True) + + +if __name__ == "__main__": + main() diff --git a/site/files/reports/JevAny_Tech_Report.pdf b/site/files/reports/JevAny_Tech_Report.pdf index de0ad58..cedc9fb 100644 Binary files a/site/files/reports/JevAny_Tech_Report.pdf and b/site/files/reports/JevAny_Tech_Report.pdf differ diff --git a/tests/test_build_choice_readout_results.py b/tests/test_build_choice_readout_results.py new file mode 100644 index 0000000..eb8863d --- /dev/null +++ b/tests/test_build_choice_readout_results.py @@ -0,0 +1,362 @@ +import copy +import json +from pathlib import Path + +import pytest + +from scripts.build_choice_readout_results import ( + DATASETS, + RUN_SPECS, + build_results, + render_svg, + write_results, +) + + +def _metric(n, correct): + accuracy = correct / n + return { + "n": n, + "correct": correct, + "accuracy": accuracy, + "cross_entropy_from_gold": 0.5, + "gold_entropy": 0.1, + "kl_from_gold": 0.4, + "brier": 0.2, + "ece": 0.03, + "mean_confidence": 0.7, + "cohort": "clean and knowable", + "excluded_non_headline_rows": 0, + "ece_bins": 15, + } + + +def _write_json(path, value): + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(json.dumps(value, indent=2) + "\n") + + +def _external_fixture(path): + value = { + "artifact_version": 5, + "provenance": {"fixture": {"sha256": "e" * 64}}, + "typed_decisions": { + "models": [ + { + "model": "Published runner-up", + "kind": "published_only", + "accuracy": 0.74, + }, + { + "model": "Published top", + "kind": "published_only", + "accuracy": 0.77, + }, + { + "model": "Jev 1.13 (OpenRouter)", + "kind": "external_api", + "accuracy": 0.73, + }, + ] + }, + "jevjudge_text": { + "models": [ + { + "model": "Jev 1.13 (OpenRouter)", + "kind": "external_api", + "accuracy": 0.65, + "answered": 724, + "requested": 724, + }, + { + "model": "Kev-4B", + "kind": "open_kev", + "accuracy": 0.54, + "answered": 724, + "requested": 724, + }, + { + "model": "Kev-27B", + "kind": "open_kev", + "accuracy": 0.64, + "answered": 724, + "requested": 724, + }, + ] + }, + } + _write_json(path, value) + + +def _run_fixture(root, spec, ordinal): + is_base = spec["weights"] == "frozen_base" + canonical_base = f"Qwen/{spec['family']}" + snapshot_revision = f"{ordinal + 1}" * 40 + encoded_repository = spec["public_repository"].replace("/", "--") + checkpoint = ( + f"/fixture/models--{encoded_repository}/snapshots/{snapshot_revision}" + if not is_base and spec["key"] != "direct4" + else f"/fixture/local-release/{spec['key']}" + ) + components = ["choice"] if is_base else ["choice", "native"] + model = f"fixture-{spec['key']}" + reports = {} + ensemble_datasets = {} + for dataset_index, (dataset, expected) in enumerate(DATASETS.items()): + n = expected["headline_n"] + choice_correct = n // 2 + ordinal + dataset_index + native_correct = choice_correct + 1 + choice_accuracy = choice_correct / n + native_accuracy = native_correct / n + reports[dataset] = { + "dataset": dataset, + "records": expected["records"], + "questions": expected["questions"], + "components_executed": components, + "components": { + "choice": {"clean": {"n": n, "acc": choice_accuracy}}, + }, + } + if not is_base: + reports[dataset]["components"]["native"] = { + "clean": {"n": n, "acc": native_accuracy} + } + choice = _metric(n, choice_correct) + choice["excluded_non_headline_rows"] = expected["questions"] - n + calibrated = copy.deepcopy(choice) + calibrated["cross_entropy_from_gold"] = 0.45 + configurations = { + "choice_t1": choice, + "transfer_dev_calibrated_choice": calibrated, + } + if not is_base: + native = _metric(n, native_correct) + fixed = _metric(n, min(n, native_correct + 1)) + tuned = _metric(n, min(n, native_correct + 2)) + for value in (native, fixed, tuned): + value["excluded_non_headline_rows"] = expected["questions"] - n + configurations.update({ + "native_shipped": native, + "fixed_blend_0_5": fixed, + "transfer_dev_tuned_blend": tuned, + }) + ensemble_datasets[dataset] = configurations + + hashes = { + "jevbench_development_sha256": "1" * 64, + "jevbench_manifest_sha256": "2" * 64, + "transfer_manifest_sha256": "3" * 64, + "transfer_development_sha256": "4" * 64, + "transfer_test_sha256": "5" * 64, + "typed_revision": "6" * 40, + "typed_test_sha256": "7" * 64, + "jevjudge_test_sha256": "8" * 64, + "jevjudge_manifest_sha256": "9" * 64, + } + manifest = { + "model": model, + "source": { + "base": spec["public_repository"] if is_base else None, + "checkpoint": None if is_base else checkpoint, + }, + "checkpoint_artifacts": None if is_base else { + "requested": f"fixture/{spec['key']}", + "resolved": f"/fixture/{spec['key']}", + "files": { + "adapter_model.safetensors": { + "sha256": f"{ordinal + 1}" * 64, + "bytes": 123, + } + }, + }, + "base_loading": { + "requested": spec["family"], + "resolved": "/fixture/base", + "canonical_base": canonical_base, + "canonical_revision": "a" * 40, + }, + "predictor": { + "method": "training-free exact choice-token projection", + "adapter_applied": not is_base, + "native_decision_mode": spec["native_decision_mode"], + "canonical_base": canonical_base, + "canonical_revision": "a" * 40, + }, + "runtime": { + "python": "3.12.0", + "torch": "2.fixture", + "code_revision": "b" * 40, + }, + "protocol": {"temperature": 1.0, "kernel_policy": "fixture exact"}, + "datasets": hashes, + "reports": reports, + } + ensemble = { + "model": model, + "protocol": { + "selection_rows": 1046, + "selection_split": "Transfer-v9 development", + }, + "selected": { + "choice_temperature": 1.2, + "native_weight": 0.0 if is_base else 0.55, + "blend_temperature": 1.1, + "selection_objective": "fixture", + }, + "datasets": ensemble_datasets, + } + directory = root / spec["key"] + _write_json(directory / "manifest.json", manifest) + _write_json(directory / "ensemble.json", ensemble) + + +@pytest.fixture +def matrix(tmp_path): + root = tmp_path / "runs" + for ordinal, spec in enumerate(RUN_SPECS): + _run_fixture(root, spec, ordinal) + external = tmp_path / "external.json" + _external_fixture(external) + return root, external + + +def test_build_results_preserves_all_five_run_types_and_protocol_groups(matrix): + root, external = matrix + + artifact = build_results(root, external) + + assert artifact["artifact"] == "choice-readout-v2" + assert [run["key"] for run in artifact["source_runs"]] == [ + "base4", "pointer4", "direct4", "base27", "pointer27" + ] + assert {run["weights"] for run in artifact["source_runs"]} == { + "frozen_base", "jevany_sft" + } + assert {run["native_readout"] for run in artifact["source_runs"]} == { + None, "pointer", "direct-token" + } + typed = artifact["datasets"]["typed_test"] + assert typed["questions"] == typed["headline_n"] == 2000 + assert typed["external_baselines"][0]["model"] == "Published top" + assert artifact["datasets"]["jevjudge_text"]["external_baselines"][1]["model"] == "Kev-27B" + assert artifact["datasets"]["jevjudge_full"] == { + "label": "JevJudge full", + "records": 3220, + "status": "unsupported", + "accuracy": None, + "reason": ( + "Training-free choice-token readout is text-only; image/video records " + "are not stripped or relabeled as a full-suite result." + ), + "native_checkpoint_results": "results/external-zero-shot-v1.json#jevjudge_full", + } + base = next(row for row in typed["runs"] if row["run"] == "base4") + direct = next(row for row in typed["runs"] if row["run"] == "direct4") + assert set(base["zero_shot"]) == {"choice_t1"} + assert set(base["transfer_dev_tuned"]) == {"transfer_dev_calibrated_choice"} + assert set(direct["zero_shot"]) == { + "choice_t1", "native_shipped", "fixed_blend_0_5" + } + assert set(direct["transfer_dev_tuned"]) == { + "transfer_dev_calibrated_choice", "transfer_dev_tuned_blend" + } + assert len(artifact["source_runs"][0]["artifacts"]["manifest"]["sha256"]) == 64 + assert artifact["source_runs"][0]["runtime"]["torch"] == "2.fixture" + + sources = {run["key"]: run["public_source"] for run in artifact["source_runs"]} + assert sources["base4"] == { + "repository": "Qwen/Qwen3.5-4B", + "url": "https://huggingface.co/Qwen/Qwen3.5-4B", + "revision": "a" * 40, + "revision_status": "verified", + "repository_evidence": "manifest.base_loading.canonical_base", + "revision_evidence": "manifest.base_loading.canonical_revision", + "evaluated_artifact_sha256": {}, + "base_model": { + "repository": "Qwen/Qwen3.5-4B", + "revision": "a" * 40, + }, + "limitation": None, + } + assert sources["pointer4"]["revision"] == "2" * 40 + assert sources["pointer4"]["revision_status"] == "verified" + assert sources["pointer27"]["revision"] == "5" * 40 + assert sources["direct4"]["repository"] == ( + "SimpleJev/JevAny-Qwen3.5-4B-Direct-Token-LoRA" + ) + assert sources["direct4"]["revision"] is None + assert sources["direct4"]["revision_status"] == "not_verified" + assert sources["direct4"]["revision_evidence"] is None + assert "do not prove an immutable public-repository revision" in ( + sources["direct4"]["limitation"] + ) + assert sources["direct4"]["evaluated_artifact_sha256"] == { + "adapter_model.safetensors": { + "sha256": "3" * 64, + "bytes": 123, + } + } + + +def test_build_results_rejects_component_count_drift(matrix): + root, external = matrix + path = root / "pointer4" / "manifest.json" + manifest = json.loads(path.read_text()) + manifest["reports"]["typed_test"]["components"]["native"]["clean"]["n"] = 1999 + _write_json(path, manifest) + + with pytest.raises(ValueError, match="native expected clean n=2000"): + build_results(root, external) + + +def test_build_results_rejects_native_component_metric_misalignment(matrix): + root, external = matrix + path = root / "direct4" / "ensemble.json" + ensemble = json.loads(path.read_text()) + native = ensemble["datasets"]["jevjudge_text"]["native_shipped"] + native["correct"] += 1 + native["accuracy"] = native["correct"] / native["n"] + _write_json(path, ensemble) + + with pytest.raises(ValueError, match="native component and ensemble are not aligned"): + build_results(root, external) + + +def test_build_results_rejects_cross_run_dataset_hash_drift(matrix): + root, external = matrix + path = root / "direct4" / "manifest.json" + manifest = json.loads(path.read_text()) + manifest["datasets"]["typed_test_sha256"] = "0" * 64 + _write_json(path, manifest) + + with pytest.raises(ValueError, match="dataset hashes differ"): + build_results(root, external) + + +def test_build_results_rejects_release_catalog_without_canonical_repositories( + matrix, tmp_path +): + root, external = matrix + catalog = tmp_path / "model-catalog.json" + _write_json(catalog, {"release": "fixture", "released_models": []}) + + with pytest.raises(ValueError, match="missing canonical repositories"): + build_results(root, external, catalog) + + +def test_synthetic_artifact_and_svg_are_written_only_to_requested_paths(matrix, tmp_path): + pytest.importorskip("matplotlib") + root, external = matrix + artifact = build_results(root, external) + json_out = tmp_path / "out" / "choice.json" + svg_out = tmp_path / "out" / "choice.svg" + + write_results(json_out, artifact) + render_svg(artifact, svg_out) + + assert json.loads(json_out.read_text())["schema_version"] == 2 + svg = svg_out.read_text() + assert "Training-free choice readout" in svg + assert "Typed Decisions" in svg + assert "JevJudge text" in svg + assert "TRANSFER-DEV-TUNED" in svg diff --git a/tests/test_build_external_report_appendix.py b/tests/test_build_external_report_appendix.py new file mode 100644 index 0000000..ef2281f --- /dev/null +++ b/tests/test_build_external_report_appendix.py @@ -0,0 +1,556 @@ +import json +import os +import shutil +from datetime import datetime, timezone +from pathlib import Path + +import pytest + +try: + import matplotlib + import pypdf +except ImportError: + if os.environ.get("JEVANY_REPORT_TESTS_REQUIRED") == "1": + raise + pytest.skip("report dependencies are installed in the report-test job", allow_module_level=True) + +import matplotlib.pyplot as plt +from matplotlib.backends.backend_pdf import PdfPages +from pypdf import PdfReader, PdfWriter + +from scripts.build_external_report_appendix import ( + CHOICE_DATASET_SPECS, + CHOICE_RUN_SPECS, + _choice_matrix_rows, + build_appendix, + load_choice_artifact, + merge_report, + page_invariants, +) +from scripts.build_choice_readout_results import ( + DATASETS as PRODUCER_DATASETS, + RUN_SPECS as PRODUCER_RUN_SPECS, + build_results, + write_results, +) + + +def _write_json(path, value): + path.write_text(json.dumps(value, indent=2) + "\n", encoding="utf-8") + + +def _external_artifact(path): + value = { + "artifact_version": 5, + "provenance": {"fixture": {"sha256": "e" * 64}}, + "typed_decisions": { + "models": [ + {"model": "JevAny-Qwen3.8-27B", "accuracy": 0.728}, + { + "model": "Jev 1.13 (OpenRouter)", + "kind": "external_api", + "accuracy": 0.727, + }, + { + "model": "Published top", + "kind": "published_only", + "accuracy": 0.770, + }, + ], + }, + "jevjudge_text": { + "models": [ + {"model": "JevAny-Qwen3.8-27B", "accuracy": 0.6644}, + { + "model": "Jev 1.13 (OpenRouter)", + "kind": "external_api", + "accuracy": 0.6506, + "answered": 724, + "requested": 724, + }, + { + "model": "Kev-27B", + "kind": "open_kev", + "accuracy": 0.6423, + "answered": 724, + "requested": 724, + }, + ], + }, + "jevjudge_full": { + "models": [ + { + "model": "JevAny-Qwen3.8-27B", + "kind": "ours", + "accuracy": 0.6227, + "skill_role": 0.3555, + "skill_role_ci_95": [0.3284, 0.3818], + "nll": 0.878, + "ece": 0.094, + }, + { + "model": "Jeff-Qwen3.5-2B", + "kind": "open_external", + "accuracy": 0.4823, + "skill_role": 0.1357, + "skill_role_ci_95": [0.1080, 0.1599], + "nll": 1.105, + "ece": 0.137, + }, + ], + }, + } + _write_json(path, value) + + +def _metric(n, correct): + return {"n": n, "correct": correct, "accuracy": correct / n} + + +def _choice_artifact(path): + source_runs = [] + for index, (key, (family, weights, native)) in enumerate(CHOICE_RUN_SPECS.items()): + source_runs.append({ + "key": key, + "model": f"fixture-{key}", + "family": family, + "weights": weights, + "native_readout": native, + "selected": { + "native_weight": 0.0 if native is None else 0.40 + index / 100, + "choice_temperature": 1.20 + index / 100, + "blend_temperature": 1.10 + index / 100, + }, + "ensemble_protocol": {"selection_rows": 1_046}, + }) + + datasets = {} + for dataset_index, (dataset, _label, records, questions, n) in enumerate( + CHOICE_DATASET_SPECS + ): + runs = [] + for run_index, (key, (family, weights, native)) in enumerate( + CHOICE_RUN_SPECS.items() + ): + choice_correct = n // 2 + dataset_index * 2 + run_index + row = { + "run": key, + "model": f"fixture-{key}", + "family": family, + "weights": weights, + "native_readout": native, + "zero_shot": {"choice_t1": _metric(n, choice_correct)}, + "transfer_dev_tuned": { + "transfer_dev_calibrated_choice": _metric(n, choice_correct), + }, + } + if native is not None: + row["zero_shot"].update({ + "native_shipped": _metric(n, choice_correct + 1), + "fixed_blend_0_5": _metric(n, choice_correct + 2), + }) + row["transfer_dev_tuned"]["transfer_dev_tuned_blend"] = _metric( + n, choice_correct + 3 + ) + runs.append(row) + datasets[dataset] = { + "records": records, + "questions": questions, + "headline_n": n, + "runs": runs, + } + datasets["jevjudge_full"] = { + "label": "JevJudge full", + "records": 3_220, + "status": "unsupported", + "accuracy": None, + "reason": "text-only fixture", + } + value = { + "schema_version": 2, + "artifact": "choice-readout-v2", + "method": { + "option_ids": "A-Z followed by a-z; at most 52 options.", + "calibration": "Scalar temperature changes probabilities but not argmax accuracy.", + }, + "source_runs": source_runs, + "datasets": datasets, + } + _write_json(path, value) + return value + + +def _producer_metric(n, correct, excluded): + return { + "n": n, + "correct": correct, + "accuracy": correct / n, + "cross_entropy_from_gold": 0.5, + "gold_entropy": 0.1, + "kl_from_gold": 0.4, + "brier": 0.2, + "ece": 0.03, + "mean_confidence": 0.7, + "cohort": "clean and knowable", + "excluded_non_headline_rows": excluded, + "ece_bins": 15, + } + + +def _producer_run(root, spec, ordinal): + is_base = spec["weights"] == "frozen_base" + canonical_base = f"Qwen/{spec['family']}" + encoded_repository = spec["public_repository"].replace("/", "--") + checkpoint = ( + f"/fixture/models--{encoded_repository}/snapshots/{str(ordinal + 1) * 40}" + if spec["key"] != "direct4" + else f"/fixture/local-release/{spec['key']}" + ) + components = ["choice"] if is_base else ["choice", "native"] + model = f"producer-{spec['key']}" + reports = {} + ensemble_datasets = {} + for dataset_index, (dataset, expected) in enumerate(PRODUCER_DATASETS.items()): + n = expected["headline_n"] + excluded = expected["questions"] - n + choice_correct = n // 2 + ordinal + dataset_index + native_correct = choice_correct + 1 + reports[dataset] = { + "dataset": dataset, + "records": expected["records"], + "questions": expected["questions"], + "components_executed": components, + "components": { + "choice": {"clean": {"n": n, "acc": choice_correct / n}}, + }, + } + if not is_base: + reports[dataset]["components"]["native"] = { + "clean": {"n": n, "acc": native_correct / n} + } + choice = _producer_metric(n, choice_correct, excluded) + calibrated = dict(choice) + calibrated["cross_entropy_from_gold"] = 0.45 + configurations = { + "choice_t1": choice, + "transfer_dev_calibrated_choice": calibrated, + } + if not is_base: + configurations.update({ + "native_shipped": _producer_metric(n, native_correct, excluded), + "fixed_blend_0_5": _producer_metric(n, native_correct + 1, excluded), + "transfer_dev_tuned_blend": _producer_metric( + n, native_correct + 2, excluded + ), + }) + ensemble_datasets[dataset] = configurations + + hashes = { + "jevbench_development_sha256": "1" * 64, + "jevbench_manifest_sha256": "2" * 64, + "transfer_manifest_sha256": "3" * 64, + "transfer_development_sha256": "4" * 64, + "transfer_test_sha256": "5" * 64, + "typed_revision": "6" * 40, + "typed_test_sha256": "7" * 64, + "jevjudge_test_sha256": "8" * 64, + "jevjudge_manifest_sha256": "9" * 64, + } + manifest = { + "model": model, + "source": { + "base": spec["public_repository"] if is_base else None, + "checkpoint": None if is_base else checkpoint, + }, + "checkpoint_artifacts": None if is_base else { + "requested": f"fixture/{spec['key']}", + "resolved": f"/fixture/{spec['key']}", + "files": { + "adapter_model.safetensors": { + "sha256": f"{ordinal + 1}" * 64, + "bytes": 123, + }, + }, + }, + "base_loading": { + "requested": spec["family"], + "resolved": "/fixture/base", + "canonical_base": canonical_base, + "canonical_revision": "a" * 40, + }, + "predictor": { + "method": "training-free exact choice-token projection", + "adapter_applied": not is_base, + "native_decision_mode": spec["native_decision_mode"], + "canonical_base": canonical_base, + "canonical_revision": "a" * 40, + }, + "runtime": { + "python": "3.13.5", + "torch": "2.fixture", + "code_revision": "b" * 40, + }, + "protocol": {"temperature": 1.0, "kernel_policy": "fixture exact"}, + "datasets": hashes, + "reports": reports, + } + ensemble = { + "model": model, + "protocol": { + "selection_rows": 1_046, + "selection_split": "Transfer-v9 development", + }, + "selected": { + "choice_temperature": 1.2, + "native_weight": 0.0 if is_base else 0.50 + ordinal / 100, + "blend_temperature": 1.0 + ordinal / 10, + "selection_objective": "fixture", + }, + "datasets": ensemble_datasets, + } + directory = root / spec["key"] + directory.mkdir(parents=True) + _write_json(directory / "manifest.json", manifest) + _write_json(directory / "ensemble.json", ensemble) + + +def _base_report(path, pages=21): + metadata = { + "Title": "Synthetic uniquely identified base report", + "CreationDate": datetime(2026, 10, 2, tzinfo=timezone.utc), + "ModDate": datetime(2026, 10, 2, tzinfo=timezone.utc), + } + with PdfPages(path, metadata=metadata) as pdf: + for page_number in range(1, pages + 1): + fig = plt.figure(figsize=(8.5, 11), facecolor="white") + fig.text( + 0.1, + 0.8, + f"UNIQUE BASE PAGE {page_number:02d}", + fontsize=16, + url=f"https://example.test/base/{page_number}", + ) + fig.text(0.1, 0.7, f"resource marker {page_number * 17}", fontsize=9) + pdf.savefig(fig) + plt.close(fig) + + +def _annotation_uris(page): + uris = [] + for reference in page.get("/Annots", []): + annotation = reference.get_object() + action = annotation.get("/A") + if action is not None: + action = action.get_object() + if action.get("/URI") is not None: + uris.append(str(action["/URI"])) + return uris + + +@pytest.fixture +def appendix_inputs(tmp_path): + external = tmp_path / "external.json" + choice = tmp_path / "choice.json" + chart = tmp_path / "chart.png" + _external_artifact(external) + _choice_artifact(choice) + plt.imsave(chart, [[0.0, 0.5], [0.75, 1.0]], cmap="viridis") + return external, chart, choice + + +@pytest.fixture +def producer_artifact(tmp_path): + root = tmp_path / "producer-runs" + for ordinal, spec in enumerate(PRODUCER_RUN_SPECS): + _producer_run(root, spec, ordinal) + external = tmp_path / "producer-external.json" + _external_artifact(external) + artifact = build_results(root, external) + choice = tmp_path / "producer-choice.json" + write_results(choice, artifact) + return external, choice + + +def test_build_appendix_renders_four_pages_directly_from_choice_json( + appendix_inputs, tmp_path +): + external, chart, choice = appendix_inputs + output = tmp_path / "appendices.pdf" + + build_appendix(external, chart, choice, output) + + pdf = PdfReader(output) + assert len(pdf.pages) == 4 + method_text = pdf.pages[2].extract_text() + results_text = pdf.pages[3].extract_text() + assert "APPENDIX M" in method_text + assert "52 exact IDs" in method_text + assert "not temperature calibration" in method_text + assert "Choice-token accuracy matrix" in results_text + assert "Matched runs" in results_text + assert "JevJudge text 724" in results_text + # base4 JevBench fixture: floor(231 / 2) = 115 => 49.78%. + assert "49.78" in results_text + + +def test_producer_artifact_flows_into_pointer_direct_and_tuned_pdf_rows( + producer_artifact, tmp_path +): + external, choice = producer_artifact + chart = tmp_path / "producer-chart.png" + output = tmp_path / "producer-appendices.pdf" + plt.imsave(chart, [[0.0, 0.5], [0.75, 1.0]], cmap="viridis") + + loaded = load_choice_artifact(choice) + source_by_key = {row["key"]: row for row in loaded["source_runs"]} + assert source_by_key["pointer4"]["native_readout"] == "pointer" + assert source_by_key["direct4"]["native_readout"] == "direct-token" + assert source_by_key["pointer4"]["selected"]["native_weight"] == 0.51 + assert source_by_key["direct4"]["selected"]["native_weight"] == 0.52 + rows = _choice_matrix_rows(loaded) + assert [(row["model"], row["readout"], row["run"]) for row in rows] == [ + ("Frozen Qwen3.5-4B", "Choice T=1", "base4"), + ("JevAny 4B Pointer", "Native", "pointer4"), + ("", "Choice T=1", "pointer4"), + ("", "Tuned stack", "pointer4"), + ("JevAny 4B Direct-Token", "Native", "direct4"), + ("", "Choice T=1", "direct4"), + ("", "Tuned stack", "direct4"), + ("Frozen Qwen3.8-27B", "Choice T=1", "base27"), + ("JevAny 27B Pointer", "Native", "pointer27"), + ("", "Choice T=1", "pointer27"), + ("", "Tuned stack", "pointer27"), + ] + matrix = {(row["run"], row["readout"]): row for row in rows} + assert matrix[("pointer4", "Native")]["metrics"][0]["correct"] == 120 + assert matrix[("pointer4", "Tuned stack")]["metrics"][0]["correct"] == 122 + assert matrix[("direct4", "Native")]["metrics"][0]["correct"] == 121 + assert matrix[("direct4", "Tuned stack")]["metrics"][0]["correct"] == 123 + build_appendix(external, chart, choice, output) + + text = PdfReader(output).pages[3].extract_text() + assert "JevAny 4B Pointer" in text + assert "JevAny 4B Direct-Token" in text + assert text.count("Tuned stack") == 3 + assert "4B Pointer: w=0.51, T=1.10" in text + assert "4B Direct-Token: w=0.52, T=1.20" in text + assert "27B Pointer: w=0.54, T=1.40" in text + # JevBench producer values: Pointer native/tuned and Direct native/tuned. + for accuracy in ("51.95", "52.81", "52.38", "53.25"): + assert accuracy in text + + +def test_appendix_and_in_place_merge_are_deterministic_and_replace_old_pages( + appendix_inputs, tmp_path +): + external, chart, choice = appendix_inputs + appendix_a = tmp_path / "appendices-a.pdf" + appendix_b = tmp_path / "appendices-b.pdf" + base = tmp_path / "base.pdf" + merged_a = tmp_path / "merged-a.pdf" + merged_b = tmp_path / "merged-b.pdf" + inplace = tmp_path / "inplace.pdf" + build_appendix(external, chart, choice, appendix_a) + build_appendix(external, chart, choice, appendix_b) + assert appendix_a.read_bytes() == appendix_b.read_bytes() + appendix_metadata = PdfReader(appendix_a).metadata + assert appendix_metadata["/CreationDate"] == "D:20261002000000Z" + assert appendix_metadata["/ModDate"] == "D:20261002000000Z" + assert appendix_metadata["/Title"] == "JevAny Technical Report โ€” Evaluation Appendices" + + _base_report(base) + original = PdfReader(base) + original_invariants = [page_invariants(original.pages[index]) for index in range(19)] + assert "UNIQUE BASE PAGE 20" in original.pages[19].extract_text() + assert "UNIQUE BASE PAGE 21" in original.pages[20].extract_text() + merge_report(base, appendix_a, merged_a, base_pages=19) + merge_report(base, appendix_b, merged_b, base_pages=19) + assert merged_a.read_bytes() == merged_b.read_bytes() + + merged = PdfReader(merged_a) + assert len(merged.pages) == 23 + for index in range(19): + assert f"UNIQUE BASE PAGE {index + 1:02d}" in merged.pages[index].extract_text() + assert page_invariants(merged.pages[index]) == original_invariants[index] + assert _annotation_uris(merged.pages[index]) == [ + f"https://example.test/base/{index + 1}" + ] + generated_text = [merged.pages[index].extract_text() for index in range(19, 23)] + assert all("UNIQUE BASE PAGE 20" not in text for text in generated_text) + assert all("UNIQUE BASE PAGE 21" not in text for text in generated_text) + assert "APPENDIX L" in generated_text[0] + assert "APPENDIX L" in generated_text[1] + assert "APPENDIX M" in generated_text[2] + assert "APPENDIX M" in generated_text[3] + assert "Training-free choice-token readout" in generated_text[2] + assert "Choice-token accuracy matrix" in generated_text[3] + merged_metadata = merged.metadata + assert merged_metadata["/CreationDate"] == "D:20261002000000Z" + assert merged_metadata["/ModDate"] == "D:20261002000000Z" + assert merged_metadata["/Title"] == ( + "JevAny: Toward General Decision Intelligence โ€” merged technical report" + ) + + shutil.copyfile(base, inplace) + merge_report(inplace, appendix_a, inplace, base_pages=19) + first_inplace = inplace.read_bytes() + assert len(PdfReader(inplace).pages) == 23 + merge_report(inplace, appendix_a, inplace, base_pages=19) + + assert inplace.read_bytes() == first_inplace == merged_a.read_bytes() + + +def test_merge_report_rejects_any_base_page_count_other_than_19( + appendix_inputs, tmp_path +): + external, chart, choice = appendix_inputs + appendix = tmp_path / "appendices.pdf" + base = tmp_path / "base.pdf" + build_appendix(external, chart, choice, appendix) + writer = PdfWriter() + writer.add_blank_page(width=612, height=792) + with base.open("wb") as stream: + writer.write(stream) + + with pytest.raises(ValueError, match="exactly 19"): + merge_report(base, appendix, tmp_path / "merged.pdf", base_pages=18) + + +def test_choice_artifact_requires_full_3220_to_be_unsupported(tmp_path): + choice = tmp_path / "choice.json" + value = _choice_artifact(choice) + value["datasets"]["jevjudge_full"]["accuracy"] = 0.5 + _write_json(choice, value) + + with pytest.raises(ValueError, match="must be explicitly unsupported"): + load_choice_artifact(choice) + + +def test_published_report_uses_one_current_27b_release(): + report = Path(__file__).resolve().parents[1] / "reports" / "JevAny_Tech_Report.pdf" + pdf = PdfReader(report) + assert len(pdf.pages) == 23 + + expected = { + 1: ("86.04", "90.04"), + 4: ("44,319", "86.04", "90.04", "0.388", "0.195", "0.026"), + 7: ("86.04",), + 8: ("44,319", "39.43", "1,261.7", "2,082"), + } + stale = ( + "85.76", "90.48", "22,160", "18.83", "602.7", "1,423", + "0.392", "0.200", "0.030", + ) + for page_number, values in expected.items(): + text = pdf.pages[page_number - 1].extract_text() or "" + assert all(value in text for value in values) + assert all(value not in text for value in stale) + compute_text = pdf.pages[7].extract_text() or "" + assert "estimated cumulative seconds-per-record timing" in compute_text + boundary_text = (pdf.pages[9].extract_text() or "").replace("-\n", "-") + assert "The merged PDF is release-synchronized." in boundary_text + assert "The original PDF is pre-" not in boundary_text + + method_text = pdf.pages[21].extract_text() or "" + assert "current 27B release checkpoint: step 44,319" in method_text + assert "matched prompt-v2/runtime reruns" in method_text diff --git a/tests/test_evaluate_choice_readout_matrix.py b/tests/test_evaluate_choice_readout_matrix.py new file mode 100644 index 0000000..4dc48c5 --- /dev/null +++ b/tests/test_evaluate_choice_readout_matrix.py @@ -0,0 +1,160 @@ +import json +import math + +import pytest + +from scripts.evaluate_choice_readout_matrix import ( + _attach_soft_gold, + _typed_record, + _typed_soft_metrics, + evaluate, + reuse_completed, +) + + +def test_typed_record_uses_ordered_argmax_when_soft_gold_is_tied(): + row = { + "id": "tie-case", + "split": "test", + "workflow": "unit", + "state": json.dumps({"context": "A tied teacher distribution"}), + "questions": json.dumps({ + "decision": { + "type": "choice", + "criteria": {"first": "First option", "second": "Second option"}, + } + }), + # The convenience label intentionally disagrees with the benchmark's + # ordered argmax rule and must not determine the converted target. + "gold": json.dumps({ + "decision": { + "label": "second", + "probabilities": {"first": 0.5, "second": 0.5}, + } + }), + } + + record = _typed_record(row) + + assert record["questions"]["decision"]["label"] == "first" + assert record["questions"]["decision"]["target"] == { + "first": 0.5, + "second": 0.5, + } + + +def test_typed_soft_metrics_normalize_distributions_and_use_hard_labels(): + rows = [ + {"p": [8.0, 2.0], "gold": [3.0, 1.0], "label": 0}, + {"p": [4.0, 6.0], "gold": [1.0, 1.0], "label": 0}, + ] + + result = _typed_soft_metrics(rows, bins=2) + + expected_kl = ( + 0.75 * math.log(0.75 / 0.8) + + 0.25 * math.log(0.25 / 0.2) + + 0.5 * math.log(0.5 / 0.4) + + 0.5 * math.log(0.5 / 0.6) + ) / 2 + assert result["n"] == 2 + assert result["correct"] == 1 + assert result["accuracy"] == pytest.approx(0.5) + assert result["kl_from_gold"] == pytest.approx(expected_kl) + assert result["brier"] == pytest.approx(0.0125) + assert result["ece"] == pytest.approx(0.2) + assert result["mean_confidence"] == pytest.approx(0.7) + + +def test_attach_soft_gold_aligns_values_to_each_prediction_rows_key_order(): + record = { + "questions": { + "choice": {"target": {"alpha": 0.1, "beta": 0.9}}, + "score": {"target": {"0": 0.2, "1": 0.3, "2": 0.5}}, + } + } + rows = [ + {"question": "choice", "keys": ["beta", "alpha"]}, + {"question": "score", "keys": ["2", "0", "1"]}, + ] + + _attach_soft_gold(rows, record) + + assert rows[0]["gold"] == [0.9, 0.1] + assert rows[1]["gold"] == [0.5, 0.2, 0.3] + + +def test_evaluate_base_only_result_saves_choice_component(tmp_path): + record = { + "state": "Choose the right option.", + "questions": { + "decision": { + "type": "choice", + "criteria": {"wrong": "Wrong", "right": "Right"}, + "label": "right", + "target": {"right": 0.8, "wrong": 0.2}, + "src": "unit/base-only", + } + }, + "_meta": { + "id": "base-only", + "group_id": "base-only", + "source": "unit", + "variant": "clean", + }, + } + + def predictor(_record): + return { + # Reverse insertion order relative to criteria to exercise keyed + # probability alignment in the full evaluation path. + "probabilities": {"decision": {"right": 0.9, "wrong": 0.1}}, + "latency_ms": 2.0, + "input_tokens": 12, + } + + destination = tmp_path / "base" + report = evaluate("typed_test", [record], predictor, destination) + + assert set(report["components"]) == {"choice"} + assert report["components"]["choice"]["clean"]["acc"] == pytest.approx(1.0) + assert not (destination / "native-rows.json").exists() + rows = json.loads((destination / "choice-rows.json").read_text()) + assert rows[0]["keys"] == ["wrong", "right"] + assert rows[0]["p"] == pytest.approx([0.1, 0.9]) + assert rows[0]["gold"] == pytest.approx([0.2, 0.8]) + prediction = json.loads( + (destination / "predictions.jsonl").read_text().strip() + ) + assert set(prediction["component_probabilities"]) == {"choice"} + + +def test_reuse_completed_validates_report_and_component_rows(tmp_path): + destination = tmp_path / "typed_test" + record = { + "state": "state", + "questions": {"q": { + "type": "choice", "criteria": {"a": None, "b": None}, + "label": "a", "target": {"a": .75, "b": .25}, "src": "unit/reuse", + }}, + "_meta": { + "id": "reuse", "group_id": "reuse", "source": "unit", "variant": "clean", + }, + } + + def predictor(_record): + return { + "probabilities": {"q": {"a": .75, "b": .25}}, + "latency_ms": 1.0, + "input_tokens": 4, + "efficient_long_context_attention_used": False, + } + + expected = evaluate("typed_test", [record], predictor, destination) + owner = type("Predictor", (), {"return_components": False})() + + assert reuse_completed("typed_test", [record], owner, destination) == expected + + (destination / "choice-rows.json").unlink() + with pytest.raises(RuntimeError, match="missing or incomplete"): + reuse_completed("typed_test", [record], owner, destination) diff --git a/tests/test_fit_choice_ensemble.py b/tests/test_fit_choice_ensemble.py new file mode 100644 index 0000000..4e54d9a --- /dev/null +++ b/tests/test_fit_choice_ensemble.py @@ -0,0 +1,65 @@ +import pytest + +from scripts.fit_choice_ensemble import ( + fit, + headline_indices, + metrics, + pooled_logits, + probabilities, +) + + +def row(identity, p, label, gold=None): + value = { + "id": identity, + "question": "decision", + "keys": ["left", "right"], + "label": label, + "p": p, + } + if gold is not None: + value["gold"] = gold + return value + + +def test_fit_selects_native_weight_on_calibration_accuracy_before_nll(): + choice = [row("a", [.9, .1], 1), row("b", [.8, .2], 1)] + native = [row("a", [.1, .9], 1), row("b", [.2, .8], 1)] + + result = fit(choice, choice, native) + + assert result["selected"]["correct"] == 2 + assert result["selected"]["native_weight"] > .5 + assert len(result["grid"]) == 101 + + +def test_soft_gold_metrics_use_ordered_labels_and_report_kl(): + rows = [row("a", [.75, .25], 0, gold=[.5, .5])] + ps = probabilities(pooled_logits(rows, None, 0), 1.0) + + result = metrics(rows, ps) + + assert result["accuracy"] == 1.0 + assert result["brier"] == pytest.approx(.125) + assert result["kl_from_gold"] == pytest.approx( + result["cross_entropy_from_gold"] - result["gold_entropy"] + ) + + +def test_choice_only_fit_has_no_native_weight(): + rows = [row("a", [.7, .3], 0), row("b", [.4, .6], 1)] + + result = fit(rows, rows, None) + + assert result["selected"]["native_weight"] == 0.0 + assert len(result["grid"]) == 1 + + +def test_transfer_selection_uses_only_clean_knowable_headline_rows(): + rows = [ + {**row("headline", [.7, .3], 0), "variant": "clean", "source": "task"}, + {**row("variant", [.7, .3], 0), "variant": "permuted", "source": "task"}, + {**row("unknown", [.7, .3], 0), "variant": "clean", "source": "unknowable"}, + ] + + assert headline_indices(rows) == [0] diff --git a/tests/test_letter_predictor.py b/tests/test_letter_predictor.py index 85cb3d3..19af8e5 100644 --- a/tests/test_letter_predictor.py +++ b/tests/test_letter_predictor.py @@ -26,6 +26,25 @@ def test_letter_predictor_rejects_native_only_load_options_before_loading(): ) +def test_letter_predictor_validates_generalized_native_arguments_before_loading(): + with pytest.raises(ValueError, match="at most one"): + LetterReadoutPredictor( + checkpoint="unused", device="cpu", native_weight=0.5, pointer_weight=0.5, + ) + with pytest.raises(ValueError, match="requires a JevAny checkpoint"): + LetterReadoutPredictor(base="unused", device="cpu", native_weight=0.5) + with pytest.raises(TypeError, match="return_components"): + LetterReadoutPredictor(checkpoint="unused", device="cpu", return_components=1) + with pytest.raises(TypeError, match="exact_kernels"): + LetterReadoutPredictor(checkpoint="unused", device="cpu", exact_kernels=1) + with pytest.raises(ValueError, match="pointer weight"): + LetterReadoutPredictor(checkpoint="unused", device="cpu", native_weight=1.01) + with pytest.raises(ValueError, match="efficient_long_context_tokens"): + LetterReadoutPredictor( + checkpoint="unused", device="cpu", efficient_long_context_tokens=1, + ) + + def test_one_row_input_ids_normalizes_template_shapes_and_mappings(): assert _one_row_input_ids(torch.tensor([1, 2])).tolist() == [[1, 2]] assert _one_row_input_ids({"input_ids": [[3, 4]]}).tolist() == [[3, 4]] @@ -90,6 +109,16 @@ def test_safe_chat_text_escapes_gemma_and_bracket_style_controls(): assert "๏ผปINST]" in escaped +def test_long_context_attention_threshold_is_cuda_only(): + predictor = object.__new__(LetterReadoutPredictor) + predictor.efficient_long_context_tokens = 16 + predictor.device = "cpu" + assert not predictor._uses_efficient_attention(16) + predictor.device = "cuda" + assert not predictor._uses_efficient_attention(15) + assert predictor._uses_efficient_attention(16) + + def test_release_unused_output_head_drops_pointer_adapter_reference(): head = object() model = type("Model", (), { @@ -113,6 +142,56 @@ def test_release_unused_output_head_drops_lm_token_aliases(): assert _release_unused_output_head(direct_token) is False +def test_release_unused_output_head_retains_direct_token_native_head(): + head = object() + direct_token = type("Model", (), { + "adapter": type("Adapter", (), {"_output_embeddings": head})(), + "lm_head": head, + })() + + assert _release_unused_output_head(direct_token, retain_native=True) is False + assert direct_token.lm_head is head + assert direct_token.adapter._output_embeddings is head + + +def test_return_components_exposes_unblended_choice_and_native_probabilities(): + predictor = object.__new__(LetterReadoutPredictor) + predictor.device = "cpu" + predictor.temperature = 1.0 + predictor.native_weight = 0.5 + predictor.pointer_weight = 0.5 + predictor.return_components = True + predictor._needs_native = True + predictor.native_decision_mode = "lm_token" + predictor.provenance = { + "method": "exact option-letter choice projection", + "adapter_applied": True, + } + predictor._letter_question = lambda state, question: ( + ["left", "right"], [0.8, 0.2], [4.0, 1.0], 11, + ) + predictor._native = lambda record: ([[0.2, 0.8]], 7) + record = { + "state": "state", + "questions": {"q": { + "type": "choice", + "instructions": "choose", + "criteria": {"left": "Left", "right": "Right"}, + "label": "left", + }}, + } + + result = predictor(record) + + assert result["component_probabilities"] == { + "choice": {"q": {"left": 0.8, "right": 0.2}}, + "native": {"q": {"left": 0.2, "right": 0.8}}, + } + assert result["probabilities"]["q"] == pytest.approx({"left": 0.5, "right": 0.5}) + assert result["input_tokens"] == 18 + assert result["readout"]["native_decision_mode"] == "lm_token" + + def test_question_options_preserve_choice_and_score_order_and_fix_noul_order(): assert question_options({ "type": "choice", diff --git a/tests/test_letter_readout.py b/tests/test_letter_readout.py index 9553ee9..c3dd823 100644 --- a/tests/test_letter_readout.py +++ b/tests/test_letter_readout.py @@ -4,6 +4,8 @@ import pytest from jevany.letter_readout import ( + CHOICE_SYMBOLS, + CYGNET_LETTERS, LETTERS, SYSTEM_PROMPT, build_prompt, @@ -45,21 +47,26 @@ def test_build_prompt_matches_cygnet_structured_rendering_and_option_order(): "A. Refund the card", "B. Issue store credit", "", - "Answer with the letter of exactly one option, and nothing else:", + "Answer with the case-sensitive ID of exactly one option, and nothing else:", ]) assert prompt == expected assert prompt.index("A. Refund") < prompt.index("B. Issue") assert "thinking" not in prompt.lower() - assert "LETTER and nothing else" in SYSTEM_PROMPT + assert "case-sensitive ID and nothing else" in SYSTEM_PROMPT -def test_prompt_accepts_all_letters_and_rejects_invalid_option_counts(): - prompt = build_prompt("state", "choose", [f"option {index}" for index in range(26)]) - assert f"{LETTERS[-1]}. option 25" in prompt +def test_prompt_preserves_cygnet_az_then_extends_to_lowercase_ids(): + prompt = build_prompt( + "state", "choose", [f"option {index}" for index in range(len(CHOICE_SYMBOLS))] + ) + assert LETTERS == CHOICE_SYMBOLS + assert f"{CYGNET_LETTERS[-1]}. option 25" in prompt + assert f"{CHOICE_SYMBOLS[26]}. option 26" in prompt + assert f"{CHOICE_SYMBOLS[-1]}. option 51" in prompt with pytest.raises(ValueError, match="at least one"): build_prompt("state", "choose", []) - with pytest.raises(ValueError, match="at most 26"): - build_prompt("state", "choose", list(range(27))) + with pytest.raises(ValueError, match="at most 52"): + build_prompt("state", "choose", list(range(53))) with pytest.raises(TypeError, match="ordered mapping"): build_prompt("state", "choose", "not an option sequence") @@ -161,8 +168,8 @@ def test_invalid_probability_distributions_are_rejected(probabilities): def test_invalid_letters_token_maps_and_logits_are_rejected(): tokenizer = FakeTokenizer(["A", "not B"]) - with pytest.raises(ValueError, match="only single uppercase"): - letter_token_ids(tokenizer, ("A", "b")) + with pytest.raises(ValueError, match="single A-Z/a-z"): + letter_token_ids(tokenizer, ("A", "1")) with pytest.raises(ValueError, match="duplicates"): letter_token_ids(tokenizer, ("A", "A")) with pytest.raises(ValueError, match="no one-token aliases for: B"): diff --git a/tests/test_playground_workflow.py b/tests/test_playground_workflow.py index 46cbf36..9058463 100644 --- a/tests/test_playground_workflow.py +++ b/tests/test_playground_workflow.py @@ -140,12 +140,19 @@ def test_connecting_reports_identity_and_capabilities_from_v1_models(app, endpoi app.set_images(True) -def test_playground_displays_letter_deployment_readout(app, endpoint): +def test_playground_displays_choice_deployment_readout(app, endpoint): + endpoint.models[0]["readout"] = "choice" + assert app.connect(endpoint.url)["served"]["readout"] == "choice" + source = (Path(__file__).parents[1] / "jevany" / "demos" / "static" / "app.js").read_text() + assert 'served.readout === "choice" || served.readout === "letter"' in source + assert 'readout ? `${readout} readout`' in source + + +def test_playground_labels_legacy_letter_descriptor_as_choice(app, endpoint): endpoint.models[0]["readout"] = "letter" assert app.connect(endpoint.url)["served"]["readout"] == "letter" source = (Path(__file__).parents[1] / "jevany" / "demos" / "static" / "app.js").read_text() - assert 'served.readout === "letter" ? "letter"' in source - assert 'readout ? `${readout} readout`' in source + assert 'served.readout === "choice" || served.readout === "letter"' in source def test_playground_accepts_older_model_descriptors_without_readout(app, endpoint): diff --git a/tests/test_readout_integration.py b/tests/test_readout_integration.py index 6f8efcf..dee9c36 100644 --- a/tests/test_readout_integration.py +++ b/tests/test_readout_integration.py @@ -5,12 +5,14 @@ import pytest +from jevany import ChoiceReadoutOptions from jevany.api import Choice, Noul, Score, SystemOneRequest from jevany.inference import InferenceOptions from jevany.letter_runtime import LetterDecisionRuntime, _unlabelled_record from jevany.readout import ( LetterReadoutOptions, add_readout_arguments, + choice_options_from_args, letter_options_from_args, ) @@ -34,15 +36,18 @@ def __init__(self): }, device_map=None, devices=["cpu"], backbone_adapter="test", temperature=1.2, ) + self.native_model = self.pointer_model self.device = "cpu" self.temperature = 1.0 self.pointer_weight = 0.25 + self.native_weight = 0.25 self.max_tokens = 2048 self.effective_max_tokens = 2048 self.tokenizer = FakeTokenizer() self.provenance = { "method": "exact option-letter alias projection", "adapter_applied": True, + "native_weight": 0.25, "pointer_weight": 0.25, } self.last_record = None @@ -72,25 +77,38 @@ def decision_request(): ) -def test_letter_cli_options_are_shared_and_native_rejects_letter_flags(): +def test_choice_cli_is_public_and_legacy_letter_flags_remain_compatible(): parser = argparse.ArgumentParser() add_readout_arguments(parser) - assert letter_options_from_args(parser.parse_args([])) is None - actual = letter_options_from_args(parser.parse_args([ + help_text = parser.format_help() + assert "--readout {native,choice}" in help_text + assert "--choice-native-weight" in help_text + assert "--letter-temperature" not in help_text + assert choice_options_from_args(parser.parse_args([])) is None + actual = choice_options_from_args(parser.parse_args([ + "--readout", "choice", "--choice-temperature", "1.5", + "--choice-native-weight", "0.25", "--choice-max-tokens", "2048", + ])) + assert actual == ChoiceReadoutOptions(temperature=1.5, native_weight=0.25, max_tokens=2048) + legacy = letter_options_from_args(parser.parse_args([ "--readout", "letter", "--letter-temperature", "1.5", "--letter-pointer-weight", "0.25", "--letter-max-tokens", "2048", ])) - assert actual == LetterReadoutOptions(temperature=1.5, pointer_weight=0.25, max_tokens=2048) - with pytest.raises(ValueError, match="require --readout letter"): - letter_options_from_args(parser.parse_args(["--letter-temperature", "2"])) + assert legacy == LetterReadoutOptions(temperature=1.5, pointer_weight=0.25, max_tokens=2048) + with pytest.raises(ValueError, match="require --readout choice"): + choice_options_from_args(parser.parse_args(["--letter-temperature", "2"])) + with pytest.raises(ValueError, match="at most one"): + choice_options_from_args(parser.parse_args([ + "--readout", "choice", "--choice-temperature", "1", "--letter-temperature", "1", + ])) for arguments, message in [ - (["--readout", "letter", "--letter-temperature", "nan"], "finite"), - (["--readout", "letter", "--letter-temperature", "0"], "positive"), - (["--readout", "letter", "--letter-pointer-weight", "1.1"], r"\[0, 1\]"), - (["--readout", "letter", "--letter-max-tokens", "1"], ">= 2"), + (["--readout", "choice", "--choice-temperature", "nan"], "finite"), + (["--readout", "choice", "--choice-temperature", "0"], "positive"), + (["--readout", "choice", "--choice-native-weight", "1.1"], r"\[0, 1\]"), + (["--readout", "choice", "--choice-max-tokens", "1"], ">= 2"), ]: with pytest.raises(ValueError, match=message): - letter_options_from_args(parser.parse_args(arguments)) + choice_options_from_args(parser.parse_args(arguments)) def test_letter_runtime_maps_requests_and_reports_effective_readout(decision_request): @@ -107,12 +125,14 @@ def test_letter_runtime_maps_requests_and_reports_effective_readout(decision_req assert predictor.last_record["state"] == {"signal": "green"} assert [q["label"] for q in predictor.last_record["questions"].values()] == ["left", False, 0] description = runtime.describe() - assert description["readout"] == "letter" - assert description["letter_readout"]["pointer_weight"] == 0.25 - assert description["letter_readout"]["pointer_temperature"] == 1.2 + assert description["readout"] == "choice" + assert description["choice_readout"]["native_weight"] == 0.25 + assert description["choice_readout"]["pointer_weight"] == 0.25 + assert description["choice_readout"]["native_temperature"] == 1.2 + assert description["letter_readout"] == description["choice_readout"] assert description["limits"] == { "state_tokens": 2048, "branch_tokens": 2048, - "packed_tokens": 2048, "choices": 26, + "packed_tokens": 2048, "choices": 52, } assert description["capabilities"]["media_types"] == [] assert not description["prefix_cache"]["enabled"] @@ -137,7 +157,7 @@ def test_unlabelled_record_adds_only_encoder_placeholders(decision_request): assert record["questions"]["score"]["label"] == 0 -def test_public_loader_dispatches_to_letter_predictor(monkeypatch): +def test_public_loader_dispatches_choice_settings_to_predictor(monkeypatch): from jevany import letter_predictor from jevany.runtime import JevModel @@ -152,34 +172,61 @@ def load(**kwargs): monkeypatch.setattr(letter_predictor, "LetterReadoutPredictor", load) local = JevModel.from_pretrained( "owner/checkpoint", device="cpu", model_name="letter-model", - readout="letter", letter_temperature=1.5, - letter_pointer_weight=0.25, letter_max_tokens=2048, + readout="choice", choice_temperature=1.5, + choice_native_weight=0.25, choice_max_tokens=2048, ) assert isinstance(local.runtime, LetterDecisionRuntime) assert seen["checkpoint"] == "owner/checkpoint" assert seen["temperature"] == 1.5 - assert seen["pointer_weight"] == 0.25 + assert seen["native_weight"] == 0.25 assert seen["max_tokens"] == 2048 - with pytest.raises(ValueError, match="require readout='letter'"): - JevModel.from_pretrained("unused", device="cpu", letter_temperature=2) + with pytest.raises(ValueError, match="require --readout choice"): + JevModel.from_pretrained("unused", device="cpu", choice_temperature=2) with pytest.raises(ValueError, match="only to native readout"): JevModel.from_pretrained( - "unused", device="cpu", readout="letter", + "unused", device="cpu", readout="choice", inference_options=InferenceOptions(), ) +def test_public_loader_accepts_legacy_letter_api(monkeypatch): + from jevany import letter_predictor + from jevany.runtime import JevModel + + predictor = FakeLetterPredictor() + predictor.checkpoint.path = "/tmp/checkpoint" + seen = {} + + def load(**kwargs): + seen.update(kwargs) + return predictor + + monkeypatch.setattr(letter_predictor, "LetterReadoutPredictor", load) + local = JevModel.from_pretrained( + "owner/checkpoint", device="cpu", model_name="letter-model", + readout="letter", letter_temperature=1.5, + letter_pointer_weight=0.25, letter_max_tokens=2048, + ) + + assert isinstance(local.runtime, LetterDecisionRuntime) + assert seen["temperature"] == 1.5 + assert seen["native_weight"] == 0.25 + assert seen["max_tokens"] == 2048 + + def test_create_app_rejects_cross_readout_settings(): pytest.importorskip("fastapi") from jevany.serve import create_app - with pytest.raises(ValueError, match="require readout='letter'"): - create_app(readout="native", letter_options=LetterReadoutOptions()) + with pytest.raises(ValueError, match="require readout='choice'"): + create_app(readout="native", choice_options=ChoiceReadoutOptions()) with pytest.raises(ValueError, match="only to native readout"): - create_app(readout="letter", inference_options=InferenceOptions()) + create_app(readout="choice", inference_options=InferenceOptions()) + # The old programmatic spelling still builds the same application path. + assert create_app(readout="letter", letter_options=LetterReadoutOptions()) -def test_decide_cli_forwards_letter_settings(tmp_path, monkeypatch, capsys): +def test_decide_cli_forwards_choice_settings(tmp_path, monkeypatch, capsys): from jevany.cli import decide_main from jevany.runtime import JevModel @@ -207,21 +254,21 @@ def load(*args, **kwargs): monkeypatch.setattr(JevModel, "from_pretrained", load) decide_main([ str(source), "--checkpoint", "owner/checkpoint", "--device", "cpu", - "--readout", "letter", "--letter-temperature", "1.5", - "--letter-pointer-weight", "0.25", "--letter-max-tokens", "2048", + "--readout", "choice", "--choice-temperature", "1.5", + "--choice-native-weight", "0.25", "--choice-max-tokens", "2048", ]) assert json.loads(capsys.readouterr().out)["answers"]["choice"]["choice"] == "right" assert seen["args"] == ("owner/checkpoint",) - assert seen["kwargs"]["readout"] == "letter" - assert seen["kwargs"]["letter_temperature"] == 1.5 - assert seen["kwargs"]["letter_pointer_weight"] == 0.25 - assert seen["kwargs"]["letter_max_tokens"] == 2048 + assert seen["kwargs"]["readout"] == "choice" + assert seen["kwargs"]["choice_temperature"] == 1.5 + assert seen["kwargs"]["choice_native_weight"] == 0.25 + assert seen["kwargs"]["choice_max_tokens"] == 2048 assert "inference_options" not in seen["kwargs"] with pytest.raises(SystemExit): decide_main([str(source), "--letter-temperature", "2"]) -def test_eval_cli_selects_letter_predictor_and_records_provenance(tmp_path, monkeypatch, capsys): +def test_eval_cli_selects_choice_predictor_and_records_provenance(tmp_path, monkeypatch, capsys): from jevany import benchmark, letter_predictor data = tmp_path / "data.jsonl" @@ -237,8 +284,9 @@ def test_eval_cli_selects_letter_predictor_and_records_provenance(tmp_path, monk class Predictor: temperature = 1.5 + native_weight = 0.25 pointer_weight = 0.25 - pointer_model = SimpleNamespace(temperature=1.2) + native_model = SimpleNamespace(temperature=1.2) provenance = {"method": "exact option-letter alias projection"} def __init__(self, **kwargs): @@ -255,18 +303,20 @@ def evaluate(records, predictor, directory, **kwargs): monkeypatch.setattr(letter_predictor, "LetterReadoutPredictor", Predictor) benchmark.main([ "--run", "owner/checkpoint", "--data", str(data), "--out", str(output), - "--device", "cpu", "--readout", "letter", "--letter-temperature", "1.5", - "--letter-pointer-weight", "0.25", "--letter-max-tokens", "2048", + "--device", "cpu", "--readout", "choice", "--choice-temperature", "1.5", + "--choice-native-weight", "0.25", "--choice-max-tokens", "2048", ]) assert json.loads(capsys.readouterr().out)["objective"] == 0.0 assert seen["checkpoint"] == "owner/checkpoint" assert seen["temperature"] == 1.5 - assert seen["pointer_weight"] == 0.25 + assert seen["native_weight"] == 0.25 assert seen["max_tokens"] == 2048 report = json.loads((output / "report.json").read_text()) - assert report["readout"] == "letter" - assert report["letter_readout"] == { - **Predictor.provenance, "pointer_temperature": 1.2, + assert report["readout"] == "choice" + assert report["choice_readout"] == { + **Predictor.provenance, "native_temperature": 1.2, "pointer_temperature": 1.2, } + assert report["letter_readout"] == report["choice_readout"] assert report["calibration_applied"] is True + assert report["calibration"]["native_temperature"] == 1.2 assert report["calibration"]["pointer_temperature"] == 1.2