diff --git a/.github/workflows/ci.yml b/.github/workflows/ci.yml
index ecdcd3e..985bb4d 100644
--- a/.github/workflows/ci.yml
+++ b/.github/workflows/ci.yml
@@ -64,3 +64,27 @@ jobs:
path: |
ci-results.xml
ci-dependencies.txt
+
+ report-test:
+ name: report appendix
+ runs-on: ubuntu-24.04
+ steps:
+ - uses: actions/checkout@v4
+ - uses: actions/setup-python@v5
+ with:
+ python-version: '3.13'
+ cache: pip
+ cache-dependency-path: reports/requirements.txt
+ - name: Install report dependencies
+ run: python -m pip install --disable-pip-version-check -r reports/requirements.txt pytest==9.1.1
+ - name: Verify report environment
+ run: |
+ python -m pip check
+ python -c "import matplotlib, pypdf, pytest; print(matplotlib.__version__, pypdf.__version__, pytest.__version__)"
+ - name: Test report artifact and PDF builders
+ env:
+ JEVANY_REPORT_TESTS_REQUIRED: '1'
+ run: >-
+ python -m pytest -q
+ tests/test_build_choice_readout_results.py
+ tests/test_build_external_report_appendix.py
diff --git a/README.md b/README.md
index 4ca72da..c298c71 100644
--- a/README.md
+++ b/README.md
@@ -140,7 +140,7 @@ where Jev helps and when to return control to the LLM.
- [๐ ๏ธ 1.2 JevAny Training](#training)
- [๐ 1.3 JevAny Deployment](#deployment)
- [๐ค 2. Pretrained Models](#pretrained-models)
- - [Training-free letter readout](#letter-readout)
+ - [Training-free choice-token readout](#choice-readout)
- [๐ 3. Benchmark Results](#evaluation)
- [โฑ๏ธ 3.1 Inference efficiency](#efficiency)
- [๐น๏ธ 4. Examples & Test Environments](#examples--test-environments)
@@ -296,30 +296,43 @@ Pointer and direct-token models share the same API. Pointer supports up to
See [readout choices](docs/TRAINING.md#pointer-and-direct-token-readouts) for
training and accuracy tradeoffs.
-### Training-free letter readout
+### Training-free choice-token readout
-The experimental letter readout presents up to 26 options as AโZ, then sums
-the frozen language model's next-token probability mass for every vocabulary
-token that decodes exactly to that uppercase letter. It can run on a base model
-without training, apply a JevAny adapter before the same readout, or combine the
-letter and native pointer distributions. The letter path is text-only and does
-not change the checkpoint metadata. Select it with `--readout letter` in
-`jevany eval`, `jevany decide`, or `jevany serve`; the checkpoint-native
-readout remains the default.
+Choice-token readout assigns the ordered options to the exact one-token IDs
+`AโZ, aโz`, scores those 52 rows at the answer position, and renormalizes over
+the available options. It needs no readout training and works on a frozen base,
+after a JevAny Pointer or Direct-Token adapter, or in a log-linear stack with
+the checkpoint-native distribution. It can change the ranking and the selected
+answer; scalar temperature calibration only changes confidence.
-[Method, limits and evaluation command](docs/LETTER_READOUT.md).
+```bash
+jevany eval --run SimpleJev/JevAny-Qwen3.5-4B-Direct-Token-LoRA \
+ --suite /path/to/suite --out runs/choice --device cuda --readout choice
+```
+
+[Method, commands, limits and complete tables](docs/CHOICE_READOUT.md) ยท
+[Machine-readable results](results/choice-readout-v2.json)
+
+[](docs/CHOICE_READOUT.md#results)
+
+- Direct-Token 4B's Transfer-dev-tuned stack reaches **79.83%** on held-out
+ Transfer, **67.65%** on Typed Decisions, and **59.25%** on JevJudge text:
+ +0.96, +0.45, and +0.83 points over its native readout.
+- Pointer 27B's tuned stack reaches **89.10%** on held-out Transfer and
+ **73.30%** on Typed, but drops from **66.44% to 64.36%** on JevJudge text.
+ The target-domain result decides whether stacking is useful.
+- The frozen 4B choice path scores 52.75% on Typed; applying the Direct-Token
+ adapter raises the same choice path to 64.80%. On JevJudge text, the same
+ adapter slightly lowers it, from 58.70% to 57.60%.
-[](docs/LETTER_READOUT.md#results)
+**Practical rule**
-- The fixed 50/50 blend improved 27B on both diagnostics: +0.43 accuracy
- points on JevBench public and +0.67 on Transfer-v9. Neither gain is
- statistically significant.
-- At 4B, the same blend gained +1.73 points on JevBench public but lost 0.10
- on Transfer-v9. Letter readout is a paired target-domain ablation, not a
- universal upgrade.
-- Letter-only measured a 1.05x/1.07x median speedup over native at 4B/27B
- on one warmed H200 panel. Blending runs both paths and cost 1.81x/1.89x the
- native median latency.
+- Use choice-token readout when the model can reason over the options and the
+ bottleneck is extracting a decision; it cannot create missing task ability.
+- Blend only when paired development errors show that native and choice rescue
+ each other, then freeze one weight before the target evaluation.
+- Use temperature for probability calibration, not accuracy: it cannot change
+ the argmax.
> **Benchmark units:** Cygnet's official **73.70** on JevBench v1.5.4 is a
> four-axis composite over 1,624 open and sealed decisions, not accuracy. Its
@@ -329,6 +342,7 @@ readout remains the default.
## ๐ 3. Benchmark Results
+The release table below uses each checkpoint's native readout.
JevAny-Qwen3.8-27B leads both benchmarks and has the lowest NLL and Brier.
Among 4B releases, direct-token leads on JevBench; pointer leads on Transfer.
@@ -357,7 +371,8 @@ NLL, Brier and ECE are measured on Transfer.
[Machine-readable results](results/model-family-v2.json) ยท
[Method and ablation report](reports/JevAny_Tech_Report.pdf)
-The external comparison uses the complete Typed Decisions test split, the full
+The external comparison below also uses checkpoint-native readouts. It uses the
+complete Typed Decisions test split, the full
3,220-record JevJudge multimodal suite, and its 724-record text slice. The same
13 models stay in the same order; `โ` means unsupported native input or no
matching result.
diff --git a/docs/CHOICE_READOUT.md b/docs/CHOICE_READOUT.md
new file mode 100644
index 0000000..7444be9
--- /dev/null
+++ b/docs/CHOICE_READOUT.md
@@ -0,0 +1,216 @@
+# Training-free choice-token readout
+
+Choice-token readout turns an ordinary causal language model into a decision
+readout without adding or training a head:
+
+1. Preserve the candidate order and assign the exact one-character IDs
+ `AโZ, aโz`.
+2. Ask for one case-sensitive option ID with thinking disabled.
+3. At the answer position, sum the mass of every vocabulary token that decodes
+ exactly to each available ID.
+4. Renormalize over the available IDs and map the distribution back to the
+ original option keys.
+
+This operation can change the ranking and the selected answer. It is therefore
+not temperature calibration. A scalar temperature changes probability values
+but cannot change argmax accuracy. A native + choice stack is a third operation:
+it pools two distributions and can change the answer.
+
+The same readout works on a frozen base model or after loading a JevAny Pointer
+or Direct-Token adapter. No additional readout training is performed.
+
+## Use it
+
+The checkpoint-native readout remains the default. Select choice-token readout
+with `--readout choice`:
+
+```bash
+jevany eval \
+ --run SimpleJev/JevAny-Qwen3.5-4B-Direct-Token-LoRA \
+ --suite /path/to/jevbench-public-v1.4.2.2 \
+ --out runs/choice-readout/qwen35-4b-direct \
+ --device cuda --readout choice --choice-temperature 1.0
+
+jevany decide request.json \
+ --checkpoint SimpleJev/JevAny-Qwen3.5-4B-LoRA \
+ --device cuda --dtype bf16 --readout choice
+
+jevany serve \
+ --checkpoint SimpleJev/JevAny-Qwen3.5-4B-LoRA \
+ --device cuda --dtype bf16 --readout choice
+```
+
+Use `--choice-native-weight` to pool the choice and checkpoint-native
+distributions. This executes both paths. `--choice-max-tokens` controls the
+choice prompt limit. The former `letter` selector and `--letter-*` flags remain
+hidden compatibility aliases.
+
+The matrix evaluator also accepts a frozen base directly:
+
+```bash
+python -m scripts.evaluate_choice_readout_matrix \
+ --base Qwen/Qwen3.5-4B \
+ --revision 851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a \
+ --base-load-path /path/to/Qwen3.5-4B \
+ --model-name Frozen-Qwen3.5-4B \
+ --out runs/choice-readout/base4 \
+ --jevbench /path/to/jevbench-public-v1.4.2.2 \
+ --transfer /path/to/transfer-v9 \
+ --typed /path/to/typed-decisions \
+ --jevjudge /path/to/jevjudge-public
+```
+
+## Limits
+
+- Text input only; JevJudge full multimodal is unsupported and is never
+ converted into a text-only score.
+- 1โ52 options per question.
+- Exact one-token aliases must exist in the tokenizer for every used ID.
+- Choice-only runs use one chat-formatted prefill per question. A stack also
+ executes the checkpoint-native prefill.
+- Matrix runs use exact math SDPA below 8,192 tokens and memory-linear fused
+ SDPA for longer prompts. All 724 JevJudge text records complete; no record is
+ truncated or rejected.
+
+## Evaluation protocol
+
+The v2 matrix uses the same prompt and data for five runs:
+
+- frozen Qwen3.5-4B;
+- JevAny Qwen3.5-4B Pointer;
+- JevAny Qwen3.5-4B Direct-Token;
+- frozen Qwen3.8-27B;
+- JevAny Qwen3.8-27B Pointer.
+
+Zero-shot rows use T=1 and no evaluation-label tuning. For tuned rows, each
+model selects one native weight on the 1,046 clean/knowable Transfer-v9
+development decisions; fitted NLL and distance from 0.5 break accuracy ties.
+One additional temperature is then fitted on the same development rows. The
+weight and temperature are frozen before Transfer test, Typed Decisions,
+JevJudge text, and the JevBench public diagnostic are scored.
+
+The native parity gates reproduce the released README results before any
+held-out panel runs:
+
+| Checkpoint | Transfer development | JevBench public |
+|---|---:|---:|
+| JevAny 4B Pointer | 823/1,046 ยท 78.68% | 185/231 ยท 80.09% |
+| JevAny 4B Direct-Token | 818/1,046 ยท 78.20% | 187/231 ยท 80.95% |
+| JevAny 27B Pointer | 900/1,046 ยท 86.04% | 208/231 ยท 90.04% |
+
+### Evaluated model sources
+
+| Run | Canonical public repository | Evaluated revision |
+|---|---|---|
+| Frozen Qwen3.5-4B | `Qwen/Qwen3.5-4B` | `851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a` |
+| JevAny 4B Pointer | `SimpleJev/JevAny-Qwen3.5-4B-LoRA` | `1c7aa9bab14ac347aeb917c0bcd757838a8a78ce` |
+| JevAny 4B Direct-Token | `SimpleJev/JevAny-Qwen3.5-4B-Direct-Token-LoRA` | Not verified for the evaluated local release |
+| Frozen Qwen3.8-27B | `Qwen/Qwen3.8-27B` | `1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0` |
+| JevAny 27B Pointer | `SimpleJev/JevAny-Qwen3.8-27B-LoRA` | `09c9e9102d5b8cc7d56558d25da1202a761b6c0d` |
+
+The Direct-Token run is still byte-identifiable: its evaluated adapter is
+`b768d54b9e5e20204b2db9e24c18ae6b10fa310b241d1931bac6b8882f127c53`
+and its native head is
+`d3909cffc156d077061114627c8aa22f60c7a8f1bdcfd20c489bb3bb009fc4a6`
+(SHA-256). The evaluation manifest and release metadata do not prove which
+immutable public-repository commit contains those exact files, so the result
+artifact records its public revision as `null` rather than guessing.
+
+## Results
+
+
+
+All values below are accuracy. JevBench is its 231-item public development
+diagnostic. Transfer test contains 1,046 clean/knowable held-out decisions.
+Typed covers 2,000 decisions in 400 test cases. JevJudge text covers all 724
+text records; it is not the 3,220-record multimodal suite.
+
+| Weights | Readout | JevBench public | Transfer test | Typed test | JevJudge text |
+|---|---|---:|---:|---:|---:|
+| Frozen Qwen3.5-4B | Choice T=1 | 79.65% | 68.45% | 52.75% | 58.70% |
+| JevAny 4B Pointer | Native shipped | 80.09% | 77.92% | **63.50%** | 58.01% |
+| | Choice T=1 | **81.39%** | 73.61% | 57.80% | 52.62% |
+| | Tuned stack ยท w=.78 | 80.95% | **78.01%** | 62.90% | **59.53%** |
+| JevAny 4B Direct-Token | Native shipped | 80.95% | 78.87% | 67.20% | 58.43% |
+| | Choice T=1 | **81.39%** | 78.59% | 64.80% | 57.60% |
+| | Tuned stack ยท w=.59 | 80.95% | **79.83%** | **67.65%** | **59.25%** |
+| Frozen Qwen3.8-27B | Choice T=1 | 87.88% | 80.69% | 68.25% | 61.74% |
+| JevAny 27B Pointer | Native shipped | **90.04%** | 87.95% | 72.80% | **66.44%** |
+| | Choice T=1 | 89.61% | 86.04% | 72.60% | 57.87% |
+| | Tuned stack ยท w=.52 | **90.04%** | **89.10%** | **73.30%** | 64.36% |
+
+The fixed, untuned 50/50 stacks remain a separate zero-shot reference:
+
+| Checkpoint | JevBench public | Transfer test | Typed test | JevJudge text |
+|---|---:|---:|---:|---:|
+| JevAny 4B Pointer | 81.39% | 77.44% | 60.70% | 59.12% |
+| JevAny 4B Direct-Token | 80.95% | 79.45% | 66.75% | 58.98% |
+| JevAny 27B Pointer | 90.04% | **89.20%** | 73.20% | 64.23% |
+
+The selected parameters differ by checkpoint:
+
+| Checkpoint | Native weight | Additional temperature |
+|---|---:|---:|
+| JevAny 4B Pointer | .78 | 1.10 |
+| JevAny 4B Direct-Token | .59 | 1.12 |
+| JevAny 27B Pointer | .52 | .84 |
+
+### External comparison
+
+| Model/readout | Typed test | JevJudge text |
+|---|---:|---:|
+| meraGPT Decider 1 ยท published | **76.80%** | โ |
+| JevAny 27B Pointer ยท native | 72.80% | **66.44%** |
+| JevAny 27B Pointer ยท tuned stack | **73.30%** | 64.36% |
+| TypeSafe Jev 1.13 | 72.70% | 65.06% |
+| Kev-27B | โ | 64.23% |
+
+The 27B stack beats its native readout by 12 held-out Transfer decisions and
+10 Typed decisions, while losing 15 JevJudge text decisions. Direct-Token 4B's
+stack improves over native on all three held-out panels. Pointer 4B improves on
+Transfer and JevJudge but loses on Typed. No single blend is universally best.
+
+### Practical principles
+
+- **Selection bottleneck:** use choice-token readout when the model can reason
+ over the candidates but the native extraction path is the bottleneck.
+- **Complementary errors:** keep a stack only when paired development results
+ show that each readout rescues errors made by the other. Select one weight on
+ development data and freeze it.
+- **Capability limit:** a different readout cannot supply knowledge, perception,
+ planning, or instruction-following ability missing from the underlying model.
+- **Calibration limit:** temperature can improve NLL, Brier, or ECE, but it
+ cannot change the selected option or accuracy.
+- **Domain shift:** target-domain validation wins over a global rule. The same
+ 27B stack helps Transfer and Typed, ties JevBench, and hurts JevJudge text.
+
+Complete metrics, hashes, runtime versions, selected weights, and the unsupported
+full-JevJudge marker are in
+[`results/choice-readout-v2.json`](../results/choice-readout-v2.json). The
+deterministic builder is
+[`scripts/build_choice_readout_results.py`](../scripts/build_choice_readout_results.py).
+
+## Benchmark units
+
+Cygnet's official `73.70` on JevBench v1.5.4 is a four-axis composite over
+1,624 open and sealed decisions, not accuracy. Its separate public-development
+result is 203/231 (87.9% accuracy). Every JevBench value on this page is
+accuracy over the same 231 public-development items and cannot be compared
+numerically with the composite.
+
+## Historical v1
+
+The original prompt-v1 experiment and one-panel latency diagnostic remain
+unchanged in [`results/letter-readout-v1.json`](../results/letter-readout-v1.json)
+and [`letter-readout-results.svg`](letter-readout-results.svg). Those runs used a
+different prompt and runtime, and their matched native reruns did not reproduce
+the canonical README rows. They are retained for audit and are not mixed into
+the v2 tables.
+
+## Attribution
+
+The Cygnet-compatible prompt and exact option-ID aggregation semantics are
+adapted from `blockbrain-ai/cygnet-recipe` commit `3cf591c`, Copyright 2026
+Nood Co and contributors, under MIT. Cygnet credits the one-token option-ID
+readout to [NInfer](https://github.com/igorls/ninfer), released under Apache-2.0.
+JevAny does not incorporate NInfer source code. See [`NOTICE`](../NOTICE).
diff --git a/docs/EVALUATION.md b/docs/EVALUATION.md
index b6e8999..a7db080 100644
--- a/docs/EVALUATION.md
+++ b/docs/EVALUATION.md
@@ -2,7 +2,7 @@
## Model family v2
-This section contains the full results for the five current LoRA SFT releases.
+This section contains checkpoint-native results for the five current LoRA SFT releases.
Transfer also informed model development and serves as a diagnostic comparison.
**Transfer** is the cross-domain evaluation built from
@@ -16,7 +16,7 @@ and ECE in the main table are Transfer metrics; every run covers every item.
Every JevBench value in this document is public-development accuracy
(`correct / 231`). It is not the official JevBench v1.5.4 composite, which
combines Intelligence, Calibration, Speed and Cost over 1,624 open and sealed
-decisions. See the [letter-readout guide](LETTER_READOUT.md#benchmark-units) for
+decisions. See the [choice-token readout guide](CHOICE_READOUT.md#benchmark-units) for
the side-by-side definitions.
| Model | Transfer โ | JevBench โ | NLL โ | Brier โ | ECE โ |
@@ -95,7 +95,7 @@ a complete local Transfer API run and JevBench's published per-tier accuracy.
- **Loss:** cross-entropy, pure InfoNCE, and mixed objectives; CE gave the best
Transfer accuracy in the loss sweep, while small contrastive terms mainly
improved calibration. The released checkpoints use CE.
-- **Readout family:** at 4B, direct-token improves JevBench (80.95% vs 80.09%),
+- **Trained checkpoint-native readout family:** at 4B, direct-token improves JevBench (80.95% vs 80.09%),
while pointer is slightly stronger on Transfer (78.68% vs 78.20%).
### Reproducibility
@@ -140,6 +140,35 @@ excluded. Gemma uses run/checkpoint timestamps because its earlier checkpoint
format did not store cumulative elapsed seconds; the other figures come from
checkpoint or terminal trainer telemetry.
+## Training-free choice-token readout
+
+Choice-token readout constrains the next token to one of 52 exact option IDs
+and renormalizes their mass. It needs no additional training and can change the
+selected answer. Scalar temperature fitting is separate: it changes confidence,
+not accuracy. A native + choice stack can change both.
+
+The stack weight and additional temperature below are selected only on 1,046
+clean/knowable Transfer-v9 development decisions and then frozen. Transfer test,
+Typed Decisions, and JevJudge text are held out. JevBench is a public diagnostic.
+All values are accuracy percentages.
+
+| Weights | Readout | JevBench public | Transfer test | Typed test | JevJudge text |
+|---|---|---:|---:|---:|---:|
+| Frozen Qwen3.5-4B | Choice T=1 | 79.65% | 68.45% | 52.75% | 58.70% |
+| JevAny 4B Pointer | Native / Choice / Tuned | 80.09 / 81.39 / 80.95 | 77.92 / 73.61 / **78.01** | **63.50** / 57.80 / 62.90 | 58.01 / 52.62 / **59.53** |
+| JevAny 4B Direct-Token | Native / Choice / Tuned | 80.95 / **81.39** / 80.95 | 78.87 / 78.59 / **79.83** | 67.20 / 64.80 / **67.65** | 58.43 / 57.60 / **59.25** |
+| Frozen Qwen3.8-27B | Choice T=1 | 87.88% | 80.69% | 68.25% | 61.74% |
+| JevAny 27B Pointer | Native / Choice / Tuned | **90.04** / 89.61 / **90.04** | 87.95 / 86.04 / **89.10** | 72.80 / 72.60 / **73.30** | **66.44** / 57.87 / 64.36 |
+
+The 4B Direct-Token stack improves over its native readout on all three held-out
+panels. The 27B stack improves Transfer and Typed, ties JevBench, and hurts
+JevJudge text. Retain a stack only when paired development errors are
+complementary and the target distribution validates it.
+
+[Full method, fixed 50/50 controls and protocol](CHOICE_READOUT.md) ยท
+[Machine-readable results](../results/choice-readout-v2.json) ยท
+[Merged 23-page technical report](../reports/JevAny_Tech_Report.pdf)
+
## Earlier releases and evaluations
The [JevBench and Kev comparison](EXTERNAL_EVALUATION.md) evaluates both released
diff --git a/docs/EXTERNAL_EVALUATION.md b/docs/EXTERNAL_EVALUATION.md
index 731db1d..004551c 100644
--- a/docs/EXTERNAL_EVALUATION.md
+++ b/docs/EXTERNAL_EVALUATION.md
@@ -152,9 +152,12 @@ and are not used in the figure.
[`external-zero-shot-v1.json`](../results/external-zero-shot-v1.json), and the
deterministic renderer is
[`plot_external_zero_shot.py`](../scripts/plot_external_zero_shot.py).
-- The two external-evaluation pages appended to the single merged technical
- report are generated by
- [`build_external_report_appendix.py`](../scripts/build_external_report_appendix.py).
+- Four evaluation pages are appended to the single merged technical report by
+ [`build_external_report_appendix.py`](../scripts/build_external_report_appendix.py):
+ two external-evaluation pages in Appendix L and two training-free
+ choice-token pages in Appendix M. Appendix M reads
+ `results/choice-readout-v2.json` directly; it does not rasterize or convert
+ the README SVG.
- The tracked artifact contains the aggregate rows needed to regenerate the
figure. Raw local run locations and SHA-256 hashes are recorded for audit but
are not included in the repository. JevAny full results come from
@@ -190,20 +193,25 @@ build/report-repro-env/bin/python scripts/plot_external_zero_shot.py
build/report-repro-env/bin/python scripts/build_external_report_appendix.py \
--data results/external-zero-shot-v1.json \
--chart docs/external-zero-shot.png \
- --output build/JevAny_Tech_Report_External_Evaluation_Appendix.pdf \
+ --choice-data results/choice-readout-v2.json \
+ --output build/JevAny_Tech_Report_Evaluation_Appendices.pdf \
--base-report reports/JevAny_Tech_Report.pdf \
--merged-output reports/JevAny_Tech_Report.pdf \
--base-pages 19
+cp reports/JevAny_Tech_Report.pdf site/files/reports/JevAny_Tech_Report.pdf
```
-`--base-pages 19` always takes the original report and agent-harness pages and
-discards any previously appended external pages before adding the newly built
-two-page appendix. The base and merged paths may therefore be identical without
-growing the report on repeated runs. Both outputs are written to temporary
-files in their destination directories and installed with `os.replace`; an
-interrupted build cannot leave a partial tracked PDF. Appendix and merged PDF
-metadata use the fixed creation and modification timestamp
-`D:20261001000000Z`. With the pinned inputs and dependencies, repeated commands
+`--base-pages 19` retains the report and agent-harness pages, synchronizes the
+superseded 27B release headline fields to step 44,319, and discards any
+previously appended evaluation pages before rebuilding Appendices L and M. The
+result is exactly 23 pages: 19 retained pages plus four generated pages. The
+base and merged paths may therefore be identical without growing the report on
+repeated runs. Both outputs are written to temporary files in their destination
+directories and installed with `os.replace`; an interrupted build cannot leave
+a partial tracked PDF. `reports/` contains only one PDFโthe merged report; the
+`site/files/` copy is its byte-identical deployment mirror. Appendix and merged
+PDF metadata use the fixed creation and modification timestamp
+`D:20261002000000Z`. With the pinned inputs and dependencies, repeated commands
produce byte-identical appendix and merged files.
### JevBench v1.5.4
diff --git a/docs/LETTER_READOUT.md b/docs/LETTER_READOUT.md
index 789f56c..136b12f 100644
--- a/docs/LETTER_READOUT.md
+++ b/docs/LETTER_READOUT.md
@@ -1,170 +1,7 @@
-# Training-free option-letter readout
+# Moved: training-free choice-token readout
-JevAny includes an experimental evaluator for a training-free, single-answer-slot
-decision readout. It adapts the prompt and probability-aggregation semantics
-from [Cygnet](https://github.com/blockbrain-ai/cygnet-recipe/tree/3cf591c692dec649f7c134449814610307c7bb3a):
+The public method name is **choice-token readout**. See
+[`CHOICE_READOUT.md`](CHOICE_READOUT.md) for the method, commands, limits,
+current results, and retained historical v1 links.
-1. Render the options in their original order as `A`, `B`, โฆ, `Z`.
-2. Ask the model for exactly one option letter with thinking disabled.
-3. At the answer position, collect every vocabulary token whose decoded surface
- is exactly each available uppercase letter. This emulates Cygnet's
- `structured_outputs.choice` mask; leading-space, punctuated, and lowercase
- forms are not admitted.
-4. Sum duplicate-token mass per letter and normalize over the available options.
-5. Optionally apply a temperature fitted for that model and evaluation domain.
-
-The model produces no explanation and the method needs no additional training.
-The same readout can be applied after loading a JevAny LoRA. For pointer
-checkpoints, the letter and native pointer distributions can also be combined
-log-linearly. The integrated commands call this knob
-`--letter-pointer-weight`; the standalone experiment script calls it
-`--pointer-weight`.
-
-## Run the public diagnostic
-
-Use `--readout letter` to score a JevAny checkpoint on any regular frozen suite
-or labelled JSONL. This example keeps temperature at 1 rather than copying a
-value fitted for another model:
-
-```bash
-jevany eval \
- --run SimpleJev/JevAny-Qwen3.5-4B-LoRA \
- --suite /path/to/jevbench-public-v1.4.2.2 \
- --out runs/letter-readout/qwen35-4b \
- --device cuda --readout letter --letter-temperature 1.0
-```
-
-The same selector works for one local decision or an HTTP deployment. Native
-checkpoint readout remains the default when `--readout` is omitted:
-
-```bash
-jevany decide request.json \
- --checkpoint SimpleJev/JevAny-Qwen3.5-4B-LoRA \
- --device cuda --dtype bf16 --readout letter
-
-jevany serve \
- --checkpoint SimpleJev/JevAny-Qwen3.5-4B-LoRA \
- --device cuda --dtype bf16 --readout letter
-```
-
-The standalone evaluator additionally supports a frozen base via `--base`,
-tier sampling, and explicit LoRA scaling:
-
-```bash
-python scripts/evaluate_letter_readout.py \
- --checkpoint SimpleJev/JevAny-Qwen3.5-4B-LoRA \
- --suite /path/to/jevbench-public-v1.4.2.2 \
- --out runs/letter-readout/qwen35-4b \
- --device cuda --dtype bf16 --temperature 1.0
-```
-
-Add `--sample-per-tier 1` for a three-record smoke test. Use
-`--lora-scale 0` with a checkpoint to evaluate its exact pinned base without
-the adapter. A nonzero `--pointer-weight` requires a native pointer checkpoint
-and adds a second, pointer-formatted model pass per request.
-
-Current limits:
-
-- text-only model input;
-- 1โ26 options per question;
-- safetensors weights for models with an untied language-model output head;
-- one letter-formatted prefill per question, plus one native prefill when
- pointer blending is enabled;
-- native inference-limit and CUDA-graph flags do not apply to the chat-formatted
- letter path; use `--letter-max-tokens` for its prompt limit.
-
-Temperature changes reported probabilities but does not change the letter-only
-argmax. Fit it on a separate calibration split for each model. Cygnet's `3.4`
-was fitted for its own Gemma configuration and is not a default for JevAny.
-
-## Benchmark units
-
-The official score and the public diagnostic answer different questions:
-
-| Result | Evaluation set | Unit |
-|---|---:|---|
-| Cygnet 73.70 | JevBench v1.5.4, 1,624 decisions: 904 open + 720 sealed | Equal-weight harmonic mean of Intelligence, Calibration, Speed and Cost |
-| Cygnet 203/231 (87.9%) | Public development set, 231 decisions | Accuracy |
-| JevAny results in the current README | Same public development set, 231 decisions | Accuracy |
-
-Cygnet ranks first among 106 systems on the official v1.5.4 composite and is a
-statistical tie with Winnow-12B Q8. The 73.70 composite is not `73.70%` and
-cannot be compared numerically with public-set accuracy. See Cygnet's
-[official-submission report](https://github.com/blockbrain-ai/cygnet-recipe/pull/4)
-and the [JevBench v1.5.4 board](https://benchmarkheaven.com/jev-models/v1.5.4).
-
-## Results
-
-All JevAny values below are local pinned-checkpoint runs at temperature 1. The
-complete metrics, suite hashes, report hashes, and paired counts are in
-[`results/letter-readout-v1.json`](../results/letter-readout-v1.json).
-These are separate matched reruns of the pinned released checkpoints for this
-ablation; they do not replace the main model-family table, whose frozen
-checkpoints and run provenance differ.
-
-
-
-JevBench public is a 231-question development diagnostic, not the sealed
-v1.5.4 board. `Base + letter` disables the JevAny adapter; the other columns use
-the released adapter. Deltas and two-sided exact McNemar p-values compare each
-adapter readout with the native pointer on the same questions.
-
-| Model | Base + letter | Native | Letter | Fixed 50/50 blend |
-|---|---:|---:|---:|---:|
-| Qwen3.5-4B | 79.65% | 80.09% | 81.39% (+1.30, p=.749) | **81.82%** (+1.73, p=.424) |
-| Qwen3.8-27B | 88.31% | 89.61% | 89.61% (+0.00, p=1.000) | **90.04%** (+0.43, p=1.000) |
-
-Transfer-v9 evaluates all 1,264 requests without rejection or truncation. Its
-accuracy headline uses the 1,046 clean knowable decisions.
-
-| Model | Native | Letter | Fixed 50/50 blend |
-|---|---:|---:|---:|
-| Qwen3.5-4B | **79.16%** | 75.72% (-3.44, p=.0028) | 79.06% (-0.10, p=1.000) |
-| Qwen3.8-27B | 86.23% | 84.23% (-2.01, p=.0375) | **86.90%** (+0.67, p=.337) |
-
-The 27B blend gets 909/1,046 decisions right versus 902/1,046 for native, but
-the seven-question gain is not statistically significant. The fixed blend is
-also not a universal improvement: at 4B it gets one fewer answer right. None of
-the positive gains in either table is significant at the 0.05 level. The
-letter-only Transfer-v9 losses show that a public-diagnostic gain does not by
-itself establish transfer.
-
-## Latency diagnostic
-
-One warmed run per mode on H200/BF16 measured the same heterogeneous 231-record
-panel with 16 warmups, one measured repeat, and concurrency 1. Values are external
-end-to-end median / p95 milliseconds; loading and network transport are
-excluded.
-
-| Model | Native | Letter | Fixed 50/50 blend |
-|---|---:|---:|---:|
-| Qwen3.5-4B | 138.43 / 200.98 | **131.86** / 201.55 (1.05x median speedup) | 250.34 / 388.31 (1.81x median latency) |
-| Qwen3.8-27B | 194.70 / 477.67 | **181.58** / 492.51 (1.07x median speedup) | 367.40 / 943.44 (1.89x median latency) |
-
-Letter-only has a modest median improvement on this panel despite using more
-logical input tokens on average (701 versus 602); p95 does not improve. The
-blend evaluates both prefills and averages 1,303 logical input tokens. This is
-not saturated server throughput or a general speed claim.
-
-## Practical principles
-
-- Treat letter readout as a model-and-domain ablation, not a drop-in upgrade.
- Keep it only after a paired evaluation on the target decision distribution.
-- Blend only when the two readouts make complementary errors. The fixed 50/50
- pool helped 27B on both diagnostics; at 4B it helped JevBench public but not
- Transfer-v9. Select the weight on a separate development split.
-- Calibrate each model, readout, and domain separately. Cygnet's fitted
- temperature `3.4` does not transfer to these Qwen checkpoints.
-- Re-measure latency in the deployment runtime and traffic mix. The one-pass
- letter path can trim median latency, while blending requires both letter and
- native passes and nearly doubles median latency here.
-
-## Attribution
-
-The Cygnet-compatible prompt and aggregation semantics are adapted from
-`blockbrain-ai/cygnet-recipe` commit
-`3cf591c692dec649f7c134449814610307c7bb3a`, Copyright 2026 Nood Co and
-contributors, under MIT. Cygnet credits the one-token option-letter readout
-method to [NInfer](https://github.com/igorls/ninfer), released under Apache-2.0.
-JevAny does not incorporate NInfer source code. See [NOTICE](../NOTICE) for the
-full Cygnet license notice.
+The former `letter` CLI/API spelling remains a compatibility alias.
diff --git a/docs/choice-readout-results.svg b/docs/choice-readout-results.svg
new file mode 100644
index 0000000..a32e254
--- /dev/null
+++ b/docs/choice-readout-results.svg
@@ -0,0 +1,928 @@
+
+
+
diff --git a/jevany/__init__.py b/jevany/__init__.py
index d999df1..a3ca0e6 100644
--- a/jevany/__init__.py
+++ b/jevany/__init__.py
@@ -5,8 +5,12 @@
from .api import Choice, Noul, Score, SystemOneRequest
from .client import JevClient
from .inference import InferenceOptions
+from .readout import ChoiceReadoutOptions
-__all__ = ["Choice", "Noul", "Score", "SystemOneRequest", "JevClient", "JevModel", "InferenceOptions"]
+__all__ = [
+ "Choice", "Noul", "Score", "SystemOneRequest", "JevClient", "JevModel",
+ "InferenceOptions", "ChoiceReadoutOptions",
+]
def __getattr__(name: str):
diff --git a/jevany/benchmark.py b/jevany/benchmark.py
index 0a361ca..5ae3a5c 100644
--- a/jevany/benchmark.py
+++ b/jevany/benchmark.py
@@ -26,7 +26,7 @@
from jevany.metrics import EPSILON, grouped_metrics, metrics, unknowable_report
from jevany.model import ContextLengthError
from jevany.predictors import LocalPredictor, RemotePredictor
-from jevany.readout import add_readout_arguments, letter_options_from_args
+from jevany.readout import add_readout_arguments, choice_options_from_args, normalize_readout
from jevany.suite import ENCODING, digest, load_split, read_manifest, record_digest, write_json
@@ -170,7 +170,7 @@ def evaluate_records(records, predictor, directory, temperature=1.0, heldout_sou
EXAMPLES = """examples:
jevany eval --run runs/my-jev --data data/starter/development.jsonl --out runs/my-jev/eval
jevany eval --run SimpleJev/JevAny-Qwen3.5-4B-LoRA --suite data/eval-suite --out runs/eval
- jevany eval --run SimpleJev/JevAny-Qwen3.5-4B-LoRA --readout letter --suite data/eval-suite --out runs/letter
+ jevany eval --run SimpleJev/JevAny-Qwen3.5-4B-LoRA --readout choice --suite data/eval-suite --out runs/choice
jevany eval --remote http://127.0.0.1:8008 --data my-labelled.jsonl --out runs/remote-eval
--data scores your own labelled JSONL (jevany.data.load_records); --suite scores a
@@ -195,12 +195,12 @@ def main(argv=None, prog=None):
ap.add_argument("--date_facts", action="store_true", help="apply jevany.api.with_date_facts to every state before scoring (the opt-in serving preprocessor); reported in report.json")
a = ap.parse_args(argv)
try:
- letter_options = letter_options_from_args(a)
+ choice_options = choice_options_from_args(a)
except ValueError as error:
ap.error(str(error))
if bool(a.run) == bool(a.remote): ap.error("give exactly one of --run or --remote")
- if a.remote and letter_options is not None:
- ap.error("--readout letter is a local model setting; configure it on the remote server")
+ if a.remote and choice_options is not None:
+ ap.error("--readout choice is a local model setting; configure it on the remote server")
if bool(a.suite) == bool(a.data): ap.error("give exactly one of --suite or --data")
if a.data:
records, heldout, split, source_hash = load_records(a.data), [], "custom", digest(Path(a.data))
@@ -212,34 +212,40 @@ def main(argv=None, prog=None):
records = [{**r, "state": with_date_facts(r["state"])} for r in records]
if a.remote:
predictor = RemotePredictor(a.remote, a.remote_model, os.environ.get("JEVANY_REMOTE_API_KEY", "local"))
- elif letter_options is not None:
+ elif choice_options is not None:
from jevany.letter_predictor import LetterReadoutPredictor
predictor = LetterReadoutPredictor(
checkpoint=a.run,
device=a.device,
options=LoadOptions.from_env(),
- temperature=letter_options.temperature,
- pointer_weight=letter_options.pointer_weight,
- max_tokens=letter_options.max_tokens,
+ temperature=choice_options.temperature,
+ native_weight=choice_options.native_weight,
+ max_tokens=choice_options.max_tokens,
)
else:
predictor = LocalPredictor(a.run, a.device, LoadOptions.from_env())
report, _ = evaluate_records(records, predictor, a.out, heldout_sources=tuple(heldout), skip_overlong=bool(a.data))
- pointer_temperature = (getattr(predictor.pointer_model, "temperature", None)
- if letter_options is not None and predictor.pointer_weight else None)
+ native_temperature = (getattr(predictor.native_model, "temperature", None)
+ if choice_options is not None and predictor.native_weight else None)
calibration_applied = (None if a.remote else predictor.temperature != 1.0
- or pointer_temperature not in (None, 1.0))
+ or native_temperature not in (None, 1.0))
report.update(suite_sha256=source_hash, data=a.data, date_facts=a.date_facts, run=a.run or a.remote, split=split,
calibration_applied=calibration_applied,
base_loading=getattr(predictor, "base_loading", None),
remote={"base_url": a.remote, "requested_model": a.remote_model, "served_model": predictor.served_model} if a.remote else None)
if not a.remote:
- report["readout"] = a.readout
- if letter_options is not None:
- report["letter_readout"] = dict(predictor.provenance)
- if pointer_temperature is not None:
- report["letter_readout"]["pointer_temperature"] = pointer_temperature
- report["calibration"]["pointer_temperature"] = pointer_temperature
+ report["readout"] = normalize_readout(a.readout)
+ if choice_options is not None:
+ choice_readout = dict(predictor.provenance)
+ if native_temperature is not None:
+ choice_readout["native_temperature"] = native_temperature
+ choice_readout["pointer_temperature"] = native_temperature
+ report["calibration"]["native_temperature"] = native_temperature
+ report["calibration"]["pointer_temperature"] = native_temperature
+ report["choice_readout"] = choice_readout
+ # Kept as an output alias for consumers of reports written before the
+ # public readout name changed to ``choice``.
+ report["letter_readout"] = dict(choice_readout)
write_json(Path(a.out) / "report.json", report)
print(json.dumps({"objective": report["objective"], "clean": report["clean"], "coverage": report["coverage"]}, indent=2))
diff --git a/jevany/cli.py b/jevany/cli.py
index 055a321..a04a0d7 100644
--- a/jevany/cli.py
+++ b/jevany/cli.py
@@ -45,7 +45,7 @@ def decide_main(argv: list[str]) -> None:
from dataclasses import fields
from .inference import InferenceOptions, add_inference_arguments, inference_options_from_args
from .placement import add_placement_arguments
- from .readout import add_readout_arguments, letter_options_from_args
+ from .readout import add_readout_arguments, choice_options_from_args
parser = argparse.ArgumentParser(prog="jevany decide")
parser.add_argument("request", help="JSON request file; - reads stdin")
@@ -63,7 +63,7 @@ def decide_main(argv: list[str]) -> None:
add_inference_arguments(parser)
args = parser.parse_args(argv)
try:
- letter_options = letter_options_from_args(args)
+ choice_options = choice_options_from_args(args)
except ValueError as error:
parser.error(str(error))
content = sys.stdin.read() if args.request == "-" else Path(args.request).read_text(encoding="utf-8")
@@ -73,12 +73,12 @@ def decide_main(argv: list[str]) -> None:
from dataclasses import replace
from .checkpoint import LoadOptions, load_options_from_args
from .runtime import JevModel
- if letter_options is not None and (
+ if choice_options is not None and (
args.cuda_graphs or args.cuda_graph_max_tokens is not None
or any(getattr(args, item.name) is not None for item in fields(InferenceOptions))
):
- parser.error("CUDA graph and native inference-limit flags cannot be used with --readout letter; "
- "use --letter-max-tokens")
+ parser.error("CUDA graph and native inference-limit flags cannot be used with --readout choice; "
+ "use --choice-max-tokens")
options = load_options_from_args(args)
if args.cuda_graphs or args.cuda_graph_max_tokens is not None:
options = options or LoadOptions.from_env()
@@ -87,11 +87,11 @@ def decide_main(argv: list[str]) -> None:
if args.cuda_graph_max_tokens is not None
else options.cuda_graph_max_tokens))
settings = ({
- "readout": "letter",
- "letter_temperature": letter_options.temperature,
- "letter_pointer_weight": letter_options.pointer_weight,
- "letter_max_tokens": letter_options.max_tokens,
- } if letter_options is not None else {
+ "readout": "choice",
+ "choice_temperature": choice_options.temperature,
+ "choice_native_weight": choice_options.native_weight,
+ "choice_max_tokens": choice_options.max_tokens,
+ } if choice_options is not None else {
"inference_options": inference_options_from_args(args),
})
client = JevModel.from_pretrained(
diff --git a/jevany/demos/static/app.js b/jevany/demos/static/app.js
index b7a51d1..2f43c48 100644
--- a/jevany/demos/static/app.js
+++ b/jevany/demos/static/app.js
@@ -217,7 +217,8 @@ async function liveStep(action) {
}
function renderConnection() {
const media = link.media || {}, served = link.served;
- const readout = served && served.readout === "letter" ? "letter" : served && served.decision_mode;
+ const readout = served && (served.readout === "choice" || served.readout === "letter")
+ ? "choice" : served && served.decision_mode;
$("connect-url").disabled = $("connect-model").disabled = !link.editable;
if (!$("connect-url").value) $("connect-url").value = link.base_url || link.default_base_url || "";
if (!$("connect-model").value && link.model && link.model !== "jevany-latest") $("connect-model").value = link.model;
diff --git a/jevany/letter_predictor.py b/jevany/letter_predictor.py
index f5c9e4b..ea9c656 100644
--- a/jevany/letter_predictor.py
+++ b/jevany/letter_predictor.py
@@ -1,14 +1,15 @@
-"""Training-free option-letter inference over a frozen base or JevAny adapter.
+"""Training-free choice-token inference over a frozen base or JevAny adapter.
-This is the in-process backend for :mod:`jevany.letter_readout`. It keeps the
-existing pointer path unchanged: a JevAny checkpoint can instead be applied to
-the chat prompt before the frozen vocabulary readout, and its pointer
-distribution can optionally be combined with the letter distribution.
+This is the in-process backend for :mod:`jevany.letter_readout`. A JevAny
+checkpoint can be applied to the chat prompt before the frozen vocabulary
+readout, and its Pointer or Direct-Token native distribution can optionally be
+combined with the choice distribution.
"""
from __future__ import annotations
from collections.abc import Mapping
+from contextlib import nullcontext
import hashlib
import json
import math
@@ -17,13 +18,14 @@
import torch
import torch.nn.functional as F
+from torch.nn.attention import SDPBackend, sdpa_kernel
from .api import SystemOneRequest, question_keys
from .checkpoint import Checkpoint, LoadOptions
from .data import api_request, materialize
from .device import sync
from .letter_readout import (
- LETTERS,
+ CHOICE_SYMBOLS,
SYSTEM_PROMPT,
build_prompt,
letter_token_ids,
@@ -61,17 +63,19 @@ def _safe_chat_text(tokenizer, text: str) -> str:
return escaped
-def _release_unused_output_head(decision_model) -> bool:
- """Drop references to the full-vocabulary head unused by letter readout."""
+def _release_unused_output_head(decision_model, *, retain_native: bool = False) -> bool:
+ """Drop the full-vocabulary head unless a native LM-token pass needs it."""
adapter = getattr(decision_model, "adapter", None)
adapter_head = getattr(adapter, "_output_embeddings", None)
native_head = getattr(decision_model, "lm_head", None)
+ if retain_native:
+ return False
if adapter_head is None and native_head is None:
return False
- # The exact letter rows are copied immediately after this call. Native
- # lm-token inference is never invoked by this predictor, so both aliases
- # can be detached even for direct-token checkpoints.
+ # Exact choice rows are copied from the tied embedding table or loaded
+ # independently from safetensors. With no LM-token native pass, both
+ # aliases can therefore be detached to release the full vocabulary head.
if native_head is not None:
decision_model.lm_head = None
if adapter_head is not None:
@@ -223,9 +227,10 @@ class LetterReadoutPredictor:
"""Benchmark predictor for an exact constrained first-token readout.
Give either ``base`` for a frozen training-free model, or ``checkpoint``
- to apply a JevAny LoRA before the same readout. ``pointer_weight > 0``
- additionally pools the checkpoint's native pointer probabilities in log
- space; this costs one extra pointer-formatted prefill per request.
+ to apply a JevAny LoRA before the same readout. ``native_weight > 0``
+ additionally pools the checkpoint's native probabilities in log space;
+ this costs one extra native-formatted prefill per request. The legacy
+ ``pointer_weight`` spelling remains an alias for ``native_weight``.
"""
def __init__(
@@ -237,7 +242,11 @@ def __init__(
options: LoadOptions | None = None,
revision: str | None = None,
temperature: float = 1.0,
- pointer_weight: float = 0.0,
+ pointer_weight: float | None = None,
+ native_weight: float | None = None,
+ return_components: bool = False,
+ exact_kernels: bool = True,
+ efficient_long_context_tokens: int | None = None,
max_tokens: int = 16_384,
) -> None:
if (base is None) == (checkpoint is None):
@@ -247,21 +256,66 @@ def __init__(
self.temperature = float(temperature)
if not math.isfinite(self.temperature) or self.temperature <= 0:
raise ValueError("temperature must be finite and positive")
+ if pointer_weight is not None and native_weight is not None:
+ raise ValueError("give at most one of native_weight or pointer_weight")
+ weight_source = (
+ "native_weight" if native_weight is not None
+ else "pointer_weight" if pointer_weight is not None
+ else "default"
+ )
+ selected_weight = (
+ native_weight if native_weight is not None
+ else pointer_weight if pointer_weight is not None
+ else 0.0
+ )
# Reuse the same validation as the actual combination path.
- geometric_blend([1.0], [1.0], pointer_weight)
- self.pointer_weight = float(pointer_weight)
- if checkpoint is None and self.pointer_weight:
- raise ValueError("pointer_weight requires a JevAny checkpoint")
+ geometric_blend([1.0], [1.0], selected_weight)
+ self.native_weight = float(selected_weight)
+ # Public compatibility alias for existing runtime/CLI integrations.
+ self.pointer_weight = self.native_weight
+ if checkpoint is None and self.native_weight:
+ raise ValueError("native_weight requires a JevAny checkpoint")
+ if type(return_components) is not bool:
+ raise TypeError("return_components must be a boolean")
+ if type(exact_kernels) is not bool:
+ raise TypeError("exact_kernels must be a boolean")
+ self.return_components = return_components
+ self.exact_kernels = exact_kernels
+ self.exact_cuda_kernels_applied = str(device).startswith("cuda") and exact_kernels
+ if (
+ efficient_long_context_tokens is not None
+ and (
+ type(efficient_long_context_tokens) is not int
+ or efficient_long_context_tokens < 2
+ )
+ ):
+ raise ValueError("efficient_long_context_tokens must be an integer >= 2 or None")
+ self.efficient_long_context_tokens = efficient_long_context_tokens
if type(max_tokens) is not int or max_tokens < 2:
raise ValueError("max_tokens must be an integer >= 2")
self.max_tokens = max_tokens
self.device = device
self.options = options or LoadOptions()
if self.options.temperature is not None:
- raise ValueError("native-head temperature is not used by letter readout; use temperature")
+ raise ValueError("native-head temperature is not used by choice readout; use temperature")
if self.options.cuda_graphs:
raise ValueError("CUDA graph capture is available only for native readout")
+ if self.exact_cuda_kernels_applied:
+ # Match LocalPredictor: benchmark evaluation disables approximate
+ # TF32 and fused SDPA kernels, while exact_kernels=False retains
+ # the serving-style PyTorch defaults.
+ torch.backends.cuda.matmul.allow_tf32 = False
+ torch.backends.cudnn.allow_tf32 = False
+ torch.backends.cuda.enable_flash_sdp(False)
+ torch.backends.cuda.enable_mem_efficient_sdp(False)
self.checkpoint: Checkpoint | None = None
+ self.native_model = None
+ self.native_decision_mode = None
+ self._needs_native = checkpoint is not None and bool(
+ self.native_weight or self.return_components
+ )
+ # Compatibility alias retained for callers predating generalized
+ # native blending.
self.pointer_model = None
if checkpoint is not None:
@@ -269,11 +323,11 @@ def __init__(
raise ValueError("revision accompanies a base; pin checkpoint revisions in owner/repo@revision")
loaded = self.checkpoint = Checkpoint(checkpoint)
if loaded.meta.special_embeddings:
- raise ValueError("letter readout is not validated for checkpoints with trained special embeddings")
- if self.pointer_weight and loaded.meta.decision_mode != "pointer":
- raise ValueError("pointer_weight requires a native pointer checkpoint")
+ raise ValueError("choice readout is not validated for checkpoints with trained special embeddings")
self.preprocessor, decision_model = loaded.load(device, self.options)
+ self.native_model = decision_model
self.pointer_model = decision_model
+ self.native_decision_mode = loaded.meta.decision_mode
self.language_model = decision_model.lm
canonical_base = loaded.meta.base
canonical_revision = loaded.meta.base_revision
@@ -305,7 +359,10 @@ def __init__(
projection_source, projection_revision = source, source_revision
checkpoint_id, adapter_scale, adapter_applied = None, 0.0, False
- released_output_head = _release_unused_output_head(decision_model)
+ released_output_head = _release_unused_output_head(
+ decision_model,
+ retain_native=(self._needs_native and self.native_decision_mode == "lm_token"),
+ )
self.language_model.eval()
override_config = (
Path(self.options.base_load_path) / "config.json"
@@ -339,7 +396,7 @@ def __init__(
hashlib.sha256(chat_template.encode()).hexdigest()
if isinstance(chat_template, str) else None
)
- self.alias_token_ids = letter_token_ids(self.tokenizer, LETTERS)
+ self.alias_token_ids = letter_token_ids(self.tokenizer, CHOICE_SYMBOLS)
flat_token_ids, alias_rows = [], {}
for letter, token_ids in self.alias_token_ids.items():
start = len(flat_token_ids)
@@ -349,7 +406,7 @@ def __init__(
if tied:
embeddings = self.language_model.get_input_embeddings().weight
if max(flat_token_ids) >= embeddings.shape[0]:
- raise ValueError("letter choice token ID exceeds the tied vocabulary projection")
+ raise ValueError("choice token ID exceeds the tied vocabulary projection")
index = torch.tensor(flat_token_ids, dtype=torch.long, device=embeddings.device)
projection = embeddings.index_select(0, index).detach().clone()
projection_bias = None
@@ -368,12 +425,12 @@ def __init__(
if not math.isfinite(self.output_softcap) or self.output_softcap <= 0:
raise ValueError("final_logit_softcapping must be finite and positive")
self.provenance = {
- "method": "exact option-letter choice projection",
- "constraint_emulation": "exact decoded uppercase letters",
- "prompt": "Cygnet-compatible",
+ "method": "training-free exact choice-token projection",
+ "constraint_emulation": "exact decoded one-character option IDs",
+ "prompt": "Cygnet-compatible A-Z; extended with a-z above 26 options",
"system_prompt_sha256": hashlib.sha256(SYSTEM_PROMPT.encode()).hexdigest(),
"chat_template_sha256": self.chat_template_sha256,
- "prompt_format_version": 1,
+ "prompt_format_version": 2,
"canonical_base": canonical_base,
"canonical_revision": canonical_revision,
"checkpoint": checkpoint_id,
@@ -381,12 +438,29 @@ def __init__(
"adapter_scale": adapter_scale,
"tied_output_embeddings": tied,
"unused_full_output_head_released": released_output_head,
+ "choice_token_rows": len(flat_token_ids),
"letter_choice_token_rows": len(flat_token_ids),
"output_softcap": self.output_softcap,
"temperature": self.temperature,
+ "native_weight": self.native_weight,
+ "native_weight_argument": weight_source,
+ "native_decision_mode": self.native_decision_mode,
"pointer_weight": self.pointer_weight,
+ "native_temperature": (
+ float(self.native_model.temperature) if self._needs_native else None
+ ),
"pointer_temperature": (
- float(self.pointer_model.temperature) if self.pointer_weight else None
+ float(self.native_model.temperature) if self._needs_native else None
+ ),
+ "return_components": self.return_components,
+ "exact_kernels": self.exact_kernels,
+ "exact_cuda_kernels_applied": self.exact_cuda_kernels_applied,
+ "efficient_long_context_tokens": self.efficient_long_context_tokens,
+ "long_context_attention": (
+ "fused SDPA (Flash/Efficient, no math fallback) at or above the "
+ "recorded threshold; exact math SDPA below it"
+ if self.efficient_long_context_tokens is not None
+ else None
),
"max_tokens": self.max_tokens,
"backbone_context_window": self.backbone_context_window,
@@ -395,6 +469,26 @@ def __init__(
"device": device,
}
+ def _uses_efficient_attention(self, token_count: int) -> bool:
+ return (
+ str(self.device).startswith("cuda")
+ and getattr(self, "efficient_long_context_tokens", None) is not None
+ and token_count >= self.efficient_long_context_tokens
+ )
+
+ def _attention_context(self, token_count: int):
+ if not self._uses_efficient_attention(token_count):
+ return nullcontext()
+ # Exact math SDPA materializes an O(sequence^2) attention matrix. A
+ # few public JevJudge text records exceed 30K tokens, which can require
+ # roughly 60 GiB for that matrix alone. Release cached short-record
+ # workspaces and force a memory-linear fused backend for those records.
+ torch.cuda.empty_cache()
+ return sdpa_kernel(
+ [SDPBackend.FLASH_ATTENTION, SDPBackend.EFFICIENT_ATTENTION],
+ set_priority=True,
+ )
+
def _chat_ids(self, prompt: str) -> torch.Tensor:
messages = [
{"role": "system", "content": SYSTEM_PROMPT},
@@ -410,22 +504,25 @@ def _chat_ids(self, prompt: str) -> torch.Tensor:
ids = _one_row_input_ids(ids)
if ids.shape[1] + 1 > self.effective_max_tokens:
raise ContextLengthError(
- f"letter prompt needs {ids.shape[1] + 1} tokens including the answer; "
+ f"choice prompt needs {ids.shape[1] + 1} tokens including the answer; "
f"limit {self.effective_max_tokens}"
)
return ids.to(self.device)
def _letter_question(self, state, question: dict) -> tuple[list[str], list[float], list[float], int]:
keys, descriptions = question_options(question)
- if len(keys) > len(LETTERS):
- raise ValueError(f"exact letter readout supports at most {len(LETTERS)} options")
+ if len(keys) > len(CHOICE_SYMBOLS):
+ raise ValueError(
+ f"exact choice-token readout supports at most {len(CHOICE_SYMBOLS)} options"
+ )
prompt = build_prompt(state, question.get("instructions") or "", descriptions)
ids = self._chat_ids(prompt)
- output = self.language_model(
- input_ids=ids,
- attention_mask=torch.ones_like(ids),
- use_cache=False,
- )
+ with self._attention_context(int(ids.shape[1])):
+ output = self.language_model(
+ input_ids=ids,
+ attention_mask=torch.ones_like(ids),
+ use_cache=False,
+ )
hidden = output.last_hidden_state[0, -1]
projection = self.letter_projection
selected_logits = F.linear(
@@ -434,15 +531,17 @@ def _letter_question(self, state, question: dict) -> tuple[list[str], list[float
if self.output_softcap is not None:
selected_logits = self.output_softcap * torch.tanh(selected_logits / self.output_softcap)
selected_logits = selected_logits.cpu()
- aliases = {letter: self.alias_rows[letter] for letter in LETTERS[:len(keys)]}
+ aliases = {
+ symbol: self.alias_rows[symbol] for symbol in CHOICE_SYMBOLS[:len(keys)]
+ }
readout = read_letter_distribution(selected_logits, aliases, self.temperature)
probabilities = [readout.calibrated_probabilities[letter] for letter in aliases]
calibrated_logits = [readout.raw_log_masses[letter] / self.temperature for letter in aliases]
return keys, probabilities, calibrated_logits, int(ids.shape[1])
- def _pointer(self, record: dict) -> tuple[list[list[float]], int]:
+ def _native(self, record: dict) -> tuple[list[list[float]], int]:
internal = materialize(record)
- encoded = self.pointer_model.encode(
+ encoded = self.native_model.encode(
self.preprocessor,
internal,
max_state=self.effective_max_tokens,
@@ -451,55 +550,83 @@ def _pointer(self, record: dict) -> tuple[list[list[float]], int]:
)
if len(encoded["ids"]) > self.effective_max_tokens:
raise ContextLengthError(
- f"pointer request needs {len(encoded['ids'])} packed tokens; "
+ f"native request needs {len(encoded['ids'])} packed tokens; "
f"limit {self.effective_max_tokens}"
)
- return [F.softmax(logits, -1).float().cpu().tolist()
- for logits in self.pointer_model.forward(encoded)], len(encoded["ids"])
+ with self._attention_context(len(encoded["ids"])):
+ rows = [F.softmax(logits, -1).float().cpu().tolist()
+ for logits in self.native_model.forward(encoded)]
+ return rows, len(encoded["ids"])
+
+ def _pointer(self, record: dict) -> tuple[list[list[float]], int]:
+ """Compatibility alias for the generalized native checkpoint pass."""
+
+ return self._native(record)
@torch.inference_mode()
def __call__(self, record: dict) -> dict:
request = api_request(record)
SystemOneRequest.model_validate(request)
if request.get("media"):
- raise ValueError("letter readout is text-only and does not accept media")
+ raise ValueError("choice readout is text-only and does not accept media")
sync(self.device)
started = time.perf_counter()
letter_rows = [
self._letter_question(request["state"], question)
for question in request["questions"].values()
]
- pointer_rows, pointer_tokens = (None, 0)
- if self.pointer_weight:
- pointer_rows, pointer_tokens = self._pointer(record)
- if len(pointer_rows) != len(letter_rows):
- raise ValueError("pointer and letter question counts differ")
+ native_rows, native_tokens = (None, 0)
+ if self._needs_native:
+ native_rows, native_tokens = self._native(record)
+ if len(native_rows) != len(letter_rows):
+ raise ValueError("native and choice question counts differ")
probabilities, logits = {}, {}
+ choice_probabilities, native_probabilities = {}, {}
for index, (question_id, (keys, letter_p, letter_z, _tokens)) in enumerate(
zip(request["questions"], letter_rows, strict=True)
):
- if pointer_rows is None:
+ choice_probabilities[question_id] = dict(zip(keys, letter_p, strict=True))
+ if native_rows is not None:
+ if len(native_rows[index]) != len(keys):
+ raise ValueError(
+ f"native and choice option counts differ for question {question_id!r}"
+ )
+ native_probabilities[question_id] = dict(
+ zip(keys, native_rows[index], strict=True)
+ )
+ if native_rows is None or not self.native_weight:
selected_p, selected_z = letter_p, letter_z
else:
selected_p, selected_z = geometric_blend(
- letter_p, pointer_rows[index], self.pointer_weight
+ letter_p, native_rows[index], self.native_weight
)
probabilities[question_id] = dict(zip(keys, selected_p, strict=True))
logits[question_id] = dict(zip(keys, selected_z, strict=True))
sync(self.device)
- return {
+ result = {
"probabilities": probabilities,
"logits": logits,
"inference_temperature": self.temperature,
"latency_ms": (time.perf_counter() - started) * 1000,
- "input_tokens": sum(row[3] for row in letter_rows) + pointer_tokens,
+ "input_tokens": sum(row[3] for row in letter_rows) + native_tokens,
+ "efficient_long_context_attention_used": (
+ any(self._uses_efficient_attention(row[3]) for row in letter_rows)
+ or self._uses_efficient_attention(native_tokens)
+ ),
"readout": {
"method": self.provenance["method"],
"adapter_applied": self.provenance["adapter_applied"],
+ "native_weight": self.native_weight,
+ "native_decision_mode": self.native_decision_mode,
"pointer_weight": self.pointer_weight,
},
}
+ if self.return_components:
+ result["component_probabilities"] = {"choice": choice_probabilities}
+ if native_rows is not None:
+ result["component_probabilities"]["native"] = native_probabilities
+ return result
__all__ = ["LetterReadoutPredictor", "geometric_blend", "question_options"]
diff --git a/jevany/letter_readout.py b/jevany/letter_readout.py
index 0b857f7..f38ac57 100644
--- a/jevany/letter_readout.py
+++ b/jevany/letter_readout.py
@@ -1,4 +1,4 @@
-"""One-token option-letter readout primitives for frozen causal LMs.
+"""One-token option-ID readout primitives for frozen causal LMs.
The Cygnet-compatible prompt and letter aggregation semantics are adapted from
the MIT-licensed ``blockbrain-ai/cygnet-recipe`` shim at commit ``3cf591c``
@@ -18,11 +18,18 @@
from typing import Any
-LETTERS = "ABCDEFGHIJKLMNOPQRSTUVWXYZ"
+CYGNET_LETTERS = "ABCDEFGHIJKLMNOPQRSTUVWXYZ"
+# Preserve Cygnet's exact A-Z assignment for its supported range, then extend
+# with distinct lowercase one-character IDs for larger decision sets. This
+# covers JevJudge's 28-30-option text requests without dropping records.
+CHOICE_SYMBOLS = CYGNET_LETTERS + "abcdefghijklmnopqrstuvwxyz"
+# Backward-compatible implementation name. Public documentation calls these
+# option IDs and the method training-free choice readout.
+LETTERS = CHOICE_SYMBOLS
SYSTEM_PROMPT = (
"You are a calibration engine. You never answer in prose. You are given a state, a question and "
"a numbered set of options, and you choose exactly one option. You reply with that option's "
- "LETTER and nothing else โ a single character, no words, no punctuation, no explanation."
+ "case-sensitive ID and nothing else โ a single character, no words, no punctuation, no explanation."
)
_NEGATIVE_INFINITY = float("-inf")
@@ -63,7 +70,7 @@ def _ordered_descriptions(options: Mapping[Any, Any] | Sequence[Any]) -> list[An
def build_prompt(state: Any, instructions: Any, options: Mapping[Any, Any] | Sequence[Any]) -> str:
- """Build Cygnet's measured user prompt, assigning options A through Z.
+ """Build the measured user prompt, assigning one-character option IDs.
Mapping insertion order or sequence order is the option order. Structured
state is rendered with ``indent=1`` as in Cygnet. This function only
@@ -78,22 +85,22 @@ def build_prompt(state: Any, instructions: Any, options: Mapping[Any, Any] | Seq
)
lines = [state_text.rstrip(), "", instruction_text.rstrip(), "", "Options:"]
lines.extend(f"{LETTERS[index]}. {description}" for index, description in enumerate(descriptions))
- lines.extend(("", "Answer with the letter of exactly one option, and nothing else:"))
+ lines.extend(("", "Answer with the case-sensitive ID of exactly one option, and nothing else:"))
return "\n".join(lines)
def _validate_letters(letters: Sequence[str]) -> tuple[str, ...]:
if isinstance(letters, (bytes, bytearray)):
- raise TypeError("letters must be a sequence of A-Z strings")
+ raise TypeError("option IDs must be a sequence of one-character strings")
selected = tuple(letters)
if not selected:
- raise ValueError("letters must not be empty")
+ raise ValueError("option IDs must not be empty")
if len(selected) > len(LETTERS):
raise ValueError(f"letters may contain at most {len(LETTERS)} entries")
- if any(type(letter) is not str or len(letter) != 1 or letter not in LETTERS for letter in selected):
- raise ValueError("letters must contain only single uppercase A-Z strings")
+ if any(type(letter) is not str or len(letter) != 1 or letter not in CHOICE_SYMBOLS for letter in selected):
+ raise ValueError("option IDs must be distinct single A-Z/a-z characters")
if len(set(selected)) != len(selected):
- raise ValueError("letters must not contain duplicates")
+ raise ValueError("option IDs must not contain duplicates")
return selected
@@ -113,12 +120,12 @@ def _decode_token(tokenizer: Any, token_id: int) -> str:
def letter_token_ids(tokenizer: Any, letters: Sequence[str]) -> dict[str, tuple[int, ...]]:
- """Scan the vocabulary for every token ID decoding to each exact letter.
+ """Scan the vocabulary for every token ID decoding to each exact option ID.
Token text is intentionally never used as a dictionary key: distinct token
IDs may decode to identical text, and every such ID contributes probability
- mass. Tokens decoding to ``" A"``, ``"A."``, or ``"a"`` are excluded:
- Cygnet's structured-output choice admits the exact uppercase string only.
+ mass. Tokens decoding to ``" A"`` or ``"A."`` are excluded. Case is
+ significant: ``A`` and ``a`` are different option IDs when both are used.
``len(tokenizer)`` must describe the full vocabulary, including added tokens.
"""
@@ -338,6 +345,8 @@ def read_letter_distribution(
__all__ = [
+ "CHOICE_SYMBOLS",
+ "CYGNET_LETTERS",
"LETTERS",
"SYSTEM_PROMPT",
"LetterReadout",
diff --git a/jevany/letter_runtime.py b/jevany/letter_runtime.py
index 2f41022..1ea0637 100644
--- a/jevany/letter_runtime.py
+++ b/jevany/letter_runtime.py
@@ -1,4 +1,4 @@
-"""System One runtime adapter for the training-free option-letter predictor."""
+"""System One runtime adapter for the training-free choice-token predictor."""
from __future__ import annotations
@@ -7,6 +7,7 @@
from typing import Any
from .api import SystemOneRequest, output_tokens, to_answers, to_record, validate_response
+from .letter_readout import CHOICE_SYMBOLS
def _unlabelled_record(request: SystemOneRequest) -> dict[str, Any]:
@@ -41,7 +42,7 @@ def __post_init__(self) -> None:
if not isinstance(self.model_id, str) or not self.model_id.strip():
raise ValueError("model_name must be a nonempty string")
if self.predictor.checkpoint is None:
- raise ValueError("serving letter readout requires a JevAny checkpoint")
+ raise ValueError("serving choice readout requires a JevAny checkpoint")
@property
def checkpoint(self):
@@ -52,25 +53,26 @@ def aliases(self) -> list[str]:
return [] if self.model_id == "jevany-latest" else ["jevany-latest"]
def clear_cache(self) -> None:
- """Letter readout currently keeps no mutable prefix cache."""
+ """Choice readout currently keeps no mutable prefix cache."""
def describe(self) -> dict[str, Any]:
- """Describe the effective letter readout and its checkpoint overlay."""
+ """Describe the effective choice readout and its checkpoint overlay."""
checkpoint = self.checkpoint
- pointer_model = self.predictor.pointer_model
- capabilities = pointer_model.inference_capabilities
+ native_model = self.predictor.native_model
+ capabilities = native_model.inference_capabilities
context_window = capabilities.context_window
effective_window = self.predictor.effective_max_tokens
- acceleration = getattr(pointer_model, "inference_acceleration", {
+ acceleration = getattr(native_model, "inference_acceleration", {
"compile_mode": None,
"lora_merged": False,
"approximate_bf16_merge": False,
"cuda_graphs": None,
})
- letter_readout = dict(self.predictor.provenance)
- if self.predictor.pointer_weight:
- letter_readout["pointer_temperature"] = pointer_model.temperature
+ choice_readout = dict(self.predictor.provenance)
+ if self.predictor.native_weight:
+ choice_readout["native_temperature"] = native_model.temperature
+ choice_readout["pointer_temperature"] = native_model.temperature
with self.lock:
return {
"id": self.model_id,
@@ -79,13 +81,15 @@ def describe(self) -> dict[str, Any]:
"base": checkpoint.meta.base,
"lora": checkpoint.meta.lora,
"device": self.predictor.device,
- "device_map": getattr(pointer_model, "device_map", None),
- "devices": getattr(pointer_model, "devices", [self.predictor.device]),
+ "device_map": getattr(native_model, "device_map", None),
+ "devices": getattr(native_model, "devices", [self.predictor.device]),
"temperature": self.predictor.temperature,
"decision_mode": checkpoint.meta.decision_mode,
- "readout": "letter",
- "letter_readout": letter_readout,
- "backbone_adapter": pointer_model.backbone_adapter,
+ "readout": "choice",
+ "choice_readout": choice_readout,
+ # Legacy descriptor alias retained for older clients.
+ "letter_readout": dict(choice_readout),
+ "backbone_adapter": native_model.backbone_adapter,
"branch_mode": "chat",
"acceleration": acceleration,
"capabilities": {
@@ -98,7 +102,7 @@ def describe(self) -> dict[str, Any]:
"state_tokens": effective_window,
"branch_tokens": effective_window,
"packed_tokens": effective_window,
- "choices": 26,
+ "choices": len(CHOICE_SYMBOLS),
},
"prefix_cache": {
"enabled": False,
@@ -111,12 +115,12 @@ def describe(self) -> dict[str, Any]:
}
def answer(self, request: SystemOneRequest) -> dict[str, Any]:
- """Run letter inference and return a validated System One response."""
+ """Run choice-token inference and return a validated System One response."""
if request.model not in (self.model_id, *self.aliases):
raise ValueError(f"unknown model {request.model!r}; this deployment serves {self.model_id!r}")
if request.media:
- raise ValueError("letter readout does not support media requests")
+ raise ValueError("choice readout does not support media requests")
_, metadata = to_record(request)
with self.lock:
prediction = self.predictor(_unlabelled_record(request))
diff --git a/jevany/readout.py b/jevany/readout.py
index 99f8b45..387a1a1 100644
--- a/jevany/readout.py
+++ b/jevany/readout.py
@@ -11,69 +11,159 @@
import math
-READOUTS = ("native", "letter")
+# ``letter`` is accepted indefinitely as the legacy spelling, but is omitted
+# from generated usage/help so new integrations consistently say ``choice``.
+READOUTS = ("native", "choice")
+_ACCEPTED_READOUTS = (*READOUTS, "letter")
+
+
+def normalize_readout(readout: str) -> str:
+ """Return the public readout name while accepting the legacy alias."""
+
+ if readout == "letter":
+ return "choice"
+ if readout not in READOUTS:
+ raise ValueError("readout must be native or choice")
+ return readout
@dataclass(frozen=True)
-class LetterReadoutOptions:
- """Runtime settings for the training-free option-letter readout."""
+class ChoiceReadoutOptions:
+ """Runtime settings for the training-free choice-token readout."""
temperature: float = 1.0
- pointer_weight: float = 0.0
+ native_weight: float = 0.0
max_tokens: int = 16_384
def __post_init__(self) -> None:
- for name in ("temperature", "pointer_weight"):
+ for name in ("temperature", "native_weight"):
value = getattr(self, name)
if isinstance(value, bool) or not isinstance(value, (int, float)) or not math.isfinite(value):
- raise ValueError(f"letter {name.replace('_', ' ')} must be finite and numeric")
+ raise ValueError(f"choice {name.replace('_', ' ')} must be finite and numeric")
if self.temperature <= 0:
- raise ValueError("letter temperature must be positive")
- if not 0 <= self.pointer_weight <= 1:
- raise ValueError("letter pointer weight must be in [0, 1]")
+ raise ValueError("choice temperature must be positive")
+ if not 0 <= self.native_weight <= 1:
+ raise ValueError("choice native weight must be in [0, 1]")
if type(self.max_tokens) is not int or self.max_tokens < 2:
- raise ValueError("letter max tokens must be an integer >= 2")
+ raise ValueError("choice max tokens must be an integer >= 2")
+
+ @property
+ def pointer_weight(self) -> float:
+ """Legacy name for :attr:`native_weight`."""
+
+ return self.native_weight
+
+
+@dataclass(frozen=True)
+class LetterReadoutOptions:
+ """Legacy spelling of :class:`ChoiceReadoutOptions`."""
+
+ temperature: float = 1.0
+ pointer_weight: float = 0.0
+ max_tokens: int = 16_384
+
+ def __post_init__(self) -> None:
+ # Validate through the public representation so both APIs have exactly
+ # the same accepted values.
+ ChoiceReadoutOptions(self.temperature, self.pointer_weight, self.max_tokens)
+
+ @property
+ def native_weight(self) -> float:
+ return self.pointer_weight
+
+ def as_choice(self) -> ChoiceReadoutOptions:
+ return ChoiceReadoutOptions(self.temperature, self.pointer_weight, self.max_tokens)
def add_readout_arguments(parser: argparse.ArgumentParser) -> None:
- """Add the same readout selector and letter-only knobs to a CLI parser."""
+ """Add shared readout arguments, with legacy flags accepted but hidden."""
parser.add_argument(
- "--readout", choices=READOUTS, default="native",
- help="native checkpoint head (default) or training-free option-letter readout",
+ "--readout", choices=_ACCEPTED_READOUTS, metavar="{native,choice}", default="native",
+ help="native checkpoint head (default) or training-free choice-token readout",
)
parser.add_argument(
- "--letter-temperature", type=float,
- help="temperature for letter probability calibration (letter readout only; default 1.0)",
+ "--choice-temperature", type=float,
+ help="probability temperature for choice readout (default 1.0)",
)
parser.add_argument(
- "--letter-pointer-weight", type=float,
- help="log-linear weight on the native pointer distribution (letter readout only; default 0)",
+ "--choice-native-weight", type=float,
+ help="log-linear weight on the checkpoint's native distribution (default 0)",
)
parser.add_argument(
- "--letter-max-tokens", type=int,
- help="maximum chat-prompt tokens including the one-letter answer (letter readout only; default 16384)",
+ "--choice-max-tokens", type=int,
+ help="maximum chat-prompt tokens including the one-token answer (default 16384)",
+ )
+ # Backward-compatible spellings. argparse.SUPPRESS keeps the preferred
+ # surface compact without removing existing scripts.
+ parser.add_argument("--letter-temperature", type=float, help=argparse.SUPPRESS)
+ parser.add_argument("--letter-pointer-weight", type=float, help=argparse.SUPPRESS)
+ parser.add_argument("--letter-max-tokens", type=int, help=argparse.SUPPRESS)
+
+
+def resolve_readout_options(
+ readout: str = "native", *,
+ choice_temperature: float | None = None,
+ choice_native_weight: float | None = None,
+ choice_max_tokens: int | None = None,
+ letter_temperature: float | None = None,
+ letter_pointer_weight: float | None = None,
+ letter_max_tokens: int | None = None,
+) -> tuple[str, ChoiceReadoutOptions | None]:
+ """Normalize aliases and resolve preferred/legacy setting spellings."""
+
+ normalized = normalize_readout(readout)
+ pairs = (
+ ("temperature", "--choice-temperature", choice_temperature,
+ "--letter-temperature", letter_temperature),
+ ("native_weight", "--choice-native-weight", choice_native_weight,
+ "--letter-pointer-weight", letter_pointer_weight),
+ ("max_tokens", "--choice-max-tokens", choice_max_tokens,
+ "--letter-max-tokens", letter_max_tokens),
+ )
+ values, used = {}, []
+ for name, preferred_flag, preferred, legacy_flag, legacy in pairs:
+ if preferred is not None and legacy is not None:
+ raise ValueError(f"give at most one of {preferred_flag} and legacy {legacy_flag}")
+ if preferred is not None:
+ values[name] = preferred
+ used.append(preferred_flag)
+ elif legacy is not None:
+ values[name] = legacy
+ used.append(legacy_flag)
+ if normalized == "native":
+ if used:
+ raise ValueError(f"{', '.join(used)} require --readout choice")
+ return normalized, None
+ return normalized, ChoiceReadoutOptions(**values)
+
+
+def choice_options_from_args(args: argparse.Namespace) -> ChoiceReadoutOptions | None:
+ """Return preferred choice settings from a shared CLI namespace."""
+
+ _, options = resolve_readout_options(
+ getattr(args, "readout", "native"),
+ choice_temperature=getattr(args, "choice_temperature", None),
+ choice_native_weight=getattr(args, "choice_native_weight", None),
+ choice_max_tokens=getattr(args, "choice_max_tokens", None),
+ letter_temperature=getattr(args, "letter_temperature", None),
+ letter_pointer_weight=getattr(args, "letter_pointer_weight", None),
+ letter_max_tokens=getattr(args, "letter_max_tokens", None),
)
+ return options
def letter_options_from_args(args: argparse.Namespace) -> LetterReadoutOptions | None:
- """Return letter settings, rejecting letter-only flags with native readout."""
-
- values = {
- "temperature": getattr(args, "letter_temperature", None),
- "pointer_weight": getattr(args, "letter_pointer_weight", None),
- "max_tokens": getattr(args, "letter_max_tokens", None),
- }
- if getattr(args, "readout", "native") == "native":
- used = ["--letter-" + name.replace("_", "-") for name, value in values.items() if value is not None]
- if used:
- raise ValueError(f"{', '.join(used)} require --readout letter")
- return None
- return LetterReadoutOptions(**{
- name: value for name, value in values.items() if value is not None
- })
+ """Legacy wrapper returning the historical options type."""
+
+ options = choice_options_from_args(args)
+ return (None if options is None else LetterReadoutOptions(
+ options.temperature, options.native_weight, options.max_tokens,
+ ))
__all__ = [
- "READOUTS", "LetterReadoutOptions", "add_readout_arguments", "letter_options_from_args",
+ "READOUTS", "ChoiceReadoutOptions", "LetterReadoutOptions",
+ "add_readout_arguments", "choice_options_from_args", "letter_options_from_args",
+ "normalize_readout", "resolve_readout_options",
]
diff --git a/jevany/runtime.py b/jevany/runtime.py
index 728f43b..1d1205e 100644
--- a/jevany/runtime.py
+++ b/jevany/runtime.py
@@ -14,6 +14,7 @@
from .device import default_device, sync
from .inference import InferenceOptions
from .model import DecisionModel
+from .readout import resolve_readout_options
DEFAULT_CHECKPOINT = "SimpleJev/JevAny-Qwen3.8-27B-LoRA"
@@ -165,7 +166,9 @@ def from_pretrained(
device: str | None = None, dtype: str | None = None,
model_name: str | None = None, options: LoadOptions | None = None,
inference_options: InferenceOptions | None = None,
- readout: str = "native", letter_temperature: float | None = None,
+ readout: str = "native", choice_temperature: float | None = None,
+ choice_native_weight: float | None = None, choice_max_tokens: int | None = None,
+ letter_temperature: float | None = None,
letter_pointer_weight: float | None = None, letter_max_tokens: int | None = None,
) -> "JevModel":
"""Load a local run or Hugging Face adapter ID (optionally ``owner/repo@revision``).
@@ -173,20 +176,25 @@ def from_pretrained(
The full backbone must fit on the selected device unless ``options.device_map``
(or JEVANY_DEVICE_MAP) splits it over the visible GPUs. ``dtype`` accepts
fp32, fp16 or bf16; omission uses the checkpoint/environment settings.
- ``readout='letter'`` replaces the checkpoint head with a training-free
- option-letter projection; its temperature, pointer blend, and prompt
- limit are deployment settings, not checkpoint metadata. Files used by
- native media requests are trusted local paths.
+ ``readout='choice'`` replaces the checkpoint head with a training-free
+ choice-token projection; its temperature, native blend, and prompt
+ limit are deployment settings, not checkpoint metadata. ``letter`` and
+ ``letter_*`` remain accepted legacy aliases. Files used by native media
+ requests are trusted local paths.
"""
import torch
- if readout not in ("native", "letter"):
- raise ValueError("readout must be native or letter")
- letter_settings = (letter_temperature, letter_pointer_weight, letter_max_tokens)
- if readout == "native" and any(value is not None for value in letter_settings):
- raise ValueError("letter readout options require readout='letter'")
- if readout == "letter" and inference_options is not None:
- raise ValueError("inference_options apply only to native readout; use letter_max_tokens")
+ readout, choice_options = resolve_readout_options(
+ readout,
+ choice_temperature=choice_temperature,
+ choice_native_weight=choice_native_weight,
+ choice_max_tokens=choice_max_tokens,
+ letter_temperature=letter_temperature,
+ letter_pointer_weight=letter_pointer_weight,
+ letter_max_tokens=letter_max_tokens,
+ )
+ if readout == "choice" and inference_options is not None:
+ raise ValueError("inference_options apply only to native readout; use choice_max_tokens")
device = default_device() if device is None else device
if device not in ("cpu", "mps", "cuda"):
@@ -203,21 +211,21 @@ def from_pretrained(
options = replace(options, dtype=dtypes[dtype])
if device == "mps" and options.attn is None:
options = replace(options, attn="sdpa")
- if readout == "letter" and options.cuda_graphs:
+ if readout == "choice" and options.cuda_graphs:
raise ValueError("CUDA graph capture is available only for native readout")
- if readout == "letter" and options.temperature is not None:
- raise ValueError("JEVANY_TEMPERATURE applies to the native head; use letter_temperature")
+ if readout == "choice" and options.temperature is not None:
+ raise ValueError("JEVANY_TEMPERATURE applies to the native head; use choice_temperature")
predictor = None
- if readout == "letter":
+ if readout == "choice":
from .letter_predictor import LetterReadoutPredictor
predictor = LetterReadoutPredictor(
checkpoint=checkpoint,
device=device,
options=options,
- temperature=1.0 if letter_temperature is None else letter_temperature,
- pointer_weight=0.0 if letter_pointer_weight is None else letter_pointer_weight,
- max_tokens=16_384 if letter_max_tokens is None else letter_max_tokens,
+ temperature=choice_options.temperature,
+ native_weight=choice_options.native_weight,
+ max_tokens=choice_options.max_tokens,
)
loaded = predictor.checkpoint
else:
diff --git a/jevany/serve.py b/jevany/serve.py
index a93d798..49944ff 100644
--- a/jevany/serve.py
+++ b/jevany/serve.py
@@ -3,7 +3,7 @@
"""FastAPI server for prefill-only decisions.
Run: uv run --extra serve python -m jevany.serve --run runs/rlcr --port 8008
-Use ``--readout letter`` for the training-free option-letter deployment path.
+Use ``--readout choice`` for the training-free choice-token deployment path.
TypeSafe-compatible: POST /v1/systemone and GET /v1/models (no auth). JEVANY_PREFIX_CACHE /
JEVANY_PREFIX_MIN_TOKENS size the state-prefix cache; JEVANY_DATE_FACTS=1 enables deterministic date preprocessing.
@@ -18,7 +18,13 @@
from .api import SystemOneRequest, with_date_facts
from .checkpoint import LoadOptions, add_placement_arguments, load_options_from_args
from .inference import InferenceOptions, add_inference_arguments, inference_options_from_args
-from .readout import LetterReadoutOptions, add_readout_arguments, letter_options_from_args
+from .readout import (
+ ChoiceReadoutOptions,
+ LetterReadoutOptions,
+ add_readout_arguments,
+ choice_options_from_args,
+ normalize_readout,
+)
from .runtime import DEFAULT_CHECKPOINT, DecisionRuntime, JevModel
DATE_FACTS = os.environ.get("JEVANY_DATE_FACTS", "0") == "1"
MEDIA_ROOT = os.environ.get("JEVANY_MEDIA_ROOT")
@@ -148,7 +154,8 @@ def create_app(
model: JevModel | None = None, device: str | None = None,
dtype: str | None = None, model_name: str | None = None,
options: LoadOptions | None = None, inference_options: InferenceOptions | None = None,
- readout: str = "native", letter_options: LetterReadoutOptions | None = None,
+ readout: str = "native", choice_options: ChoiceReadoutOptions | None = None,
+ letter_options: LetterReadoutOptions | None = None,
) -> FastAPI:
"""Build an isolated app, loading one checkpoint during ASGI startup.
@@ -156,18 +163,21 @@ def create_app(
and cache with Python callers. Loading options cannot accompany an injected
model. Each worker loads its own full model; use one worker per device.
"""
- if model is not None and (readout != "native" or letter_options is not None or any(
+ readout = normalize_readout(readout)
+ if choice_options is not None and letter_options is not None:
+ raise ValueError("give at most one of choice_options and legacy letter_options")
+ if letter_options is not None:
+ choice_options = letter_options.as_choice()
+ if model is not None and (readout != "native" or choice_options is not None or any(
value is not None for value in (checkpoint, device, dtype, model_name, options, inference_options)
)):
raise ValueError("pass either a loaded model or checkpoint loading options")
- if readout not in ("native", "letter"):
- raise ValueError("readout must be native or letter")
- if readout == "native" and letter_options is not None:
- raise ValueError("letter_options require readout='letter'")
- if readout == "letter" and inference_options is not None:
- raise ValueError("inference_options apply only to native readout; use letter_options.max_tokens")
- if readout == "letter" and letter_options is None:
- letter_options = LetterReadoutOptions()
+ if readout == "native" and choice_options is not None:
+ raise ValueError("choice_options require readout='choice'")
+ if readout == "choice" and inference_options is not None:
+ raise ValueError("inference_options apply only to native readout; use choice_options.max_tokens")
+ if readout == "choice" and choice_options is None:
+ choice_options = ChoiceReadoutOptions()
@asynccontextmanager
async def lifespan(application: FastAPI):
@@ -175,11 +185,11 @@ async def lifespan(application: FastAPI):
local = model
else:
settings = ({
- "readout": "letter",
- "letter_temperature": letter_options.temperature,
- "letter_pointer_weight": letter_options.pointer_weight,
- "letter_max_tokens": letter_options.max_tokens,
- } if letter_options is not None else {
+ "readout": "choice",
+ "choice_temperature": choice_options.temperature,
+ "choice_native_weight": choice_options.native_weight,
+ "choice_max_tokens": choice_options.max_tokens,
+ } if choice_options is not None else {
"inference_options": inference_options,
})
local = JevModel.from_pretrained(
@@ -221,15 +231,15 @@ def main(argv=None):
ap.add_argument("--port", type=int, default=8008)
a = ap.parse_args(argv)
try:
- letter_options = letter_options_from_args(a)
+ choice_options = choice_options_from_args(a)
except ValueError as error:
ap.error(str(error))
- if letter_options is not None and (
+ if choice_options is not None and (
a.cuda_graphs or a.cuda_graph_max_tokens is not None
or any(getattr(a, item.name) is not None for item in fields(InferenceOptions))
):
- ap.error("CUDA graph and native inference-limit flags cannot be used with --readout letter; "
- "use --letter-max-tokens")
+ ap.error("CUDA graph and native inference-limit flags cannot be used with --readout choice; "
+ "use --choice-max-tokens")
options = load_options_from_args(a)
if a.cuda_graphs or a.cuda_graph_max_tokens is not None:
options = options or LoadOptions.from_env()
@@ -239,8 +249,8 @@ def main(argv=None):
application = create_app(a.run, device=a.device, dtype=a.dtype, model_name=a.model_name,
options=options,
inference_options=(inference_options_from_args(a)
- if letter_options is None else None),
- readout=a.readout, letter_options=letter_options)
+ if choice_options is None else None),
+ readout=normalize_readout(a.readout), choice_options=choice_options)
import uvicorn
uvicorn.run(application, host=a.host, port=a.port)
diff --git a/reports/JevAny_Tech_Report.pdf b/reports/JevAny_Tech_Report.pdf
index de0ad58..cedc9fb 100644
Binary files a/reports/JevAny_Tech_Report.pdf and b/reports/JevAny_Tech_Report.pdf differ
diff --git a/results/choice-readout-v2.json b/results/choice-readout-v2.json
new file mode 100644
index 0000000..05d9a87
--- /dev/null
+++ b/results/choice-readout-v2.json
@@ -0,0 +1,2672 @@
+{
+ "schema_version": 2,
+ "artifact": "choice-readout-v2",
+ "title": "Training-free choice-token readout evaluation",
+ "generated_by": "scripts/build_choice_readout_results.py",
+ "method": {
+ "readout": "Constrain the next token to exact one-character option IDs and renormalize their probability mass; this can change argmax decisions.",
+ "training": "No additional training for the choice-token readout.",
+ "option_ids": "A-Z followed by a-z; at most 52 options.",
+ "calibration": "Scalar temperature changes probabilities but not argmax accuracy.",
+ "ensemble": "Native + choice configurations use log-linear/geometric pooling."
+ },
+ "protocol_groups": {
+ "zero_shot": {
+ "meaning": "No Typed, JevJudge, JevBench, or Transfer-test labels tune the readout configuration.",
+ "configurations": [
+ "choice_t1",
+ "native_shipped",
+ "fixed_blend_0_5"
+ ]
+ },
+ "transfer_dev_tuned": {
+ "meaning": "Weight and additional temperature are selected only on Transfer-v9 development, then frozen for every displayed evaluation panel.",
+ "configurations": [
+ "transfer_dev_calibrated_choice",
+ "transfer_dev_tuned_blend"
+ ],
+ "accuracy_note": "Temperature-only calibrated choice has the same hard accuracy as choice_t1; only a changed blend weight can change its argmax."
+ }
+ },
+ "expected_counts": {
+ "transfer_calibration": {
+ "label": "Transfer-v9 development",
+ "role": "selection_only",
+ "records": 1264,
+ "questions": 1264,
+ "headline_n": 1046
+ },
+ "transfer_test": {
+ "label": "Transfer-v9 test",
+ "role": "held_out_evaluation",
+ "records": 1264,
+ "questions": 1264,
+ "headline_n": 1046
+ },
+ "typed_test": {
+ "label": "Typed Decisions test",
+ "role": "held_out_external_evaluation",
+ "records": 400,
+ "questions": 2000,
+ "headline_n": 2000
+ },
+ "jevbench_public": {
+ "label": "JevBench public development",
+ "role": "public_diagnostic_not_used_for_selection",
+ "records": 231,
+ "questions": 231,
+ "headline_n": 231
+ },
+ "jevjudge_text": {
+ "label": "JevJudge text",
+ "role": "held_out_external_evaluation",
+ "records": 724,
+ "questions": 724,
+ "headline_n": 724
+ }
+ },
+ "dataset_hashes": {
+ "jevbench_development_sha256": "b6d6eb34fdbbf46ec11be8657da5e513e3538990d7a00ca5582c55d8bd7941c6",
+ "jevbench_manifest_sha256": "e356ad5a2c1dd3917c126e74bdee58940391568de013502abcdcd902bde3c2a5",
+ "transfer_manifest_sha256": "3c4f0be94509a3612678bfd3a30fd99a8d0ca3c47ddfe7318075d95b2fa365e4",
+ "transfer_development_sha256": "53872ba045564315cf3c3dd8fb656e76441c6c514b1a7ff985ff0db3955710cf",
+ "transfer_test_sha256": "048abea5e12c73c3a622ce5836898c00fa43766e0bf48fb2459e77ee9873497f",
+ "typed_revision": "d0e2f0c42fef86cc15d1688d25a19f5ba7c85b18",
+ "typed_test_sha256": "4f294f218ea1da27f3efef936359389c62ea4d3973a41457732990f1d31b647c",
+ "jevjudge_test_sha256": "5c50f89e9ff1beb2107648b56f2a3cbc19c946e379519caf59afc191700e83cb",
+ "jevjudge_manifest_sha256": "ee52a15f82629742093a89910a8f72e7ca97248b197aad6ec801f984c2a15b68"
+ },
+ "code_revision": "ae891ed97147b57b41a78a4bb4960b018db3bf42",
+ "source_runs": [
+ {
+ "key": "base4",
+ "model": "Frozen-Qwen3.5-4B",
+ "family": "Qwen3.5-4B",
+ "weights": "frozen_base",
+ "native_readout": null,
+ "artifacts": {
+ "manifest": {
+ "path": "/lustre-storage/fsx/tianxinwei/JevAny/runs/choice-readout-v2-20261002-r1/base4/manifest.json",
+ "sha256": "49e57b70600869d5235a3363a3238161b753b3082f9d348cbd25e810a562418b"
+ },
+ "ensemble": {
+ "path": "/lustre-storage/fsx/tianxinwei/JevAny/runs/choice-readout-v2-20261002-r1/base4/ensemble.json",
+ "sha256": "388cf7b00de00c5c5090f27c8436981cadb1a71986f6855f77806c58840b03d2"
+ }
+ },
+ "source": {
+ "base": "Qwen/Qwen3.5-4B",
+ "checkpoint": null
+ },
+ "public_source": {
+ "repository": "Qwen/Qwen3.5-4B",
+ "url": "https://huggingface.co/Qwen/Qwen3.5-4B",
+ "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a",
+ "revision_status": "verified",
+ "repository_evidence": "manifest.base_loading.canonical_base",
+ "revision_evidence": "manifest.base_loading.canonical_revision",
+ "evaluated_artifact_sha256": {},
+ "base_model": {
+ "repository": "Qwen/Qwen3.5-4B",
+ "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a"
+ },
+ "limitation": null
+ },
+ "checkpoint_artifacts": null,
+ "base_loading": {
+ "canonical_base": "Qwen/Qwen3.5-4B",
+ "canonical_base_is_local": false,
+ "canonical_revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a",
+ "override_used": true,
+ "override_config_sha256": "ddc63e1c717afa86c865bb5e01313d89d72bb53b97ad4a8a03ba8510c0621670"
+ },
+ "predictor": {
+ "method": "training-free exact choice-token projection",
+ "constraint_emulation": "exact decoded one-character option IDs",
+ "prompt": "Cygnet-compatible A-Z; extended with a-z above 26 options",
+ "system_prompt_sha256": "e2f0ac9f62eb446d7826b3abc3f7909287f4075175e9b58e4cd3ca72718671fe",
+ "chat_template_sha256": "a4aee8afcf2e0711942cf848899be66016f8d14a889ff9ede07bca099c28f715",
+ "prompt_format_version": 2,
+ "canonical_base": "Qwen/Qwen3.5-4B",
+ "canonical_revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a",
+ "checkpoint": null,
+ "adapter_applied": false,
+ "adapter_scale": 0.0,
+ "tied_output_embeddings": true,
+ "unused_full_output_head_released": false,
+ "letter_choice_token_rows": 52,
+ "output_softcap": null,
+ "temperature": 1.0,
+ "native_weight": 0.0,
+ "native_weight_argument": "native_weight",
+ "native_decision_mode": null,
+ "pointer_weight": 0.0,
+ "native_temperature": null,
+ "pointer_temperature": null,
+ "return_components": false,
+ "exact_kernels": true,
+ "exact_cuda_kernels_applied": true,
+ "efficient_long_context_tokens": 8192,
+ "long_context_attention": "fused SDPA (Flash/Efficient, no math fallback) at or above the recorded threshold; exact math SDPA below it",
+ "max_tokens": 65536,
+ "backbone_context_window": 262144,
+ "effective_max_tokens": 65536,
+ "dtype": "torch.bfloat16",
+ "device": "cuda"
+ },
+ "runtime": {
+ "python": "3.12.14",
+ "torch": "2.8.0+cu128",
+ "cuda": "12.8",
+ "transformers": "5.17.0",
+ "peft": "0.21.0",
+ "numpy": "2.5.3",
+ "pyarrow": "25.0.1",
+ "safetensors": "0.8.0",
+ "code_revision": "ae891ed97147b57b41a78a4bb4960b018db3bf42"
+ },
+ "dataset_hashes": {
+ "jevbench_development_sha256": "b6d6eb34fdbbf46ec11be8657da5e513e3538990d7a00ca5582c55d8bd7941c6",
+ "jevbench_manifest_sha256": "e356ad5a2c1dd3917c126e74bdee58940391568de013502abcdcd902bde3c2a5",
+ "transfer_manifest_sha256": "3c4f0be94509a3612678bfd3a30fd99a8d0ca3c47ddfe7318075d95b2fa365e4",
+ "transfer_development_sha256": "53872ba045564315cf3c3dd8fb656e76441c6c514b1a7ff985ff0db3955710cf",
+ "transfer_test_sha256": "048abea5e12c73c3a622ce5836898c00fa43766e0bf48fb2459e77ee9873497f",
+ "typed_revision": "d0e2f0c42fef86cc15d1688d25a19f5ba7c85b18",
+ "typed_test_sha256": "4f294f218ea1da27f3efef936359389c62ea4d3973a41457732990f1d31b647c",
+ "jevjudge_test_sha256": "5c50f89e9ff1beb2107648b56f2a3cbc19c946e379519caf59afc191700e83cb",
+ "jevjudge_manifest_sha256": "ee52a15f82629742093a89910a8f72e7ca97248b197aad6ec801f984c2a15b68"
+ },
+ "evaluation_protocol": {
+ "kernel_policy": "LocalPredictor-compatible exact CUDA policy",
+ "dtype": "bf16",
+ "attn": "sdpa",
+ "temperature": 1.0,
+ "blend_selection": "not performed here; components are saved for fitting on Transfer-v9 development only",
+ "efficiency_scope": "base calls execute choice only; checkpoint calls execute choice and native. Do not compare component efficiency from these end-to-end timings",
+ "long_context_attention": "exact math SDPA below the configured threshold; memory-linear fused SDPA without a math fallback at or above it",
+ "efficient_long_context_tokens": 8192,
+ "resume": {
+ "enabled": true,
+ "reused_datasets": [
+ "jevbench_public",
+ "transfer_calibration",
+ "transfer_test",
+ "typed_test"
+ ],
+ "preexisting_code_revision": "256001ca806f9d7e302bea967850db214f3d8e4e",
+ "dataset_code_revisions": {
+ "jevbench_public": "256001ca806f9d7e302bea967850db214f3d8e4e",
+ "transfer_calibration": "256001ca806f9d7e302bea967850db214f3d8e4e",
+ "transfer_test": "256001ca806f9d7e302bea967850db214f3d8e4e",
+ "typed_test": "256001ca806f9d7e302bea967850db214f3d8e4e",
+ "jevjudge_text": "ae891ed97147b57b41a78a4bb4960b018db3bf42"
+ }
+ }
+ },
+ "ensemble_protocol": {
+ "selection_split": "Transfer-v9 development clean/knowable accuracy cohort; never any reported test panel",
+ "selection_rows": 1046,
+ "weight_grid": "0.00 to 1.00 inclusive, step 0.01",
+ "pool": "log-linear/geometric",
+ "temperature_fit": "bounded scalar NLL minimization, T in [0.05, 20]",
+ "ece": "15 equal-width bins for Typed Decisions; 10 for every other dataset",
+ "selection_objective": "accuracy; calibrated NLL then distance from 0.5 break ties",
+ "transfer": "the selected weight and temperature are frozen for every held-out dataset",
+ "native_shipped": "the checkpoint's native probabilities, including its shipped inference temperature; no additional temperature is applied",
+ "fixed_blend_0_5": "equal log-linear pooling of T=1 choice probabilities and shipped native probabilities; no additional temperature is applied",
+ "transfer_dev_tuned_blend": "the Transfer-dev-selected native weight followed by the reported scalar additional temperature"
+ },
+ "selected": {
+ "choice_temperature": 1.8514751919432642,
+ "native_weight": 0.0,
+ "blend_temperature": 1.8514751919432642,
+ "selection_objective": "hard-label accuracy on the Transfer-v9 development clean/knowable cohort; calibrated NLL and distance from 0.5 break ties"
+ }
+ },
+ {
+ "key": "pointer4",
+ "model": "JevAny-Qwen3.5-4B-Pointer",
+ "family": "Qwen3.5-4B",
+ "weights": "jevany_sft",
+ "native_readout": "pointer",
+ "artifacts": {
+ "manifest": {
+ "path": "/lustre-storage/fsx/tianxinwei/JevAny/runs/choice-readout-v2-20261002-r1/pointer4/manifest.json",
+ "sha256": "11b64050dc31f834c64dbe54c975df1fb365b9eb9465a4089e7de98ba00242b1"
+ },
+ "ensemble": {
+ "path": "/lustre-storage/fsx/tianxinwei/JevAny/runs/choice-readout-v2-20261002-r1/pointer4/ensemble.json",
+ "sha256": "0e18398586de8760719ea6a3e8956f221a69e341e75d3364e22fda7df9cb6488"
+ }
+ },
+ "source": {
+ "base": null,
+ "checkpoint": "/storage/home/tianxinwei/.cache/huggingface/hub/models--SimpleJev--JevAny-Qwen3.5-4B-LoRA/snapshots/1c7aa9bab14ac347aeb917c0bcd757838a8a78ce"
+ },
+ "public_source": {
+ "repository": "SimpleJev/JevAny-Qwen3.5-4B-LoRA",
+ "url": "https://huggingface.co/SimpleJev/JevAny-Qwen3.5-4B-LoRA",
+ "revision": "1c7aa9bab14ac347aeb917c0bcd757838a8a78ce",
+ "revision_status": "verified",
+ "repository_evidence": "manifest.source.checkpoint",
+ "revision_evidence": "manifest.source.checkpoint",
+ "evaluated_artifact_sha256": {
+ "adapter_model.safetensors": {
+ "sha256": "5acbaec8d554517250bf9c626b3f4b9e2ec31fc67957f3486aa72e2c91ee907f",
+ "bytes": 64993432
+ },
+ "head.pt": {
+ "sha256": "e5ea982efd06ecaf71b493045ebccdbfdca39f6fe128de164bf636d73152cd44",
+ "bytes": 5513315
+ },
+ "adapter_config.json": {
+ "sha256": "8a138b2c091d2ba463a8107e9e19182202395188750742047f4a12e8b59e1eda",
+ "bytes": 1265
+ }
+ },
+ "base_model": {
+ "repository": "Qwen/Qwen3.5-4B",
+ "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a"
+ },
+ "limitation": null
+ },
+ "checkpoint_artifacts": {
+ "requested": "/storage/home/tianxinwei/.cache/huggingface/hub/models--SimpleJev--JevAny-Qwen3.5-4B-LoRA/snapshots/1c7aa9bab14ac347aeb917c0bcd757838a8a78ce",
+ "resolved": "/storage/home/tianxinwei/.cache/huggingface/hub/models--SimpleJev--JevAny-Qwen3.5-4B-LoRA/snapshots/1c7aa9bab14ac347aeb917c0bcd757838a8a78ce",
+ "files": {
+ "adapter_model.safetensors": {
+ "sha256": "5acbaec8d554517250bf9c626b3f4b9e2ec31fc67957f3486aa72e2c91ee907f",
+ "bytes": 64993432
+ },
+ "head.pt": {
+ "sha256": "e5ea982efd06ecaf71b493045ebccdbfdca39f6fe128de164bf636d73152cd44",
+ "bytes": 5513315
+ },
+ "adapter_config.json": {
+ "sha256": "8a138b2c091d2ba463a8107e9e19182202395188750742047f4a12e8b59e1eda",
+ "bytes": 1265
+ }
+ }
+ },
+ "base_loading": {
+ "canonical_base": "Qwen/Qwen3.5-4B",
+ "canonical_base_is_local": false,
+ "canonical_revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a",
+ "override_used": true,
+ "override_config_sha256": "ddc63e1c717afa86c865bb5e01313d89d72bb53b97ad4a8a03ba8510c0621670"
+ },
+ "predictor": {
+ "method": "training-free exact choice-token projection",
+ "constraint_emulation": "exact decoded one-character option IDs",
+ "prompt": "Cygnet-compatible A-Z; extended with a-z above 26 options",
+ "system_prompt_sha256": "e2f0ac9f62eb446d7826b3abc3f7909287f4075175e9b58e4cd3ca72718671fe",
+ "chat_template_sha256": "a4aee8afcf2e0711942cf848899be66016f8d14a889ff9ede07bca099c28f715",
+ "prompt_format_version": 2,
+ "canonical_base": "Qwen/Qwen3.5-4B",
+ "canonical_revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a",
+ "checkpoint": "/storage/home/tianxinwei/.cache/huggingface/hub/models--SimpleJev--JevAny-Qwen3.5-4B-LoRA/snapshots/1c7aa9bab14ac347aeb917c0bcd757838a8a78ce",
+ "adapter_applied": true,
+ "adapter_scale": 1.0,
+ "tied_output_embeddings": true,
+ "unused_full_output_head_released": false,
+ "letter_choice_token_rows": 52,
+ "output_softcap": null,
+ "temperature": 1.0,
+ "native_weight": 0.5,
+ "native_weight_argument": "native_weight",
+ "native_decision_mode": "pointer",
+ "pointer_weight": 0.5,
+ "native_temperature": 1.0,
+ "pointer_temperature": 1.0,
+ "return_components": true,
+ "exact_kernels": true,
+ "exact_cuda_kernels_applied": true,
+ "efficient_long_context_tokens": 8192,
+ "long_context_attention": "fused SDPA (Flash/Efficient, no math fallback) at or above the recorded threshold; exact math SDPA below it",
+ "max_tokens": 65536,
+ "backbone_context_window": 262144,
+ "effective_max_tokens": 65536,
+ "dtype": "torch.bfloat16",
+ "device": "cuda"
+ },
+ "runtime": {
+ "python": "3.12.14",
+ "torch": "2.8.0+cu128",
+ "cuda": "12.8",
+ "transformers": "5.17.0",
+ "peft": "0.21.0",
+ "numpy": "2.5.3",
+ "pyarrow": "25.0.1",
+ "safetensors": "0.8.0",
+ "code_revision": "ae891ed97147b57b41a78a4bb4960b018db3bf42"
+ },
+ "dataset_hashes": {
+ "jevbench_development_sha256": "b6d6eb34fdbbf46ec11be8657da5e513e3538990d7a00ca5582c55d8bd7941c6",
+ "jevbench_manifest_sha256": "e356ad5a2c1dd3917c126e74bdee58940391568de013502abcdcd902bde3c2a5",
+ "transfer_manifest_sha256": "3c4f0be94509a3612678bfd3a30fd99a8d0ca3c47ddfe7318075d95b2fa365e4",
+ "transfer_development_sha256": "53872ba045564315cf3c3dd8fb656e76441c6c514b1a7ff985ff0db3955710cf",
+ "transfer_test_sha256": "048abea5e12c73c3a622ce5836898c00fa43766e0bf48fb2459e77ee9873497f",
+ "typed_revision": "d0e2f0c42fef86cc15d1688d25a19f5ba7c85b18",
+ "typed_test_sha256": "4f294f218ea1da27f3efef936359389c62ea4d3973a41457732990f1d31b647c",
+ "jevjudge_test_sha256": "5c50f89e9ff1beb2107648b56f2a3cbc19c946e379519caf59afc191700e83cb",
+ "jevjudge_manifest_sha256": "ee52a15f82629742093a89910a8f72e7ca97248b197aad6ec801f984c2a15b68"
+ },
+ "evaluation_protocol": {
+ "kernel_policy": "LocalPredictor-compatible exact CUDA policy",
+ "dtype": "bf16",
+ "attn": "sdpa",
+ "temperature": 1.0,
+ "blend_selection": "not performed here; components are saved for fitting on Transfer-v9 development only",
+ "efficiency_scope": "base calls execute choice only; checkpoint calls execute choice and native. Do not compare component efficiency from these end-to-end timings",
+ "long_context_attention": "exact math SDPA below the configured threshold; memory-linear fused SDPA without a math fallback at or above it",
+ "efficient_long_context_tokens": 8192,
+ "resume": {
+ "enabled": true,
+ "reused_datasets": [
+ "jevbench_public",
+ "transfer_calibration",
+ "transfer_test",
+ "typed_test"
+ ],
+ "preexisting_code_revision": "256001ca806f9d7e302bea967850db214f3d8e4e",
+ "dataset_code_revisions": {
+ "jevbench_public": "256001ca806f9d7e302bea967850db214f3d8e4e",
+ "transfer_calibration": "256001ca806f9d7e302bea967850db214f3d8e4e",
+ "transfer_test": "256001ca806f9d7e302bea967850db214f3d8e4e",
+ "typed_test": "256001ca806f9d7e302bea967850db214f3d8e4e",
+ "jevjudge_text": "ae891ed97147b57b41a78a4bb4960b018db3bf42"
+ }
+ }
+ },
+ "ensemble_protocol": {
+ "selection_split": "Transfer-v9 development clean/knowable accuracy cohort; never any reported test panel",
+ "selection_rows": 1046,
+ "weight_grid": "0.00 to 1.00 inclusive, step 0.01",
+ "pool": "log-linear/geometric",
+ "temperature_fit": "bounded scalar NLL minimization, T in [0.05, 20]",
+ "ece": "15 equal-width bins for Typed Decisions; 10 for every other dataset",
+ "selection_objective": "accuracy; calibrated NLL then distance from 0.5 break ties",
+ "transfer": "the selected weight and temperature are frozen for every held-out dataset",
+ "native_shipped": "the checkpoint's native probabilities, including its shipped inference temperature; no additional temperature is applied",
+ "fixed_blend_0_5": "equal log-linear pooling of T=1 choice probabilities and shipped native probabilities; no additional temperature is applied",
+ "transfer_dev_tuned_blend": "the Transfer-dev-selected native weight followed by the reported scalar additional temperature"
+ },
+ "selected": {
+ "choice_temperature": 0.8505257300716687,
+ "native_weight": 0.78,
+ "blend_temperature": 1.0977784557618349,
+ "selection_objective": "hard-label accuracy on the Transfer-v9 development clean/knowable cohort; calibrated NLL and distance from 0.5 break ties"
+ }
+ },
+ {
+ "key": "direct4",
+ "model": "JevAny-Qwen3.5-4B-Direct-Token",
+ "family": "Qwen3.5-4B",
+ "weights": "jevany_sft",
+ "native_readout": "direct-token",
+ "artifacts": {
+ "manifest": {
+ "path": "/lustre-storage/fsx/tianxinwei/JevAny/runs/choice-readout-v2-20261002-r1/direct4/manifest.json",
+ "sha256": "244ede1e22b445634be5792977e19e594bfaeb2fe783f905a6ccf2a0424078e5"
+ },
+ "ensemble": {
+ "path": "/lustre-storage/fsx/tianxinwei/JevAny/runs/choice-readout-v2-20261002-r1/direct4/ensemble.json",
+ "sha256": "e3a2aee6ca34b98a0780807b6b49c86dcf261b55e5559da8fc61dd17b9f03db7"
+ }
+ },
+ "source": {
+ "base": null,
+ "checkpoint": "/lustre-storage/fsx/tianxinwei/JevAny/releases/model-family-v2-20260929/JevAny-Qwen3.5-4B-Direct-Token-LoRA"
+ },
+ "public_source": {
+ "repository": "SimpleJev/JevAny-Qwen3.5-4B-Direct-Token-LoRA",
+ "url": "https://huggingface.co/SimpleJev/JevAny-Qwen3.5-4B-Direct-Token-LoRA",
+ "revision": null,
+ "revision_status": "not_verified",
+ "repository_evidence": "results/model-family-v2.json#released_models",
+ "revision_evidence": null,
+ "evaluated_artifact_sha256": {
+ "adapter_model.safetensors": {
+ "sha256": "b768d54b9e5e20204b2db9e24c18ae6b10fa310b241d1931bac6b8882f127c53",
+ "bytes": 64993432
+ },
+ "head.pt": {
+ "sha256": "d3909cffc156d077061114627c8aa22f60c7a8f1bdcfd20c489bb3bb009fc4a6",
+ "bytes": 4699
+ },
+ "adapter_config.json": {
+ "sha256": "f84527483265b514d0a5f68d9c44e3d7952825c61d3659a858af8342e6ce9ff2",
+ "bytes": 1265
+ }
+ },
+ "base_model": {
+ "repository": "Qwen/Qwen3.5-4B",
+ "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a"
+ },
+ "limitation": "The evaluated checkpoint came from a local release. Its exact files are pinned below by SHA-256, but the run manifest and release metadata do not prove an immutable public-repository revision for those bytes."
+ },
+ "checkpoint_artifacts": {
+ "requested": "/lustre-storage/fsx/tianxinwei/JevAny/releases/model-family-v2-20260929/JevAny-Qwen3.5-4B-Direct-Token-LoRA",
+ "resolved": "/lustre-storage/fsx/tianxinwei/JevAny/releases/model-family-v2-20260929/JevAny-Qwen3.5-4B-Direct-Token-LoRA",
+ "files": {
+ "adapter_model.safetensors": {
+ "sha256": "b768d54b9e5e20204b2db9e24c18ae6b10fa310b241d1931bac6b8882f127c53",
+ "bytes": 64993432
+ },
+ "head.pt": {
+ "sha256": "d3909cffc156d077061114627c8aa22f60c7a8f1bdcfd20c489bb3bb009fc4a6",
+ "bytes": 4699
+ },
+ "adapter_config.json": {
+ "sha256": "f84527483265b514d0a5f68d9c44e3d7952825c61d3659a858af8342e6ce9ff2",
+ "bytes": 1265
+ }
+ }
+ },
+ "base_loading": {
+ "canonical_base": "Qwen/Qwen3.5-4B",
+ "canonical_base_is_local": false,
+ "canonical_revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a",
+ "override_used": true,
+ "override_config_sha256": "ddc63e1c717afa86c865bb5e01313d89d72bb53b97ad4a8a03ba8510c0621670"
+ },
+ "predictor": {
+ "method": "training-free exact choice-token projection",
+ "constraint_emulation": "exact decoded one-character option IDs",
+ "prompt": "Cygnet-compatible A-Z; extended with a-z above 26 options",
+ "system_prompt_sha256": "e2f0ac9f62eb446d7826b3abc3f7909287f4075175e9b58e4cd3ca72718671fe",
+ "chat_template_sha256": "a4aee8afcf2e0711942cf848899be66016f8d14a889ff9ede07bca099c28f715",
+ "prompt_format_version": 2,
+ "canonical_base": "Qwen/Qwen3.5-4B",
+ "canonical_revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a",
+ "checkpoint": "/lustre-storage/fsx/tianxinwei/JevAny/releases/model-family-v2-20260929/JevAny-Qwen3.5-4B-Direct-Token-LoRA",
+ "adapter_applied": true,
+ "adapter_scale": 1.0,
+ "tied_output_embeddings": true,
+ "unused_full_output_head_released": false,
+ "letter_choice_token_rows": 52,
+ "output_softcap": null,
+ "temperature": 1.0,
+ "native_weight": 0.5,
+ "native_weight_argument": "native_weight",
+ "native_decision_mode": "lm_token",
+ "pointer_weight": 0.5,
+ "native_temperature": 1.0,
+ "pointer_temperature": 1.0,
+ "return_components": true,
+ "exact_kernels": true,
+ "exact_cuda_kernels_applied": true,
+ "efficient_long_context_tokens": 8192,
+ "long_context_attention": "fused SDPA (Flash/Efficient, no math fallback) at or above the recorded threshold; exact math SDPA below it",
+ "max_tokens": 65536,
+ "backbone_context_window": 262144,
+ "effective_max_tokens": 65536,
+ "dtype": "torch.bfloat16",
+ "device": "cuda"
+ },
+ "runtime": {
+ "python": "3.12.14",
+ "torch": "2.8.0+cu128",
+ "cuda": "12.8",
+ "transformers": "5.17.0",
+ "peft": "0.21.0",
+ "numpy": "2.5.3",
+ "pyarrow": "25.0.1",
+ "safetensors": "0.8.0",
+ "code_revision": "ae891ed97147b57b41a78a4bb4960b018db3bf42"
+ },
+ "dataset_hashes": {
+ "jevbench_development_sha256": "b6d6eb34fdbbf46ec11be8657da5e513e3538990d7a00ca5582c55d8bd7941c6",
+ "jevbench_manifest_sha256": "e356ad5a2c1dd3917c126e74bdee58940391568de013502abcdcd902bde3c2a5",
+ "transfer_manifest_sha256": "3c4f0be94509a3612678bfd3a30fd99a8d0ca3c47ddfe7318075d95b2fa365e4",
+ "transfer_development_sha256": "53872ba045564315cf3c3dd8fb656e76441c6c514b1a7ff985ff0db3955710cf",
+ "transfer_test_sha256": "048abea5e12c73c3a622ce5836898c00fa43766e0bf48fb2459e77ee9873497f",
+ "typed_revision": "d0e2f0c42fef86cc15d1688d25a19f5ba7c85b18",
+ "typed_test_sha256": "4f294f218ea1da27f3efef936359389c62ea4d3973a41457732990f1d31b647c",
+ "jevjudge_test_sha256": "5c50f89e9ff1beb2107648b56f2a3cbc19c946e379519caf59afc191700e83cb",
+ "jevjudge_manifest_sha256": "ee52a15f82629742093a89910a8f72e7ca97248b197aad6ec801f984c2a15b68"
+ },
+ "evaluation_protocol": {
+ "kernel_policy": "LocalPredictor-compatible exact CUDA policy",
+ "dtype": "bf16",
+ "attn": "sdpa",
+ "temperature": 1.0,
+ "blend_selection": "not performed here; components are saved for fitting on Transfer-v9 development only",
+ "efficiency_scope": "base calls execute choice only; checkpoint calls execute choice and native. Do not compare component efficiency from these end-to-end timings",
+ "long_context_attention": "exact math SDPA below the configured threshold; memory-linear fused SDPA without a math fallback at or above it",
+ "efficient_long_context_tokens": 8192,
+ "resume": {
+ "enabled": true,
+ "reused_datasets": [
+ "jevbench_public",
+ "transfer_calibration",
+ "transfer_test",
+ "typed_test"
+ ],
+ "preexisting_code_revision": "256001ca806f9d7e302bea967850db214f3d8e4e",
+ "dataset_code_revisions": {
+ "jevbench_public": "256001ca806f9d7e302bea967850db214f3d8e4e",
+ "transfer_calibration": "256001ca806f9d7e302bea967850db214f3d8e4e",
+ "transfer_test": "256001ca806f9d7e302bea967850db214f3d8e4e",
+ "typed_test": "256001ca806f9d7e302bea967850db214f3d8e4e",
+ "jevjudge_text": "ae891ed97147b57b41a78a4bb4960b018db3bf42"
+ }
+ }
+ },
+ "ensemble_protocol": {
+ "selection_split": "Transfer-v9 development clean/knowable accuracy cohort; never any reported test panel",
+ "selection_rows": 1046,
+ "weight_grid": "0.00 to 1.00 inclusive, step 0.01",
+ "pool": "log-linear/geometric",
+ "temperature_fit": "bounded scalar NLL minimization, T in [0.05, 20]",
+ "ece": "15 equal-width bins for Typed Decisions; 10 for every other dataset",
+ "selection_objective": "accuracy; calibrated NLL then distance from 0.5 break ties",
+ "transfer": "the selected weight and temperature are frozen for every held-out dataset",
+ "native_shipped": "the checkpoint's native probabilities, including its shipped inference temperature; no additional temperature is applied",
+ "fixed_blend_0_5": "equal log-linear pooling of T=1 choice probabilities and shipped native probabilities; no additional temperature is applied",
+ "transfer_dev_tuned_blend": "the Transfer-dev-selected native weight followed by the reported scalar additional temperature"
+ },
+ "selected": {
+ "choice_temperature": 1.246278988841639,
+ "native_weight": 0.59,
+ "blend_temperature": 1.1240763000234277,
+ "selection_objective": "hard-label accuracy on the Transfer-v9 development clean/knowable cohort; calibrated NLL and distance from 0.5 break ties"
+ }
+ },
+ {
+ "key": "base27",
+ "model": "Frozen-Qwen3.8-27B",
+ "family": "Qwen3.8-27B",
+ "weights": "frozen_base",
+ "native_readout": null,
+ "artifacts": {
+ "manifest": {
+ "path": "/lustre-storage/fsx/tianxinwei/JevAny/runs/choice-readout-v2-20261002-r1/base27/manifest.json",
+ "sha256": "5f1d1c6a755bdb8f8ee5d37f011345213cbfa205256457622d08a88f88d81fd2"
+ },
+ "ensemble": {
+ "path": "/lustre-storage/fsx/tianxinwei/JevAny/runs/choice-readout-v2-20261002-r1/base27/ensemble.json",
+ "sha256": "c18a65303093f9673f22cc7a9fcbd7de89a73cd46372b04311e935f01277e6fb"
+ }
+ },
+ "source": {
+ "base": "Qwen/Qwen3.8-27B",
+ "checkpoint": null
+ },
+ "public_source": {
+ "repository": "Qwen/Qwen3.8-27B",
+ "url": "https://huggingface.co/Qwen/Qwen3.8-27B",
+ "revision": "1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0",
+ "revision_status": "verified",
+ "repository_evidence": "manifest.base_loading.canonical_base",
+ "revision_evidence": "manifest.base_loading.canonical_revision",
+ "evaluated_artifact_sha256": {},
+ "base_model": {
+ "repository": "Qwen/Qwen3.8-27B",
+ "revision": "1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0"
+ },
+ "limitation": null
+ },
+ "checkpoint_artifacts": null,
+ "base_loading": {
+ "canonical_base": "Qwen/Qwen3.8-27B",
+ "canonical_base_is_local": false,
+ "canonical_revision": "1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0",
+ "override_used": true,
+ "override_config_sha256": "191e0af232104ed8b65258cf3fb2b842e288008baca7633c11b82a1ac7203aab"
+ },
+ "predictor": {
+ "method": "training-free exact choice-token projection",
+ "constraint_emulation": "exact decoded one-character option IDs",
+ "prompt": "Cygnet-compatible A-Z; extended with a-z above 26 options",
+ "system_prompt_sha256": "e2f0ac9f62eb446d7826b3abc3f7909287f4075175e9b58e4cd3ca72718671fe",
+ "chat_template_sha256": "c3cf9e34abf4f9e36c2d72165aa9c132d3e2a725b6c2586aaa3a8af9d7a81041",
+ "prompt_format_version": 2,
+ "canonical_base": "Qwen/Qwen3.8-27B",
+ "canonical_revision": "1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0",
+ "checkpoint": null,
+ "adapter_applied": false,
+ "adapter_scale": 0.0,
+ "tied_output_embeddings": false,
+ "unused_full_output_head_released": false,
+ "letter_choice_token_rows": 52,
+ "output_softcap": null,
+ "temperature": 1.0,
+ "native_weight": 0.0,
+ "native_weight_argument": "native_weight",
+ "native_decision_mode": null,
+ "pointer_weight": 0.0,
+ "native_temperature": null,
+ "pointer_temperature": null,
+ "return_components": false,
+ "exact_kernels": true,
+ "exact_cuda_kernels_applied": true,
+ "efficient_long_context_tokens": 8192,
+ "long_context_attention": "fused SDPA (Flash/Efficient, no math fallback) at or above the recorded threshold; exact math SDPA below it",
+ "max_tokens": 65536,
+ "backbone_context_window": 262144,
+ "effective_max_tokens": 65536,
+ "dtype": "torch.bfloat16",
+ "device": "cuda"
+ },
+ "runtime": {
+ "python": "3.12.14",
+ "torch": "2.8.0+cu128",
+ "cuda": "12.8",
+ "transformers": "5.17.0",
+ "peft": "0.21.0",
+ "numpy": "2.5.3",
+ "pyarrow": "25.0.1",
+ "safetensors": "0.8.0",
+ "code_revision": "ae891ed97147b57b41a78a4bb4960b018db3bf42"
+ },
+ "dataset_hashes": {
+ "jevbench_development_sha256": "b6d6eb34fdbbf46ec11be8657da5e513e3538990d7a00ca5582c55d8bd7941c6",
+ "jevbench_manifest_sha256": "e356ad5a2c1dd3917c126e74bdee58940391568de013502abcdcd902bde3c2a5",
+ "transfer_manifest_sha256": "3c4f0be94509a3612678bfd3a30fd99a8d0ca3c47ddfe7318075d95b2fa365e4",
+ "transfer_development_sha256": "53872ba045564315cf3c3dd8fb656e76441c6c514b1a7ff985ff0db3955710cf",
+ "transfer_test_sha256": "048abea5e12c73c3a622ce5836898c00fa43766e0bf48fb2459e77ee9873497f",
+ "typed_revision": "d0e2f0c42fef86cc15d1688d25a19f5ba7c85b18",
+ "typed_test_sha256": "4f294f218ea1da27f3efef936359389c62ea4d3973a41457732990f1d31b647c",
+ "jevjudge_test_sha256": "5c50f89e9ff1beb2107648b56f2a3cbc19c946e379519caf59afc191700e83cb",
+ "jevjudge_manifest_sha256": "ee52a15f82629742093a89910a8f72e7ca97248b197aad6ec801f984c2a15b68"
+ },
+ "evaluation_protocol": {
+ "kernel_policy": "LocalPredictor-compatible exact CUDA policy",
+ "dtype": "bf16",
+ "attn": "sdpa",
+ "temperature": 1.0,
+ "blend_selection": "not performed here; components are saved for fitting on Transfer-v9 development only",
+ "efficiency_scope": "base calls execute choice only; checkpoint calls execute choice and native. Do not compare component efficiency from these end-to-end timings",
+ "long_context_attention": "exact math SDPA below the configured threshold; memory-linear fused SDPA without a math fallback at or above it",
+ "efficient_long_context_tokens": 8192,
+ "resume": {
+ "enabled": true,
+ "reused_datasets": [
+ "jevbench_public",
+ "transfer_calibration",
+ "transfer_test",
+ "typed_test"
+ ],
+ "preexisting_code_revision": "256001ca806f9d7e302bea967850db214f3d8e4e",
+ "dataset_code_revisions": {
+ "jevbench_public": "256001ca806f9d7e302bea967850db214f3d8e4e",
+ "transfer_calibration": "256001ca806f9d7e302bea967850db214f3d8e4e",
+ "transfer_test": "256001ca806f9d7e302bea967850db214f3d8e4e",
+ "typed_test": "256001ca806f9d7e302bea967850db214f3d8e4e",
+ "jevjudge_text": "ae891ed97147b57b41a78a4bb4960b018db3bf42"
+ }
+ }
+ },
+ "ensemble_protocol": {
+ "selection_split": "Transfer-v9 development clean/knowable accuracy cohort; never any reported test panel",
+ "selection_rows": 1046,
+ "weight_grid": "0.00 to 1.00 inclusive, step 0.01",
+ "pool": "log-linear/geometric",
+ "temperature_fit": "bounded scalar NLL minimization, T in [0.05, 20]",
+ "ece": "15 equal-width bins for Typed Decisions; 10 for every other dataset",
+ "selection_objective": "accuracy; calibrated NLL then distance from 0.5 break ties",
+ "transfer": "the selected weight and temperature are frozen for every held-out dataset",
+ "native_shipped": "the checkpoint's native probabilities, including its shipped inference temperature; no additional temperature is applied",
+ "fixed_blend_0_5": "equal log-linear pooling of T=1 choice probabilities and shipped native probabilities; no additional temperature is applied",
+ "transfer_dev_tuned_blend": "the Transfer-dev-selected native weight followed by the reported scalar additional temperature"
+ },
+ "selected": {
+ "choice_temperature": 1.718360083457113,
+ "native_weight": 0.0,
+ "blend_temperature": 1.718360083457113,
+ "selection_objective": "hard-label accuracy on the Transfer-v9 development clean/knowable cohort; calibrated NLL and distance from 0.5 break ties"
+ }
+ },
+ {
+ "key": "pointer27",
+ "model": "JevAny-Qwen3.8-27B-Pointer",
+ "family": "Qwen3.8-27B",
+ "weights": "jevany_sft",
+ "native_readout": "pointer",
+ "artifacts": {
+ "manifest": {
+ "path": "/lustre-storage/fsx/tianxinwei/JevAny/runs/choice-readout-v2-20261002-r1/pointer27/manifest.json",
+ "sha256": "c2d6dd8bb4b9f2b9e5b9c57088f673e0839af5435acb88b561559ae0a824367a"
+ },
+ "ensemble": {
+ "path": "/lustre-storage/fsx/tianxinwei/JevAny/runs/choice-readout-v2-20261002-r1/pointer27/ensemble.json",
+ "sha256": "2270c1a3309b20fbe77bf6e905e2677ba29ecc3a3c9d4251b9932ee887e64826"
+ }
+ },
+ "source": {
+ "base": null,
+ "checkpoint": "/storage/home/tianxinwei/.cache/huggingface/hub/models--SimpleJev--JevAny-Qwen3.8-27B-LoRA/snapshots/09c9e9102d5b8cc7d56558d25da1202a761b6c0d"
+ },
+ "public_source": {
+ "repository": "SimpleJev/JevAny-Qwen3.8-27B-LoRA",
+ "url": "https://huggingface.co/SimpleJev/JevAny-Qwen3.8-27B-LoRA",
+ "revision": "09c9e9102d5b8cc7d56558d25da1202a761b6c0d",
+ "revision_status": "verified",
+ "repository_evidence": "manifest.source.checkpoint",
+ "revision_evidence": "manifest.source.checkpoint",
+ "evaluated_artifact_sha256": {
+ "adapter_model.safetensors": {
+ "sha256": "f4b835f4f7be78a819067dac226feaad2fc7b9eb0b968c5e6e665ad5328e6731",
+ "bytes": 233584648
+ },
+ "head.pt": {
+ "sha256": "f78fa283f7e979ce418c5b137d24459350fecaea73e5766894532c1cd99684c7",
+ "bytes": 10756195
+ },
+ "adapter_config.json": {
+ "sha256": "31d8e262a21a45b80fd409f08d47930effd6cd9bc12156c59ad61e627f50a275",
+ "bytes": 1266
+ }
+ },
+ "base_model": {
+ "repository": "Qwen/Qwen3.8-27B",
+ "revision": "1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0"
+ },
+ "limitation": null
+ },
+ "checkpoint_artifacts": {
+ "requested": "/storage/home/tianxinwei/.cache/huggingface/hub/models--SimpleJev--JevAny-Qwen3.8-27B-LoRA/snapshots/09c9e9102d5b8cc7d56558d25da1202a761b6c0d",
+ "resolved": "/storage/home/tianxinwei/.cache/huggingface/hub/models--SimpleJev--JevAny-Qwen3.8-27B-LoRA/snapshots/09c9e9102d5b8cc7d56558d25da1202a761b6c0d",
+ "files": {
+ "adapter_model.safetensors": {
+ "sha256": "f4b835f4f7be78a819067dac226feaad2fc7b9eb0b968c5e6e665ad5328e6731",
+ "bytes": 233584648
+ },
+ "head.pt": {
+ "sha256": "f78fa283f7e979ce418c5b137d24459350fecaea73e5766894532c1cd99684c7",
+ "bytes": 10756195
+ },
+ "adapter_config.json": {
+ "sha256": "31d8e262a21a45b80fd409f08d47930effd6cd9bc12156c59ad61e627f50a275",
+ "bytes": 1266
+ }
+ }
+ },
+ "base_loading": {
+ "canonical_base": "Qwen/Qwen3.8-27B",
+ "canonical_base_is_local": false,
+ "canonical_revision": "1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0",
+ "override_used": true,
+ "override_config_sha256": "191e0af232104ed8b65258cf3fb2b842e288008baca7633c11b82a1ac7203aab"
+ },
+ "predictor": {
+ "method": "training-free exact choice-token projection",
+ "constraint_emulation": "exact decoded one-character option IDs",
+ "prompt": "Cygnet-compatible A-Z; extended with a-z above 26 options",
+ "system_prompt_sha256": "e2f0ac9f62eb446d7826b3abc3f7909287f4075175e9b58e4cd3ca72718671fe",
+ "chat_template_sha256": "c3cf9e34abf4f9e36c2d72165aa9c132d3e2a725b6c2586aaa3a8af9d7a81041",
+ "prompt_format_version": 2,
+ "canonical_base": "Qwen/Qwen3.8-27B",
+ "canonical_revision": "1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0",
+ "checkpoint": "/storage/home/tianxinwei/.cache/huggingface/hub/models--SimpleJev--JevAny-Qwen3.8-27B-LoRA/snapshots/09c9e9102d5b8cc7d56558d25da1202a761b6c0d",
+ "adapter_applied": true,
+ "adapter_scale": 1.0,
+ "tied_output_embeddings": false,
+ "unused_full_output_head_released": false,
+ "letter_choice_token_rows": 52,
+ "output_softcap": null,
+ "temperature": 1.0,
+ "native_weight": 0.5,
+ "native_weight_argument": "native_weight",
+ "native_decision_mode": "pointer",
+ "pointer_weight": 0.5,
+ "native_temperature": 1.0,
+ "pointer_temperature": 1.0,
+ "return_components": true,
+ "exact_kernels": true,
+ "exact_cuda_kernels_applied": true,
+ "efficient_long_context_tokens": 8192,
+ "long_context_attention": "fused SDPA (Flash/Efficient, no math fallback) at or above the recorded threshold; exact math SDPA below it",
+ "max_tokens": 65536,
+ "backbone_context_window": 262144,
+ "effective_max_tokens": 65536,
+ "dtype": "torch.bfloat16",
+ "device": "cuda"
+ },
+ "runtime": {
+ "python": "3.12.14",
+ "torch": "2.8.0+cu128",
+ "cuda": "12.8",
+ "transformers": "5.17.0",
+ "peft": "0.21.0",
+ "numpy": "2.5.3",
+ "pyarrow": "25.0.1",
+ "safetensors": "0.8.0",
+ "code_revision": "ae891ed97147b57b41a78a4bb4960b018db3bf42"
+ },
+ "dataset_hashes": {
+ "jevbench_development_sha256": "b6d6eb34fdbbf46ec11be8657da5e513e3538990d7a00ca5582c55d8bd7941c6",
+ "jevbench_manifest_sha256": "e356ad5a2c1dd3917c126e74bdee58940391568de013502abcdcd902bde3c2a5",
+ "transfer_manifest_sha256": "3c4f0be94509a3612678bfd3a30fd99a8d0ca3c47ddfe7318075d95b2fa365e4",
+ "transfer_development_sha256": "53872ba045564315cf3c3dd8fb656e76441c6c514b1a7ff985ff0db3955710cf",
+ "transfer_test_sha256": "048abea5e12c73c3a622ce5836898c00fa43766e0bf48fb2459e77ee9873497f",
+ "typed_revision": "d0e2f0c42fef86cc15d1688d25a19f5ba7c85b18",
+ "typed_test_sha256": "4f294f218ea1da27f3efef936359389c62ea4d3973a41457732990f1d31b647c",
+ "jevjudge_test_sha256": "5c50f89e9ff1beb2107648b56f2a3cbc19c946e379519caf59afc191700e83cb",
+ "jevjudge_manifest_sha256": "ee52a15f82629742093a89910a8f72e7ca97248b197aad6ec801f984c2a15b68"
+ },
+ "evaluation_protocol": {
+ "kernel_policy": "LocalPredictor-compatible exact CUDA policy",
+ "dtype": "bf16",
+ "attn": "sdpa",
+ "temperature": 1.0,
+ "blend_selection": "not performed here; components are saved for fitting on Transfer-v9 development only",
+ "efficiency_scope": "base calls execute choice only; checkpoint calls execute choice and native. Do not compare component efficiency from these end-to-end timings",
+ "long_context_attention": "exact math SDPA below the configured threshold; memory-linear fused SDPA without a math fallback at or above it",
+ "efficient_long_context_tokens": 8192,
+ "resume": {
+ "enabled": true,
+ "reused_datasets": [
+ "jevbench_public",
+ "transfer_calibration",
+ "transfer_test",
+ "typed_test"
+ ],
+ "preexisting_code_revision": "256001ca806f9d7e302bea967850db214f3d8e4e",
+ "dataset_code_revisions": {
+ "jevbench_public": "256001ca806f9d7e302bea967850db214f3d8e4e",
+ "transfer_calibration": "256001ca806f9d7e302bea967850db214f3d8e4e",
+ "transfer_test": "256001ca806f9d7e302bea967850db214f3d8e4e",
+ "typed_test": "256001ca806f9d7e302bea967850db214f3d8e4e",
+ "jevjudge_text": "ae891ed97147b57b41a78a4bb4960b018db3bf42"
+ }
+ }
+ },
+ "ensemble_protocol": {
+ "selection_split": "Transfer-v9 development clean/knowable accuracy cohort; never any reported test panel",
+ "selection_rows": 1046,
+ "weight_grid": "0.00 to 1.00 inclusive, step 0.01",
+ "pool": "log-linear/geometric",
+ "temperature_fit": "bounded scalar NLL minimization, T in [0.05, 20]",
+ "ece": "15 equal-width bins for Typed Decisions; 10 for every other dataset",
+ "selection_objective": "accuracy; calibrated NLL then distance from 0.5 break ties",
+ "transfer": "the selected weight and temperature are frozen for every held-out dataset",
+ "native_shipped": "the checkpoint's native probabilities, including its shipped inference temperature; no additional temperature is applied",
+ "fixed_blend_0_5": "equal log-linear pooling of T=1 choice probabilities and shipped native probabilities; no additional temperature is applied",
+ "transfer_dev_tuned_blend": "the Transfer-dev-selected native weight followed by the reported scalar additional temperature"
+ },
+ "selected": {
+ "choice_temperature": 0.5969253208700946,
+ "native_weight": 0.52,
+ "blend_temperature": 0.8352476617411964,
+ "selection_objective": "hard-label accuracy on the Transfer-v9 development clean/knowable cohort; calibrated NLL and distance from 0.5 break ties"
+ }
+ }
+ ],
+ "external_source": {
+ "path": "results/external-zero-shot-v1.json",
+ "sha256": "8a589a7e5d0d5ae5232fbdf30a71cdce5399cb5f5aacb1fcaf974057be48b3b9",
+ "artifact_version": 5,
+ "provenance": {
+ "typed_existing": {
+ "path": "results/external-decision-evals-20261001/typed-decisions.json",
+ "sha256": "a9fbcc06a5e957630e83f18c5dc01d41ff7b3e2289b0d816f2e1459ef814341e"
+ },
+ "typed_open_summary": {
+ "local_run_path": "runs/external-open-baselines-20261001/summary.json",
+ "sha256": "4b4f06f8f02ce6cb4a610f110450469c46b917fd9aa1d571077d0c054320eab2"
+ },
+ "jevjudge_full_context_scores": {
+ "local_run_path": "runs/jevjudge-open-baselines-20261001/scores-full-context",
+ "protocol": "native inputs, no truncation, 65,536-token shared context ceiling",
+ "main_cohort_sha256_by_file": {
+ "bongard-mini.json": "f3fa601f617ed2e70b31d98fe2c589ee8787ee47624233f3fbd8b3d02630daa9",
+ "jeff-gemma4-e2b.json": "78fd3e26bf01b5630e23656049f2a33e3a77e81a0381d6106c2ad3afce4d0491",
+ "jeff-qwen35-08b.json": "ea2dc271c6d23c11278e6fdcdf56a9f45a1938404577d98ce101ec6669f58dbd",
+ "jeff-qwen35-2b.json": "8d76f061b8cd67811d48b1eb5467e9274c1648e47cb3c835974c8b32c647a404",
+ "jevany-gemma-4b-step2771.json": "88561342bec2909101bb8e3bb90a7f9a73626a35056468fae0d27f4e0389c39a",
+ "jevany-muse-glimmer-30b-step3324.json": "8022d3ff2cb33c602bf6b3aca55cdc10840f1d0735bd523c53dacb0fb966ba00",
+ "jevany-qwen35-4b-direct-release.json": "db0b29b095e73e6fe6f0795632686a05e0057f03f7b27eae46121db6854eee02",
+ "jevany-qwen35-4b-pointer-step13850.json": "c7c70a9b2adda41a597bbdd2c354320f5caf2fe4ce09d24c3fb5731615ff5cc0",
+ "jevany-qwen38-27b-step44319.json": "fcbdda6cb097f84b8bade8b7272f588dd049d32eb0ccd1f59cb6a76850954661",
+ "opendecider-small.json": "438b89596d7fe7aa593d1d024ef6576082d074e0b488eebc1c4336ee65a28032"
+ },
+ "supplemental_sha256_by_file": {
+ "kev-27b.json": "056a42c2fef59f5ecffb40e342690ac4467b273f374cbf5d59ec9dbbe3df0cae",
+ "kev-4b.json": "c6985eef36f04b962c17bc72ca3bd6d2235bcf0619904db6dda135a649291db1",
+ "laya.json": "abde9d49d8c100b8c6c09ec6b3aaea442f0f9b2f1350e34701d1b615e1651b90"
+ },
+ "supplemental_protocol": "Mixed audit protocols; these hashes are not covered by the no-truncation main-cohort claim."
+ },
+ "jevjudge_full_jevany": {
+ "local_run_path": "runs/jevjudge-full-current-20261001",
+ "summary_sha256": "fcc3025cd29ec057d869067edc5d45f5d154b8596f3281aabc4d0c83996b6609",
+ "final_audit_sha256": "ad0055609f1a3dd87210738acb45a8fce113c2ad0fc9b507af8ddce8d644d773",
+ "official_scorer_sha256": "477b39c76dd6e6200f0919fcb56e80045ffd29b66445496d5e6e70c3d1537141",
+ "bootstrap_scorer_sha256": "8f7413d70290bd1009acd398478bb26cbc98f89ebd19af298fc0930b7a763133",
+ "protocol": "native text/image/video inputs; 65,536-token ceiling; no token, option, or media truncation; 3,220/3,220 per model"
+ },
+ "jevjudge_full_open": {
+ "local_run_path": "worktrees/JevAny/jevjudge-full-open-20261001/runs/jevjudge-full-open-20261001-v2",
+ "original_summary_sha256": "f34fc4265b6cf5847d4f1bd46a5601111cc927f57c4e6178acbf61e496477c51",
+ "corrected_final_summary_sha256": "3c348c94f0494815ae4236ffb2524e92d6cdd3d237812d21465e64e18ddc00b3",
+ "compatibility_matrix_sha256": "128cc4d00520fa0515f3f5d5172bc4d704c08ebea63ba8b60f93f225db4ac98e",
+ "note": "Published headline values use returned-probability NLL and the final source-stratified group bootstrap; the original summary retained a raw-logit NLL diagnostic."
+ },
+ "jevjudge_kev_text": {
+ "local_run_path": "worktrees/JevAny/kev-jevjudge-20261001/runs/jevjudge-kev-full-context-20261001",
+ "summary_sha256": "4e1f448503983a13a14df9fc877cd876906c09915ef8173bc676778d52e06598",
+ "protocol": "724/724 text records; 65,536-token ceiling; 4,096-token chunked-KV prefill; no input truncation"
+ },
+ "jevjudge_jev_113_openrouter": {
+ "source_commit": "2ff9d3019bd7f910486543ce76513321537a3c6f",
+ "protocol": "OpenRouter full-context rerun; 724/724 JevJudge text-only records",
+ "note": "Aggregate metrics were added directly to the public evaluation table; item-level outputs are not tracked in this artifact."
+ }
+ }
+ },
+ "repository_catalog": {
+ "path": "results/model-family-v2.json",
+ "sha256": "a49f03bd9371f1846d78a09682f000cba9f64004d8d5a222b3868f11ccb48d8c",
+ "release": "model-family-v2"
+ },
+ "datasets": {
+ "transfer_calibration": {
+ "label": "Transfer-v9 development",
+ "role": "selection_only",
+ "records": 1264,
+ "questions": 1264,
+ "headline_n": 1046,
+ "runs": [
+ {
+ "run": "base4",
+ "model": "Frozen-Qwen3.5-4B",
+ "family": "Qwen3.5-4B",
+ "weights": "frozen_base",
+ "native_readout": null,
+ "zero_shot": {
+ "choice_t1": {
+ "n": 1046,
+ "correct": 725,
+ "accuracy": 0.6931166347992351,
+ "cross_entropy_from_gold": 0.8799398258908452,
+ "gold_entropy": 0.0,
+ "kl_from_gold": 0.8799398258908452,
+ "brier": 0.43925607559848767,
+ "ece": 0.12305476565507475,
+ "mean_confidence": 0.8154458555896062,
+ "cohort": "clean and knowable",
+ "excluded_non_headline_rows": 218,
+ "ece_bins": 10
+ }
+ },
+ "transfer_dev_tuned": {
+ "transfer_dev_calibrated_choice": {
+ "n": 1046,
+ "correct": 725,
+ "accuracy": 0.6931166347992351,
+ "cross_entropy_from_gold": 0.7684499162138351,
+ "gold_entropy": 0.0,
+ "kl_from_gold": 0.7684499162138351,
+ "brier": 0.40380428797912027,
+ "ece": 0.0442424982205075,
+ "mean_confidence": 0.7062354250490573,
+ "cohort": "clean and knowable",
+ "excluded_non_headline_rows": 218,
+ "ece_bins": 10
+ }
+ }
+ },
+ {
+ "run": "pointer4",
+ "model": "JevAny-Qwen3.5-4B-Pointer",
+ "family": "Qwen3.5-4B",
+ "weights": "jevany_sft",
+ "native_readout": "pointer",
+ "zero_shot": {
+ "choice_t1": {
+ "n": 1046,
+ "correct": 784,
+ "accuracy": 0.7495219885277247,
+ "cross_entropy_from_gold": 0.6801893631325125,
+ "gold_entropy": 0.0,
+ "kl_from_gold": 0.6801893631325125,
+ "brier": 0.3338033580246269,
+ "ece": 0.054371329145242946,
+ "mean_confidence": 0.6968489272720685,
+ "cohort": "clean and knowable",
+ "excluded_non_headline_rows": 218,
+ "ece_bins": 10
+ },
+ "native_shipped": {
+ "n": 1046,
+ "correct": 823,
+ "accuracy": 0.7868068833652008,
+ "cross_entropy_from_gold": 0.5868780830556992,
+ "gold_entropy": 0.0,
+ "kl_from_gold": 0.5868780830556992,
+ "brier": 0.2965183466659115,
+ "ece": 0.03479245011625781,
+ "mean_confidence": 0.805131298589329,
+ "cohort": "clean and knowable",
+ "excluded_non_headline_rows": 218,
+ "ece_bins": 10
+ },
+ "fixed_blend_0_5": {
+ "n": 1046,
+ "correct": 822,
+ "accuracy": 0.7858508604206501,
+ "cross_entropy_from_gold": 0.5717414381657305,
+ "gold_entropy": 0.0,
+ "kl_from_gold": 0.5717414381657305,
+ "brier": 0.2919156848114442,
+ "ece": 0.03811670878840213,
+ "mean_confidence": 0.7688884953818541,
+ "cohort": "clean and knowable",
+ "excluded_non_headline_rows": 218,
+ "ece_bins": 10
+ }
+ },
+ "transfer_dev_tuned": {
+ "transfer_dev_calibrated_choice": {
+ "n": 1046,
+ "correct": 784,
+ "accuracy": 0.7495219885277247,
+ "cross_entropy_from_gold": 0.675438096110365,
+ "gold_entropy": 0.0,
+ "kl_from_gold": 0.675438096110365,
+ "brier": 0.3287201057208544,
+ "ece": 0.030612288328262394,
+ "mean_confidence": 0.7238557515998861,
+ "cohort": "clean and knowable",
+ "excluded_non_headline_rows": 218,
+ "ece_bins": 10
+ },
+ "transfer_dev_tuned_blend": {
+ "n": 1046,
+ "correct": 828,
+ "accuracy": 0.7915869980879541,
+ "cross_entropy_from_gold": 0.569720964551106,
+ "gold_entropy": 0.0,
+ "kl_from_gold": 0.569720964551106,
+ "brier": 0.2929198506354004,
+ "ece": 0.05359065555472778,
+ "mean_confidence": 0.7774203199582407,
+ "cohort": "clean and knowable",
+ "excluded_non_headline_rows": 218,
+ "ece_bins": 10
+ }
+ }
+ },
+ {
+ "run": "direct4",
+ "model": "JevAny-Qwen3.5-4B-Direct-Token",
+ "family": "Qwen3.5-4B",
+ "weights": "jevany_sft",
+ "native_readout": "direct-token",
+ "zero_shot": {
+ "choice_t1": {
+ "n": 1046,
+ "correct": 796,
+ "accuracy": 0.7609942638623327,
+ "cross_entropy_from_gold": 0.6010867296757888,
+ "gold_entropy": 0.0,
+ "kl_from_gold": 0.6010867296757888,
+ "brier": 0.3064220873561066,
+ "ece": 0.04138187687571737,
+ "mean_confidence": 0.796631394719685,
+ "cohort": "clean and knowable",
+ "excluded_non_headline_rows": 218,
+ "ece_bins": 10
+ },
+ "native_shipped": {
+ "n": 1046,
+ "correct": 818,
+ "accuracy": 0.7820267686424475,
+ "cross_entropy_from_gold": 0.5642005514977974,
+ "gold_entropy": 0.0,
+ "kl_from_gold": 0.5642005514977974,
+ "brier": 0.2908559938309828,
+ "ece": 0.029163506342423436,
+ "mean_confidence": 0.8005854296720307,
+ "cohort": "clean and knowable",
+ "excluded_non_headline_rows": 218,
+ "ece_bins": 10
+ },
+ "fixed_blend_0_5": {
+ "n": 1046,
+ "correct": 821,
+ "accuracy": 0.7848948374760994,
+ "cross_entropy_from_gold": 0.5635679825981634,
+ "gold_entropy": 0.0,
+ "kl_from_gold": 0.5635679825981634,
+ "brier": 0.28919182401119004,
+ "ece": 0.03032435006136549,
+ "mean_confidence": 0.79621407300774,
+ "cohort": "clean and knowable",
+ "excluded_non_headline_rows": 218,
+ "ece_bins": 10
+ }
+ },
+ "transfer_dev_tuned": {
+ "transfer_dev_calibrated_choice": {
+ "n": 1046,
+ "correct": 796,
+ "accuracy": 0.7609942638623327,
+ "cross_entropy_from_gold": 0.5914027552210456,
+ "gold_entropy": 0.0,
+ "kl_from_gold": 0.5914027552210456,
+ "brier": 0.3043378646694837,
+ "ece": 0.03285794273637177,
+ "mean_confidence": 0.765631487860452,
+ "cohort": "clean and knowable",
+ "excluded_non_headline_rows": 218,
+ "ece_bins": 10
+ },
+ "transfer_dev_tuned_blend": {
+ "n": 1046,
+ "correct": 826,
+ "accuracy": 0.7896749521988528,
+ "cross_entropy_from_gold": 0.5583414368363817,
+ "gold_entropy": 0.0,
+ "kl_from_gold": 0.5583414368363817,
+ "brier": 0.2874056312419155,
+ "ece": 0.027602128315946592,
+ "mean_confidence": 0.7807455671293496,
+ "cohort": "clean and knowable",
+ "excluded_non_headline_rows": 218,
+ "ece_bins": 10
+ }
+ }
+ },
+ {
+ "run": "base27",
+ "model": "Frozen-Qwen3.8-27B",
+ "family": "Qwen3.8-27B",
+ "weights": "frozen_base",
+ "native_readout": null,
+ "zero_shot": {
+ "choice_t1": {
+ "n": 1046,
+ "correct": 820,
+ "accuracy": 0.7839388145315488,
+ "cross_entropy_from_gold": 0.6672937347936473,
+ "gold_entropy": 0.0,
+ "kl_from_gold": 0.6672937347936473,
+ "brier": 0.31577326969120223,
+ "ece": 0.08097099874334637,
+ "mean_confidence": 0.8596910291171252,
+ "cohort": "clean and knowable",
+ "excluded_non_headline_rows": 218,
+ "ece_bins": 10
+ }
+ },
+ "transfer_dev_tuned": {
+ "transfer_dev_calibrated_choice": {
+ "n": 1046,
+ "correct": 820,
+ "accuracy": 0.7839388145315488,
+ "cross_entropy_from_gold": 0.5895245884799086,
+ "gold_entropy": 0.0,
+ "kl_from_gold": 0.5895245884799086,
+ "brier": 0.3023594256806041,
+ "ece": 0.029495602615825022,
+ "mean_confidence": 0.7806446298690051,
+ "cohort": "clean and knowable",
+ "excluded_non_headline_rows": 218,
+ "ece_bins": 10
+ }
+ }
+ },
+ {
+ "run": "pointer27",
+ "model": "JevAny-Qwen3.8-27B-Pointer",
+ "family": "Qwen3.8-27B",
+ "weights": "jevany_sft",
+ "native_readout": "pointer",
+ "zero_shot": {
+ "choice_t1": {
+ "n": 1046,
+ "correct": 885,
+ "accuracy": 0.8460803059273423,
+ "cross_entropy_from_gold": 0.5085522464835983,
+ "gold_entropy": 0.0,
+ "kl_from_gold": 0.5085522464835983,
+ "brier": 0.2411187763545078,
+ "ece": 0.10827608077576636,
+ "mean_confidence": 0.7390643592114102,
+ "cohort": "clean and knowable",
+ "excluded_non_headline_rows": 218,
+ "ece_bins": 10
+ },
+ "native_shipped": {
+ "n": 1046,
+ "correct": 900,
+ "accuracy": 0.8604206500956023,
+ "cross_entropy_from_gold": 0.3878296713265552,
+ "gold_entropy": 0.0,
+ "kl_from_gold": 0.3878296713265552,
+ "brier": 0.19524463105141282,
+ "ece": 0.026304156956950205,
+ "mean_confidence": 0.8776426213905197,
+ "cohort": "clean and knowable",
+ "excluded_non_headline_rows": 218,
+ "ece_bins": 10
+ },
+ "fixed_blend_0_5": {
+ "n": 1046,
+ "correct": 907,
+ "accuracy": 0.8671128107074569,
+ "cross_entropy_from_gold": 0.3773854700281526,
+ "gold_entropy": 0.0,
+ "kl_from_gold": 0.3773854700281526,
+ "brier": 0.19080039228840473,
+ "ece": 0.030952348286441715,
+ "mean_confidence": 0.8386108687417092,
+ "cohort": "clean and knowable",
+ "excluded_non_headline_rows": 218,
+ "ece_bins": 10
+ }
+ },
+ "transfer_dev_tuned": {
+ "transfer_dev_calibrated_choice": {
+ "n": 1046,
+ "correct": 885,
+ "accuracy": 0.8460803059273423,
+ "cross_entropy_from_gold": 0.4457110992469971,
+ "gold_entropy": 0.0,
+ "kl_from_gold": 0.4457110992469971,
+ "brier": 0.21999786383788478,
+ "ece": 0.023323136487035337,
+ "mean_confidence": 0.8447172036358691,
+ "cohort": "clean and knowable",
+ "excluded_non_headline_rows": 218,
+ "ece_bins": 10
+ },
+ "transfer_dev_tuned_blend": {
+ "n": 1046,
+ "correct": 909,
+ "accuracy": 0.8690248565965584,
+ "cross_entropy_from_gold": 0.3698933820845643,
+ "gold_entropy": 0.0,
+ "kl_from_gold": 0.3698933820845643,
+ "brier": 0.18873769003914254,
+ "ece": 0.020429096141296538,
+ "mean_confidence": 0.8663359693950322,
+ "cohort": "clean and knowable",
+ "excluded_non_headline_rows": 218,
+ "ece_bins": 10
+ }
+ }
+ }
+ ]
+ },
+ "transfer_test": {
+ "label": "Transfer-v9 test",
+ "role": "held_out_evaluation",
+ "records": 1264,
+ "questions": 1264,
+ "headline_n": 1046,
+ "runs": [
+ {
+ "run": "base4",
+ "model": "Frozen-Qwen3.5-4B",
+ "family": "Qwen3.5-4B",
+ "weights": "frozen_base",
+ "native_readout": null,
+ "zero_shot": {
+ "choice_t1": {
+ "n": 1046,
+ "correct": 716,
+ "accuracy": 0.6845124282982792,
+ "cross_entropy_from_gold": 0.8349623273118204,
+ "gold_entropy": 0.0,
+ "kl_from_gold": 0.8349623273118204,
+ "brier": 0.42656058074391484,
+ "ece": 0.1303798195616654,
+ "mean_confidence": 0.8139445762365763,
+ "cohort": "clean and knowable",
+ "excluded_non_headline_rows": 218,
+ "ece_bins": 10
+ }
+ },
+ "transfer_dev_tuned": {
+ "transfer_dev_calibrated_choice": {
+ "n": 1046,
+ "correct": 716,
+ "accuracy": 0.6845124282982792,
+ "cross_entropy_from_gold": 0.7410651656609946,
+ "gold_entropy": 0.0,
+ "kl_from_gold": 0.7410651656609946,
+ "brier": 0.3904408201413376,
+ "ece": 0.04134673815393105,
+ "mean_confidence": 0.7051530768940606,
+ "cohort": "clean and knowable",
+ "excluded_non_headline_rows": 218,
+ "ece_bins": 10
+ }
+ }
+ },
+ {
+ "run": "pointer4",
+ "model": "JevAny-Qwen3.5-4B-Pointer",
+ "family": "Qwen3.5-4B",
+ "weights": "jevany_sft",
+ "native_readout": "pointer",
+ "zero_shot": {
+ "choice_t1": {
+ "n": 1046,
+ "correct": 770,
+ "accuracy": 0.7361376673040153,
+ "cross_entropy_from_gold": 0.6902431349522291,
+ "gold_entropy": 0.0,
+ "kl_from_gold": 0.6902431349522291,
+ "brier": 0.348668135528129,
+ "ece": 0.04630103938423926,
+ "mean_confidence": 0.701833119833475,
+ "cohort": "clean and knowable",
+ "excluded_non_headline_rows": 218,
+ "ece_bins": 10
+ },
+ "native_shipped": {
+ "n": 1046,
+ "correct": 815,
+ "accuracy": 0.7791586998087954,
+ "cross_entropy_from_gold": 0.5717563825431627,
+ "gold_entropy": 0.0,
+ "kl_from_gold": 0.5717563825431627,
+ "brier": 0.29765821801656817,
+ "ece": 0.04239090732927413,
+ "mean_confidence": 0.8129759944331685,
+ "cohort": "clean and knowable",
+ "excluded_non_headline_rows": 218,
+ "ece_bins": 10
+ },
+ "fixed_blend_0_5": {
+ "n": 1046,
+ "correct": 810,
+ "accuracy": 0.7743785850860421,
+ "cross_entropy_from_gold": 0.5666923610469762,
+ "gold_entropy": 0.0,
+ "kl_from_gold": 0.5666923610469762,
+ "brier": 0.2939336464727867,
+ "ece": 0.04874813932892946,
+ "mean_confidence": 0.7776535459474749,
+ "cohort": "clean and knowable",
+ "excluded_non_headline_rows": 218,
+ "ece_bins": 10
+ }
+ },
+ "transfer_dev_tuned": {
+ "transfer_dev_calibrated_choice": {
+ "n": 1046,
+ "correct": 770,
+ "accuracy": 0.7361376673040153,
+ "cross_entropy_from_gold": 0.6888466560394513,
+ "gold_entropy": 0.0,
+ "kl_from_gold": 0.6888466560394513,
+ "brier": 0.34649580012970715,
+ "ece": 0.03269551339782706,
+ "mean_confidence": 0.7287518287657895,
+ "cohort": "clean and knowable",
+ "excluded_non_headline_rows": 218,
+ "ece_bins": 10
+ },
+ "transfer_dev_tuned_blend": {
+ "n": 1046,
+ "correct": 816,
+ "accuracy": 0.780114722753346,
+ "cross_entropy_from_gold": 0.5578570106435864,
+ "gold_entropy": 0.0,
+ "kl_from_gold": 0.5578570106435864,
+ "brier": 0.2914736076482713,
+ "ece": 0.04496393628178446,
+ "mean_confidence": 0.7852072596268804,
+ "cohort": "clean and knowable",
+ "excluded_non_headline_rows": 218,
+ "ece_bins": 10
+ }
+ }
+ },
+ {
+ "run": "direct4",
+ "model": "JevAny-Qwen3.5-4B-Direct-Token",
+ "family": "Qwen3.5-4B",
+ "weights": "jevany_sft",
+ "native_readout": "direct-token",
+ "zero_shot": {
+ "choice_t1": {
+ "n": 1046,
+ "correct": 822,
+ "accuracy": 0.7858508604206501,
+ "cross_entropy_from_gold": 0.5697635579703885,
+ "gold_entropy": 0.0,
+ "kl_from_gold": 0.5697635579703885,
+ "brier": 0.29077517389394053,
+ "ece": 0.0470678026592797,
+ "mean_confidence": 0.8092660785307848,
+ "cohort": "clean and knowable",
+ "excluded_non_headline_rows": 218,
+ "ece_bins": 10
+ },
+ "native_shipped": {
+ "n": 1046,
+ "correct": 825,
+ "accuracy": 0.7887189292543021,
+ "cross_entropy_from_gold": 0.5304680016857037,
+ "gold_entropy": 0.0,
+ "kl_from_gold": 0.5304680016857037,
+ "brier": 0.2727158680694733,
+ "ece": 0.040601108067036665,
+ "mean_confidence": 0.8199335179349373,
+ "cohort": "clean and knowable",
+ "excluded_non_headline_rows": 218,
+ "ece_bins": 10
+ },
+ "fixed_blend_0_5": {
+ "n": 1046,
+ "correct": 831,
+ "accuracy": 0.7944550669216062,
+ "cross_entropy_from_gold": 0.5309958800935447,
+ "gold_entropy": 0.0,
+ "kl_from_gold": 0.5309958800935447,
+ "brier": 0.2725904993128285,
+ "ece": 0.038203239429849746,
+ "mean_confidence": 0.8132871591572012,
+ "cohort": "clean and knowable",
+ "excluded_non_headline_rows": 218,
+ "ece_bins": 10
+ }
+ },
+ "transfer_dev_tuned": {
+ "transfer_dev_calibrated_choice": {
+ "n": 1046,
+ "correct": 822,
+ "accuracy": 0.7858508604206501,
+ "cross_entropy_from_gold": 0.5628608839489921,
+ "gold_entropy": 0.0,
+ "kl_from_gold": 0.5628608839489921,
+ "brier": 0.28883932397403955,
+ "ece": 0.030719110557561026,
+ "mean_confidence": 0.7766618482775124,
+ "cohort": "clean and knowable",
+ "excluded_non_headline_rows": 218,
+ "ece_bins": 10
+ },
+ "transfer_dev_tuned_blend": {
+ "n": 1046,
+ "correct": 835,
+ "accuracy": 0.7982791586998088,
+ "cross_entropy_from_gold": 0.5262143707831856,
+ "gold_entropy": 0.0,
+ "kl_from_gold": 0.5262143707831856,
+ "brier": 0.26999002433749075,
+ "ece": 0.027780702772704318,
+ "mean_confidence": 0.7973056353797929,
+ "cohort": "clean and knowable",
+ "excluded_non_headline_rows": 218,
+ "ece_bins": 10
+ }
+ }
+ },
+ {
+ "run": "base27",
+ "model": "Frozen-Qwen3.8-27B",
+ "family": "Qwen3.8-27B",
+ "weights": "frozen_base",
+ "native_readout": null,
+ "zero_shot": {
+ "choice_t1": {
+ "n": 1046,
+ "correct": 844,
+ "accuracy": 0.8068833652007649,
+ "cross_entropy_from_gold": 0.6110678829313702,
+ "gold_entropy": 0.0,
+ "kl_from_gold": 0.6110678829313702,
+ "brier": 0.279721544152333,
+ "ece": 0.06893200869919826,
+ "mean_confidence": 0.8617564454102284,
+ "cohort": "clean and knowable",
+ "excluded_non_headline_rows": 218,
+ "ece_bins": 10
+ }
+ },
+ "transfer_dev_tuned": {
+ "transfer_dev_calibrated_choice": {
+ "n": 1046,
+ "correct": 844,
+ "accuracy": 0.8068833652007649,
+ "cross_entropy_from_gold": 0.5494987252614483,
+ "gold_entropy": 0.0,
+ "kl_from_gold": 0.5494987252614483,
+ "brier": 0.2709760690574176,
+ "ece": 0.027612650998887305,
+ "mean_confidence": 0.7874674967417389,
+ "cohort": "clean and knowable",
+ "excluded_non_headline_rows": 218,
+ "ece_bins": 10
+ }
+ }
+ },
+ {
+ "run": "pointer27",
+ "model": "JevAny-Qwen3.8-27B-Pointer",
+ "family": "Qwen3.8-27B",
+ "weights": "jevany_sft",
+ "native_readout": "pointer",
+ "zero_shot": {
+ "choice_t1": {
+ "n": 1046,
+ "correct": 900,
+ "accuracy": 0.8604206500956023,
+ "cross_entropy_from_gold": 0.47077064689010384,
+ "gold_entropy": 0.0,
+ "kl_from_gold": 0.47077064689010384,
+ "brier": 0.2114754899927417,
+ "ece": 0.11231046843501219,
+ "mean_confidence": 0.7539620937273849,
+ "cohort": "clean and knowable",
+ "excluded_non_headline_rows": 218,
+ "ece_bins": 10
+ },
+ "native_shipped": {
+ "n": 1046,
+ "correct": 920,
+ "accuracy": 0.8795411089866156,
+ "cross_entropy_from_gold": 0.3191183572882807,
+ "gold_entropy": 0.0,
+ "kl_from_gold": 0.3191183572882807,
+ "brier": 0.15983357286839844,
+ "ece": 0.02547180335097676,
+ "mean_confidence": 0.8916975436746409,
+ "cohort": "clean and knowable",
+ "excluded_non_headline_rows": 218,
+ "ece_bins": 10
+ },
+ "fixed_blend_0_5": {
+ "n": 1046,
+ "correct": 933,
+ "accuracy": 0.8919694072657743,
+ "cross_entropy_from_gold": 0.32150313724237106,
+ "gold_entropy": 0.0,
+ "kl_from_gold": 0.32150313724237106,
+ "brier": 0.1550705009584002,
+ "ece": 0.043199084741469725,
+ "mean_confidence": 0.8548688707373127,
+ "cohort": "clean and knowable",
+ "excluded_non_headline_rows": 218,
+ "ece_bins": 10
+ }
+ },
+ "transfer_dev_tuned": {
+ "transfer_dev_calibrated_choice": {
+ "n": 1046,
+ "correct": 900,
+ "accuracy": 0.8604206500956023,
+ "cross_entropy_from_gold": 0.40058733865111656,
+ "gold_entropy": 0.0,
+ "kl_from_gold": 0.40058733865111656,
+ "brier": 0.1880387608542386,
+ "ece": 0.02034085988544015,
+ "mean_confidence": 0.8604184603403277,
+ "cohort": "clean and knowable",
+ "excluded_non_headline_rows": 218,
+ "ece_bins": 10
+ },
+ "transfer_dev_tuned_blend": {
+ "n": 1046,
+ "correct": 932,
+ "accuracy": 0.8910133843212237,
+ "cross_entropy_from_gold": 0.3094830481103647,
+ "gold_entropy": 0.0,
+ "kl_from_gold": 0.3094830481103647,
+ "brier": 0.15280535574202198,
+ "ece": 0.034827378233759226,
+ "mean_confidence": 0.8807493062406941,
+ "cohort": "clean and knowable",
+ "excluded_non_headline_rows": 218,
+ "ece_bins": 10
+ }
+ }
+ }
+ ]
+ },
+ "typed_test": {
+ "label": "Typed Decisions test",
+ "role": "held_out_external_evaluation",
+ "records": 400,
+ "questions": 2000,
+ "headline_n": 2000,
+ "runs": [
+ {
+ "run": "base4",
+ "model": "Frozen-Qwen3.5-4B",
+ "family": "Qwen3.5-4B",
+ "weights": "frozen_base",
+ "native_readout": null,
+ "zero_shot": {
+ "choice_t1": {
+ "n": 2000,
+ "correct": 1055,
+ "accuracy": 0.5275,
+ "cross_entropy_from_gold": 1.485304034737827,
+ "gold_entropy": 0.7673585556243288,
+ "kl_from_gold": 0.7179454791134983,
+ "brier": 0.32209870787613476,
+ "ece": 0.21683605197556294,
+ "mean_confidence": 0.7434521192146655,
+ "cohort": "clean and knowable",
+ "excluded_non_headline_rows": 0,
+ "ece_bins": 15
+ }
+ },
+ "transfer_dev_tuned": {
+ "transfer_dev_calibrated_choice": {
+ "n": 2000,
+ "correct": 1055,
+ "accuracy": 0.5275,
+ "cross_entropy_from_gold": 1.160641706252667,
+ "gold_entropy": 0.7673585556243288,
+ "kl_from_gold": 0.3932831506283383,
+ "brier": 0.22234376212257462,
+ "ece": 0.0900516370077249,
+ "mean_confidence": 0.6134237235271172,
+ "cohort": "clean and knowable",
+ "excluded_non_headline_rows": 0,
+ "ece_bins": 15
+ }
+ }
+ },
+ {
+ "run": "pointer4",
+ "model": "JevAny-Qwen3.5-4B-Pointer",
+ "family": "Qwen3.5-4B",
+ "weights": "jevany_sft",
+ "native_readout": "pointer",
+ "zero_shot": {
+ "choice_t1": {
+ "n": 2000,
+ "correct": 1156,
+ "accuracy": 0.578,
+ "cross_entropy_from_gold": 1.1689837218988208,
+ "gold_entropy": 0.7673585556243288,
+ "kl_from_gold": 0.4016251662744921,
+ "brier": 0.21958081006971433,
+ "ece": 0.11804038821625913,
+ "mean_confidence": 0.6283756741526655,
+ "cohort": "clean and knowable",
+ "excluded_non_headline_rows": 0,
+ "ece_bins": 15
+ },
+ "native_shipped": {
+ "n": 2000,
+ "correct": 1270,
+ "accuracy": 0.635,
+ "cross_entropy_from_gold": 1.2461303849511813,
+ "gold_entropy": 0.7673585556243288,
+ "kl_from_gold": 0.4787718293268526,
+ "brier": 0.21769431550066473,
+ "ece": 0.11542953593789118,
+ "mean_confidence": 0.7131901462821837,
+ "cohort": "clean and knowable",
+ "excluded_non_headline_rows": 0,
+ "ece_bins": 15
+ },
+ "fixed_blend_0_5": {
+ "n": 2000,
+ "correct": 1214,
+ "accuracy": 0.607,
+ "cross_entropy_from_gold": 1.1474089185651393,
+ "gold_entropy": 0.7673585556243288,
+ "kl_from_gold": 0.3800503629408105,
+ "brier": 0.20385335585187295,
+ "ece": 0.08838893066306003,
+ "mean_confidence": 0.6710193141005254,
+ "cohort": "clean and knowable",
+ "excluded_non_headline_rows": 0,
+ "ece_bins": 15
+ }
+ },
+ "transfer_dev_tuned": {
+ "transfer_dev_calibrated_choice": {
+ "n": 2000,
+ "correct": 1156,
+ "accuracy": 0.578,
+ "cross_entropy_from_gold": 1.2283795069938794,
+ "gold_entropy": 0.7673585556243288,
+ "kl_from_gold": 0.4610209513695507,
+ "brier": 0.23872296531625706,
+ "ece": 0.13219447566436857,
+ "mean_confidence": 0.6614429063056575,
+ "cohort": "clean and knowable",
+ "excluded_non_headline_rows": 0,
+ "ece_bins": 15
+ },
+ "transfer_dev_tuned_blend": {
+ "n": 2000,
+ "correct": 1258,
+ "accuracy": 0.629,
+ "cross_entropy_from_gold": 1.1499057537598978,
+ "gold_entropy": 0.7673585556243288,
+ "kl_from_gold": 0.38254719813556903,
+ "brier": 0.19860375584665835,
+ "ece": 0.09685551065650508,
+ "mean_confidence": 0.6767929224204166,
+ "cohort": "clean and knowable",
+ "excluded_non_headline_rows": 0,
+ "ece_bins": 15
+ }
+ }
+ },
+ {
+ "run": "direct4",
+ "model": "JevAny-Qwen3.5-4B-Direct-Token",
+ "family": "Qwen3.5-4B",
+ "weights": "jevany_sft",
+ "native_readout": "direct-token",
+ "zero_shot": {
+ "choice_t1": {
+ "n": 2000,
+ "correct": 1296,
+ "accuracy": 0.648,
+ "cross_entropy_from_gold": 1.157150279237127,
+ "gold_entropy": 0.7673585556243288,
+ "kl_from_gold": 0.38979172361279835,
+ "brier": 0.18515075572006146,
+ "ece": 0.07029919279872957,
+ "mean_confidence": 0.6765339723262027,
+ "cohort": "clean and knowable",
+ "excluded_non_headline_rows": 0,
+ "ece_bins": 15
+ },
+ "native_shipped": {
+ "n": 2000,
+ "correct": 1344,
+ "accuracy": 0.672,
+ "cross_entropy_from_gold": 1.2022614330547874,
+ "gold_entropy": 0.7673585556243288,
+ "kl_from_gold": 0.43490287743045863,
+ "brier": 0.17877494146152656,
+ "ece": 0.08278732481632818,
+ "mean_confidence": 0.7379338187164135,
+ "cohort": "clean and knowable",
+ "excluded_non_headline_rows": 0,
+ "ece_bins": 15
+ },
+ "fixed_blend_0_5": {
+ "n": 2000,
+ "correct": 1335,
+ "accuracy": 0.6675,
+ "cross_entropy_from_gold": 1.1350838501828007,
+ "gold_entropy": 0.7673585556243288,
+ "kl_from_gold": 0.3677252945584719,
+ "brier": 0.1702376402804755,
+ "ece": 0.06198979447604478,
+ "mean_confidence": 0.7066262108749651,
+ "cohort": "clean and knowable",
+ "excluded_non_headline_rows": 0,
+ "ece_bins": 15
+ }
+ },
+ "transfer_dev_tuned": {
+ "transfer_dev_calibrated_choice": {
+ "n": 2000,
+ "correct": 1296,
+ "accuracy": 0.648,
+ "cross_entropy_from_gold": 1.083382488931363,
+ "gold_entropy": 0.7673585556243288,
+ "kl_from_gold": 0.31602393330703427,
+ "brier": 0.16593211527952617,
+ "ece": 0.07601165366691312,
+ "mean_confidence": 0.6328419411033379,
+ "cohort": "clean and knowable",
+ "excluded_non_headline_rows": 0,
+ "ece_bins": 15
+ },
+ "transfer_dev_tuned_blend": {
+ "n": 2000,
+ "correct": 1353,
+ "accuracy": 0.6765,
+ "cross_entropy_from_gold": 1.0910947225952166,
+ "gold_entropy": 0.7673585556243288,
+ "kl_from_gold": 0.3237361669708878,
+ "brier": 0.15773140989073764,
+ "ece": 0.06194953110709811,
+ "mean_confidence": 0.6893724371900853,
+ "cohort": "clean and knowable",
+ "excluded_non_headline_rows": 0,
+ "ece_bins": 15
+ }
+ }
+ },
+ {
+ "run": "base27",
+ "model": "Frozen-Qwen3.8-27B",
+ "family": "Qwen3.8-27B",
+ "weights": "frozen_base",
+ "native_readout": null,
+ "zero_shot": {
+ "choice_t1": {
+ "n": 2000,
+ "correct": 1365,
+ "accuracy": 0.6825,
+ "cross_entropy_from_gold": 1.4486833046492098,
+ "gold_entropy": 0.7673585556243288,
+ "kl_from_gold": 0.681324749024881,
+ "brier": 0.23332950942149694,
+ "ece": 0.12575488372079707,
+ "mean_confidence": 0.8052338907255023,
+ "cohort": "clean and knowable",
+ "excluded_non_headline_rows": 0,
+ "ece_bins": 15
+ }
+ },
+ "transfer_dev_tuned": {
+ "transfer_dev_calibrated_choice": {
+ "n": 2000,
+ "correct": 1365,
+ "accuracy": 0.6825,
+ "cross_entropy_from_gold": 1.0976128315315092,
+ "gold_entropy": 0.7673585556243288,
+ "kl_from_gold": 0.3302542759071805,
+ "brier": 0.16553442637117138,
+ "ece": 0.07868470788264477,
+ "mean_confidence": 0.7000954995537738,
+ "cohort": "clean and knowable",
+ "excluded_non_headline_rows": 0,
+ "ece_bins": 15
+ }
+ }
+ },
+ {
+ "run": "pointer27",
+ "model": "JevAny-Qwen3.8-27B-Pointer",
+ "family": "Qwen3.8-27B",
+ "weights": "jevany_sft",
+ "native_readout": "pointer",
+ "zero_shot": {
+ "choice_t1": {
+ "n": 2000,
+ "correct": 1452,
+ "accuracy": 0.726,
+ "cross_entropy_from_gold": 0.9201029180583041,
+ "gold_entropy": 0.7673585556243288,
+ "kl_from_gold": 0.15274436243397538,
+ "brier": 0.0785204257840767,
+ "ece": 0.10839391090387629,
+ "mean_confidence": 0.6176060890961234,
+ "cohort": "clean and knowable",
+ "excluded_non_headline_rows": 0,
+ "ece_bins": 15
+ },
+ "native_shipped": {
+ "n": 2000,
+ "correct": 1456,
+ "accuracy": 0.728,
+ "cross_entropy_from_gold": 1.0602046256660616,
+ "gold_entropy": 0.7673585556243288,
+ "kl_from_gold": 0.2928460700417328,
+ "brier": 0.13141516555965035,
+ "ece": 0.05304485014212312,
+ "mean_confidence": 0.7527105045858369,
+ "cohort": "clean and knowable",
+ "excluded_non_headline_rows": 0,
+ "ece_bins": 15
+ },
+ "fixed_blend_0_5": {
+ "n": 2000,
+ "correct": 1464,
+ "accuracy": 0.732,
+ "cross_entropy_from_gold": 0.9466132871046253,
+ "gold_entropy": 0.7673585556243288,
+ "kl_from_gold": 0.17925473148029658,
+ "brier": 0.09467929687142354,
+ "ece": 0.034120016906592165,
+ "mean_confidence": 0.6998051537601903,
+ "cohort": "clean and knowable",
+ "excluded_non_headline_rows": 0,
+ "ece_bins": 15
+ }
+ },
+ "transfer_dev_tuned": {
+ "transfer_dev_calibrated_choice": {
+ "n": 2000,
+ "correct": 1452,
+ "accuracy": 0.726,
+ "cross_entropy_from_gold": 1.0152256258087546,
+ "gold_entropy": 0.7673585556243288,
+ "kl_from_gold": 0.2478670701844259,
+ "brier": 0.11614111181065959,
+ "ece": 0.020460563111670046,
+ "mean_confidence": 0.7331494836433007,
+ "cohort": "clean and knowable",
+ "excluded_non_headline_rows": 0,
+ "ece_bins": 15
+ },
+ "transfer_dev_tuned_blend": {
+ "n": 2000,
+ "correct": 1466,
+ "accuracy": 0.733,
+ "cross_entropy_from_gold": 1.004976949424594,
+ "gold_entropy": 0.7673585556243288,
+ "kl_from_gold": 0.23761839380026517,
+ "brier": 0.11476036996840079,
+ "ece": 0.02738445525330449,
+ "mean_confidence": 0.740111940631487,
+ "cohort": "clean and knowable",
+ "excluded_non_headline_rows": 0,
+ "ece_bins": 15
+ }
+ }
+ }
+ ],
+ "external_baselines": [
+ {
+ "model": "meraGPT Decider 1",
+ "kind": "published_only",
+ "status": "dataset-card result; not locally rerun",
+ "accuracy": 0.768
+ },
+ {
+ "model": "Jev 1.13 (OpenRouter)",
+ "canonical_model": "TypeSafe Jev 1.13.0",
+ "kind": "external_api",
+ "result_source": "published_only",
+ "status": "dataset-card result; not locally rerun",
+ "accuracy": 0.727
+ }
+ ]
+ },
+ "jevbench_public": {
+ "label": "JevBench public development",
+ "role": "public_diagnostic_not_used_for_selection",
+ "records": 231,
+ "questions": 231,
+ "headline_n": 231,
+ "runs": [
+ {
+ "run": "base4",
+ "model": "Frozen-Qwen3.5-4B",
+ "family": "Qwen3.5-4B",
+ "weights": "frozen_base",
+ "native_readout": null,
+ "zero_shot": {
+ "choice_t1": {
+ "n": 231,
+ "correct": 184,
+ "accuracy": 0.7965367965367965,
+ "cross_entropy_from_gold": 0.46503985941725867,
+ "gold_entropy": 0.0,
+ "kl_from_gold": 0.46503985941725867,
+ "brier": 0.26023857002415907,
+ "ece": 0.042706500279636225,
+ "mean_confidence": 0.8277960488256922,
+ "cohort": "clean and knowable",
+ "excluded_non_headline_rows": 0,
+ "ece_bins": 10
+ }
+ },
+ "transfer_dev_tuned": {
+ "transfer_dev_calibrated_choice": {
+ "n": 231,
+ "correct": 184,
+ "accuracy": 0.7965367965367965,
+ "cross_entropy_from_gold": 0.5034935201937853,
+ "gold_entropy": 0.0,
+ "kl_from_gold": 0.5034935201937853,
+ "brier": 0.2731813432864743,
+ "ece": 0.08846006305082621,
+ "mean_confidence": 0.7245256388535549,
+ "cohort": "clean and knowable",
+ "excluded_non_headline_rows": 0,
+ "ece_bins": 10
+ }
+ }
+ },
+ {
+ "run": "pointer4",
+ "model": "JevAny-Qwen3.5-4B-Pointer",
+ "family": "Qwen3.5-4B",
+ "weights": "jevany_sft",
+ "native_readout": "pointer",
+ "zero_shot": {
+ "choice_t1": {
+ "n": 231,
+ "correct": 188,
+ "accuracy": 0.8138528138528138,
+ "cross_entropy_from_gold": 0.47901846432733636,
+ "gold_entropy": 0.0,
+ "kl_from_gold": 0.47901846432733636,
+ "brier": 0.2561308970527479,
+ "ece": 0.056124141655517386,
+ "mean_confidence": 0.7692551472534949,
+ "cohort": "clean and knowable",
+ "excluded_non_headline_rows": 0,
+ "ece_bins": 10
+ },
+ "native_shipped": {
+ "n": 231,
+ "correct": 185,
+ "accuracy": 0.8008658008658008,
+ "cross_entropy_from_gold": 0.4554565575008672,
+ "gold_entropy": 0.0,
+ "kl_from_gold": 0.4554565575008672,
+ "brier": 0.2585838755503942,
+ "ece": 0.03718147002370739,
+ "mean_confidence": 0.8308520442419092,
+ "cohort": "clean and knowable",
+ "excluded_non_headline_rows": 0,
+ "ece_bins": 10
+ },
+ "fixed_blend_0_5": {
+ "n": 231,
+ "correct": 188,
+ "accuracy": 0.8138528138528138,
+ "cross_entropy_from_gold": 0.4142501461307885,
+ "gold_entropy": 0.0,
+ "kl_from_gold": 0.4142501461307885,
+ "brier": 0.23413224994049325,
+ "ece": 0.02272572239740828,
+ "mean_confidence": 0.807826264488957,
+ "cohort": "clean and knowable",
+ "excluded_non_headline_rows": 0,
+ "ece_bins": 10
+ }
+ },
+ "transfer_dev_tuned": {
+ "transfer_dev_calibrated_choice": {
+ "n": 231,
+ "correct": 188,
+ "accuracy": 0.8138528138528138,
+ "cross_entropy_from_gold": 0.46979954158584164,
+ "gold_entropy": 0.0,
+ "kl_from_gold": 0.46979954158584164,
+ "brier": 0.2515032873679713,
+ "ece": 0.03081010600115095,
+ "mean_confidence": 0.7956829590845715,
+ "cohort": "clean and knowable",
+ "excluded_non_headline_rows": 0,
+ "ece_bins": 10
+ },
+ "transfer_dev_tuned_blend": {
+ "n": 231,
+ "correct": 187,
+ "accuracy": 0.8095238095238095,
+ "cross_entropy_from_gold": 0.4300189742220573,
+ "gold_entropy": 0.0,
+ "kl_from_gold": 0.4300189742220573,
+ "brier": 0.24497083465715366,
+ "ece": 0.03200768487813679,
+ "mean_confidence": 0.8105343163775357,
+ "cohort": "clean and knowable",
+ "excluded_non_headline_rows": 0,
+ "ece_bins": 10
+ }
+ }
+ },
+ {
+ "run": "direct4",
+ "model": "JevAny-Qwen3.5-4B-Direct-Token",
+ "family": "Qwen3.5-4B",
+ "weights": "jevany_sft",
+ "native_readout": "direct-token",
+ "zero_shot": {
+ "choice_t1": {
+ "n": 231,
+ "correct": 188,
+ "accuracy": 0.8138528138528138,
+ "cross_entropy_from_gold": 0.4119059642596153,
+ "gold_entropy": 0.0,
+ "kl_from_gold": 0.4119059642596153,
+ "brier": 0.23630630313186377,
+ "ece": 0.06493498624236703,
+ "mean_confidence": 0.8264788235127298,
+ "cohort": "clean and knowable",
+ "excluded_non_headline_rows": 0,
+ "ece_bins": 10
+ },
+ "native_shipped": {
+ "n": 231,
+ "correct": 187,
+ "accuracy": 0.8095238095238095,
+ "cross_entropy_from_gold": 0.43291971264790263,
+ "gold_entropy": 0.0,
+ "kl_from_gold": 0.43291971264790263,
+ "brier": 0.25567839211017723,
+ "ece": 0.05126707211644283,
+ "mean_confidence": 0.8459940130356174,
+ "cohort": "clean and knowable",
+ "excluded_non_headline_rows": 0,
+ "ece_bins": 10
+ },
+ "fixed_blend_0_5": {
+ "n": 231,
+ "correct": 187,
+ "accuracy": 0.8095238095238095,
+ "cross_entropy_from_gold": 0.4013583529376732,
+ "gold_entropy": 0.0,
+ "kl_from_gold": 0.4013583529376732,
+ "brier": 0.23610741709317634,
+ "ece": 0.04227451187274651,
+ "mean_confidence": 0.8360951760741018,
+ "cohort": "clean and knowable",
+ "excluded_non_headline_rows": 0,
+ "ece_bins": 10
+ }
+ },
+ "transfer_dev_tuned": {
+ "transfer_dev_calibrated_choice": {
+ "n": 231,
+ "correct": 188,
+ "accuracy": 0.8138528138528138,
+ "cross_entropy_from_gold": 0.419037743510702,
+ "gold_entropy": 0.0,
+ "kl_from_gold": 0.419037743510702,
+ "brier": 0.2361960585013741,
+ "ece": 0.04983130926953957,
+ "mean_confidence": 0.7972245036382399,
+ "cohort": "clean and knowable",
+ "excluded_non_headline_rows": 0,
+ "ece_bins": 10
+ },
+ "transfer_dev_tuned_blend": {
+ "n": 231,
+ "correct": 187,
+ "accuracy": 0.8095238095238095,
+ "cross_entropy_from_gold": 0.40445010913159385,
+ "gold_entropy": 0.0,
+ "kl_from_gold": 0.40445010913159385,
+ "brier": 0.23725278784385523,
+ "ece": 0.04395161391787564,
+ "mean_confidence": 0.8242603879406017,
+ "cohort": "clean and knowable",
+ "excluded_non_headline_rows": 0,
+ "ece_bins": 10
+ }
+ }
+ },
+ {
+ "run": "base27",
+ "model": "Frozen-Qwen3.8-27B",
+ "family": "Qwen3.8-27B",
+ "weights": "frozen_base",
+ "native_readout": null,
+ "zero_shot": {
+ "choice_t1": {
+ "n": 231,
+ "correct": 203,
+ "accuracy": 0.8787878787878788,
+ "cross_entropy_from_gold": 0.2864628223866121,
+ "gold_entropy": 0.0,
+ "kl_from_gold": 0.2864628223866121,
+ "brier": 0.1604073338283613,
+ "ece": 0.022674652746122903,
+ "mean_confidence": 0.8936596611320321,
+ "cohort": "clean and knowable",
+ "excluded_non_headline_rows": 0,
+ "ece_bins": 10
+ }
+ },
+ "transfer_dev_tuned": {
+ "transfer_dev_calibrated_choice": {
+ "n": 231,
+ "correct": 203,
+ "accuracy": 0.8787878787878788,
+ "cross_entropy_from_gold": 0.32118852252593943,
+ "gold_entropy": 0.0,
+ "kl_from_gold": 0.32118852252593943,
+ "brier": 0.16739722641469756,
+ "ece": 0.09116510874430303,
+ "mean_confidence": 0.821229989485614,
+ "cohort": "clean and knowable",
+ "excluded_non_headline_rows": 0,
+ "ece_bins": 10
+ }
+ }
+ },
+ {
+ "run": "pointer27",
+ "model": "JevAny-Qwen3.8-27B-Pointer",
+ "family": "Qwen3.8-27B",
+ "weights": "jevany_sft",
+ "native_readout": "pointer",
+ "zero_shot": {
+ "choice_t1": {
+ "n": 231,
+ "correct": 207,
+ "accuracy": 0.8961038961038961,
+ "cross_entropy_from_gold": 0.34460785579024505,
+ "gold_entropy": 0.0,
+ "kl_from_gold": 0.34460785579024505,
+ "brier": 0.16729832168502917,
+ "ece": 0.12307461523799648,
+ "mean_confidence": 0.7818628405160418,
+ "cohort": "clean and knowable",
+ "excluded_non_headline_rows": 0,
+ "ece_bins": 10
+ },
+ "native_shipped": {
+ "n": 231,
+ "correct": 208,
+ "accuracy": 0.9004329004329005,
+ "cross_entropy_from_gold": 0.2650469579600223,
+ "gold_entropy": 0.0,
+ "kl_from_gold": 0.2650469579600223,
+ "brier": 0.14556199524820898,
+ "ece": 0.02122154456035358,
+ "mean_confidence": 0.884113426683873,
+ "cohort": "clean and knowable",
+ "excluded_non_headline_rows": 0,
+ "ece_bins": 10
+ },
+ "fixed_blend_0_5": {
+ "n": 231,
+ "correct": 208,
+ "accuracy": 0.9004329004329005,
+ "cross_entropy_from_gold": 0.26310190901965735,
+ "gold_entropy": 0.0,
+ "kl_from_gold": 0.26310190901965735,
+ "brier": 0.1406077982333171,
+ "ece": 0.05751149881179743,
+ "mean_confidence": 0.8560666856708948,
+ "cohort": "clean and knowable",
+ "excluded_non_headline_rows": 0,
+ "ece_bins": 10
+ }
+ },
+ "transfer_dev_tuned": {
+ "transfer_dev_calibrated_choice": {
+ "n": 231,
+ "correct": 207,
+ "accuracy": 0.8961038961038961,
+ "cross_entropy_from_gold": 0.26532298607460764,
+ "gold_entropy": 0.0,
+ "kl_from_gold": 0.26532298607460764,
+ "brier": 0.14331108028016937,
+ "ece": 0.042115349045515435,
+ "mean_confidence": 0.8714557776895534,
+ "cohort": "clean and knowable",
+ "excluded_non_headline_rows": 0,
+ "ece_bins": 10
+ },
+ "transfer_dev_tuned_blend": {
+ "n": 231,
+ "correct": 208,
+ "accuracy": 0.9004329004329005,
+ "cross_entropy_from_gold": 0.24856632495280404,
+ "gold_entropy": 0.0,
+ "kl_from_gold": 0.24856632495280404,
+ "brier": 0.13696196671048083,
+ "ece": 0.05151408138171153,
+ "mean_confidence": 0.8781699403131888,
+ "cohort": "clean and knowable",
+ "excluded_non_headline_rows": 0,
+ "ece_bins": 10
+ }
+ }
+ }
+ ]
+ },
+ "jevjudge_text": {
+ "label": "JevJudge text",
+ "role": "held_out_external_evaluation",
+ "records": 724,
+ "questions": 724,
+ "headline_n": 724,
+ "runs": [
+ {
+ "run": "base4",
+ "model": "Frozen-Qwen3.5-4B",
+ "family": "Qwen3.5-4B",
+ "weights": "frozen_base",
+ "native_readout": null,
+ "zero_shot": {
+ "choice_t1": {
+ "n": 724,
+ "correct": 425,
+ "accuracy": 0.5870165745856354,
+ "cross_entropy_from_gold": 0.9990413838295111,
+ "gold_entropy": 0.0,
+ "kl_from_gold": 0.9990413838295111,
+ "brier": 0.5424665612791818,
+ "ece": 0.124790188004864,
+ "mean_confidence": 0.7118067625904997,
+ "cohort": "clean and knowable",
+ "excluded_non_headline_rows": 0,
+ "ece_bins": 10
+ }
+ },
+ "transfer_dev_tuned": {
+ "transfer_dev_calibrated_choice": {
+ "n": 724,
+ "correct": 425,
+ "accuracy": 0.5870165745856354,
+ "cross_entropy_from_gold": 0.9193370649594155,
+ "gold_entropy": 0.0,
+ "kl_from_gold": 0.9193370649594155,
+ "brier": 0.5121791926918943,
+ "ece": 0.030915955188617703,
+ "mean_confidence": 0.5929232826389871,
+ "cohort": "clean and knowable",
+ "excluded_non_headline_rows": 0,
+ "ece_bins": 10
+ }
+ }
+ },
+ {
+ "run": "pointer4",
+ "model": "JevAny-Qwen3.5-4B-Pointer",
+ "family": "Qwen3.5-4B",
+ "weights": "jevany_sft",
+ "native_readout": "pointer",
+ "zero_shot": {
+ "choice_t1": {
+ "n": 724,
+ "correct": 381,
+ "accuracy": 0.5262430939226519,
+ "cross_entropy_from_gold": 1.027944601965405,
+ "gold_entropy": 0.0,
+ "kl_from_gold": 1.027944601965405,
+ "brier": 0.5701372991529206,
+ "ece": 0.05964150883024191,
+ "mean_confidence": 0.5692400844020692,
+ "cohort": "clean and knowable",
+ "excluded_non_headline_rows": 0,
+ "ece_bins": 10
+ },
+ "native_shipped": {
+ "n": 724,
+ "correct": 420,
+ "accuracy": 0.580110497237569,
+ "cross_entropy_from_gold": 0.9966502097264682,
+ "gold_entropy": 0.0,
+ "kl_from_gold": 0.9966502097264682,
+ "brier": 0.5538821469971053,
+ "ece": 0.07444449081079427,
+ "mean_confidence": 0.621924240558197,
+ "cohort": "clean and knowable",
+ "excluded_non_headline_rows": 0,
+ "ece_bins": 10
+ },
+ "fixed_blend_0_5": {
+ "n": 724,
+ "correct": 428,
+ "accuracy": 0.5911602209944752,
+ "cross_entropy_from_gold": 0.9621095236741931,
+ "gold_entropy": 0.0,
+ "kl_from_gold": 0.9621095236741931,
+ "brier": 0.5413174973396784,
+ "ece": 0.06904408118946087,
+ "mean_confidence": 0.5874577811451718,
+ "cohort": "clean and knowable",
+ "excluded_non_headline_rows": 0,
+ "ece_bins": 10
+ }
+ },
+ "transfer_dev_tuned": {
+ "transfer_dev_calibrated_choice": {
+ "n": 724,
+ "correct": 381,
+ "accuracy": 0.5262430939226519,
+ "cross_entropy_from_gold": 1.048869475648102,
+ "gold_entropy": 0.0,
+ "kl_from_gold": 1.048869475648102,
+ "brier": 0.5765749498377344,
+ "ece": 0.0728368732066297,
+ "mean_confidence": 0.5930575646329146,
+ "cohort": "clean and knowable",
+ "excluded_non_headline_rows": 0,
+ "ece_bins": 10
+ },
+ "transfer_dev_tuned_blend": {
+ "n": 724,
+ "correct": 431,
+ "accuracy": 0.5953038674033149,
+ "cross_entropy_from_gold": 0.9665595951901557,
+ "gold_entropy": 0.0,
+ "kl_from_gold": 0.9665595951901557,
+ "brier": 0.5427219958129672,
+ "ece": 0.0883382849075786,
+ "mean_confidence": 0.5909347225687963,
+ "cohort": "clean and knowable",
+ "excluded_non_headline_rows": 0,
+ "ece_bins": 10
+ }
+ }
+ },
+ {
+ "run": "direct4",
+ "model": "JevAny-Qwen3.5-4B-Direct-Token",
+ "family": "Qwen3.5-4B",
+ "weights": "jevany_sft",
+ "native_readout": "direct-token",
+ "zero_shot": {
+ "choice_t1": {
+ "n": 724,
+ "correct": 417,
+ "accuracy": 0.5759668508287292,
+ "cross_entropy_from_gold": 0.9595839832215673,
+ "gold_entropy": 0.0,
+ "kl_from_gold": 0.9595839832215673,
+ "brier": 0.5411152452677374,
+ "ece": 0.10266762314870827,
+ "mean_confidence": 0.6575552343312288,
+ "cohort": "clean and knowable",
+ "excluded_non_headline_rows": 0,
+ "ece_bins": 10
+ },
+ "native_shipped": {
+ "n": 724,
+ "correct": 423,
+ "accuracy": 0.5842541436464088,
+ "cross_entropy_from_gold": 0.9150453634780736,
+ "gold_entropy": 0.0,
+ "kl_from_gold": 0.9150453634780736,
+ "brier": 0.5160204686928853,
+ "ece": 0.09170853327092285,
+ "mean_confidence": 0.6676634582809498,
+ "cohort": "clean and knowable",
+ "excluded_non_headline_rows": 0,
+ "ece_bins": 10
+ },
+ "fixed_blend_0_5": {
+ "n": 724,
+ "correct": 427,
+ "accuracy": 0.5897790055248618,
+ "cross_entropy_from_gold": 0.9182029376845215,
+ "gold_entropy": 0.0,
+ "kl_from_gold": 0.9182029376845215,
+ "brier": 0.5196712705778425,
+ "ece": 0.07987836608665412,
+ "mean_confidence": 0.6592116906583616,
+ "cohort": "clean and knowable",
+ "excluded_non_headline_rows": 0,
+ "ece_bins": 10
+ }
+ },
+ "transfer_dev_tuned": {
+ "transfer_dev_calibrated_choice": {
+ "n": 724,
+ "correct": 417,
+ "accuracy": 0.5759668508287292,
+ "cross_entropy_from_gold": 0.9368946832422624,
+ "gold_entropy": 0.0,
+ "kl_from_gold": 0.9368946832422624,
+ "brier": 0.5297156601921297,
+ "ece": 0.06832163591814601,
+ "mean_confidence": 0.6163866223639445,
+ "cohort": "clean and knowable",
+ "excluded_non_headline_rows": 0,
+ "ece_bins": 10
+ },
+ "transfer_dev_tuned_blend": {
+ "n": 724,
+ "correct": 429,
+ "accuracy": 0.5925414364640884,
+ "cross_entropy_from_gold": 0.905162432888216,
+ "gold_entropy": 0.0,
+ "kl_from_gold": 0.905162432888216,
+ "brier": 0.5125121544257433,
+ "ece": 0.06810038720073028,
+ "mean_confidence": 0.6374560089027138,
+ "cohort": "clean and knowable",
+ "excluded_non_headline_rows": 0,
+ "ece_bins": 10
+ }
+ }
+ },
+ {
+ "run": "base27",
+ "model": "Frozen-Qwen3.8-27B",
+ "family": "Qwen3.8-27B",
+ "weights": "frozen_base",
+ "native_readout": null,
+ "zero_shot": {
+ "choice_t1": {
+ "n": 724,
+ "correct": 447,
+ "accuracy": 0.6174033149171271,
+ "cross_entropy_from_gold": 1.0435471379379775,
+ "gold_entropy": 0.0,
+ "kl_from_gold": 1.0435471379379775,
+ "brier": 0.5477779993257101,
+ "ece": 0.18642668688209224,
+ "mean_confidence": 0.7965262996294791,
+ "cohort": "clean and knowable",
+ "excluded_non_headline_rows": 0,
+ "ece_bins": 10
+ }
+ },
+ "transfer_dev_tuned": {
+ "transfer_dev_calibrated_choice": {
+ "n": 724,
+ "correct": 447,
+ "accuracy": 0.6174033149171271,
+ "cross_entropy_from_gold": 0.8776519496796663,
+ "gold_entropy": 0.0,
+ "kl_from_gold": 0.8776519496796663,
+ "brier": 0.4981210685244041,
+ "ece": 0.08864165712545208,
+ "mean_confidence": 0.6911587312881367,
+ "cohort": "clean and knowable",
+ "excluded_non_headline_rows": 0,
+ "ece_bins": 10
+ }
+ }
+ },
+ {
+ "run": "pointer27",
+ "model": "JevAny-Qwen3.8-27B-Pointer",
+ "family": "Qwen3.8-27B",
+ "weights": "jevany_sft",
+ "native_readout": "pointer",
+ "zero_shot": {
+ "choice_t1": {
+ "n": 724,
+ "correct": 419,
+ "accuracy": 0.5787292817679558,
+ "cross_entropy_from_gold": 0.9221648550274847,
+ "gold_entropy": 0.0,
+ "kl_from_gold": 0.9221648550274847,
+ "brier": 0.5384293459629866,
+ "ece": 0.12300846658221797,
+ "mean_confidence": 0.654141167166119,
+ "cohort": "clean and knowable",
+ "excluded_non_headline_rows": 0,
+ "ece_bins": 10
+ },
+ "native_shipped": {
+ "n": 724,
+ "correct": 481,
+ "accuracy": 0.664364640883978,
+ "cross_entropy_from_gold": 0.8240838957036632,
+ "gold_entropy": 0.0,
+ "kl_from_gold": 0.8240838957036632,
+ "brier": 0.4761621807919303,
+ "ece": 0.10150582984409334,
+ "mean_confidence": 0.7093057903641973,
+ "cohort": "clean and knowable",
+ "excluded_non_headline_rows": 0,
+ "ece_bins": 10
+ },
+ "fixed_blend_0_5": {
+ "n": 724,
+ "correct": 465,
+ "accuracy": 0.6422651933701657,
+ "cross_entropy_from_gold": 0.8340868983096242,
+ "gold_entropy": 0.0,
+ "kl_from_gold": 0.8340868983096242,
+ "brier": 0.48649946917918646,
+ "ece": 0.09293531151742315,
+ "mean_confidence": 0.6762654332459023,
+ "cohort": "clean and knowable",
+ "excluded_non_headline_rows": 0,
+ "ece_bins": 10
+ }
+ },
+ "transfer_dev_tuned": {
+ "transfer_dev_calibrated_choice": {
+ "n": 724,
+ "correct": 419,
+ "accuracy": 0.5787292817679558,
+ "cross_entropy_from_gold": 1.0389937594613554,
+ "gold_entropy": 0.0,
+ "kl_from_gold": 1.0389937594613554,
+ "brier": 0.5859846798375742,
+ "ece": 0.18095024427280532,
+ "mean_confidence": 0.7592297383784663,
+ "cohort": "clean and knowable",
+ "excluded_non_headline_rows": 0,
+ "ece_bins": 10
+ },
+ "transfer_dev_tuned_blend": {
+ "n": 724,
+ "correct": 466,
+ "accuracy": 0.643646408839779,
+ "cross_entropy_from_gold": 0.8472248615476166,
+ "gold_entropy": 0.0,
+ "kl_from_gold": 0.8472248615476166,
+ "brier": 0.49231968730390446,
+ "ece": 0.10423989831022704,
+ "mean_confidence": 0.7135438249208921,
+ "cohort": "clean and knowable",
+ "excluded_non_headline_rows": 0,
+ "ece_bins": 10
+ }
+ }
+ }
+ ],
+ "external_baselines": [
+ {
+ "model": "Jev 1.13 (OpenRouter)",
+ "canonical_model": "TypeSafe Jev 1.13.0",
+ "kind": "external_api",
+ "result_source": "api_rerun",
+ "status": "OpenRouter full-context rerun",
+ "accuracy": 0.6505524861878453,
+ "nll": 0.836,
+ "brier": 0.461,
+ "ece": 0.076,
+ "answered": 724,
+ "requested": 724
+ },
+ {
+ "model": "Kev-27B",
+ "kind": "open_kev",
+ "status": "local full-context 4,096-token chunked-KV rerun",
+ "accuracy": 0.6422651933701657,
+ "nll": 0.8774529517226382,
+ "brier": 0.48251125757984015,
+ "ece": 0.07166661519467314,
+ "answered": 724,
+ "requested": 724,
+ "repository": "jaredpalmer/kev-27b",
+ "revision": "01b81998019be550f0ae858727df49bac9511195"
+ }
+ ]
+ },
+ "jevjudge_full": {
+ "label": "JevJudge full",
+ "records": 3220,
+ "status": "unsupported",
+ "accuracy": null,
+ "reason": "Training-free choice-token readout is text-only; image/video records are not stripped or relabeled as a full-suite result.",
+ "native_checkpoint_results": "results/external-zero-shot-v1.json#jevjudge_full"
+ }
+ },
+ "legacy_results": {
+ "status": "retained unchanged",
+ "artifact": "results/letter-readout-v1.json"
+ }
+}
diff --git a/scripts/build_choice_readout_results.py b/scripts/build_choice_readout_results.py
new file mode 100644
index 0000000..0a37b3c
--- /dev/null
+++ b/scripts/build_choice_readout_results.py
@@ -0,0 +1,985 @@
+#!/usr/bin/env python3
+"""Validate, summarize, and plot the five-run choice-readout matrix.
+
+The input directory must contain ``base4``, ``pointer4``, ``direct4``,
+``base27``, and ``pointer27`` subdirectories. Each subdirectory contains the
+``manifest.json`` emitted by ``evaluate_choice_readout_matrix.py`` and the
+``ensemble.json`` emitted by ``fit_choice_ensemble.py``.
+
+The builder fails closed on incomplete panels, component/count drift, or
+cross-run dataset-hash drift. It never substitutes partial coverage or a
+media-stripped score for JevJudge full.
+"""
+
+from __future__ import annotations
+
+import argparse
+import copy
+import hashlib
+import json
+import math
+import re
+from pathlib import Path
+from typing import Any
+
+
+ROOT = Path(__file__).resolve().parents[1]
+DEFAULT_EXTERNAL = ROOT / "results/external-zero-shot-v1.json"
+DEFAULT_MODEL_CATALOG = ROOT / "results/model-family-v2.json"
+DEFAULT_JSON_OUT = ROOT / "results/choice-readout-v2.json"
+DEFAULT_SVG_OUT = ROOT / "docs/choice-readout-results.svg"
+
+INK = "#213248"
+MUTED = "#64748B"
+RULE = "#E4E9EF"
+CHOICE = "#278577"
+NATIVE = "#8493A6"
+TUNED = "#165F55"
+EXTERNAL = "#A17BB7"
+TYPESAFE = "#C08A42"
+
+RUN_SPECS = (
+ {
+ "key": "base4",
+ "family": "Qwen3.5-4B",
+ "public_repository": "Qwen/Qwen3.5-4B",
+ "weights": "frozen_base",
+ "native_readout": None,
+ "native_decision_mode": None,
+ },
+ {
+ "key": "pointer4",
+ "family": "Qwen3.5-4B",
+ "public_repository": "SimpleJev/JevAny-Qwen3.5-4B-LoRA",
+ "weights": "jevany_sft",
+ "native_readout": "pointer",
+ "native_decision_mode": "pointer",
+ },
+ {
+ "key": "direct4",
+ "family": "Qwen3.5-4B",
+ "public_repository": "SimpleJev/JevAny-Qwen3.5-4B-Direct-Token-LoRA",
+ "weights": "jevany_sft",
+ "native_readout": "direct-token",
+ "native_decision_mode": "lm_token",
+ },
+ {
+ "key": "base27",
+ "family": "Qwen3.8-27B",
+ "public_repository": "Qwen/Qwen3.8-27B",
+ "weights": "frozen_base",
+ "native_readout": None,
+ "native_decision_mode": None,
+ },
+ {
+ "key": "pointer27",
+ "family": "Qwen3.8-27B",
+ "public_repository": "SimpleJev/JevAny-Qwen3.8-27B-LoRA",
+ "weights": "jevany_sft",
+ "native_readout": "pointer",
+ "native_decision_mode": "pointer",
+ },
+)
+
+DATASETS = {
+ "transfer_calibration": {
+ "label": "Transfer-v9 development",
+ "role": "selection_only",
+ "records": 1_264,
+ "questions": 1_264,
+ "headline_n": 1_046,
+ },
+ "transfer_test": {
+ "label": "Transfer-v9 test",
+ "role": "held_out_evaluation",
+ "records": 1_264,
+ "questions": 1_264,
+ "headline_n": 1_046,
+ },
+ "typed_test": {
+ "label": "Typed Decisions test",
+ "role": "held_out_external_evaluation",
+ "records": 400,
+ "questions": 2_000,
+ "headline_n": 2_000,
+ },
+ "jevbench_public": {
+ "label": "JevBench public development",
+ "role": "public_diagnostic_not_used_for_selection",
+ "records": 231,
+ "questions": 231,
+ "headline_n": 231,
+ },
+ "jevjudge_text": {
+ "label": "JevJudge text",
+ "role": "held_out_external_evaluation",
+ "records": 724,
+ "questions": 724,
+ "headline_n": 724,
+ },
+}
+
+ZERO_SHOT_CONFIGS = (
+ "choice_t1",
+ "native_shipped",
+ "fixed_blend_0_5",
+)
+TUNED_CONFIGS = (
+ "transfer_dev_calibrated_choice",
+ "transfer_dev_tuned_blend",
+)
+BASE_CONFIGS = {"choice_t1", "transfer_dev_calibrated_choice"}
+CHECKPOINT_CONFIGS = BASE_CONFIGS | {
+ "native_shipped",
+ "fixed_blend_0_5",
+ "transfer_dev_tuned_blend",
+}
+
+
+def _read_object(path: Path) -> dict:
+ try:
+ value = json.loads(path.read_text(encoding="utf-8"))
+ except FileNotFoundError as error:
+ raise ValueError(f"missing input artifact: {path}") from error
+ except json.JSONDecodeError as error:
+ raise ValueError(f"invalid JSON artifact: {path}: {error}") from error
+ if not isinstance(value, dict):
+ raise ValueError(f"JSON artifact must be an object: {path}")
+ return value
+
+
+def _sha256(path: Path) -> str:
+ digest = hashlib.sha256()
+ with path.open("rb") as stream:
+ for block in iter(lambda: stream.read(1024 * 1024), b""):
+ digest.update(block)
+ return digest.hexdigest()
+
+
+def _display_path(path: Path) -> str:
+ resolved = path.resolve()
+ try:
+ return str(resolved.relative_to(ROOT.resolve()))
+ except ValueError:
+ return str(resolved)
+
+
+def _finite(value: Any) -> bool:
+ return (
+ not isinstance(value, bool)
+ and isinstance(value, (int, float))
+ and math.isfinite(value)
+ )
+
+
+def _probability(value: Any, context: str) -> float:
+ if not _finite(value) or not 0 <= value <= 1:
+ raise ValueError(f"{context} must be a finite number in [0, 1]")
+ return float(value)
+
+
+def _positive(value: Any, context: str) -> float:
+ if not _finite(value) or value <= 0:
+ raise ValueError(f"{context} must be a finite positive number")
+ return float(value)
+
+
+def _metric(metric: Any, expected_n: int, context: str) -> dict:
+ if not isinstance(metric, dict):
+ raise ValueError(f"{context} must be an object")
+ if metric.get("n") != expected_n:
+ raise ValueError(
+ f"{context} expected n={expected_n}, got {metric.get('n')!r}"
+ )
+ correct = metric.get("correct")
+ if isinstance(correct, bool) or not isinstance(correct, int) or not 0 <= correct <= expected_n:
+ raise ValueError(f"{context}.correct must be an integer in [0, {expected_n}]")
+ accuracy = _probability(metric.get("accuracy"), f"{context}.accuracy")
+ if not math.isclose(accuracy, correct / expected_n, rel_tol=0, abs_tol=1e-12):
+ raise ValueError(f"{context}.accuracy does not equal correct / n")
+ for key in (
+ "cross_entropy_from_gold",
+ "gold_entropy",
+ "kl_from_gold",
+ "brier",
+ "ece",
+ "mean_confidence",
+ ):
+ if key in metric and not _finite(metric[key]):
+ raise ValueError(f"{context}.{key} must be finite")
+ return copy.deepcopy(metric)
+
+
+def _manifest_component_accuracy(report: dict, component: str, expected_n: int, context: str) -> float:
+ try:
+ clean = report["components"][component]["clean"]
+ except (KeyError, TypeError) as error:
+ raise ValueError(f"{context}: missing {component} clean report") from error
+ if clean.get("n") != expected_n:
+ raise ValueError(
+ f"{context}: {component} expected clean n={expected_n}, "
+ f"got {clean.get('n')!r}"
+ )
+ return _probability(clean.get("acc"), f"{context}.{component}.clean.acc")
+
+
+def _validate_source(spec: dict, manifest: dict, context: str) -> None:
+ source = manifest.get("source")
+ predictor = manifest.get("predictor")
+ if not isinstance(source, dict) or not isinstance(predictor, dict):
+ raise ValueError(f"{context}: missing source or predictor provenance")
+ is_base = spec["weights"] == "frozen_base"
+ if is_base:
+ if not source.get("base") or source.get("checkpoint") is not None:
+ raise ValueError(f"{context}: frozen-base run must use source.base only")
+ if predictor.get("adapter_applied") is not False:
+ raise ValueError(f"{context}: frozen-base run unexpectedly applied an adapter")
+ if manifest.get("checkpoint_artifacts") is not None:
+ raise ValueError(f"{context}: frozen-base run has checkpoint artifacts")
+ else:
+ if not source.get("checkpoint") or source.get("base") is not None:
+ raise ValueError(f"{context}: checkpoint run must use source.checkpoint only")
+ if predictor.get("adapter_applied") is not True:
+ raise ValueError(f"{context}: checkpoint run did not apply its adapter")
+ artifacts = manifest.get("checkpoint_artifacts")
+ if not isinstance(artifacts, dict) or not isinstance(artifacts.get("files"), dict):
+ raise ValueError(f"{context}: checkpoint artifact hashes are missing")
+ if predictor.get("native_decision_mode") != spec["native_decision_mode"]:
+ raise ValueError(
+ f"{context}: expected native_decision_mode={spec['native_decision_mode']!r}, "
+ f"got {predictor.get('native_decision_mode')!r}"
+ )
+
+
+def _immutable_revision(value: Any) -> str | None:
+ if isinstance(value, str) and re.fullmatch(r"[0-9a-f]{40}", value):
+ return value
+ return None
+
+
+def _snapshot_revision(path: Any, repository: str) -> str | None:
+ """Return a Hub snapshot revision only when the path names this repository."""
+
+ if not isinstance(path, str):
+ return None
+ encoded_repository = repository.replace("/", "--")
+ match = re.search(
+ rf"(?:^|/)models--{re.escape(encoded_repository)}/snapshots/([0-9a-f]{{40}})(?:/|$)",
+ path,
+ )
+ return match.group(1) if match else None
+
+
+def _public_source(spec: dict, manifest: dict) -> dict:
+ """Expose reproducible public identity separately from local load paths.
+
+ A local release directory is not silently treated as a public Hub revision.
+ In that case the evaluated files remain identified by their SHA-256 values
+ and the missing public revision is explicit.
+ """
+
+ repository = spec["public_repository"]
+ base_loading = manifest.get("base_loading")
+ base_loading = base_loading if isinstance(base_loading, dict) else {}
+ base_repository = base_loading.get("canonical_base")
+ base_revision = _immutable_revision(base_loading.get("canonical_revision"))
+ if not isinstance(base_repository, str) or not base_repository:
+ raise ValueError(f"{spec['key']}: canonical base repository is missing")
+ if base_revision is None:
+ raise ValueError(f"{spec['key']}: immutable canonical base revision is missing")
+
+ artifacts = manifest.get("checkpoint_artifacts")
+ artifact_hashes = {}
+ if isinstance(artifacts, dict) and isinstance(artifacts.get("files"), dict):
+ for filename, metadata in artifacts["files"].items():
+ if not isinstance(metadata, dict) or not re.fullmatch(
+ r"[0-9a-f]{64}", str(metadata.get("sha256", ""))
+ ):
+ raise ValueError(
+ f"{spec['key']}: checkpoint artifact {filename!r} lacks SHA-256"
+ )
+ artifact_hashes[filename] = {
+ "sha256": metadata["sha256"],
+ "bytes": metadata.get("bytes"),
+ }
+
+ if spec["weights"] == "frozen_base":
+ if repository != base_repository:
+ raise ValueError(
+ f"{spec['key']}: public base repository does not match canonical base"
+ )
+ revision = base_revision
+ revision_evidence = "manifest.base_loading.canonical_revision"
+ repository_evidence = "manifest.base_loading.canonical_base"
+ limitation = None
+ else:
+ source = manifest["source"]["checkpoint"]
+ revision = _snapshot_revision(source, repository)
+ revision_evidence = "manifest.source.checkpoint" if revision else None
+ repository_evidence = (
+ "manifest.source.checkpoint"
+ if revision
+ else "results/model-family-v2.json#released_models"
+ )
+ limitation = None if revision else (
+ "The evaluated checkpoint came from a local release. Its exact files "
+ "are pinned below by SHA-256, but the run manifest and release metadata "
+ "do not prove an immutable public-repository revision for those bytes."
+ )
+
+ return {
+ "repository": repository,
+ "url": f"https://huggingface.co/{repository}",
+ "revision": revision,
+ "revision_status": "verified" if revision else "not_verified",
+ "repository_evidence": repository_evidence,
+ "revision_evidence": revision_evidence,
+ "evaluated_artifact_sha256": artifact_hashes,
+ "base_model": {
+ "repository": base_repository,
+ "revision": base_revision,
+ },
+ "limitation": limitation,
+ }
+
+
+def _load_run(run_root: Path, spec: dict) -> dict:
+ directory = run_root / spec["key"]
+ manifest_path = directory / "manifest.json"
+ ensemble_path = directory / "ensemble.json"
+ manifest = _read_object(manifest_path)
+ ensemble = _read_object(ensemble_path)
+ context = spec["key"]
+
+ model = manifest.get("model")
+ if not isinstance(model, str) or not model.strip():
+ raise ValueError(f"{context}: manifest model must be a non-empty string")
+ if ensemble.get("model") != model:
+ raise ValueError(f"{context}: manifest and ensemble model names differ")
+ _validate_source(spec, manifest, context)
+
+ expected_components = {"choice"}
+ if spec["native_readout"] is not None:
+ expected_components.add("native")
+ reports = manifest.get("reports")
+ datasets = ensemble.get("datasets")
+ if not isinstance(reports, dict) or not isinstance(datasets, dict):
+ raise ValueError(f"{context}: missing report or ensemble datasets")
+ if set(DATASETS) - set(reports) or set(DATASETS) - set(datasets):
+ raise ValueError(f"{context}: one or more required datasets are missing")
+
+ expected_configs = BASE_CONFIGS if spec["native_readout"] is None else CHECKPOINT_CONFIGS
+ summarized = {}
+ for dataset, expected in DATASETS.items():
+ report = reports[dataset]
+ if report.get("records") != expected["records"]:
+ raise ValueError(
+ f"{context}/{dataset}: expected {expected['records']} records, "
+ f"got {report.get('records')!r}"
+ )
+ if report.get("questions") != expected["questions"]:
+ raise ValueError(
+ f"{context}/{dataset}: expected {expected['questions']} questions, "
+ f"got {report.get('questions')!r}"
+ )
+ components = set(report.get("components", {}))
+ executed = set(report.get("components_executed", components))
+ if components != expected_components or executed != expected_components:
+ raise ValueError(
+ f"{context}/{dataset}: expected components {sorted(expected_components)}, "
+ f"got reports={sorted(components)}, executed={sorted(executed)}"
+ )
+ component_accuracy = {
+ component: _manifest_component_accuracy(
+ report, component, expected["headline_n"], f"{context}/{dataset}"
+ )
+ for component in sorted(expected_components)
+ }
+
+ configurations = datasets[dataset]
+ if not isinstance(configurations, dict) or set(configurations) != expected_configs:
+ raise ValueError(
+ f"{context}/{dataset}: expected ensemble configs "
+ f"{sorted(expected_configs)}, got "
+ f"{sorted(configurations) if isinstance(configurations, dict) else configurations!r}"
+ )
+ checked = {
+ name: _metric(
+ values,
+ expected["headline_n"],
+ f"{context}/{dataset}/{name}",
+ )
+ for name, values in configurations.items()
+ }
+ expected_excluded = expected["questions"] - expected["headline_n"]
+ for name, values in checked.items():
+ if values.get("excluded_non_headline_rows") != expected_excluded:
+ raise ValueError(
+ f"{context}/{dataset}/{name}: expected "
+ f"excluded_non_headline_rows={expected_excluded}"
+ )
+ if not math.isclose(
+ checked["choice_t1"]["accuracy"], component_accuracy["choice"],
+ rel_tol=0, abs_tol=1e-12,
+ ):
+ raise ValueError(f"{context}/{dataset}: choice component and ensemble are not aligned")
+ if checked["choice_t1"]["correct"] != checked["transfer_dev_calibrated_choice"]["correct"]:
+ raise ValueError(
+ f"{context}/{dataset}: temperature-only calibration changed hard decisions"
+ )
+ if "native" in expected_components and not math.isclose(
+ checked["native_shipped"]["accuracy"], component_accuracy["native"],
+ rel_tol=0, abs_tol=1e-12,
+ ):
+ raise ValueError(f"{context}/{dataset}: native component and ensemble are not aligned")
+
+ summarized[dataset] = {
+ "zero_shot": {
+ key: checked[key] for key in ZERO_SHOT_CONFIGS if key in checked
+ },
+ "transfer_dev_tuned": {
+ key: checked[key] for key in TUNED_CONFIGS if key in checked
+ },
+ }
+
+ protocol = ensemble.get("protocol")
+ selected = ensemble.get("selected")
+ if not isinstance(protocol, dict) or protocol.get("selection_rows") != 1_046:
+ raise ValueError(f"{context}: ensemble must select on 1,046 Transfer-dev rows")
+ if not isinstance(selected, dict):
+ raise ValueError(f"{context}: missing selected ensemble settings")
+ choice_temperature = _positive(
+ selected.get("choice_temperature"), f"{context}.choice_temperature"
+ )
+ blend_temperature = _positive(
+ selected.get("blend_temperature"), f"{context}.blend_temperature"
+ )
+ native_weight = _probability(
+ selected.get("native_weight"), f"{context}.native_weight"
+ )
+ if spec["native_readout"] is None and native_weight != 0:
+ raise ValueError(f"{context}: frozen base selected a native weight")
+
+ dataset_hashes = manifest.get("datasets")
+ runtime = manifest.get("runtime")
+ if not isinstance(dataset_hashes, dict) or not dataset_hashes:
+ raise ValueError(f"{context}: dataset hashes are missing")
+ if not isinstance(runtime, dict) or not runtime.get("code_revision"):
+ raise ValueError(f"{context}: runtime/code revision is missing")
+
+ return {
+ "key": spec["key"],
+ "model": model,
+ "family": spec["family"],
+ "weights": spec["weights"],
+ "native_readout": spec["native_readout"],
+ "artifacts": {
+ "manifest": {
+ "path": _display_path(manifest_path),
+ "sha256": _sha256(manifest_path),
+ },
+ "ensemble": {
+ "path": _display_path(ensemble_path),
+ "sha256": _sha256(ensemble_path),
+ },
+ },
+ "source": copy.deepcopy(manifest.get("source")),
+ "public_source": _public_source(spec, manifest),
+ "checkpoint_artifacts": copy.deepcopy(manifest.get("checkpoint_artifacts")),
+ "base_loading": copy.deepcopy(manifest.get("base_loading")),
+ "predictor": copy.deepcopy(manifest.get("predictor")),
+ "runtime": copy.deepcopy(runtime),
+ "dataset_hashes": copy.deepcopy(dataset_hashes),
+ "evaluation_protocol": copy.deepcopy(manifest.get("protocol")),
+ "ensemble_protocol": copy.deepcopy(protocol),
+ "selected": {
+ **copy.deepcopy(selected),
+ "choice_temperature": choice_temperature,
+ "blend_temperature": blend_temperature,
+ "native_weight": native_weight,
+ },
+ "datasets": summarized,
+ }
+
+
+def _find_model(rows: Any, name: str, context: str) -> dict:
+ if not isinstance(rows, list):
+ raise ValueError(f"{context}: models must be a list")
+ matches = [row for row in rows if isinstance(row, dict) and row.get("model") == name]
+ if len(matches) != 1:
+ raise ValueError(f"{context}: expected exactly one {name!r} row")
+ _probability(matches[0].get("accuracy"), f"{context}/{name}.accuracy")
+ return copy.deepcopy(matches[0])
+
+
+def _load_external(path: Path) -> dict:
+ external = _read_object(path)
+ typed = external.get("typed_decisions")
+ text = external.get("jevjudge_text")
+ if not isinstance(typed, dict) or not isinstance(text, dict):
+ raise ValueError("external artifact is missing Typed Decisions or JevJudge text")
+
+ typed_rows = typed.get("models")
+ if not isinstance(typed_rows, list):
+ raise ValueError("external Typed Decisions models must be a list")
+ published = [
+ row for row in typed_rows
+ if isinstance(row, dict)
+ and row.get("kind") == "published_only"
+ and _finite(row.get("accuracy"))
+ ]
+ if not published:
+ raise ValueError("external artifact has no published Typed Decisions baseline")
+ typed_top = copy.deepcopy(max(published, key=lambda row: (row["accuracy"], row["model"])))
+ _probability(typed_top.get("accuracy"), "Typed published top accuracy")
+ typed_typesafe = _find_model(
+ typed_rows, "Jev 1.13 (OpenRouter)", "external/typed_decisions"
+ )
+
+ text_rows = text.get("models")
+ text_typesafe = _find_model(
+ text_rows, "Jev 1.13 (OpenRouter)", "external/jevjudge_text"
+ )
+ kev_rows = [
+ row for row in text_rows
+ if isinstance(row, dict)
+ and row.get("kind") == "open_kev"
+ and _finite(row.get("accuracy"))
+ ] if isinstance(text_rows, list) else []
+ if not kev_rows:
+ raise ValueError("external artifact has no scored JevJudge-text Kev baseline")
+ text_kev = copy.deepcopy(max(kev_rows, key=lambda row: (row["accuracy"], row["model"])))
+ for row in (text_typesafe, text_kev):
+ if (row.get("answered"), row.get("requested")) != (724, 724):
+ raise ValueError(
+ f"external/jevjudge_text/{row['model']}: expected 724/724 coverage"
+ )
+
+ return {
+ "artifact": {
+ "path": _display_path(path),
+ "sha256": _sha256(path),
+ "artifact_version": external.get("artifact_version"),
+ "provenance": copy.deepcopy(external.get("provenance")),
+ },
+ "typed_decisions": [typed_top, typed_typesafe],
+ "jevjudge_text": [text_typesafe, text_kev],
+ }
+
+
+def _load_repository_catalog(path: Path) -> dict:
+ catalog = _read_object(path)
+ released = catalog.get("released_models")
+ repositories = {
+ row.get("repository")
+ for row in released
+ if isinstance(row, dict) and isinstance(row.get("repository"), str)
+ } if isinstance(released, list) else set()
+ required = {
+ spec["public_repository"]
+ for spec in RUN_SPECS
+ if spec["weights"] != "frozen_base"
+ }
+ missing = required - repositories
+ if missing:
+ raise ValueError(
+ "model release catalog is missing canonical repositories: "
+ + ", ".join(sorted(missing))
+ )
+ return {
+ "path": _display_path(path),
+ "sha256": _sha256(path),
+ "release": catalog.get("release"),
+ }
+
+
+def build_results(
+ run_root: Path,
+ external_path: Path = DEFAULT_EXTERNAL,
+ model_catalog_path: Path = DEFAULT_MODEL_CATALOG,
+) -> dict:
+ """Build a validated, JSON-serializable result artifact without writing it."""
+
+ run_root = Path(run_root)
+ repository_catalog = _load_repository_catalog(Path(model_catalog_path))
+ runs = [_load_run(run_root, spec) for spec in RUN_SPECS]
+ if len({run["model"] for run in runs}) != len(runs):
+ raise ValueError("the five run manifests must have unique model labels")
+
+ reference_hashes = runs[0]["dataset_hashes"]
+ for run in runs[1:]:
+ if run["dataset_hashes"] != reference_hashes:
+ raise ValueError(
+ f"{run['key']}: dataset hashes differ from {runs[0]['key']}"
+ )
+ revisions = {run["runtime"].get("code_revision") for run in runs}
+ if len(revisions) != 1:
+ raise ValueError("all five runs must use one code revision")
+
+ external = _load_external(Path(external_path))
+ result_datasets = {}
+ for dataset, expected in DATASETS.items():
+ result_datasets[dataset] = {
+ **copy.deepcopy(expected),
+ "runs": [
+ {
+ "run": run["key"],
+ "model": run["model"],
+ "family": run["family"],
+ "weights": run["weights"],
+ "native_readout": run["native_readout"],
+ **copy.deepcopy(run["datasets"][dataset]),
+ }
+ for run in runs
+ ],
+ }
+ result_datasets["typed_test"]["external_baselines"] = copy.deepcopy(
+ external["typed_decisions"]
+ )
+ result_datasets["jevjudge_text"]["external_baselines"] = copy.deepcopy(
+ external["jevjudge_text"]
+ )
+ result_datasets["jevjudge_full"] = {
+ "label": "JevJudge full",
+ "records": 3_220,
+ "status": "unsupported",
+ "accuracy": None,
+ "reason": (
+ "Training-free choice-token readout is text-only; image/video records "
+ "are not stripped or relabeled as a full-suite result."
+ ),
+ "native_checkpoint_results": "results/external-zero-shot-v1.json#jevjudge_full",
+ }
+
+ public_runs = []
+ for run in runs:
+ public_runs.append({key: copy.deepcopy(value) for key, value in run.items() if key != "datasets"})
+
+ return {
+ "schema_version": 2,
+ "artifact": "choice-readout-v2",
+ "title": "Training-free choice-token readout evaluation",
+ "generated_by": "scripts/build_choice_readout_results.py",
+ "method": {
+ "readout": (
+ "Constrain the next token to exact one-character option IDs and "
+ "renormalize their probability mass; this can change argmax decisions."
+ ),
+ "training": "No additional training for the choice-token readout.",
+ "option_ids": "A-Z followed by a-z; at most 52 options.",
+ "calibration": (
+ "Scalar temperature changes probabilities but not argmax accuracy."
+ ),
+ "ensemble": (
+ "Native + choice configurations use log-linear/geometric pooling."
+ ),
+ },
+ "protocol_groups": {
+ "zero_shot": {
+ "meaning": (
+ "No Typed, JevJudge, JevBench, or Transfer-test labels tune the "
+ "readout configuration."
+ ),
+ "configurations": list(ZERO_SHOT_CONFIGS),
+ },
+ "transfer_dev_tuned": {
+ "meaning": (
+ "Weight and additional temperature are selected only on Transfer-v9 "
+ "development, then frozen for every displayed evaluation panel."
+ ),
+ "configurations": list(TUNED_CONFIGS),
+ "accuracy_note": (
+ "Temperature-only calibrated choice has the same hard accuracy as "
+ "choice_t1; only a changed blend weight can change its argmax."
+ ),
+ },
+ },
+ "expected_counts": copy.deepcopy(DATASETS),
+ "dataset_hashes": copy.deepcopy(reference_hashes),
+ "code_revision": next(iter(revisions)),
+ "source_runs": public_runs,
+ "external_source": external["artifact"],
+ "repository_catalog": repository_catalog,
+ "datasets": result_datasets,
+ "legacy_results": {
+ "status": "retained unchanged",
+ "artifact": "results/letter-readout-v1.json",
+ },
+ }
+
+
+def write_results(path: Path, artifact: dict) -> None:
+ path = Path(path)
+ path.parent.mkdir(parents=True, exist_ok=True)
+ payload = json.dumps(
+ artifact, indent=2, ensure_ascii=False, allow_nan=False
+ ) + "\n"
+ path.write_text(payload, encoding="utf-8")
+
+
+def _run_result(artifact: dict, dataset: str, run_key: str) -> dict:
+ matches = [
+ row for row in artifact["datasets"][dataset]["runs"]
+ if row["run"] == run_key
+ ]
+ if len(matches) != 1:
+ raise ValueError(f"{dataset}: expected exactly one {run_key} result")
+ return matches[0]
+
+
+def _plot_rows(artifact: dict, dataset: str) -> list[dict]:
+ def accuracy(run: str, group: str, config: str) -> float:
+ return _run_result(artifact, dataset, run)[group][config]["accuracy"]
+
+ rows = [
+ {
+ "label": "Frozen Qwen3.5 4B ยท Choice",
+ "group": "zero_shot",
+ "kind": "choice",
+ "accuracy": accuracy("base4", "zero_shot", "choice_t1"),
+ },
+ {
+ "label": "JevAny 4B Pointer ยท Native",
+ "group": "zero_shot",
+ "kind": "native",
+ "accuracy": accuracy("pointer4", "zero_shot", "native_shipped"),
+ },
+ {
+ "label": "JevAny 4B Pointer ยท Choice",
+ "group": "zero_shot",
+ "kind": "choice",
+ "accuracy": accuracy("pointer4", "zero_shot", "choice_t1"),
+ },
+ {
+ "label": "JevAny 4B Direct-Token ยท Native",
+ "group": "zero_shot",
+ "kind": "native",
+ "accuracy": accuracy("direct4", "zero_shot", "native_shipped"),
+ },
+ {
+ "label": "JevAny 4B Direct-Token ยท Choice",
+ "group": "zero_shot",
+ "kind": "choice",
+ "accuracy": accuracy("direct4", "zero_shot", "choice_t1"),
+ },
+ {
+ "label": "Frozen Qwen3.8 27B ยท Choice",
+ "group": "zero_shot",
+ "kind": "choice",
+ "accuracy": accuracy("base27", "zero_shot", "choice_t1"),
+ },
+ {
+ "label": "JevAny 27B Pointer ยท Native",
+ "group": "zero_shot",
+ "kind": "native",
+ "accuracy": accuracy("pointer27", "zero_shot", "native_shipped"),
+ },
+ {
+ "label": "JevAny 27B Pointer ยท Choice",
+ "group": "zero_shot",
+ "kind": "choice",
+ "accuracy": accuracy("pointer27", "zero_shot", "choice_t1"),
+ },
+ {
+ "label": "JevAny 4B Pointer ยท Tuned blend",
+ "group": "transfer_dev_tuned",
+ "kind": "tuned",
+ "accuracy": accuracy(
+ "pointer4", "transfer_dev_tuned", "transfer_dev_tuned_blend"
+ ),
+ },
+ {
+ "label": "JevAny 4B Direct-Token ยท Tuned blend",
+ "group": "transfer_dev_tuned",
+ "kind": "tuned",
+ "accuracy": accuracy(
+ "direct4", "transfer_dev_tuned", "transfer_dev_tuned_blend"
+ ),
+ },
+ {
+ "label": "JevAny 27B Pointer ยท Tuned blend",
+ "group": "transfer_dev_tuned",
+ "kind": "tuned",
+ "accuracy": accuracy(
+ "pointer27", "transfer_dev_tuned", "transfer_dev_tuned_blend"
+ ),
+ },
+ ]
+ baselines = artifact["datasets"][dataset]["external_baselines"]
+ for baseline in baselines:
+ is_typesafe = baseline["model"] == "Jev 1.13 (OpenRouter)"
+ label = (
+ "TypeSafe Jev 1.13"
+ if is_typesafe
+ else f"{baseline['model']} ยท "
+ + ("published" if dataset == "typed_test" else "open")
+ )
+ rows.append({
+ "label": label,
+ "group": "external",
+ "kind": "typesafe" if is_typesafe else "external",
+ "accuracy": baseline["accuracy"],
+ "published": dataset == "typed_test" and not is_typesafe,
+ })
+ return rows
+
+
+def render_svg(artifact: dict, output: Path) -> None:
+ """Render the compact two-panel README figure from a validated artifact."""
+
+ try:
+ import matplotlib
+ except ImportError as error:
+ raise RuntimeError("matplotlib is required to render the SVG") from error
+
+ matplotlib.use("Agg")
+ import matplotlib.pyplot as plt
+ from matplotlib.patches import Patch
+
+ plt.rcParams.update({
+ "font.family": ["DejaVu Sans", "sans-serif"],
+ "svg.fonttype": "none",
+ "svg.hashsalt": "jevany-choice-readout-v2",
+ "text.color": INK,
+ "hatch.linewidth": 0.7,
+ })
+ panels = (
+ ("typed_test", "Typed Decisions", "2,000 decisions ยท accuracy (%)"),
+ ("jevjudge_text", "JevJudge text", "724 records ยท accuracy (%)"),
+ )
+ fig, axes = plt.subplots(1, 2, figsize=(18, 9.3), dpi=110, facecolor="white")
+ fig.subplots_adjust(left=0.19, right=0.988, bottom=0.12, top=0.78, wspace=0.30)
+ fig.text(0.035, 0.955, "Training-free choice readout", fontsize=24, weight="bold")
+ fig.text(
+ 0.035,
+ 0.914,
+ "Zero-shot readouts and Transfer-dev-tuned blends ยท no Typed or JevJudge labels used for tuning",
+ fontsize=12.2,
+ color=MUTED,
+ )
+ fig.legend(
+ handles=[
+ Patch(facecolor=CHOICE, label="Choice T=1 ยท zero-shot"),
+ Patch(facecolor=NATIVE, label="Native shipped ยท zero-shot"),
+ Patch(facecolor=TUNED, hatch="///", label="Transfer-dev-tuned blend"),
+ Patch(facecolor=TYPESAFE, label="TypeSafe Jev"),
+ Patch(facecolor=EXTERNAL, hatch="////", label="Published / open baseline"),
+ ],
+ loc="upper right",
+ bbox_to_anchor=(0.988, 0.982),
+ ncol=3,
+ frameon=False,
+ fontsize=10.1,
+ handlelength=1.2,
+ columnspacing=1.25,
+ )
+
+ description_parts = []
+ colors = {
+ "choice": CHOICE,
+ "native": NATIVE,
+ "tuned": TUNED,
+ "external": EXTERNAL,
+ "typesafe": TYPESAFE,
+ }
+ positions = list(range(8)) + list(range(9, 12)) + list(range(13, 15))
+ for ax, (dataset, title, scope) in zip(axes, panels, strict=True):
+ rows = _plot_rows(artifact, dataset)
+ if len(rows) != len(positions):
+ raise ValueError(f"{dataset}: plot requires exactly 15 rows")
+ values = [row["accuracy"] * 100 for row in rows]
+ xmax = min(100, max(70, int(math.ceil((max(values) + 4) / 10) * 10)))
+ ax.axhspan(-0.7, 7.65, color="#F7FAFC", zorder=0)
+ ax.axhspan(8.35, 11.65, color="#EDF7F4", zorder=0)
+ ax.axhspan(12.35, 14.65, color="#FAF7FC", zorder=0)
+ for y, row, value in zip(positions, rows, values, strict=True):
+ bars = ax.barh(
+ y,
+ value,
+ height=0.64,
+ color=colors[row["kind"]],
+ edgecolor="#805A2B" if row.get("published") else "none",
+ linewidth=0.7 if row.get("published") else 0,
+ zorder=3,
+ )
+ if row["kind"] == "tuned":
+ bars[0].set_hatch("///")
+ elif row.get("published"):
+ bars[0].set_hatch("////")
+ ax.text(
+ min(value + xmax * 0.015, xmax * 0.985),
+ y,
+ f"{value:.1f}",
+ ha="right" if value > xmax * 0.91 else "left",
+ va="center",
+ fontsize=9.2,
+ color=INK,
+ weight="bold" if row["group"] != "external" else "normal",
+ )
+ ax.text(0, -0.78, "ZERO-SHOT", fontsize=9.0, color=MUTED, weight="bold")
+ ax.text(0, 8.22, "TRANSFER-DEV-TUNED", fontsize=9.0, color=TUNED, weight="bold")
+ ax.text(0, 12.22, "EXTERNAL REFERENCE", fontsize=9.0, color=MUTED, weight="bold")
+ ax.set_title(title, loc="left", fontsize=17, color=INK, weight="bold", pad=28)
+ ax.text(0, 1.015, scope, transform=ax.transAxes, fontsize=10.2, color=MUTED)
+ ax.set_xlim(0, xmax)
+ ticks = list(range(0, xmax + 1, 20))
+ ax.set_xticks(ticks)
+ ax.set_xticklabels([str(value) for value in ticks], fontsize=9.2, color=MUTED)
+ ax.set_yticks(positions)
+ ax.set_yticklabels([row["label"] for row in rows], fontsize=9.0, color=INK)
+ for tick, row in zip(ax.get_yticklabels(), rows, strict=True):
+ if row["group"] != "external":
+ tick.set_weight("bold" if row["kind"] == "tuned" else "normal")
+ ax.set_ylim(15.0, -1.1)
+ ax.tick_params(axis="x", length=0, pad=7)
+ ax.tick_params(axis="y", length=0, pad=7)
+ ax.set_axisbelow(True)
+ ax.grid(axis="x", color=RULE, linewidth=0.8)
+ for spine in ax.spines.values():
+ spine.set_visible(False)
+ ax.axvline(0, color=RULE, linewidth=1)
+ description_parts.append(
+ f"{title}: " + ", ".join(
+ f"{row['label']} {row['accuracy'] * 100:.2f}%" for row in rows
+ )
+ )
+
+ fig.text(
+ 0.035,
+ 0.044,
+ "JevJudge full is unsupported for this text-only readout. Choice temperature calibration is omitted here because it cannot change accuracy.",
+ fontsize=10.1,
+ color=MUTED,
+ )
+ output = Path(output)
+ output.parent.mkdir(parents=True, exist_ok=True)
+ metadata = {
+ "Date": None,
+ "Title": "Training-free choice readout accuracy",
+ "Description": " ".join(description_parts),
+ }
+ fig.savefig(output, metadata=metadata)
+ plt.close(fig)
+ output.write_text(
+ "\n".join(line.rstrip() for line in output.read_text().splitlines()) + "\n",
+ encoding="utf-8",
+ )
+
+
+def main(argv: list[str] | None = None) -> None:
+ parser = argparse.ArgumentParser(description=__doc__)
+ parser.add_argument("--run-root", type=Path, required=True)
+ parser.add_argument("--external", type=Path, default=DEFAULT_EXTERNAL)
+ parser.add_argument("--model-catalog", type=Path, default=DEFAULT_MODEL_CATALOG)
+ parser.add_argument("--json-out", type=Path, default=DEFAULT_JSON_OUT)
+ parser.add_argument("--svg-out", type=Path, default=DEFAULT_SVG_OUT)
+ args = parser.parse_args(argv)
+
+ artifact = build_results(args.run_root, args.external, args.model_catalog)
+ write_results(args.json_out, artifact)
+ render_svg(artifact, args.svg_out)
+ print(f"Wrote {_display_path(args.json_out)} and {_display_path(args.svg_out)}")
+
+
+if __name__ == "__main__":
+ main()
diff --git a/scripts/build_external_report_appendix.py b/scripts/build_external_report_appendix.py
index 435746f..341bad4 100644
--- a/scripts/build_external_report_appendix.py
+++ b/scripts/build_external_report_appendix.py
@@ -1,10 +1,12 @@
#!/usr/bin/env python3
-"""Build and reproducibly merge the two-page external-evaluation appendix."""
+"""Build and reproducibly merge the external and choice-readout appendices."""
import argparse
from contextlib import contextmanager
from datetime import datetime, timezone
+import hashlib
import json
+import math
import os
from pathlib import Path
import tempfile
@@ -17,6 +19,14 @@
from matplotlib.backends.backend_pdf import PdfPages
from matplotlib.patches import FancyBboxPatch
from pypdf import PdfReader, PdfWriter
+from pypdf._cmap import get_encoding
+from pypdf.generic import (
+ ArrayObject,
+ ContentStream,
+ IndirectObject,
+ NumberObject,
+ TextStringObject,
+)
ROOT = Path(__file__).resolve().parents[1]
@@ -25,27 +35,78 @@
RULE = "#DCE4EC"
TEAL = "#278577"
PALE = "#F3F7FA"
-PDF_TIMESTAMP = datetime(2026, 10, 1, tzinfo=timezone.utc)
-PDF_DATE_LITERAL = "D:20261001000000Z"
+PDF_TIMESTAMP = datetime(2026, 10, 2, tzinfo=timezone.utc)
+PDF_DATE_LITERAL = "D:20261002000000Z"
+EXTERNAL_APPENDIX_PAGES = 2
+CHOICE_APPENDIX_PAGES = 2
+APPENDIX_PAGES = EXTERNAL_APPENDIX_PAGES + CHOICE_APPENDIX_PAGES
+REPORT_BASE_PAGES = 19
+MERGED_REPORT_PAGES = REPORT_BASE_PAGES + APPENDIX_PAGES
APPENDIX_METADATA = {
- "Title": "JevAny Technical Report โ External Evaluation Appendix",
+ "Title": "JevAny Technical Report โ Evaluation Appendices",
"Author": "SimpleJev",
- "Subject": "Typed Decisions and JevJudge-Public v0.3 evaluation",
+ "Subject": "External evaluation and training-free choice-token readout",
"Creator": "scripts/build_external_report_appendix.py",
"Producer": "Matplotlib PDF backend",
"CreationDate": PDF_TIMESTAMP,
"ModDate": PDF_TIMESTAMP,
}
MERGED_METADATA = {
- "/Title": "JevAny: Toward General Decision Intelligence โ with Agent-Harness and External-Evaluation Appendices",
+ "/Title": "JevAny: Toward General Decision Intelligence โ merged technical report",
"/Author": "SimpleJev",
- "/Subject": "JevAny model, bounded decision harness, and external decision-suite evaluation",
+ "/Subject": (
+ "JevAny model, bounded decision harness, external evaluation, and "
+ "training-free choice-token readout"
+ ),
"/Creator": "scripts/build_external_report_appendix.py",
"/Producer": "pypdf",
"/CreationDate": PDF_DATE_LITERAL,
"/ModDate": PDF_DATE_LITERAL,
}
+# The source report predates the final 27B release checkpoint. Synchronize the
+# handful of release-headline fields while merging so the one published PDF
+# does not call two different checkpoints "current". The original layout,
+# resources, links, and every unrelated result remain untouched.
+CURRENT_RELEASE_TEXT = {
+ 1: (("85.76", "86.04", 2), ("90.48", "90.04", 2)),
+ 4: (
+ ("22,160", "44,319", 1),
+ ("85.76", "86.04", 1),
+ ("90.48", "90.04", 1),
+ ("0.392", "0.388", 1),
+ ("0.200", "0.195", 1),
+ ("0.030", "0.026", 1),
+ ),
+ 7: (("85.76", "86.04", 1),),
+ 8: (
+ ("22,160", "44,319", 1),
+ ("18.83", "39.43", 1),
+ ("602.7", "1,261.7", 1),
+ ("1,423", "2,082", 1),
+ ),
+}
+
+CURRENT_COMPUTE_PROVENANCE = (
+ (
+ "onds;itsvalueusestherecordedrunstartandcheckpointtimestamp."
+ "Direct-token,Qwen27B,and",
+ "onds; its value uses the recorded run start and checkpoint timestamp. "
+ "Direct-token and Muse use",
+ ),
+ (
+ "Museusecheckpoint-nativeelapsedtime;Qwen4Bpointerusesterminaltrainertelemetrybecause",
+ "checkpoint-native elapsed time; Qwen 4B pointer uses terminal trainer telemetry. "
+ "Qwen 27B uses",
+ ),
+ (
+ "thereleasedcheckpointisthecompletedstep13,850run."
+ "Peakwithin-runparallelismwas40GPUs.",
+ "estimated cumulative seconds-per-record timing. "
+ "Peak within-run parallelism was 40 GPUs.",
+ ),
+)
+
@contextmanager
def atomic_destination(destination: Path):
@@ -67,15 +128,24 @@ def atomic_destination(destination: Path):
temporary.unlink(missing_ok=True)
-def header(fig, title: str, subtitle: str, page: int) -> None:
- fig.text(0.065, 0.955, "SimpleJev ยท TECHNICAL REPORT ยท APPENDIX L", color=TEAL,
+def header(
+ fig,
+ title: str,
+ subtitle: str,
+ page: int,
+ *,
+ appendix: str = "L",
+ section: str = "External evaluation",
+ pages: int = EXTERNAL_APPENDIX_PAGES,
+) -> None:
+ fig.text(0.065, 0.955, f"SimpleJev ยท TECHNICAL REPORT ยท APPENDIX {appendix}", color=TEAL,
fontsize=8.5, weight="bold")
fig.text(0.065, 0.905, title, color=INK, fontsize=22, weight="bold")
fig.text(0.065, 0.872, subtitle, color=MUTED, fontsize=9.5)
fig.lines.append(plt.Line2D([0.065, 0.935], [0.845, 0.845], transform=fig.transFigure,
color=RULE, linewidth=1))
fig.text(0.065, 0.035, "JevAny ยท Toward General Decision Intelligence", color=MUTED, fontsize=7.5)
- fig.text(0.935, 0.035, f"External evaluation ยท {page}/2", color=MUTED,
+ fig.text(0.935, 0.035, f"{section} ยท {page}/{pages}", color=MUTED,
fontsize=7.5, ha="right")
@@ -92,26 +162,48 @@ def overview_page(pdf: PdfPages, data: dict, chart: Path) -> None:
fig = plt.figure(figsize=(8.5, 11), facecolor="white")
header(fig, "External decision-suite evaluation",
"Accuracy on Typed Decisions ยท JevJudge-Public v0.3 full multimodal suite ยท text-only subset", 1)
- ax = fig.add_axes([0.055, 0.42, 0.89, 0.36])
- ax.imshow(mpimg.imread(chart))
- ax.axis("off")
+ chart_image = mpimg.imread(chart)
+ # The source asset is a wide, three-panel README graphic. Rendering it as
+ # one image makes its labels roughly four points on a portrait report page.
+ # Crop the panels (including their titles and axes) and use a 1 + 2 layout
+ # so every label is materially larger while the source pixels stay intact.
+ height, width = chart_image.shape[:2]
+ if width >= 3 and height >= 3:
+ y_start, y_stop = int(height * 0.15), int(height * 0.91)
+ panels = (
+ (chart_image[y_start:y_stop, int(width * 0.02):int(width * 0.43)],
+ [0.065, 0.500, 0.870, 0.300]),
+ (chart_image[y_start:y_stop, int(width * 0.42):int(width * 0.71)],
+ [0.065, 0.115, 0.420, 0.335]),
+ (chart_image[y_start:y_stop, int(width * 0.70):int(width * 0.995)],
+ [0.515, 0.115, 0.420, 0.335]),
+ )
+ for panel, bounds in panels:
+ ax = fig.add_axes(bounds)
+ ax.imshow(panel)
+ ax.axis("off")
+ else:
+ # Keep tiny fixture images usable in report-generation tests.
+ ax = fig.add_axes([0.065, 0.115, 0.870, 0.685])
+ ax.imshow(chart_image)
+ ax.axis("off")
full = {row["model"]: row for row in data["jevjudge_full"]["models"]}
text = {row["model"]: row for row in data["jevjudge_text"]["models"]}
typed = {row["model"]: row for row in data["typed_decisions"]["models"]}
- qwen = full["JevAny-Qwen3.8-27B"]
- jeff = full["Jeff-Qwen3.5-2B"]
- card(fig, 0.065, "Full-suite accuracy",
- f"Qwen3.8-27B: {qwen['accuracy'] * 100:.2f}%\n"
- f"Jeff-2B: {jeff['accuracy'] * 100:.2f}%\n"
- f"Gap: {(qwen['accuracy'] - jeff['accuracy']) * 100:.2f} points")
- card(fig, 0.365, "Typed accuracy",
- f"JevAny Qwen27: {typed['JevAny-Qwen3.8-27B']['accuracy'] * 100:.2f}%\n"
- f"Jev 1.13: {typed['Jev 1.13 (OpenRouter)']['accuracy'] * 100:.2f}%\n"
- "Typed result published ยท gap 0.10 pt")
- card(fig, 0.665, "Text-only accuracy",
- f"JevAny Qwen27: {text['JevAny-Qwen3.8-27B']['accuracy'] * 100:.2f}%\n"
- f"Jev 1.13: {text['Jev 1.13 (OpenRouter)']['accuracy'] * 100:.2f}%\n"
- f"Kev-27B: {text['Kev-27B']['accuracy'] * 100:.2f}%")
+ qwen_full = full["JevAny-Qwen3.8-27B"]
+ qwen_typed = typed["JevAny-Qwen3.8-27B"]
+ qwen_text = text["JevAny-Qwen3.8-27B"]
+ fig.text(
+ 0.065, 0.818,
+ "Same 13-model order ยท โ unsupported/no result ยท teal JevAny ยท gold Jev API ยท purple Kev ยท gray other open ยท hatch published",
+ color=MUTED, fontsize=7.8,
+ )
+ fig.text(
+ 0.065, 0.795,
+ f"Qwen3.8-27B topline ยท Typed {qwen_typed['accuracy'] * 100:.2f}% ยท "
+ f"JevJudge full {qwen_full['accuracy'] * 100:.2f}% ยท text-only {qwen_text['accuracy'] * 100:.2f}%",
+ color=TEAL, fontsize=8.2, weight="bold",
+ )
pdf.savefig(fig)
plt.close(fig)
@@ -173,47 +265,699 @@ def results_page(pdf: PdfPages, data: dict) -> None:
plt.close(fig)
-def build_appendix(data_path: Path, chart_path: Path, output: Path) -> None:
- data = json.loads(data_path.read_text())
+CHOICE_RUN_SPECS = {
+ "base4": ("Qwen3.5-4B", "frozen_base", None),
+ "pointer4": ("Qwen3.5-4B", "jevany_sft", "pointer"),
+ "direct4": ("Qwen3.5-4B", "jevany_sft", "direct-token"),
+ "base27": ("Qwen3.8-27B", "frozen_base", None),
+ "pointer27": ("Qwen3.8-27B", "jevany_sft", "pointer"),
+}
+CHOICE_DATASET_SPECS = (
+ ("jevbench_public", "JevBench\npublic", 231, 231, 231),
+ ("transfer_test", "Transfer-v9\ntest", 1_264, 1_264, 1_046),
+ ("typed_test", "Typed\ntest", 400, 2_000, 2_000),
+ ("jevjudge_text", "JevJudge\ntext", 724, 724, 724),
+)
+
+
+def _read_json_object(path: Path, name: str) -> dict:
+ try:
+ value = json.loads(path.read_text(encoding="utf-8"))
+ except FileNotFoundError as error:
+ raise ValueError(f"missing {name}: {path}") from error
+ except json.JSONDecodeError as error:
+ raise ValueError(f"invalid {name}: {path}: {error}") from error
+ if not isinstance(value, dict):
+ raise ValueError(f"{name} must be a JSON object: {path}")
+ return value
+
+
+def _choice_metric(row: dict, group: str, config: str, n: int, context: str) -> dict:
+ try:
+ metric = row[group][config]
+ except (KeyError, TypeError) as error:
+ raise ValueError(f"{context}: missing {group}/{config}") from error
+ if not isinstance(metric, dict) or metric.get("n") != n:
+ raise ValueError(f"{context}/{group}/{config}: expected n={n}")
+ correct = metric.get("correct")
+ accuracy = metric.get("accuracy")
+ if isinstance(correct, bool) or not isinstance(correct, int) or not 0 <= correct <= n:
+ raise ValueError(f"{context}/{group}/{config}: invalid correct count")
+ if (
+ isinstance(accuracy, bool)
+ or not isinstance(accuracy, (int, float))
+ or not math.isfinite(accuracy)
+ or not math.isclose(float(accuracy), correct / n, rel_tol=0, abs_tol=1e-12)
+ ):
+ raise ValueError(f"{context}/{group}/{config}: accuracy does not equal correct / n")
+ return metric
+
+
+def load_choice_artifact(path: Path) -> dict:
+ """Load the validated v2 matrix consumed by Appendix M."""
+
+ artifact = _read_json_object(path, "choice-readout artifact")
+ if artifact.get("schema_version") != 2 or artifact.get("artifact") != "choice-readout-v2":
+ raise ValueError("choice-readout artifact must use schema_version=2")
+ method = artifact.get("method")
+ if not isinstance(method, dict) or method.get("option_ids") != (
+ "A-Z followed by a-z; at most 52 options."
+ ):
+ raise ValueError("choice-readout artifact does not declare the v2 52-ID method")
+
+ source_rows = artifact.get("source_runs")
+ if not isinstance(source_rows, list):
+ raise ValueError("choice-readout artifact has no source_runs list")
+ source_by_key = {
+ row.get("key"): row for row in source_rows if isinstance(row, dict)
+ }
+ if set(source_by_key) != set(CHOICE_RUN_SPECS) or len(source_rows) != len(source_by_key):
+ raise ValueError("choice-readout artifact must contain the five required source runs")
+ for key, (family, weights, native_readout) in CHOICE_RUN_SPECS.items():
+ row = source_by_key[key]
+ if (row.get("family"), row.get("weights"), row.get("native_readout")) != (
+ family, weights, native_readout
+ ):
+ raise ValueError(f"choice-readout source run {key!r} has inconsistent identity")
+ selected = row.get("selected")
+ if not isinstance(selected, dict):
+ raise ValueError(f"choice-readout source run {key!r} has no selected settings")
+ for field in ("native_weight", "choice_temperature", "blend_temperature"):
+ value = selected.get(field)
+ if (
+ isinstance(value, bool)
+ or not isinstance(value, (int, float))
+ or not math.isfinite(value)
+ ):
+ raise ValueError(f"choice-readout source run {key!r} has invalid {field}")
+ if not 0 <= selected["native_weight"] <= 1:
+ raise ValueError(f"choice-readout source run {key!r} has invalid native_weight")
+ if selected["choice_temperature"] <= 0 or selected["blend_temperature"] <= 0:
+ raise ValueError(f"choice-readout source run {key!r} has invalid temperature")
+ if weights == "frozen_base" and selected["native_weight"] != 0:
+ raise ValueError(f"frozen-base run {key!r} selected a native weight")
+ ensemble_protocol = row.get("ensemble_protocol")
+ if not isinstance(ensemble_protocol, dict) or (
+ ensemble_protocol.get("selection_rows") != 1_046
+ ):
+ raise ValueError(f"choice-readout source run {key!r} has invalid selection rows")
+
+ datasets = artifact.get("datasets")
+ if not isinstance(datasets, dict):
+ raise ValueError("choice-readout artifact has no datasets object")
+ for dataset, _label, records, questions, n in CHOICE_DATASET_SPECS:
+ panel = datasets.get(dataset)
+ if not isinstance(panel, dict) or (
+ panel.get("records"), panel.get("questions"), panel.get("headline_n")
+ ) != (records, questions, n):
+ raise ValueError(
+ f"choice-readout dataset {dataset!r} must declare "
+ f"records/questions/headline_n={records}/{questions}/{n}"
+ )
+ rows = panel.get("runs")
+ if not isinstance(rows, list):
+ raise ValueError(f"choice-readout dataset {dataset!r} has no runs list")
+ by_key = {row.get("run"): row for row in rows if isinstance(row, dict)}
+ if set(by_key) != set(CHOICE_RUN_SPECS) or len(rows) != len(by_key):
+ raise ValueError(f"choice-readout dataset {dataset!r} must contain five runs")
+ for key, (family, weights, native_readout) in CHOICE_RUN_SPECS.items():
+ row = by_key[key]
+ if (row.get("family"), row.get("weights"), row.get("native_readout")) != (
+ family, weights, native_readout
+ ):
+ raise ValueError(f"{dataset}/{key}: inconsistent run identity")
+ _choice_metric(row, "zero_shot", "choice_t1", n, f"{dataset}/{key}")
+ if native_readout is not None:
+ _choice_metric(row, "zero_shot", "native_shipped", n, f"{dataset}/{key}")
+ _choice_metric(
+ row,
+ "transfer_dev_tuned",
+ "transfer_dev_tuned_blend",
+ n,
+ f"{dataset}/{key}",
+ )
+
+ full = datasets.get("jevjudge_full")
+ if not isinstance(full, dict) or (
+ full.get("records"), full.get("status"), full.get("accuracy")
+ ) != (3_220, "unsupported", None):
+ raise ValueError("choice-token JevJudge full must be explicitly unsupported")
+ return artifact
+
+
+def _rounded_box(fig, x, y, width, height, *, facecolor=PALE, edgecolor=RULE):
+ patch = FancyBboxPatch(
+ (x, y), width, height, transform=fig.transFigure,
+ boxstyle="round,pad=0.010,rounding_size=0.010",
+ facecolor=facecolor, edgecolor=edgecolor, linewidth=0.8,
+ )
+ fig.patches.append(patch)
+
+
+def choice_method_page(pdf: PdfPages, data: dict) -> None:
+ fig = plt.figure(figsize=(8.5, 11), facecolor="white")
+ header(
+ fig,
+ "Training-free choice-token readout",
+ "A constrained readout, probability calibration, and stacking are different operations",
+ 1,
+ appendix="M",
+ section="Choice-token readout",
+ pages=CHOICE_APPENDIX_PAGES,
+ )
+ fig.text(
+ 0.065, 0.815,
+ "The frozen language model scores one exact option ID at the answer position. "
+ "No explanation is generated and no additional training is required for this readout.",
+ color=INK, fontsize=10.0, wrap=True, linespacing=1.45,
+ )
+
+ pipeline = (
+ ("1", "Ordered options", "Preserve candidate order"),
+ ("2", "52 exact IDs", "AโZ followed by aโz"),
+ ("3", "One-token projection", "Read next-token mass"),
+ ("4", "Decision distribution", "Renormalize over IDs"),
+ )
+ for index, (number, title, body) in enumerate(pipeline):
+ x = 0.065 + index * 0.222
+ _rounded_box(fig, x, 0.685, 0.185, 0.095, facecolor="#F5F9F8")
+ fig.text(x + 0.014, 0.750, number, color=TEAL, fontsize=9, weight="bold")
+ fig.text(x + 0.042, 0.750, title, color=INK, fontsize=9.2, weight="bold")
+ fig.text(x + 0.014, 0.712, body, color=MUTED, fontsize=7.8)
+ if index < len(pipeline) - 1:
+ fig.text(x + 0.198, 0.730, "โ", color=TEAL, fontsize=14, weight="bold")
+
+ semantic_cards = (
+ (
+ "Choice-token ยท T=1",
+ "Training-free constrained\nprojection. This is a readout,\n"
+ "not temperature calibration;\nit can disagree with native.",
+ ),
+ (
+ "Temperature calibration",
+ "Fit scalar T on separate\ndevelopment rows. T changes\n"
+ "probabilities; it cannot change\nargmax or accuracy.",
+ ),
+ (
+ "Native + choice stack",
+ "Log-linear pooling can change\nrankings. Select native weight\n"
+ "and additional temperature on\nTransfer-v9 development only.",
+ ),
+ )
+ for index, (title, body) in enumerate(semantic_cards):
+ x = 0.065 + index * 0.299
+ _rounded_box(fig, x, 0.475, 0.270, 0.150)
+ fig.text(x + 0.016, 0.590, title, color=INK, fontsize=9.6, weight="bold")
+ fig.text(
+ x + 0.016, 0.558, body, color=MUTED, fontsize=7.8,
+ va="top", linespacing=1.35,
+ )
+
+ sources = {row["key"]: row for row in data["source_runs"]}
+ table = (
+ ("Frozen base", "โ", "Choice-token", f"{sources['base4']['family']}, {sources['base27']['family']}"),
+ ("JevAny SFT", "Pointer", "Choice-token", f"{sources['pointer4']['family']}, {sources['pointer27']['family']}"),
+ ("JevAny SFT", "Direct-Token", "Choice-token", sources["direct4"]["family"]),
+ )
+ fig.text(0.065, 0.425, "Five-run comparison", color=INK, fontsize=12, weight="bold")
+ columns = (0.065, 0.245, 0.420, 0.610)
+ for x, label in zip(columns, ("Weights", "Native reference", "Training-free path", "Backbone family")):
+ fig.text(x, 0.390, label, color=MUTED, fontsize=8.0, weight="bold")
+ fig.lines.append(plt.Line2D([0.065, 0.935], [0.377, 0.377], transform=fig.transFigure,
+ color=RULE, linewidth=1))
+ for index, values in enumerate(table):
+ y = 0.346 - index * 0.052
+ for x, value in zip(columns, values):
+ fig.text(x, y, value, color=INK, fontsize=8.4)
+ fig.lines.append(plt.Line2D([0.065, 0.935], [y - 0.019, y - 0.019],
+ transform=fig.transFigure, color="#EEF2F6", linewidth=0.7))
+
+ _rounded_box(
+ fig, 0.065, 0.123, 0.870, 0.075,
+ facecolor="#FFF8E8", edgecolor="#E7C46A",
+ )
+ fig.text(0.083, 0.176, "Release sync and protocol boundary", color="#8A5A00", fontsize=9.5, weight="bold")
+ fig.text(
+ 0.083, 0.151,
+ "The merged report and README identify the current 27B release checkpoint: step 44,319.",
+ color=INK, fontsize=8.0,
+ )
+ fig.text(
+ 0.083, 0.132,
+ "Appendix M native rows are matched prompt-v2/runtime reruns; Appendix L's earlier external-study row can differ by one decision.",
+ color=INK, fontsize=8.0,
+ )
+ _rounded_box(fig, 0.065, 0.061, 0.870, 0.047, facecolor="#F7FAFC")
+ fig.text(0.083, 0.091, "Scope", color=INK, fontsize=8.7, weight="bold")
+ fig.text(
+ 0.137, 0.091,
+ "Text only ยท up to 52 IDs ยท JevJudge text 724 supported ยท full multimodal unsupported ยท agent tasks unchanged",
+ color=MUTED, fontsize=7.6,
+ )
+ pdf.savefig(fig)
+ plt.close(fig)
+
+
+def _dataset_run(data: dict, dataset: str, run: str) -> dict:
+ matches = [row for row in data["datasets"][dataset]["runs"] if row["run"] == run]
+ if len(matches) != 1:
+ raise ValueError(f"{dataset}: expected one {run} row")
+ return matches[0]
+
+
+def _choice_matrix_rows(data: dict) -> list[dict]:
+ sources = {row["key"]: row for row in data["source_runs"]}
+ definitions = (
+ ("Frozen Qwen3.5-4B", "Choice T=1", "base4", "zero_shot", "choice_t1", "โ"),
+ ("JevAny 4B Pointer", "Native", "pointer4", "zero_shot", "native_shipped", "shipped"),
+ ("", "Choice T=1", "pointer4", "zero_shot", "choice_t1", "โ"),
+ ("", "Tuned stack", "pointer4", "transfer_dev_tuned", "transfer_dev_tuned_blend", None),
+ ("JevAny 4B Direct-Token", "Native", "direct4", "zero_shot", "native_shipped", "shipped"),
+ ("", "Choice T=1", "direct4", "zero_shot", "choice_t1", "โ"),
+ ("", "Tuned stack", "direct4", "transfer_dev_tuned", "transfer_dev_tuned_blend", None),
+ ("Frozen Qwen3.8-27B", "Choice T=1", "base27", "zero_shot", "choice_t1", "โ"),
+ ("JevAny 27B Pointer", "Native", "pointer27", "zero_shot", "native_shipped", "shipped"),
+ ("", "Choice T=1", "pointer27", "zero_shot", "choice_t1", "โ"),
+ ("", "Tuned stack", "pointer27", "transfer_dev_tuned", "transfer_dev_tuned_blend", None),
+ )
+ rows = []
+ for model, readout, run, group, config, weight in definitions:
+ if weight is None:
+ weight = f"{sources[run]['selected']['native_weight']:.2f}"
+ metrics = []
+ for dataset, _label, _records, _questions, n in CHOICE_DATASET_SPECS:
+ metric = _choice_metric(
+ _dataset_run(data, dataset, run), group, config, n, f"{dataset}/{run}"
+ )
+ metrics.append(metric)
+ rows.append({
+ "model": model,
+ "readout": readout,
+ "run": run,
+ "weight": weight,
+ "metrics": metrics,
+ "tuned": readout == "Tuned stack",
+ "native": readout == "Native",
+ })
+ return rows
+
+
+def choice_results_page(pdf: PdfPages, data: dict) -> None:
+ fig = plt.figure(figsize=(8.5, 11), facecolor="white")
+ header(
+ fig,
+ "Choice-token accuracy matrix",
+ "Matched runs ยท current 27B release step 44,319 ยท Transfer-dev settings frozen before evaluation",
+ 2,
+ appendix="M",
+ section="Choice-token readout",
+ pages=CHOICE_APPENDIX_PAGES,
+ )
+ rows = _choice_matrix_rows(data)
+ ax = fig.add_axes([0.055, 0.285, 0.89, 0.530])
+ ax.axis("off")
+ x_model, x_readout, x_weight = 0.00, 0.315, 0.495
+ x_metrics = (0.635, 0.755, 0.875, 0.995)
+ ax.text(x_model, 1.035, "Model", transform=ax.transAxes, fontsize=8.0,
+ color=MUTED, weight="bold")
+ ax.text(x_readout, 1.035, "Readout", transform=ax.transAxes, fontsize=8.0,
+ color=MUTED, weight="bold")
+ ax.text(x_weight, 1.035, "w_native", transform=ax.transAxes, fontsize=8.0,
+ color=MUTED, weight="bold", ha="right")
+ for x, (_dataset, label, _records, _questions, n) in zip(
+ x_metrics, CHOICE_DATASET_SPECS, strict=True
+ ):
+ ax.text(x, 1.035, f"{label}\n(n={n:,})", transform=ax.transAxes,
+ fontsize=7.7, color=MUTED, weight="bold", ha="right", va="bottom")
+ ax.plot([0, 1], [0.995, 0.995], transform=ax.transAxes, color=RULE, linewidth=1)
+
+ group_spans = ((0, 0), (1, 3), (4, 6), (7, 7), (8, 10))
+ step, first_y = 0.082, 0.930
+ for group_index, (start, end) in enumerate(group_spans):
+ top = first_y - start * step + 0.034
+ bottom = first_y - end * step - 0.034
+ if group_index % 2:
+ patch = FancyBboxPatch(
+ (-0.012, bottom), 1.022, top - bottom,
+ transform=ax.transAxes, boxstyle="round,pad=0.004,rounding_size=0.005",
+ facecolor="#F7FAFC", edgecolor="none", zorder=0,
+ )
+ ax.add_patch(patch)
+
+ maxima = [max(row["metrics"][index]["accuracy"] for row in rows) for index in range(4)]
+ for row_index, row in enumerate(rows):
+ y = first_y - row_index * step
+ color = TEAL if row["tuned"] else INK
+ ax.text(x_model, y, row["model"], transform=ax.transAxes, fontsize=8.0,
+ color=INK, va="center", weight="bold" if row["model"] else "normal")
+ ax.text(x_readout, y, row["readout"], transform=ax.transAxes, fontsize=8.0,
+ color=color, va="center", weight="bold" if row["tuned"] else "normal")
+ ax.text(x_weight, y, row["weight"], transform=ax.transAxes, fontsize=8.0,
+ color=MUTED, va="center", ha="right")
+ for metric_index, (x, metric) in enumerate(zip(x_metrics, row["metrics"], strict=True)):
+ best = math.isclose(metric["accuracy"], maxima[metric_index], rel_tol=0, abs_tol=1e-12)
+ ax.text(
+ x, y, f"{metric['accuracy'] * 100:.2f}", transform=ax.transAxes,
+ fontsize=8.1, color=TEAL if best else color, va="center", ha="right",
+ weight="bold" if best or row["tuned"] else "normal",
+ )
+ ax.plot([0, 1], [y - 0.040, y - 0.040], transform=ax.transAxes,
+ color="#EEF2F6", linewidth=0.6, zorder=1)
+
+ sources = {row["key"]: row for row in data["source_runs"]}
+ tuned = []
+ for key, label in (("pointer4", "4B Pointer"), ("direct4", "4B Direct-Token"),
+ ("pointer27", "27B Pointer")):
+ selected = sources[key]["selected"]
+ tuned.append(
+ f"{label}: w={selected['native_weight']:.2f}, T={selected['blend_temperature']:.2f}"
+ )
+ fig.text(0.065, 0.235, "Transfer-dev selection", color=INK, fontsize=10.5, weight="bold")
+ fig.text(0.065, 0.205, " ยท ".join(tuned), color=TEAL, fontsize=8.4, weight="bold")
+ fig.text(
+ 0.065, 0.169,
+ "Weights and additional temperatures use only 1,046 clean/knowable Transfer-v9 development rows.\n"
+ "Scalar T changes probabilities, not the accuracy values above.",
+ color=MUTED, fontsize=8.2, linespacing=1.35,
+ )
+ fig.text(
+ 0.065, 0.125,
+ "Native = matched native rerun at the checkpoint's shipped inference temperature.\n"
+ "JevBench is a public development diagnostic; Transfer, Typed, and JevJudge text are held-out panels.",
+ color=MUTED, fontsize=8.2, linespacing=1.35,
+ )
+ fig.text(
+ 0.065, 0.082,
+ "Choice-token is text-only: JevJudge text 724 is reported; full JevJudge 3,220 is unsupported.\n"
+ "Source and complete calibration metrics: results/choice-readout-v2.json.",
+ color=MUTED, fontsize=8.2, linespacing=1.35,
+ )
+ pdf.savefig(fig)
+ plt.close(fig)
+
+
+def build_appendix(
+ data_path: Path,
+ chart_path: Path,
+ choice_data_path: Path,
+ output: Path,
+) -> None:
+ data = _read_json_object(data_path, "external-evaluation artifact")
+ choice_data = load_choice_artifact(choice_data_path)
with atomic_destination(output) as temporary:
with PdfPages(temporary, metadata=APPENDIX_METADATA) as pdf:
overview_page(pdf, data, chart_path)
results_page(pdf, data)
+ choice_method_page(pdf, choice_data)
+ choice_results_page(pdf, choice_data)
appendix = PdfReader(temporary)
- if len(appendix.pages) != 2:
- raise RuntimeError(f"expected a two-page appendix, got {len(appendix.pages)} pages")
+ if len(appendix.pages) != APPENDIX_PAGES:
+ raise RuntimeError(
+ f"expected a {APPENDIX_PAGES}-page appendix, got {len(appendix.pages)} pages"
+ )
+
+
+def _canonical_pdf_object(value, active: set[int] | None = None):
+ """Resolve PDF references into a stable, object-number-independent value."""
+
+ if isinstance(value, IndirectObject):
+ return _canonical_pdf_object(value.get_object(), active)
+ if active is None:
+ active = set()
+ if isinstance(value, dict):
+ marker = id(value)
+ if marker in active:
+ return ("cycle",)
+ active.add(marker)
+ try:
+ items = tuple(sorted(
+ (
+ str(key),
+ _canonical_pdf_object(item, active),
+ )
+ for key, item in value.items()
+ if str(key) != "/Length"
+ ))
+ stream = None
+ if hasattr(value, "get_data"):
+ stream = hashlib.sha256(value.get_data()).hexdigest()
+ return ("dict", items, stream)
+ finally:
+ active.remove(marker)
+ if isinstance(value, (list, tuple)):
+ return tuple(_canonical_pdf_object(item, active) for item in value)
+ if isinstance(value, bytes):
+ return ("bytes", hashlib.sha256(value).hexdigest())
+ if value is None or isinstance(value, (bool, int, float, str)):
+ return value
+ return str(value)
def page_invariants(page) -> tuple:
"""Return page properties that must survive the merge unchanged."""
contents = page.get_contents()
content_bytes = contents.get_data() if contents is not None else b""
+ annotations = []
+ for reference in page.get("/Annots", []):
+ annotation = reference.get_object()
+ # /P is a back-reference to the owning page and would make the
+ # otherwise stable annotation target depend on PDF object numbers.
+ annotations.append(_canonical_pdf_object({
+ key: value for key, value in annotation.items() if str(key) != "/P"
+ }))
return (
tuple(float(value) for value in page.mediabox),
tuple(float(value) for value in page.cropbox),
page.rotation,
content_bytes,
page.extract_text() or "",
- len(page.get("/Annots", [])),
+ _canonical_pdf_object(page.get("/Resources")),
+ tuple(annotations),
)
+def _font_text_maps(page) -> dict[str, tuple[dict[str, str], dict[str, str]]]:
+ maps = {}
+ fonts = page["/Resources"]["/Font"]
+ for name, reference in fonts.items():
+ _encoding, character_map = get_encoding(reference.get_object())
+ reverse = {
+ unicode_text: glyph
+ for glyph, unicode_text in character_map.items()
+ if isinstance(glyph, str)
+ and isinstance(unicode_text, str)
+ and len(unicode_text) == 1
+ }
+ maps[str(name)] = (character_map, reverse)
+ return maps
+
+
+def _decode_text_object(value, character_map: dict[str, str]) -> str:
+ return "".join(character_map.get(glyph, glyph) for glyph in str(value))
+
+
+def _encode_text_object(
+ text: str,
+ reverse_map: dict[str, str],
+ prototype,
+) -> TextStringObject:
+ try:
+ glyphs = "".join(reverse_map[character] for character in text)
+ except KeyError as error:
+ raise RuntimeError(f"release-sync font cannot encode {error.args[0]!r}") from error
+ original_length = len(prototype.original_bytes)
+ glyph_count = len(str(prototype))
+ if glyph_count == 0 or original_length % glyph_count:
+ raise RuntimeError("release-sync text object has an unsupported encoding width")
+ byte_width = original_length // glyph_count
+ if byte_width not in {1, 2}:
+ raise RuntimeError(f"release-sync text uses unsupported {byte_width}-byte glyphs")
+ raw = b"".join(ord(glyph).to_bytes(byte_width, "big") for glyph in glyphs)
+ result = TextStringObject(glyphs)
+ result._original_bytes = raw
+ return result
+
+
+def _word_array(text: str, reverse_map: dict[str, str], prototype) -> ArrayObject:
+ words = text.split()
+ result = ArrayObject()
+ for index, word in enumerate(words):
+ if index:
+ result.append(NumberObject(-300))
+ result.append(_encode_text_object(word, reverse_map, prototype))
+ return result
+
+
+def synchronize_current_release(page, page_number: int) -> bool:
+ """Replace the superseded 27B headline fields in the retained report core."""
+
+ replacements = CURRENT_RELEASE_TEXT.get(page_number)
+ if replacements is None:
+ return False
+ extracted = page.extract_text() or ""
+ legacy_counts = {old: extracted.count(old) for old, _new, _count in replacements}
+ if not any(legacy_counts.values()):
+ return False
+ for old, _new, expected_count in replacements:
+ if legacy_counts[old] != expected_count:
+ raise RuntimeError(
+ f"base report page {page_number}: expected {expected_count} occurrences "
+ f"of legacy value {old!r}, found {legacy_counts[old]}"
+ )
+
+ text_maps = _font_text_maps(page)
+ content = ContentStream(page.get_contents(), page.indirect_reference.pdf)
+ active_font = None
+ replaced_counts = {old: 0 for old, _new, _count in replacements}
+ provenance_replaced = 0
+ for operands, operator in content.operations:
+ if operator == b"Tf":
+ active_font = str(operands[0])
+ continue
+ if operator not in {b"Tj", b"TJ"} or active_font is None:
+ continue
+ character_map, reverse_map = text_maps[active_font]
+ values = operands[0] if operator == b"TJ" else ArrayObject([operands[0]])
+ decoded_line = "".join(
+ _decode_text_object(value, character_map)
+ for value in values
+ if hasattr(value, "original_bytes")
+ )
+
+ if page_number == 8:
+ provenance = next(
+ (replacement for legacy, replacement in CURRENT_COMPUTE_PROVENANCE
+ if decoded_line == legacy),
+ None,
+ )
+ if provenance is not None:
+ prototype = next(
+ value for value in values if hasattr(value, "original_bytes")
+ )
+ operands[0] = _word_array(provenance, reverse_map, prototype)
+ provenance_replaced += 1
+ continue
+
+ for index, value in enumerate(values):
+ if not hasattr(value, "original_bytes"):
+ continue
+ decoded = _decode_text_object(value, character_map)
+ updated = decoded
+ for old, new, _expected_count in replacements:
+ count = updated.count(old)
+ if count:
+ updated = updated.replace(old, new)
+ replaced_counts[old] += count
+ if updated != decoded:
+ values[index] = _encode_text_object(updated, reverse_map, value)
+ if operator == b"Tj":
+ operands[0] = values[0]
+
+ for old, _new, expected_count in replacements:
+ if replaced_counts[old] != expected_count:
+ raise RuntimeError(
+ f"base report page {page_number}: replaced {replaced_counts[old]} of "
+ f"{expected_count} expected {old!r} values"
+ )
+ if page_number == 8 and provenance_replaced != len(CURRENT_COMPUTE_PROVENANCE):
+ raise RuntimeError(
+ "base report page 8: could not synchronize the 27B compute provenance"
+ )
+ # In-place list edits do not invalidate ContentStream's raw-byte cache.
+ content.operations = content.operations
+ synchronized_parts = []
+ active_font = None
+ for operands, operator in content.operations:
+ if operator == b"Tf":
+ active_font = str(operands[0])
+ continue
+ if operator not in {b"Tj", b"TJ"} or active_font is None:
+ continue
+ character_map, _reverse_map = text_maps[active_font]
+ values = operands[0] if operator == b"TJ" else [operands[0]]
+ synchronized_parts.extend(
+ _decode_text_object(value, character_map)
+ for value in values
+ if hasattr(value, "original_bytes")
+ )
+ synchronized = "".join(synchronized_parts)
+ for old, new, expected_count in replacements:
+ if old in synchronized or synchronized.count(new) < expected_count:
+ raise RuntimeError(
+ f"base report page {page_number}: failed to synchronize {old!r} to {new!r}; "
+ f"old={synchronized.count(old)}, new={synchronized.count(new)}"
+ )
+ page.replace_contents(content)
+ return True
+
+
+def synchronize_report_boundary(page, page_number: int) -> bool:
+ """Update the old appendix claim that the merged core is byte-unchanged."""
+
+ if page_number != 10:
+ return False
+ extracted = page.extract_text() or ""
+ if "The original PDF is pre-" not in extracted:
+ return False
+ replacements = {
+ "original ": "merged ",
+ "pre-": "release-",
+ "serv": "synch",
+ "ed unchanged.": "ronized.",
+ }
+ counts = {old: 0 for old in replacements}
+ content = ContentStream(page.get_contents(), page.indirect_reference.pdf)
+ for operands, operator in content.operations:
+ if operator not in {b"Tj", b"TJ"}:
+ continue
+ values = operands[0] if operator == b"TJ" else ArrayObject([operands[0]])
+ for index, value in enumerate(values):
+ if not isinstance(value, TextStringObject):
+ continue
+ updated = str(value)
+ for old, new in replacements.items():
+ count = updated.count(old)
+ if count:
+ updated = updated.replace(old, new)
+ counts[old] += count
+ if updated != str(value):
+ values[index] = TextStringObject(updated)
+ if operator == b"Tj":
+ operands[0] = values[0]
+ if any(count != 1 for count in counts.values()):
+ raise RuntimeError(
+ f"base report page 10: could not synchronize appendix boundary: {counts}"
+ )
+ content.operations = content.operations
+ page.replace_contents(content)
+ return True
+
+
def merge_report(base_report: Path, appendix_report: Path, merged_output: Path, base_pages: int) -> None:
- if base_pages < 1:
- raise ValueError("--base-pages must be positive")
+ if base_pages != REPORT_BASE_PAGES:
+ raise ValueError(f"--base-pages must be exactly {REPORT_BASE_PAGES}")
base = PdfReader(base_report)
appendix = PdfReader(appendix_report)
if len(base.pages) < base_pages:
raise ValueError(
f"base report has {len(base.pages)} pages, fewer than --base-pages={base_pages}"
)
- if len(appendix.pages) != 2:
- raise ValueError(f"appendix must have exactly two pages, got {len(appendix.pages)}")
+ if len(appendix.pages) != APPENDIX_PAGES:
+ raise ValueError(
+ f"appendix must have exactly {APPENDIX_PAGES} pages, got {len(appendix.pages)}"
+ )
writer = PdfWriter()
writer.pdf_header = "%PDF-1.7"
+ synchronized_pages = set()
+ expected_base_invariants = {}
for index in range(base_pages):
writer.add_page(base.pages[index])
+ changed = synchronize_current_release(writer.pages[-1], index + 1)
+ changed = synchronize_report_boundary(writer.pages[-1], index + 1) or changed
+ if changed:
+ synchronized_pages.add(index)
+ expected_base_invariants[index] = page_invariants(writer.pages[-1])
for page in appendix.pages:
writer.add_page(page)
writer.add_metadata(MERGED_METADATA)
@@ -222,18 +966,54 @@ def merge_report(base_report: Path, appendix_report: Path, merged_output: Path,
with temporary.open("wb") as stream:
writer.write(stream)
merged = PdfReader(temporary)
- expected_pages = base_pages + len(appendix.pages)
+ expected_pages = MERGED_REPORT_PAGES
if len(merged.pages) != expected_pages:
raise RuntimeError(f"expected {expected_pages} merged pages, got {len(merged.pages)}")
for index in range(base_pages):
- if page_invariants(base.pages[index]) != page_invariants(merged.pages[index]):
+ expected = (
+ expected_base_invariants[index]
+ if index in synchronized_pages
+ else page_invariants(base.pages[index])
+ )
+ actual = page_invariants(merged.pages[index])
+ if index in synchronized_pages:
+ # A writer-attached page temporarily extracts its glyph IDs,
+ # while the serialized page resolves them through ToUnicode.
+ # Compare every invariant except that transient text view.
+ expected = expected[:4] + expected[5:]
+ actual = actual[:4] + actual[5:]
+ if expected != actual:
raise RuntimeError(f"base page {index + 1} changed during merge")
+ for page_number, replacements in CURRENT_RELEASE_TEXT.items():
+ page_text = merged.pages[page_number - 1].extract_text() or ""
+ is_release_page = (page_number - 1) in synchronized_pages or any(
+ page_text.count(new) >= expected_count
+ for _old, new, expected_count in replacements
+ )
+ if not is_release_page:
+ continue
+ for old, new, expected_count in replacements:
+ if old in page_text or page_text.count(new) < expected_count:
+ raise RuntimeError(
+ f"serialized base page {page_number} does not contain the current "
+ f"release value {new!r}"
+ )
+ boundary_text = merged.pages[9].extract_text() or ""
+ boundary_text = boundary_text.replace("-\n", "-")
+ if "The original PDF is pre-" in boundary_text or (
+ "The merged PDF is release-synchronized." not in boundary_text
+ and 9 in synchronized_pages
+ ):
+ raise RuntimeError("serialized base page 10 has a stale appendix boundary")
-def main() -> None:
+def main(argv: list[str] | None = None) -> None:
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument("--data", type=Path, default=ROOT / "results/external-zero-shot-v1.json")
parser.add_argument("--chart", type=Path, default=ROOT / "docs/external-zero-shot.png")
+ parser.add_argument(
+ "--choice-data", type=Path, default=ROOT / "results/choice-readout-v2.json"
+ )
parser.add_argument("--output", type=Path, required=True)
parser.add_argument(
"--base-report",
@@ -245,12 +1025,12 @@ def main() -> None:
type=Path,
help="merged report; may be the same path as --base-report",
)
- parser.add_argument("--base-pages", type=int, default=19)
- args = parser.parse_args()
+ parser.add_argument("--base-pages", type=int, default=REPORT_BASE_PAGES)
+ args = parser.parse_args(argv)
if (args.base_report is None) != (args.merged_output is None):
parser.error("--base-report and --merged-output must be provided together")
- if args.base_pages < 1:
- parser.error("--base-pages must be positive")
+ if args.base_pages != REPORT_BASE_PAGES:
+ parser.error(f"--base-pages must be exactly {REPORT_BASE_PAGES}")
if args.base_report is not None:
appendix_path = args.output.expanduser().absolute()
base_path = args.base_report.expanduser().absolute()
@@ -258,12 +1038,12 @@ def main() -> None:
if appendix_path in {base_path, merged_path}:
parser.error("--output must differ from --base-report and --merged-output")
- build_appendix(args.data, args.chart, args.output)
+ build_appendix(args.data, args.chart, args.choice_data, args.output)
print(f"Wrote reproducible appendix: {args.output}")
if args.base_report is not None:
merge_report(args.base_report, args.output, args.merged_output, args.base_pages)
print(
- f"Wrote reproducible {args.base_pages + 2}-page merged report: "
+ f"Wrote reproducible {MERGED_REPORT_PAGES}-page merged report: "
f"{args.merged_output}"
)
diff --git a/scripts/evaluate_choice_readout_matrix.py b/scripts/evaluate_choice_readout_matrix.py
new file mode 100644
index 0000000..fba7833
--- /dev/null
+++ b/scripts/evaluate_choice_readout_matrix.py
@@ -0,0 +1,447 @@
+#!/usr/bin/env python3
+"""Evaluate training-free choice readout and a checkpoint's native readout.
+
+The script loads one model once, then scores the same frozen model on JevBench,
+Transfer-v9, Typed Decisions and JevJudge's text-only subset. Checkpoint runs
+save both readout components so ensemble weights can be selected offline on a
+separate calibration split without repeating inference or looking at test
+labels. Frozen-base runs save the training-free choice component only.
+"""
+
+from __future__ import annotations
+
+import argparse
+import json
+import platform
+import statistics
+import subprocess
+import time
+from collections import defaultdict
+from pathlib import Path
+import math
+
+import numpy as np
+import pyarrow
+import pyarrow.parquet as pq
+import peft
+import safetensors
+import torch
+import transformers
+
+from jevany.api import question_keys
+from jevany.benchmark import prediction_rows, summarize
+from jevany.checkpoint import LoadOptions
+from jevany.data import api_request
+from jevany.letter_predictor import LetterReadoutPredictor
+from jevany.suite import (
+ ENCODING,
+ digest,
+ load_split,
+ read_json,
+ record_digest,
+ write_json,
+)
+
+
+def _typed_rows(dataset: Path, split: str) -> list[dict]:
+ path = dataset / "all" / f"{split}-00000-of-00001.parquet"
+ return pq.read_table(path).to_pylist()
+
+
+def _typed_record(row: dict) -> dict:
+ state = json.loads(row["state"])
+ questions = json.loads(row["questions"])
+ gold = json.loads(row["gold"])
+ converted = {}
+ for qid, question in questions.items():
+ answer = gold[qid]
+ value = dict(question)
+ keys = question_keys(value["type"], value.get("criteria"))
+ # Typed Decisions accuracy is ordered argmax agreement with the soft
+ # teacher distribution. The dataset's convenience `label` differs on
+ # tied rows, so derive the target exactly as the benchmark scorer does.
+ label_key = max(keys, key=lambda key: float(answer["probabilities"][key]))
+ if value["type"] == "choice":
+ value["label"] = label_key
+ elif value["type"] == "noul":
+ value["label"] = label_key == "true"
+ else:
+ value["label"] = int(label_key)
+ value["target"] = answer["probabilities"]
+ value["src"] = f"typed_decisions/{row['workflow']}/{qid}"
+ converted[qid] = value
+ return {
+ "state": state,
+ "questions": converted,
+ "_meta": {
+ "id": row["id"],
+ "group_id": row["id"],
+ "source": "LocalLLaMA/typed-decisions",
+ "variant": "clean",
+ "split": row["split"],
+ "workflow": row["workflow"],
+ },
+ }
+
+
+def _attach_soft_gold(rows: list[dict], record: dict) -> None:
+ questions = record["questions"]
+ for row in rows:
+ target = questions[row["question"]].get("target")
+ if target is not None:
+ row["gold"] = [float(target[key]) for key in row["keys"]]
+
+
+def _component_prediction(probabilities: dict) -> dict:
+ return {"probabilities": probabilities}
+
+
+def _typed_soft_metrics(rows: list[dict], bins: int = 15) -> dict:
+ totals = [0] * bins
+ correct = [0] * bins
+ confidence = [0.0] * bins
+ kl, brier = 0.0, 0.0
+ for row in rows:
+ p = [float(value) for value in row["p"]]
+ gold = [float(value) for value in row["gold"]]
+ p_total, gold_total = sum(p), sum(gold)
+ p = [value / p_total for value in p]
+ gold = [value / gold_total for value in gold]
+ kl += sum(g * math.log(max(g, 1e-12) / max(value, 1e-12))
+ for g, value in zip(gold, p, strict=True))
+ brier += sum((value - g) ** 2 for value, g in zip(p, gold, strict=True))
+ predicted = max(range(len(p)), key=p.__getitem__)
+ conf = max(p)
+ bucket = min(bins - 1, int(conf * bins))
+ totals[bucket] += 1
+ correct[bucket] += int(predicted == row["label"])
+ confidence[bucket] += conf
+ n = len(rows)
+ return {
+ "n": n,
+ "correct": sum(correct),
+ "accuracy": sum(correct) / n,
+ "kl_from_gold": kl / n,
+ "brier": brier / n,
+ "ece": sum(
+ totals[index] / n * abs(
+ correct[index] / totals[index] - confidence[index] / totals[index]
+ )
+ for index in range(bins) if totals[index]
+ ),
+ "mean_confidence": sum(confidence) / n,
+ }
+
+
+def evaluate(name: str, records: list[dict], predictor, destination: Path) -> dict:
+ destination.mkdir(parents=True, exist_ok=False)
+ component_rows: dict[str, list[dict]] = defaultdict(list)
+ latencies = []
+ efficient_long_context_records = 0
+ started = time.time()
+ with (destination / "predictions.jsonl").open("w", encoding=ENCODING) as stream:
+ for index, record in enumerate(records, 1):
+ result = predictor(record)
+ efficient_long_context_records += int(
+ result.get("efficient_long_context_attention_used", False)
+ )
+ components = result.get("component_probabilities") or {
+ "choice": result["probabilities"]
+ }
+ saved = {}
+ for component, probabilities in components.items():
+ if probabilities is None:
+ continue
+ rows = prediction_rows(record, _component_prediction(probabilities))
+ _attach_soft_gold(rows, record)
+ component_rows[component].extend(rows)
+ saved[component] = probabilities
+ stream.write(json.dumps({
+ "id": record["_meta"]["id"],
+ "request_sha256": record_digest(api_request(record)),
+ "component_probabilities": saved,
+ "latency_ms": result["latency_ms"],
+ "input_tokens": result["input_tokens"],
+ }, allow_nan=False) + "\n")
+ stream.flush()
+ latencies.append(float(result["latency_ms"]))
+ if index % 25 == 0:
+ print(f"{name}: {index}/{len(records)}", flush=True)
+ reports = {}
+ for component, rows in component_rows.items():
+ write_json(destination / f"{component}-rows.json", rows)
+ reports[component] = summarize(rows)
+ if name == "typed_test":
+ reports[component]["typed_soft_metrics"] = _typed_soft_metrics(rows)
+ report = {
+ "dataset": name,
+ "records": len(records),
+ "questions": sum(len(record["questions"]) for record in records),
+ "components": reports,
+ "components_executed": sorted(component_rows),
+ "efficient_long_context_records": efficient_long_context_records,
+ "latency_scope": (
+ "end-to-end predictor call; checkpoint calls execute both choice and native "
+ "components, so these values are not per-component efficiency measurements"
+ ),
+ "latency_ms": {
+ "median": statistics.median(latencies),
+ "p95": sorted(latencies)[min(len(latencies) - 1, int(.95 * len(latencies)))],
+ },
+ "wall_seconds": time.time() - started,
+ }
+ write_json(destination / "report.json", report)
+ return report
+
+
+def reuse_completed(
+ name: str,
+ records: list[dict],
+ predictor,
+ destination: Path,
+) -> dict:
+ """Load and validate one completed dataset from an interrupted matrix run."""
+
+ report_path = destination / "report.json"
+ if not report_path.is_file():
+ raise RuntimeError(
+ f"cannot resume incomplete dataset directory: {destination}; preserve or "
+ "move that directory aside before retrying"
+ )
+ report = read_json(report_path)
+ expected_questions = sum(len(record["questions"]) for record in records)
+ expected_components = {"choice", "native"} if predictor.return_components else {"choice"}
+ actual_components = set(report.get("components", {}))
+ if (
+ report.get("dataset") != name
+ or report.get("records") != len(records)
+ or report.get("questions") != expected_questions
+ or actual_components != expected_components
+ ):
+ raise RuntimeError(f"completed report does not match requested resume dataset: {destination}")
+ for component in expected_components:
+ rows_path = destination / f"{component}-rows.json"
+ if not rows_path.is_file() or len(read_json(rows_path)) != expected_questions:
+ raise RuntimeError(f"completed component rows are missing or incomplete: {rows_path}")
+ return report
+
+
+def _correct(report: dict, component: str) -> int:
+ clean = report["components"][component]["clean"]
+ return round(clean["n"] * clean["acc"])
+
+
+def _code_revision() -> str | None:
+ root = Path(__file__).resolve().parents[1]
+ result = subprocess.run(
+ ["git", "rev-parse", "HEAD"], cwd=root, text=True,
+ stdout=subprocess.PIPE, stderr=subprocess.DEVNULL, check=False,
+ )
+ return result.stdout.strip() if result.returncode == 0 else None
+
+
+def _checkpoint_artifacts(predictor) -> dict | None:
+ checkpoint = predictor.checkpoint
+ if checkpoint is None:
+ return None
+ root = Path(checkpoint.path)
+ files = {}
+ for name in ("adapter_model.safetensors", "head.pt", "adapter_config.json"):
+ path = root / name
+ if path.is_file():
+ files[name] = {"sha256": digest(path), "bytes": path.stat().st_size}
+ return {"requested": checkpoint.requested, "resolved": str(root), "files": files}
+
+
+def main() -> None:
+ parser = argparse.ArgumentParser(description=__doc__)
+ source = parser.add_mutually_exclusive_group(required=True)
+ source.add_argument("--base")
+ source.add_argument("--checkpoint")
+ parser.add_argument("--base-load-path", required=True)
+ parser.add_argument("--revision")
+ parser.add_argument("--model-name", required=True)
+ parser.add_argument("--out", type=Path, required=True)
+ parser.add_argument("--jevbench", type=Path, required=True)
+ parser.add_argument("--transfer", type=Path, required=True)
+ parser.add_argument("--typed", type=Path, required=True)
+ parser.add_argument("--jevjudge", type=Path, required=True)
+ parser.add_argument("--expected-jevbench-native-correct", type=int)
+ parser.add_argument("--expected-transfer-native-correct", type=int)
+ parser.add_argument("--device", default="cuda")
+ parser.add_argument("--dtype", choices=("fp32", "bf16"), default="bf16")
+ parser.add_argument("--attn", choices=("eager", "sdpa"), default="sdpa")
+ parser.add_argument("--max-tokens", type=int, default=65_536)
+ parser.add_argument(
+ "--efficient-long-context-tokens", type=int, default=8_192,
+ help=(
+ "use memory-linear fused SDPA at or above this many tokens; short "
+ "parity panels retain exact math SDPA"
+ ),
+ )
+ parser.add_argument(
+ "--resume", action="store_true",
+ help="reuse validated completed dataset directories in an interrupted run",
+ )
+ parser.add_argument(
+ "--preexisting-code-revision",
+ help="code revision that produced reused reports (required when any are reused)",
+ )
+ args = parser.parse_args()
+ if args.revision and args.checkpoint:
+ parser.error("--revision accompanies --base, not --checkpoint")
+ if args.base and (args.expected_jevbench_native_correct is not None
+ or args.expected_transfer_native_correct is not None):
+ parser.error("native parity assertions require --checkpoint")
+
+ if args.resume:
+ if not args.out.is_dir():
+ parser.error("--resume requires an existing --out directory")
+ else:
+ if args.preexisting_code_revision:
+ parser.error("--preexisting-code-revision requires --resume")
+ args.out.mkdir(parents=True, exist_ok=False)
+ options = LoadOptions(
+ dtype={"fp32": torch.float32, "bf16": torch.bfloat16}[args.dtype],
+ merge=False,
+ temperature=None,
+ base_load_path=args.base_load_path,
+ attn=args.attn,
+ )
+ predictor = LetterReadoutPredictor(
+ base=args.base,
+ checkpoint=args.checkpoint,
+ revision=args.revision,
+ device=args.device,
+ options=options,
+ temperature=1.0,
+ native_weight=0.5 if args.checkpoint else 0.0,
+ return_components=bool(args.checkpoint),
+ exact_kernels=True,
+ efficient_long_context_tokens=args.efficient_long_context_tokens,
+ max_tokens=args.max_tokens,
+ )
+
+ current_code_revision = _code_revision()
+ dataset_code_revisions = {}
+ reused_datasets = []
+
+ def run_dataset(name: str, records: list[dict]) -> dict:
+ destination = args.out / name
+ if args.resume and destination.exists():
+ report = reuse_completed(name, records, predictor, destination)
+ reused_datasets.append(name)
+ dataset_code_revisions[name] = args.preexisting_code_revision
+ return report
+ report = evaluate(name, records, predictor, destination)
+ dataset_code_revisions[name] = current_code_revision
+ return report
+
+ parity_datasets = {
+ "jevbench_public": load_split(args.jevbench, "development"),
+ "transfer_calibration": load_split(args.transfer, "development"),
+ }
+ reports = {
+ name: run_dataset(name, records)
+ for name, records in parity_datasets.items()
+ }
+ if args.checkpoint:
+ checks = (
+ ("jevbench_public", args.expected_jevbench_native_correct),
+ ("transfer_calibration", args.expected_transfer_native_correct),
+ )
+ for dataset, expected in checks:
+ if expected is not None:
+ actual = _correct(reports[dataset], "native")
+ if actual != expected:
+ raise RuntimeError(
+ f"{dataset} native parity failed: expected {expected}, got {actual}"
+ )
+ heldout_datasets = {
+ "transfer_test": load_split(args.transfer, "test", allow_test=True),
+ "typed_test": [_typed_record(row) for row in _typed_rows(args.typed, "test")],
+ "jevjudge_text": [
+ record for record in load_split(args.jevjudge, "test", allow_test=True)
+ if record["_meta"]["modality"] == "text"
+ ],
+ }
+ reports.update({
+ name: run_dataset(name, records)
+ for name, records in heldout_datasets.items()
+ })
+ if reused_datasets and not args.preexisting_code_revision:
+ parser.error("--preexisting-code-revision is required when reports are reused")
+
+ typed_manifest = json.loads((args.typed / ".hf-fetch.json").read_text())
+ manifest = {
+ "model": args.model_name,
+ "source": {"base": args.base, "checkpoint": args.checkpoint},
+ "checkpoint_artifacts": _checkpoint_artifacts(predictor),
+ "base_loading": predictor.base_loading,
+ "predictor": predictor.provenance,
+ "runtime": {
+ "python": platform.python_version(),
+ "torch": torch.__version__,
+ "cuda": torch.version.cuda,
+ "transformers": transformers.__version__,
+ "peft": peft.__version__,
+ "numpy": np.__version__,
+ "pyarrow": pyarrow.__version__,
+ "safetensors": safetensors.__version__,
+ "code_revision": current_code_revision,
+ },
+ "protocol": {
+ "kernel_policy": "LocalPredictor-compatible exact CUDA policy",
+ "dtype": args.dtype,
+ "attn": args.attn,
+ "temperature": 1.0,
+ "blend_selection": (
+ "not performed here; components are saved for fitting on Transfer-v9 "
+ "development only"
+ ),
+ "efficiency_scope": (
+ "base calls execute choice only; checkpoint calls execute choice and native. "
+ "Do not compare component efficiency from these end-to-end timings"
+ ),
+ "long_context_attention": (
+ "exact math SDPA below the configured threshold; memory-linear fused SDPA "
+ "without a math fallback at or above it"
+ ),
+ "efficient_long_context_tokens": args.efficient_long_context_tokens,
+ "resume": {
+ "enabled": args.resume,
+ "reused_datasets": reused_datasets,
+ "preexisting_code_revision": args.preexisting_code_revision,
+ "dataset_code_revisions": dataset_code_revisions,
+ },
+ },
+ "datasets": {
+ "jevbench_development_sha256": digest(args.jevbench / "development.jsonl"),
+ "jevbench_manifest_sha256": digest(args.jevbench / "manifest.json"),
+ "transfer_manifest_sha256": digest(args.transfer / "manifest.json"),
+ "transfer_development_sha256": digest(args.transfer / "development.jsonl"),
+ "transfer_test_sha256": digest(args.transfer / "test.jsonl"),
+ "typed_revision": typed_manifest["resolved_revision"],
+ "typed_test_sha256": digest(
+ args.typed / "all" / "test-00000-of-00001.parquet"
+ ),
+ "jevjudge_test_sha256": digest(args.jevjudge / "test.jsonl"),
+ "jevjudge_manifest_sha256": digest(args.jevjudge / "manifest.json"),
+ },
+ "reports": reports,
+ }
+ write_json(args.out / "manifest.json", manifest)
+ print(json.dumps({
+ "model": args.model_name,
+ "results": {
+ name: {
+ component: report["clean"]["acc"]
+ for component, report in value["components"].items()
+ }
+ for name, value in reports.items()
+ },
+ }, indent=2), flush=True)
+
+
+if __name__ == "__main__":
+ main()
diff --git a/scripts/fit_choice_ensemble.py b/scripts/fit_choice_ensemble.py
new file mode 100644
index 0000000..c4ae147
--- /dev/null
+++ b/scripts/fit_choice_ensemble.py
@@ -0,0 +1,298 @@
+#!/usr/bin/env python3
+"""Fit choice/native pooling on Transfer-v9 development and score frozen tests.
+
+The blend weight and temperature are selected only from the deterministic
+Transfer-v9 development partition emitted by ``evaluate_choice_readout_matrix``.
+Those fixed values are then applied unchanged to every held-out dataset.
+"""
+
+from __future__ import annotations
+
+import argparse
+import json
+import math
+from pathlib import Path
+
+import numpy as np
+
+from jevany.suite import read_json, write_json
+
+
+EPS = 1e-12
+
+
+def load_components(directory: Path) -> tuple[list[dict], list[dict] | None]:
+ choice = read_json(directory / "choice-rows.json")
+ native_path = directory / "native-rows.json"
+ native = read_json(native_path) if native_path.exists() else None
+ if native is not None:
+ left = [(row["id"], row["question"], row["keys"], row["label"]) for row in choice]
+ right = [(row["id"], row["question"], row["keys"], row["label"]) for row in native]
+ if left != right:
+ raise ValueError(f"component rows are not aligned: {directory}")
+ return choice, native
+
+
+def headline_indices(rows: list[dict]) -> list[int]:
+ """Indices in the benchmark's accuracy cohort.
+
+ Transfer-v9 also contains non-clean robustness variants and unknowable
+ questions. The benchmark never counts either in headline accuracy, so
+ neither may influence weight or temperature selection.
+ """
+
+ return [
+ index for index, row in enumerate(rows)
+ if row.get("variant", "clean") == "clean"
+ and row.get("source") != "unknowable"
+ ]
+
+
+def select_indices(rows: list[dict], indices: list[int]) -> list[dict]:
+ return [rows[index] for index in indices]
+
+
+def targets(rows: list[dict]) -> list[np.ndarray]:
+ values = []
+ for row in rows:
+ if "gold" in row:
+ target = np.asarray(row["gold"], dtype=np.float64)
+ target /= target.sum()
+ else:
+ target = np.zeros(len(row["p"]), dtype=np.float64)
+ target[row["label"]] = 1.0
+ values.append(target)
+ return values
+
+
+def pooled_logits(choice: list[dict], native: list[dict] | None, weight: float) -> list[np.ndarray]:
+ if native is None and weight:
+ raise ValueError("native weight requires native component rows")
+ result = []
+ for index, row in enumerate(choice):
+ left = np.log(np.maximum(np.asarray(row["p"], dtype=np.float64), EPS))
+ if native is None:
+ result.append(left)
+ else:
+ right = np.log(np.maximum(np.asarray(native[index]["p"], dtype=np.float64), EPS))
+ result.append((1 - weight) * left + weight * right)
+ return result
+
+
+def probabilities(logits: list[np.ndarray], temperature: float) -> list[np.ndarray]:
+ result = []
+ for values in logits:
+ scaled = values / temperature
+ scaled -= scaled.max()
+ p = np.exp(scaled)
+ result.append(p / p.sum())
+ return result
+
+
+def soft_nll(rows: list[dict], ps: list[np.ndarray]) -> float:
+ gold = targets(rows)
+ return float(np.mean([
+ -float(np.sum(target * np.log(np.maximum(p, EPS))))
+ for target, p in zip(gold, ps, strict=True)
+ ]))
+
+
+def fit_temperature(rows: list[dict], logits: list[np.ndarray]) -> tuple[float, float]:
+ """Golden-section search in log-temperature space with fixed bounds."""
+
+ lo, hi = math.log(.05), math.log(20.0)
+ ratio = (math.sqrt(5) - 1) / 2
+ x1, x2 = hi - ratio * (hi - lo), lo + ratio * (hi - lo)
+
+ def objective(log_temperature: float) -> float:
+ return soft_nll(rows, probabilities(logits, math.exp(log_temperature)))
+
+ f1, f2 = objective(x1), objective(x2)
+ for _ in range(80):
+ if f1 <= f2:
+ hi, x2, f2 = x2, x1, f1
+ x1 = hi - ratio * (hi - lo)
+ f1 = objective(x1)
+ else:
+ lo, x1, f1 = x1, x2, f2
+ x2 = lo + ratio * (hi - lo)
+ f2 = objective(x2)
+ point = (lo + hi) / 2
+ return math.exp(point), objective(point)
+
+
+def fit(rows: list[dict], choice: list[dict], native: list[dict] | None) -> dict:
+ candidates = [0.0] if native is None else [index / 100 for index in range(101)]
+ best = None
+ trace = []
+ for weight in candidates:
+ logits = pooled_logits(choice, native, weight)
+ temperature, nll = fit_temperature(rows, logits)
+ t1 = probabilities(logits, 1.0)
+ correct = sum(
+ int(int(p.argmax()) == row["label"])
+ for row, p in zip(rows, t1, strict=True)
+ )
+ item = {
+ "native_weight": weight,
+ "temperature": temperature,
+ "correct": correct,
+ "accuracy": correct / len(rows),
+ "nll": nll,
+ }
+ trace.append(item)
+ if best is None or (
+ -correct, nll, abs(weight - .5)
+ ) < (
+ -best["correct"], best["nll"], abs(best["native_weight"] - .5)
+ ):
+ best = item
+ return {"selected": best, "grid": trace}
+
+
+def ece(rows: list[dict], ps: list[np.ndarray], bins: int = 15) -> float:
+ totals = np.zeros(bins)
+ correct = np.zeros(bins)
+ confidence = np.zeros(bins)
+ for row, p in zip(rows, ps, strict=True):
+ conf = float(p.max())
+ bucket = min(bins - 1, int(conf * bins))
+ totals[bucket] += 1
+ correct[bucket] += int(int(p.argmax()) == row["label"])
+ confidence[bucket] += conf
+ n = len(rows)
+ return float(sum(
+ totals[index] / n * abs(correct[index] / totals[index] - confidence[index] / totals[index])
+ for index in range(bins) if totals[index]
+ ))
+
+
+def metrics(rows: list[dict], ps: list[np.ndarray], *, ece_bins: int = 10) -> dict:
+ gold = targets(rows)
+ cross_entropy = soft_nll(rows, ps)
+ gold_entropy = float(np.mean([
+ -float(np.sum(target * np.log(np.maximum(target, EPS)))) for target in gold
+ ]))
+ return {
+ "n": len(rows),
+ "correct": sum(int(int(p.argmax()) == row["label"]) for row, p in zip(rows, ps, strict=True)),
+ "accuracy": float(np.mean([
+ int(int(p.argmax()) == row["label"])
+ for row, p in zip(rows, ps, strict=True)
+ ])),
+ "cross_entropy_from_gold": cross_entropy,
+ "gold_entropy": gold_entropy,
+ "kl_from_gold": cross_entropy - gold_entropy,
+ "brier": float(np.mean([
+ float(np.sum((p - target) ** 2))
+ for target, p in zip(gold, ps, strict=True)
+ ])),
+ "ece": ece(rows, ps, bins=ece_bins),
+ "mean_confidence": float(np.mean([p.max() for p in ps])),
+ }
+
+
+def evaluate(directory: Path, selected: dict) -> dict:
+ choice, native = load_components(directory)
+ configurations = {
+ "choice_t1": (0.0, 1.0),
+ "transfer_dev_calibrated_choice": (0.0, selected["choice_temperature"]),
+ }
+ if native is not None:
+ configurations.update({
+ "native_shipped": (1.0, 1.0),
+ "fixed_blend_0_5": (.5, 1.0),
+ "transfer_dev_tuned_blend": (
+ selected["native_weight"], selected["blend_temperature"]
+ ),
+ })
+ indices = headline_indices(choice)
+ ece_bins = 15 if directory.name == "typed_test" else 10
+ results = {}
+ for name, (weight, temperature) in configurations.items():
+ ps = probabilities(pooled_logits(choice, native, weight), temperature)
+ headline = metrics(
+ select_indices(choice, indices),
+ [ps[index] for index in indices],
+ ece_bins=ece_bins,
+ )
+ headline["cohort"] = "clean and knowable"
+ headline["excluded_non_headline_rows"] = len(choice) - len(indices)
+ headline["ece_bins"] = ece_bins
+ results[name] = headline
+ return results
+
+
+def main() -> None:
+ parser = argparse.ArgumentParser(description=__doc__)
+ parser.add_argument("--run", type=Path, required=True)
+ parser.add_argument("--out", type=Path, required=True)
+ args = parser.parse_args()
+ manifest = read_json(args.run / "manifest.json")
+ calibration_dir = args.run / "transfer_calibration"
+ choice, native = load_components(calibration_dir)
+ indices = headline_indices(choice)
+ calibration_choice = select_indices(choice, indices)
+ calibration_native = select_indices(native, indices) if native is not None else None
+ choice_fit = fit(calibration_choice, calibration_choice, None)
+ blend_fit = fit(calibration_choice, calibration_choice, calibration_native)
+ selected = {
+ "choice_temperature": choice_fit["selected"]["temperature"],
+ "native_weight": blend_fit["selected"]["native_weight"],
+ "blend_temperature": blend_fit["selected"]["temperature"],
+ "selection_objective": (
+ "hard-label accuracy on the Transfer-v9 development clean/knowable "
+ "cohort; calibrated NLL and distance from 0.5 break ties"
+ ),
+ }
+ datasets = {}
+ for name in (
+ "transfer_calibration", "transfer_test", "typed_test", "jevbench_public",
+ "jevjudge_text"
+ ):
+ datasets[name] = evaluate(args.run / name, selected)
+ result = {
+ "model": manifest["model"],
+ "source_manifest": str(args.run / "manifest.json"),
+ "protocol": {
+ "selection_split": (
+ "Transfer-v9 development clean/knowable accuracy cohort; never any "
+ "reported test panel"
+ ),
+ "selection_rows": len(calibration_choice),
+ "weight_grid": "0.00 to 1.00 inclusive, step 0.01",
+ "pool": "log-linear/geometric",
+ "temperature_fit": "bounded scalar NLL minimization, T in [0.05, 20]",
+ "ece": "15 equal-width bins for Typed Decisions; 10 for every other dataset",
+ "selection_objective": "accuracy; calibrated NLL then distance from 0.5 break ties",
+ "transfer": "the selected weight and temperature are frozen for every held-out dataset",
+ "native_shipped": (
+ "the checkpoint's native probabilities, including its shipped inference "
+ "temperature; no additional temperature is applied"
+ ),
+ "fixed_blend_0_5": (
+ "equal log-linear pooling of T=1 choice probabilities and shipped native "
+ "probabilities; no additional temperature is applied"
+ ),
+ "transfer_dev_tuned_blend": (
+ "the Transfer-dev-selected native weight followed by the reported scalar "
+ "additional temperature"
+ ),
+ },
+ "selected": selected,
+ "selection": {"choice": choice_fit, "blend": blend_fit},
+ "datasets": datasets,
+ }
+ write_json(args.out, result)
+ print(json.dumps({
+ "model": result["model"],
+ "selected": selected,
+ "test_accuracy": {
+ dataset: {name: values["accuracy"] for name, values in scores.items()}
+ for dataset, scores in datasets.items() if dataset != "transfer_calibration"
+ },
+ }, indent=2), flush=True)
+
+
+if __name__ == "__main__":
+ main()
diff --git a/site/files/reports/JevAny_Tech_Report.pdf b/site/files/reports/JevAny_Tech_Report.pdf
index de0ad58..cedc9fb 100644
Binary files a/site/files/reports/JevAny_Tech_Report.pdf and b/site/files/reports/JevAny_Tech_Report.pdf differ
diff --git a/tests/test_build_choice_readout_results.py b/tests/test_build_choice_readout_results.py
new file mode 100644
index 0000000..eb8863d
--- /dev/null
+++ b/tests/test_build_choice_readout_results.py
@@ -0,0 +1,362 @@
+import copy
+import json
+from pathlib import Path
+
+import pytest
+
+from scripts.build_choice_readout_results import (
+ DATASETS,
+ RUN_SPECS,
+ build_results,
+ render_svg,
+ write_results,
+)
+
+
+def _metric(n, correct):
+ accuracy = correct / n
+ return {
+ "n": n,
+ "correct": correct,
+ "accuracy": accuracy,
+ "cross_entropy_from_gold": 0.5,
+ "gold_entropy": 0.1,
+ "kl_from_gold": 0.4,
+ "brier": 0.2,
+ "ece": 0.03,
+ "mean_confidence": 0.7,
+ "cohort": "clean and knowable",
+ "excluded_non_headline_rows": 0,
+ "ece_bins": 15,
+ }
+
+
+def _write_json(path, value):
+ path.parent.mkdir(parents=True, exist_ok=True)
+ path.write_text(json.dumps(value, indent=2) + "\n")
+
+
+def _external_fixture(path):
+ value = {
+ "artifact_version": 5,
+ "provenance": {"fixture": {"sha256": "e" * 64}},
+ "typed_decisions": {
+ "models": [
+ {
+ "model": "Published runner-up",
+ "kind": "published_only",
+ "accuracy": 0.74,
+ },
+ {
+ "model": "Published top",
+ "kind": "published_only",
+ "accuracy": 0.77,
+ },
+ {
+ "model": "Jev 1.13 (OpenRouter)",
+ "kind": "external_api",
+ "accuracy": 0.73,
+ },
+ ]
+ },
+ "jevjudge_text": {
+ "models": [
+ {
+ "model": "Jev 1.13 (OpenRouter)",
+ "kind": "external_api",
+ "accuracy": 0.65,
+ "answered": 724,
+ "requested": 724,
+ },
+ {
+ "model": "Kev-4B",
+ "kind": "open_kev",
+ "accuracy": 0.54,
+ "answered": 724,
+ "requested": 724,
+ },
+ {
+ "model": "Kev-27B",
+ "kind": "open_kev",
+ "accuracy": 0.64,
+ "answered": 724,
+ "requested": 724,
+ },
+ ]
+ },
+ }
+ _write_json(path, value)
+
+
+def _run_fixture(root, spec, ordinal):
+ is_base = spec["weights"] == "frozen_base"
+ canonical_base = f"Qwen/{spec['family']}"
+ snapshot_revision = f"{ordinal + 1}" * 40
+ encoded_repository = spec["public_repository"].replace("/", "--")
+ checkpoint = (
+ f"/fixture/models--{encoded_repository}/snapshots/{snapshot_revision}"
+ if not is_base and spec["key"] != "direct4"
+ else f"/fixture/local-release/{spec['key']}"
+ )
+ components = ["choice"] if is_base else ["choice", "native"]
+ model = f"fixture-{spec['key']}"
+ reports = {}
+ ensemble_datasets = {}
+ for dataset_index, (dataset, expected) in enumerate(DATASETS.items()):
+ n = expected["headline_n"]
+ choice_correct = n // 2 + ordinal + dataset_index
+ native_correct = choice_correct + 1
+ choice_accuracy = choice_correct / n
+ native_accuracy = native_correct / n
+ reports[dataset] = {
+ "dataset": dataset,
+ "records": expected["records"],
+ "questions": expected["questions"],
+ "components_executed": components,
+ "components": {
+ "choice": {"clean": {"n": n, "acc": choice_accuracy}},
+ },
+ }
+ if not is_base:
+ reports[dataset]["components"]["native"] = {
+ "clean": {"n": n, "acc": native_accuracy}
+ }
+ choice = _metric(n, choice_correct)
+ choice["excluded_non_headline_rows"] = expected["questions"] - n
+ calibrated = copy.deepcopy(choice)
+ calibrated["cross_entropy_from_gold"] = 0.45
+ configurations = {
+ "choice_t1": choice,
+ "transfer_dev_calibrated_choice": calibrated,
+ }
+ if not is_base:
+ native = _metric(n, native_correct)
+ fixed = _metric(n, min(n, native_correct + 1))
+ tuned = _metric(n, min(n, native_correct + 2))
+ for value in (native, fixed, tuned):
+ value["excluded_non_headline_rows"] = expected["questions"] - n
+ configurations.update({
+ "native_shipped": native,
+ "fixed_blend_0_5": fixed,
+ "transfer_dev_tuned_blend": tuned,
+ })
+ ensemble_datasets[dataset] = configurations
+
+ hashes = {
+ "jevbench_development_sha256": "1" * 64,
+ "jevbench_manifest_sha256": "2" * 64,
+ "transfer_manifest_sha256": "3" * 64,
+ "transfer_development_sha256": "4" * 64,
+ "transfer_test_sha256": "5" * 64,
+ "typed_revision": "6" * 40,
+ "typed_test_sha256": "7" * 64,
+ "jevjudge_test_sha256": "8" * 64,
+ "jevjudge_manifest_sha256": "9" * 64,
+ }
+ manifest = {
+ "model": model,
+ "source": {
+ "base": spec["public_repository"] if is_base else None,
+ "checkpoint": None if is_base else checkpoint,
+ },
+ "checkpoint_artifacts": None if is_base else {
+ "requested": f"fixture/{spec['key']}",
+ "resolved": f"/fixture/{spec['key']}",
+ "files": {
+ "adapter_model.safetensors": {
+ "sha256": f"{ordinal + 1}" * 64,
+ "bytes": 123,
+ }
+ },
+ },
+ "base_loading": {
+ "requested": spec["family"],
+ "resolved": "/fixture/base",
+ "canonical_base": canonical_base,
+ "canonical_revision": "a" * 40,
+ },
+ "predictor": {
+ "method": "training-free exact choice-token projection",
+ "adapter_applied": not is_base,
+ "native_decision_mode": spec["native_decision_mode"],
+ "canonical_base": canonical_base,
+ "canonical_revision": "a" * 40,
+ },
+ "runtime": {
+ "python": "3.12.0",
+ "torch": "2.fixture",
+ "code_revision": "b" * 40,
+ },
+ "protocol": {"temperature": 1.0, "kernel_policy": "fixture exact"},
+ "datasets": hashes,
+ "reports": reports,
+ }
+ ensemble = {
+ "model": model,
+ "protocol": {
+ "selection_rows": 1046,
+ "selection_split": "Transfer-v9 development",
+ },
+ "selected": {
+ "choice_temperature": 1.2,
+ "native_weight": 0.0 if is_base else 0.55,
+ "blend_temperature": 1.1,
+ "selection_objective": "fixture",
+ },
+ "datasets": ensemble_datasets,
+ }
+ directory = root / spec["key"]
+ _write_json(directory / "manifest.json", manifest)
+ _write_json(directory / "ensemble.json", ensemble)
+
+
+@pytest.fixture
+def matrix(tmp_path):
+ root = tmp_path / "runs"
+ for ordinal, spec in enumerate(RUN_SPECS):
+ _run_fixture(root, spec, ordinal)
+ external = tmp_path / "external.json"
+ _external_fixture(external)
+ return root, external
+
+
+def test_build_results_preserves_all_five_run_types_and_protocol_groups(matrix):
+ root, external = matrix
+
+ artifact = build_results(root, external)
+
+ assert artifact["artifact"] == "choice-readout-v2"
+ assert [run["key"] for run in artifact["source_runs"]] == [
+ "base4", "pointer4", "direct4", "base27", "pointer27"
+ ]
+ assert {run["weights"] for run in artifact["source_runs"]} == {
+ "frozen_base", "jevany_sft"
+ }
+ assert {run["native_readout"] for run in artifact["source_runs"]} == {
+ None, "pointer", "direct-token"
+ }
+ typed = artifact["datasets"]["typed_test"]
+ assert typed["questions"] == typed["headline_n"] == 2000
+ assert typed["external_baselines"][0]["model"] == "Published top"
+ assert artifact["datasets"]["jevjudge_text"]["external_baselines"][1]["model"] == "Kev-27B"
+ assert artifact["datasets"]["jevjudge_full"] == {
+ "label": "JevJudge full",
+ "records": 3220,
+ "status": "unsupported",
+ "accuracy": None,
+ "reason": (
+ "Training-free choice-token readout is text-only; image/video records "
+ "are not stripped or relabeled as a full-suite result."
+ ),
+ "native_checkpoint_results": "results/external-zero-shot-v1.json#jevjudge_full",
+ }
+ base = next(row for row in typed["runs"] if row["run"] == "base4")
+ direct = next(row for row in typed["runs"] if row["run"] == "direct4")
+ assert set(base["zero_shot"]) == {"choice_t1"}
+ assert set(base["transfer_dev_tuned"]) == {"transfer_dev_calibrated_choice"}
+ assert set(direct["zero_shot"]) == {
+ "choice_t1", "native_shipped", "fixed_blend_0_5"
+ }
+ assert set(direct["transfer_dev_tuned"]) == {
+ "transfer_dev_calibrated_choice", "transfer_dev_tuned_blend"
+ }
+ assert len(artifact["source_runs"][0]["artifacts"]["manifest"]["sha256"]) == 64
+ assert artifact["source_runs"][0]["runtime"]["torch"] == "2.fixture"
+
+ sources = {run["key"]: run["public_source"] for run in artifact["source_runs"]}
+ assert sources["base4"] == {
+ "repository": "Qwen/Qwen3.5-4B",
+ "url": "https://huggingface.co/Qwen/Qwen3.5-4B",
+ "revision": "a" * 40,
+ "revision_status": "verified",
+ "repository_evidence": "manifest.base_loading.canonical_base",
+ "revision_evidence": "manifest.base_loading.canonical_revision",
+ "evaluated_artifact_sha256": {},
+ "base_model": {
+ "repository": "Qwen/Qwen3.5-4B",
+ "revision": "a" * 40,
+ },
+ "limitation": None,
+ }
+ assert sources["pointer4"]["revision"] == "2" * 40
+ assert sources["pointer4"]["revision_status"] == "verified"
+ assert sources["pointer27"]["revision"] == "5" * 40
+ assert sources["direct4"]["repository"] == (
+ "SimpleJev/JevAny-Qwen3.5-4B-Direct-Token-LoRA"
+ )
+ assert sources["direct4"]["revision"] is None
+ assert sources["direct4"]["revision_status"] == "not_verified"
+ assert sources["direct4"]["revision_evidence"] is None
+ assert "do not prove an immutable public-repository revision" in (
+ sources["direct4"]["limitation"]
+ )
+ assert sources["direct4"]["evaluated_artifact_sha256"] == {
+ "adapter_model.safetensors": {
+ "sha256": "3" * 64,
+ "bytes": 123,
+ }
+ }
+
+
+def test_build_results_rejects_component_count_drift(matrix):
+ root, external = matrix
+ path = root / "pointer4" / "manifest.json"
+ manifest = json.loads(path.read_text())
+ manifest["reports"]["typed_test"]["components"]["native"]["clean"]["n"] = 1999
+ _write_json(path, manifest)
+
+ with pytest.raises(ValueError, match="native expected clean n=2000"):
+ build_results(root, external)
+
+
+def test_build_results_rejects_native_component_metric_misalignment(matrix):
+ root, external = matrix
+ path = root / "direct4" / "ensemble.json"
+ ensemble = json.loads(path.read_text())
+ native = ensemble["datasets"]["jevjudge_text"]["native_shipped"]
+ native["correct"] += 1
+ native["accuracy"] = native["correct"] / native["n"]
+ _write_json(path, ensemble)
+
+ with pytest.raises(ValueError, match="native component and ensemble are not aligned"):
+ build_results(root, external)
+
+
+def test_build_results_rejects_cross_run_dataset_hash_drift(matrix):
+ root, external = matrix
+ path = root / "direct4" / "manifest.json"
+ manifest = json.loads(path.read_text())
+ manifest["datasets"]["typed_test_sha256"] = "0" * 64
+ _write_json(path, manifest)
+
+ with pytest.raises(ValueError, match="dataset hashes differ"):
+ build_results(root, external)
+
+
+def test_build_results_rejects_release_catalog_without_canonical_repositories(
+ matrix, tmp_path
+):
+ root, external = matrix
+ catalog = tmp_path / "model-catalog.json"
+ _write_json(catalog, {"release": "fixture", "released_models": []})
+
+ with pytest.raises(ValueError, match="missing canonical repositories"):
+ build_results(root, external, catalog)
+
+
+def test_synthetic_artifact_and_svg_are_written_only_to_requested_paths(matrix, tmp_path):
+ pytest.importorskip("matplotlib")
+ root, external = matrix
+ artifact = build_results(root, external)
+ json_out = tmp_path / "out" / "choice.json"
+ svg_out = tmp_path / "out" / "choice.svg"
+
+ write_results(json_out, artifact)
+ render_svg(artifact, svg_out)
+
+ assert json.loads(json_out.read_text())["schema_version"] == 2
+ svg = svg_out.read_text()
+ assert "Training-free choice readout" in svg
+ assert "Typed Decisions" in svg
+ assert "JevJudge text" in svg
+ assert "TRANSFER-DEV-TUNED" in svg
diff --git a/tests/test_build_external_report_appendix.py b/tests/test_build_external_report_appendix.py
new file mode 100644
index 0000000..ef2281f
--- /dev/null
+++ b/tests/test_build_external_report_appendix.py
@@ -0,0 +1,556 @@
+import json
+import os
+import shutil
+from datetime import datetime, timezone
+from pathlib import Path
+
+import pytest
+
+try:
+ import matplotlib
+ import pypdf
+except ImportError:
+ if os.environ.get("JEVANY_REPORT_TESTS_REQUIRED") == "1":
+ raise
+ pytest.skip("report dependencies are installed in the report-test job", allow_module_level=True)
+
+import matplotlib.pyplot as plt
+from matplotlib.backends.backend_pdf import PdfPages
+from pypdf import PdfReader, PdfWriter
+
+from scripts.build_external_report_appendix import (
+ CHOICE_DATASET_SPECS,
+ CHOICE_RUN_SPECS,
+ _choice_matrix_rows,
+ build_appendix,
+ load_choice_artifact,
+ merge_report,
+ page_invariants,
+)
+from scripts.build_choice_readout_results import (
+ DATASETS as PRODUCER_DATASETS,
+ RUN_SPECS as PRODUCER_RUN_SPECS,
+ build_results,
+ write_results,
+)
+
+
+def _write_json(path, value):
+ path.write_text(json.dumps(value, indent=2) + "\n", encoding="utf-8")
+
+
+def _external_artifact(path):
+ value = {
+ "artifact_version": 5,
+ "provenance": {"fixture": {"sha256": "e" * 64}},
+ "typed_decisions": {
+ "models": [
+ {"model": "JevAny-Qwen3.8-27B", "accuracy": 0.728},
+ {
+ "model": "Jev 1.13 (OpenRouter)",
+ "kind": "external_api",
+ "accuracy": 0.727,
+ },
+ {
+ "model": "Published top",
+ "kind": "published_only",
+ "accuracy": 0.770,
+ },
+ ],
+ },
+ "jevjudge_text": {
+ "models": [
+ {"model": "JevAny-Qwen3.8-27B", "accuracy": 0.6644},
+ {
+ "model": "Jev 1.13 (OpenRouter)",
+ "kind": "external_api",
+ "accuracy": 0.6506,
+ "answered": 724,
+ "requested": 724,
+ },
+ {
+ "model": "Kev-27B",
+ "kind": "open_kev",
+ "accuracy": 0.6423,
+ "answered": 724,
+ "requested": 724,
+ },
+ ],
+ },
+ "jevjudge_full": {
+ "models": [
+ {
+ "model": "JevAny-Qwen3.8-27B",
+ "kind": "ours",
+ "accuracy": 0.6227,
+ "skill_role": 0.3555,
+ "skill_role_ci_95": [0.3284, 0.3818],
+ "nll": 0.878,
+ "ece": 0.094,
+ },
+ {
+ "model": "Jeff-Qwen3.5-2B",
+ "kind": "open_external",
+ "accuracy": 0.4823,
+ "skill_role": 0.1357,
+ "skill_role_ci_95": [0.1080, 0.1599],
+ "nll": 1.105,
+ "ece": 0.137,
+ },
+ ],
+ },
+ }
+ _write_json(path, value)
+
+
+def _metric(n, correct):
+ return {"n": n, "correct": correct, "accuracy": correct / n}
+
+
+def _choice_artifact(path):
+ source_runs = []
+ for index, (key, (family, weights, native)) in enumerate(CHOICE_RUN_SPECS.items()):
+ source_runs.append({
+ "key": key,
+ "model": f"fixture-{key}",
+ "family": family,
+ "weights": weights,
+ "native_readout": native,
+ "selected": {
+ "native_weight": 0.0 if native is None else 0.40 + index / 100,
+ "choice_temperature": 1.20 + index / 100,
+ "blend_temperature": 1.10 + index / 100,
+ },
+ "ensemble_protocol": {"selection_rows": 1_046},
+ })
+
+ datasets = {}
+ for dataset_index, (dataset, _label, records, questions, n) in enumerate(
+ CHOICE_DATASET_SPECS
+ ):
+ runs = []
+ for run_index, (key, (family, weights, native)) in enumerate(
+ CHOICE_RUN_SPECS.items()
+ ):
+ choice_correct = n // 2 + dataset_index * 2 + run_index
+ row = {
+ "run": key,
+ "model": f"fixture-{key}",
+ "family": family,
+ "weights": weights,
+ "native_readout": native,
+ "zero_shot": {"choice_t1": _metric(n, choice_correct)},
+ "transfer_dev_tuned": {
+ "transfer_dev_calibrated_choice": _metric(n, choice_correct),
+ },
+ }
+ if native is not None:
+ row["zero_shot"].update({
+ "native_shipped": _metric(n, choice_correct + 1),
+ "fixed_blend_0_5": _metric(n, choice_correct + 2),
+ })
+ row["transfer_dev_tuned"]["transfer_dev_tuned_blend"] = _metric(
+ n, choice_correct + 3
+ )
+ runs.append(row)
+ datasets[dataset] = {
+ "records": records,
+ "questions": questions,
+ "headline_n": n,
+ "runs": runs,
+ }
+ datasets["jevjudge_full"] = {
+ "label": "JevJudge full",
+ "records": 3_220,
+ "status": "unsupported",
+ "accuracy": None,
+ "reason": "text-only fixture",
+ }
+ value = {
+ "schema_version": 2,
+ "artifact": "choice-readout-v2",
+ "method": {
+ "option_ids": "A-Z followed by a-z; at most 52 options.",
+ "calibration": "Scalar temperature changes probabilities but not argmax accuracy.",
+ },
+ "source_runs": source_runs,
+ "datasets": datasets,
+ }
+ _write_json(path, value)
+ return value
+
+
+def _producer_metric(n, correct, excluded):
+ return {
+ "n": n,
+ "correct": correct,
+ "accuracy": correct / n,
+ "cross_entropy_from_gold": 0.5,
+ "gold_entropy": 0.1,
+ "kl_from_gold": 0.4,
+ "brier": 0.2,
+ "ece": 0.03,
+ "mean_confidence": 0.7,
+ "cohort": "clean and knowable",
+ "excluded_non_headline_rows": excluded,
+ "ece_bins": 15,
+ }
+
+
+def _producer_run(root, spec, ordinal):
+ is_base = spec["weights"] == "frozen_base"
+ canonical_base = f"Qwen/{spec['family']}"
+ encoded_repository = spec["public_repository"].replace("/", "--")
+ checkpoint = (
+ f"/fixture/models--{encoded_repository}/snapshots/{str(ordinal + 1) * 40}"
+ if spec["key"] != "direct4"
+ else f"/fixture/local-release/{spec['key']}"
+ )
+ components = ["choice"] if is_base else ["choice", "native"]
+ model = f"producer-{spec['key']}"
+ reports = {}
+ ensemble_datasets = {}
+ for dataset_index, (dataset, expected) in enumerate(PRODUCER_DATASETS.items()):
+ n = expected["headline_n"]
+ excluded = expected["questions"] - n
+ choice_correct = n // 2 + ordinal + dataset_index
+ native_correct = choice_correct + 1
+ reports[dataset] = {
+ "dataset": dataset,
+ "records": expected["records"],
+ "questions": expected["questions"],
+ "components_executed": components,
+ "components": {
+ "choice": {"clean": {"n": n, "acc": choice_correct / n}},
+ },
+ }
+ if not is_base:
+ reports[dataset]["components"]["native"] = {
+ "clean": {"n": n, "acc": native_correct / n}
+ }
+ choice = _producer_metric(n, choice_correct, excluded)
+ calibrated = dict(choice)
+ calibrated["cross_entropy_from_gold"] = 0.45
+ configurations = {
+ "choice_t1": choice,
+ "transfer_dev_calibrated_choice": calibrated,
+ }
+ if not is_base:
+ configurations.update({
+ "native_shipped": _producer_metric(n, native_correct, excluded),
+ "fixed_blend_0_5": _producer_metric(n, native_correct + 1, excluded),
+ "transfer_dev_tuned_blend": _producer_metric(
+ n, native_correct + 2, excluded
+ ),
+ })
+ ensemble_datasets[dataset] = configurations
+
+ hashes = {
+ "jevbench_development_sha256": "1" * 64,
+ "jevbench_manifest_sha256": "2" * 64,
+ "transfer_manifest_sha256": "3" * 64,
+ "transfer_development_sha256": "4" * 64,
+ "transfer_test_sha256": "5" * 64,
+ "typed_revision": "6" * 40,
+ "typed_test_sha256": "7" * 64,
+ "jevjudge_test_sha256": "8" * 64,
+ "jevjudge_manifest_sha256": "9" * 64,
+ }
+ manifest = {
+ "model": model,
+ "source": {
+ "base": spec["public_repository"] if is_base else None,
+ "checkpoint": None if is_base else checkpoint,
+ },
+ "checkpoint_artifacts": None if is_base else {
+ "requested": f"fixture/{spec['key']}",
+ "resolved": f"/fixture/{spec['key']}",
+ "files": {
+ "adapter_model.safetensors": {
+ "sha256": f"{ordinal + 1}" * 64,
+ "bytes": 123,
+ },
+ },
+ },
+ "base_loading": {
+ "requested": spec["family"],
+ "resolved": "/fixture/base",
+ "canonical_base": canonical_base,
+ "canonical_revision": "a" * 40,
+ },
+ "predictor": {
+ "method": "training-free exact choice-token projection",
+ "adapter_applied": not is_base,
+ "native_decision_mode": spec["native_decision_mode"],
+ "canonical_base": canonical_base,
+ "canonical_revision": "a" * 40,
+ },
+ "runtime": {
+ "python": "3.13.5",
+ "torch": "2.fixture",
+ "code_revision": "b" * 40,
+ },
+ "protocol": {"temperature": 1.0, "kernel_policy": "fixture exact"},
+ "datasets": hashes,
+ "reports": reports,
+ }
+ ensemble = {
+ "model": model,
+ "protocol": {
+ "selection_rows": 1_046,
+ "selection_split": "Transfer-v9 development",
+ },
+ "selected": {
+ "choice_temperature": 1.2,
+ "native_weight": 0.0 if is_base else 0.50 + ordinal / 100,
+ "blend_temperature": 1.0 + ordinal / 10,
+ "selection_objective": "fixture",
+ },
+ "datasets": ensemble_datasets,
+ }
+ directory = root / spec["key"]
+ directory.mkdir(parents=True)
+ _write_json(directory / "manifest.json", manifest)
+ _write_json(directory / "ensemble.json", ensemble)
+
+
+def _base_report(path, pages=21):
+ metadata = {
+ "Title": "Synthetic uniquely identified base report",
+ "CreationDate": datetime(2026, 10, 2, tzinfo=timezone.utc),
+ "ModDate": datetime(2026, 10, 2, tzinfo=timezone.utc),
+ }
+ with PdfPages(path, metadata=metadata) as pdf:
+ for page_number in range(1, pages + 1):
+ fig = plt.figure(figsize=(8.5, 11), facecolor="white")
+ fig.text(
+ 0.1,
+ 0.8,
+ f"UNIQUE BASE PAGE {page_number:02d}",
+ fontsize=16,
+ url=f"https://example.test/base/{page_number}",
+ )
+ fig.text(0.1, 0.7, f"resource marker {page_number * 17}", fontsize=9)
+ pdf.savefig(fig)
+ plt.close(fig)
+
+
+def _annotation_uris(page):
+ uris = []
+ for reference in page.get("/Annots", []):
+ annotation = reference.get_object()
+ action = annotation.get("/A")
+ if action is not None:
+ action = action.get_object()
+ if action.get("/URI") is not None:
+ uris.append(str(action["/URI"]))
+ return uris
+
+
+@pytest.fixture
+def appendix_inputs(tmp_path):
+ external = tmp_path / "external.json"
+ choice = tmp_path / "choice.json"
+ chart = tmp_path / "chart.png"
+ _external_artifact(external)
+ _choice_artifact(choice)
+ plt.imsave(chart, [[0.0, 0.5], [0.75, 1.0]], cmap="viridis")
+ return external, chart, choice
+
+
+@pytest.fixture
+def producer_artifact(tmp_path):
+ root = tmp_path / "producer-runs"
+ for ordinal, spec in enumerate(PRODUCER_RUN_SPECS):
+ _producer_run(root, spec, ordinal)
+ external = tmp_path / "producer-external.json"
+ _external_artifact(external)
+ artifact = build_results(root, external)
+ choice = tmp_path / "producer-choice.json"
+ write_results(choice, artifact)
+ return external, choice
+
+
+def test_build_appendix_renders_four_pages_directly_from_choice_json(
+ appendix_inputs, tmp_path
+):
+ external, chart, choice = appendix_inputs
+ output = tmp_path / "appendices.pdf"
+
+ build_appendix(external, chart, choice, output)
+
+ pdf = PdfReader(output)
+ assert len(pdf.pages) == 4
+ method_text = pdf.pages[2].extract_text()
+ results_text = pdf.pages[3].extract_text()
+ assert "APPENDIX M" in method_text
+ assert "52 exact IDs" in method_text
+ assert "not temperature calibration" in method_text
+ assert "Choice-token accuracy matrix" in results_text
+ assert "Matched runs" in results_text
+ assert "JevJudge text 724" in results_text
+ # base4 JevBench fixture: floor(231 / 2) = 115 => 49.78%.
+ assert "49.78" in results_text
+
+
+def test_producer_artifact_flows_into_pointer_direct_and_tuned_pdf_rows(
+ producer_artifact, tmp_path
+):
+ external, choice = producer_artifact
+ chart = tmp_path / "producer-chart.png"
+ output = tmp_path / "producer-appendices.pdf"
+ plt.imsave(chart, [[0.0, 0.5], [0.75, 1.0]], cmap="viridis")
+
+ loaded = load_choice_artifact(choice)
+ source_by_key = {row["key"]: row for row in loaded["source_runs"]}
+ assert source_by_key["pointer4"]["native_readout"] == "pointer"
+ assert source_by_key["direct4"]["native_readout"] == "direct-token"
+ assert source_by_key["pointer4"]["selected"]["native_weight"] == 0.51
+ assert source_by_key["direct4"]["selected"]["native_weight"] == 0.52
+ rows = _choice_matrix_rows(loaded)
+ assert [(row["model"], row["readout"], row["run"]) for row in rows] == [
+ ("Frozen Qwen3.5-4B", "Choice T=1", "base4"),
+ ("JevAny 4B Pointer", "Native", "pointer4"),
+ ("", "Choice T=1", "pointer4"),
+ ("", "Tuned stack", "pointer4"),
+ ("JevAny 4B Direct-Token", "Native", "direct4"),
+ ("", "Choice T=1", "direct4"),
+ ("", "Tuned stack", "direct4"),
+ ("Frozen Qwen3.8-27B", "Choice T=1", "base27"),
+ ("JevAny 27B Pointer", "Native", "pointer27"),
+ ("", "Choice T=1", "pointer27"),
+ ("", "Tuned stack", "pointer27"),
+ ]
+ matrix = {(row["run"], row["readout"]): row for row in rows}
+ assert matrix[("pointer4", "Native")]["metrics"][0]["correct"] == 120
+ assert matrix[("pointer4", "Tuned stack")]["metrics"][0]["correct"] == 122
+ assert matrix[("direct4", "Native")]["metrics"][0]["correct"] == 121
+ assert matrix[("direct4", "Tuned stack")]["metrics"][0]["correct"] == 123
+ build_appendix(external, chart, choice, output)
+
+ text = PdfReader(output).pages[3].extract_text()
+ assert "JevAny 4B Pointer" in text
+ assert "JevAny 4B Direct-Token" in text
+ assert text.count("Tuned stack") == 3
+ assert "4B Pointer: w=0.51, T=1.10" in text
+ assert "4B Direct-Token: w=0.52, T=1.20" in text
+ assert "27B Pointer: w=0.54, T=1.40" in text
+ # JevBench producer values: Pointer native/tuned and Direct native/tuned.
+ for accuracy in ("51.95", "52.81", "52.38", "53.25"):
+ assert accuracy in text
+
+
+def test_appendix_and_in_place_merge_are_deterministic_and_replace_old_pages(
+ appendix_inputs, tmp_path
+):
+ external, chart, choice = appendix_inputs
+ appendix_a = tmp_path / "appendices-a.pdf"
+ appendix_b = tmp_path / "appendices-b.pdf"
+ base = tmp_path / "base.pdf"
+ merged_a = tmp_path / "merged-a.pdf"
+ merged_b = tmp_path / "merged-b.pdf"
+ inplace = tmp_path / "inplace.pdf"
+ build_appendix(external, chart, choice, appendix_a)
+ build_appendix(external, chart, choice, appendix_b)
+ assert appendix_a.read_bytes() == appendix_b.read_bytes()
+ appendix_metadata = PdfReader(appendix_a).metadata
+ assert appendix_metadata["/CreationDate"] == "D:20261002000000Z"
+ assert appendix_metadata["/ModDate"] == "D:20261002000000Z"
+ assert appendix_metadata["/Title"] == "JevAny Technical Report โ Evaluation Appendices"
+
+ _base_report(base)
+ original = PdfReader(base)
+ original_invariants = [page_invariants(original.pages[index]) for index in range(19)]
+ assert "UNIQUE BASE PAGE 20" in original.pages[19].extract_text()
+ assert "UNIQUE BASE PAGE 21" in original.pages[20].extract_text()
+ merge_report(base, appendix_a, merged_a, base_pages=19)
+ merge_report(base, appendix_b, merged_b, base_pages=19)
+ assert merged_a.read_bytes() == merged_b.read_bytes()
+
+ merged = PdfReader(merged_a)
+ assert len(merged.pages) == 23
+ for index in range(19):
+ assert f"UNIQUE BASE PAGE {index + 1:02d}" in merged.pages[index].extract_text()
+ assert page_invariants(merged.pages[index]) == original_invariants[index]
+ assert _annotation_uris(merged.pages[index]) == [
+ f"https://example.test/base/{index + 1}"
+ ]
+ generated_text = [merged.pages[index].extract_text() for index in range(19, 23)]
+ assert all("UNIQUE BASE PAGE 20" not in text for text in generated_text)
+ assert all("UNIQUE BASE PAGE 21" not in text for text in generated_text)
+ assert "APPENDIX L" in generated_text[0]
+ assert "APPENDIX L" in generated_text[1]
+ assert "APPENDIX M" in generated_text[2]
+ assert "APPENDIX M" in generated_text[3]
+ assert "Training-free choice-token readout" in generated_text[2]
+ assert "Choice-token accuracy matrix" in generated_text[3]
+ merged_metadata = merged.metadata
+ assert merged_metadata["/CreationDate"] == "D:20261002000000Z"
+ assert merged_metadata["/ModDate"] == "D:20261002000000Z"
+ assert merged_metadata["/Title"] == (
+ "JevAny: Toward General Decision Intelligence โ merged technical report"
+ )
+
+ shutil.copyfile(base, inplace)
+ merge_report(inplace, appendix_a, inplace, base_pages=19)
+ first_inplace = inplace.read_bytes()
+ assert len(PdfReader(inplace).pages) == 23
+ merge_report(inplace, appendix_a, inplace, base_pages=19)
+
+ assert inplace.read_bytes() == first_inplace == merged_a.read_bytes()
+
+
+def test_merge_report_rejects_any_base_page_count_other_than_19(
+ appendix_inputs, tmp_path
+):
+ external, chart, choice = appendix_inputs
+ appendix = tmp_path / "appendices.pdf"
+ base = tmp_path / "base.pdf"
+ build_appendix(external, chart, choice, appendix)
+ writer = PdfWriter()
+ writer.add_blank_page(width=612, height=792)
+ with base.open("wb") as stream:
+ writer.write(stream)
+
+ with pytest.raises(ValueError, match="exactly 19"):
+ merge_report(base, appendix, tmp_path / "merged.pdf", base_pages=18)
+
+
+def test_choice_artifact_requires_full_3220_to_be_unsupported(tmp_path):
+ choice = tmp_path / "choice.json"
+ value = _choice_artifact(choice)
+ value["datasets"]["jevjudge_full"]["accuracy"] = 0.5
+ _write_json(choice, value)
+
+ with pytest.raises(ValueError, match="must be explicitly unsupported"):
+ load_choice_artifact(choice)
+
+
+def test_published_report_uses_one_current_27b_release():
+ report = Path(__file__).resolve().parents[1] / "reports" / "JevAny_Tech_Report.pdf"
+ pdf = PdfReader(report)
+ assert len(pdf.pages) == 23
+
+ expected = {
+ 1: ("86.04", "90.04"),
+ 4: ("44,319", "86.04", "90.04", "0.388", "0.195", "0.026"),
+ 7: ("86.04",),
+ 8: ("44,319", "39.43", "1,261.7", "2,082"),
+ }
+ stale = (
+ "85.76", "90.48", "22,160", "18.83", "602.7", "1,423",
+ "0.392", "0.200", "0.030",
+ )
+ for page_number, values in expected.items():
+ text = pdf.pages[page_number - 1].extract_text() or ""
+ assert all(value in text for value in values)
+ assert all(value not in text for value in stale)
+ compute_text = pdf.pages[7].extract_text() or ""
+ assert "estimated cumulative seconds-per-record timing" in compute_text
+ boundary_text = (pdf.pages[9].extract_text() or "").replace("-\n", "-")
+ assert "The merged PDF is release-synchronized." in boundary_text
+ assert "The original PDF is pre-" not in boundary_text
+
+ method_text = pdf.pages[21].extract_text() or ""
+ assert "current 27B release checkpoint: step 44,319" in method_text
+ assert "matched prompt-v2/runtime reruns" in method_text
diff --git a/tests/test_evaluate_choice_readout_matrix.py b/tests/test_evaluate_choice_readout_matrix.py
new file mode 100644
index 0000000..4dc48c5
--- /dev/null
+++ b/tests/test_evaluate_choice_readout_matrix.py
@@ -0,0 +1,160 @@
+import json
+import math
+
+import pytest
+
+from scripts.evaluate_choice_readout_matrix import (
+ _attach_soft_gold,
+ _typed_record,
+ _typed_soft_metrics,
+ evaluate,
+ reuse_completed,
+)
+
+
+def test_typed_record_uses_ordered_argmax_when_soft_gold_is_tied():
+ row = {
+ "id": "tie-case",
+ "split": "test",
+ "workflow": "unit",
+ "state": json.dumps({"context": "A tied teacher distribution"}),
+ "questions": json.dumps({
+ "decision": {
+ "type": "choice",
+ "criteria": {"first": "First option", "second": "Second option"},
+ }
+ }),
+ # The convenience label intentionally disagrees with the benchmark's
+ # ordered argmax rule and must not determine the converted target.
+ "gold": json.dumps({
+ "decision": {
+ "label": "second",
+ "probabilities": {"first": 0.5, "second": 0.5},
+ }
+ }),
+ }
+
+ record = _typed_record(row)
+
+ assert record["questions"]["decision"]["label"] == "first"
+ assert record["questions"]["decision"]["target"] == {
+ "first": 0.5,
+ "second": 0.5,
+ }
+
+
+def test_typed_soft_metrics_normalize_distributions_and_use_hard_labels():
+ rows = [
+ {"p": [8.0, 2.0], "gold": [3.0, 1.0], "label": 0},
+ {"p": [4.0, 6.0], "gold": [1.0, 1.0], "label": 0},
+ ]
+
+ result = _typed_soft_metrics(rows, bins=2)
+
+ expected_kl = (
+ 0.75 * math.log(0.75 / 0.8)
+ + 0.25 * math.log(0.25 / 0.2)
+ + 0.5 * math.log(0.5 / 0.4)
+ + 0.5 * math.log(0.5 / 0.6)
+ ) / 2
+ assert result["n"] == 2
+ assert result["correct"] == 1
+ assert result["accuracy"] == pytest.approx(0.5)
+ assert result["kl_from_gold"] == pytest.approx(expected_kl)
+ assert result["brier"] == pytest.approx(0.0125)
+ assert result["ece"] == pytest.approx(0.2)
+ assert result["mean_confidence"] == pytest.approx(0.7)
+
+
+def test_attach_soft_gold_aligns_values_to_each_prediction_rows_key_order():
+ record = {
+ "questions": {
+ "choice": {"target": {"alpha": 0.1, "beta": 0.9}},
+ "score": {"target": {"0": 0.2, "1": 0.3, "2": 0.5}},
+ }
+ }
+ rows = [
+ {"question": "choice", "keys": ["beta", "alpha"]},
+ {"question": "score", "keys": ["2", "0", "1"]},
+ ]
+
+ _attach_soft_gold(rows, record)
+
+ assert rows[0]["gold"] == [0.9, 0.1]
+ assert rows[1]["gold"] == [0.5, 0.2, 0.3]
+
+
+def test_evaluate_base_only_result_saves_choice_component(tmp_path):
+ record = {
+ "state": "Choose the right option.",
+ "questions": {
+ "decision": {
+ "type": "choice",
+ "criteria": {"wrong": "Wrong", "right": "Right"},
+ "label": "right",
+ "target": {"right": 0.8, "wrong": 0.2},
+ "src": "unit/base-only",
+ }
+ },
+ "_meta": {
+ "id": "base-only",
+ "group_id": "base-only",
+ "source": "unit",
+ "variant": "clean",
+ },
+ }
+
+ def predictor(_record):
+ return {
+ # Reverse insertion order relative to criteria to exercise keyed
+ # probability alignment in the full evaluation path.
+ "probabilities": {"decision": {"right": 0.9, "wrong": 0.1}},
+ "latency_ms": 2.0,
+ "input_tokens": 12,
+ }
+
+ destination = tmp_path / "base"
+ report = evaluate("typed_test", [record], predictor, destination)
+
+ assert set(report["components"]) == {"choice"}
+ assert report["components"]["choice"]["clean"]["acc"] == pytest.approx(1.0)
+ assert not (destination / "native-rows.json").exists()
+ rows = json.loads((destination / "choice-rows.json").read_text())
+ assert rows[0]["keys"] == ["wrong", "right"]
+ assert rows[0]["p"] == pytest.approx([0.1, 0.9])
+ assert rows[0]["gold"] == pytest.approx([0.2, 0.8])
+ prediction = json.loads(
+ (destination / "predictions.jsonl").read_text().strip()
+ )
+ assert set(prediction["component_probabilities"]) == {"choice"}
+
+
+def test_reuse_completed_validates_report_and_component_rows(tmp_path):
+ destination = tmp_path / "typed_test"
+ record = {
+ "state": "state",
+ "questions": {"q": {
+ "type": "choice", "criteria": {"a": None, "b": None},
+ "label": "a", "target": {"a": .75, "b": .25}, "src": "unit/reuse",
+ }},
+ "_meta": {
+ "id": "reuse", "group_id": "reuse", "source": "unit", "variant": "clean",
+ },
+ }
+
+ def predictor(_record):
+ return {
+ "probabilities": {"q": {"a": .75, "b": .25}},
+ "latency_ms": 1.0,
+ "input_tokens": 4,
+ "efficient_long_context_attention_used": False,
+ }
+
+ expected = evaluate("typed_test", [record], predictor, destination)
+ owner = type("Predictor", (), {"return_components": False})()
+
+ assert reuse_completed("typed_test", [record], owner, destination) == expected
+
+ (destination / "choice-rows.json").unlink()
+ with pytest.raises(RuntimeError, match="missing or incomplete"):
+ reuse_completed("typed_test", [record], owner, destination)
diff --git a/tests/test_fit_choice_ensemble.py b/tests/test_fit_choice_ensemble.py
new file mode 100644
index 0000000..4e54d9a
--- /dev/null
+++ b/tests/test_fit_choice_ensemble.py
@@ -0,0 +1,65 @@
+import pytest
+
+from scripts.fit_choice_ensemble import (
+ fit,
+ headline_indices,
+ metrics,
+ pooled_logits,
+ probabilities,
+)
+
+
+def row(identity, p, label, gold=None):
+ value = {
+ "id": identity,
+ "question": "decision",
+ "keys": ["left", "right"],
+ "label": label,
+ "p": p,
+ }
+ if gold is not None:
+ value["gold"] = gold
+ return value
+
+
+def test_fit_selects_native_weight_on_calibration_accuracy_before_nll():
+ choice = [row("a", [.9, .1], 1), row("b", [.8, .2], 1)]
+ native = [row("a", [.1, .9], 1), row("b", [.2, .8], 1)]
+
+ result = fit(choice, choice, native)
+
+ assert result["selected"]["correct"] == 2
+ assert result["selected"]["native_weight"] > .5
+ assert len(result["grid"]) == 101
+
+
+def test_soft_gold_metrics_use_ordered_labels_and_report_kl():
+ rows = [row("a", [.75, .25], 0, gold=[.5, .5])]
+ ps = probabilities(pooled_logits(rows, None, 0), 1.0)
+
+ result = metrics(rows, ps)
+
+ assert result["accuracy"] == 1.0
+ assert result["brier"] == pytest.approx(.125)
+ assert result["kl_from_gold"] == pytest.approx(
+ result["cross_entropy_from_gold"] - result["gold_entropy"]
+ )
+
+
+def test_choice_only_fit_has_no_native_weight():
+ rows = [row("a", [.7, .3], 0), row("b", [.4, .6], 1)]
+
+ result = fit(rows, rows, None)
+
+ assert result["selected"]["native_weight"] == 0.0
+ assert len(result["grid"]) == 1
+
+
+def test_transfer_selection_uses_only_clean_knowable_headline_rows():
+ rows = [
+ {**row("headline", [.7, .3], 0), "variant": "clean", "source": "task"},
+ {**row("variant", [.7, .3], 0), "variant": "permuted", "source": "task"},
+ {**row("unknown", [.7, .3], 0), "variant": "clean", "source": "unknowable"},
+ ]
+
+ assert headline_indices(rows) == [0]
diff --git a/tests/test_letter_predictor.py b/tests/test_letter_predictor.py
index 85cb3d3..19af8e5 100644
--- a/tests/test_letter_predictor.py
+++ b/tests/test_letter_predictor.py
@@ -26,6 +26,25 @@ def test_letter_predictor_rejects_native_only_load_options_before_loading():
)
+def test_letter_predictor_validates_generalized_native_arguments_before_loading():
+ with pytest.raises(ValueError, match="at most one"):
+ LetterReadoutPredictor(
+ checkpoint="unused", device="cpu", native_weight=0.5, pointer_weight=0.5,
+ )
+ with pytest.raises(ValueError, match="requires a JevAny checkpoint"):
+ LetterReadoutPredictor(base="unused", device="cpu", native_weight=0.5)
+ with pytest.raises(TypeError, match="return_components"):
+ LetterReadoutPredictor(checkpoint="unused", device="cpu", return_components=1)
+ with pytest.raises(TypeError, match="exact_kernels"):
+ LetterReadoutPredictor(checkpoint="unused", device="cpu", exact_kernels=1)
+ with pytest.raises(ValueError, match="pointer weight"):
+ LetterReadoutPredictor(checkpoint="unused", device="cpu", native_weight=1.01)
+ with pytest.raises(ValueError, match="efficient_long_context_tokens"):
+ LetterReadoutPredictor(
+ checkpoint="unused", device="cpu", efficient_long_context_tokens=1,
+ )
+
+
def test_one_row_input_ids_normalizes_template_shapes_and_mappings():
assert _one_row_input_ids(torch.tensor([1, 2])).tolist() == [[1, 2]]
assert _one_row_input_ids({"input_ids": [[3, 4]]}).tolist() == [[3, 4]]
@@ -90,6 +109,16 @@ def test_safe_chat_text_escapes_gemma_and_bracket_style_controls():
assert "๏ผปINST]" in escaped
+def test_long_context_attention_threshold_is_cuda_only():
+ predictor = object.__new__(LetterReadoutPredictor)
+ predictor.efficient_long_context_tokens = 16
+ predictor.device = "cpu"
+ assert not predictor._uses_efficient_attention(16)
+ predictor.device = "cuda"
+ assert not predictor._uses_efficient_attention(15)
+ assert predictor._uses_efficient_attention(16)
+
+
def test_release_unused_output_head_drops_pointer_adapter_reference():
head = object()
model = type("Model", (), {
@@ -113,6 +142,56 @@ def test_release_unused_output_head_drops_lm_token_aliases():
assert _release_unused_output_head(direct_token) is False
+def test_release_unused_output_head_retains_direct_token_native_head():
+ head = object()
+ direct_token = type("Model", (), {
+ "adapter": type("Adapter", (), {"_output_embeddings": head})(),
+ "lm_head": head,
+ })()
+
+ assert _release_unused_output_head(direct_token, retain_native=True) is False
+ assert direct_token.lm_head is head
+ assert direct_token.adapter._output_embeddings is head
+
+
+def test_return_components_exposes_unblended_choice_and_native_probabilities():
+ predictor = object.__new__(LetterReadoutPredictor)
+ predictor.device = "cpu"
+ predictor.temperature = 1.0
+ predictor.native_weight = 0.5
+ predictor.pointer_weight = 0.5
+ predictor.return_components = True
+ predictor._needs_native = True
+ predictor.native_decision_mode = "lm_token"
+ predictor.provenance = {
+ "method": "exact option-letter choice projection",
+ "adapter_applied": True,
+ }
+ predictor._letter_question = lambda state, question: (
+ ["left", "right"], [0.8, 0.2], [4.0, 1.0], 11,
+ )
+ predictor._native = lambda record: ([[0.2, 0.8]], 7)
+ record = {
+ "state": "state",
+ "questions": {"q": {
+ "type": "choice",
+ "instructions": "choose",
+ "criteria": {"left": "Left", "right": "Right"},
+ "label": "left",
+ }},
+ }
+
+ result = predictor(record)
+
+ assert result["component_probabilities"] == {
+ "choice": {"q": {"left": 0.8, "right": 0.2}},
+ "native": {"q": {"left": 0.2, "right": 0.8}},
+ }
+ assert result["probabilities"]["q"] == pytest.approx({"left": 0.5, "right": 0.5})
+ assert result["input_tokens"] == 18
+ assert result["readout"]["native_decision_mode"] == "lm_token"
+
+
def test_question_options_preserve_choice_and_score_order_and_fix_noul_order():
assert question_options({
"type": "choice",
diff --git a/tests/test_letter_readout.py b/tests/test_letter_readout.py
index 9553ee9..c3dd823 100644
--- a/tests/test_letter_readout.py
+++ b/tests/test_letter_readout.py
@@ -4,6 +4,8 @@
import pytest
from jevany.letter_readout import (
+ CHOICE_SYMBOLS,
+ CYGNET_LETTERS,
LETTERS,
SYSTEM_PROMPT,
build_prompt,
@@ -45,21 +47,26 @@ def test_build_prompt_matches_cygnet_structured_rendering_and_option_order():
"A. Refund the card",
"B. Issue store credit",
"",
- "Answer with the letter of exactly one option, and nothing else:",
+ "Answer with the case-sensitive ID of exactly one option, and nothing else:",
])
assert prompt == expected
assert prompt.index("A. Refund") < prompt.index("B. Issue")
assert "thinking" not in prompt.lower()
- assert "LETTER and nothing else" in SYSTEM_PROMPT
+ assert "case-sensitive ID and nothing else" in SYSTEM_PROMPT
-def test_prompt_accepts_all_letters_and_rejects_invalid_option_counts():
- prompt = build_prompt("state", "choose", [f"option {index}" for index in range(26)])
- assert f"{LETTERS[-1]}. option 25" in prompt
+def test_prompt_preserves_cygnet_az_then_extends_to_lowercase_ids():
+ prompt = build_prompt(
+ "state", "choose", [f"option {index}" for index in range(len(CHOICE_SYMBOLS))]
+ )
+ assert LETTERS == CHOICE_SYMBOLS
+ assert f"{CYGNET_LETTERS[-1]}. option 25" in prompt
+ assert f"{CHOICE_SYMBOLS[26]}. option 26" in prompt
+ assert f"{CHOICE_SYMBOLS[-1]}. option 51" in prompt
with pytest.raises(ValueError, match="at least one"):
build_prompt("state", "choose", [])
- with pytest.raises(ValueError, match="at most 26"):
- build_prompt("state", "choose", list(range(27)))
+ with pytest.raises(ValueError, match="at most 52"):
+ build_prompt("state", "choose", list(range(53)))
with pytest.raises(TypeError, match="ordered mapping"):
build_prompt("state", "choose", "not an option sequence")
@@ -161,8 +168,8 @@ def test_invalid_probability_distributions_are_rejected(probabilities):
def test_invalid_letters_token_maps_and_logits_are_rejected():
tokenizer = FakeTokenizer(["A", "not B"])
- with pytest.raises(ValueError, match="only single uppercase"):
- letter_token_ids(tokenizer, ("A", "b"))
+ with pytest.raises(ValueError, match="single A-Z/a-z"):
+ letter_token_ids(tokenizer, ("A", "1"))
with pytest.raises(ValueError, match="duplicates"):
letter_token_ids(tokenizer, ("A", "A"))
with pytest.raises(ValueError, match="no one-token aliases for: B"):
diff --git a/tests/test_playground_workflow.py b/tests/test_playground_workflow.py
index 46cbf36..9058463 100644
--- a/tests/test_playground_workflow.py
+++ b/tests/test_playground_workflow.py
@@ -140,12 +140,19 @@ def test_connecting_reports_identity_and_capabilities_from_v1_models(app, endpoi
app.set_images(True)
-def test_playground_displays_letter_deployment_readout(app, endpoint):
+def test_playground_displays_choice_deployment_readout(app, endpoint):
+ endpoint.models[0]["readout"] = "choice"
+ assert app.connect(endpoint.url)["served"]["readout"] == "choice"
+ source = (Path(__file__).parents[1] / "jevany" / "demos" / "static" / "app.js").read_text()
+ assert 'served.readout === "choice" || served.readout === "letter"' in source
+ assert 'readout ? `${readout} readout`' in source
+
+
+def test_playground_labels_legacy_letter_descriptor_as_choice(app, endpoint):
endpoint.models[0]["readout"] = "letter"
assert app.connect(endpoint.url)["served"]["readout"] == "letter"
source = (Path(__file__).parents[1] / "jevany" / "demos" / "static" / "app.js").read_text()
- assert 'served.readout === "letter" ? "letter"' in source
- assert 'readout ? `${readout} readout`' in source
+ assert 'served.readout === "choice" || served.readout === "letter"' in source
def test_playground_accepts_older_model_descriptors_without_readout(app, endpoint):
diff --git a/tests/test_readout_integration.py b/tests/test_readout_integration.py
index 6f8efcf..dee9c36 100644
--- a/tests/test_readout_integration.py
+++ b/tests/test_readout_integration.py
@@ -5,12 +5,14 @@
import pytest
+from jevany import ChoiceReadoutOptions
from jevany.api import Choice, Noul, Score, SystemOneRequest
from jevany.inference import InferenceOptions
from jevany.letter_runtime import LetterDecisionRuntime, _unlabelled_record
from jevany.readout import (
LetterReadoutOptions,
add_readout_arguments,
+ choice_options_from_args,
letter_options_from_args,
)
@@ -34,15 +36,18 @@ def __init__(self):
},
device_map=None, devices=["cpu"], backbone_adapter="test", temperature=1.2,
)
+ self.native_model = self.pointer_model
self.device = "cpu"
self.temperature = 1.0
self.pointer_weight = 0.25
+ self.native_weight = 0.25
self.max_tokens = 2048
self.effective_max_tokens = 2048
self.tokenizer = FakeTokenizer()
self.provenance = {
"method": "exact option-letter alias projection",
"adapter_applied": True,
+ "native_weight": 0.25,
"pointer_weight": 0.25,
}
self.last_record = None
@@ -72,25 +77,38 @@ def decision_request():
)
-def test_letter_cli_options_are_shared_and_native_rejects_letter_flags():
+def test_choice_cli_is_public_and_legacy_letter_flags_remain_compatible():
parser = argparse.ArgumentParser()
add_readout_arguments(parser)
- assert letter_options_from_args(parser.parse_args([])) is None
- actual = letter_options_from_args(parser.parse_args([
+ help_text = parser.format_help()
+ assert "--readout {native,choice}" in help_text
+ assert "--choice-native-weight" in help_text
+ assert "--letter-temperature" not in help_text
+ assert choice_options_from_args(parser.parse_args([])) is None
+ actual = choice_options_from_args(parser.parse_args([
+ "--readout", "choice", "--choice-temperature", "1.5",
+ "--choice-native-weight", "0.25", "--choice-max-tokens", "2048",
+ ]))
+ assert actual == ChoiceReadoutOptions(temperature=1.5, native_weight=0.25, max_tokens=2048)
+ legacy = letter_options_from_args(parser.parse_args([
"--readout", "letter", "--letter-temperature", "1.5",
"--letter-pointer-weight", "0.25", "--letter-max-tokens", "2048",
]))
- assert actual == LetterReadoutOptions(temperature=1.5, pointer_weight=0.25, max_tokens=2048)
- with pytest.raises(ValueError, match="require --readout letter"):
- letter_options_from_args(parser.parse_args(["--letter-temperature", "2"]))
+ assert legacy == LetterReadoutOptions(temperature=1.5, pointer_weight=0.25, max_tokens=2048)
+ with pytest.raises(ValueError, match="require --readout choice"):
+ choice_options_from_args(parser.parse_args(["--letter-temperature", "2"]))
+ with pytest.raises(ValueError, match="at most one"):
+ choice_options_from_args(parser.parse_args([
+ "--readout", "choice", "--choice-temperature", "1", "--letter-temperature", "1",
+ ]))
for arguments, message in [
- (["--readout", "letter", "--letter-temperature", "nan"], "finite"),
- (["--readout", "letter", "--letter-temperature", "0"], "positive"),
- (["--readout", "letter", "--letter-pointer-weight", "1.1"], r"\[0, 1\]"),
- (["--readout", "letter", "--letter-max-tokens", "1"], ">= 2"),
+ (["--readout", "choice", "--choice-temperature", "nan"], "finite"),
+ (["--readout", "choice", "--choice-temperature", "0"], "positive"),
+ (["--readout", "choice", "--choice-native-weight", "1.1"], r"\[0, 1\]"),
+ (["--readout", "choice", "--choice-max-tokens", "1"], ">= 2"),
]:
with pytest.raises(ValueError, match=message):
- letter_options_from_args(parser.parse_args(arguments))
+ choice_options_from_args(parser.parse_args(arguments))
def test_letter_runtime_maps_requests_and_reports_effective_readout(decision_request):
@@ -107,12 +125,14 @@ def test_letter_runtime_maps_requests_and_reports_effective_readout(decision_req
assert predictor.last_record["state"] == {"signal": "green"}
assert [q["label"] for q in predictor.last_record["questions"].values()] == ["left", False, 0]
description = runtime.describe()
- assert description["readout"] == "letter"
- assert description["letter_readout"]["pointer_weight"] == 0.25
- assert description["letter_readout"]["pointer_temperature"] == 1.2
+ assert description["readout"] == "choice"
+ assert description["choice_readout"]["native_weight"] == 0.25
+ assert description["choice_readout"]["pointer_weight"] == 0.25
+ assert description["choice_readout"]["native_temperature"] == 1.2
+ assert description["letter_readout"] == description["choice_readout"]
assert description["limits"] == {
"state_tokens": 2048, "branch_tokens": 2048,
- "packed_tokens": 2048, "choices": 26,
+ "packed_tokens": 2048, "choices": 52,
}
assert description["capabilities"]["media_types"] == []
assert not description["prefix_cache"]["enabled"]
@@ -137,7 +157,7 @@ def test_unlabelled_record_adds_only_encoder_placeholders(decision_request):
assert record["questions"]["score"]["label"] == 0
-def test_public_loader_dispatches_to_letter_predictor(monkeypatch):
+def test_public_loader_dispatches_choice_settings_to_predictor(monkeypatch):
from jevany import letter_predictor
from jevany.runtime import JevModel
@@ -152,34 +172,61 @@ def load(**kwargs):
monkeypatch.setattr(letter_predictor, "LetterReadoutPredictor", load)
local = JevModel.from_pretrained(
"owner/checkpoint", device="cpu", model_name="letter-model",
- readout="letter", letter_temperature=1.5,
- letter_pointer_weight=0.25, letter_max_tokens=2048,
+ readout="choice", choice_temperature=1.5,
+ choice_native_weight=0.25, choice_max_tokens=2048,
)
assert isinstance(local.runtime, LetterDecisionRuntime)
assert seen["checkpoint"] == "owner/checkpoint"
assert seen["temperature"] == 1.5
- assert seen["pointer_weight"] == 0.25
+ assert seen["native_weight"] == 0.25
assert seen["max_tokens"] == 2048
- with pytest.raises(ValueError, match="require readout='letter'"):
- JevModel.from_pretrained("unused", device="cpu", letter_temperature=2)
+ with pytest.raises(ValueError, match="require --readout choice"):
+ JevModel.from_pretrained("unused", device="cpu", choice_temperature=2)
with pytest.raises(ValueError, match="only to native readout"):
JevModel.from_pretrained(
- "unused", device="cpu", readout="letter",
+ "unused", device="cpu", readout="choice",
inference_options=InferenceOptions(),
)
+def test_public_loader_accepts_legacy_letter_api(monkeypatch):
+ from jevany import letter_predictor
+ from jevany.runtime import JevModel
+
+ predictor = FakeLetterPredictor()
+ predictor.checkpoint.path = "/tmp/checkpoint"
+ seen = {}
+
+ def load(**kwargs):
+ seen.update(kwargs)
+ return predictor
+
+ monkeypatch.setattr(letter_predictor, "LetterReadoutPredictor", load)
+ local = JevModel.from_pretrained(
+ "owner/checkpoint", device="cpu", model_name="letter-model",
+ readout="letter", letter_temperature=1.5,
+ letter_pointer_weight=0.25, letter_max_tokens=2048,
+ )
+
+ assert isinstance(local.runtime, LetterDecisionRuntime)
+ assert seen["temperature"] == 1.5
+ assert seen["native_weight"] == 0.25
+ assert seen["max_tokens"] == 2048
+
+
def test_create_app_rejects_cross_readout_settings():
pytest.importorskip("fastapi")
from jevany.serve import create_app
- with pytest.raises(ValueError, match="require readout='letter'"):
- create_app(readout="native", letter_options=LetterReadoutOptions())
+ with pytest.raises(ValueError, match="require readout='choice'"):
+ create_app(readout="native", choice_options=ChoiceReadoutOptions())
with pytest.raises(ValueError, match="only to native readout"):
- create_app(readout="letter", inference_options=InferenceOptions())
+ create_app(readout="choice", inference_options=InferenceOptions())
+ # The old programmatic spelling still builds the same application path.
+ assert create_app(readout="letter", letter_options=LetterReadoutOptions())
-def test_decide_cli_forwards_letter_settings(tmp_path, monkeypatch, capsys):
+def test_decide_cli_forwards_choice_settings(tmp_path, monkeypatch, capsys):
from jevany.cli import decide_main
from jevany.runtime import JevModel
@@ -207,21 +254,21 @@ def load(*args, **kwargs):
monkeypatch.setattr(JevModel, "from_pretrained", load)
decide_main([
str(source), "--checkpoint", "owner/checkpoint", "--device", "cpu",
- "--readout", "letter", "--letter-temperature", "1.5",
- "--letter-pointer-weight", "0.25", "--letter-max-tokens", "2048",
+ "--readout", "choice", "--choice-temperature", "1.5",
+ "--choice-native-weight", "0.25", "--choice-max-tokens", "2048",
])
assert json.loads(capsys.readouterr().out)["answers"]["choice"]["choice"] == "right"
assert seen["args"] == ("owner/checkpoint",)
- assert seen["kwargs"]["readout"] == "letter"
- assert seen["kwargs"]["letter_temperature"] == 1.5
- assert seen["kwargs"]["letter_pointer_weight"] == 0.25
- assert seen["kwargs"]["letter_max_tokens"] == 2048
+ assert seen["kwargs"]["readout"] == "choice"
+ assert seen["kwargs"]["choice_temperature"] == 1.5
+ assert seen["kwargs"]["choice_native_weight"] == 0.25
+ assert seen["kwargs"]["choice_max_tokens"] == 2048
assert "inference_options" not in seen["kwargs"]
with pytest.raises(SystemExit):
decide_main([str(source), "--letter-temperature", "2"])
-def test_eval_cli_selects_letter_predictor_and_records_provenance(tmp_path, monkeypatch, capsys):
+def test_eval_cli_selects_choice_predictor_and_records_provenance(tmp_path, monkeypatch, capsys):
from jevany import benchmark, letter_predictor
data = tmp_path / "data.jsonl"
@@ -237,8 +284,9 @@ def test_eval_cli_selects_letter_predictor_and_records_provenance(tmp_path, monk
class Predictor:
temperature = 1.5
+ native_weight = 0.25
pointer_weight = 0.25
- pointer_model = SimpleNamespace(temperature=1.2)
+ native_model = SimpleNamespace(temperature=1.2)
provenance = {"method": "exact option-letter alias projection"}
def __init__(self, **kwargs):
@@ -255,18 +303,20 @@ def evaluate(records, predictor, directory, **kwargs):
monkeypatch.setattr(letter_predictor, "LetterReadoutPredictor", Predictor)
benchmark.main([
"--run", "owner/checkpoint", "--data", str(data), "--out", str(output),
- "--device", "cpu", "--readout", "letter", "--letter-temperature", "1.5",
- "--letter-pointer-weight", "0.25", "--letter-max-tokens", "2048",
+ "--device", "cpu", "--readout", "choice", "--choice-temperature", "1.5",
+ "--choice-native-weight", "0.25", "--choice-max-tokens", "2048",
])
assert json.loads(capsys.readouterr().out)["objective"] == 0.0
assert seen["checkpoint"] == "owner/checkpoint"
assert seen["temperature"] == 1.5
- assert seen["pointer_weight"] == 0.25
+ assert seen["native_weight"] == 0.25
assert seen["max_tokens"] == 2048
report = json.loads((output / "report.json").read_text())
- assert report["readout"] == "letter"
- assert report["letter_readout"] == {
- **Predictor.provenance, "pointer_temperature": 1.2,
+ assert report["readout"] == "choice"
+ assert report["choice_readout"] == {
+ **Predictor.provenance, "native_temperature": 1.2, "pointer_temperature": 1.2,
}
+ assert report["letter_readout"] == report["choice_readout"]
assert report["calibration_applied"] is True
+ assert report["calibration"]["native_temperature"] == 1.2
assert report["calibration"]["pointer_temperature"] == 1.2