Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
24 changes: 24 additions & 0 deletions .github/workflows/ci.yml
Original file line number Diff line number Diff line change
Expand Up @@ -64,3 +64,27 @@ jobs:
path: |
ci-results.xml
ci-dependencies.txt

report-test:
name: report appendix
runs-on: ubuntu-24.04
steps:
- uses: actions/checkout@v4
- uses: actions/setup-python@v5
with:
python-version: '3.13'
cache: pip
cache-dependency-path: reports/requirements.txt
- name: Install report dependencies
run: python -m pip install --disable-pip-version-check -r reports/requirements.txt pytest==9.1.1
- name: Verify report environment
run: |
python -m pip check
python -c "import matplotlib, pypdf, pytest; print(matplotlib.__version__, pypdf.__version__, pytest.__version__)"
- name: Test report artifact and PDF builders
env:
JEVANY_REPORT_TESTS_REQUIRED: '1'
run: >-
python -m pytest -q
tests/test_build_choice_readout_results.py
tests/test_build_external_report_appendix.py
59 changes: 37 additions & 22 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -140,7 +140,7 @@ where Jev helps and when to return control to the LLM.
- [🛠️ 1.2 JevAny Training](#training)
- [🚀 1.3 JevAny Deployment](#deployment)
- [🤗 2. Pretrained Models](#pretrained-models)
- [Training-free letter readout](#letter-readout)
- [Training-free choice-token readout](#choice-readout)
- [📊 3. Benchmark Results](#evaluation)
- [⏱️ 3.1 Inference efficiency](#efficiency)
- [🕹️ 4. Examples & Test Environments](#examples--test-environments)
Expand Down Expand Up @@ -296,30 +296,43 @@ Pointer and direct-token models share the same API. Pointer supports up to
See [readout choices](docs/TRAINING.md#pointer-and-direct-token-readouts) for
training and accuracy tradeoffs.

### Training-free letter readout <a name="letter-readout"></a>
### Training-free choice-token readout <a name="choice-readout"></a><a name="letter-readout"></a>

The experimental letter readout presents up to 26 options as A–Z, then sums
the frozen language model's next-token probability mass for every vocabulary
token that decodes exactly to that uppercase letter. It can run on a base model
without training, apply a JevAny adapter before the same readout, or combine the
letter and native pointer distributions. The letter path is text-only and does
not change the checkpoint metadata. Select it with `--readout letter` in
`jevany eval`, `jevany decide`, or `jevany serve`; the checkpoint-native
readout remains the default.
Choice-token readout assigns the ordered options to the exact one-token IDs
`A–Z, a–z`, scores those 52 rows at the answer position, and renormalizes over
the available options. It needs no readout training and works on a frozen base,
after a JevAny Pointer or Direct-Token adapter, or in a log-linear stack with
the checkpoint-native distribution. It can change the ranking and the selected
answer; scalar temperature calibration only changes confidence.

[Method, limits and evaluation command](docs/LETTER_READOUT.md).
```bash
jevany eval --run SimpleJev/JevAny-Qwen3.5-4B-Direct-Token-LoRA \
--suite /path/to/suite --out runs/choice --device cuda --readout choice
```

[Method, commands, limits and complete tables](docs/CHOICE_READOUT.md) ·
[Machine-readable results](results/choice-readout-v2.json)

[![Training-free choice-token, checkpoint-native, Transfer-dev-tuned blend, and external baseline accuracy on Typed Decisions and JevJudge text](docs/choice-readout-results.svg)](docs/CHOICE_READOUT.md#results)

- Direct-Token 4B's Transfer-dev-tuned stack reaches **79.83%** on held-out
Transfer, **67.65%** on Typed Decisions, and **59.25%** on JevJudge text:
+0.96, +0.45, and +0.83 points over its native readout.
- Pointer 27B's tuned stack reaches **89.10%** on held-out Transfer and
**73.30%** on Typed, but drops from **66.44% to 64.36%** on JevJudge text.
The target-domain result decides whether stacking is useful.
- The frozen 4B choice path scores 52.75% on Typed; applying the Direct-Token
adapter raises the same choice path to 64.80%. On JevJudge text, the same
adapter slightly lowers it, from 58.70% to 57.60%.

[![Accuracy change from native pointer for letter readout and the fixed blend](docs/letter-readout-results.svg)](docs/LETTER_READOUT.md#results)
**Practical rule**

- The fixed 50/50 blend improved 27B on both diagnostics: +0.43 accuracy
points on JevBench public and +0.67 on Transfer-v9. Neither gain is
statistically significant.
- At 4B, the same blend gained +1.73 points on JevBench public but lost 0.10
on Transfer-v9. Letter readout is a paired target-domain ablation, not a
universal upgrade.
- Letter-only measured a 1.05x/1.07x median speedup over native at 4B/27B
on one warmed H200 panel. Blending runs both paths and cost 1.81x/1.89x the
native median latency.
- Use choice-token readout when the model can reason over the options and the
bottleneck is extracting a decision; it cannot create missing task ability.
- Blend only when paired development errors show that native and choice rescue
each other, then freeze one weight before the target evaluation.
- Use temperature for probability calibration, not accuracy: it cannot change
the argmax.

> **Benchmark units:** Cygnet's official **73.70** on JevBench v1.5.4 is a
> four-axis composite over 1,624 open and sealed decisions, not accuracy. Its
Expand All @@ -329,6 +342,7 @@ readout remains the default.

## 📊 3. Benchmark Results <a name="evaluation"></a>

The release table below uses each checkpoint's native readout.
JevAny-Qwen3.8-27B leads both benchmarks and has the lowest NLL and Brier.
Among 4B releases, direct-token leads on JevBench; pointer leads on Transfer.

Expand Down Expand Up @@ -357,7 +371,8 @@ NLL, Brier and ECE are measured on Transfer.
[Machine-readable results](results/model-family-v2.json) ·
[Method and ablation report](reports/JevAny_Tech_Report.pdf)

The external comparison uses the complete Typed Decisions test split, the full
The external comparison below also uses checkpoint-native readouts. It uses the
complete Typed Decisions test split, the full
3,220-record JevJudge multimodal suite, and its 724-record text slice. The same
13 models stay in the same order; `—` means unsupported native input or no
matching result.
Expand Down
216 changes: 216 additions & 0 deletions docs/CHOICE_READOUT.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,216 @@
# Training-free choice-token readout

Choice-token readout turns an ordinary causal language model into a decision
readout without adding or training a head:

1. Preserve the candidate order and assign the exact one-character IDs
`A–Z, a–z`.
2. Ask for one case-sensitive option ID with thinking disabled.
3. At the answer position, sum the mass of every vocabulary token that decodes
exactly to each available ID.
4. Renormalize over the available IDs and map the distribution back to the
original option keys.

This operation can change the ranking and the selected answer. It is therefore
not temperature calibration. A scalar temperature changes probability values
but cannot change argmax accuracy. A native + choice stack is a third operation:
it pools two distributions and can change the answer.

The same readout works on a frozen base model or after loading a JevAny Pointer
or Direct-Token adapter. No additional readout training is performed.

## Use it

The checkpoint-native readout remains the default. Select choice-token readout
with `--readout choice`:

```bash
jevany eval \
--run SimpleJev/JevAny-Qwen3.5-4B-Direct-Token-LoRA \
--suite /path/to/jevbench-public-v1.4.2.2 \
--out runs/choice-readout/qwen35-4b-direct \
--device cuda --readout choice --choice-temperature 1.0

jevany decide request.json \
--checkpoint SimpleJev/JevAny-Qwen3.5-4B-LoRA \
--device cuda --dtype bf16 --readout choice

jevany serve \
--checkpoint SimpleJev/JevAny-Qwen3.5-4B-LoRA \
--device cuda --dtype bf16 --readout choice
```

Use `--choice-native-weight` to pool the choice and checkpoint-native
distributions. This executes both paths. `--choice-max-tokens` controls the
choice prompt limit. The former `letter` selector and `--letter-*` flags remain
hidden compatibility aliases.

The matrix evaluator also accepts a frozen base directly:

```bash
python -m scripts.evaluate_choice_readout_matrix \
--base Qwen/Qwen3.5-4B \
--revision 851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a \
--base-load-path /path/to/Qwen3.5-4B \
--model-name Frozen-Qwen3.5-4B \
--out runs/choice-readout/base4 \
--jevbench /path/to/jevbench-public-v1.4.2.2 \
--transfer /path/to/transfer-v9 \
--typed /path/to/typed-decisions \
--jevjudge /path/to/jevjudge-public
```

## Limits

- Text input only; JevJudge full multimodal is unsupported and is never
converted into a text-only score.
- 1–52 options per question.
- Exact one-token aliases must exist in the tokenizer for every used ID.
- Choice-only runs use one chat-formatted prefill per question. A stack also
executes the checkpoint-native prefill.
- Matrix runs use exact math SDPA below 8,192 tokens and memory-linear fused
SDPA for longer prompts. All 724 JevJudge text records complete; no record is
truncated or rejected.

## Evaluation protocol

The v2 matrix uses the same prompt and data for five runs:

- frozen Qwen3.5-4B;
- JevAny Qwen3.5-4B Pointer;
- JevAny Qwen3.5-4B Direct-Token;
- frozen Qwen3.8-27B;
- JevAny Qwen3.8-27B Pointer.

Zero-shot rows use T=1 and no evaluation-label tuning. For tuned rows, each
model selects one native weight on the 1,046 clean/knowable Transfer-v9
development decisions; fitted NLL and distance from 0.5 break accuracy ties.
One additional temperature is then fitted on the same development rows. The
weight and temperature are frozen before Transfer test, Typed Decisions,
JevJudge text, and the JevBench public diagnostic are scored.

The native parity gates reproduce the released README results before any
held-out panel runs:

| Checkpoint | Transfer development | JevBench public |
|---|---:|---:|
| JevAny 4B Pointer | 823/1,046 · 78.68% | 185/231 · 80.09% |
| JevAny 4B Direct-Token | 818/1,046 · 78.20% | 187/231 · 80.95% |
| JevAny 27B Pointer | 900/1,046 · 86.04% | 208/231 · 90.04% |

### Evaluated model sources

| Run | Canonical public repository | Evaluated revision |
|---|---|---|
| Frozen Qwen3.5-4B | `Qwen/Qwen3.5-4B` | `851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a` |
| JevAny 4B Pointer | `SimpleJev/JevAny-Qwen3.5-4B-LoRA` | `1c7aa9bab14ac347aeb917c0bcd757838a8a78ce` |
| JevAny 4B Direct-Token | `SimpleJev/JevAny-Qwen3.5-4B-Direct-Token-LoRA` | Not verified for the evaluated local release |
| Frozen Qwen3.8-27B | `Qwen/Qwen3.8-27B` | `1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0` |
| JevAny 27B Pointer | `SimpleJev/JevAny-Qwen3.8-27B-LoRA` | `09c9e9102d5b8cc7d56558d25da1202a761b6c0d` |

The Direct-Token run is still byte-identifiable: its evaluated adapter is
`b768d54b9e5e20204b2db9e24c18ae6b10fa310b241d1931bac6b8882f127c53`
and its native head is
`d3909cffc156d077061114627c8aa22f60c7a8f1bdcfd20c489bb3bb009fc4a6`
(SHA-256). The evaluation manifest and release metadata do not prove which
immutable public-repository commit contains those exact files, so the result
artifact records its public revision as `null` rather than guessing.

## Results

![Training-free choice-token, checkpoint-native, Transfer-dev-tuned blend, and external baseline accuracy on Typed Decisions and JevJudge text](choice-readout-results.svg)

All values below are accuracy. JevBench is its 231-item public development
diagnostic. Transfer test contains 1,046 clean/knowable held-out decisions.
Typed covers 2,000 decisions in 400 test cases. JevJudge text covers all 724
text records; it is not the 3,220-record multimodal suite.

| Weights | Readout | JevBench public | Transfer test | Typed test | JevJudge text |
|---|---|---:|---:|---:|---:|
| Frozen Qwen3.5-4B | Choice T=1 | 79.65% | 68.45% | 52.75% | 58.70% |
| JevAny 4B Pointer | Native shipped | 80.09% | 77.92% | **63.50%** | 58.01% |
| | Choice T=1 | **81.39%** | 73.61% | 57.80% | 52.62% |
| | Tuned stack · w=.78 | 80.95% | **78.01%** | 62.90% | **59.53%** |
| JevAny 4B Direct-Token | Native shipped | 80.95% | 78.87% | 67.20% | 58.43% |
| | Choice T=1 | **81.39%** | 78.59% | 64.80% | 57.60% |
| | Tuned stack · w=.59 | 80.95% | **79.83%** | **67.65%** | **59.25%** |
| Frozen Qwen3.8-27B | Choice T=1 | 87.88% | 80.69% | 68.25% | 61.74% |
| JevAny 27B Pointer | Native shipped | **90.04%** | 87.95% | 72.80% | **66.44%** |
| | Choice T=1 | 89.61% | 86.04% | 72.60% | 57.87% |
| | Tuned stack · w=.52 | **90.04%** | **89.10%** | **73.30%** | 64.36% |

The fixed, untuned 50/50 stacks remain a separate zero-shot reference:

| Checkpoint | JevBench public | Transfer test | Typed test | JevJudge text |
|---|---:|---:|---:|---:|
| JevAny 4B Pointer | 81.39% | 77.44% | 60.70% | 59.12% |
| JevAny 4B Direct-Token | 80.95% | 79.45% | 66.75% | 58.98% |
| JevAny 27B Pointer | 90.04% | **89.20%** | 73.20% | 64.23% |

The selected parameters differ by checkpoint:

| Checkpoint | Native weight | Additional temperature |
|---|---:|---:|
| JevAny 4B Pointer | .78 | 1.10 |
| JevAny 4B Direct-Token | .59 | 1.12 |
| JevAny 27B Pointer | .52 | .84 |

### External comparison

| Model/readout | Typed test | JevJudge text |
|---|---:|---:|
| meraGPT Decider 1 · published | **76.80%** | — |
| JevAny 27B Pointer · native | 72.80% | **66.44%** |
| JevAny 27B Pointer · tuned stack | **73.30%** | 64.36% |
| TypeSafe Jev 1.13 | 72.70% | 65.06% |
| Kev-27B | — | 64.23% |

The 27B stack beats its native readout by 12 held-out Transfer decisions and
10 Typed decisions, while losing 15 JevJudge text decisions. Direct-Token 4B's
stack improves over native on all three held-out panels. Pointer 4B improves on
Transfer and JevJudge but loses on Typed. No single blend is universally best.

### Practical principles

- **Selection bottleneck:** use choice-token readout when the model can reason
over the candidates but the native extraction path is the bottleneck.
- **Complementary errors:** keep a stack only when paired development results
show that each readout rescues errors made by the other. Select one weight on
development data and freeze it.
- **Capability limit:** a different readout cannot supply knowledge, perception,
planning, or instruction-following ability missing from the underlying model.
- **Calibration limit:** temperature can improve NLL, Brier, or ECE, but it
cannot change the selected option or accuracy.
- **Domain shift:** target-domain validation wins over a global rule. The same
27B stack helps Transfer and Typed, ties JevBench, and hurts JevJudge text.

Complete metrics, hashes, runtime versions, selected weights, and the unsupported
full-JevJudge marker are in
[`results/choice-readout-v2.json`](../results/choice-readout-v2.json). The
deterministic builder is
[`scripts/build_choice_readout_results.py`](../scripts/build_choice_readout_results.py).

## Benchmark units

Cygnet's official `73.70` on JevBench v1.5.4 is a four-axis composite over
1,624 open and sealed decisions, not accuracy. Its separate public-development
result is 203/231 (87.9% accuracy). Every JevBench value on this page is
accuracy over the same 231 public-development items and cannot be compared
numerically with the composite.

## Historical v1

The original prompt-v1 experiment and one-panel latency diagnostic remain
unchanged in [`results/letter-readout-v1.json`](../results/letter-readout-v1.json)
and [`letter-readout-results.svg`](letter-readout-results.svg). Those runs used a
different prompt and runtime, and their matched native reruns did not reproduce
the canonical README rows. They are retained for audit and are not mixed into
the v2 tables.

## Attribution

The Cygnet-compatible prompt and exact option-ID aggregation semantics are
adapted from `blockbrain-ai/cygnet-recipe` commit `3cf591c`, Copyright 2026
Nood Co and contributors, under MIT. Cygnet credits the one-token option-ID
readout to [NInfer](https://github.com/igorls/ninfer), released under Apache-2.0.
JevAny does not incorporate NInfer source code. See [`NOTICE`](../NOTICE).
35 changes: 32 additions & 3 deletions docs/EVALUATION.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,7 +2,7 @@

## Model family v2

This section contains the full results for the five current LoRA SFT releases.
This section contains checkpoint-native results for the five current LoRA SFT releases.
Transfer also informed model development and serves as a diagnostic comparison.

**Transfer** is the cross-domain evaluation built from
Expand All @@ -16,7 +16,7 @@ and ECE in the main table are Transfer metrics; every run covers every item.
Every JevBench value in this document is public-development accuracy
(`correct / 231`). It is not the official JevBench v1.5.4 composite, which
combines Intelligence, Calibration, Speed and Cost over 1,624 open and sealed
decisions. See the [letter-readout guide](LETTER_READOUT.md#benchmark-units) for
decisions. See the [choice-token readout guide](CHOICE_READOUT.md#benchmark-units) for
the side-by-side definitions.

| Model | Transfer ↑ | JevBench ↑ | NLL ↓ | Brier ↓ | ECE ↓ |
Expand Down Expand Up @@ -95,7 +95,7 @@ a complete local Transfer API run and JevBench's published per-tier accuracy.
- **Loss:** cross-entropy, pure InfoNCE, and mixed objectives; CE gave the best
Transfer accuracy in the loss sweep, while small contrastive terms mainly
improved calibration. The released checkpoints use CE.
- **Readout family:** at 4B, direct-token improves JevBench (80.95% vs 80.09%),
- **Trained checkpoint-native readout family:** at 4B, direct-token improves JevBench (80.95% vs 80.09%),
while pointer is slightly stronger on Transfer (78.68% vs 78.20%).

### Reproducibility
Expand Down Expand Up @@ -140,6 +140,35 @@ excluded. Gemma uses run/checkpoint timestamps because its earlier checkpoint
format did not store cumulative elapsed seconds; the other figures come from
checkpoint or terminal trainer telemetry.

## Training-free choice-token readout

Choice-token readout constrains the next token to one of 52 exact option IDs
and renormalizes their mass. It needs no additional training and can change the
selected answer. Scalar temperature fitting is separate: it changes confidence,
not accuracy. A native + choice stack can change both.

The stack weight and additional temperature below are selected only on 1,046
clean/knowable Transfer-v9 development decisions and then frozen. Transfer test,
Typed Decisions, and JevJudge text are held out. JevBench is a public diagnostic.
All values are accuracy percentages.

| Weights | Readout | JevBench public | Transfer test | Typed test | JevJudge text |
|---|---|---:|---:|---:|---:|
| Frozen Qwen3.5-4B | Choice T=1 | 79.65% | 68.45% | 52.75% | 58.70% |
| JevAny 4B Pointer | Native / Choice / Tuned | 80.09 / 81.39 / 80.95 | 77.92 / 73.61 / **78.01** | **63.50** / 57.80 / 62.90 | 58.01 / 52.62 / **59.53** |
| JevAny 4B Direct-Token | Native / Choice / Tuned | 80.95 / **81.39** / 80.95 | 78.87 / 78.59 / **79.83** | 67.20 / 64.80 / **67.65** | 58.43 / 57.60 / **59.25** |
| Frozen Qwen3.8-27B | Choice T=1 | 87.88% | 80.69% | 68.25% | 61.74% |
| JevAny 27B Pointer | Native / Choice / Tuned | **90.04** / 89.61 / **90.04** | 87.95 / 86.04 / **89.10** | 72.80 / 72.60 / **73.30** | **66.44** / 57.87 / 64.36 |

The 4B Direct-Token stack improves over its native readout on all three held-out
panels. The 27B stack improves Transfer and Typed, ties JevBench, and hurts
JevJudge text. Retain a stack only when paired development errors are
complementary and the target distribution validates it.

[Full method, fixed 50/50 controls and protocol](CHOICE_READOUT.md) ·
[Machine-readable results](../results/choice-readout-v2.json) ·
[Merged 23-page technical report](../reports/JevAny_Tech_Report.pdf)

## Earlier releases and evaluations

The [JevBench and Kev comparison](EXTERNAL_EVALUATION.md) evaluates both released
Expand Down
Loading
Loading