Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 2 additions & 0 deletions ACKNOWLEDGEMENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -14,6 +14,8 @@ The Phi-4 vision conversion helper adapts configuration and weight-name mappings

The native Phi-4 Reasoning Vision adapter follows Microsoft's [published model layout and image preprocessing](https://huggingface.co/microsoft/Phi-4-reasoning-vision-15B/tree/c3e4fac79ddace21976ced56fbf1564b8bd8c89f), released under MIT. Its weights are downloaded separately from Microsoft.

The experimental training-free option-letter readout adapts prompt and probability-aggregation semantics from [Cygnet](https://github.com/blockbrain-ai/cygnet-recipe) at commit `3cf591c692dec649f7c134449814610307c7bb3a`, Copyright 2026 Nood Co and contributors, under MIT. Cygnet credits the one-token option-letter readout method to [NInfer](https://github.com/igorls/ninfer), released under Apache-2.0. JevAny implements its own model integration and does not incorporate NInfer source code. The full Cygnet notice is in `NOTICE`.

The README teaser uses Lucide icons. The [source records](docs/icons/sources.json) and [license notices](docs/icons/LICENSE) accompany the editable SVG.

The supported-model cards use logos from [Lobe Icons](https://github.com/lobehub/lobe-icons) under MIT. Their [source records](docs/model-logos/sources.json) and [license](docs/model-logos/LICENSE) accompany the SVG. Model and publisher marks belong to their respective owners.
33 changes: 33 additions & 0 deletions NOTICE
Original file line number Diff line number Diff line change
Expand Up @@ -12,6 +12,39 @@ adapted from Hugging Face Transformers, Copyright 2025 The HuggingFace Inc. team
under the Apache License 2.0:
https://github.com/huggingface/transformers

The training-free option-letter prompt and probability aggregation include
logic adapted from Cygnet at commit
3cf591c692dec649f7c134449814610307c7bb3a:
https://github.com/blockbrain-ai/cygnet-recipe

Cygnet is distributed under the MIT License:

Copyright (c) 2026 Nood Co (github.com/blockbrain-ai) and contributors

Permission is hereby granted, free of charge, to any person obtaining a copy
of this software and associated documentation files (the "Software"), to deal
in the Software without restriction, including without limitation the rights
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
copies of the Software, and to permit persons to whom the Software is
furnished to do so, subject to the following conditions:

The above copyright notice and this permission notice shall be included in all
copies or substantial portions of the Software.

THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
SOFTWARE.

Cygnet credits the one-token option-letter readout method to NInfer, released
under the Apache License 2.0:
https://github.com/igorls/ninfer

JevAny does not incorporate NInfer source code.

The released adapters require Qwen3.8-27B by the Qwen team. The base model is
distributed separately under its own Apache License 2.0 terms.

Expand Down
32 changes: 32 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -140,6 +140,7 @@ where Jev helps and when to return control to the LLM.
- [🛠️ 1.2 JevAny Training](#training)
- [🚀 1.3 JevAny Deployment](#deployment)
- [🤗 2. Pretrained Models](#pretrained-models)
- [Training-free letter readout](#letter-readout)
- [📊 3. Benchmark Results](#evaluation)
- [⏱️ 3.1 Inference efficiency](#efficiency)
- [🕹️ 4. Examples & Test Environments](#examples--test-environments)
Expand Down Expand Up @@ -295,6 +296,37 @@ Pointer and direct-token models share the same API. Pointer supports up to
See [readout choices](docs/TRAINING.md#pointer-and-direct-token-readouts) for
training and accuracy tradeoffs.

### Training-free letter readout <a name="letter-readout"></a>

The experimental letter readout presents up to 26 options as A–Z, then sums
the frozen language model's next-token probability mass for every vocabulary
token that decodes exactly to that uppercase letter. It can run on a base model
without training, apply a JevAny adapter before the same readout, or combine the
letter and native pointer distributions. The letter path is text-only and does
not change the checkpoint metadata. Select it with `--readout letter` in
`jevany eval`, `jevany decide`, or `jevany serve`; the checkpoint-native
readout remains the default.

[Method, limits and evaluation command](docs/LETTER_READOUT.md).

[![Accuracy change from native pointer for letter readout and the fixed blend](docs/letter-readout-results.svg)](docs/LETTER_READOUT.md#results)

- The fixed 50/50 blend improved 27B on both diagnostics: +0.43 accuracy
points on JevBench public and +0.67 on Transfer-v9. Neither gain is
statistically significant.
- At 4B, the same blend gained +1.73 points on JevBench public but lost 0.10
on Transfer-v9. Letter readout is a paired target-domain ablation, not a
universal upgrade.
- Letter-only measured a 1.05x/1.07x median speedup over native at 4B/27B
on one warmed H200 panel. Blending runs both paths and cost 1.81x/1.89x the
native median latency.

> **Benchmark units:** Cygnet's official **73.70** on JevBench v1.5.4 is a
> four-axis composite over 1,624 open and sealed decisions, not accuracy. Its
> separate public-development result is **203/231 (87.9% accuracy)**. The
> JevBench values in this section are also accuracy on those 231 public
> development items, so they must not be compared directly with 73.70.

## 📊 3. Benchmark Results <a name="evaluation"></a>

JevAny-Qwen3.8-27B leads both benchmarks and has the lowest NLL and Brier.
Expand Down
6 changes: 6 additions & 0 deletions docs/EVALUATION.md
Original file line number Diff line number Diff line change
Expand Up @@ -13,6 +13,12 @@ MMLU, MMLU-Pro, SciQ, and four robustness slices. **JevBench** is accuracy acros
all 231 public development items. NLL, Brier,
and ECE in the main table are Transfer metrics; every run covers every item.

Every JevBench value in this document is public-development accuracy
(`correct / 231`). It is not the official JevBench v1.5.4 composite, which
combines Intelligence, Calibration, Speed and Cost over 1,624 open and sealed
decisions. See the [letter-readout guide](LETTER_READOUT.md#benchmark-units) for
the side-by-side definitions.

| Model | Transfer ↑ | JevBench ↑ | NLL ↓ | Brier ↓ | ECE ↓ |
|---|---:|---:|---:|---:|---:|
| Kev-4B | 74.19% | 75.32% | 0.858 | 0.380 | 0.125 |
Expand Down
170 changes: 170 additions & 0 deletions docs/LETTER_READOUT.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,170 @@
# Training-free option-letter readout

JevAny includes an experimental evaluator for a training-free, single-answer-slot
decision readout. It adapts the prompt and probability-aggregation semantics
from [Cygnet](https://github.com/blockbrain-ai/cygnet-recipe/tree/3cf591c692dec649f7c134449814610307c7bb3a):

1. Render the options in their original order as `A`, `B`, …, `Z`.
2. Ask the model for exactly one option letter with thinking disabled.
3. At the answer position, collect every vocabulary token whose decoded surface
is exactly each available uppercase letter. This emulates Cygnet's
`structured_outputs.choice` mask; leading-space, punctuated, and lowercase
forms are not admitted.
4. Sum duplicate-token mass per letter and normalize over the available options.
5. Optionally apply a temperature fitted for that model and evaluation domain.

The model produces no explanation and the method needs no additional training.
The same readout can be applied after loading a JevAny LoRA. For pointer
checkpoints, the letter and native pointer distributions can also be combined
log-linearly. The integrated commands call this knob
`--letter-pointer-weight`; the standalone experiment script calls it
`--pointer-weight`.

## Run the public diagnostic

Use `--readout letter` to score a JevAny checkpoint on any regular frozen suite
or labelled JSONL. This example keeps temperature at 1 rather than copying a
value fitted for another model:

```bash
jevany eval \
--run SimpleJev/JevAny-Qwen3.5-4B-LoRA \
--suite /path/to/jevbench-public-v1.4.2.2 \
--out runs/letter-readout/qwen35-4b \
--device cuda --readout letter --letter-temperature 1.0
```

The same selector works for one local decision or an HTTP deployment. Native
checkpoint readout remains the default when `--readout` is omitted:

```bash
jevany decide request.json \
--checkpoint SimpleJev/JevAny-Qwen3.5-4B-LoRA \
--device cuda --dtype bf16 --readout letter

jevany serve \
--checkpoint SimpleJev/JevAny-Qwen3.5-4B-LoRA \
--device cuda --dtype bf16 --readout letter
```

The standalone evaluator additionally supports a frozen base via `--base`,
tier sampling, and explicit LoRA scaling:

```bash
python scripts/evaluate_letter_readout.py \
--checkpoint SimpleJev/JevAny-Qwen3.5-4B-LoRA \
--suite /path/to/jevbench-public-v1.4.2.2 \
--out runs/letter-readout/qwen35-4b \
--device cuda --dtype bf16 --temperature 1.0
```

Add `--sample-per-tier 1` for a three-record smoke test. Use
`--lora-scale 0` with a checkpoint to evaluate its exact pinned base without
the adapter. A nonzero `--pointer-weight` requires a native pointer checkpoint
and adds a second, pointer-formatted model pass per request.

Current limits:

- text-only model input;
- 1–26 options per question;
- safetensors weights for models with an untied language-model output head;
- one letter-formatted prefill per question, plus one native prefill when
pointer blending is enabled;
- native inference-limit and CUDA-graph flags do not apply to the chat-formatted
letter path; use `--letter-max-tokens` for its prompt limit.

Temperature changes reported probabilities but does not change the letter-only
argmax. Fit it on a separate calibration split for each model. Cygnet's `3.4`
was fitted for its own Gemma configuration and is not a default for JevAny.

## Benchmark units

The official score and the public diagnostic answer different questions:

| Result | Evaluation set | Unit |
|---|---:|---|
| Cygnet 73.70 | JevBench v1.5.4, 1,624 decisions: 904 open + 720 sealed | Equal-weight harmonic mean of Intelligence, Calibration, Speed and Cost |
| Cygnet 203/231 (87.9%) | Public development set, 231 decisions | Accuracy |
| JevAny results in the current README | Same public development set, 231 decisions | Accuracy |

Cygnet ranks first among 106 systems on the official v1.5.4 composite and is a
statistical tie with Winnow-12B Q8. The 73.70 composite is not `73.70%` and
cannot be compared numerically with public-set accuracy. See Cygnet's
[official-submission report](https://github.com/blockbrain-ai/cygnet-recipe/pull/4)
and the [JevBench v1.5.4 board](https://benchmarkheaven.com/jev-models/v1.5.4).

## Results

All JevAny values below are local pinned-checkpoint runs at temperature 1. The
complete metrics, suite hashes, report hashes, and paired counts are in
[`results/letter-readout-v1.json`](../results/letter-readout-v1.json).
These are separate matched reruns of the pinned released checkpoints for this
ablation; they do not replace the main model-family table, whose frozen
checkpoints and run provenance differ.

![Accuracy-point change from native pointer for letter and fixed-blend readouts](letter-readout-results.svg)

JevBench public is a 231-question development diagnostic, not the sealed
v1.5.4 board. `Base + letter` disables the JevAny adapter; the other columns use
the released adapter. Deltas and two-sided exact McNemar p-values compare each
adapter readout with the native pointer on the same questions.

| Model | Base + letter | Native | Letter | Fixed 50/50 blend |
|---|---:|---:|---:|---:|
| Qwen3.5-4B | 79.65% | 80.09% | 81.39% (+1.30, p=.749) | **81.82%** (+1.73, p=.424) |
| Qwen3.8-27B | 88.31% | 89.61% | 89.61% (+0.00, p=1.000) | **90.04%** (+0.43, p=1.000) |

Transfer-v9 evaluates all 1,264 requests without rejection or truncation. Its
accuracy headline uses the 1,046 clean knowable decisions.

| Model | Native | Letter | Fixed 50/50 blend |
|---|---:|---:|---:|
| Qwen3.5-4B | **79.16%** | 75.72% (-3.44, p=.0028) | 79.06% (-0.10, p=1.000) |
| Qwen3.8-27B | 86.23% | 84.23% (-2.01, p=.0375) | **86.90%** (+0.67, p=.337) |

The 27B blend gets 909/1,046 decisions right versus 902/1,046 for native, but
the seven-question gain is not statistically significant. The fixed blend is
also not a universal improvement: at 4B it gets one fewer answer right. None of
the positive gains in either table is significant at the 0.05 level. The
letter-only Transfer-v9 losses show that a public-diagnostic gain does not by
itself establish transfer.

## Latency diagnostic

One warmed run per mode on H200/BF16 measured the same heterogeneous 231-record
panel with 16 warmups, one measured repeat, and concurrency 1. Values are external
end-to-end median / p95 milliseconds; loading and network transport are
excluded.

| Model | Native | Letter | Fixed 50/50 blend |
|---|---:|---:|---:|
| Qwen3.5-4B | 138.43 / 200.98 | **131.86** / 201.55 (1.05x median speedup) | 250.34 / 388.31 (1.81x median latency) |
| Qwen3.8-27B | 194.70 / 477.67 | **181.58** / 492.51 (1.07x median speedup) | 367.40 / 943.44 (1.89x median latency) |

Letter-only has a modest median improvement on this panel despite using more
logical input tokens on average (701 versus 602); p95 does not improve. The
blend evaluates both prefills and averages 1,303 logical input tokens. This is
not saturated server throughput or a general speed claim.

## Practical principles

- Treat letter readout as a model-and-domain ablation, not a drop-in upgrade.
Keep it only after a paired evaluation on the target decision distribution.
- Blend only when the two readouts make complementary errors. The fixed 50/50
pool helped 27B on both diagnostics; at 4B it helped JevBench public but not
Transfer-v9. Select the weight on a separate development split.
- Calibrate each model, readout, and domain separately. Cygnet's fitted
temperature `3.4` does not transfer to these Qwen checkpoints.
- Re-measure latency in the deployment runtime and traffic mix. The one-pass
letter path can trim median latency, while blending requires both letter and
native passes and nearly doubles median latency here.

## Attribution

The Cygnet-compatible prompt and aggregation semantics are adapted from
`blockbrain-ai/cygnet-recipe` commit
`3cf591c692dec649f7c134449814610307c7bb3a`, Copyright 2026 Nood Co and
contributors, under MIT. Cygnet credits the one-token option-letter readout
method to [NInfer](https://github.com/igorls/ninfer), released under Apache-2.0.
JevAny does not incorporate NInfer source code. See [NOTICE](../NOTICE) for the
full Cygnet license notice.
38 changes: 38 additions & 0 deletions docs/letter-readout-results.svg
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Loading
Loading