A cheap, fail-open semantic edge layer for Jev (TypeSafe's System One model), served through the OpenCode Zen gateway.
Jev is not an LLM. It can't write text, code, or images. It takes a state (any messy context) plus typed questions and returns typed decisions with calibrated confidence in ~70–500ms — a choice from up to 255 options, a score on your rubric, or yes/no. This package turns that primitive into a reusable decision service with a confidence policy, per-lane decision logging, a blind eval harness, and a skill router over large skill rosters.
Verified live (2026-09-19/20, jev-1.13-free, cost $0 across all runs):
| test | result |
|---|---|
| 3-question smoke (noul/choice/score mixed) | all correct, 1 call |
| Blind site QA vs labeled ground truth (samuraiblaque.com) | 4/5 honest agreement, ~0.6s |
| VLM-describe-once → Jev rubric pass (8 checks) | 8/8, 0.64s wall |
| Skill router, 1,248 skills, 3-leg fusion | correct picks at 0.99–1.0 confidence, 1.0–1.3s |
| Second-signal ablation | Jev demoted a keyword-tied noise skill 2.00 → 0.18 fused |
| Embedding leg (nomic-embed-text, local ollama) | closed the vocabulary-divergence recall gap; 1,248 vectors built in 17s |
| Gauntlet preflight checklist (unlabeled) | 3 clear / 0 fail / 3 uncertain on a live page, exit-code routing |
Independent early evidence (RouterArena-style ablations, a 33K-skill retrieval study, a 2,000-email phishing benchmark) points the same direction:
- Jev alone on a routing task ≈ no-Jev baseline. The retrieval evidence, not the model, did the work.
- Jev fused with a first signal (keyword/embedding retrieval) beat either alone.
- One giant verdict ("is this phishing?") scored 62.6%; the same model's outputs decomposed into narrow signals + code reached ~95%.
So the architecture here is: deterministic facts in code → Jev as a second semantic signal → policy on confidence → escalation is your choice. Jev is replaceable infrastructure — swap the backend, keep the service. Never the only signal on expensive decisions, never the chat model.
# get an OpenCode Zen key (https://opencode.ai/zen) or a TypeSafe key (https://typesafe.ai)
export OPENCODE_ZEN_API_KEY=... # or TYPESAFE_API_KEY with the TypeSafe endpoint
# no dependencies beyond the Python stdlib (3.9+)Models: jev-1.13-free (limited-time free tier) or jev-1.13 ($0.042/M input, output free).
from fastloop import evaluate, summarize
rec = evaluate(
state="Support email text: My payments have failed for three days...",
questions={
"is_urgent": {"type": "noul", "instructions": "Does this need attention right now?"},
"department": {"type": "choice", "instructions": "Which team handles this?",
"criteria": {"billing": "...", "shipping": "...", "returns": "..."}},
"frustration": {"type": "score", "instructions": "How frustrated?", "criteria": ["Calm", "Frustrated", "Very angry"]},
},
lane="inbox-triage",
)
print(summarize(rec)) # [inbox-triage] jev-1.13-free 0.5s cost=0 :: is_urgent=0.88, ...
print(rec["routing"]) # per-question: action / escalate / abstain by confidenceEvery decision appends to ~/.hermes/fastloop/decisions.jsonl (or set log=False): ts, lane, model, cost, latency, input_hash, questions, answers, routing, thresholds. That log is the eval database — label outcomes later, tune thresholds from evidence, never from vibes.
Fail-open: backend errors raise FastLoopError — catch it and choose your own safe fallback (usually: just do the expensive thing).
python3 eval_preflight.py ground_truth.jsonGround truth comes first from deterministic checks; Jev sees only the state — never your labels. Scores agreement + keeps dissent visible.
python3 build_skill_index.py # index a skill roster (name + description)
python3 embed_retrieve.py --build # optional: local embeddings via ollama (nomic-embed-text)
python3 skill_router.py "judge my renders against a rubric"
# ABSTAIN when the first signal is weak - Jev can only rerank what retrieval surfacesThree legs, all fail-open: keyword (deterministic), embedding (local ollama, catches vocabulary-divergent queries), Jev (semantic rerank). Any leg going down degrades gracefully. Abstain only when the first signal is weak AND no embedding leg exists.
python3 preflight_checklist.py checklist.json # exit 0 clear / 1 fails / 2 uncertainFeed a deliverable's state plus defect-shaped questions; get per-check pass/fail/uncertain with exit-code routing. High-confidence failures get fixed before the expensive jury wakes up; uncertain ones escalate. Costs ~$0 and ~0.7s per pass.
| type | shape | returns |
|---|---|---|
noul |
{"type":"noul","instructions":"..."} |
yes/no probability (0–1) |
choice |
{"type":"choice","instructions":"...","criteria":{...}} |
choice + confidence + distribution |
score |
{"type":"score","instructions":"...","criteria":[...]} |
rubric index + confidence + legend |
All three mix in one call; questions evaluate in parallel and in isolation. Decompose multi-factor judgments into atomic questions and combine with code.
Fits: recognizing, discriminating, scoring, ranking, selecting among things already present — triage, routing, reranking, QA gates, date-vs-stated-anchor comparisons, supervision signals.
Doesn't: inventing, synthesizing, planning, arithmetic/counting (compute in code), free-form generation (impossible by design), image/video/audio perception (feed it VLM observations instead).
Gotcha: the Zen gateway WAF 403-blocks the default Python-urllib user-agent — jev_ask.py sends a browser UA; keep that header on any hand-rolled client.
MIT