Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
21 changes: 18 additions & 3 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -534,6 +534,24 @@ docs/architecture/plan-integration/*
!docs/architecture/plan-integration/proposals/
!docs/architecture/plan-integration/*.architecture.json
!docs/architecture/plan-integration/*.md
# Un-ignored 2026-09-19 — jev-judge programme record (Jev / TypeSafe System One as an
# optional judgement provider): stage receipts, council briefs and seat reviews.
# Same shape as plan-integration above: small Markdown/JSON only.
!docs/architecture/jev-judge/
docs/architecture/jev-judge/*
!docs/architecture/jev-judge/council/
!docs/architecture/jev-judge/receipts/
!docs/architecture/jev-judge/*.architecture.json
!docs/architecture/jev-judge/*.md
# Un-ignored 2026-09-19 — learning-tier programme record (feed the mentor's writers on all
# six harnesses, then measure one judgement model on the struggle ruler): council brief,
# seat reviews, arbitration, pre-registration and stage receipts. Small Markdown/JSON only.
!docs/architecture/learning-tier/
docs/architecture/learning-tier/*
!docs/architecture/learning-tier/council/
!docs/architecture/learning-tier/receipts/
!docs/architecture/learning-tier/*.architecture.json
!docs/architecture/learning-tier/*.md

# Demo recordings (large, local-only)
demos/
Expand Down Expand Up @@ -621,9 +639,6 @@ MagicMock/
.pi/
agent/
skills-lock.json
#
# openspec files
openspec/
# dotfiles/dirs
.graphify-labels.json
.graphifyignore
Expand Down
20 changes: 0 additions & 20 deletions .graphify-labels.json

This file was deleted.

40 changes: 0 additions & 40 deletions .graphifyignore

This file was deleted.

6 changes: 5 additions & 1 deletion .pre-commit-config.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -41,11 +41,15 @@ repos:
# sha256 and the UAT redacted summary carries the rubric hash and the
# private bundle's manifest digest (council D-22) -- every one a digest the
# redaction rules allow out precisely because it identifies without revealing.
# Learning-tier council manifests and receipts (2026-09-19) are the same two
# classes: brief/system-prompt digests and pre-registration fingerprints.
exclude: |
(?x)^(
packages/studyloop/tests/acceptance/uat/data/.*_registry\.json|
docs/architecture/plan-integration/council/.*/manifest.*\.json|
docs/architecture/plan-integration/receipts/.*\.json
docs/architecture/plan-integration/receipts/.*\.json|
docs/architecture/learning-tier/council/.*/manifest.*\.json|
docs/architecture/learning-tier/receipts/.*\.json
)$
- repo: https://github.com/PyCQA/bandit
rev: 1.8.3
Expand Down
4 changes: 2 additions & 2 deletions .secrets.baseline

Some generated files are not rendered by default. Learn more about how customized files appear on GitHub.

33 changes: 33 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -29,6 +29,39 @@ experience may change before `1.0.0`.
a hand-off during a live session, or rejoining another session never shows
it under different material. Proposal only — the companion says nothing
about it. The no-plan payload is byte-identical.
- A standalone `studyloop-study-notes` Agent Skill with per-lesson Markdown and
section-overview templates, source/enrichment attribution, and explicit
Obsidian-only, xTiles-only, and linked dual-destination workflows. Installed
separately through the skills CLI, not by `studyloop install agents`; its
guide is published on the docs site as *Lesson Study Notes*
(`docs/study-notes-skill.md`).

### Changed

- The `openspec/` tree is no longer listed in `.gitignore`. Seventy-eight
tracked, load-bearing files lived under an ignored path, so every new spec
or archive file was invisible to `git status` and skipped by `git add -A`.
The `herdr-ghostty-multiplexer-transport` change is archived as deferred
(the owner's 2026-09-05 decision, unchanged: tmux stays the production
default, herdr an experimental opt-in) with its five open tasks closed as
not built and its spec deltas deliberately not merged — they modified
requirements no main spec contains and describe a wterm selector and a ttyd
fallback the tree has since retired. The July e2e/MCP archive's four open
tasks are reconciled against the tree (all four were built the same day
their "not done" notes were written), so `openspec validate --archived
--all` is green for the first time since the release guard was added.

### Fixed

- The `now` engine counts every "last seen N day(s) ago" from the one clock
it reads per plan. The due-progress collector took its day count from a
second clock inside `history/progress.py`, so under a frozen test clock the
count — and the score built from it — drifted with the real date (the
medium-energy screen in receipt `now-rubric-2026-09-16` printed 8 days for
a struggle planted 3 days back). Production always read one wall clock, so
no learner saw the split; `spaced_repetition_due()` gains a keyword-only
`now` for callers that already hold an instant, defaulting to the wall
clock for everyone else.

## [0.5.0] - 2026-09-21

Expand Down
7 changes: 4 additions & 3 deletions Justfile
Original file line number Diff line number Diff line change
Expand Up @@ -218,9 +218,10 @@ release-consistency:
# change with commits since the last tag must be archived or carry a
# `deferred: <reason>`, and archive entries ADDED since the last tag must pass
# `openspec validate` (soft-skipped when the CLI is absent, same convention as
# spec-check; not `--archived --all`, because a July archive predating this
# guard has unticked tasks nobody has evidence to reconcile, and re-failing
# every future release on it would teach people to ignore the gate).
# spec-check; not `--archived --all`, so that a historical archive nobody is
# working on can never re-fail a future release and teach people to ignore
# the gate — the July archive that motivated this was reconciled against the
# tree on 2026-09-23 and `--archived --all` is green today).
# Deliberately NOT part of preflight: open changes are legal during a cycle;
# only shipping one is not. Both guards would have fired on the 0.2.0 cut
# (2026-09-04 review, Q5).
Expand Down
100 changes: 100 additions & 0 deletions docs/architecture/jev-judge/receipts/stage0-access-2026-09-19.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,100 @@
# Stage 0 — Jev access spike (2026-09-19)

**Programme:** `feat/jev-judge` — evaluate TypeSafe's Jev (a "System One" judgement model:
typed questions against a text state, returns typed answers with probabilities and
confidence, no text generation) as an *optional judgement provider* for StudyLoop's
unfilled judgement slots. Out of scope by design: the decision engine and completion
review (`learning/decision.py`, `planning/views.py`) — those count and compare dates,
which the vendor's own jaggedness page says Jev cannot do.

**Question this stage answers:** does the pre-release account answer at all, does the SDK
match its docs, and do the five 1–4 teach-back dimensions map onto Jev's Score legend?
It is an access and shape check, **not** a measurement of accuracy (n = 1 synthetic state).

## Method

- Script: `scripts/eval/jev_stage0_spike.py` (disposable; reads `TYPESAFE_API_KEY` from
the environment, never prints it). SDK `typesafe-sdk==0.7.0`, Python 3.12.8.
- Model pinned to `jev-1.13.0` — never the `jev-latest` alias, because any threshold
tuned against an alias silently moves when the alias does (vendor's advice, adopted).
- State: one synthetic teach-back (1,873 chars) — a networking-background learner
explaining Python decorators with a middlebox/NAT analogy and a retry-decorator
transfer example. **One factual error planted on purpose**: "the wrapping happens every
time you call the decorated function, not when it's defined". A human marks that
Accuracy 2 ("mostly correct, minor gaps") at best.
- Questions in one call: five `Score` questions whose four criteria are the verbatim level
descriptors from `agents/shared/teach-back-protocol.md` (Recitation → Teaching), plus
two `Noul` controls — positive (`has_factual_error`) and negative (`not_english`).
- Three identical calls to measure spread.

## Results

| dimension | jev score ×3 (0–3) | StudyLoop 1–4 (mean) | confidence ×3 | range | human expectation |
|---|---|---|---|---|---|
| accuracy | 2.09, 2.10, 2.15 | **3.11** | 0.52, 0.52, 0.54 | 0.06 | **2** (planted error) |
| own_words | 2.98, 2.98, 2.98 | 3.98 | 0.98, 0.98, 0.98 | 0.00 | 4 (novel analogies) |
| structure | 2.90, 2.88, 2.87 | 3.88 | 0.90, 0.88, 0.87 | 0.03 | 3–4 |
| depth | 2.03, 2.03, 2.03 | 3.03 | 0.96, 0.96, 0.97 | 0.00 | 3 (WHAT/HOW/WHY, no tradeoffs) |
| transfer | 2.67, 2.65, 2.68 | 3.67 | 0.67, 0.65, 0.68 | 0.03 | 3–4 |
| noul: has_factual_error | 0.48, 0.48, 0.43 | – | – | 0.05 | high |
| noul: not_english | 0.01, 0.01, 0.01 | – | – | 0.00 | ≈ 0 |

Per-level probabilities are returned for every Score (accuracy call 1:
`{0: 0.02, 1: 0.16, 2: 0.54, 3: 0.28}`), so a caller can take the argmax level and gate on
its mass instead of rounding the expected-value float — the docs say score levels are weak
in numerical calibration and warn against interpolating between levels.

Access: `model_answered = jev-1.13.0`. Latency 1,779 ms cold, then 521 / 534 ms. Usage
973 input / 107 output tokens per call → about $0.00004 per teach-back at $0.042 / Mtok.
Cost is not a factor in any later decision.

## Findings

1. **Access works and the SDK matches its docs.** `TypeSafeClient()`, `Score`, `Noul`,
`client.system_one(state=, questions=, model=)`; answers carry `score`, `confidence`,
`probabilities`, `legend`.
2. **Consistent, not deterministic — confirmed by probe, not by prose.** Across identical
calls the Score range was ≤ 0.06 on a 0–3 scale (two dimensions identical to three
decimals), Noul range ≤ 0.05. Any Stage 1/2 test must assert a band or a level, never
an exact float, and CI must replay recorded fixtures rather than call live.
3. **The pedagogical dimensions landed where a human would put them.** Own words, structure
and depth scored within the human expectation with confidence ≥ 0.87; depth 3.03 is
exactly the "WHAT, HOW and WHY, but no tradeoffs" reading of the text.
4. **Accuracy was blind to the planted error.** 82 % of the mass sat on the two "accurate"
levels and the error Noul stayed at 0.43–0.48, i.e. "don't know". This is the vendor's
documented *literal reading* / *no technical precision* edge: Jev is a common-sense
judge, not a Python-semantics checker. **But the confidence signal worked** — accuracy
was the least-confident dimension (0.52 vs ≥ 0.65 elsewhere), so a confidence gate
would have routed it to "ask the learner" rather than recording a 3.
5. Negative control held at 0.01: the Noul is not agreeing with everything.

## Design consequences carried into Stage 1 / Stage 2

- **Split by strength.** Correctness is domain knowledge; pedagogy is calibrated
judgement. Stage 2 should have the mentor agent (harness LLM, which does know Python
semantics) supply the *accuracy* judgement, and Jev score the four pedagogical
dimensions — own words, structure, depth, transfer — where it excelled and where a
generative model is weakest at calibration. Do **not** ship Jev as a sole accuracy judge.
- **Argmax + gate, not rounding.** Map a Score to StudyLoop's 1–4 as `argmax(probabilities)
+ 1`, and record it only when the argmax mass clears a threshold pinned to `jev-1.13.0`;
below it, the dimension is "not scored — ask the learner" (the same never-fabricate rule
the parked Bedrock extractor taught: `history/teachback.py` must never receive a guess).
- **Hypothesis for the Stage 2 gold set, not a result:** the two least-confident dimensions
here (accuracy 0.52, transfer 0.67) were the two a human would hesitate on. Whether
confidence tracks human disagreement is the first thing the labelled set should test.

## Limits

One synthetic state, three repeats, no gold labels. Nothing above is an accuracy figure;
it is a shape check that surfaced one hard constraint (finding 4) early enough to design
around it.

## Reproduce

```bash
set -a; . /path/to/studyloop/.env; set +a # TYPESAFE_API_KEY only; never committed
uv run --no-project --python 3.12 --with typesafe-sdk==0.7.0 python scripts/eval/jev_stage0_spike.py
```

Writes `stage0-receipt.json` beside this file (the committed copy is the run described
above). Numbers will differ slightly on re-run — see finding 2.
Loading
Loading