Skip to content

Housekeeping: skill guide published, jev plan on main, openspec tree un-ignored and archived, one-clock fix - #35

Merged
NetDevAutomate merged 12 commits into
mainfrom
chore/housekeeping-0.5.x
Sep 23, 2026
Merged

NetDevAutomate merged 12 commits into
mainfrom
chore/housekeeping-0.5.x

Conversation

@NetDevAutomate

Copy link
Copy Markdown
Owner

Summary

The housekeeping pass that follows #32: every loose end the tracker did not hold, each with a countable finish, as one pure fast-forward of main. Twelve commits, one per concern.

# Concern Finish Commits
1 The studyloop-study-notes skill work that sat uncommitted on the main checkout On the branch unchanged (cherry-picked from feat/study-notes-skill, file-by-file identical; only the CHANGELOG merges under one ### Added), and its guide published on the docs site as Lesson Study Notes — one exclude line, one nav entry, the SKILL.md link now a GitHub URL so --strict accepts it a82fa2fd 2fcdb510 42d28de3
2 The two graphify files (.graphify-labels.json, .graphifyignore) — deletions the owner confirmed deliberate Untracked; graphify_refine.py's two docstrings no longer claim the labels file is tracked 2fcdb510
3 The unpushed feat/jev-judge planning branch (Stage 0 receipt + council-validated plan) Both commits cherry-picked, content verified file-by-file against the originals (only .secrets.baseline's generated_at differs — regenerated by a whole-repo detect-secrets scan, results file count 72 → 72). Plan §7's five owner decisions filed as #34 (decision 5, merge order, is resolved by this very cherry-pick). Worktree removed by a script that gated the removal on status --porcelain empty and every ignored entry being a known cache 728600c1 ce78bf98
4 openspec/ ignored in .gitignore while 78 tracked files live under it Rule removed. Found because git add -A openspec/changes refused the new archive; git status --ignored showed nothing else hidden under the path 7ed7ae3d
5 The herdr-ghostty-multiplexer-transport change: deferred: since 2026-09-05, never archived, five open boxes Archived as 2026-09-23-… with --skip-specs. Each open box carries a closed 2026-09-23 — NOT built, deferred with the change disposition and a reopen condition. Deltas not merged, verified against the tree: all eight MODIFIED headers exist in no main spec; the web-ui delta preserves the wterm selector e9cb5656 removed and a bootstrap e4b17d46 replaced; the session-transports delta routes a ttyd fallback the PTY start path calls retired. Main specs byte-identical (sha256 of all 23). openspec/changes now holds only archive/ fb999c53
6 The frozen-clock gap recorded in receipt now-rubric-2026-09-16 row 3c (e) RED → GREEN. spaced_repetition_due() gains keyword-only now; _due_progress_candidates takes the engine's instant; build_now_plan passes its one read. The RED patches history.progress.datetime with a decoy (2031) so it fails on any date if the count comes from anywhere but the engine. Deliberate test change: six collector stubs accept **_. One RED expectation corrected at GREEN (score = 138 + PLAN_RELATED_BIAS, asserted from the constant). Receipt row 3c records the fix beside the gap 3b6e8e00 e126973d
7 The July archive 2026-07-12-complete-e2e-harness-and-desktop-mcp: four open boxes, the one red in openspec validate --archived --all since the release guard was added All four reconciled against the tree — each was built the same day its "not done" note was written (fake agent tests/_fake_agent.py; journey phases across two e2e files; session-end/review rows asserted; MCP explorer tools in 9f09cf08). --archived --all 8/8 for the first time. The Justfile comment and validate_new_archives docstring that cited the July case now state the design reason and record it closed 6d3c43d6 d277286f
8 CHANGELOG [Unreleased] Added/Changed/Fixed for the above c1a28de1

Verification (all on this head)

  • ruff check clean · ruff format --check 1044 files · pyright 0 errors · bandit passed
  • mkdocs build --strict exit 0 (the study-notes page present in the built site)
  • openspec validate --specs --all 23/23 · --archived --all 8/8 · scripts/check-release-consistency.py --skip-wheel --release --pre-tag passes (validates the herdr archive as a new archive)
  • JS node --test 164/164
  • Engine/decision/history/recap/web-now/web-history 153 passed (RED flipped); docs contract 31; docs drift + counts + second-brain 318; release-consistency script tests 16
  • Golden now_plan_no_active.json ec451ce8 byte-identical (the no-plan payload is untouched)
  • Full pytest -rfE: see the comment below (run against the committed environmental id set from the item-4 control receipt)

Owner steps after merge

  1. git merge --ff-only origin/chore/housekeeping-0.5.x && git push origin main (pure fast-forward of 12 commits).
  2. Delete two local branches that this PR supersedes and whose commits it carries with different SHAs: git branch -D feat/study-notes-skill feat/jev-judge — -D because the cherry-picks changed the SHAs, so -d's merged-check refuses; content proven present above.
  3. Say the word and I delete the three remote branches (feat/body-double-first-move, feat/study-notes-skill, chore/housekeeping-0.5.x) in one ruleset window, proving the deletion command is permitted against the live ruleset first.

Not in this PR, on purpose

One Markdown note per lesson plus a linked section overview, with the
seven-part structure of the comprehensive study notes kept intact and
explicit Obsidian-only, xTiles-only, dual-destination and portable
workflows. The package is skills/studyloop-study-notes/ (SKILL.md, two
templates, four references) and has no runtime dependency on StudyLoop;
it is installed separately through the skills CLI.

This work was written on the main checkout while PR #32 was open and is
moved here unchanged so main can fast-forward to the PR head: main and
the branch both rewrite the CHANGELOG [Unreleased] section, and the
uncommitted copy would have refused the merge. The guide is source-only
for now, as mkdocs.yml's documentation contract requires until a page is
rooted and added to nav.
Both files were already listed in .gitignore; deleting the tracked copies
completes that decision (owner, 2026-09-22: deliberate). They are local
graphify tuning — curated community labels and an ignore list — that
belong to one machine's graph build, not the repository.

scripts/graphify_refine.py already tolerates the labels file being absent
(load_curated_labels returns {} when it does not exist); only its docstrings
still claimed the file was tracked in git, so they now say local and
optional. Nothing else in the tree references either file.
The guide moved onto this branch source-only: mkdocs.yml excludes every
page not whitelisted, so it built into nothing until it was both excluded
and rooted in nav. This is the skill's finish as the housekeeping plan
named it: one exclude line, one nav entry beside Obsidian Export, and the
repo-relative link to SKILL.md becomes the GitHub URL because a published
page may not link outside the docs tree under --strict. The page itself is
unchanged and still says the skill is installed with the skills CLI, not by
studyloop install agents.

mkdocs build --strict clean, the page present in the built site;
test_docs_drift + test_docs_no_hardcoded_test_counts + test_second_brain_docs
318 passed, 3 skipped.
… dimension blind

Evaluate TypeSafe's Jev (jev-1.13.0, a System One judgement model) as an
optional judgement provider for StudyLoop's unfilled judgement slots. Stage 0
is an access-and-shape check only: one synthetic teach-back with a planted
factual error, five Score questions whose criteria are the verbatim 1-4 level
descriptors from agents/shared/teach-back-protocol.md, two Noul controls,
three identical calls.

Why this is recorded rather than thrown away: the probe surfaced a hard
constraint early. Own words / structure / depth landed where a human would
put them (confidence >= 0.87), but accuracy put 82% of its mass on the two
"accurate" levels despite the planted error, and the error Noul sat at
0.43-0.48. Jev is a common-sense judge, not a Python-semantics checker.
Its confidence signal did flag accuracy as the least certain dimension
(0.52), so a confidence gate would have routed it to "ask the learner".
Design consequence carried into Stage 2: the harness LLM supplies accuracy,
Jev scores the four pedagogical dimensions, argmax-level + mass gate instead
of rounding the float.

Identical calls moved <= 0.06 on a 0-3 scale: consistent, not deterministic,
so later tests assert bands/levels and CI replays fixtures, never calls live.

Out of scope by design: learning/decision.py and the completion review
(counting + dates), which the vendor's jaggedness page says Jev cannot do.

.gitignore gains a jev-judge allowlist block in the same shape as
plan-integration (small Markdown/JSON only). The spike reads
TYPESAFE_API_KEY from the environment; the key lives in an ignored .env.
…sure one judge

Planning round for two work items the owner asked for after the Jev Stage 0
spike: (1) make the mentor agent actually write learning signals on every
supported harness, proven by programmatic simulation of a learner; (2) one
pre-registered, directional measurement of Jev as a filter over harness-
proposed struggle candidates on the frozen 13-session gold. Brief, three
seats (gpt-6-astra, qwen3-coder, grok-4.6 re-run), arbitration and the
resulting plan are all on the record here.

Why the plan differs from the brief the seats reviewed: verifying their
claims against source overturned four of them and sharpened the rest.
The largest: the gold labels (6c7176e, 2026-05-31) were authored against
the pre-rebuild corpus, and 9 of the 10 gold sessions present in both
databases now have DIFFERENT transcripts in the live file (the 2026-09-12
clean start stripped tool echoes; one Claude session 195 -> 49 messages).
So the ruler is the cold archive, all 13, mode=ro -- not the live-union the
brief proposed. Also verified: both DBs are schema 48 (Grok B4 falsified),
claude_code is in STUDY_SOURCES (B3 falsified), eval_runner already opens
read-only (B2 partly falsified), C1's recall varied 1/9..7/9 across the 83
ledger rows and was never run held-out (Astra F05 confirmed; corrects the
coordinator's own "stuck at 1/9" claim), all six harness binaries are
installed here (Astra F01 achievable), held-out has 2 negatives (0/2 FP has
a 77.6% one-sided upper bound -- directional only).

Decisions carried into the plan: MCP record_teachback with a validator
shared with the CLI; writer set split W_auto (additive, pre-approved) vs
W_srs (SRS mutators accept unverified card hashes -- stay prompt-per-call);
one YAML trigger table projected to all six harnesses with CI byte-identity;
definition of done states exactly which harnesses passed and counts no skip
as a pass; Jev measured only as the delta between matched candidate lists
with and without its filter; no adopt/reject clause; two branches, two PRs
(deviates from the owner's "one branch" ask -- flagged for decision).

Grok's first seat looped for 24,000 tokens announcing repo inspections it
could not perform (kept as seat-grok-4.6.INVALID-tool-loop.md, truncated);
the re-run with scripts/council/system-seat.md is the valid seat.

.gitignore: learning-tier allow-list block, same shape as plan-integration.
.pre-commit-config.yaml + .secrets.baseline: council manifest/receipt JSON
under learning-tier excluded from the hex-entropy detector, same class and
rationale as plan-integration (sha256 digests of committed public text).
Baseline regenerated whole-repo with the pinned v1.5.0: results 72 -> 72,
only the filter pattern changed.
openspec/ was ignored at .gitignore:644 while 78 files under it are tracked
and load-bearing (the release gate reads openspec/changes, CI validates the
specs). That is the tracked-but-ignored trap this file's own comment warns
about: every new archive or spec file was invisible to `git status` and
skipped by `git add -A`, so archives had to be force-added. Nothing else
under openspec/ was being ignored (git status --ignored: only the archive
created this morning), so removing the rule surfaces exactly the files that
should have been visible all along.
…ed, deltas not merged

The change has said `deferred:` in its .openspec.yaml since 2026-09-05 (the
owner kept tmux as the production default; herdr stays an experimental
opt-in, ghostty dev-only), but it sat un-archived with five open boxes, so
openspec/changes held a live directory nobody was working on. Each of the
five boxes now carries an explicit "closed 2026-09-23 - NOT built, deferred
with the change" disposition naming its reopen condition (T1.2 dedicated
TmuxBackend suite, T1.3 settings.py mention, T1.4 old-format round-trip
test, T1.6 test_session_start mock target, T4.2 T6 herdr detach xfail), so
the archive counts complete without claiming work that was never done.

Archived with --skip-specs. Verified against the tree, not recalled: all
eight MODIFIED requirement headers exist in no main spec (grep of
openspec/specs for each header: no file), so they were modifications of
requirements never written; the web-ui delta preserves the wterm selector
e9cb565 removed and a bootstrap e4b17d4 replaced; the session-transports
delta routes a ttyd fallback the PTY start path calls retired. The delta
files stay in the archive as the record. Main specs byte-identical (sha256
of all 23 spec.md before/after).

openspec/changes now holds only archive/. `openspec validate --archived
--all`: the new archive passes (7 passed; the one failure is the July
archive that fails identically on main). `--specs --all` 23/23. The release
gate (validate_new_archives) will validate this archive at 0.5.1 and it
passes today.
…history.progress's own

Receipt now-rubric-2026-09-16, row 3c reading (e) (2026-09-21): the medium
screen printed "last seen 8 day(s) ago" for a struggle planted three days
before the frozen date. spaced_repetition_due() reads datetime.now(UTC) in
history/progress.py; the guidance tests freeze decision.datetime only. The
count drifts with the real date, and with it the due copy's score
(100 + min(days, 30) + 35), so a pinned screen can re-rank on a later day.

The test patches progress.datetime with a decoy (2031) so it fails on any
wall-clock date unless the count comes from the engine's instant. Fails now:
days_ago 1570 where 3 is planted.
…s own

build_now_plan reads datetime.now(UTC) once; the due-progress collector
passed nothing on and spaced_repetition_due() read a second clock inside
history/progress.py. Under the guidance tests, which freeze decision.datetime
only, the due copy of a struggle planted three days before the frozen date
printed "last seen 8 day(s) ago" on 2026-09-21 and 10 on 2026-09-23 — and
its score (100 + min(days, 30) + 35) drifted with it, so a pinned screen
could re-rank on a later date. In production both reads were the same wall
clock microseconds apart, so no learner saw the split (receipt
now-rubric-2026-09-16, row 3c (e)).

spaced_repetition_due gains a keyword-only `now` defaulting to the wall
clock (recap, review CLI, planning evaluation and authoring are unchanged);
_due_progress_candidates takes the engine's instant and hands it on;
build_now_plan passes its one read. Deliberate test change: the six stubs of
_due_progress_candidates in test_learning_decision.py and
test_now_plan_guidance.py accept **_ so the seam's new keyword reaches them.
One RED expectation corrected at GREEN: the primary's score is the
collector's 138 plus PLAN_RELATED_BIAS (rule 5), asserted from the constant.

The struggle collector's own today = datetime.now(UTC).date() is left as is:
it lives in decision.py, inside the frozen module, so it has no gap.

RED flipped; guidance + decision + history + recap + web-now + web-history
153 passed; ruff/format/pyright clean; golden ec451ce8 byte-identical;
docs contract green. Receipt row 3c records the fix beside the gap.
… four boxes, all built

openspec validate --archived --all has reported one failure on every run
since the release guard was added: 2026-07-12-complete-e2e-harness-and-
desktop-mcp with four open boxes, which the Justfile excused as "unticked
tasks nobody has evidence to reconcile". The evidence exists; the notes
were written before the same-day commits that built the items:

  1.4 fake-agent PTY binary — tests/_fake_agent.py (reads the persona,
      asks the Socratic bank), test_fake_agent.py, spawned by the journey
      and body-double/ghostty e2e files.
  4.1 journey through generation/review/session end — landed across
      test_representative_user_journey.py (fake-agent walk: start, turn,
      end, durable row) and test_journey_generate_review.py (phases 2-4).
  4.4 session-end export asserted — study_sessions row after
      /api/session/end; card_reviews/review_sessions rows after review.
  5.1 MCP get_lesson_tree/read_lesson/search_lessons — 9f09cf0, dated the
      same day as the "confirmed absent" note; tested in test_mcp_tools.py.

The original notes are kept and each disposition is appended with its
pointer, as the herdr archive does. openspec validate --archived --all:
8 passed, 0 failed.
…ly archive it cited is reconciled

The Justfile comment and validate_new_archives docstring said the July archive "has unticked tasks nobody has evidence to reconcile". As of the previous commit it does not. The scoping to NEW archives stays — a historical archive nobody is working on must never re-fail a release — so the comments now state the design reason and record that the motivating case is closed.
…d, herdr archived as deferred, July archive reconciled, one-clock fix

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Unresolved moderate findings remain in graphify exclusions, clock propagation, Jev model pinning, and lesson-template serialization.

Get a fresh assessment by requesting another Copilot review.

Review effort: Lite
Findings: 3 Medium severity · 2 Low severity

Open (5)
What changed in this PR

This housekeeping PR publishes the study-notes skill, records Jev planning evidence, reconciles OpenSpec archives, fixes frozen-clock handling, and cleans up graphify configuration.

Changes:

  • Adds study-notes documentation, templates, and navigation.
  • Archives and reconciles deferred OpenSpec work.
  • Threads the engine clock through due-progress calculations.
  • Updates planning, release checks, changelog, and repository metadata.
File Summary / review note
skills/​studyloop-study-notes/​SKILL.md Adds the study-notes workflow.
skills/​studyloop-study-notes/​references/​xtiles.md Documents xTiles presentation.
skills/​studyloop-study-notes/​references/​obsidian.md Documents Obsidian workflows.
skills/​studyloop-study-notes/​references/​evidence-and-updates.md Defines evidence and update rules.
skills/​studyloop-study-notes/​assets/​section-overview.md Adds the overview template.
skills/​studyloop-study-notes/​assets/​lesson.md Adds the lesson template. Moderate (1 vote): quote the sequence placeholder so values such as 001 retain their exact text.
scripts/​graphify_refine.py Updates graphify handling. Moderate (2 votes): preserve a pre-scan exclusion for the vendored browser tree before removing .graphifyignore.
scripts/​eval/​jev_stage0_spike.py Adds the Jev access spike. Nit (1 vote): record the uncertain factual-error control as unsuccessful/attempted. Moderate (2 votes): do not retry without the pinned model.
scripts/​check-release-consistency.py Documents archive validation. Nit (1 vote): clarify that it runs openspec validate --archived --all while filtering newly added archives.
packages/​studyloop/​tests/​test_now_plan_guidance.py Adds frozen-clock regression coverage.
packages/​studyloop/​tests/​test_learning_decision.py Updates collector test stubs.
packages/​studyloop/​src/​studyloop/​learning/​decision.py Forwards the engine instant. Moderate (3 votes): thread the instant through all date-derived collectors, including energy-demand and struggle classification.
packages/​studyloop/​src/​studyloop/​history/​progress.py Supports an explicit scheduling instant.
openspec/​changes/​archive/​2026-09-23-herdr-ghostty-multiplexer-transport/​tasks.md Archives deferred transport tasks. Nit (1 vote): accurately retain the unbuilt test gap and reopen condition.
openspec/​changes/​archive/​2026-09-23-herdr-ghostty-multiplexer-transport/​specs/​web-ui/​spec.md Preserves the deferred web UI delta.
openspec/​changes/​archive/​2026-09-23-herdr-ghostty-multiplexer-transport/​specs/​session-transports/​spec.md Preserves the deferred transport delta.
openspec/​changes/​archive/​2026-09-23-herdr-ghostty-multiplexer-transport/​specs/​live-session-orchestration/​spec.md Preserves the deferred orchestration delta.
openspec/​changes/​archive/​2026-09-23-herdr-ghostty-multiplexer-transport/​proposal.md Archives the deferred proposal.
openspec/​changes/​archive/​2026-09-23-herdr-ghostty-multiplexer-transport/​design.md Archives the deferred design.
openspec/​changes/​archive/​2026-09-23-herdr-ghostty-multiplexer-transport/​.openspec.yaml Records deferral metadata.
openspec/​changes/​archive/​2026-07-12-complete-e2e-harness-and-desktop-mcp/​tasks.md Reconciles historical archive tasks.
mkdocs.yml Publishes the study-notes guide in navigation.
Justfile Updates release-check documentation. Nit (1 vote): describe the actual all-archive invocation and filtering.
docs/​study-notes-skill.md Adds the public skill guide.
docs/​architecture/​learning-tier/​plan-2026-09-19.md Adds the learning-tier plan. Nit (2 votes): mark decision 5 resolved and leave only decisions 1–4 pending.
docs/​architecture/​learning-tier/​council/​spend-round1.json Records council usage.
docs/​architecture/​learning-tier/​council/​seat-qwen3-coder.md Adds council review.
docs/​architecture/​learning-tier/​council/​seat-openai.gpt-6-astra.md Adds council review.
docs/​architecture/​learning-tier/​council/​seat-grok-4.6.INVALID-tool-loop.md Records the invalid council run.
docs/​architecture/​learning-tier/​council/​plan-grok-rerun/​seat-grok-4.6.md Adds the valid rerun review.
docs/​architecture/​learning-tier/​council/​plan-grok-rerun/​manifest.json Records rerun metadata.
docs/​architecture/​learning-tier/​council/​brief-plan-2026-09-19.md Adds the council brief.
docs/​architecture/​learning-tier/​council/​arbitration-plan-2026-09-19.md Records arbitration decisions.
docs/​architecture/​jev-judge/​receipts/​stage0-receipt.json Adds raw Jev receipt data.
docs/​architecture/​jev-judge/​receipts/​stage0-access-2026-09-19.md Documents the Jev spike. Nit (2 votes): correct the state_chars count to match the receipt.
CHANGELOG.md Records the housekeeping changes.
.secrets.baseline Regenerates secret-scan metadata.
.pre-commit-config.yaml Excludes council receipt manifests.
.graphifyignore Removes tracked graphify exclusions. Moderate (1 vote): retain an equivalent pre-scan ignore for the vendored browser tree.
.graphify-labels.json Removes tracked curated labels.
.gitignore Un-ignores OpenSpec and planning records.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment on lines +397 to +402
``now`` is the engine's one clock read (``build_now_plan``): the collector
counts ``days_ago`` from it rather than from a second read inside
``history.progress``, so the day count on a screen — and the score built
from it — cannot disagree with the rest of the plan (receipt
``now-rubric-2026-09-16``, row 3c (e): a frozen engine printed a drifting
"last seen 8 day(s) ago" for a struggle planted three days back).
Comment on lines +129 to +133
try:
resp = client.system_one(state=STATE, questions=QUESTIONS, model=MODEL)
except TypeError:
# SDK may not take model= on this call; fall back and record what answered.
resp = client.system_one(state=STATE, questions=QUESTIONS)
Comment on lines +28 to +30
Curated labels live in ``.graphify-labels.json`` at the repo root — a local,
gitignored file (untracked since 2026-09-22); the script runs without it and
applies no curated labels when it is absent.
the environment, never prints it). SDK `typesafe-sdk==0.7.0`, Python 3.12.8.
- Model pinned to `jev-1.13.0` — never the `jev-latest` alias, because any threshold
tuned against an alias silently moves when the alias does (vendor's advice, adopted).
- State: one synthetic teach-back (1,873 chars) — a networking-background learner
Comment on lines +3 to +5
**Status:** council-validated planning round (3 seats, all ACCEPT-WITH-CORRECTIONS; every BLOCKING/MAJOR
finding verified against source — see `council/arbitration-plan-2026-09-19.md`). **Owner decisions
outstanding:** §7. **Base:** `main` @ `4f8e3e0f`.
@NetDevAutomate

Copy link
Copy Markdown
Owner Author

Full suite on this head (c1a28de1)

uv run --group dev pytest -q -p no:cacheprovider -rfE on the macOS sandbox host: 7316 passed, 31 failed, 14 errors, 17 skipped in 16m07s.

Failing/erroring ids (45) compared against the committed environmental set in docs/architecture/plan-integration/receipts/full-suite-control-item4-2026-09-18.md (51 ids = 44 environmental + item 4's seven then-REDs):

  • run − environmental set = one id: packages/agent-session-tools/tests/test_sync_conversation_integrity.py::test_concatenated_remote_dump_with_existing_archive — pytest-timeout at 60 s inside subprocess.run(["sqlite3", "-bail", …]).communicate(). It fails identically on a clean main control worktree with the same venv, and this branch does not touch packages/agent-session-tools (git diff --stat main..HEAD -- packages/agent-session-tools is empty). Pre-existing on this host; not a regression; not fixed here — it needs its own look (the host sqlite3 CLI never returning to the pipe).
  • environmental set − run = the seven item-4 REDs, green since 2026-09-18, as expected.

So: zero regressions attributable to this branch; the 44 shared ids are the known sandbox set.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants