Skip to content

Header labels and six doctor rows that gave wrong advice - #50

Merged
NetDevAutomate merged 14 commits into
mainfrom
fix/header-and-doctor-advice
Sep 28, 2026
Merged

NetDevAutomate merged 14 commits into
mainfrom
fix/header-and-doctor-advice

Conversation

@NetDevAutomate

@NetDevAutomate NetDevAutomate commented Sep 27, 2026 •

Copy link
Copy Markdown
Owner

Andy's 2026-09-27 report: a screenshot of the web header, plus six studyloop doctor rows, each with a question. Every one of them turned out to be the doctor or the page saying something wrong or incomplete. Built test-first: RED 05b3d026 (13 failing plus one guard), then one GREEN commit per concern.

Header

First pass (bdc6a44a): the voice-engine badge ("System voices") never had a CSS rule, so it rendered as large bold header text; Voice and Theme got visible labels like Font and Size.

Reported again on 2026-09-28 ("still not neat/central/aligned"). The first pass had only been checked by parsing HTML and CSS; measured in a browser it was still wrong, so it was redone against geometry tests (RED 4aad237e, GREEN 99df223e):

  • The three kinds of control were three heights (voice picker 26px, buttons 30.5px, pickers 32px). Now one 32px height; the voice picker takes the other pickers' font and shape.
  • Labels took part in layout, so the row could not be aligned for every item. They now sit above their control out of flow, and every control is on the brand's centre line (measured: all controls centred at 36px, brand at 36px, in the user's state with "Kokoro (server)" and Catppuccin Mocha).
  • The engine badge is a coloured dot and small text, not a pill: the pill put VOICE 3px below the other labels.
  • At a 1024px tablet width with the semantic chip showing, the controls ran 75px past the window and under the brand. They now wrap.

Tests measure the rendered header at 1440px and 1024px for every label the badge can show: one control height, one centre line per row, labels inside the header on a shared bottom edge, no overlaps, no overflow.

CSS and JS were never revalidated (20a52746 RED, 93f30789 GREEN)

Found while checking the header: after style.css changed, a reload still drew the old header. / is no-store, but the static mount sent CSS and JS with no Cache-Control, so a browser applies heuristic freshness and can reuse a weeks-old stylesheet for days after an update. The mount now adds Cache-Control: no-cache (a 304 when unchanged).

Doctor

  • 556d2443 project_aliases is not inert: agent-session-tools' build_project_filter reads it. Deleting it as advised would have cut 732 sessions out of one project's --project search. A class test now scans every direct load_config() read.
  • f6c656a3 Grok Build is checked against the repo-root AGENTS.md it shares with Codex (codex/AGENTS.md), instead of "No manifest entry for grok".
  • 36d10fa8 The xTiles row says whether the wind-down skill is installed and which harness configs register an xtiles MCP server. New read-only installers.xtiles_mcp_harnesses().
  • 942d253a A stale export_freshness row lists each harness's own age. It also detects kiro-cli's session store changing after the last Kiro export, which is Kiro sessions since 5 Sep are not exported: kiro-cli moved to ~/.kiro/sessions/cli #49: 22 days of Kiro sessions not exported. On the real database it reads "kiro_cli's newest export is 22 d old, but kiro-cli wrote to ~/.kiro/sessions/cli under 1 h ago … By harness: codex 2 d, pi 12 d, claude_code 12 d, kiro_cli 22 d, opencode 196 d".
  • 6b489e59 The Kokoro model-files row appears only for tts.backend: kokoro, the one backend that reads those files (about 354 MB, downloaded on first use).
  • 0a69b67a "Obsidian export disabled" names obsidian.export_enabled, says it is independent of second_brain, and gives the folder it would write to.

Tests

  • New and affected suites pass: header geometry 2/2 plus the rest of the layout-regression file (26/26 with the static header tests), static cache 5/5, web app 28, static and web 258, settings and docs-drift 488, agents 84, second brain, MCP and xTiles 103, exporter 48, voice 41. JS 167/167.
  • test_rows_vault_missing_warns fails on this machine before and after this branch. It is in the recorded machine-specific set (full-suite-control-item4-2026-09-18.md).
  • Full-suite comparison: see the comments below.

Not in this PR

Andy sent a screenshot of the web header and six doctor rows with a question
each. Thirteen tests fail for the reason the report gives; one is a guard.

- header: the voice-engine badge has no CSS rule for any state the script
  assigns, and the voice and theme pickers have no visible label while Font
  and Size do;
- project_aliases is reported as an inert key although agent-session-tools'
  build_project_filter reads it (checked as a class: every direct load_config()
  read must name a known key);
- grok is reported as "No manifest entry" although it reads the repo-root
  AGENTS.md that the codex/AGENTS.md entry tracks;
- the xTiles row calls the wind-down skill "opt-in" and says nothing about
  where an xtiles MCP server is registered;
- export freshness reports only the newest message across every harness,
  which hid a 22-day Kiro gap, and cannot see kiro-cli's new session store
  (the no-false-positive case is a guard and passes now);
- the Kokoro model row appears for backends that never read those files;
- "Obsidian export disabled" gives no reason and no next step.

Two signature stubs keep the RED type-checked so it fails on behaviour:
installers.xtiles_mcp_harnesses() returns [] and check_export_freshness()
accepts and ignores kiro_sessions_dir.
Reported with a screenshot: "System voices" rendered as large bold text
between the voice and theme pickers, and those two pickers had no visible
label while Font and Size did.

The badge's classes (tts-engine-badge + pending/ok/degraded) never had a CSS
rule, so the span inherited the header's text styles. It is now a small pill
in the voice field's label row, using the .bd-chip colours: green when the
host's Kokoro is speaking, yellow when something else is (the "degraded"
state the badge exists to make noticeable), italic while detecting.

The voice and theme pickers get the same visible label as Font and Size, so
the header reads VOICE / THEME / FONT / SIZE. The voice field hides as a whole
when voice is off (x-show plus x-cloak), so no label is left alone. The
controls now share one bottom edge (align-items: flex-end); stretch had pulled
the unlabelled buttons and pickers up to the labelled fields' height.

Checked in a browser against this tree: labels VOICE, THEME, FONT, SIZE; badge
10.2 px, weight 600, yellow; all seven visible controls end at the same y.
Static and web suites 258 passed, JS 167/167.
The unknown-key check said project_aliases was "silently inert" and told the
learner to remove it. It is not inert: agent-session-tools' build_project_filter
reads it from the same config.yaml to let session search find a project's
history under old paths, usernames and worktrees. On the machine that reported
it, following the advice would have cut `session-query search --project` off
from 732 of that project's sessions.

It is read with load_config().get() rather than through DEFAULT_CONFIG, which
is why the known-key list missed it. The RED's class-level test now scans every
direct load_config() read in agent-session-tools, so the next such key fails a
test instead of reaching a learner's doctor output.
…sing

"No manifest entry for grok" was a row about a file that cannot exist. Grok
Build has no definition of its own by design: its installer links the same
repo-root AGENTS.md as Codex, which the codex/AGENTS.md manifest entry tracks.
The check looked only for a grok/ key, found none, and left the file Grok
actually reads unchecked.

A harness with no key of its own is now checked against the manifest entries
of the files installers._TOOL_LINKS links for it, and the row says so:
"grok agent definition current (shares codex/AGENTS.md)". A stale shared file
warns exactly as a stale own file does.
"Obsidian export disabled" gave no reason and no next step, and a learner on
second_brain.provider: xtiles read it as "off because Obsidian is not my
second brain". The two are independent: this is the session-memory export,
switched by obsidian.export_enabled; second_brain only chooses where study
notes go.

The row now names the switch, says it is independent of second_brain, and
gives the folder it would write to:
  To write session notes into ~/Obsidian/Personal/AgentMemory, set
  export_enabled: true under the obsidian: section of config.yaml
"Kokoro model files are not pre-warmed" appeared for every backend and pointed
at a manual download. Only study-speak's local kokoro backend reads
~/.cache/kokoro-onnx, and it downloads the files itself on first use. With
tts.backend: openvox the row asked for about 354 MB (325.5 MB model plus
28.2 MB voices, measured from the release URLs) that nothing on the machine
would read.

The row now appears only for backend kokoro, gives the size, and says the
download happens on first use, with the docs pointer for fetching early.
"no programmatic backend; prompts and an opt-in assistant skill" read as "the
skill still needs installing". It does not: `studyloop install agents` links
the wind-down skill for every harness, and "opt-in" is the learner's yes at
wind-down. What decides whether the offer can ever appear is an MCP server
named xtiles in the session, and the row never said where one was.

The row now says whether the skill is installed and which harness configs
register an xtiles MCP server, or that none does. New read-only helper
installers.xtiles_mcp_harnesses() reads the Claude, Codex, Kiro and OpenCode
MCP configs that _mcp_config_path already names; Grok Build registers through
its own CLI and is not read. On the reporting machine the answer is "codex"
only, which is why no Kiro or Claude session could make the offer.
export_freshness reported the newest message across every harness. On the
reporting machine that was Codex at 48 h, which hid that Kiro, the harness in
daily use, had exported nothing for 22 days.

The cause is #49: kiro-cli 2.x keeps sessions in ~/.kiro/sessions/cli/ and has
not written to the data.sqlite3 tables the exporter reads since 5 Sep. The
4-hourly sweep still exits 0 with "added: 0", so nothing else reports it.

The row now does the investigation itself:
- a stale warning lists every harness's own age, newest first;
- if kiro-cli's session store changed well after the newest kiro_cli export,
  the row says exactly that and points at #49. This warns even when another
  harness is fresh, because the fresh one is what hid it. The probe is one
  stat of the directory, and the constant says to retire it when #49 lands.

Against the real database it now reads: "kiro_cli's newest export is 22 d old,
but kiro-cli wrote to ~/.kiro/sessions/cli under 1 h ago ... By harness:
codex 2 d, pi 12 d, claude_code 12 d, kiro_cli 22 d, opencode 196 d".
…vice

One Fixed entry for the header and one for the doctor rows, with #49 named for the Kiro session store the exporter does not read yet.
Copilot AI lite review requested due to automatic review settings September 27, 2026 14:01

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Three moderate findings remain regarding session-file mtimes, OpenCode timestamp handling, and Grok xTiles detection.

Review effort: Lite
Findings: 1 Medium severity

Open (1)
What changed in this PR

Improves header usability and corrects several studyloop doctor diagnostics with regression tests and changelog updates.

Changes:

  • Labels and aligns header controls and styles the voice badge.
  • Corrects configuration, agent, voice, xTiles, Obsidian, and export diagnostics.
  • Adds targeted regression coverage.
File Summary
packages/​studyloop/​tests/​test_web_header_controls.py Tests header labels and layout.
packages/​studyloop/​tests/​test_voice_backends.py Tests backend-specific voice diagnostics.
packages/​studyloop/​tests/​test_settings_custom.py Tests project_aliases recognition.
packages/​studyloop/​tests/​test_mcp_registration.py Tests xTiles MCP detection.
packages/​studyloop/​tests/​test_doctor_second_brain.py Tests xTiles doctor messaging.
packages/​studyloop/​tests/​test_doctor_exporter.py Tests export freshness diagnostics.
packages/​studyloop/​tests/​test_doctor_config.py Tests Obsidian diagnostics.
packages/​studyloop/​tests/​test_doctor_agents.py Tests shared agent definitions.
packages/​studyloop/​src/​studyloop/​web/​static/​style.css Styles and aligns header controls.
packages/​studyloop/​src/​studyloop/​web/​static/​index.html Adds header labels and voice grouping.
packages/​studyloop/​src/​studyloop/​settings.py Recognizes project_aliases.
packages/​studyloop/​src/​studyloop/​installers.py Detects xTiles MCP registrations.
packages/​studyloop/​src/​studyloop/​doctor/​voice.py Limits Kokoro checks to the Kokoro backend.
packages/​studyloop/​src/​studyloop/​doctor/​exporter.py Adds per-harness freshness and Kiro-store diagnostics; follow-up is needed for session-file mtimes and OpenCode timestamps.
packages/​studyloop/​src/​studyloop/​doctor/​config.py Improves Obsidian and xTiles guidance.
packages/​studyloop/​src/​studyloop/​doctor/​agents.py Handles shared agent definitions.
CHANGELOG.md Documents the fixes.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment on lines +280 to +283
try:
store_changed = datetime.fromtimestamp(store.stat().st_mtime, tz=UTC)
except OSError:
store_changed = None
@NetDevAutomate

Copy link
Copy Markdown
Owner Author

Full suite on head e528180c, run in the branch's own worktree venv (studyloop.__file__ resolves to the worktree): 7392 passed, 31 failed, 14 errors, 16 skipped (972 s).

The 45 failing and erroring ids were compared with the recorded machine-specific set (docs/architecture/plan-integration/receipts/full-suite-control-item4-2026-09-18.md, 51 ids):

CI 15/15 green.

Reported again on 2026-09-28 with a screenshot: "the top panel of the web ui
is still not neat/central/aligned". The first fix on this branch was checked by
parsing index.html and style.css, which cannot see alignment. Measured in a
browser it is still wrong:

* the three kinds of control are three heights -- the voice picker 26px, the
  buttons 30.5px, the other pickers 32px -- so the voice picker's top edge sits
  below its neighbours';
* the engine pill makes the voice label row taller, so VOICE sits 3px lower
  than THEME, FONT and SIZE;
* the controls are bottom-aligned under labels that take part in layout, so
  the row sits 6.5px below the brand's centre line;
* at a 1024px tablet width with the semantic chip showing, the controls run
  75px past the window and under the brand.

Two tests, laptop (1440) and tablet (1024), each run for every label the
engine badge can show, with voice on and the semantic chip visible so every
optional item is present. They assert one control height, one centre line per
row (the brand's, when there is one row), labels inside the header sharing a
bottom edge, no overlaps, and no horizontal overflow.

Both fail on this commit for the reasons above: heights [26, 30.5, 32] at
1440px, and 75px of overflow at 1024px.
The header's controls could not be aligned while labels took part in layout:
a labelled field was taller than the buttons beside it, so aligning tops,
bottoms or middles was wrong for one kind of item every time. On top of that
the three kinds of control were three heights (26px, 30.5px, 32px).

* Field labels sit above their control but out of flow (absolute, in the
  header's top padding, now 20px), so every item in the row is one 32px
  control and align-items:center puts them all on the brand's centre line.
* One 32px height for header buttons and pickers; the voice picker takes the
  theme/font/size pickers' font, colour and shape instead of a smaller,
  dimmer 12px style of its own.
* The engine badge is a coloured dot and small text, not a pill. The pill was
  taller than a label, which put VOICE 3px below THEME, FONT and SIZE, and it
  was the loudest thing in a row of quiet labels. Degraded keeps the warn
  colour.
* The controls wrap instead of overflowing, with 24px between rows for a
  wrapped row's labels, and the brand no longer shrinks: at a 1024px tablet
  width with the semantic chip showing they ran 75px past the window.

Measured in the user's state (voice on, "Kokoro (server)", Catppuccin Mocha)
at 1440px: every control 32px with its centre at 36px, the brand's centre at
36px; every label spans 6.1-17px.

Both header geometry tests pass; the rest of the layout-regression file and
the static header tests pass (26/26).
Found while checking the header fix: after style.css changed on disk, a reload still drew the old header. / is no-store, but the static mount sends style.css, components.js and every js/ module with only an ETag and Last-Modified, so a browser applies heuristic freshness (commonly a tenth of the file's age) and can reuse a weeks-old stylesheet for days. After git pull the learner gets new HTML beside old CSS and JS.

Four assets fail for that reason (Cache-Control is None); the 304 guard passes, so revalidation stays cheap once the header is added.
The static mount is now a StaticFiles subclass that adds Cache-Control: no-cache to every response. The browser keeps its copy and asks first; an unchanged file costs a 304 against the ETag Starlette already sends. Without it, an update could pair the new no-store index.html with a stylesheet or script the browser had decided was still fresh for days -- which is what the header fix on this branch would have looked like after a merge.

test_web_static_cache: 5/5. test_web_app, test_web_dev_engines, the wheel smoke: green. ruff, format, pyright clean.
The first entry described bottom-aligned controls and a pill badge, which this branch replaced after measuring them; it also gains the static-asset Cache-Control fix.
@NetDevAutomate

Copy link
Copy Markdown
Owner Author

Full-suite comparison for this branch's header and cache commits: run on #51's tree, which contains all of #50 — 44 of 45 failing ids in the recorded machine-specific set, the 45th the known agent-session-tools sqlite3 timeout. Details on #51.

@NetDevAutomate
NetDevAutomate merged commit 66577cb into main Sep 28, 2026
15 checks passed
@NetDevAutomate
NetDevAutomate deleted the fix/header-and-doctor-advice branch September 28, 2026 14:59
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants