Header labels and six doctor rows that gave wrong advice - #50
Conversation
Andy sent a screenshot of the web header and six doctor rows with a question each. Thirteen tests fail for the reason the report gives; one is a guard. - header: the voice-engine badge has no CSS rule for any state the script assigns, and the voice and theme pickers have no visible label while Font and Size do; - project_aliases is reported as an inert key although agent-session-tools' build_project_filter reads it (checked as a class: every direct load_config() read must name a known key); - grok is reported as "No manifest entry" although it reads the repo-root AGENTS.md that the codex/AGENTS.md entry tracks; - the xTiles row calls the wind-down skill "opt-in" and says nothing about where an xtiles MCP server is registered; - export freshness reports only the newest message across every harness, which hid a 22-day Kiro gap, and cannot see kiro-cli's new session store (the no-false-positive case is a guard and passes now); - the Kokoro model row appears for backends that never read those files; - "Obsidian export disabled" gives no reason and no next step. Two signature stubs keep the RED type-checked so it fails on behaviour: installers.xtiles_mcp_harnesses() returns [] and check_export_freshness() accepts and ignores kiro_sessions_dir.
Reported with a screenshot: "System voices" rendered as large bold text between the voice and theme pickers, and those two pickers had no visible label while Font and Size did. The badge's classes (tts-engine-badge + pending/ok/degraded) never had a CSS rule, so the span inherited the header's text styles. It is now a small pill in the voice field's label row, using the .bd-chip colours: green when the host's Kokoro is speaking, yellow when something else is (the "degraded" state the badge exists to make noticeable), italic while detecting. The voice and theme pickers get the same visible label as Font and Size, so the header reads VOICE / THEME / FONT / SIZE. The voice field hides as a whole when voice is off (x-show plus x-cloak), so no label is left alone. The controls now share one bottom edge (align-items: flex-end); stretch had pulled the unlabelled buttons and pickers up to the labelled fields' height. Checked in a browser against this tree: labels VOICE, THEME, FONT, SIZE; badge 10.2 px, weight 600, yellow; all seven visible controls end at the same y. Static and web suites 258 passed, JS 167/167.
The unknown-key check said project_aliases was "silently inert" and told the learner to remove it. It is not inert: agent-session-tools' build_project_filter reads it from the same config.yaml to let session search find a project's history under old paths, usernames and worktrees. On the machine that reported it, following the advice would have cut `session-query search --project` off from 732 of that project's sessions. It is read with load_config().get() rather than through DEFAULT_CONFIG, which is why the known-key list missed it. The RED's class-level test now scans every direct load_config() read in agent-session-tools, so the next such key fails a test instead of reaching a learner's doctor output.
…sing "No manifest entry for grok" was a row about a file that cannot exist. Grok Build has no definition of its own by design: its installer links the same repo-root AGENTS.md as Codex, which the codex/AGENTS.md manifest entry tracks. The check looked only for a grok/ key, found none, and left the file Grok actually reads unchecked. A harness with no key of its own is now checked against the manifest entries of the files installers._TOOL_LINKS links for it, and the row says so: "grok agent definition current (shares codex/AGENTS.md)". A stale shared file warns exactly as a stale own file does.
"Obsidian export disabled" gave no reason and no next step, and a learner on second_brain.provider: xtiles read it as "off because Obsidian is not my second brain". The two are independent: this is the session-memory export, switched by obsidian.export_enabled; second_brain only chooses where study notes go. The row now names the switch, says it is independent of second_brain, and gives the folder it would write to: To write session notes into ~/Obsidian/Personal/AgentMemory, set export_enabled: true under the obsidian: section of config.yaml
"Kokoro model files are not pre-warmed" appeared for every backend and pointed at a manual download. Only study-speak's local kokoro backend reads ~/.cache/kokoro-onnx, and it downloads the files itself on first use. With tts.backend: openvox the row asked for about 354 MB (325.5 MB model plus 28.2 MB voices, measured from the release URLs) that nothing on the machine would read. The row now appears only for backend kokoro, gives the size, and says the download happens on first use, with the docs pointer for fetching early.
"no programmatic backend; prompts and an opt-in assistant skill" read as "the skill still needs installing". It does not: `studyloop install agents` links the wind-down skill for every harness, and "opt-in" is the learner's yes at wind-down. What decides whether the offer can ever appear is an MCP server named xtiles in the session, and the row never said where one was. The row now says whether the skill is installed and which harness configs register an xtiles MCP server, or that none does. New read-only helper installers.xtiles_mcp_harnesses() reads the Claude, Codex, Kiro and OpenCode MCP configs that _mcp_config_path already names; Grok Build registers through its own CLI and is not read. On the reporting machine the answer is "codex" only, which is why no Kiro or Claude session could make the offer.
export_freshness reported the newest message across every harness. On the reporting machine that was Codex at 48 h, which hid that Kiro, the harness in daily use, had exported nothing for 22 days. The cause is #49: kiro-cli 2.x keeps sessions in ~/.kiro/sessions/cli/ and has not written to the data.sqlite3 tables the exporter reads since 5 Sep. The 4-hourly sweep still exits 0 with "added: 0", so nothing else reports it. The row now does the investigation itself: - a stale warning lists every harness's own age, newest first; - if kiro-cli's session store changed well after the newest kiro_cli export, the row says exactly that and points at #49. This warns even when another harness is fresh, because the fresh one is what hid it. The probe is one stat of the directory, and the constant says to retire it when #49 lands. Against the real database it now reads: "kiro_cli's newest export is 22 d old, but kiro-cli wrote to ~/.kiro/sessions/cli under 1 h ago ... By harness: codex 2 d, pi 12 d, claude_code 12 d, kiro_cli 22 d, opencode 196 d".
…vice One Fixed entry for the header and one for the doctor rows, with #49 named for the Kiro session store the exporter does not read yet.
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
Three moderate findings remain regarding session-file mtimes, OpenCode timestamp handling, and Grok xTiles detection.
Review effort: Lite
Findings: 1
Open (1)
What changed in this PR
Improves header usability and corrects several studyloop doctor diagnostics with regression tests and changelog updates.
Changes:
- Labels and aligns header controls and styles the voice badge.
- Corrects configuration, agent, voice, xTiles, Obsidian, and export diagnostics.
- Adds targeted regression coverage.
| File | Summary |
|---|---|
packages/studyloop/tests/test_web_header_controls.py |
Tests header labels and layout. |
packages/studyloop/tests/test_voice_backends.py |
Tests backend-specific voice diagnostics. |
packages/studyloop/tests/test_settings_custom.py |
Tests project_aliases recognition. |
packages/studyloop/tests/test_mcp_registration.py |
Tests xTiles MCP detection. |
packages/studyloop/tests/test_doctor_second_brain.py |
Tests xTiles doctor messaging. |
packages/studyloop/tests/test_doctor_exporter.py |
Tests export freshness diagnostics. |
packages/studyloop/tests/test_doctor_config.py |
Tests Obsidian diagnostics. |
packages/studyloop/tests/test_doctor_agents.py |
Tests shared agent definitions. |
packages/studyloop/src/studyloop/web/static/style.css |
Styles and aligns header controls. |
packages/studyloop/src/studyloop/web/static/index.html |
Adds header labels and voice grouping. |
packages/studyloop/src/studyloop/settings.py |
Recognizes project_aliases. |
packages/studyloop/src/studyloop/installers.py |
Detects xTiles MCP registrations. |
packages/studyloop/src/studyloop/doctor/voice.py |
Limits Kokoro checks to the Kokoro backend. |
packages/studyloop/src/studyloop/doctor/exporter.py |
Adds per-harness freshness and Kiro-store diagnostics; follow-up is needed for session-file mtimes and OpenCode timestamps. |
packages/studyloop/src/studyloop/doctor/config.py |
Improves Obsidian and xTiles guidance. |
packages/studyloop/src/studyloop/doctor/agents.py |
Handles shared agent definitions. |
CHANGELOG.md |
Documents the fixes. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
| try: | ||
| store_changed = datetime.fromtimestamp(store.stat().st_mtime, tz=UTC) | ||
| except OSError: | ||
| store_changed = None |
|
Full suite on head The 45 failing and erroring ids were compared with the recorded machine-specific set (
CI 15/15 green. |
Reported again on 2026-09-28 with a screenshot: "the top panel of the web ui is still not neat/central/aligned". The first fix on this branch was checked by parsing index.html and style.css, which cannot see alignment. Measured in a browser it is still wrong: * the three kinds of control are three heights -- the voice picker 26px, the buttons 30.5px, the other pickers 32px -- so the voice picker's top edge sits below its neighbours'; * the engine pill makes the voice label row taller, so VOICE sits 3px lower than THEME, FONT and SIZE; * the controls are bottom-aligned under labels that take part in layout, so the row sits 6.5px below the brand's centre line; * at a 1024px tablet width with the semantic chip showing, the controls run 75px past the window and under the brand. Two tests, laptop (1440) and tablet (1024), each run for every label the engine badge can show, with voice on and the semantic chip visible so every optional item is present. They assert one control height, one centre line per row (the brand's, when there is one row), labels inside the header sharing a bottom edge, no overlaps, and no horizontal overflow. Both fail on this commit for the reasons above: heights [26, 30.5, 32] at 1440px, and 75px of overflow at 1024px.
The header's controls could not be aligned while labels took part in layout: a labelled field was taller than the buttons beside it, so aligning tops, bottoms or middles was wrong for one kind of item every time. On top of that the three kinds of control were three heights (26px, 30.5px, 32px). * Field labels sit above their control but out of flow (absolute, in the header's top padding, now 20px), so every item in the row is one 32px control and align-items:center puts them all on the brand's centre line. * One 32px height for header buttons and pickers; the voice picker takes the theme/font/size pickers' font, colour and shape instead of a smaller, dimmer 12px style of its own. * The engine badge is a coloured dot and small text, not a pill. The pill was taller than a label, which put VOICE 3px below THEME, FONT and SIZE, and it was the loudest thing in a row of quiet labels. Degraded keeps the warn colour. * The controls wrap instead of overflowing, with 24px between rows for a wrapped row's labels, and the brand no longer shrinks: at a 1024px tablet width with the semantic chip showing they ran 75px past the window. Measured in the user's state (voice on, "Kokoro (server)", Catppuccin Mocha) at 1440px: every control 32px with its centre at 36px, the brand's centre at 36px; every label spans 6.1-17px. Both header geometry tests pass; the rest of the layout-regression file and the static header tests pass (26/26).
Found while checking the header fix: after style.css changed on disk, a reload still drew the old header. / is no-store, but the static mount sends style.css, components.js and every js/ module with only an ETag and Last-Modified, so a browser applies heuristic freshness (commonly a tenth of the file's age) and can reuse a weeks-old stylesheet for days. After git pull the learner gets new HTML beside old CSS and JS. Four assets fail for that reason (Cache-Control is None); the 304 guard passes, so revalidation stays cheap once the header is added.
The static mount is now a StaticFiles subclass that adds Cache-Control: no-cache to every response. The browser keeps its copy and asks first; an unchanged file costs a 304 against the ETag Starlette already sends. Without it, an update could pair the new no-store index.html with a stylesheet or script the browser had decided was still fresh for days -- which is what the header fix on this branch would have looked like after a merge. test_web_static_cache: 5/5. test_web_app, test_web_dev_engines, the wheel smoke: green. ruff, format, pyright clean.
The first entry described bottom-aligned controls and a pill badge, which this branch replaced after measuring them; it also gains the static-asset Cache-Control fix.

Andy's 2026-09-27 report: a screenshot of the web header, plus six
studyloop doctorrows, each with a question. Every one of them turned out to be the doctor or the page saying something wrong or incomplete. Built test-first: RED05b3d026(13 failing plus one guard), then one GREEN commit per concern.Header
First pass (
bdc6a44a): the voice-engine badge ("System voices") never had a CSS rule, so it rendered as large bold header text; Voice and Theme got visible labels like Font and Size.Reported again on 2026-09-28 ("still not neat/central/aligned"). The first pass had only been checked by parsing HTML and CSS; measured in a browser it was still wrong, so it was redone against geometry tests (RED
4aad237e, GREEN99df223e):Tests measure the rendered header at 1440px and 1024px for every label the badge can show: one control height, one centre line per row, labels inside the header on a shared bottom edge, no overlaps, no overflow.
CSS and JS were never revalidated (
20a52746RED,93f30789GREEN)Found while checking the header: after
style.csschanged, a reload still drew the old header./isno-store, but the static mount sent CSS and JS with no Cache-Control, so a browser applies heuristic freshness and can reuse a weeks-old stylesheet for days after an update. The mount now addsCache-Control: no-cache(a 304 when unchanged).Doctor
556d2443project_aliasesis not inert: agent-session-tools'build_project_filterreads it. Deleting it as advised would have cut 732 sessions out of one project's--projectsearch. A class test now scans every directload_config()read.f6c656a3Grok Build is checked against the repo-rootAGENTS.mdit shares with Codex (codex/AGENTS.md), instead of "No manifest entry for grok".36d10fa8The xTiles row says whether the wind-down skill is installed and which harness configs register anxtilesMCP server. New read-onlyinstallers.xtiles_mcp_harnesses().942d253aA staleexport_freshnessrow lists each harness's own age. It also detects kiro-cli's session store changing after the last Kiro export, which is Kiro sessions since 5 Sep are not exported: kiro-cli moved to ~/.kiro/sessions/cli #49: 22 days of Kiro sessions not exported. On the real database it reads "kiro_cli's newest export is 22 d old, but kiro-cli wrote to ~/.kiro/sessions/cli under 1 h ago … By harness: codex 2 d, pi 12 d, claude_code 12 d, kiro_cli 22 d, opencode 196 d".6b489e59The Kokoro model-files row appears only fortts.backend: kokoro, the one backend that reads those files (about 354 MB, downloaded on first use).0a69b67a"Obsidian export disabled" namesobsidian.export_enabled, says it is independent ofsecond_brain, and gives the folder it would write to.Tests
test_rows_vault_missing_warnsfails on this machine before and after this branch. It is in the recorded machine-specific set (full-suite-control-item4-2026-09-18.md).Not in this PR