Skip to content

Body Double for anything; a live session is the agent's screen - #51

Merged
NetDevAutomate merged 16 commits into
mainfrom
fix/body-double-calm
Sep 28, 2026
Merged

NetDevAutomate merged 16 commits into
mainfrom
fix/body-double-calm

Conversation

@NetDevAutomate

@NetDevAutomate NetDevAutomate commented Sep 28, 2026 •

Copy link
Copy Markdown
Owner

Andy's 2026-09-28 report, with two screenshots: "I should be able to run a body double session for anything I'm doing, no restrictions or list of Focus areas. If I start one any way it allows me but I end up with a very 'busy' screen with the agent not being a clear central focus."

Stacked on #50: this branch contains #50's commits, and targets main because CI only runs for PRs into main. Review #50 first; fast-forwarding main to this branch lands both.

Test-first: RED 7dde69cd (six failing for the reported reasons), GREEN 8fb7c807.

No Focus card on Body Double

It listed the three most recent pending study topics (on Andy's machine, auto-captured questions from other sessions) with an "at capacity" chip, above the picker and above the terminal. The activity field was always free text, so it restricted nothing; it only looked as if it did, on a surface whose own spec says body doubling is not a study thread (ADR-0003).

  • The three-topic rule is unchanged where study threads start (the Study picker's park-first check). GET/POST /api/body-double/focus and their unit tests stay; the web UI no longer calls them, and committed focus is managed with studyloop focus.
  • Notes taken in Body Double are filed under the activity, not offered a list of study topics.

A live session is the agent's screen

Measured before, at 1440x900: the view's heading, the 25:00 timer block and the Focus card put the terminal 432px down the content area, its bottom fell below the window, Capture sat wholly below the fold, and the Park-a-thought button covered the terminal's last line.

  • The view goes full-height while live, like the Study view (absolute, inset 0 inside .content-area). Its heading and big timer step aside; the session strip carries the activity, the Pomodoro (time, Start/Pause/Resume, Break) and End session.
  • The console fills the rest of the window: 148-827px at 1440x900 (82% of the content area), 79% at 1024x768.
  • Capture folds to one row under it while live (one $watch on sessionActive, so start, reattach and end agree). Note and Park open it; ending gives back the learner's own idle choice.
  • The floating Park-a-thought button and Pomodoro widget step aside during a live Body Double session. The P shortcut still parks.
  • Also: the console's status line and dot had no CSS ("Connected · kiro" was large bold text over the terminal), and the idle 25:00 sat hard left under a centred settings row.

Two defects found by measuring this change before calling it done, each now pinned by a test:

  • With Capture open at 1024x768, the console spilled 90px under the note form (its 240px floor sat on the console, not its section).
  • Starting the Pomodoro put the floating widget on End session; only a hit test (elementFromPoint) showed it.

Tests

  • Body Double e2e (workspace + journey): 38/38.
  • Retired with the card, their behaviour gone: the Focus-pane fold, the four committed-focus tests, the focus-slot note topic.
  • Updated: the fold-reload test (Capture only), the Pomodoro end test (starts it from the strip), the post-end test (opens the folded Capture first), journey phases 1-3 and 9 (no card; the rule of three read from the API; the activity typed, not picked).
  • Session lifecycle, first move, remaining surface, layout regression, smoke browser, session recovery, header, static cache, notes-and-focus: 129 passed, 2 skipped. JS 167/167. mkdocs build --strict clean.
  • Full-suite comparison: see the comment below.

Not in this PR

  • Web Body Double and Study sessions still start with the launcher's generic fallback persona (mode focus has no personas/focus.md), which tells the agent to drive the session. That is the next fix, not this one.

Andy sent a screenshot of the web header and six doctor rows with a question
each. Thirteen tests fail for the reason the report gives; one is a guard.

- header: the voice-engine badge has no CSS rule for any state the script
  assigns, and the voice and theme pickers have no visible label while Font
  and Size do;
- project_aliases is reported as an inert key although agent-session-tools'
  build_project_filter reads it (checked as a class: every direct load_config()
  read must name a known key);
- grok is reported as "No manifest entry" although it reads the repo-root
  AGENTS.md that the codex/AGENTS.md entry tracks;
- the xTiles row calls the wind-down skill "opt-in" and says nothing about
  where an xtiles MCP server is registered;
- export freshness reports only the newest message across every harness,
  which hid a 22-day Kiro gap, and cannot see kiro-cli's new session store
  (the no-false-positive case is a guard and passes now);
- the Kokoro model row appears for backends that never read those files;
- "Obsidian export disabled" gives no reason and no next step.

Two signature stubs keep the RED type-checked so it fails on behaviour:
installers.xtiles_mcp_harnesses() returns [] and check_export_freshness()
accepts and ignores kiro_sessions_dir.
Reported with a screenshot: "System voices" rendered as large bold text
between the voice and theme pickers, and those two pickers had no visible
label while Font and Size did.

The badge's classes (tts-engine-badge + pending/ok/degraded) never had a CSS
rule, so the span inherited the header's text styles. It is now a small pill
in the voice field's label row, using the .bd-chip colours: green when the
host's Kokoro is speaking, yellow when something else is (the "degraded"
state the badge exists to make noticeable), italic while detecting.

The voice and theme pickers get the same visible label as Font and Size, so
the header reads VOICE / THEME / FONT / SIZE. The voice field hides as a whole
when voice is off (x-show plus x-cloak), so no label is left alone. The
controls now share one bottom edge (align-items: flex-end); stretch had pulled
the unlabelled buttons and pickers up to the labelled fields' height.

Checked in a browser against this tree: labels VOICE, THEME, FONT, SIZE; badge
10.2 px, weight 600, yellow; all seven visible controls end at the same y.
Static and web suites 258 passed, JS 167/167.
The unknown-key check said project_aliases was "silently inert" and told the
learner to remove it. It is not inert: agent-session-tools' build_project_filter
reads it from the same config.yaml to let session search find a project's
history under old paths, usernames and worktrees. On the machine that reported
it, following the advice would have cut `session-query search --project` off
from 732 of that project's sessions.

It is read with load_config().get() rather than through DEFAULT_CONFIG, which
is why the known-key list missed it. The RED's class-level test now scans every
direct load_config() read in agent-session-tools, so the next such key fails a
test instead of reaching a learner's doctor output.
…sing

"No manifest entry for grok" was a row about a file that cannot exist. Grok
Build has no definition of its own by design: its installer links the same
repo-root AGENTS.md as Codex, which the codex/AGENTS.md manifest entry tracks.
The check looked only for a grok/ key, found none, and left the file Grok
actually reads unchecked.

A harness with no key of its own is now checked against the manifest entries
of the files installers._TOOL_LINKS links for it, and the row says so:
"grok agent definition current (shares codex/AGENTS.md)". A stale shared file
warns exactly as a stale own file does.
"Obsidian export disabled" gave no reason and no next step, and a learner on
second_brain.provider: xtiles read it as "off because Obsidian is not my
second brain". The two are independent: this is the session-memory export,
switched by obsidian.export_enabled; second_brain only chooses where study
notes go.

The row now names the switch, says it is independent of second_brain, and
gives the folder it would write to:
  To write session notes into ~/Obsidian/Personal/AgentMemory, set
  export_enabled: true under the obsidian: section of config.yaml
"Kokoro model files are not pre-warmed" appeared for every backend and pointed
at a manual download. Only study-speak's local kokoro backend reads
~/.cache/kokoro-onnx, and it downloads the files itself on first use. With
tts.backend: openvox the row asked for about 354 MB (325.5 MB model plus
28.2 MB voices, measured from the release URLs) that nothing on the machine
would read.

The row now appears only for backend kokoro, gives the size, and says the
download happens on first use, with the docs pointer for fetching early.
"no programmatic backend; prompts and an opt-in assistant skill" read as "the
skill still needs installing". It does not: `studyloop install agents` links
the wind-down skill for every harness, and "opt-in" is the learner's yes at
wind-down. What decides whether the offer can ever appear is an MCP server
named xtiles in the session, and the row never said where one was.

The row now says whether the skill is installed and which harness configs
register an xtiles MCP server, or that none does. New read-only helper
installers.xtiles_mcp_harnesses() reads the Claude, Codex, Kiro and OpenCode
MCP configs that _mcp_config_path already names; Grok Build registers through
its own CLI and is not read. On the reporting machine the answer is "codex"
only, which is why no Kiro or Claude session could make the offer.
export_freshness reported the newest message across every harness. On the
reporting machine that was Codex at 48 h, which hid that Kiro, the harness in
daily use, had exported nothing for 22 days.

The cause is #49: kiro-cli 2.x keeps sessions in ~/.kiro/sessions/cli/ and has
not written to the data.sqlite3 tables the exporter reads since 5 Sep. The
4-hourly sweep still exits 0 with "added: 0", so nothing else reports it.

The row now does the investigation itself:
- a stale warning lists every harness's own age, newest first;
- if kiro-cli's session store changed well after the newest kiro_cli export,
  the row says exactly that and points at #49. This warns even when another
  harness is fresh, because the fresh one is what hid it. The probe is one
  stat of the directory, and the constant says to retire it when #49 lands.

Against the real database it now reads: "kiro_cli's newest export is 22 d old,
but kiro-cli wrote to ~/.kiro/sessions/cli under 1 h ago ... By harness:
codex 2 d, pi 12 d, claude_code 12 d, kiro_cli 22 d, opencode 196 d".
…vice

One Fixed entry for the header and one for the doctor rows, with #49 named for the Kiro session store the exporter does not read yet.
Reported again on 2026-09-28 with a screenshot: "the top panel of the web ui
is still not neat/central/aligned". The first fix on this branch was checked by
parsing index.html and style.css, which cannot see alignment. Measured in a
browser it is still wrong:

* the three kinds of control are three heights -- the voice picker 26px, the
  buttons 30.5px, the other pickers 32px -- so the voice picker's top edge sits
  below its neighbours';
* the engine pill makes the voice label row taller, so VOICE sits 3px lower
  than THEME, FONT and SIZE;
* the controls are bottom-aligned under labels that take part in layout, so
  the row sits 6.5px below the brand's centre line;
* at a 1024px tablet width with the semantic chip showing, the controls run
  75px past the window and under the brand.

Two tests, laptop (1440) and tablet (1024), each run for every label the
engine badge can show, with voice on and the semantic chip visible so every
optional item is present. They assert one control height, one centre line per
row (the brand's, when there is one row), labels inside the header sharing a
bottom edge, no overlaps, and no horizontal overflow.

Both fail on this commit for the reasons above: heights [26, 30.5, 32] at
1440px, and 75px of overflow at 1024px.
The header's controls could not be aligned while labels took part in layout:
a labelled field was taller than the buttons beside it, so aligning tops,
bottoms or middles was wrong for one kind of item every time. On top of that
the three kinds of control were three heights (26px, 30.5px, 32px).

* Field labels sit above their control but out of flow (absolute, in the
  header's top padding, now 20px), so every item in the row is one 32px
  control and align-items:center puts them all on the brand's centre line.
* One 32px height for header buttons and pickers; the voice picker takes the
  theme/font/size pickers' font, colour and shape instead of a smaller,
  dimmer 12px style of its own.
* The engine badge is a coloured dot and small text, not a pill. The pill was
  taller than a label, which put VOICE 3px below THEME, FONT and SIZE, and it
  was the loudest thing in a row of quiet labels. Degraded keeps the warn
  colour.
* The controls wrap instead of overflowing, with 24px between rows for a
  wrapped row's labels, and the brand no longer shrinks: at a 1024px tablet
  width with the semantic chip showing they ran 75px past the window.

Measured in the user's state (voice on, "Kokoro (server)", Catppuccin Mocha)
at 1440px: every control 32px with its centre at 36px, the brand's centre at
36px; every label spans 6.1-17px.

Both header geometry tests pass; the rest of the layout-regression file and
the static header tests pass (26/26).
Found while checking the header fix: after style.css changed on disk, a reload still drew the old header. / is no-store, but the static mount sends style.css, components.js and every js/ module with only an ETag and Last-Modified, so a browser applies heuristic freshness (commonly a tenth of the file's age) and can reuse a weeks-old stylesheet for days. After git pull the learner gets new HTML beside old CSS and JS.

Four assets fail for that reason (Cache-Control is None); the 304 guard passes, so revalidation stays cheap once the header is added.
The static mount is now a StaticFiles subclass that adds Cache-Control: no-cache to every response. The browser keeps its copy and asks first; an unchanged file costs a 304 against the ETag Starlette already sends. Without it, an update could pair the new no-store index.html with a stylesheet or script the browser had decided was still fresh for days -- which is what the header fix on this branch would have looked like after a merge.

test_web_static_cache: 5/5. test_web_app, test_web_dev_engines, the wheel smoke: green. ruff, format, pyright clean.
The first entry described bottom-aligned controls and a pill badge, which this branch replaced after measuring them; it also gains the static-asset Cache-Control fix.
…nt's screen

Reported 2026-09-28 with two screenshots: "I should be able to run a body
double session for anything I'm doing, no restrictions or list of Focus
areas. If I start one any way it allows me but I end up with a very 'busy'
screen with the agent not being a clear central focus."

The Focus card listed the three most recent pending study topics with an
"at capacity" chip, above the picker and above the live terminal. The
activity field was always free text, so the card restricted nothing; it only
looked as if it did, on a surface whose spec already says body doubling is
not a new study thread (ADR-0003). Measured live at 1440x900: the view's
heading, the 25:00 timer block and the Focus card put the terminal 432px down
the content area, its bottom fell below the window, the Capture card sat
wholly below the fold, and the Park-a-thought button covered the terminal's
bottom-right corner.

New tests, all red on this commit for those reasons:
* no Focus card, no study topic anywhere on the view (every slot filled
  first so a surviving card would have to show), note topics offer none;
* live at 1440x900 and 1024x768: no heading or timer block, one strip, then a
  console that reaches the window's bottom with >= 60% of the content area,
  covered by nothing;
* the strip carries the Pomodoro time and its control beside End;
* Capture folds for the session, a tab opens it without squeezing the agent
  below 240px, and ending gives back the learner's idle layout;
* a note is filed under what you are working on, not a study topic.

The bd_page fixture now waits for the picker rather than the Focus card.
Reported 2026-09-28: "I should be able to run a body double session for
anything I'm doing, no restrictions or list of Focus areas ... I end up with
a very 'busy' screen with the agent not being a clear central focus."

No Focus card. It listed the three most recent pending study topics (on the
learner's machine, auto-captured questions from other sessions) with an "at
capacity" chip, above the picker and above the terminal. The activity field
was always free text, so it restricted nothing; it only looked as if it did,
on a surface whose spec already says body doubling is not a study thread
(ADR-0003). The three-topic rule is unchanged where study threads start (the
Study picker's park-first check), GET /api/body-double/focus and its unit
tests stay, and committed focus is managed with `studyloop focus`. The note
composer offers the activity instead of study topics.

A live session is the agent's screen, the way the Study view already is
(absolute, inset 0 inside .content-area, so the column has a height to fill):
* the view's heading and big timer step aside; the strip carries the
  activity, the Pomodoro (time, Start/Pause/Resume, Break) and End session;
* the console fills the rest of the window;
* Capture folds to one row under it while live (one $watch on sessionActive,
  so start, reattach and end all agree); Note and Park open it, and ending
  gives back the learner's own idle choice;
* the floating Park-a-thought button and Pomodoro widget step aside while a
  Body Double session is live: they covered the terminal's last line and,
  once the strip moved up, the End button itself.
Also: the console's status line and dot had no CSS ("Connected · kiro" was
large bold text; the dot was invisible), and the idle 25:00 sat hard left
under a centred settings row.

Two defects found by measuring this change before calling it done, each now
pinned: with Capture open at 1024x768 the console kept a 240px floor its
section did not, so it spilled 90px under the note form (the floor moved to
the section); and starting the Pomodoro put the floating widget on End, which
only a hit test showed (the click failed after 30s with "intercepts pointer
events").

Measured at 1440x900: strip 85-136px, console 148-827px (82% of the content
area), Capture row 839-888px; at 1024x768 the console gets 79%. With Capture
open the console keeps 270px (laptop) / 257px (tablet) and nothing overlaps.

Tests retired with the card, their behaviour gone: the Focus-pane fold, the
four committed-focus tests, the focus-slot note topic. Updated: the fold
reload test (Capture only), the Pomodoro end test (starts it from the strip),
the post-end test (opens the folded Capture first), journey phases 1-3 and 9
(no card; the rule of three read from the API; the activity typed, not
picked). Body Double e2e 38/38, related web suites 129 passed, JS 167/167,
mkdocs --strict clean.
Copilot AI lite review requested due to automatic review settings September 28, 2026 13:49
@NetDevAutomate
NetDevAutomate changed the base branch from fix/header-and-doctor-advice to main September 28, 2026 13:49

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Moderate unresolved issues remain in exporter freshness detection, Grok xTiles detection, live-session reload state, stale activity selection, and navigation-scoped controls.

Review effort: Lite
Findings: 3 Medium severity

Open (3)
What changed in this PR

This PR broadens Body Double to arbitrary activities and makes live sessions console-first, while including header, caching, and doctor-diagnostic updates.

Changes:

  • Removes the Body Double Focus card and files notes under activities.
  • Reworks live sessions around the agent console with compact controls.
  • Improves static caching, diagnostics, documentation, and regression coverage.
File Summary
packages/​studyloop/​tests/​test_web_static_cache.py Static asset revalidation tests
packages/​studyloop/​tests/​test_web_layout_regression.py Header geometry tests
packages/​studyloop/​tests/​test_web_header_controls.py Header control tests
packages/​studyloop/​tests/​test_voice_backends.py Voice diagnostic tests
packages/​studyloop/​tests/​test_settings_custom.py Configuration-key tests
packages/​studyloop/​tests/​test_mcp_registration.py xTiles MCP detection tests
packages/​studyloop/​tests/​test_doctor_second_brain.py xTiles diagnostic tests
packages/​studyloop/​tests/​test_doctor_exporter.py Export freshness tests
packages/​studyloop/​tests/​test_doctor_config.py Obsidian diagnostic tests
packages/​studyloop/​tests/​test_doctor_agents.py Shared agent definition tests
packages/​studyloop/​tests/​e2e/​test_body_double_workspace.py Body Double layout and lifecycle tests
packages/​studyloop/​tests/​e2e/​test_body_double_journey.py Body Double journey coverage
packages/​studyloop/​src/​studyloop/​web/​static/​style.css Header and live-session styling
packages/​studyloop/​src/​studyloop/​web/​static/​index.html Header and Body Double markup
packages/​studyloop/​src/​studyloop/​web/​static/​components.js Live-session state and activity notes
packages/​studyloop/​src/​studyloop/​web/​app.py Static asset revalidation
packages/​studyloop/​src/​studyloop/​settings.py project_aliases recognition
packages/​studyloop/​src/​studyloop/​installers.py xTiles MCP registration detection
packages/​studyloop/​src/​studyloop/​doctor/​voice.py Backend-specific voice checks
packages/​studyloop/​src/​studyloop/​doctor/​exporter.py Per-harness freshness diagnostics
packages/​studyloop/​src/​studyloop/​doctor/​config.py Obsidian and xTiles diagnostics
packages/​studyloop/​src/​studyloop/​doctor/​agents.py Shared agent definition checks
docs/​web-ui-guide.md Revised Body Double workflow
CHANGELOG.md User-facing change documentation

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

store = kiro_sessions_dir or KIRO_SESSIONS_DIR
kiro_newest = by_source.get("kiro_cli")
try:
store_changed = datetime.fromtimestamp(store.stat().st_mtime, tz=UTC)
Comment on lines +3325 to +3326
const topic = (this.liveActivity || this.noteTopic || '').trim();
return topic ? [topic] : [];
Comment on lines +5322 to +5323
body:has(.body-double-view.bd-live) .pomodoro {
display: none;
@NetDevAutomate

Copy link
Copy Markdown
Owner Author

Full suite on this branch's tree (head 8fb7c807, which contains #50): 7397 passed, 31 failed, 14 errors, 16 skipped in 16 min. Of the 45 failing ids, 44 are in the recorded machine-specific set (docs/architecture/plan-integration/receipts/full-suite-control-item4-2026-09-18.md); the other is packages/agent-session-tools/tests/test_sync_conversation_integrity.py::test_concatenated_remote_dump_with_existing_archive, the known sqlite3 subprocess timeout in a package this branch does not touch. No failing id is new.

@NetDevAutomate
NetDevAutomate merged commit 8fb7c80 into main Sep 28, 2026
17 checks passed
@NetDevAutomate
NetDevAutomate deleted the fix/body-double-calm branch September 28, 2026 14:59
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants