Repository navigation
Conversation
When a Composio key is configured the Claude driver mounts the Connect MCP server and pre-allows its tools, but the persona never mentions it — so the model does not link "can you read my email?" to the composio tools it is holding, and answers that it has no access. Add the line, gated the way the computer and agents hints already are: a composioMcp capability the driver declares, which gates the integration itself. A key in the config says the user has those connections, not that this engine can reach them — Codex, Grok and the ACP drivers ignore integrations.composio, and must not be told otherwise. Tests: Claude mounts and claims it; ACP claims neither.
…ials file
The driver decided auth state by testing whether ~/.claude/.credentials.json
exists. On macOS, Claude Code keeps its OAuth tokens in the login Keychain
(generic-password item `Claude Code-credentials`), so that file never exists
and every signed-in Mac user was reported as signed out.
That also disabled the model picker, which is gated on the same snapshot flag
in the renderer — affected users were silently locked to MODELS.default with
no indication the two were related.
Ask the CLI instead: `claude auth status` prints {"loggedIn":true,...} on
stdout. That is storage-agnostic, so it keeps working if the credential store
changes again, and it also covers API-key and Bedrock/Vertex users the file
check never saw. The existing file check stays as a fallback for CLIs
predating the subcommand.
Fixes milind-soni#108
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The sidebar is declared w-[320px] shrink-0, so on a 360px screen it leaves 40px for the conversation. Below md it now sits out of flow and slides in over the chat behind a scrim, reached from a menu button in the top-left. Every mobile class is scoped with max-md: rather than applied at the base and cancelled with md:. That is deliberate. Tailwind v4 emits the native `translate` property, and any value other than `none` makes the element a containing block for its `fixed` descendants — so md:translate-x-0 would not cancel anything, it would silently reparent NewRoomPanel's overlay and the "+" menu backdrop on desktop. With max-md:, none of these properties are emitted above the breakpoint at all: the desktop layout is unchanged by absence, not by cancellation. View headers reserve room for the overlaid button the same way they already do for the Windows window controls. Verified: desktop rendering at 1280x800 is pixel-identical to main.
Four ways out, because a drawer you cannot dismiss is worse than no drawer: tap the scrim, press Escape, press the button again, or pick a conversation. The last one watches state rather than patching call sites. Picking a bot is not the only thing that puts content over the chat: the reducer flips activeView without touching selectedId when you re-pick the bot you are already on, and Plugins and bot settings open panels that sit below the drawer's stacking level. Watching all four values catches every one of them in a single place, and stays correct when new triggers appear. Escape restores focus to the menu button, matching ApiKeys.tsx. The scrim and selection paths deliberately do not — a pointer gesture should not move focus.
The drawer still moves, it just arrives instantly. Added to the existing reduced-motion block rather than a new one.
Adds `droidAgent`, backed by `droid exec -o acp`, plus the registration and default-instance entry. Droid is the one ACP harness here that does not take session settings from argv. `droid exec --help` states it under "Stream JSON-RPC Mode": CLI flags do not configure JSON-RPC sessions, so a `-m` passed through spawnArgs is validated and then ignored, and the session silently runs whatever ~/.factory/settings.json selected. Two small hooks on AcpSupport cover it: configureSession() applies model and autonomy over the wire between session/new (or session/load) and the first prompt. Both are always explicit, so a mode or model pinned in settings.json can never decide how a turn runs. A rejected setting fails the turn with a message naming the engine, the setting and the method. resolveModels() reads the user-local catalog, since droid's real model list is half per-machine: `custom:` providers, modelFavorites ordering, and sessionDefaultSettings.model all live in settings.json. Unreadable settings fall back to the static built-in slice. Sign-in detection checks all three credential filenames droid can write (auth.v2.file, auth.v2.loginkeychain, auth.v2.keyring), since which one exists depends on the account's secure_auth_storage flag, and honours FACTORY_HOME_OVERRIDE, which replaces the CLI's HOME rather than its data root. windowsKnownDirs() gains ~/bin, where the Windows installer puts droid.exe, so the app finds it without a restart.
opencode's ACP subcommand takes no -m, so the model has to be set with session/set_config_option before prompting. The hook is opt-in: harnesses that pass -m on the command line are untouched. The requested model must be confirmed by the agent, or the turn aborts. An agent that acknowledges the call but keeps its old model would burn a paid turn on something other than what the picker shows.
Neither the -32602 abort nor the session/load branch of the model hook had a test, so a refactor of that block had nothing to catch a regression.
opencode 1.18.18 puts {inputTokens, outputTokens} at the root of the
session/prompt result. Reading only _meta silently dropped the count.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01A52XvU63rgJypGD8U9cmNP
…thenticated A harness whose child needs a policy composed from its own config could not reach that config, and one whose readiness depends on asking the CLI could not answer synchronously. Both are additive: existing supports ignore the new parameter and keep returning a boolean.
The test passed `resumeCursor: "fake-acp-session"`, which is the same id the
fake returns from session/new. core.ts sets `sessionId = cursor` only on a
successful load, so if session/load threw and the code fell back to
session/new, the emitted sessionId would have been byte-identical and the
assertion would still have passed. The test claimed to lock the resume path and
locked nothing.
Proved rather than argued: making the fake's session/load return a JSON-RPC
error left the old test GREEN. With a distinct cursor the same break turns it
RED —
- "sessionId": "resumed-thread-1"
+ "sessionId": "fake-acp-session"
— which is exactly the silent fallback it is supposed to catch. The temporary
break was reverted; fake-acp-cli.ts is untouched by this commit.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01A52XvU63rgJypGD8U9cmNP
The existing unadvertised-model test rides the fake CLI's -32602, so it
settles inside `request()` and never reaches the confirmation guard in
core.ts. The guard's own case, an agent that answers OK and quietly keeps
its old model, had no test at all. It is also the case the guard was
written for: an error is loud, this one is silent.
Proved rather than argued. With the guard neutered, the new test does not
merely fail, it reports `ok: true` on a turn that ran `m-one` while
`m-two` was asked for:
- "ok": false
+ "ok": true
That is exactly the failure the guard prevents, a paid turn spent on the
wrong model with nothing to show for it.
core.ts is untouched. This is coverage for behaviour that already shipped
earlier in this branch.
Raised by CodeRabbit on milind-soni#122.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01A52XvU63rgJypGD8U9cmNP
…chain `claudeSignedIn` looked only for ~/.claude/.credentials.json. That is where Claude Code keeps credentials on Linux and WSL; on macOS it uses the Keychain and writes no such file. So every signed-in Mac reported as signed out, and the engine sat greyed out in the model picker as "sign-in required" no matter how many times you signed in. Fall back to the Keychain when the file is absent on darwin. The probe is `security find-generic-password` for the `Claude Code-credentials` service *without* -w: that asks whether the item exists and never reads the secret, so it neither needs nor triggers an authorisation prompt. The platform and the probe are injected so the decision is testable off a Mac, and without depending on whoever runs the suite being signed into Claude.
Three changes that any client benefits from, the desktop app included. **The SSE stream is not resumable.** Broadcast frames carry no sequence number and /api/events honours no cursor, so a client that loses its connection for two seconds has no way to ask for what it missed — its only recovery is to refetch the world. Every frame now carries a monotonic `seq` stamped at broadcast, emitted as the SSE `id:` field, and `?since=` (or `Last-Event-ID`) replays from there. A reconnect after a blip is now a replay of a handful of frames instead of a full re-hydration. **Hydration is all-or-nothing.** GET /api/bots returns every bot's entire transcript. That is the right shape over loopback and a poor one anywhere else — a long-running fleet ships megabytes on every cold start. `?messages=n` returns the newest n per thread with a `hasMore` flag, and GET /api/threads/:id/messages?before= walks backwards from there. Screen captures are reduced to a flag in that shape and fetched individually from GET /api/threads/:id/messages/:id/image, so a base64 desktop capture is never inlined into a hydration payload. **BotRecord.notifications is a dead switch.** The per-bot toggle already ships in the UI and nothing reads it. `buildNotification` does, and returns null when it is off; the resulting `notify` frame is what a client shows. The summary strips code fences, because a notification whose body is a diff is not a notification. server/testing/sse.ts is a small helper for driving a real event stream in tests — the frames are asserted on the wire, not on an internal bus.
A review bot flagged the test as writing and then deleting the real ~/.claude/.credentials.json, signing the developer out. It does not: server/testing/setup.ts repoints HOME and USERPROFILE at a fresh temp dir before any test module loads, and os.homedir() reads those at call time, so the file written and removed was always $TMPDIR/omb-test-home-*/.claude/. Still worth changing. The safety was action at a distance — nothing in the test file says the path it writes to is fake, and a reader checking whether this suite is safe to run has to know that setup.ts exists and what it does. Injecting the third probe alongside the platform and the keychain makes the test hermetic by construction: it now touches no filesystem at all, so the question cannot come up again, from a reviewer or from a future runner configured without that setup file.
…adge **The image route materialised threads it was asked about.** `messagesFor` creates and caches a ThreadState for whatever id it is handed, and the sibling page route guards for exactly that three lines above — the new image route did not. A client asking for images on ids that were never real grew the thread map for as long as it kept asking. Same guard, same 404. **Every cold start hydrated twice.** The eager `loadAll()` runs, then the EventSource connects with no `Last-Event-ID`, so the server answers `hello` with `resumed: false`, and that ran `loadAll()` again — eight API calls and two full transcript downloads per page load. This predates the change (it was `onopen` doing the second load before), but leaving it in a commit whose whole purpose is to stop re-downloading transcripts would be absurd. The first hello is now never a reason to re-hydrate; only a reconnect the server could not replay is. Kept the eager load rather than deferring everything to `hello`, which was the other way to fix it: a page that cannot open an EventSource at all should still show what the API can tell it, rather than nothing. **Opening a bot from its notification left the badge on.** The `notify` handler used `rawDispatch`, which clears `unread` in local state but does not PATCH it back — so the badge returned on the next hydration. The wrapped `dispatch` is the one that tells the server. A notification target is unread by definition, so this was every time, not an edge case.
feat(acp): opt-in model selection through a session config option
…acos-signin Detect a signed-in Claude on macOS, where credentials live in the Key…
…i#192) * Harden the turn engine, gate Composio per bot, and give every bot a workspace with file memory - TurnWatchdog: activity-based stall detection for dispatched turns (1:1 and room) — interrupts and settles a turn whose thread has been silent for OMB_TURN_STALL_MS (default 20m); turns parked on a human approval are exempt - Rooms now respect and set bot.busy: a bot can no longer run a 1:1 turn and a room turn concurrently, and interrupting a bot reaches its room turn - Per-bot Composio gate (bot.composio): the workspace key no longer reaches every bot unconditionally; imported team members start with it off - Teams import: no seeded greeting; persona fields bounded at PATCH /api/bots (100/200/4000) matching manifest caps; chief-of-staff roster clips name/role/about and caps the roster at 40 bots - Per-bot workspaces (~/.openmausbot/workspaces/<botId>) as turn cwd for CLI engines, with plain-file memory: MEMORY.md injected into the system prompt under a 200-line/24KB budget, memory/ topic files read on demand - Token usage: thread.token-usage.updated events now fold into a per-task usage tally (input/output/turns) exposed over the API - timingSafeEqual for the internal comms bearer token Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Harden room ownership and expose connected-app grants * SQLite message store + transcript search/export + per-bot Composio toggle UI - server/message-db.ts (node:sqlite, built into Node >=23.4 — no new dependency, nothing to bundle): messages persist as per-mutation deltas (one INSERT per append, one UPDATE per patch) instead of rewriting the whole messages-<threadId>.json on every message; WAL mode; legacy JSON thread files import lazily on first read and are renamed .imported as a one-time backup - GET /api/search?q= — case-insensitive substring search over text messages across every transcript (LIKE scan; local scale needs no FTS index), hits resolved to their bot/task or room - GET /api/threads/:id/export?format=markdown|json — the visible branch as a downloadable transcript, screen-frame pixels stripped - Sidebar search now surfaces transcript hits under 'In conversations'; clicking one opens the conversation (and switches to the task that holds it) - Bot settings gains a 'Connected apps' toggle wired to the per-bot composio gate (shown when a Composio key is configured) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Harden transcript migration and persistence * Release SQLite handles between tests * Close reloaded SQLite handles on Windows --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The sidecar was reviewed three times in parallel — on its own PR, under the toggle PR, and under the iOS PR — and each line fixed what its review found. This folds all three into one, keeping the strongest version wherever two lines fixed the same thing differently: From the iOS line: CRLF/bare-CR-tolerant SSE framing; the mDNS on-link check derived from interface netmasks rather than an RFC 1918 prefix guess, plus byte-budget clamping for TXT names; device-record normalization that rejects zero and negative timestamps; named state-file modes; the byte-pipe branch destroying the response when the harness dies mid-image; a 16 MiB ceiling on `tailscale status` output. From the toggle line: fail-closed scrubbing — a response that parses but cannot be scrubbed is a 502, never forwarded raw; pairing that rolls back and reports when the write fails, and a lastSeenAt write failure that no longer signs a phone out; `originIsLoopback` on the control plane; a runnable bin (shebang, exec bit, restored bin entry); the self-scheduling control-page poll. Kept from this line where others regressed it: `+json` structured-suffix scrubbing; the fail-closed SSE ceiling (the toggle line's cap silently dropped a frame); the goodbye-datagram flush; the headers-phase deadline with its 504/502 distinction; the pairing code checked before the device cap, so a wrong guess cannot probe fleet state. One bearer parser everywhere, case-insensitive per RFC 7235.
The coverage each review round produced, folded into one suite: the CRLF framing trio and the on-link and byte-clamp suites from the iOS line; the control-plane suite, the fail-closed response suite, and the failing-disk device cases from the toggle line; this line's verified-free-port harness kept as the skeleton throughout. Two assertions changed meaning on purpose, both because the reconciled control plane keeps the strictest of the three origin policies — only the exact addressed authority passes. A cross-origin GET is now refused (a safe-method list is a list that goes stale the day a read starts leaking), and a loopback origin on any other port is refused with it. The toggle line's RemoteListener suite is deliberately absent: the class it tests was deleted with its last caller. 9 files, 113 tests.
The toggle layer lands on top of the unioned companion/ from the sidecar branch, which already folds in every fix this line made to the sidecar — its own copies resolve wholesale to the union. Kept from this line: the electron toggle itself, the Companion settings section, and the newer test teardown primitives (waitForExit and removeTempDir), which all four spawning suites now use; the sidecar branch's stopAndClean and teardown.ts retire in their favour. Kept from the sidecar branch: the explicit OMB_COMPANION_DIR redirect in test setup, and the Windows dying-stdin guard in procs.ts. The RemoteListener suite goes with the class it tested — deleted with its last caller. The README regains the Settings → Companion paragraph, which belongs at this layer where the toggle exists.
The iOS layer now sits on the reconciled sidecar and toggle layers, so every companion/ and toggle-layer file resolves wholesale to the layer that owns it — the fixes this line carried for those layers are all in the union below, and its own copies retire. What this layer keeps is what is genuinely its own: ios/, the fixture capture script, the testing runbook, and the companion docs.
…d-soni#198) * Bundle the server's dependencies so the packaged app can start 0.1.24 built, signed, notarized and installed cleanly, then died on every launch: ERR_MODULE_NOT_FOUND: Cannot find package 'zod' imported from Resources/server/config.js The packaged app ships no node_modules by design (electron-builder.yml line 21 calls the three pieces self-contained, line 33 excludes them). build:server was plain tsc, which transpiles without bundling, so the `zod` import milind-soni#194 introduced survived verbatim into a tree with nothing to resolve it against. zod was the first bare import the server ever had, so the invariant had never been tested. Bundle every entry point with esbuild after tsc, mirroring scripts/bundle-updater.mjs which already vendors electron-updater for the same reason. All six are bundled, not just index.ts: the proxies run as their own processes and import nothing external today, but the next one that does would fail the same silent way. Entry points keep their relative paths, which the proxy lookups depend on. Both gates that should have caught this were blind to it. The unit suite runs inside the repo, where a bare import resolves from ./node_modules; the Windows packaging check asserts index.js exists but never runs it. So the new smoke test copies dist-server OUT of the repo before starting it, and CI now starts the real packaged copy on the runner. Verified by mutation: reverting to plain tsc output fails the smoke test with the original ERR_MODULE_NOT_FOUND. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * Build the server before smoke-testing it The smoke test assumed dist-server was already on disk. It is gitignored (milind-soni#190), so a fresh CI checkout has never built it and the test died on ENOENT before it could prove anything. It passed locally only because a build happened to be sitting there. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * Never let scratch cleanup fail the smoke test Windows holds file handles briefly after the process that owned them dies, so removing the scratch dir right after the kill raised EPERM and failed a run whose server had actually started fine. Linux raises EACCES the same way (f66d30f). Cleanup is housekeeping; the assertion is the test. Mutation re-checked: plain tsc output still fails. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
testDecodesThePagedFleet unwraps a group with three messages and a page boundary, and the capture script never created one — the committed fixture's room was an accident of whichever harness the fixtures were last captured against, and the first regeneration on a clean machine failed the test. The script now creates the room itself: directly on the harness, because room creation is deliberately not on the sidecar's allowlist and so is setup the phone cannot perform — then five messages through the sidecar, captured at messages=3, which is what makes the pinned count and hasMore=true properties of the capture rather than of history.
* Working folder: pin the default for cloud runs, clear via an empty field Follow-ups from CodeRabbit's review of milind-soni#183 (merged): a cloud run now pins task.cwd = null so the header chip never shows the bot's host folder for a task that runs on the box; the non-desktop text field sends null when emptied, which is what the server takes as "clear" — it sent "" and was rejected. (The third finding, the ~ path boundary, was already fixed on main in src/lib/short-path.ts.) Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * Make cloud task folder pin authoritative --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> Co-authored-by: milind-soni <milindsoni201@gmail.com>
# Conflicts: # package.json
* Let the store announce its own writes, and give bots a real activity state 2.6 — emit inside the store. Every write to the store used to need a matching broadcast at the call site, and there were 31 of them across four files. Miss one and the UI drifts from disk; broadcast without writing and a restart loses what the user just watched. The Store now emits a typed StoreChange after each write (message, message.patch, thread, bot, bot.deleted, group, group.deleted) and index.ts maps those onto SSE frames in ONE subscriber. The mirror broadcasts are gone. The few endpoints whose callers need a full transcript on the wire (task create/switch, team import) still send their richer payload on top. 2.7 — activity state. busy could not tell working from waiting-on-you from a stalled engine. BotRecord.activity is one of working | waiting-on-you | idle | no-signal | dead; busy is DERIVED from it in store.setActivity, the only place runtime state changes now, so all existing busy readers keep working unchanged. Transitions: dispatch → working; a card that actually reaches a human → waiting-on-you; answered → working; settled → idle; a setup error (the engine could not start) → dead until the next dispatch. no-signal is the slot the liveness clock (milind-soni#193) fills. The sidebar preview says "Waiting for you…" when that is what is happening. Items 2.6 and 2.7 of docs/plans/agent-harness-upgrades-v2.md. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * store emits: drop the last manual group broadcasts; persist the transient reset Review follow-ups (CodeRabbit on milind-soni#199): - every patchGroup was followed by a manual broadcastGroup, so each room update sent two identical frames now that the store emits. The manual calls are gone and broadcastGroup with them — the store's group change is the one source. - the load-time reset of busy/activity was in-memory only: a bots.json left by a process that died mid-turn kept saying busy on every load. The reset now marks the file for a save when it changed anything. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * Complete store-owned update delivery --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> Co-authored-by: milind-soni <milindsoni201@gmail.com>
A turn ends across three server frames: the settled reply message, turn.completed, then the bot patch that flips busy off. The tail derived its "working" dots from busy && !streaming alone, so in that window the dots popped back in under the reply that just landed, then vanished a beat later — and the pinned scroll re-anchored around each height change. That grow-shrink-jump was the end-of-stream jitter. showWorkingDots (src/lib/turn-tail.ts) now decides the tail: a settled bot text at the end of the transcript means there is nothing to wait for, so the dots stay hidden until something actually new starts — a tool chip, the next user prompt, or (in rooms) a different speaker taking the floor. Rooms gate per speaker: a reply from a previous speaker doesn't cover the bot now on the floor, and a message with no attribution fails toward showing the indicator. Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2.3 — durable delegations. The per-thread handoff queue lived in a Map and died with the process: a delegation queued right before a restart never ran, silently. It is now written to ~/.openmausbot/delegations.json on queue, drain, and discard, loaded at boot, and drained through the same path a settled turn uses — target and approvePeerComms are still re-checked at drain time; a source bot that no longer exists is skipped. Provider permissions still die with the process (nobody can answer for an unattended bot); queued work is not a permission. 2.4 — secret redaction, rescoped. The harness never stores tool RESULTS — CLI drivers run tools inside the CLI and the harness only sees names and outcomes — so there was no tool-result body to redact. What it does store, and now replays into every rebuild, is the bot's reply text, activity chip titles (an ACP engine's title can be the whole command line), and permission card summaries (the command being approved). Those can carry keys. - redactSecretsInText: high-precision content patterns only — known key prefixes (sk-, ghp_/github_pat_, xox*-, AKIA, AIza, npm_), JWTs, PEM private-key blocks, Bearer tokens, and secret-shaped key=value. No generic hex/base64 heuristics: those rewrite real code. - Applied at the single store write for role "bot" (text, tool.name, card title/subtitle/summary). What the user typed stays as typed — pasting your own .env for the bot to use is your call. - redactSecrets (the native tee) gains the same content pass over string values, and the events NDJSON now goes through it too. - Live streaming deltas are not redacted (tokens split across events); the settled message that replaces the bubble is, and that is what is stored and replayed. Items 2.3 and 2.4 of docs/plans/agent-harness-upgrades-v2.md. Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> Co-authored-by: milind-soni <milindsoni201@gmail.com>
* Answer approvals with a typed outcome, and notice a bot repeating itself
2.5 — fail-closed approvals. respondToRequest used to resolve void or
THROW when the ask was gone (turn ended, broker died, engine has no
asks): the user's Allow click became a 500 and the card sat open
forever. It now resolves a typed RequestOutcome — allowed-once |
rejected | answered | unavailable — across all six drivers, and the
harness branches on it: `unavailable` settles the card as dismissed and
leaves a chip ("Couldn't deliver that answer — the request is no longer
open, so the action was not run"), returning 200 {outcome}. The auto-
approve path hands the ask back to a human on `unavailable` instead of
on a throw. request.resolved is typed too: behavior allow|deny|answer,
source user|auto|timeout|system|unavailable|peer — codex used to stamp
every resolution "user", timeouts included; ACP and codex now say
timeout/system where that is what happened. Kept as-is on purpose: the
timeout NOTE text (post-decision guidance to the model, the decision
itself was already typed) and request.opened's own tool/summary.
2.8 — repeat-call detection (observe only). server/repeat-detector.ts
counts identical calls per turn keyed on tool + arguments — from every
permission ask's summary and from ACP item titles; a bare tool name is
never counted (five "Bash" may be five commands, and Claude's
item.started carries only the name). At 5, 10 and 20 a chip says "Same
call repeated N× — Bash: git status — it may be stuck". No auto-stop:
the human has Stop; cutting a call off needs the harness to own the
call (3.2/3.3).
Items 2.5 and 2.8 of docs/plans/agent-harness-upgrades-v2.md.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Harden approval outcomes and repeat tracking
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: milind-soni <milindsoni201@gmail.com>
The server has tallied per-task usage since tasks landed, and wireTask
already carries it — the client just never typed or rendered it. Task
rows now show a quiet combined total ("12.3k") next to the timestamp,
with the input/output split on hover; the picker button keeps its shape
and carries the open task's tally in its hover title instead.
Formatting lives in src/lib/format-tokens: zero hides the chip, sub-1k
spells out "tokens", then k/M with one decimal. Rounding runs on integer
tenths so 999,950 promotes to "1M" instead of float-rounding down.
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
…#203) While a reply streams, a code block now highlights after its content has been unchanged for 250ms instead of waiting for the stream to settle. The result is written to the content-hashed highlight cache, so the settled bubble (a fresh component instance) mounts highlighted from cache instead of popping from plain <pre> to Shiki output a beat later. The debounce rides the effect lifecycle: every content change re-runs the effect and its cleanup clears the pending timer, so a growing block never tokenizes per token. Partial fences may cache under their own hash — harmless, since the final content gets a distinct key and the cap evicts strays. Non-streaming rendering is unchanged. Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Reviewed and integrated from milind-soni#161, updated for the SQLite-backed server, hardened at the paired-device boundary, and verified on desktop CI plus an iOS simulator build.
⌘K / Ctrl+K from anywhere opens a centered switcher: bots and rooms answer instantly from local state, transcript hits ride the existing /api/search endpoint a 150ms debounce later (same stale-response guard as the sidebar search). Empty query is pure switcher mode — every bot and room, no message section. Ranking lives in src/lib/palette-rank.ts as a pure helper: prefix matches outrank substring matches, matching is case-insensitive, and each tier preserves the caller's order so pinned-first lists survive. Arrow keys move one flat cursor across the sections, Enter selects (message hits open the owning bot or room, switching to the hit's task like the sidebar does), Esc closes. Row hover tracks mousemove rather than mouseenter so hits arriving under a resting pointer can't steal the keyboard selection. Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Long computer-use threads mount hundreds of rows — inline base64 screenshots included — so the DOM stays heavy even though the memoized list bails out of re-renders. ChatView and GroupView now mount only the last 120 messages; a quiet "Show earlier messages (X more)" pill at the top of the transcript pulls the boundary back by another 120 per click. The boundary is a pure decision in src/lib/transcript-window.ts: it is anchored per bot+task (render-phase reset on bot.id/threadId switch), so appends grow the window instead of sliding rows out from under the reader, and a thread that shrinks beneath a stale boundary — branch switch, edit rewind — falls back to a fresh tail window rather than blanking the transcript. Expanding must not yank the viewport: the click captures scrollHeight, then a layout effect shifts scrollTop by the growth before paint (browser scroll anchoring is disabled on these containers). The bottom-follow scrollTo keys on the FULL list's length, so expansion never re-triggers it, and expanding breaks follow the way any scrollback reading does. Tail-derived logic — working dots, lastBotTextId, regenerate — still computes from the full list. Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
A bot's MEMORY.md rides into every turn's system prompt and its memory/*.md topic files steer what it believes, yet none of it was visible anywhere in the UI — the user had to know the workspace path and open the files by hand to see (or fix) what a bot had decided to remember. Three routes expose the files the workspace already owns: GET /api/bots/:id/memory returns the WHOLE MEMORY.md (not the load budget's cut — an editor must show everything, with `truncated` flagging what the prompt loader would drop) plus a name+size listing of memory/ topics; PUT writes it back, string-checked at the boundary and capped at 256KB so a runaway paste fails with an explanation; GET /memory/topics/:name serves one topic read-only. Topic names pass one strict gate (single plain-markdown path segment, decoded before it is judged) in both the route and the workspace helper, so no coat of percent-encoding turns the route into an arbitrary-file read — the tests plant real files at the traversal targets and prove every encoding of ../ dies as a 400 before the filesystem is touched. The Settings panel gets a Memory card in the existing card language: collapsed by default (most visits never look, and expanding re-reads so mid-session notes appear), a monospace editor with save + budget warning, and the topic list opening into a read-only viewer. Keyed by bot id so switching bots can never show one bot's notes under another's name. Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
The phone's chat field was type-only. The desktop composer already has a mic — press to talk, press to stop, partials land in the box so you can edit before sending — and the architecture note for this app was that SFSpeechRecognizer is on every iPhone. Composer dictation is the smaller half of that, and it is the half you actually use while walking around. Same engine as electron/resources/speech-helper.swift: an AVAudioEngine tap into SFSpeechAudioBufferRecognitionRequest, on-device when the recognizer supports it, locales from the user's preferred languages rather than a hardcoded en-US. Composer mode, not call mode — no silence endpointing. Tap the mic to stop. The last partial is what you send. The join (typed text + live transcript) lives in CompanionCore so it can be tested without a phone. Partials replace each other after the text that was already in the field; they never stack. The mic stays next to send. Hiding it once text arrives is the desktop pattern, where Escape stops listening and the toolbar only has room for one action. A phone has neither — this is how you stop, and how you add another sentence by voice after the first one. Backgrounding or an audio interruption stops the session. NSMicrophoneUsageDescription and NSSpeechRecognitionUsageDescription are in project.yml. Without them the first tap crashes rather than prompting. Co-authored-by: Computomatix <mnthr7@users.noreply.github.com>
Cancel an in-flight start instead of racing a second tap during the permission prompt. Fail closed when the locale has no recognizer. Tear the audio tap down before endAudio so a late buffer cannot fail the recognition task. Treat opening the computer panel as leaving chat (NavigationStack keeps ChatView mounted). Co-authored-by: Computomatix <mnthr7@users.noreply.github.com>
Cancel an in-flight start when send is tapped during the permission prompt. Drop the unused SwiftUI import. Label the README tree fence (MD040). Document the speech Info.plist keys and that a denial is shown on that same attempt. Co-authored-by: Computomatix <mnthr7@users.noreply.github.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What changed
Voice input in the iOS companion chat composer: tap the mic, talk, tap to stop, then send or edit. Partials stream into the field as you speak.
This is composer dictation, not call mode. Same engine as the desktop helper (
SFSpeechRecognizeron anAVAudioEnginetap), on-device when the phone supports it, using the user's preferred language rather than hardcoded English.The join between already-typed text and the live transcript lives in
CompanionCore(Dictation.draft) so it can be tested without a phone. Partials replace each other after a frozen base; they never stack.The mic stays next to send so you can stop without an Escape key and add another sentence by voice. Backgrounding, an audio interruption, leaving the chat, or opening the computer panel ends the session. Send during the permission prompt cancels the in-flight start.
NSMicrophoneUsageDescriptionandNSSpeechRecognitionUsageDescriptionare inios/project.yml. Without them the first tap crashes rather than prompting.Follows milind-soni#161 / milind-soni#204. No harness or sidecar changes.
Why
The desktop composer already has a mic. The iOS README called that out as missing on purpose ("no affordance without a feature behind it"). The phone is the better of the two devices for this, and walking around with a bot is the reason the companion exists.
Call mode and spoken replies are still later. This is the half you use in the chat window.
How it was verified
ios/Tests/CompanionCoreTests/DictationTests.swift.milind-soni/OpenMausBotmain(post-Add ios/: the SwiftUI companion app milind-soni/OpenMausBot#161:visibleTranscript, unread-while-open, last-id auto-scroll) rather than opening the OpenMausMobile companion-stack branch.swift testor the simulator; those still need a Mac (cd ios && swift test, thenxcodegen generatebecauseSpeechDictation.swiftis new inApp/).server//src//electron/changes, sopnpm typecheck/pnpm testare unchanged by this diff.ios/TESTING.mdstage 4 step 6 (not run here).Reviewed against CodeRabbit comments from milind-soni#161 before opening: MD040 on README fences, teardown order (tap off before
endAudio), ChatView last-id auto-scroll left in place, computer-panel push does not disappear ChatView so dictation stops on the destination'sonAppear.Screenshots (UI changes)
Composer now has a mic to the left of send. While listening the icon fills, pulses, and turns red; the placeholder reads "Listening…". No device screenshot from this environment — add one from a simulator or phone before or after opening.
Checklist
pnpm typecheckandpnpm testpass locallyserver/is unchangeddist-server/edits (it's build output)shell: true/ cmd.exe string-building — this is iOS App target + Foundation-only CompanionCore