Skip to content

ios: dictate into the chat composer - #2

Closed
mnthr7 wants to merge 175 commits into
mainfrom
cursor/ios-composer-dictation-2f83
Closed

mnthr7 wants to merge 175 commits into
mainfrom
cursor/ios-composer-dictation-2f83

Conversation

@mnthr7

@mnthr7 mnthr7 commented Aug 18, 2026

Copy link
Copy Markdown
Owner

What changed

Voice input in the iOS companion chat composer: tap the mic, talk, tap to stop, then send or edit. Partials stream into the field as you speak.

This is composer dictation, not call mode. Same engine as the desktop helper (SFSpeechRecognizer on an AVAudioEngine tap), on-device when the phone supports it, using the user's preferred language rather than hardcoded English.

The join between already-typed text and the live transcript lives in CompanionCore (Dictation.draft) so it can be tested without a phone. Partials replace each other after a frozen base; they never stack.

The mic stays next to send so you can stop without an Escape key and add another sentence by voice. Backgrounding, an audio interruption, leaving the chat, or opening the computer panel ends the session. Send during the permission prompt cancels the in-flight start.

NSMicrophoneUsageDescription and NSSpeechRecognitionUsageDescription are in ios/project.yml. Without them the first tap crashes rather than prompting.

Follows milind-soni#161 / milind-soni#204. No harness or sidecar changes.

Why

The desktop composer already has a mic. The iOS README called that out as missing on purpose ("no affordance without a feature behind it"). The phone is the better of the two devices for this, and walking around with a bot is the reason the companion exists.

Call mode and spoken replies are still later. This is the half you use in the chat window.

How it was verified

  • Dictation join and locale-candidate tests in ios/Tests/CompanionCoreTests/DictationTests.swift.
  • Replayed onto current milind-soni/OpenMausBot main (post-Add ios/: the SwiftUI companion app milind-soni/OpenMausBot#161: visibleTranscript, unread-while-open, last-id auto-scroll) rather than opening the OpenMausMobile companion-stack branch.
  • This environment has no Swift toolchain and cannot run swift test or the simulator; those still need a Mac (cd ios && swift test, then xcodegen generate because SpeechDictation.swift is new in App/).
  • No server/ / src/ / electron/ changes, so pnpm typecheck / pnpm test are unchanged by this diff.
  • End-to-end on a phone is ios/TESTING.md stage 4 step 6 (not run here).

Reviewed against CodeRabbit comments from milind-soni#161 before opening: MD040 on README fences, teardown order (tap off before endAudio), ChatView last-id auto-scroll left in place, computer-panel push does not disappear ChatView so dictation stops on the destination's onAppear.

Screenshots (UI changes)

Composer now has a mic to the left of send. While listening the icon fills, pulses, and turns red; the placeholder reads "Listening…". No device screenshot from this environment — add one from a simulator or phone before or after opening.

Checklist

  • pnpm typecheck and pnpm test pass locally
  • Server behavior changes come with tests (see CONTRIBUTING.md → Tests) — server/ is unchanged
  • No dist-server/ edits (it's build output)
  • macOS-only code is platform-gated; no shell: true / cmd.exe string-building — this is iOS App target + Foundation-only CompanionCore
  • No secrets in logs, responses, events, or argv
Open in Web Open in Cursor 

milind-soni and others added 30 commits August 14, 2026 16:33
When a Composio key is configured the Claude driver mounts the Connect
MCP server and pre-allows its tools, but the persona never mentions it —
so the model does not link "can you read my email?" to the composio tools
it is holding, and answers that it has no access.

Add the line, gated the way the computer and agents hints already are: a
composioMcp capability the driver declares, which gates the integration
itself. A key in the config says the user has those connections, not that
this engine can reach them — Codex, Grok and the ACP drivers ignore
integrations.composio, and must not be told otherwise.

Tests: Claude mounts and claims it; ACP claims neither.
…ials file

The driver decided auth state by testing whether ~/.claude/.credentials.json
exists. On macOS, Claude Code keeps its OAuth tokens in the login Keychain
(generic-password item `Claude Code-credentials`), so that file never exists
and every signed-in Mac user was reported as signed out.

That also disabled the model picker, which is gated on the same snapshot flag
in the renderer — affected users were silently locked to MODELS.default with
no indication the two were related.

Ask the CLI instead: `claude auth status` prints {"loggedIn":true,...} on
stdout. That is storage-agnostic, so it keeps working if the credential store
changes again, and it also covers API-key and Bedrock/Vertex users the file
check never saw. The existing file check stays as a fallback for CLIs
predating the subcommand.

Fixes milind-soni#108

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The sidebar is declared w-[320px] shrink-0, so on a 360px screen it leaves
40px for the conversation. Below md it now sits out of flow and slides in
over the chat behind a scrim, reached from a menu button in the top-left.

Every mobile class is scoped with max-md: rather than applied at the base
and cancelled with md:. That is deliberate. Tailwind v4 emits the native
`translate` property, and any value other than `none` makes the element a
containing block for its `fixed` descendants — so md:translate-x-0 would
not cancel anything, it would silently reparent NewRoomPanel's overlay and
the "+" menu backdrop on desktop. With max-md:, none of these properties
are emitted above the breakpoint at all: the desktop layout is unchanged
by absence, not by cancellation.

View headers reserve room for the overlaid button the same way they
already do for the Windows window controls.

Verified: desktop rendering at 1280x800 is pixel-identical to main.
Four ways out, because a drawer you cannot dismiss is worse than no
drawer: tap the scrim, press Escape, press the button again, or pick a
conversation.

The last one watches state rather than patching call sites. Picking a bot
is not the only thing that puts content over the chat: the reducer flips
activeView without touching selectedId when you re-pick the bot you are
already on, and Plugins and bot settings open panels that sit below the
drawer's stacking level. Watching all four values catches every one of
them in a single place, and stays correct when new triggers appear.

Escape restores focus to the menu button, matching ApiKeys.tsx. The scrim
and selection paths deliberately do not — a pointer gesture should not
move focus.
The drawer still moves, it just arrives instantly. Added to the existing
reduced-motion block rather than a new one.
Adds `droidAgent`, backed by `droid exec -o acp`, plus the registration and
default-instance entry.

Droid is the one ACP harness here that does not take session settings from
argv. `droid exec --help` states it under "Stream JSON-RPC Mode": CLI flags
do not configure JSON-RPC sessions, so a `-m` passed through spawnArgs is
validated and then ignored, and the session silently runs whatever
~/.factory/settings.json selected. Two small hooks on AcpSupport cover it:

  configureSession() applies model and autonomy over the wire between
  session/new (or session/load) and the first prompt. Both are always
  explicit, so a mode or model pinned in settings.json can never decide how
  a turn runs. A rejected setting fails the turn with a message naming the
  engine, the setting and the method.

  resolveModels() reads the user-local catalog, since droid's real model
  list is half per-machine: `custom:` providers, modelFavorites ordering,
  and sessionDefaultSettings.model all live in settings.json. Unreadable
  settings fall back to the static built-in slice.

Sign-in detection checks all three credential filenames droid can write
(auth.v2.file, auth.v2.loginkeychain, auth.v2.keyring), since which one
exists depends on the account's secure_auth_storage flag, and honours
FACTORY_HOME_OVERRIDE, which replaces the CLI's HOME rather than its data
root. windowsKnownDirs() gains ~/bin, where the Windows installer puts
droid.exe, so the app finds it without a restart.
opencode's ACP subcommand takes no -m, so the model has to be set with
session/set_config_option before prompting. The hook is opt-in: harnesses
that pass -m on the command line are untouched.

The requested model must be confirmed by the agent, or the turn aborts.
An agent that acknowledges the call but keeps its old model would burn a
paid turn on something other than what the picker shows.
Neither the -32602 abort nor the session/load branch of the model hook had
a test, so a refactor of that block had nothing to catch a regression.
opencode 1.18.18 puts {inputTokens, outputTokens} at the root of the
session/prompt result. Reading only _meta silently dropped the count.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01A52XvU63rgJypGD8U9cmNP
…thenticated

A harness whose child needs a policy composed from its own config could not
reach that config, and one whose readiness depends on asking the CLI could
not answer synchronously. Both are additive: existing supports ignore the
new parameter and keep returning a boolean.
The test passed `resumeCursor: "fake-acp-session"`, which is the same id the
fake returns from session/new. core.ts sets `sessionId = cursor` only on a
successful load, so if session/load threw and the code fell back to
session/new, the emitted sessionId would have been byte-identical and the
assertion would still have passed. The test claimed to lock the resume path and
locked nothing.

Proved rather than argued: making the fake's session/load return a JSON-RPC
error left the old test GREEN. With a distinct cursor the same break turns it
RED —

    - "sessionId": "resumed-thread-1"
    + "sessionId": "fake-acp-session"

— which is exactly the silent fallback it is supposed to catch. The temporary
break was reverted; fake-acp-cli.ts is untouched by this commit.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01A52XvU63rgJypGD8U9cmNP
The existing unadvertised-model test rides the fake CLI's -32602, so it
settles inside `request()` and never reaches the confirmation guard in
core.ts. The guard's own case, an agent that answers OK and quietly keeps
its old model, had no test at all. It is also the case the guard was
written for: an error is loud, this one is silent.

Proved rather than argued. With the guard neutered, the new test does not
merely fail, it reports `ok: true` on a turn that ran `m-one` while
`m-two` was asked for:

    - "ok": false
    + "ok": true

That is exactly the failure the guard prevents, a paid turn spent on the
wrong model with nothing to show for it.

core.ts is untouched. This is coverage for behaviour that already shipped
earlier in this branch.

Raised by CodeRabbit on milind-soni#122.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01A52XvU63rgJypGD8U9cmNP
…chain

`claudeSignedIn` looked only for ~/.claude/.credentials.json. That is where
Claude Code keeps credentials on Linux and WSL; on macOS it uses the Keychain
and writes no such file. So every signed-in Mac reported as signed out, and
the engine sat greyed out in the model picker as "sign-in required" no matter
how many times you signed in.

Fall back to the Keychain when the file is absent on darwin. The probe is
`security find-generic-password` for the `Claude Code-credentials` service
*without* -w: that asks whether the item exists and never reads the secret,
so it neither needs nor triggers an authorisation prompt.

The platform and the probe are injected so the decision is testable off a
Mac, and without depending on whoever runs the suite being signed into
Claude.
Three changes that any client benefits from, the desktop app included.

**The SSE stream is not resumable.** Broadcast frames carry no sequence
number and /api/events honours no cursor, so a client that loses its
connection for two seconds has no way to ask for what it missed — its only
recovery is to refetch the world. Every frame now carries a monotonic `seq`
stamped at broadcast, emitted as the SSE `id:` field, and `?since=` (or
`Last-Event-ID`) replays from there. A reconnect after a blip is now a
replay of a handful of frames instead of a full re-hydration.

**Hydration is all-or-nothing.** GET /api/bots returns every bot's entire
transcript. That is the right shape over loopback and a poor one anywhere
else — a long-running fleet ships megabytes on every cold start.
`?messages=n` returns the newest n per thread with a `hasMore` flag, and
GET /api/threads/:id/messages?before= walks backwards from there. Screen
captures are reduced to a flag in that shape and fetched individually from
GET /api/threads/:id/messages/:id/image, so a base64 desktop capture is
never inlined into a hydration payload.

**BotRecord.notifications is a dead switch.** The per-bot toggle already
ships in the UI and nothing reads it. `buildNotification` does, and returns
null when it is off; the resulting `notify` frame is what a client shows.
The summary strips code fences, because a notification whose body is a
diff is not a notification.

server/testing/sse.ts is a small helper for driving a real event stream in
tests — the frames are asserted on the wire, not on an internal bus.
A review bot flagged the test as writing and then deleting the real
~/.claude/.credentials.json, signing the developer out. It does not:
server/testing/setup.ts repoints HOME and USERPROFILE at a fresh temp dir
before any test module loads, and os.homedir() reads those at call time, so
the file written and removed was always $TMPDIR/omb-test-home-*/.claude/.

Still worth changing. The safety was action at a distance — nothing in the
test file says the path it writes to is fake, and a reader checking whether
this suite is safe to run has to know that setup.ts exists and what it does.
Injecting the third probe alongside the platform and the keychain makes the
test hermetic by construction: it now touches no filesystem at all, so the
question cannot come up again, from a reviewer or from a future runner
configured without that setup file.
…adge

**The image route materialised threads it was asked about.** `messagesFor`
creates and caches a ThreadState for whatever id it is handed, and the
sibling page route guards for exactly that three lines above — the new image
route did not. A client asking for images on ids that were never real grew
the thread map for as long as it kept asking. Same guard, same 404.

**Every cold start hydrated twice.** The eager `loadAll()` runs, then the
EventSource connects with no `Last-Event-ID`, so the server answers
`hello` with `resumed: false`, and that ran `loadAll()` again — eight API
calls and two full transcript downloads per page load. This predates the
change (it was `onopen` doing the second load before), but leaving it in a
commit whose whole purpose is to stop re-downloading transcripts would be
absurd. The first hello is now never a reason to re-hydrate; only a
reconnect the server could not replay is.

Kept the eager load rather than deferring everything to `hello`, which was
the other way to fix it: a page that cannot open an EventSource at all
should still show what the API can tell it, rather than nothing.

**Opening a bot from its notification left the badge on.** The `notify`
handler used `rawDispatch`, which clears `unread` in local state but does
not PATCH it back — so the badge returned on the next hydration. The
wrapped `dispatch` is the one that tells the server. A notification target
is unread by definition, so this was every time, not an edge case.
feat(acp): opt-in model selection through a session config option
…acos-signin

Detect a signed-in Claude on macOS, where credentials live in the Key…
mnthr7 and others added 29 commits August 17, 2026 18:05
…i#192)

* Harden the turn engine, gate Composio per bot, and give every bot a workspace with file memory

- TurnWatchdog: activity-based stall detection for dispatched turns (1:1 and
  room) — interrupts and settles a turn whose thread has been silent for
  OMB_TURN_STALL_MS (default 20m); turns parked on a human approval are exempt
- Rooms now respect and set bot.busy: a bot can no longer run a 1:1 turn and a
  room turn concurrently, and interrupting a bot reaches its room turn
- Per-bot Composio gate (bot.composio): the workspace key no longer reaches
  every bot unconditionally; imported team members start with it off
- Teams import: no seeded greeting; persona fields bounded at PATCH /api/bots
  (100/200/4000) matching manifest caps; chief-of-staff roster clips
  name/role/about and caps the roster at 40 bots
- Per-bot workspaces (~/.openmausbot/workspaces/<botId>) as turn cwd for CLI
  engines, with plain-file memory: MEMORY.md injected into the system prompt
  under a 200-line/24KB budget, memory/ topic files read on demand
- Token usage: thread.token-usage.updated events now fold into a per-task
  usage tally (input/output/turns) exposed over the API
- timingSafeEqual for the internal comms bearer token

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Harden room ownership and expose connected-app grants

* SQLite message store + transcript search/export + per-bot Composio toggle UI

- server/message-db.ts (node:sqlite, built into Node >=23.4 — no new dependency,
  nothing to bundle): messages persist as per-mutation deltas (one INSERT per
  append, one UPDATE per patch) instead of rewriting the whole
  messages-<threadId>.json on every message; WAL mode; legacy JSON thread files
  import lazily on first read and are renamed .imported as a one-time backup
- GET /api/search?q= — case-insensitive substring search over text messages
  across every transcript (LIKE scan; local scale needs no FTS index), hits
  resolved to their bot/task or room
- GET /api/threads/:id/export?format=markdown|json — the visible branch as a
  downloadable transcript, screen-frame pixels stripped
- Sidebar search now surfaces transcript hits under 'In conversations';
  clicking one opens the conversation (and switches to the task that holds it)
- Bot settings gains a 'Connected apps' toggle wired to the per-bot composio
  gate (shown when a Composio key is configured)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Harden transcript migration and persistence

* Release SQLite handles between tests

* Close reloaded SQLite handles on Windows

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The sidecar was reviewed three times in parallel — on its own PR, under the
toggle PR, and under the iOS PR — and each line fixed what its review found.
This folds all three into one, keeping the strongest version wherever two
lines fixed the same thing differently:

From the iOS line: CRLF/bare-CR-tolerant SSE framing; the mDNS on-link check
derived from interface netmasks rather than an RFC 1918 prefix guess, plus
byte-budget clamping for TXT names; device-record normalization that rejects
zero and negative timestamps; named state-file modes; the byte-pipe branch
destroying the response when the harness dies mid-image; a 16 MiB ceiling on
`tailscale status` output.

From the toggle line: fail-closed scrubbing — a response that parses but
cannot be scrubbed is a 502, never forwarded raw; pairing that rolls back and
reports when the write fails, and a lastSeenAt write failure that no longer
signs a phone out; `originIsLoopback` on the control plane; a runnable bin
(shebang, exec bit, restored bin entry); the self-scheduling control-page
poll.

Kept from this line where others regressed it: `+json` structured-suffix
scrubbing; the fail-closed SSE ceiling (the toggle line's cap silently
dropped a frame); the goodbye-datagram flush; the headers-phase deadline
with its 504/502 distinction; the pairing code checked before the device
cap, so a wrong guess cannot probe fleet state.

One bearer parser everywhere, case-insensitive per RFC 7235.
The coverage each review round produced, folded into one suite: the CRLF
framing trio and the on-link and byte-clamp suites from the iOS line; the
control-plane suite, the fail-closed response suite, and the failing-disk
device cases from the toggle line; this line's verified-free-port harness
kept as the skeleton throughout.

Two assertions changed meaning on purpose, both because the reconciled
control plane keeps the strictest of the three origin policies — only the
exact addressed authority passes. A cross-origin GET is now refused (a
safe-method list is a list that goes stale the day a read starts leaking),
and a loopback origin on any other port is refused with it.

The toggle line's RemoteListener suite is deliberately absent: the class it
tests was deleted with its last caller.

9 files, 113 tests.
The toggle layer lands on top of the unioned companion/ from the sidecar
branch, which already folds in every fix this line made to the sidecar —
its own copies resolve wholesale to the union.

Kept from this line: the electron toggle itself, the Companion settings
section, and the newer test teardown primitives (waitForExit and
removeTempDir), which all four spawning suites now use; the sidecar
branch's stopAndClean and teardown.ts retire in their favour. Kept from
the sidecar branch: the explicit OMB_COMPANION_DIR redirect in test setup,
and the Windows dying-stdin guard in procs.ts.

The RemoteListener suite goes with the class it tested — deleted with its
last caller. The README regains the Settings → Companion paragraph, which
belongs at this layer where the toggle exists.
The iOS layer now sits on the reconciled sidecar and toggle layers, so
every companion/ and toggle-layer file resolves wholesale to the layer
that owns it — the fixes this line carried for those layers are all in
the union below, and its own copies retire. What this layer keeps is what
is genuinely its own: ios/, the fixture capture script, the testing
runbook, and the companion docs.
…d-soni#198)

* Bundle the server's dependencies so the packaged app can start

0.1.24 built, signed, notarized and installed cleanly, then died on
every launch:

  ERR_MODULE_NOT_FOUND: Cannot find package 'zod'
    imported from Resources/server/config.js

The packaged app ships no node_modules by design (electron-builder.yml
line 21 calls the three pieces self-contained, line 33 excludes them).
build:server was plain tsc, which transpiles without bundling, so the
`zod` import milind-soni#194 introduced survived verbatim into a tree with nothing
to resolve it against. zod was the first bare import the server ever
had, so the invariant had never been tested.

Bundle every entry point with esbuild after tsc, mirroring
scripts/bundle-updater.mjs which already vendors electron-updater for
the same reason. All six are bundled, not just index.ts: the proxies run
as their own processes and import nothing external today, but the next
one that does would fail the same silent way. Entry points keep their
relative paths, which the proxy lookups depend on.

Both gates that should have caught this were blind to it. The unit suite
runs inside the repo, where a bare import resolves from ./node_modules;
the Windows packaging check asserts index.js exists but never runs it.
So the new smoke test copies dist-server OUT of the repo before starting
it, and CI now starts the real packaged copy on the runner. Verified by
mutation: reverting to plain tsc output fails the smoke test with the
original ERR_MODULE_NOT_FOUND.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* Build the server before smoke-testing it

The smoke test assumed dist-server was already on disk. It is gitignored
(milind-soni#190), so a fresh CI checkout has never built it and the test died on
ENOENT before it could prove anything. It passed locally only because a
build happened to be sitting there.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* Never let scratch cleanup fail the smoke test

Windows holds file handles briefly after the process that owned them
dies, so removing the scratch dir right after the kill raised EPERM and
failed a run whose server had actually started fine. Linux raises EACCES
the same way (f66d30f).

Cleanup is housekeeping; the assertion is the test. Mutation re-checked:
plain tsc output still fails.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
testDecodesThePagedFleet unwraps a group with three messages and a page
boundary, and the capture script never created one — the committed fixture's
room was an accident of whichever harness the fixtures were last captured
against, and the first regeneration on a clean machine failed the test.

The script now creates the room itself: directly on the harness, because
room creation is deliberately not on the sidecar's allowlist and so is setup
the phone cannot perform — then five messages through the sidecar, captured
at messages=3, which is what makes the pinned count and hasMore=true
properties of the capture rather than of history.
* Working folder: pin the default for cloud runs, clear via an empty field

Follow-ups from CodeRabbit's review of milind-soni#183 (merged): a cloud run now
pins task.cwd = null so the header chip never shows the bot's host
folder for a task that runs on the box; the non-desktop text field
sends null when emptied, which is what the server takes as "clear" —
it sent "" and was rejected. (The third finding, the ~ path boundary,
was already fixed on main in src/lib/short-path.ts.)

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* Make cloud task folder pin authoritative

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: milind-soni <milindsoni201@gmail.com>
* Let the store announce its own writes, and give bots a real activity state

2.6 — emit inside the store. Every write to the store used to need a
matching broadcast at the call site, and there were 31 of them across
four files. Miss one and the UI drifts from disk; broadcast without
writing and a restart loses what the user just watched. The Store now
emits a typed StoreChange after each write (message, message.patch,
thread, bot, bot.deleted, group, group.deleted) and index.ts maps those
onto SSE frames in ONE subscriber. The mirror broadcasts are gone. The
few endpoints whose callers need a full transcript on the wire (task
create/switch, team import) still send their richer payload on top.

2.7 — activity state. busy could not tell working from waiting-on-you
from a stalled engine. BotRecord.activity is one of working |
waiting-on-you | idle | no-signal | dead; busy is DERIVED from it in
store.setActivity, the only place runtime state changes now, so all
existing busy readers keep working unchanged. Transitions: dispatch →
working; a card that actually reaches a human → waiting-on-you; answered
→ working; settled → idle; a setup error (the engine could not start) →
dead until the next dispatch. no-signal is the slot the liveness clock
(milind-soni#193) fills. The sidebar preview says "Waiting for you…" when that is
what is happening.

Items 2.6 and 2.7 of docs/plans/agent-harness-upgrades-v2.md.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* store emits: drop the last manual group broadcasts; persist the transient reset

Review follow-ups (CodeRabbit on milind-soni#199):
- every patchGroup was followed by a manual broadcastGroup, so each room
  update sent two identical frames now that the store emits. The manual
  calls are gone and broadcastGroup with them — the store's group change
  is the one source.
- the load-time reset of busy/activity was in-memory only: a bots.json
  left by a process that died mid-turn kept saying busy on every load. The
  reset now marks the file for a save when it changed anything.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* Complete store-owned update delivery

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: milind-soni <milindsoni201@gmail.com>
A turn ends across three server frames: the settled reply message,
turn.completed, then the bot patch that flips busy off. The tail derived
its "working" dots from busy && !streaming alone, so in that window the
dots popped back in under the reply that just landed, then vanished a
beat later — and the pinned scroll re-anchored around each height
change. That grow-shrink-jump was the end-of-stream jitter.

showWorkingDots (src/lib/turn-tail.ts) now decides the tail: a settled
bot text at the end of the transcript means there is nothing to wait
for, so the dots stay hidden until something actually new starts — a
tool chip, the next user prompt, or (in rooms) a different speaker
taking the floor. Rooms gate per speaker: a reply from a previous
speaker doesn't cover the bot now on the floor, and a message with no
attribution fails toward showing the indicator.

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2.3 — durable delegations. The per-thread handoff queue lived in a Map
and died with the process: a delegation queued right before a restart
never ran, silently. It is now written to ~/.openmausbot/delegations.json
on queue, drain, and discard, loaded at boot, and drained through the
same path a settled turn uses — target and approvePeerComms are still
re-checked at drain time; a source bot that no longer exists is skipped.
Provider permissions still die with the process (nobody can answer for an
unattended bot); queued work is not a permission.

2.4 — secret redaction, rescoped. The harness never stores tool RESULTS
— CLI drivers run tools inside the CLI and the harness only sees names
and outcomes — so there was no tool-result body to redact. What it does
store, and now replays into every rebuild, is the bot's reply text,
activity chip titles (an ACP engine's title can be the whole command
line), and permission card summaries (the command being approved). Those
can carry keys.

- redactSecretsInText: high-precision content patterns only — known key
  prefixes (sk-, ghp_/github_pat_, xox*-, AKIA, AIza, npm_), JWTs, PEM
  private-key blocks, Bearer tokens, and secret-shaped key=value. No
  generic hex/base64 heuristics: those rewrite real code.
- Applied at the single store write for role "bot" (text, tool.name,
  card title/subtitle/summary). What the user typed stays as typed —
  pasting your own .env for the bot to use is your call.
- redactSecrets (the native tee) gains the same content pass over string
  values, and the events NDJSON now goes through it too.
- Live streaming deltas are not redacted (tokens split across events);
  the settled message that replaces the bubble is, and that is what is
  stored and replayed.

Items 2.3 and 2.4 of docs/plans/agent-harness-upgrades-v2.md.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: milind-soni <milindsoni201@gmail.com>
* Answer approvals with a typed outcome, and notice a bot repeating itself

2.5 — fail-closed approvals. respondToRequest used to resolve void or
THROW when the ask was gone (turn ended, broker died, engine has no
asks): the user's Allow click became a 500 and the card sat open
forever. It now resolves a typed RequestOutcome — allowed-once |
rejected | answered | unavailable — across all six drivers, and the
harness branches on it: `unavailable` settles the card as dismissed and
leaves a chip ("Couldn't deliver that answer — the request is no longer
open, so the action was not run"), returning 200 {outcome}. The auto-
approve path hands the ask back to a human on `unavailable` instead of
on a throw. request.resolved is typed too: behavior allow|deny|answer,
source user|auto|timeout|system|unavailable|peer — codex used to stamp
every resolution "user", timeouts included; ACP and codex now say
timeout/system where that is what happened. Kept as-is on purpose: the
timeout NOTE text (post-decision guidance to the model, the decision
itself was already typed) and request.opened's own tool/summary.

2.8 — repeat-call detection (observe only). server/repeat-detector.ts
counts identical calls per turn keyed on tool + arguments — from every
permission ask's summary and from ACP item titles; a bare tool name is
never counted (five "Bash" may be five commands, and Claude's
item.started carries only the name). At 5, 10 and 20 a chip says "Same
call repeated N× — Bash: git status — it may be stuck". No auto-stop:
the human has Stop; cutting a call off needs the harness to own the
call (3.2/3.3).

Items 2.5 and 2.8 of docs/plans/agent-harness-upgrades-v2.md.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* Harden approval outcomes and repeat tracking

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: milind-soni <milindsoni201@gmail.com>
The server has tallied per-task usage since tasks landed, and wireTask
already carries it — the client just never typed or rendered it. Task
rows now show a quiet combined total ("12.3k") next to the timestamp,
with the input/output split on hover; the picker button keeps its shape
and carries the open task's tally in its hover title instead.

Formatting lives in src/lib/format-tokens: zero hides the chip, sub-1k
spells out "tokens", then k/M with one decimal. Rounding runs on integer
tenths so 999,950 promotes to "1M" instead of float-rounding down.

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
…#203)

While a reply streams, a code block now highlights after its content has
been unchanged for 250ms instead of waiting for the stream to settle.
The result is written to the content-hashed highlight cache, so the
settled bubble (a fresh component instance) mounts highlighted from
cache instead of popping from plain <pre> to Shiki output a beat later.

The debounce rides the effect lifecycle: every content change re-runs
the effect and its cleanup clears the pending timer, so a growing block
never tokenizes per token. Partial fences may cache under their own
hash — harmless, since the final content gets a distinct key and the
cap evicts strays. Non-streaming rendering is unchanged.

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Reviewed and integrated from milind-soni#161, updated for the SQLite-backed server, hardened at the paired-device boundary, and verified on desktop CI plus an iOS simulator build.
⌘K / Ctrl+K from anywhere opens a centered switcher: bots and rooms
answer instantly from local state, transcript hits ride the existing
/api/search endpoint a 150ms debounce later (same stale-response guard
as the sidebar search). Empty query is pure switcher mode — every bot
and room, no message section.

Ranking lives in src/lib/palette-rank.ts as a pure helper: prefix
matches outrank substring matches, matching is case-insensitive, and
each tier preserves the caller's order so pinned-first lists survive.
Arrow keys move one flat cursor across the sections, Enter selects
(message hits open the owning bot or room, switching to the hit's task
like the sidebar does), Esc closes. Row hover tracks mousemove rather
than mouseenter so hits arriving under a resting pointer can't steal
the keyboard selection.

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Long computer-use threads mount hundreds of rows — inline base64
screenshots included — so the DOM stays heavy even though the memoized
list bails out of re-renders. ChatView and GroupView now mount only the
last 120 messages; a quiet "Show earlier messages (X more)" pill at the
top of the transcript pulls the boundary back by another 120 per click.

The boundary is a pure decision in src/lib/transcript-window.ts: it is
anchored per bot+task (render-phase reset on bot.id/threadId switch), so
appends grow the window instead of sliding rows out from under the
reader, and a thread that shrinks beneath a stale boundary — branch
switch, edit rewind — falls back to a fresh tail window rather than
blanking the transcript.

Expanding must not yank the viewport: the click captures scrollHeight,
then a layout effect shifts scrollTop by the growth before paint
(browser scroll anchoring is disabled on these containers). The
bottom-follow scrollTo keys on the FULL list's length, so expansion
never re-triggers it, and expanding breaks follow the way any
scrollback reading does. Tail-derived logic — working dots,
lastBotTextId, regenerate — still computes from the full list.

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
A bot's MEMORY.md rides into every turn's system prompt and its
memory/*.md topic files steer what it believes, yet none of it was
visible anywhere in the UI — the user had to know the workspace path
and open the files by hand to see (or fix) what a bot had decided to
remember.

Three routes expose the files the workspace already owns:
GET /api/bots/:id/memory returns the WHOLE MEMORY.md (not the load
budget's cut — an editor must show everything, with `truncated`
flagging what the prompt loader would drop) plus a name+size listing
of memory/ topics; PUT writes it back, string-checked at the boundary
and capped at 256KB so a runaway paste fails with an explanation; GET
/memory/topics/:name serves one topic read-only. Topic names pass one
strict gate (single plain-markdown path segment, decoded before it is
judged) in both the route and the workspace helper, so no coat of
percent-encoding turns the route into an arbitrary-file read — the
tests plant real files at the traversal targets and prove every
encoding of ../ dies as a 400 before the filesystem is touched.

The Settings panel gets a Memory card in the existing card language:
collapsed by default (most visits never look, and expanding re-reads
so mid-session notes appear), a monospace editor with save + budget
warning, and the topic list opening into a read-only viewer. Keyed by
bot id so switching bots can never show one bot's notes under
another's name.

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
The phone's chat field was type-only. The desktop composer already has a
mic — press to talk, press to stop, partials land in the box so you can
edit before sending — and the architecture note for this app was that
SFSpeechRecognizer is on every iPhone. Composer dictation is the smaller
half of that, and it is the half you actually use while walking around.

Same engine as electron/resources/speech-helper.swift: an AVAudioEngine
tap into SFSpeechAudioBufferRecognitionRequest, on-device when the
recognizer supports it, locales from the user's preferred languages
rather than a hardcoded en-US. Composer mode, not call mode — no silence
endpointing. Tap the mic to stop. The last partial is what you send.

The join (typed text + live transcript) lives in CompanionCore so it can
be tested without a phone. Partials replace each other after the text
that was already in the field; they never stack.

The mic stays next to send. Hiding it once text arrives is the desktop
pattern, where Escape stops listening and the toolbar only has room for
one action. A phone has neither — this is how you stop, and how you add
another sentence by voice after the first one. Backgrounding or an audio
interruption stops the session.

NSMicrophoneUsageDescription and NSSpeechRecognitionUsageDescription are
in project.yml. Without them the first tap crashes rather than prompting.

Co-authored-by: Computomatix <mnthr7@users.noreply.github.com>
Cancel an in-flight start instead of racing a second tap during the
permission prompt. Fail closed when the locale has no recognizer.
Tear the audio tap down before endAudio so a late buffer cannot
fail the recognition task. Treat opening the computer panel as
leaving chat (NavigationStack keeps ChatView mounted).

Co-authored-by: Computomatix <mnthr7@users.noreply.github.com>
Cancel an in-flight start when send is tapped during the permission
prompt. Drop the unused SwiftUI import. Label the README tree fence
(MD040). Document the speech Info.plist keys and that a denial is
shown on that same attempt.

Co-authored-by: Computomatix <mnthr7@users.noreply.github.com>
@mnthr7 mnthr7 closed this Aug 18, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.