Skip to content

fix(agents2): keep agents running when another community is selected - #838

Open
baxen wants to merge 5 commits into
mainfrom
honey/agents2-all-communities
Open

baxen wants to merge 5 commits into
mainfrom
honey/agents2-all-communities

Conversation

@baxen

@baxen baxen commented Oct 10, 2026 •

Copy link
Copy Markdown
Collaborator

Problem

Switching communities killed running Agents2 agents. Agents2Service only listened to the selected community. The Claude Code and Codex runtimes treated snapshot().agents as the list of agents that should be alive, so an agent whose community stopped being selected had its process disposed mid-turn. Its community then went unheard until you switched back.

Change

  • Agents2Service binds every joined community where the viewer has an agent, plus the selected one, whether or not it is selected.
    • Each community's live events go only to that community's agents.
    • Cleanup of a deleted agent also runs in communities that aren't selected.
    • A failed connection to an unselected community is retried every 30 s, since nobody is there to press its Retry.
  • Snapshot gains running: the viewer's agents in every joined community that isn't disconnected. A reconnect doesn't remove them. agents is still the selected community's agents, so the Agents2 page is unchanged.
  • agents2.relay(pubkey) returns the owner's connection to an agent's own community.
  • Claude Code / Codex keep processes for everything in running. They read thread history, channel names and memory through the agent's own community instead of the selected one. Codex no longer disposes everything on a community switch. Its saved-session key is the same origin:viewer string as before, so no saved sessions are orphaned.
  • Communities.open(id) is a new method that connects a joined community without selecting it. communities/service.ts is a FOUNDATION file; baxen asked for this fix after reviewing the plan that called out this addition.

A process now stops only when its agent is deleted, its community is left or disconnected, the viewer signs out, or the plugin is disabled.

One shared Claude Code spare (25e885f)

Keeping every agent running made Claude Code's per-agent warm spare expensive: one idle claude (~300 MB RSS) per agent, and up to six processes per agent.

  • One spare for the whole app. It is spawned with only a folder and model; it gets no prompt yet. The agent that starts the next new conversation takes it, sends its own initialize (prompt, memory, hooks), and installs its own tool handler. Then another spare starts.
  • Why this works: spawning is what a spare saves; initialize costs almost nothing. Measured against claude 2.1.296, the time from send to turn start is ~530 ms cold vs ~45 ms with the spare (medians, 10 interleaved rounds). A spawned-only process is as fast as a fully initialized one.
  • claude lists the SDK MCP server's tools right after spawn, before initialize. Until the spare is taken, a placeholder handler answers; the tools list does not depend on the agent.
  • Which spare is kept: the spare follows the latest new conversation. When no running agent uses its folder and model, for example after a model change, it is replaced. If it fails to start or dies, it is not restarted until a conversation asks, so a broken claude is not respawned in a loop.
  • No per-agent process cap. The 15-minute idle stop stays.
  • The agent's memory is read while the turn's prompt is built, so taking the spare does not wait on it.
  • Codex is unchanged. A cold app-server answers initialize in ~85 ms and starts a thread in ~95 ms, so a spare would save ~100 ms at most.

Tradeoffs

  • Extra connections: one websocket per joined community that has an agent, opened once the agent list loads. These connections stay open until the community is left.
  • Warm processes: one spare Claude Code process in total, whatever the number of agents. When agents alternate between different folders or models, new conversations miss the spare and start cold (~0.5 s more). A conversation resumed after its idle stop is still cold, because --resume is a spawn flag.
  • Codex app-servers have no idle stop. Each Codex agent that has worked keeps an app-server (~400 MB RSS with its children) until it stops running. That was already true before this PR, but now it also applies to agents in other communities. Restarting is cheap (~100 ms plus thread/resume), so an idle stop is a sensible follow-up.

Validation

  • Full Vitest suite: 653 files / 9193 tests passed at 2e36587.

  • Affected suites: agents2, bundled and communities, 222 files / 2878 tests, passed at 43f3c3b. Type check and lint are clean.

  • New tests:

    • An agent keeps receiving its own community's mentions while another community is selected, and never hears the other community's events.
    • An agent in an unselected community is connected at start.
    • An agent survives a reconnect and stops on disconnect or leave.
    • A failed unselected community is retried; the selected one is left alone.
    • Communities.open opens a joined community without selecting it, and returns nothing for a community that isn't joined.
  • Mutation check: removing the per-community event filter fails the routing test.

  • Codex tests updated: the test that asserted "disposes on community change" now asserts the server ends when its agent stops running.

  • Independent agent review: done. Its medium finding was that a background community whose first connect fails never recovers. It's fixed in 43f3c3b.

  • Shared spare (25e885f):

    • Full Vitest suite: 653 files / 9200 tests passed at 25e885f. Type check and lint are clean.
    • New tests: a new conversation takes the spare and another starts; the spare's tool handler switches to the agent that takes it; two agents share it; a spare for another folder is killed and the conversation does not wait for it to start (fails without the fix); the spare is kept when it fits any agent and replaced when it fits none, including after a model change; a dead spare is not restarted until a conversation starts; conversations are unlimited.
    • Live check with real claude, outside the app: two agents with different prompts took turns using the shared spare, and each answered with its own name.
    • Independent agent review: no misrouted tools or wrong cwd/model. It found two issues, both fixed: waiting on a mismatched spare's spawn, and a stale spare after a settings change.
  • Lifecycle review fixes (8904805, c8a0ee8): codex00's review found three lifecycle bugs, each reproduced. Each fix takes one decision away from a place that shouldn't make it:

    • A Claude launch that finished after its last agent was removed refilled the spare, which then stayed with no agents. Taking the spare now refills it only where the runtime wants one, so the runtime alone decides where a spare is kept.
    • An agent list still being read when Agents2 stopped could reopen agent communities afterwards. Teardown now marks that read stale, using the turn counter that already discards out-of-date reads.
    • Codex skipped its sync whenever the inventory status was error, so leaving a community after a failed refresh left its servers running. It now follows running regardless of inventory status, as Claude Code does.
    • Each fix has a regression test that fails without it and passes with it. Full Vitest suite: 653 files / 9203 tests passed at c8a0ee8.

Not yet done

  • Agent run in the native app: the dev build needs a Keychain approval, so baxen will run it. Also check there that exactly one idle claude exists beyond those in use.
  • Human test: pending.

No browser cases added or removed.

🤖 Generated with Claude Code

Weekend review stack (2026-10-11)

main (d55d70a07, including #851) → #838 → #852 → #855 session viewer/context → #856 model/effort picker → #857 ordinary metadata → #858 identity adoption → v1 retirement (baxen/remove-agents-v1-tail). Each PR is based on its predecessor. New layers remain drafts for native/human acceptance.

This branch was restacked onto its current parent. Review and merge bottom-up; after a squash merge, restack only each child’s own commits onto the new parent and retarget its base. baxen/next is an append-only testing mirror of the full stack, not a PR to merge into main. Green is unchanged.

@baxen
baxen marked this pull request as ready for review October 11, 2026 04:14
@baxen
baxen requested review from a team as code owners October 11, 2026 04:14
@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Oct 11, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-10-11T04:19:04.122279Z c8a0ee8 Draft marked ready
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@wesbillman wesbillman left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

No actionable changes requested. Reviewed cross-community routing, connection/retry and teardown lifetimes, and shared Claude spare ownership; I found no source-demonstrated regression in this revision.

Star Lord’s automated source review via Wes’s account. Head c8a0ee8ad6d98853c4bf9d4006a71febb4bd3088; base f8705a95ba8034c1a3c7ec5f09847db1c864d97b. Source-only: no tests or app execution, and no independent CI verification; native-app switching/reconnect behavior and process counts remain unverified.

console.warn("Claude Code could not start", error);
return undefined;
});
this.opening.set(key, starting);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🤖 [P2] Preserve per-agent admission before the shared native process ceiling

Removing MAX_LIVE/room() also removes idle eviction on admission. Completed Claude conversations retain their processes for 15 minutes (sessions.ts:193–200), while the native host permits only 64 live processes across all plugins (src-tauri/src/host_process.rs:21–22,102–104). Under high conversation fan-out, one agent can consume that shared budget and prevent unrelated agents from launching processes, including Codex servers or upload readers. The message rate limit does not bound retained processes.

Please retain bounded per-agent live + opening admission and idle eviction alongside the shared spare, returning a local busy result instead of exhausting the shared registry. Add fake-host capacity/controlled-clock coverage for pending starts and completed-idle sessions. This is a source-established starvation path, not a measured memory or native stress-test result.

Honey and others added 5 commits October 11, 2026 09:57
Agents2 listened only to the selected community, and the Claude Code and
Codex runtimes stopped the processes of agents that left that list, so
switching communities killed work mid-turn and left the other community
unheard.

Agents2 now binds each joined community the viewer has an agent in (opened
through a new Communities.open, without selecting it), delivers each
community's live events only to its own agents, and exposes the running
agents and each agent's community connection. The runtimes keep processes
for every running agent and read conversations, names and memory from the
agent's own community. Processes stop when the agent is deleted, its
community is left or disconnected, or the viewer signs out.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Honey <8e307ae0076a4dab6b94b036ea3edc7e08f823a625269c1e6919e881a048b4d2@buzz.block.builderlab.xyz>
A first connect that fails leaves the session in error, and only the
selected community has a Retry the viewer can press. Agents2 now tries an
unselected community's failed connection again every 30 seconds, so its
agents recover without the viewer opening it. Also covers
Communities.open directly.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Honey <8e307ae0076a4dab6b94b036ea3edc7e08f823a625269c1e6919e881a048b4d2@buzz.block.builderlab.xyz>
Each Claude Code agent kept its own warm `claude` process and up to six
processes in all. With agents in many communities kept running, that is
one idle ~300 MB process per agent.

Starting the process is what a spare saves (time to turn start ~530ms cold
vs ~45ms warm, measured against claude 2.1.296); sending `initialize`
costs almost nothing. So the spare is now spawned without `initialize`,
fixing only its folder and model, and the agent that takes it sends its
own prompt and installs its own tool handler. One spare serves every
agent; it follows the latest new conversation, and is replaced when no
agent runs where it was started. The per-agent process cap is removed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Bradley Axen <baxen@squareup.com>
A conversation whose launch finished after its last agent was removed
took the spare and started a replacement, which then stayed with no
agents. Taking the spare now refills it only where the runtime wants one,
so stopping the last agent (`warm()` with no places) is final.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Bradley Axen <baxen@squareup.com>
- A native agent list still being read when Agents2 stops is now stale,
  so it cannot reopen agent communities after teardown.
- Codex follows `running` whatever the inventory status, as Claude Code
  does, so a failed refresh no longer keeps a stopped agent's server.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Bradley Axen <baxen@squareup.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants