Skip to content

Server resumes a thread into whichever provider thread its latest event names, with no cross-thread ownership check #3461

Description

@Willhong

Summary

turn.submit and thread resume take the provider thread id from the thread's most recent stored event that carries one (getLastStoredProviderThreadId). Any single mis-stamped event therefore redirects the whole thread into another thread's provider session, and nothing on the server refuses a resume when a different thread in the same environment already owns that provider thread. On 2026-09-11 this turned one daemon-side identity mix-up (#2327, fixed in #3460) into a bb thread running a turn inside another thread's Codex rollout and the owner being locked out with already has an active writer.

Versions and environment

  • bb 0.42.0 run from source at 4a5e7557f (server, host daemon on the same macOS host); behavior verified unchanged on main at 8d32c61e by reading the code paths linked below.
  • macOS 25.6.0, Node 25.9.0, Codex CLI 0.154.0.
  • Unmanaged local environment env_3rm6y2khyw, provider codex.

Steps to reproduce

Live sequence (from the events table; ids are real):

  1. Have two Codex threads on one environment bridge: thr_fh9dq7fvsu (Codex 01a08d8d) and thr_r6xinwrds3 (Codex 01a08ddd).
  2. Get one event of thr_fh9dq7fvsu stamped with providerThreadId: 01a08ddd. On main before Retry transient Codex writer-lock refusals and cover the writer handoff #3460, a message edit on thr_fh9dq7fvsu followed by spawning thr_r6xinwrds3 does this (see turn/completed can precede command completion and wedge later turns between active-writer and no-session errors #2327 comment for the mechanism); its 09:32:30 turn/completed carried 01a08ddd.
  3. Send any message to thr_fh9dq7fvsu later (bb thread tell thr_fh9dq7fvsu "잠시 꺼줘" at 10:12:46).
  4. Send any message to thr_r6xinwrds3 (10:19:28).

Closest isolated repro without the daemon bug (not yet written as a test): append a turn/completed event to thread A with thread B's providerThreadId, then call prepareTurnSubmitCommandPayload for A and observe providerThreadId equals B's.

Expected vs actual

Actual, step 3: thr_fh9dq7fvsu emitted thread/identity {providerThreadId: 01a08ddd} (event seq 4658, 10:12:47) and the turn ran inside 01a08ddd's rollout; ~/.codex/sessions/2026/09/11/rollout-2026-09-11T09-28-44-01a08ddd-….jsonl contains the 01:12:51Z user message that thr_r6xinwrds3 never sent.

Actual, step 4 (bb thread log thr_r6xinwrds3):

── Error ───────────────────────────────────────────────────
  Command turn.submit failed
  thread 01a08ddd-e5ae-7771-be6c-c3a202882b11 already has an active writer

Expected: the server resumes a thread only into the provider thread that thread was started or last legitimately re-identified with, and refuses (or at least logs and rejects) a resume whose provider thread is currently owned by a different thread in the same environment, before any bridge command is sent.

Evidence

  • Resume target selection:
    export function getLastStoredProviderThreadId(
    db: DbQueryConnection,
    threadId: string,
    ): string | null {
    const latestProviderRow = db
    .select({ providerThreadId: events.providerThreadId })
    .from(events)
    .where(
    sql`${events.threadId} = ${threadId}
    AND ${events.providerThreadId} IS NOT NULL
    AND ${events.sequence} > COALESCE((
    SELECT MAX(context_clear.sequence)
    FROM events AS context_clear
    WHERE context_clear.thread_id = ${threadId}
    AND context_clear.type = 'system/operation'
    AND json_extract(context_clear.data, '$.operation') = ${THREAD_CONTEXT_CLEAR_OPERATION}
    AND json_extract(context_clear.data, '$.status') = 'completed'
    ), 0)`,
    )
    .orderBy(sql`${events.sequence} DESC`)
    .limit(1)
    .get();
    if (!latestProviderRow?.providerThreadId) {
    return null;
    }
    return latestProviderRow.providerThreadId;
    (latest provider_thread_id of any event type after the last context clear).
  • Callers:
    const providerThreadId = requireProviderThreadId(
    args.providerThreadId ?? getLastProviderThreadId(deps, args.thread.id),
    args.thread.id,
    );
    and
    const providerThreadId = getLastProviderThreadId(deps, thread.id);
    .
  • No query joins events.provider_thread_id across threads before resume; bb.db on the affected machine shows thr_fh9dq7fvsu and thr_r6xinwrds3 both carrying 01a08ddd in events.provider_thread_id from 09:28:46 onward.
  • Daemon log: ~/.bb/logs/host-daemon.38.log lines 1783–1784 (turn.submit → JsonRpcResponseError … already has an active writer, code -32000, recovery: null).
  • Investigation thread: bb thr_2q5raz5a83 on host host_erd6f98r4i.

What you ruled out

Suggested priority and effort (optional)

Medium. It needs the daemon bug (or any future stamping bug) to trigger, but when it does the user's messages land in another thread's session and the original thread wedges with no recovery path. Effort small: derive the resume target from thread/identity events (plus fork/rewind identities) instead of any event, and add a uniqueness check on (environment_id, provider_thread_id) among live threads before dispatching thread.resume / turn.submit.

AGENT GENERATED

No activity

Activity on this issue will appear here.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    confirmed-reproBug reproduced again from a clean trusted checkout; see linked reportthreadsTurns, timeline, messaging, forks

    Type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions