fix(agent): unify goal activation and harden long-running execution - #3211
Merged
Merged
Conversation
bobleer
marked this pull request as ready for review
September 23, 2026 06:31
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Plain
/goal <objective>messages now activate the persistent goal on the executing host, including mobile/IM relay, peer-host and CLI submissions. Previously the Web UI intercepted the command while other clients sent an ordinary message, and the runtime excluded that prompt from goal accounting and continuation.The lifecycle audit also found that headless jobs could report success after the first turn, sessions shared one accounting state, and budget wrap-up could repeat. This PR fixes those behaviors and aligns first-turn, edited-objective and continuation prompts around the same evidence-based completion contract.
Behavior
/goalthrough CLI startup, chat, steering and the action registry; keep existing control words reserved.execand Detached Dispatch observing matching goal continuations. An intermediate successful turn does not finish the job; incomplete stopped goals return errors. Persist the dispatch job's current continuation turn and keep CLI cancellation targeted to the observed turn.No wire fields, persisted shapes, Cargo features or model settings change. Old hosts retain their old behavior; upgrading a controller cannot upgrade a target's goal implementation.
Verification
AI-assisted implementation and full-flow review. Focused local checks:
cargo test --locked -p openbitfun-agent-runtime --no-default-features --features agent-runtime --test agent_long_horizon_contractscargo test --locked -p openbitfun-agent-runtime --no-default-features --features agent-runtime --test agent_session_contracts scheduler_contracts::cargo test --locked -p openbitfun-core --no-default-features --features agent-runtime,git,remote-workspace --lib -- thread_goal_ host_queue_ agentic::goal_mode:: agentic::execution::round_executor::testscargo test --locked -p openbitfun-cli --bin openbitfun -- modes::exec::tests:: dispatch::worker::tests:: dispatch::store::tests::goal_continuation_ actions:: modes::chat::tests::goal_prompts_ modes::chat::tests::pending_session_operation ui::command_menu:: ui::command_palette::cargo check --locked -p openbitfun-desktop,cargo build --locked -p openbitfun-desktop, andcargo test --locked -p openbitfun-desktop --lib(468 passed, 11 ignored)pnpm run fmt:rs,pnpm run check:repo-hygiene, andgit diff --checkCoverage includes parser boundaries, legacy metadata, remote storage identity, repeated steering, duplicate queue promotion, concurrent usage reads, terminal and failed-turn accounting, stale mutations/retries/queues, budget wrap-up, blocked resume, headless goal-run transitions, prompt escaping, dispatch continuation-turn persistence and stale-owner rejection, and existing CLI output/settlement/dispatch contracts. A scheduler regression also exercises an unrelated session completing while a goal continuation is retrying.
Remote workspace identity/storage and host-owned queue execution are exercised with fixtures. Physical mobile relay, IM, SSH and Peer Device sessions, real-provider multi-turn execution and a disconnected remote Detached Dispatch run have not been exercised end to end. Headless continuation decisions and existing worker/exec projection paths have focused tests; this is not a claim of real-provider task completion.
Windows CI hit the existing narrowly tolerated Tauri desktop test loader failure (
0xc0000139 / STATUS_ENTRYPOINT_NOT_FOUND) before tests executed. Windows desktop library tests therefore did not run; Windows compilation, Core, CLI and other contract checks passed. Linux CI and local macOS desktop library tests passed. This PR does not change that CI exception.Remaining boundaries
The existing 100-continuation safety stop remains; explicit resume resets that window. Optional token accounting covers main-session non-cached input plus output, not provider-wide or child-session billing, and is a soft limit. Host restart recovery still requires explicit session inspection/resumption. Prompt instructions strengthen completion auditing but cannot prove arbitrary user objectives independently of model judgment. These limits are documented in the CLI README.
Checklist