diff --git a/.agents/skills/testing-workflow/SKILL.md b/.agents/skills/testing-workflow/SKILL.md index b37177ab..e082ffd3 100644 --- a/.agents/skills/testing-workflow/SKILL.md +++ b/.agents/skills/testing-workflow/SKILL.md @@ -34,6 +34,7 @@ suites are outside this distribution's verification. | Sandbox on macOS | `pnpm test:sandbox` | | Source-sync, workflow and release tools | `pnpm test:release-tools` | | npm release archive installation | `MCODE_RELEASE_TAG=vX.Y.Z MCODE_RELEASE_ARCHIVE=/path/to/package.tar.gz pnpm verify --profile package` | +| TUI source and test lint | `pnpm lint:tui` | | Types and standalone build boundary | `pnpm typecheck`, `pnpm build`, `pnpm check:standalone` | | Published files and generated paths | `pnpm check:source`, `pnpm check:tsconfig` | diff --git a/CONTRIBUTING.md b/CONTRIBUTING.md index cddfe258..8aa6b387 100644 --- a/CONTRIBUTING.md +++ b/CONTRIBUTING.md @@ -39,7 +39,7 @@ pnpm install --frozen-lockfile pnpm verify ``` -`pnpm verify` runs the complete gate list in the same order as GitHub CI. Normal PR and main-branch checks use Node.js 24 on Linux and macOS. The Linux job runs the full profile; the macOS job uses `pnpm verify --profile platform`, which omits only the duplicate TypeScript compiler check. Windows runs the focused `pnpm verify --profile windows` contract on PRs; the profile is Windows-only and fails closed elsewhere. It checks source inventory, release tooling, build boundaries, artifacts, and Windows-specific tests without running the full capability suite. Gates that depend on platform behaviour are selected by platform rather than skipped silently; run `pnpm verify --list`, `pnpm verify --profile platform --list`, or `pnpm verify --profile windows --list` to inspect each plan. Individual gates remain available as their own scripts, such as `pnpm typecheck` or `pnpm test:byok`, while you iterate. +`pnpm verify` runs the complete gate list in the same order as GitHub CI. Normal PR and main-branch checks use Node.js 24 on Linux and macOS. The Linux job runs the full profile; the macOS job uses `pnpm verify --profile platform`, which omits only the duplicate TypeScript compiler check. Windows runs the focused `pnpm verify --profile windows` contract on PRs; the profile is Windows-only and fails closed elsewhere. It checks source inventory, release tooling, build boundaries, artifacts, and Windows-specific tests without running the full capability suite. Gates that depend on platform behaviour are selected by platform rather than skipped silently; run `pnpm verify --list`, `pnpm verify --profile platform --list`, or `pnpm verify --profile windows --list` to inspect each plan. Individual gates remain available as their own scripts, such as `pnpm typecheck` or `pnpm test:byok`, while you iterate. This distribution ships no ESLint stack or lint gate; the source repository's `lint:tui` step and its Airbnb/TypeScript/import configuration are not part of the public source. CI writes per-gate timing and exit metadata to the Job Summary and a seven-day `verification--node--` artifact. For a local report, set `MCODE_VERIFY_REPORT_DIR` to a directory outside the repository. Reports distinguish `PASS`, `FAIL`, intentional `SKIP`, and `NOT_RUN` after a failure. JSON is checkpointed before and after each gate; a cancelled run may leave `RUNNING`, which is not a pass. If installation fails before verification starts, no verification report is available. Reports do not collect command output, environment variables, or runtime data; read the corresponding gate's job log for failure details, including the existing bounded BYOK timeout diagnostics. CI jobs have a 15-minute verification limit and a 10-minute release-audit limit. diff --git a/README.md b/README.md index 76d7ad3d..1281f200 100644 --- a/README.md +++ b/README.md @@ -280,6 +280,18 @@ The Web UI originated as the community **mcode-webui** plugin and was migrated i This repository also hosts issue reporting for the MiniMax Code desktop app. The published source covers the terminal TUI, headless CLI, and ACP; it does not include the desktop application's source. Select the affected product when filing an issue. For a desktop bug, include the app version, operating system, and a log upload ID if available from **Settings → General → Upload logs**. For a CLI bug, include `mcode --version`, your interface, and a minimal reproduction. Remove credentials and private project content from reports. +## Feedback and contact + +| Channel | Use it for | +| --- | --- | +| [GitHub Issues](https://github.com/MiniMax-AI/minimax-code/issues/new/choose) | Public bug reports, feature requests, and questions about the CLI or desktop app. | +| [MiniMaxCode@minimax.io](mailto:MiniMaxCode@minimax.io) | General feedback and support inquiries. | +| [security.mcode@minimax.io](mailto:security.mcode@minimax.io) | Private vulnerability reports. Send reproduction details and redacted evidence here; see [Security](SECURITY.md). | +| [Discord](https://minimax.io/discord) | Community discussion and feedback. | +| [Feishu feedback group QR code](https://cdn.hailuoai.com/hailuo-video-web/public_assets/minimax_code_feishu_group_url.png) | Chinese-language community feedback. Scan with Feishu, or find the QR code in the Chinese desktop app under the user menu → **Contact us → Feishu**. | + +Follow [MiniMax on X](https://x.com/MiniMaxAgent) for updates. Keep vulnerability details, credentials, and private project content out of public issues and community chats. + ## License First-party code defaults to [MIT](LICENSE). Existing file-level and package-level licenses remain in place. See [third-party notices](THIRD_PARTY_NOTICES.md) and [license status](LICENSE-STATUS.md) for dependencies, assets, and `mcode-tools`. diff --git a/README_ZH.md b/README_ZH.md index 579f6b33..6113d319 100644 --- a/README_ZH.md +++ b/README_ZH.md @@ -279,6 +279,18 @@ Web UI 源自社区的 **mcode-webui** 插件,现已作为一等公民包迁 本仓库也承接 MiniMax Code 桌面版的问题反馈。公开源码范围为终端 TUI、Headless CLI 和 ACP,不包含桌面应用源码。提交 Issue 时请选择对应产品。桌面版问题请注明应用版本、操作系统,以及「设置 → 通用 → 上传日志」生成的日志上传 ID(如可用);CLI 问题请注明 `mcode --version`、运行入口与最小复现。报告中请移除凭据和私人项目内容。 +## 反馈与联系我们 + +| 渠道 | 适用场景 | +| --- | --- | +| [GitHub Issues](https://github.com/MiniMax-AI/minimax-code/issues/new/choose) | 公开报告 CLI 或桌面版的 Bug、提出功能建议与使用问题。 | +| [MiniMaxCode@minimax.io](mailto:MiniMaxCode@minimax.io) | 一般反馈与支持咨询。 | +| [security.mcode@minimax.io](mailto:security.mcode@minimax.io) | 私密报告安全漏洞。请通过此邮箱发送复现步骤和脱敏证据,详见[安全报告指南](SECURITY.md)。 | +| [Discord](https://minimax.io/discord) | 社区交流与反馈。 | +| [飞书反馈群二维码](https://cdn.hailuoai.com/hailuo-video-web/public_assets/minimax_code_feishu_group_url.png) | 中文社区反馈。使用飞书扫码,或在中文版桌面应用的用户菜单 → **联系我们 → 飞书** 中查看二维码。 | + +也可关注 [MiniMax 的 X 账号](https://x.com/MiniMaxAgent) 获取动态。请勿在公开 Issue 或社区聊天中发布漏洞细节、凭据和私人项目内容。 + ## 许可 第一方代码默认采用 [MIT](LICENSE);文件或子包已有独立声明时保留原许可。依赖、资源与 `mcode-tools` 的许可分别见 [第三方声明](THIRD_PARTY_NOTICES.md) 和 [许可状态](LICENSE-STATUS.md)。 diff --git a/SECURITY.md b/SECURITY.md index 06d84460..0b7ba676 100644 --- a/SECURITY.md +++ b/SECURITY.md @@ -2,9 +2,11 @@ This project is a source preview. Maintainers prioritize security issues on the default branch; no support period for older versions or response SLA has been committed. -Report vulnerabilities privately through **Security → Advisories → Report a vulnerability** on GitHub. If that entry is not enabled, open an issue without vulnerability details asking maintainers for a private channel. Share reproduction details only after that channel is available. Do not put credentials, exploit details, or real user data in public issues. +Report vulnerabilities privately by emailing [security.mcode@minimax.io](mailto:security.mcode@minimax.io). You can send reproduction details and redacted evidence directly to this address without opening a public issue first. Do not put credentials, exploit details, or real user data in public issues or community chats. -The release coordinator, @hetaoBackend, coordinates security triage; see [Maintainers](docs/maintainers.md). GitHub private vulnerability reporting is not currently enabled, and no public fallback security email is listed. Until a private channel is available, open an issue without vulnerability details as described above. No response SLA is currently promised. +If **Security → Advisories → Report a vulnerability** is available on GitHub, you can also use that private reporting channel. If it is unavailable, use the security email above. + +The release coordinator, @hetaoBackend, coordinates security triage; see [Maintainers](docs/maintainers.md). No response SLA is currently promised. Include the affected version, operating system and Node.js version, a minimal reproduction, expected and actual permission boundaries, and necessary redacted evidence. Use synthetic files and dedicated test accounts; do not test other people's accounts or infrastructure. diff --git a/docs/examples.md b/docs/examples.md index 83e4147d..44fe04b9 100644 --- a/docs/examples.md +++ b/docs/examples.md @@ -28,6 +28,8 @@ Use `/model` in the interactive TUI to select a model or choose **+ Add 3rd-part Preset IDs come from models.dev and do not select entries in the bundled inference registry. Onboarding saves the chosen URL under `custom_provider`; subsequent requests use that saved URL. +In **Custom provider**, enter the name, Base URL, protocol, and API key first. Then choose **Import models from /models** to fetch the list using that key. Search and select a model to test; a successful test saves all imported models and selects the chosen one. Importing alone does not save configuration. If discovery fails or returns no models, retry, press **Esc** to edit the key, or choose **Enter a model ID manually**. Saved connections also support **refresh models** in `/provider`. + Before adding a custom provider, set a key in your current shell rather than putting it in command arguments or source: ```bash diff --git a/docs/source-sync-0.5.5.md b/docs/source-sync-0.5.5.md new file mode 100644 index 00000000..fd99d1fa --- /dev/null +++ b/docs/source-sync-0.5.5.md @@ -0,0 +1,56 @@ +# Reviewed source update for 0.5.5 + +This update ports the public-compatible 0.5.5 behavior from source revision +`d28ae33c08a907f5e1ef17d70725eea0d0c667a0` onto the existing public distribution. +It preserves public fixes and distribution adaptations rather than replacing the +repository with a product build's source tree. + +## Included behavior + +- Managed foreground Bash defaults to a one-hour total timeout and caps larger + requests at one hour. The 60-second foreground yield preserves the running + process and its original deadline. Explicit background tasks use a one-hour + runtime watchdog by default; explicit command timeouts remain separate. +- Automatic context compaction reserves output space before reaching the context + limit. Checkpoint output scales with the model window; reasoning-only output + exhaustion advances to a smaller candidate. Provider errors get one bounded + logical retry. Automatic failures retain the original history; valid manual + checkpoints are committed even when a local next-request estimate rejects them. +- Update downloads reuse the TUI's proxy and loopback-bypass policy and load + their network dependency lazily. Existing public installation ownership and + registry selection remain intact. +- Memory-tool results have a 16 KiB model-facing limit with a head/tail preview + and guidance for reading more. This does not upload memory or change its storage. + +The root and TUI source versions are 0.5.5. This source update does not republish +the existing npm package or move an existing release tag. + +## Privacy and publication decisions + +| Surface | Decision | +| --- | --- | +| Usage, metrics and diagnostics | Keep the public independent opt-ins, disabled defaults, and `DO_NOT_TRACK` / `MCODE_DISABLE_TELEMETRY` overrides. | +| Evaluation capture and data contribution | Do not import automatic capture wiring, evaluation payload/transport expansions, or default-enabled contribution behavior. | +| Workspace collection and indexing | Keep snapshot collection, archive creation, background upload/retry and semantic-index activation excluded. | +| Feedback and automatic error reports | Keep the public reviewed-text/count-only feedback projection and allowlisted diagnostic schemas; no raw conversations, tool output or workspace files. | +| Compaction observations | Import only local content-free count/budget/outcome facts; preserve the existing public telemetry consent boundary. | +| Managed account, BYOK, plugins, connectors, search and user-requested deployments | Preserve supported public clients and behavior. No private endpoints, generated service contracts or new cloud authorization dependencies are introduced. | + +The reviewed update is selective. The older `release/extraction.json` +`sourceRevision` remains the base for future three-way comparisons: changing it +would incorrectly mark the remaining runtime ownership migrations, service +integrations and source-tree moves as synchronized. Those changes need their own +public dependency and privacy review. Private candidate reports remain outside +this repository and are not publication artifacts. + +## Validation boundary + +Imported tests exercise compaction failure/recovery, budget boundaries and real +Bash execution using synthetic data. Public privacy tests inspect the outgoing +telemetry/diagnostic data, including the final feedback archive. The repository's +full verifier additionally checks source export, types, build boundaries, TUI, +headless BYOK, ACP, permissions and platform-applicable sandbox behavior. + +Offline tests are not evidence of live-service ingestion, model quality or +Windows/Linux runtime acceptance. Actual check results belong in the pull +request's validation record. diff --git a/docs/source-sync.md b/docs/source-sync.md index 990548d2..ae2b6d99 100644 --- a/docs/source-sync.md +++ b/docs/source-sync.md @@ -24,3 +24,11 @@ Output includes `report.json`, `candidates/`, and `.private-review`. Candidates 5. Scan complete history and current source before creating the public PR. Include public changes, validation results, and capability descriptions, never private review reports. After accepted public changes are ported back, subsequent three-way comparisons should show them as synchronized or cleanly mergeable while retaining standalone adaptations. Synchronization is not a blind overwrite: conflicts, missing source, and new files require maintainer judgment. + +## Selective release updates + +A bounded release update can port reviewed behavior without adopting unrelated +runtime ownership migrations or private service integrations. Keep the existing +three-way baseline until the entire target revision has been reviewed; record +the selected revision, included behavior and excluded boundaries separately. +See the [0.5.5 review](source-sync-0.5.5.md) for the current selective update. diff --git a/docs/tui-capabilities.md b/docs/tui-capabilities.md index cc26b7aa..be2a3b6a 100644 --- a/docs/tui-capabilities.md +++ b/docs/tui-capabilities.md @@ -2,6 +2,65 @@ The current capability target is **TUI 0.4.12**; see [version and evidence baseline](open-source-status.md#version-and-evidence-baseline) for the separate workspace and embedded-tool versions. “Restored” below describes implementation and assembly, not acceptance of every account or online service. +## Output speed + +The activity line and completed-turn summary show provider output tokens divided by +model generation time. Timing starts at the first nonempty text, reasoning, or tool +token and ends when the model finishes, excluding first-token wait and tool execution. +Multiple responses in a turn use total tokens divided by total generation time. + +Speed appears after the first response with both provider usage and generation timing; +streaming text is not estimated. Confirmed speed stays visible during later requests +and tools, duplicate messages do not count twice, and a new turn resets the samples. +Messages without generation timing are excluded. Values below 10 tok/s use one decimal +place; higher values are rounded to integers. Batched provider events measure the +observed generation window, not server hardware throughput. + +## Terminal titles and notifications + +Terminal titles show the current state, session name and MCode, for example +`Needs approval | Fix login | MCode`. Renaming or switching a session updates the +title. Unnamed sessions use the project name and a short session ID. Titles are +cleared when MCode exits or suspends and reapplied when it resumes. + +Configure these presentation settings in the MCode data directory's `config.yaml`: + +```yaml +tui: + terminalTitle: [status, session-name, app-name] + notifications: + when: unfocused + method: auto + events: [turn-complete, turn-failed, permission-required, question-required] +``` + +Title items can be ordered or omitted; `project-name` is also available. Set +`terminalTitle` to `null` or `[]` to disable title updates. Unknown items are ignored. +Notification `when` accepts `unfocused`, `always` or `never`; `method` accepts +`auto`, `osc9`, `osc777` or `bel`. Omitting `events` enables all four events; `[]` +disables them. Apply configuration changes by restarting MCode. + +Notifications identify the session and suppress duplicates. Completion waits for +the session's queue to finish; failed turns and requests for input can notify +independently. Known foreground focus suppresses notifications by default. When +focus is unknown, delivery is best-effort; cmux manages its own surface focus. +Automatic delivery uses the detected terminal's notification protocol or falls +back to a bell. The existing Windows toast bridge is restricted to local Windows +or WSL interop. Terminal settings and OS notification permissions still apply. + +VS Code normally displays a process name in its terminal tabs. To display MCode's +session titles, use this VS Code setting: + +```json +"terminal.integrated.tabs.title": "${sequence}" +``` + +A manually assigned tab title overrides automatic titles. VS Code's bell is a +terminal-tab indicator, not a guarantee of a desktop notification. See the +[VS Code terminal appearance documentation](https://code.visualstudio.com/docs/terminal/appearance#_tab-text). +Inside tmux, OSC notifications require passthrough and support from the outer +terminal; use `method: bel` for a bell fallback. + The evidence column summarizes the historical TUI 0.3.11 restoration record from 2026-09-11. It does not claim fresh TUI 0.4.12 live-service acceptance. Use [current verification status](verification.md#current-source-verification-status) for checks run against the updated source and explicit NOT RUN boundaries. ## Capability matrix @@ -25,6 +84,33 @@ Status legend: ✅ supported · ⚠ partial · ❌ unsupported · 🚧 requires | Files, shell, subagents, sessions, headless, ACP | Actual runtime retained | BYOK, file reads, session resume, ACP, sandbox, and status protocol tests | | Built-in skills, MCP, plugin tools | Original TUI assets and activation conditions retained | Asset build, plugin, and MCP tests; no claim that every skill has passed a real task | +## Local Bash execution + +When the current turn includes native `task_output`, foreground Bash waits up to +60 seconds before returning the same command's background task ID. Its total +command timeout defaults to 600 seconds and is capped at 600 seconds; a shorter +requested timeout applies. Backgrounding and output reads preserve the original +deadline. Without native `task_output`, Bash stays in the foreground with a +120-second default and a 300-second cap, and its schema omits `run_in_background`. +Explicit background commands use the requested timeout; when omitted, the +existing 30-minute runtime watchdog applies. + +Only exit code zero is success. Results retain available exit, signal, timeout, +cancellation, and partial-output facts. Large output keeps its original beginning +and end within a 24 KiB first-response text budget, with a full-log reference when +persistence succeeds. `task_output` reads use byte offsets; a successful read can +report a failed command. Stop failures and incomplete logs are reported separately. +An optional `description` supplies the TUI summary while execution and permission +checks continue to use the original command. + +## Skill directory links + +Workspace `.agents/skills`, `.claude/skills`, and `.minimax/skills` support +directory symlinks, both for the entire skill root and for individual skill +directories. Targets may live outside the workspace. Existing external-source +enable settings and duplicate-name priority still apply. Linked directories are +watched for `SKILL.md` creation and edits; broken links are skipped. `SKILL.md` +itself must remain a regular file. ## Slash-command parity (TUI ↔ Web UI) The webui talks to the same engine, but it only handles a **strict subset @@ -180,10 +266,16 @@ and initial prompt are not applied. In regular mode, independent feature panels occupy the complete visible terminal area, including short Rewind previews and scope pickers. Closing a panel restores -the current conversation. Closing a full-viewport interaction rebuilds the chat -screen so its temporary rows do not leave a large blank area above the conversation. -When running content shrinks entirely within the current screen, the renderer -keeps native scrollback and the Composer position stable. +the current conversation. Closing, replacing or shrinking a transient region +restores the exposed chat rows. This includes inline selectors such as `/theme`, +completion menus, multi-line drafts, image previews, queued messages, task and +Goal summaries, welcome notices and status rows. Short documents refresh in place; +history is reconstructed only when the smaller layout needs to bring scrolled +rows back into view. This rule follows the rendered layout, including asynchronous +updates, rather than requiring each close handler to request a special redraw. +When background running content shrinks entirely within the current screen and +the transient layout stays unchanged, the renderer keeps native scrollback and +the Composer position stable. Freed rows temporarily remain blank at the top of the active screen and subsequent output reuses them. This avoids resetting the host's scroll position when a turn finishes. Redundant resize notifications with unchanged dimensions do not rebuild @@ -221,3 +313,26 @@ existing diagnostic-counts projection; raw error text, stacks and session IDs are not added to the uploaded ZIP. Offline tests cover persisted tool histories, archives, concurrent parent output, side-session cleanup and local diagnostics; this does not establish native-terminal or live-model acceptance. + +## Select a plugin for a message + +Type `@` in the Composer to search files and installed, enabled plugins. Plugin +candidates show their source so packages with the same display name can be +selected independently. Choose a plugin with Tab or Enter, then describe the task. +The Composer shows `@Name` and retains the plugin identity through editing, undo, +prompt history, queued-message recovery, saved drafts, and `/edit` after a +message is sent. Ctrl+C clearing/restoration and external-editor edits retain +unchanged plugin bindings. If external edits make duplicate labels ambiguous, +reselect those plugins in the Composer. Displayed messages remain readable; Runtime retains +the original input separately when needed to recover the plugin identity for editing. + +Selection applies to that message. Runtime checks the plugin's effective Skills, +MCP tools, and App tools again for the turn and asks the Agent to prefer relevant +capabilities. Selecting a plugin does not install or enable it. An unavailable +selection is reported to the Agent rather than redirected to a same-named package. + +Exec and ACP text prompts can use the durable linked form, for example +`[@Notes](plugin://notes%40local) summarize these files`. The ID is the package +name plus its `local` or `official` source; display labels do not determine the +selection. Legacy whitespace-delimited `@package-name` text remains supported +when it identifies exactly one effective plugin. diff --git a/package.json b/package.json index 15a237f8..c10d7b7e 100644 --- a/package.json +++ b/package.json @@ -1,6 +1,6 @@ { "name": "minimax-code", - "version": "0.5.2", + "version": "0.5.7", "private": true, "type": "module", "description": "Standalone MiniMax Code TUI with managed accounts, BYOK models, cloud tools, plugins and ACP.", diff --git a/packages/agent-core/src/event-bridge/bridge.ts b/packages/agent-core/src/event-bridge/bridge.ts index da320a85..1d4e4b43 100644 --- a/packages/agent-core/src/event-bridge/bridge.ts +++ b/packages/agent-core/src/event-bridge/bridge.ts @@ -164,6 +164,8 @@ export class EventBridge { private awaitingMessageId: boolean = false; /** Wall-clock ms when the first thinking_delta of the current msg arrived. */ private thinkingStartMs: number | undefined; + /** First nonempty text/thinking/tool token; excludes first-token wait from throughput. */ + private firstTokenMs: number | undefined; /** Wall-clock ms when thinking ended (first non-thinking delta or message_end). */ private thinkingEndMs: number | undefined; /** Whether the current message has already produced a non-thinking delta. */ @@ -361,6 +363,7 @@ export class EventBridge { this.activeAssistantMessageId = undefined; this.chunkIndex = 0; this.thinkingStartMs = undefined; + this.firstTokenMs = undefined; this.thinkingEndMs = undefined; this.thinkingClosed = false; this.awaitingMessageId = true; @@ -398,6 +401,14 @@ export class EventBridge { ], }; } + if ( + this.firstTokenMs === undefined && + (('delta' in update && update.delta !== '') || + ((update.kind === 'toolcall_start' || update.kind === 'toolcall_delta') && + update.toolName !== undefined)) + ) { + this.firstTokenMs = this.now(); + } if ( update.kind === 'toolcall_start' || update.kind === 'toolcall_delta' || @@ -587,11 +598,18 @@ export class EventBridge { this.ctx.includeDetailedUsage === true ? this.ctx.requestDurationMs?.(event.message) : undefined; + const decodeDurationMs = + this.ctx.includeDetailedUsage === true && this.firstTokenMs !== undefined + ? Math.max(0, this.now() - this.firstTokenMs) + : undefined; const usage: TokenUsage = { ...(extractAssistantUsage(event.message, this.contextWindow) ?? { total_tokens: 0, context_window: this.contextWindow, }), + ...(decodeDurationMs !== undefined && Number.isFinite(decodeDurationMs) + ? { decode_duration_ms: decodeDurationMs } + : {}), ...(typeof requestDurationMs === 'number' && Number.isFinite(requestDurationMs) && requestDurationMs > 0 diff --git a/packages/agent-core/src/protocol/agent-message.ts b/packages/agent-core/src/protocol/agent-message.ts index daacf836..f33222d6 100644 --- a/packages/agent-core/src/protocol/agent-message.ts +++ b/packages/agent-core/src/protocol/agent-message.ts @@ -127,12 +127,10 @@ export interface TokenUsage { input_tokens?: number; /** Provider-reported generated tokens for this physical request. */ output_tokens?: number; - /** - * Physical provider request duration measured at the stream wrapper. - * Consumers use the elapsed duration directly instead of reconstructing - * throughput from downstream chunk timestamps, which may be buffered. - */ + /** Full physical provider request duration, including first-token wait. */ request_duration_ms?: number; + /** First nonempty text/thinking/tool token to message_end, excluding tool execution. */ + decode_duration_ms?: number; /** * Cached prompt tokens reused across turns (provider-specific). Optional so * legacy producers/consumers stay compatible; Pi providers surface this via diff --git a/packages/agent-core/test/unit/event-bridge/decode-timing.test.ts b/packages/agent-core/test/unit/event-bridge/decode-timing.test.ts new file mode 100644 index 00000000..19198d52 --- /dev/null +++ b/packages/agent-core/test/unit/event-bridge/decode-timing.test.ts @@ -0,0 +1,163 @@ +import { describe, expect, it } from "vitest"; +import type { AgentEvent } from "@earendil-works/pi-agent-core"; +import type { AssistantMessage } from "@earendil-works/pi-ai"; +import { EventBridge } from "../../../src/event-bridge/bridge.js"; +import type { RuntimeEvent } from "@mavis/protocol"; + +function assistant( + content: AssistantMessage["content"] = [], +): AssistantMessage { + return { + role: "assistant", + content, + api: "openai-completions", + provider: "openai", + model: "test-model", + stopReason: content.some((part) => part.type === "toolCall") + ? "toolUse" + : "stop", + timestamp: 0, + usage: { + input: 10, + output: 100, + cacheRead: 0, + cacheWrite: 0, + totalTokens: 110, + cost: { input: 0, output: 0, cacheRead: 0, cacheWrite: 0, total: 0 }, + }, + }; +} + +function messageUsage(events: RuntimeEvent[]) { + return events + .map((event) => event.payload.stream_resp) + .filter((raw): raw is string => typeof raw === "string") + .map((raw) => JSON.parse(raw).agent_message?.usage) + .find((usage) => usage !== undefined); +} + +function fixture(detailed = true) { + let now = 0; + let sequence = 0; + const bridge = new EventBridge({ + sessionId: "session-test", + turnId: "turn-test", + eventIdGenerator: (kind) => `${kind}-${++sequence}`, + runtimeSeqGenerator: () => ++sequence, + nowMs: () => now, + includeDetailedUsage: detailed, + requestDurationMs: () => 7_000, + }); + return { + bridge, + setTime: (value: number) => { + now = value; + }, + }; +} + +function delta( + type: "text_delta" | "thinking_delta" | "toolcall_delta", + value: string, +): AgentEvent { + return { + type: "message_update", + message: assistant(), + assistantMessageEvent: { + type, + contentIndex: 0, + delta: value, + partial: assistant(), + }, + }; +} + +describe("EventBridge generation timing", () => { + it.each(["text", "thinking", "tool"] as const)( + "starts at the first nonempty %s token and excludes delayed tool completion", + async (kind) => { + const { bridge, setTime } = fixture(); + await bridge.processEvent({ + type: "message_start", + message: assistant(), + }); + bridge.setActiveAssistantMessageId("message-1"); + setTime(100); + for (const type of [ + "text_delta", + "thinking_delta", + "toolcall_delta", + ] as const) { + await bridge.processEvent(delta(type, "")); + } + const tool = { + type: "toolCall" as const, + id: "tool-1", + name: "bash", + arguments: {}, + }; + setTime(5_000); + await bridge.processEvent( + kind === "tool" + ? { + type: "message_update", + message: assistant(), + assistantMessageEvent: { + type: "toolcall_start", + contentIndex: 0, + partial: assistant([tool]), + }, + } + : delta(kind === "text" ? "text_delta" : "thinking_delta", "hello"), + ); + setTime(7_000); + const finished = await bridge.processEvent({ + type: "message_end", + message: assistant([tool]), + }); + expect(messageUsage(finished.events)).toBeUndefined(); + setTime(20_000); + const completed = await bridge.processEvent({ + type: "tool_execution_end", + toolCallId: "tool-1", + toolName: "bash", + result: { content: [{ type: "text", text: "done" }] }, + isError: false, + }); + expect(messageUsage(completed.events)).toMatchObject({ + decode_duration_ms: 2_000, + request_duration_ms: 7_000, + }); + await bridge.processEvent({ + type: "message_start", + message: assistant(), + }); + bridge.setActiveAssistantMessageId("message-2"); + const buffered = await bridge.processEvent({ + type: "message_end", + message: assistant([{ type: "text", text: "buffered" }]), + }); + expect(messageUsage(buffered.events)?.decode_duration_ms).toBeUndefined(); + }, + ); + + it.each([true, false])( + "preserves zero-duration samples only with detailed usage enabled: %s", + async (detailed) => { + const { bridge } = fixture(detailed); + await bridge.processEvent({ + type: "message_start", + message: assistant(), + }); + bridge.setActiveAssistantMessageId("instant"); + await bridge.processEvent(delta("text_delta", "hello")); + const result = await bridge.processEvent({ + type: "message_end", + message: assistant([{ type: "text", text: "hello" }]), + }); + expect(messageUsage(result.events)?.decode_duration_ms).toBe( + detailed ? 0 : undefined, + ); + }, + ); +}); diff --git a/packages/agent-modules/context-manager/src/provider-budget.ts b/packages/agent-modules/context-manager/src/provider-budget.ts index 234185b3..c8bce00c 100644 --- a/packages/agent-modules/context-manager/src/provider-budget.ts +++ b/packages/agent-modules/context-manager/src/provider-budget.ts @@ -1,6 +1,8 @@ import { DEFAULT_CONTEXT_MANAGER_SETTINGS } from './settings.js'; const PROVIDER_INPUT_RATIO = 0.95; +/** Integer percent keeps the trigger exact (0.95 is not exact in binary floating point). */ +const AUTOMATIC_TRIGGER_PERCENT = 95; export function resolveDynamicMaxTokens(input: { readonly contextWindow: number; @@ -20,9 +22,13 @@ export function resolveCompactionTokenBudget(input: { }): { readonly providerInputLimit: number; readonly automaticTriggerAt: number } { const configuredOutput = positive(input.configuredMaxOutputTokens); const { reserveTokens, safetyMarginTokens } = DEFAULT_CONTEXT_MANAGER_SETTINGS; - const fullOutputInputBudget = input.contextWindow - configuredOutput - safetyMarginTokens; - const effectiveOutput = - fullOutputInputBudget >= reserveTokens ? configuredOutput : outputFloor(configuredOutput); + // Admission mirrors resolveDynamicMaxTokens: a real Provider request shrinks + // its output budget down to the dynamic floor when input is large, so the + // hard input limit reserves the floor instead of the full configured output. + // Reserving the full output here rejects requests the Provider would accept + // (e.g. 200K-window models with 128K configured output collapsed the input + // budget to ~70K and wedged sessions whose fixed prompt exceeded it). + const effectiveOutput = outputFloor(configuredOutput); const providerInputLimit = Math.max( 1, Math.min( @@ -31,13 +37,18 @@ export function resolveCompactionTokenBudget(input: { input.contextWindow - effectiveOutput - safetyMarginTokens, ), ); - const proactiveReserve = Math.min(reserveTokens * 2, Math.floor(input.contextWindow / 4)); + // The automatic trigger reserves the configured output, capped at a quarter + // of the window, and compacts at 95% of the remainder. Large windows then + // start compaction while the main and checkpoint requests still have their + // configured output room; small windows are not collapsed by a large output + // limit. providerInputLimit stays the upper bound. + const outputReserve = Math.min(configuredOutput, Math.floor(input.contextWindow / 4)); + const outputReservedTrigger = Math.floor( + ((input.contextWindow - outputReserve) * AUTOMATIC_TRIGGER_PERCENT) / 100, + ); return { providerInputLimit, - automaticTriggerAt: Math.min( - providerInputLimit, - Math.max(1, input.contextWindow - proactiveReserve), - ), + automaticTriggerAt: Math.min(providerInputLimit, Math.max(1, outputReservedTrigger)), }; } diff --git a/packages/agent-modules/context-manager/test/provider-budget.test.ts b/packages/agent-modules/context-manager/test/provider-budget.test.ts new file mode 100644 index 00000000..fa932acd --- /dev/null +++ b/packages/agent-modules/context-manager/test/provider-budget.test.ts @@ -0,0 +1,105 @@ +import { describe, expect, it } from 'vitest'; + +import { + resolveCompactionTokenBudget, + resolveDynamicMaxTokens, +} from '../src/provider-budget.js'; + +describe('resolveDynamicMaxTokens', () => { + it('shrinks a Spark output budget from the complete Provider input footprint', () => { + expect( + resolveDynamicMaxTokens({ + contextWindow: 128_000, + configuredMaxTokens: 128_000, + estimatedContextTokens: 100_000, + }), + ).toBe(25_952); + }); +}); + +describe('resolveCompactionTokenBudget', () => { + it.each([ + [200_000, 181_568], + [400_000, 380_000], + [512_000, 486_400], + [1_000_000, 950_000], + ] as const)('caps a %i-token context at the conservative Provider limit', (contextWindow, expected) => { + expect( + resolveCompactionTokenBudget({ + contextWindow, + configuredMaxOutputTokens: 16_384, + }).providerInputLimit, + ).toBe(expected); + }); + + it.each([ + ['1M context', 1_000_000, 128_000, 950_000, 828_400], + ['512K context', 512_000, 128_000, 486_400, 364_800], + ['roomy 400K context', 400_000, 128_000, 380_000, 285_000], + ['gpt-5.4 output floor', 272_000, 128_000, 253_568, 193_800], + ['MiniMax-M2.7 window', 200_000, 128_000, 181_568, 142_500], + ['256K relay window', 256_000, 128_000, 237_568, 182_400], + ['Spark dynamic fallback', 128_000, 128_000, 109_568, 91_200], + ['1M context with a small output limit', 1_000_000, 32_000, 950_000, 919_600], + ['64K context capped by the hard limit', 64_000, 128_000, 45_568, 45_568], + ['caller-limited output', 65_536, 8_192, 49_152, 49_152], + ] as const)( + 'resolves the shared input limit and automatic trigger for %s', + (_case, contextWindow, configuredMaxOutputTokens, providerInputLimit, automaticTriggerAt) => { + expect( + resolveCompactionTokenBudget({ contextWindow, configuredMaxOutputTokens }), + ).toEqual({ providerInputLimit, automaticTriggerAt }); + }, + ); + + it.each([ + ['65K context', 65_536, 46_694], + ['Spark', 128_000, 106_035], + ['large context', 272_000, 242_835], + ] as const)( + 'reserves a small configured output below the window-quarter cap for %s', + (_case, contextWindow, automaticTriggerAt) => { + expect( + resolveCompactionTokenBudget({ + contextWindow, + configuredMaxOutputTokens: 16_384, + }).automaticTriggerAt, + ).toBe(automaticTriggerAt); + }, + ); + + it('leaves the configured output room when a large window reaches the trigger', () => { + const contextWindow = 1_000_000; + const configuredMaxTokens = 128_000; + const { automaticTriggerAt } = resolveCompactionTokenBudget({ + contextWindow, + configuredMaxOutputTokens: configuredMaxTokens, + }); + + expect( + resolveDynamicMaxTokens({ + contextWindow, + configuredMaxTokens, + estimatedContextTokens: automaticTriggerAt, + }), + ).toBe(configuredMaxTokens); + }); + + it.each([Number.NaN, 0])('ignores an invalid configured output limit: %s', (configuredMaxOutputTokens) => { + expect( + resolveCompactionTokenBudget({ + contextWindow: 65_536, + configuredMaxOutputTokens, + }).providerInputLimit, + ).toBe(49_152); + }); + + it('clamps an invalidly small context to one token', () => { + expect( + resolveCompactionTokenBudget({ + contextWindow: 1, + configuredMaxOutputTokens: 128_000, + }), + ).toEqual({ providerInputLimit: 1, automaticTriggerAt: 1 }); + }); +}); diff --git a/packages/agent-modules/goal/src/tool-impls.ts b/packages/agent-modules/goal/src/tool-impls.ts index ea669580..2568dd83 100644 --- a/packages/agent-modules/goal/src/tool-impls.ts +++ b/packages/agent-modules/goal/src/tool-impls.ts @@ -246,6 +246,13 @@ export class UpdateGoalTool implements ToolImpl< objectiveDigest: digestThreadGoalObjective(existing.objective), ...(summary ? { summary } : {}), }); + if (collection === 'not_a_goal_turn') { + return err( + UpdateGoalToolDef.name, + 'cannot update goal because this turn is not bound to the goal; continue handling the user request in this turn without updating goal status. This tool cannot resume a goal', + { reason: collection, currentGoal: serializeGoal(existing) }, + ); + } if (collection === 'stale') { return { ...err( diff --git a/packages/agent-modules/goal/src/types.ts b/packages/agent-modules/goal/src/types.ts index 2470a4af..fce100d3 100644 --- a/packages/agent-modules/goal/src/types.ts +++ b/packages/agent-modules/goal/src/types.ts @@ -80,7 +80,12 @@ export interface GoalTurnSignal { readonly summary?: string; } -export type ThreadGoalSignalCollectionResult = 'accepted' | 'stale' | 'no_goal' | 'paused'; +export type ThreadGoalSignalCollectionResult = + | 'accepted' + | 'not_a_goal_turn' + | 'stale' + | 'no_goal' + | 'paused'; export const THREAD_GOAL_STATUS_REASONS = [ 'complete(worker_proposal)', diff --git a/packages/agent-modules/goal/test/unit/thread-goal/tool-impls.test.ts b/packages/agent-modules/goal/test/unit/thread-goal/tool-impls.test.ts index a20f14af..50895e9f 100644 --- a/packages/agent-modules/goal/test/unit/thread-goal/tool-impls.test.ts +++ b/packages/agent-modules/goal/test/unit/thread-goal/tool-impls.test.ts @@ -563,14 +563,16 @@ describe("thread-goal tool impls", () => { const before = await store.getBySession(ctx.sessionId); const patch = vi.spyOn(store, "patch"); - // An unbound proposal: the Turn carries no Goal binding, so the host - // rejects it. Both the stale and the unbound refusals share `err()`. const stale = await new UpdateGoalTool( store, signalCollector("stale"), ).execute(ctx, { status: "complete", }); + const unbound = await new UpdateGoalTool( + store, + signalCollector("not_a_goal_turn"), + ).execute(ctx, { status: "complete" }); const missingMode = await new UpdateGoalTool( store, signalCollector(), @@ -580,6 +582,8 @@ describe("thread-goal tool impls", () => { }); expect(stale.isError).toBe(true); + expect(unbound.isError).toBe(true); + expect(unbound.terminate).toBeUndefined(); expect(missingMode.isError).toBe(true); expect(patch).not.toHaveBeenCalled(); expect(await store.getBySession(ctx.sessionId)).toEqual(before); diff --git a/packages/agent-modules/skills/src/registry.ts b/packages/agent-modules/skills/src/registry.ts index 50c469b4..97277ce6 100644 --- a/packages/agent-modules/skills/src/registry.ts +++ b/packages/agent-modules/skills/src/registry.ts @@ -226,7 +226,7 @@ export class SkillRegistry { }; for (const root of this.roots) { collectPath(root.rootPath, root); - for (const skillDir of listWatchableSkillDirs(root.rootPath)) { + for (const skillDir of listWatchableSkillDirs(root)) { collectPath(skillDir, root); } } @@ -472,7 +472,10 @@ function existingWatchTarget(targetPath: string): SkillWatchTarget | undefined { } function allowsDirectorySymlinkOutsideRoot(root: SkillSourceRoot): boolean { - return ['agent', 'global', 'user', 'builtin'].includes(root.kind); + return ( + root.allowDirectorySymlinksOutsideRoot === true || + ['agent', 'global', 'user', 'builtin'].includes(root.kind) + ); } async function listSkillFiles( @@ -811,10 +814,20 @@ function sameSkillFileStat(expected: SkillFileStat, actual: SkillFileStat): bool ); } -function listWatchableSkillDirs(rootPath: string): string[] { +function listWatchableSkillDirs(root: SkillSourceRoot): string[] { + const { rootPath } = root; try { return readdirSync(rootPath, { withFileTypes: true }) - .filter((child) => child.isDirectory()) + .filter((child) => { + if (child.isDirectory()) return true; + if (!child.isSymbolicLink() || !allowsDirectorySymlinkOutsideRoot(root)) return false; + try { + // Watch linked directories even before they contain a SKILL.md. + return statSync(path.join(rootPath, child.name)).isDirectory(); + } catch { + return false; + } + }) .map((child) => path.join(rootPath, child.name)) .sort(); } catch { diff --git a/packages/agent-modules/skills/src/types.ts b/packages/agent-modules/skills/src/types.ts index 150617e0..841c158e 100644 --- a/packages/agent-modules/skills/src/types.ts +++ b/packages/agent-modules/skills/src/types.ts @@ -7,6 +7,8 @@ export interface SkillSourceRoot { rootPath: string; priority?: number; external?: boolean; + /** Follow directory links outside the root for configured compatibility sources. */ + allowDirectorySymlinksOutsideRoot?: boolean; } export type SkillDiagnosticLevel = 'warning' | 'error'; diff --git a/packages/agent-modules/system-reminder/package.json b/packages/agent-modules/system-reminder/package.json index 1ee0840c..0ddefcd9 100644 --- a/packages/agent-modules/system-reminder/package.json +++ b/packages/agent-modules/system-reminder/package.json @@ -14,7 +14,8 @@ "@mavis/config": "workspace:^", "@types/node": "^20", "typescript": "^5", - "vitest": "^2" + "vitest": "^2", + "@mavis/shared": "workspace:^" }, "private": true, "license": "MIT" diff --git a/packages/agent-modules/system-reminder/src/plugin-reference.ts b/packages/agent-modules/system-reminder/src/plugin-reference.ts index 5878a928..2ad38865 100644 --- a/packages/agent-modules/system-reminder/src/plugin-reference.ts +++ b/packages/agent-modules/system-reminder/src/plugin-reference.ts @@ -1,12 +1,15 @@ +import { parsePluginMentions } from '@mavis/shared/plugin-mention'; + export interface EffectivePluginToolGroup { readonly source: string; readonly tools: readonly string[]; - /** Deferred App tools must be discovered before they can be invoked. */ + /** Deferred tools must be discovered before they can be invoked. */ readonly access?: 'direct' | 'tool_search'; } export interface EffectivePluginCapabilityInventory { readonly name: string; + readonly pluginId?: string; readonly appTools: readonly EffectivePluginToolGroup[]; readonly mcpTools: readonly EffectivePluginToolGroup[]; readonly skills: readonly string[]; @@ -33,7 +36,7 @@ export function detectPluginReferencesForMessages { + if (plugins.length === 0 && unavailablePluginIds.length === 0) return undefined; + const blocks = plugins.slice(0, 8).map((plugin) => { const name = inlineCode(plugin.name); - return [ + const detail = [ ``, `Referenced as \`@${name}\`.`, '', @@ -66,20 +70,46 @@ export function buildPluginReferenceReminder( '', '', ].join('\n'); + return detail.length <= 4096 + ? detail + : `\nCapability details omitted to fit the context limit. Use this Plugin's tool provenance and Skill namespace in the available catalogs.\n`; }); + const bounded: string[] = []; + let remaining = 12_288; + let omitted = plugins.length > 8 || unavailablePluginIds.length > 8; + for (const block of [ + ...blocks, + ...unavailablePluginIds + .slice(0, 8) + .map( + (id) => + `Selected Plugin \`${inlineCode(id)}\` is unavailable for this turn. Tell the user it could not be used; do not substitute a same-named Plugin or enable/install it automatically.`, + ), + ]) { + if (block.length + 2 > remaining) { + omitted = true; + continue; + } + bounded.push(block); + remaining -= block.length + 2; + } + return [ '', 'The user explicitly selected the following Plugin capabilities for this request.', 'Prefer them when relevant; other tools remain available if needed.', '', - blocks.join('\n\n'), + bounded.join('\n\n'), + ...(omitted ? ['Additional selected capabilities omitted to fit the context limit.'] : []), '', - 'Only the capabilities listed above are effective for this turn.', + 'The listed capabilities are drawn from the effective inventory for this turn; lists may be truncated.', 'Do not invent or claim unavailable Plugin capabilities.', - ...(plugins.some((plugin) => plugin.appTools.some((group) => group.access === 'tool_search')) + ...(plugins.some((plugin) => + [...plugin.appTools, ...plugin.mcpTools].some((group) => group.access === 'tool_search'), + ) ? [ - 'For App tools marked `via tool_search + mcp_invoke`, discover the exact tool with `tool_search` before calling it through `mcp_invoke`.', + 'For tools marked `via tool_search + mcp_invoke`, discover the exact tool with `tool_search` before calling it through `mcp_invoke`.', ] : []), 'Before following a listed Skill, call the `skill` tool with its exact name.', @@ -91,6 +121,7 @@ export function buildPluginReferenceReminder( function formatAppToolGroups(groups: readonly EffectivePluginToolGroup[]): string { if (groups.length === 0) return 'none'; return [...groups] + .slice(0, 16) .sort((left, right) => { const sourceOrder = normalizedName(left.source).localeCompare(normalizedName(right.source)); if (sourceOrder !== 0) return sourceOrder; @@ -108,11 +139,24 @@ function detectReferencesInText( text: string, plugins: readonly T[], ): T[] { - const normalizedText = text.normalize('NFKC'); + const linked = parsePluginMentions(text); + let plainText = text; + for (const mention of [...linked].reverse()) + plainText = `${plainText.slice(0, mention.start)}${' '.repeat(mention.end - mention.start)}${plainText.slice(mention.end)}`; + const normalizedText = plainText.normalize('NFKC'); const matches: Array<{ index: number; plugin: T }> = []; for (const plugin of plugins) { + const explicit = linked.find((mention) => mention.pluginId === plugin.pluginId); + if (explicit) { + matches.push({ index: explicit.start, plugin }); + continue; + } const name = normalizedName(plugin.name); - if (!name) continue; + if ( + !name || + plugins.filter((candidate) => normalizedName(candidate.name) === name).length !== 1 + ) + continue; const pattern = new RegExp(`(^|\\s)@${escapeRegExp(name)}(?=\\s|$)`, 'giu'); const match = pattern.exec(normalizedText); if (!match) continue; @@ -128,16 +172,19 @@ function formatToolGroups( ): string { if (groups.length === 0) return 'none'; return [...groups] + .slice(0, 16) .sort((left, right) => normalizedName(left.source).localeCompare(normalizedName(right.source))) .map((group) => { const tools = formatToolNames(group.tools); - return `- ${label} \`${inlineCode(group.source)}\`: ${tools || 'none'}`; + const access = group.access === 'tool_search' ? ' via `tool_search` + `mcp_invoke`' : ''; + return `- ${label} \`${inlineCode(group.source)}\`${access}: ${tools || 'none'}`; }) .join('\n'); } function formatToolNames(tools: readonly string[]): string { return [...new Set(tools)] + .slice(0, 32) .sort((left, right) => normalizedName(left).localeCompare(normalizedName(right))) .map((tool) => `\`${inlineCode(tool)}\``) .join(', '); @@ -146,6 +193,7 @@ function formatToolNames(tools: readonly string[]): string { function formatSkills(skills: readonly string[]): string { if (skills.length === 0) return 'none'; return [...new Set(skills)] + .slice(0, 32) .sort((left, right) => normalizedName(left).localeCompare(normalizedName(right))) .map((skill) => `- \`${inlineCode(skill)}\``) .join('\n'); diff --git a/packages/agent-tools/src/desktop/builtin-defs.ts b/packages/agent-tools/src/desktop/builtin-defs.ts index ebc13739..b50749a5 100644 --- a/packages/agent-tools/src/desktop/builtin-defs.ts +++ b/packages/agent-tools/src/desktop/builtin-defs.ts @@ -4,6 +4,7 @@ import type { ToolDefinition } from '@mavis/agent-core/tools'; import { createMavisOperationClassifier } from '../shared/mavis-operation-classifier.js'; import { LOCAL_MAVIS_COMMANDS } from './local-mavis-commands.js'; +import { BashDescriptionSchema } from './local-bash-input.js'; import { prepareAskUserArguments } from './prepare-ask-user-arguments.js'; import { @@ -102,54 +103,39 @@ export const LocalEditToolDef = { } as const satisfies ToolDefinition; export type LocalEditToolInput = Static; -const _isWindows = process.platform === 'win32'; -const MAX_BACKGROUND_BASH_TIMEOUT_SECONDS = 2_147_483; - export const LocalBashToolDef = { name: 'bash', executionMode: 'sequential', description: [ - 'Executes a command in a fresh local shell and returns its output.', + 'Executes a shell command and returns its output.', '', - '- Each call starts in the session workspace. `cd` and shell state do not persist between calls; use absolute paths or change directories within the same command.', - '- Use dedicated `read`, `write`, `edit`, `grep`, and `glob` tools for file operations. Do not use shell commands for file reading, searching, or modification unless the user explicitly requests it or you have verified that the dedicated tools cannot perform the required operation. Use `bash` for processes, git, package managers, builds, tests, and necessary pipelines.', - '- The shell is non-interactive: no TTY or stdin prompts. Run non-interactive commands yourself; check `--help` for suitable flags before asking the user to run one. Leave physical authorization (OAuth consent, MFA, hardware keys) to the user.', - '- Use `run_in_background` for long commands. Do not increase timeouts to mask hung commands. A returned task id refers to the original process; do not rerun it. Use `task_query` or `task_output` to inspect progress and `task_stop` to cancel.', + '- Each call starts in the session workspace. Changes to the working directory and shell state (variables and functions) do not persist between calls. If a command requires a different working directory, change directories within the same call.', + '- Use dedicated `read`, `write`, `edit`, `grep`, and `glob` tools for file operations. Do not use shell commands for file reading, searching, or modification unless the user explicitly requests it or you have verified that the dedicated tools cannot perform the required operation. Use `bash` for processes, git, package managers, builds, tests, and pipelines.', + '- The shell is non-interactive: no TTY or stdin prompts. For commands that require interaction, check `--help` for non-interactive options before asking the user to run them. Leave physical authorization (OAuth consent, MFA, hardware keys) to the user.', '- For file or directory deletion, use one top-level `rm -- ...`; the local runtime routes it through recoverable deletion.', '- Do not bypass recoverable deletion with absolute paths to deletion commands or inline scripts. If it fails, report the failure instead of falling back to permanent deletion. Permission checks still apply.', - ...(_isWindows - ? [ - '', - 'Windows shell:', - '- PowerShell is preferred; Bash is a fallback when PowerShell is unavailable. Match the selected shell syntax.', - '- For PowerShell, use `$env:VAR`; do not use `export`, `/dev/null`, heredocs, `sed -i`, or `&&`. Do not wrap commands in `powershell -Command`.', - "- Start multi-statement PowerShell scripts with `$ErrorActionPreference = 'Stop'`. Use single quotes for literals and regex patterns. Do NOT use Bash-style backslash escapes.", - '- For native program failures in PowerShell, check `$LASTEXITCODE` or enable `$PSNativeCommandUseErrorActionPreference` when available.', - '- Avoid `Get-Content | ... | Set-Content` file-editing pipelines. Prefer file tools or a script with explicit UTF-8 encoding for batch changes. If Get-Content/Set-Content are needed, specify `-Encoding UTF8`; PowerShell 5.1 writes a BOM.', - '- If a CLI is missing, check `.cmd` or `.ps1` wrappers. If `bash` opens WSL or produces garbled output, switch to direct `node`/`python` execution or a verified Git Bash path. After two failures with the same approach, change strategy.', - ] - : []), '', '# Git', '- Interactive flags (`-i`, e.g. `git rebase -i`, `git add -i`) are not supported in this environment.', '- Use the `gh` CLI for GitHub operations (PRs, issues, API).', - '- Commit or push only when the user asks. If on the default branch, branch first.', + '- Commit or push only when the user asks.', ].join('\n'), schema: Type.Object({ command: Type.String({ - description: 'Shell command line to execute locally.', + description: 'The command to execute', }), + description: BashDescriptionSchema, timeout: Type.Optional( Type.Number({ - maximum: MAX_BACKGROUND_BASH_TIMEOUT_SECONDS, + exclusiveMinimum: 0, description: - 'Timeout in seconds. Foreground: default 120s, max 300s. Use run_in_background for longer commands. Background tasks use the requested timeout, subject to a runtime watchdog.', + 'Total command timeout in seconds. Foreground-only: default 120s, max 300s. Foreground with automatic backgrounding: default/max 3600s total, including foreground time. Explicit background: uses the specified timeout, or a 1-hour runtime limit if omitted. Backgrounding does not reset the timeout.', }), ), run_in_background: Type.Optional( Type.Boolean({ description: - 'Start in the background and return a task id immediately. Foreground commands may also return a task id after 15s without restarting the process.', + 'Start in the background and return a task id immediately. Foreground commands may also return a task id after 60s without restarting the process.', }), ), }), diff --git a/packages/agent-tools/src/desktop/index.ts b/packages/agent-tools/src/desktop/index.ts index 5b289f55..7d8b8875 100644 --- a/packages/agent-tools/src/desktop/index.ts +++ b/packages/agent-tools/src/desktop/index.ts @@ -49,6 +49,10 @@ export * from './local-ask-user.js'; export * from './local-feature-enable.js'; export * from './local-browser.js'; export * from './local-pi-tools.js'; +export * from './local-bash-result.js'; +export * from './local-bash-timing.js'; +export * from './local-bash-input.js'; +export * from './local-bash-contract.js'; export { executeLocalHostTrashIfRequested } from './host-trash-executor.js'; export type { LocalHostTrashRuntime } from './host-trash-executor.js'; export * from './local-glob.js'; @@ -198,3 +202,4 @@ export { CloudSessionReadError, type CloudSessionReaderOptions, } from './cloud-session-reader.js'; +export { DESKTOP_BASH_PREVIEW_BYTES } from './output-limit.js'; diff --git a/packages/agent-tools/src/desktop/local-bash-contract.ts b/packages/agent-tools/src/desktop/local-bash-contract.ts new file mode 100644 index 00000000..0f5d5470 --- /dev/null +++ b/packages/agent-tools/src/desktop/local-bash-contract.ts @@ -0,0 +1,83 @@ +import { getShellConfig } from '@earendil-works/pi-coding-agent/shell'; +import type { ToolDefinition } from '@mavis/agent-core/tools'; +import { Clone, Type } from '@sinclair/typebox'; + +import { LocalBashToolDef } from './builtin-defs.js'; +import { + DEFAULT_FOREGROUND_BASH_SOFT_YIELD_MS, + DEFAULT_FOREGROUND_BASH_TIMEOUT_SECONDS, + MAX_FOREGROUND_BASH_TIMEOUT_SECONDS, + MAX_MANAGED_BASH_TIMEOUT_SECONDS, +} from './local-bash-timing.js'; + +export interface LocalBashTurnCapabilities { + readonly background: boolean; + readonly shell: ReturnType['type'] | 'unavailable'; +} + +export function resolveLocalBashShell(): LocalBashTurnCapabilities['shell'] { + try { + return getShellConfig().type; + } catch { + return 'unavailable'; + } +} + +/** A fresh definition for the final admitted Turn inventory. */ +export function createLocalBashToolDefinition( + capabilities: LocalBashTurnCapabilities, +): ToolDefinition { + const base = Clone(LocalBashToolDef.schema); + const schema = capabilities.background ? base : Type.Omit(base, ['run_in_background']); + schema.properties.timeout.description = capabilities.background + ? `Total command timeout in seconds. Foreground with automatic backgrounding: default/max ${MAX_MANAGED_BASH_TIMEOUT_SECONDS}s, including foreground time. Explicit background: uses the specified timeout, or a 1-hour runtime limit if omitted.` + : `Timeout in seconds. Foreground: default ${DEFAULT_FOREGROUND_BASH_TIMEOUT_SECONDS}s; values above ${MAX_FOREGROUND_BASH_TIMEOUT_SECONDS}s are capped.`; + return { + ...LocalBashToolDef, + schema, + description: `${LocalBashToolDef.description}\n\n${renderBashUsage(capabilities)}`, + }; +} + +function renderBashUsage(capabilities: LocalBashTurnCapabilities): string { + return [ + ...(capabilities.background + ? [ + `- Foreground calls wait up to ${DEFAULT_FOREGROUND_BASH_SOFT_YIELD_MS / 1000}s, then return a task_id if unfinished. The same command continues in the background; do not start it again.`, + '- Set run_in_background=true to return a task_id without waiting for completion.', + '- Background completion notifies you automatically. Continue independent work; use task_output for current output and task_stop, when available, to cancel.', + '- Backgrounding and reading output do not reset the command timeout.', + ] + : ['- This Turn supports foreground execution only.']), + "- Large output preserves the beginning and end. Follow the result's instructions to read omitted output.", + ...shellRules(capabilities.shell), + ].join('\n'); +} + +function shellRules(shell: LocalBashTurnCapabilities['shell']): string[] { + if (shell === 'unavailable') + return [ + '- No local shell was resolved. Do not assume Bash or PowerShell syntax is executable.', + ]; + if (shell === 'bash') + return [ + '- Selected shell: Bash. Use Bash syntax, including on Windows when Bash is the selected fallback.', + ]; + if (shell === 'sh') + return [ + '- Selected shell: POSIX sh. Use POSIX shell syntax; Bash arrays, [[ ... ]] and process substitution are unavailable.', + ]; + return [ + shell === 'pwsh' + ? '- Selected shell: PowerShell 7 (pwsh). Pipeline-chain operators && and || are supported; use PowerShell quoting and variable syntax.' + : '- Selected shell: Windows PowerShell 5.1. The && and || operators are unsupported; use separate statements and explicit success checks.', + '- Use $env:VAR and single quotes for literals/regex. Do not use Bash export, /dev/null, heredocs, sed -i, or backslash quoting. Do not wrap ordinary commands in powershell -Command or cmd /c.', + "- Start multi-statement scripts with $ErrorActionPreference = 'Stop'. Check $LASTEXITCODE after native programs, or enable $PSNativeCommandUseErrorActionPreference when available.", + `- Prefer dedicated file tools. If shell file I/O is necessary, specify UTF-8 and avoid Get-Content | ... | Set-Content editing pipelines.${ + shell === 'powershell' + ? ' PowerShell 5.1 -Encoding UTF8 writes a BOM.' + : ' PowerShell 7 defaults to UTF-8 without BOM.' + }`, + '- If a CLI is missing, check .cmd/.ps1 wrappers. Change strategy after repeated syntax failures.', + ]; +} diff --git a/packages/agent-tools/src/desktop/local-bash-input.ts b/packages/agent-tools/src/desktop/local-bash-input.ts new file mode 100644 index 00000000..aba2d6f8 --- /dev/null +++ b/packages/agent-tools/src/desktop/local-bash-input.ts @@ -0,0 +1,14 @@ +import { Type } from '@sinclair/typebox'; + +export const BashDescriptionSchema = Type.Optional( + Type.String({ + description: 'A short description of what this command does.', + }), +); + +export function resolveBashDescription(description: unknown, command: string): string { + if (description !== undefined && typeof description !== 'string') { + throw new Error('Bash description must be a string.'); + } + return description || command; +} diff --git a/packages/agent-tools/src/desktop/local-bash-result.ts b/packages/agent-tools/src/desktop/local-bash-result.ts new file mode 100644 index 00000000..ca29c048 --- /dev/null +++ b/packages/agent-tools/src/desktop/local-bash-result.ts @@ -0,0 +1,99 @@ +import { + BashExecutionError, + type BashExecutionOutcome, + type BashToolDetails, +} from '@earendil-works/pi-coding-agent/tools'; +import type { AgentToolResult } from '@earendil-works/pi-agent-core'; + +import type { LocalBackgroundBashExecutorResult } from './types.js'; + +export interface LocalBashExecutionOutcome extends Omit { + reason: BashExecutionOutcome['reason'] | 'preflight_failed' | 'watchdog_timeout'; +} + +export function localBashResultFromPi( + result: AgentToolResult, +): LocalBackgroundBashExecutorResult { + const details = (result.details ?? {}) as Record; + const execution = readBashExecutionOutcome(details); + return { + text: result.content + .filter((part): part is { type: 'text'; text: string } => part.type === 'text') + .map((part) => part.text) + .join('\n'), + details, + ...(execution && execution.status !== 'succeeded' ? { isError: true } : {}), + }; +} + +export function localBashResultFromError( + error: unknown, + signal?: AbortSignal, +): LocalBackgroundBashExecutorResult { + const details: BashToolDetails = error instanceof BashExecutionError ? error.details : {}; + const cause = error instanceof BashExecutionError ? error.cause : error; + const causeRecord = + cause && typeof cause === 'object' ? (cause as Record) : undefined; + const errorCode = typeof causeRecord?.code === 'string' ? causeRecord.code : undefined; + const preflight = + errorCode?.startsWith('SANDBOX_') && + (causeRecord?.stage === undefined || + ['init', 'parse', 'invocation', 'pre-spawn'].includes(String(causeRecord.stage))); + const text = error instanceof Error ? error.message : String(error); + const execution: LocalBashExecutionOutcome = { + status: signal?.aborted ? 'canceled' : 'failed', + reason: signal?.aborted ? 'canceled' : 'unknown', + exitCode: null, + message: text, + ...(errorCode ? { errorCode } : {}), + ...details.execution, + ...(preflight ? { reason: 'preflight_failed' as const } : {}), + }; + if (execution.status === 'canceled' && !execution.cancellationReason && signal?.aborted) { + execution.cancellationReason = + signal.reason instanceof Error + ? signal.reason.message + : typeof signal.reason === 'string' + ? signal.reason + : undefined; + } + return { text, details: { ...details, execution }, isError: true }; +} + +export function readBashExecutionOutcome(details: unknown): LocalBashExecutionOutcome | undefined { + if (!details || typeof details !== 'object' || !('execution' in details)) return undefined; + const execution = details.execution; + if (!execution || typeof execution !== 'object') return undefined; + const value = execution as Record; + if ( + typeof value.status !== 'string' || + !['succeeded', 'failed', 'canceled'].includes(value.status) || + typeof value.reason !== 'string' || + ![ + 'exited', + 'signaled', + 'command_timeout', + 'canceled', + 'spawn_failed', + 'unknown', + 'preflight_failed', + 'watchdog_timeout', + ].includes(value.reason) || + !( + value.exitCode === null || + (typeof value.exitCode === 'number' && Number.isInteger(value.exitCode)) + ) || + [value.signal, value.errorCode, value.message, value.cancellationReason].some( + (field) => field !== undefined && typeof field !== 'string', + ) + ) + return undefined; + return execution as LocalBashExecutionOutcome; +} + +export function formatBashExecutionOutcome(execution: LocalBashExecutionOutcome): string { + if (execution.message) return execution.message; + if (execution.signal) return `Command terminated by signal ${execution.signal}`; + if (execution.exitCode !== null) return `Command exited with code ${execution.exitCode}`; + return `Command ${execution.status} (${execution.reason})`; +} diff --git a/packages/agent-tools/src/desktop/local-bash-timing.ts b/packages/agent-tools/src/desktop/local-bash-timing.ts new file mode 100644 index 00000000..58a19750 --- /dev/null +++ b/packages/agent-tools/src/desktop/local-bash-timing.ts @@ -0,0 +1,73 @@ +import type { ToolResult } from '@mavis/agent-core/tools'; + +export const DEFAULT_FOREGROUND_BASH_TIMEOUT_SECONDS = 120; +export const MAX_FOREGROUND_BASH_TIMEOUT_SECONDS = 300; +export const MAX_MANAGED_BASH_TIMEOUT_SECONDS = 3_600; +export const MAX_BASH_TIMEOUT_SECONDS = 2_147_483; +export const DEFAULT_FOREGROUND_BASH_SOFT_YIELD_MS = 60_000; + +export interface LocalBashTiming { + commandTimeoutSeconds?: number; + requestedTimeoutSeconds?: number; + commandTimerStartedAt?: number; + commandDeadlineAt?: number; +} + +export function resolveLocalBashTiming( + timeout: number | undefined, + mode: 'direct_foreground' | 'managed_foreground' | 'explicit_background', +): LocalBashTiming { + const normalizedTimeout = + mode === 'managed_foreground' && + timeout !== undefined && + (!Number.isFinite(timeout) || timeout <= 0) + ? undefined + : timeout; + if ( + normalizedTimeout !== undefined && + (!Number.isFinite(normalizedTimeout) || + normalizedTimeout <= 0 || + (mode !== 'managed_foreground' && normalizedTimeout > MAX_BASH_TIMEOUT_SECONDS)) + ) { + throw new Error( + `timeout must be a finite positive number of seconds, at most ${MAX_BASH_TIMEOUT_SECONDS}.`, + ); + } + const commandTimeoutSeconds = + mode === 'explicit_background' + ? normalizedTimeout + : mode === 'managed_foreground' + ? Math.min( + normalizedTimeout ?? MAX_MANAGED_BASH_TIMEOUT_SECONDS, + MAX_MANAGED_BASH_TIMEOUT_SECONDS, + ) + : Math.min( + normalizedTimeout ?? DEFAULT_FOREGROUND_BASH_TIMEOUT_SECONDS, + MAX_FOREGROUND_BASH_TIMEOUT_SECONDS, + ); + return { + commandTimeoutSeconds, + ...(normalizedTimeout !== undefined && normalizedTimeout !== commandTimeoutSeconds + ? { requestedTimeoutSeconds: normalizedTimeout } + : {}), + }; +} + +export function validateBashCommand(command: string): void { + if (typeof command !== 'string' || command.trim().length === 0) { + throw new Error('command must be a non-empty shell command.'); + } +} + +export function withLocalBashTiming(result: ToolResult, timing: LocalBashTiming): ToolResult { + const text = + timing.requestedTimeoutSeconds !== undefined + ? `${result.text}\n\n[bash timeout: requested ${timing.requestedTimeoutSeconds}s; effective ${timing.commandTimeoutSeconds}s command limit.]` + : result.text; + return { + ...result, + text, + content: [{ type: 'text', text }], + details: { ...result.details, timing: { ...(result.details?.timing as object), ...timing } }, + }; +} diff --git a/packages/agent-tools/src/desktop/local-memory.ts b/packages/agent-tools/src/desktop/local-memory.ts index c22957b9..ad132790 100644 --- a/packages/agent-tools/src/desktop/local-memory.ts +++ b/packages/agent-tools/src/desktop/local-memory.ts @@ -1,6 +1,11 @@ import { bindTool, type ToolImpl, type ToolResult } from '@mavis/agent-core/tools'; +import { MEMORY_TOOL_OUTPUT_MAX_BYTES } from '@mavis/shared/memory-limits'; import { LocalMemoryToolDef, type LocalMemoryToolInput } from './builtin-defs.js'; +import { + applyDesktopTextLimit, + limitDesktopHeadTailLines, +} from './output-limit.js'; import type { LocalMemoryAdapter, LocalRuntimeToolContext } from './types.js'; @bindTool(LocalMemoryToolDef) @@ -17,11 +22,19 @@ export class LocalMemoryTool implements ToolImpl< ): Promise { if (signal?.aborted) throw new Error('Operation aborted'); const result = await this.adapter.execute(ctx, input, signal); - return { + const toolResult: ToolResult = { tool_name: LocalMemoryToolDef.name, text: result.text, content: [{ type: 'text', text: result.text }], details: { kind: 'memory', ...(result.details ?? {}) }, }; + return applyDesktopTextLimit( + toolResult, + limitDesktopHeadTailLines(result.text, { + maxBytes: MEMORY_TOOL_OUTPUT_MAX_BYTES, + notice: () => + '[Memory output truncated. Narrow the memory search query or use the read tool with a file path and offset/limit to inspect more.]', + }), + ); } } diff --git a/packages/agent-tools/src/desktop/local-pi-tools.ts b/packages/agent-tools/src/desktop/local-pi-tools.ts index 81c0c3ed..d044552e 100644 --- a/packages/agent-tools/src/desktop/local-pi-tools.ts +++ b/packages/agent-tools/src/desktop/local-pi-tools.ts @@ -21,6 +21,19 @@ import { isAbsolute, resolve as resolvePath } from 'node:path'; import { homedir, release } from 'node:os'; import { fileURLToPath } from 'node:url'; import { randomUUID } from 'node:crypto'; +import { resolveBashDescription } from './local-bash-input.js'; +import { + DEFAULT_FOREGROUND_BASH_SOFT_YIELD_MS, + resolveLocalBashTiming, + validateBashCommand, + withLocalBashTiming, + type LocalBashTiming, +} from './local-bash-timing.js'; +import { + localBashResultFromError, + localBashResultFromPi, + readBashExecutionOutcome, +} from './local-bash-result.js'; import { LocalBashToolDef, @@ -65,6 +78,7 @@ import { buildWriteToolResult, createCapturingWriteOperations } from '../shared/ import { applyDesktopTextLimit, DESKTOP_BASH_MAX_BYTES, + DESKTOP_BASH_PREVIEW_BYTES, DESKTOP_READ_TEXT_MAX_BYTES, limitDesktopHeadTailLines, limitDesktopPrefixLines, @@ -372,6 +386,11 @@ export class LocalBashTool implements ToolImpl< // The background executor and the pi-turn-runner fallback apply the same // hook — keep the three spawn sites in sync. this.tool = createBashTool(workspaceRoot, { + output: { + strategy: 'head_tail', + maxBytes: DESKTOP_BASH_PREVIEW_BYTES, + maxLines: Number.MAX_SAFE_INTEGER, + }, spawnHook: (ctx) => { const { env, removed } = sanitizeBashSubprocessEnv(ctx.env, this.envPolicy); this.lastEnvRemoved = removed; @@ -386,6 +405,49 @@ export class LocalBashTool implements ToolImpl< signal?: AbortSignal, ): Promise { if (signal?.aborted) throw new Error('Operation aborted'); + let timing: LocalBashTiming; + let description: string; + try { + description = resolveBashDescription(input.description, input.command); + validateBashCommand(input.command); + timing = resolveLocalBashTiming( + input.timeout, + input.run_in_background + ? 'explicit_background' + : ctx.canConsumeBackgroundBashOutput === true && + ctx.allowBashAutoPromotion !== false && + this.backgroundAdapter?.runManagedForeground + ? 'managed_foreground' + : 'direct_foreground', + ); + } catch (error) { + const text = error instanceof Error ? error.message : String(error); + return { + tool_name: LocalBashToolDef.name, + text, + content: [{ type: 'text', text }], + isError: true, + details: { status: 'invalid_input', error_code: 'BASH_INVALID_INPUT' }, + }; + } + try { + const result = await this.executeValidated(ctx, { ...input, description }, timing, signal); + return { ...result, details: { ...result.details, description } }; + } catch (error) { + if (error instanceof Error) { + const details = (error as Error & { details?: Record }).details; + Object.assign(error, { details: { ...details, description } }); + } + throw error; + } + } + + private async executeValidated( + ctx: LocalRuntimeToolContext, + input: LocalBashToolInput, + timing: LocalBashTiming, + signal?: AbortSignal, + ): Promise { if (input.run_in_background === true && ctx.canConsumeBackgroundBashOutput !== true) { return unavailableBackgroundBashOutputResult(); } @@ -394,7 +456,7 @@ export class LocalBashTool implements ToolImpl< workspaceRoot: this.workspaceRoot, runtime: this.hostTrashRuntime, envPolicy: this.envPolicy, - timeoutSeconds: resolveForegroundTimeout(input), + timeoutSeconds: resolveForegroundTimeout({ timeout: timing.commandTimeoutSeconds }), ...(signal ? { signal } : {}), }); if (hostTrashResult) { @@ -426,13 +488,21 @@ export class LocalBashTool implements ToolImpl< return backgroundBashToolResult( started.taskId, 'started', - this.consumeEnvSanitizationHint(started.details), + this.consumeEnvSanitizationHint({ + ...started.details, + description: input.description, + timing: { ...timing, ...(started.details?.timing as object) }, + }), ); } const { run_in_background: _runInBackground, ...rest } = input; - const foregroundInput = { ...rest, timeout: resolveForegroundTimeout(rest) }; - if (ctx.allowBashAutoPromotion !== false && this.backgroundAdapter?.runManagedForeground) { + const foregroundInput = { ...rest, timeout: timing.commandTimeoutSeconds }; + if ( + ctx.canConsumeBackgroundBashOutput === true && + ctx.allowBashAutoPromotion !== false && + this.backgroundAdapter?.runManagedForeground + ) { const managed = await this.backgroundAdapter.runManagedForeground( ctx, foregroundInput, @@ -444,7 +514,11 @@ export class LocalBashTool implements ToolImpl< return backgroundBashToolResult( managed.taskId, 'auto_promoted', - this.consumeEnvSanitizationHint(managed.details), + this.consumeEnvSanitizationHint({ + ...managed.details, + description: input.description, + timing: { ...timing, ...(managed.details?.timing as object) }, + }), ); } const text = `${managed.errorMessage ?? 'Local managed bash failed to start'}`; @@ -460,27 +534,36 @@ export class LocalBashTool implements ToolImpl< typeof managed.details?.fullOutputPath === 'string' ? managed.details.fullOutputPath : undefined; - const limited = limitDesktopBashOutput(managed.text, managedFullOutputPath); + const managedDetails = this.consumeEnvSanitizationHint(managed.details); - return applyDesktopTextLimit( - { + const result = withLocalBashTiming( + withCompatibleBashToolResponseFromPiDetails({ tool_name: LocalBashToolDef.name, text: managed.text, content: [{ type: 'text', text: managed.text }], details: { ...managedDetails, - status: 'completed', + status: readBashExecutionOutcome(managedDetails)?.status ?? 'completed', task_id: managed.taskId, }, ...(managed.isError ? { isError: true } : {}), - }, - limited, + }), + timing, + ); + return applyDesktopTextLimit( + result, + limitDesktopBashOutput(result.text, managedFullOutputPath, result.details, managed.taskId), ); } let result: ToolResult; try { const tool = this.sandboxOperationsFactory ? createBashTool(this.workspaceRoot, { + output: { + strategy: 'head_tail', + maxBytes: DESKTOP_BASH_PREVIEW_BYTES, + maxLines: Number.MAX_SAFE_INTEGER, + }, operations: this.sandboxOperationsFactory.create({ identity: directForegroundIdentity(ctx), workspaceRoot: this.workspaceRoot, @@ -493,28 +576,29 @@ export class LocalBashTool implements ToolImpl< }) : this.tool; const res = await tool.execute('', foregroundInput, signal); - result = withCompatibleBashToolResponseFromPiDetails( - toToolResult(LocalBashToolDef.name, res), - ); + const completed = localBashResultFromPi(res); + result = withCompatibleBashToolResponseFromPiDetails({ + tool_name: LocalBashToolDef.name, + ...completed, + content: [{ type: 'text', text: completed.text }], + }); } catch (error) { if (signal?.aborted || (error instanceof Error && error.name === 'AbortError')) { throw error; } - const text = error instanceof Error ? error.message : String(error); - const fullOutputPath = extractBashFullOutputPath(text); - result = { + const failed = localBashResultFromError(error, signal); + result = withCompatibleBashToolResponseFromPiDetails({ tool_name: LocalBashToolDef.name, - text, - content: [{ type: 'text', text }], - ...(fullOutputPath ? { details: { fullOutputPath } } : {}), - isError: true, - }; + ...failed, + content: [{ type: 'text', text: failed.text }], + }); } + result = withLocalBashTiming(result, timing); const fullOutputPath = typeof result.details?.fullOutputPath === 'string' ? result.details.fullOutputPath : undefined; - const limited = limitDesktopBashOutput(result.text, fullOutputPath); + const limited = limitDesktopBashOutput(result.text, fullOutputPath, result.details); result = applyDesktopTextLimit(result, limited); result.details = this.consumeEnvSanitizationHint({ ...(result.details ?? {}), @@ -550,57 +634,67 @@ function directForegroundIdentity(ctx: LocalRuntimeToolContext): LocalSandboxInv }; } -function limitDesktopBashOutput(text: string, fullOutputPath?: string) { - const limited = limitDesktopHeadTailLines(text, { +function limitDesktopBashOutput( + text: string, + fullOutputPath?: string, + details?: Record, + taskId?: string, +) { + const output = details?.output as { rawBytes?: number; persistenceError?: string } | undefined; + let previewText = text; + if (!taskId) { + if ((details?.truncation as { truncated?: boolean } | undefined)?.truncated) { + previewText += `\n\n[desktop bash output truncated: original_bytes=${output?.rawBytes}; original head+tail shown.${fullOutputPath ? ` Full output: ${fullOutputPath}` : ''}]`; + } + if (output?.persistenceError) { + previewText += `\n\n[Output persistence failed; complete output is unavailable: ${output.persistenceError}]`; + } + } + const limited = limitDesktopHeadTailLines(previewText, { maxBytes: DESKTOP_BASH_MAX_BYTES, notice: ({ originalBytes }) => fullOutputPath ? `[desktop bash output truncated: original_bytes=${originalBytes}; head+tail shown; boundary lines may be UTF-8 byte-truncated; Full output: ${fullOutputPath}]` - : `[desktop bash output truncated: original_bytes=${originalBytes}; head+tail shown; boundary lines may be UTF-8 byte-truncated; rerun with a narrower command or redirect output to a file and use read.]`, + : `[desktop bash output truncated: original_bytes=${originalBytes}; head+tail shown; boundary lines may be UTF-8 byte-truncated.]`, }); + if ( + limited.truncation || + (details?.truncation as { truncated?: boolean } | undefined)?.truncated + ) { + limited.truncation = { + truncated: true, + has_more: true, + strategy: 'head_tail_lines', + original_bytes: + output?.rawBytes ?? limited.truncation?.original_bytes ?? Buffer.byteLength(text, 'utf8'), + returned_bytes: Buffer.byteLength(limited.text, 'utf8'), + max_bytes: DESKTOP_BASH_MAX_BYTES, + }; + } return withDesktopOutputContinuation(limited, { - continuation_hint: fullOutputPath + continuation_hint: taskId ? { - tool: 'read', - preserve_args: ['path'], - instruction: `Read the full output at ${fullOutputPath}; continue long reads with the returned next_offset.`, + tool: 'task_output', + preserve_args: ['task_id'], + instruction: `Read task ${taskId} with task_output; continue with its next_offset (UTF-8 log file bytes). Check output persistence status before treating the log as complete.`, } - : { - tool: 'bash', - preserve_args: [], - instruction: - 'Do not blindly rerun a potentially side-effecting command; narrow it or redirect output to a file, then use read.', - }, + : fullOutputPath + ? { + tool: 'read', + preserve_args: ['path'], + instruction: `Read the full output at ${fullOutputPath}; continue long reads with the returned next_offset.`, + } + : { + tool: 'bash', + preserve_args: [], + instruction: + 'Do not blindly rerun a potentially side-effecting command; narrow it or redirect output to a file, then use read.', + }, }); } -function extractBashFullOutputPath(text: string): string | undefined { - const matches = [...text.matchAll(/Full output:\s*(.+?)\](?:\n|$)/gi)]; - const path = matches.at(-1)?.[1]?.trim(); - return path || undefined; -} - -/** - * Default foreground timeout (seconds) applied when the model omits `timeout`. - * Restores main's historical 120s bound so interactive/blocking foreground - * commands (e.g. `lark-cli auth login`) cannot hang forever. Background tasks - * (`run_in_background: true`) are intentionally NOT bounded by this. - */ -export const DEFAULT_FOREGROUND_BASH_TIMEOUT_SECONDS = 120; -export const MAX_FOREGROUND_BASH_TIMEOUT_SECONDS = 300; -export const DEFAULT_FOREGROUND_BASH_SOFT_YIELD_MS = 15_000; - -/** - * Keep interactive foreground commands bounded. Missing, non-finite and - * non-positive values use the 120s default; larger values are capped at 5m. - * Longer commands should opt into background execution explicitly. - */ export function resolveForegroundTimeout(input: { timeout?: number }): number { - const requested = input.timeout; - if (requested === undefined || !Number.isFinite(requested) || requested <= 0) { - return DEFAULT_FOREGROUND_BASH_TIMEOUT_SECONDS; - } - return Math.min(requested, MAX_FOREGROUND_BASH_TIMEOUT_SECONDS); + return resolveLocalBashTiming(input.timeout, 'direct_foreground').commandTimeoutSeconds!; } function backgroundBashToolResult( @@ -611,8 +705,20 @@ function backgroundBashToolResult( const lead = status === 'auto_promoted' ? 'The command is still running and was yielded to a managed background task without restarting it.' - : 'Started local background bash task.'; - const text = `\n${lead} The owning conversation will automatically resume when it finishes; use task_query/task_output to inspect incremental output and task_stop to cancel it.\n`; + : 'Background Bash task accepted; command startup may still be in progress.'; + const timing = details.timing as LocalBashTiming | undefined; + const timingText = + timing?.commandTimeoutSeconds !== undefined + ? ` Command limit: ${timing.commandTimeoutSeconds}s total; yielding does not reset it.` + : ''; + const purpose = + typeof details.description === 'string' + ? `Purpose: ${limitDesktopHeadTailLines(details.description, { + maxBytes: 1024, + notice: () => '[purpose truncated]', + }).text}\n` + : ''; + const text = `\n${purpose}${lead}${timingText} The owning conversation will automatically resume when it finishes; use task_output to inspect incremental output, and task_stop to cancel if that tool is available.\n`; return { tool_name: LocalBashToolDef.name, text, @@ -735,3 +841,9 @@ async function pathExists(filePath: string): Promise { function isFsNotFoundError(err: unknown): boolean { return typeof err === 'object' && err !== null && (err as { code?: unknown }).code === 'ENOENT'; } + +export { + DEFAULT_FOREGROUND_BASH_TIMEOUT_SECONDS, + MAX_FOREGROUND_BASH_TIMEOUT_SECONDS, + DEFAULT_FOREGROUND_BASH_SOFT_YIELD_MS, +} from './local-bash-timing.js'; diff --git a/packages/agent-tools/src/desktop/local-task-control.ts b/packages/agent-tools/src/desktop/local-task-control.ts index 48f7f18b..cb05921f 100644 --- a/packages/agent-tools/src/desktop/local-task-control.ts +++ b/packages/agent-tools/src/desktop/local-task-control.ts @@ -14,6 +14,7 @@ import type { LocalTaskControlAdapter, LocalTaskOutputReadResult, } from './types.js'; +import { formatBashExecutionOutcome, readBashExecutionOutcome } from './local-bash-result.js'; const MAX_TASK_OUTPUT_WAIT_MS = 30_000; @@ -82,6 +83,13 @@ export class LocalTaskOutputTool implements ToolImpl< ...(signal ? { signal } : {}), }); const status = read.status ?? task.status; + const snapshot = read.task ?? task; + const execution = readBashExecutionOutcome(snapshot.metadata?.bashDetails); + const output = (snapshot.metadata?.bashDetails as Record | undefined)?.output; + const persistenceHint = + (output as { persistence?: string } | undefined)?.persistence === 'incomplete' + ? 'Output persistence is incomplete; this log may be missing command output.\n' + : ''; const body = read.content || '(no output yet)'; const cursor = [ ...(read.nextOffset === undefined ? [] : [`next_offset="${read.nextOffset}"`]), @@ -94,10 +102,18 @@ export class LocalTaskOutputTool implements ToolImpl< ? '' : 'Your requested wait_ms exceeded 30000 ms and was capped at 30000 ms (30 seconds). This output read reached its wait limit; the background task was not stopped. Use wait_ms=30000 or less for future reads.\n'; const hint = signal?.aborted ? '' : this.pollingHint(ctx, input.task_id, { ...read, status }); - const text = `\n${receipt}${waitLimitHint}${hint}${body}\n`; - return ok(LocalTaskOutputToolDef.name, text, { + const outcome = execution + ? `${formatBashExecutionOutcome(execution)}\n` + : ''; + const text = `\n${receipt}${waitLimitHint}${persistenceHint}${hint}${outcome}${body}\n`; + const result = ok(LocalTaskOutputToolDef.name, text, { task_id: input.task_id, status, + ...(execution ? { execution } : {}), + ...(output ? { output } : {}), + timing: + (snapshot.metadata?.bashDetails as Record | undefined)?.timing ?? + snapshot.metadata?.timing, effective_wait_ms: effectiveWaitMs, wait_ms_clamped: waitMsClamped, ...(read.timedOut === undefined ? {} : { timed_out: read.timedOut }), @@ -105,6 +121,7 @@ export class LocalTaskOutputTool implements ToolImpl< ...(read.truncated === undefined ? {} : { truncated: read.truncated }), ...(read.summary ? { summary: read.summary } : {}), }); + return result; } private pollingHint( @@ -152,8 +169,21 @@ export class LocalTaskStopTool implements ToolImpl< async execute(ctx: LocalRuntimeToolContext, input: LocalTaskStopToolInput): Promise { const task = await this.adapter.stop(ctx, input.task_id, input.reason); if (!task) return notFound(LocalTaskStopToolDef.name, input.task_id); - const text = `\nStop requested for ${input.task_id}.\n`; - return ok(LocalTaskStopToolDef.name, text, { task_id: input.task_id, status: task.status }); + const stopError = + task.metadata?.stopError ?? + (task.lastError?.code === 'TASK_STOP_FAILED' ? task.lastError : undefined); + const message = stopError + ? `Stop failed for ${input.task_id}: ${(stopError as { message: string }).message}. Process cleanup could not be confirmed.` + : `Task ${input.task_id} is ${task.status}.`; + const text = `\n${message}\n`; + return { + ...ok(LocalTaskStopToolDef.name, text, { + task_id: input.task_id, + status: task.status, + ...(stopError ? { stopError } : {}), + }), + ...(stopError ? { isError: true } : {}), + }; } } @@ -167,16 +197,28 @@ function renderTask(task: BackgroundTask): string { ...(mode ? [`execution_mode=${mode}`] : []), ]; const lineageText = lineage.length === 0 ? '' : ` (${lineage.join(' ')})`; - return `- ${task.taskId} [${task.kind}/${task.status}]${description}${lineageText}`; + const execution = readBashExecutionOutcome(task.metadata?.bashDetails); + const outcome = execution ? ` — ${formatBashExecutionOutcome(execution)}` : ''; + const output = (task.metadata?.bashDetails as { output?: { persistence?: string } } | undefined) + ?.output; + const persistence = output?.persistence === 'incomplete' ? ' — output log incomplete' : ''; + return `- ${task.taskId} [${task.kind}/${task.status}]${description}${lineageText}${outcome}${persistence}`; } function summarizeTask(task: BackgroundTask): Record { const sessionId = childSessionId(task); const mode = executionMode(task); + const execution = readBashExecutionOutcome(task.metadata?.bashDetails); return { task_id: task.taskId, kind: task.kind, status: task.status, + ...(execution ? { execution } : {}), + output: (task.metadata?.bashDetails as Record | undefined)?.output, + timing: + (task.metadata?.bashDetails as Record | undefined)?.timing ?? + task.metadata?.timing, + ...(task.metadata?.stopError ? { stopError: task.metadata.stopError } : {}), // Handle recovery after the first receipt scrolled out of context. Only a // subagent task that really owns a child Session gets a session handle. ...(sessionId ? { session_id: sessionId } : {}), diff --git a/packages/agent-tools/src/desktop/output-limit.ts b/packages/agent-tools/src/desktop/output-limit.ts index b7225c83..eaf2e971 100644 --- a/packages/agent-tools/src/desktop/output-limit.ts +++ b/packages/agent-tools/src/desktop/output-limit.ts @@ -3,12 +3,14 @@ import type { ToolResult, ToolResultContent } from '@mavis/agent-core/tools'; export const DESKTOP_GREP_CONTENT_MAX_BYTES = 16 * 1024; export const DESKTOP_READ_TEXT_MAX_BYTES = 24 * 1024; export const DESKTOP_BASH_MAX_BYTES = 24 * 1024; +// Leave room for exit/timing receipts and the readable output reference. +export const DESKTOP_BASH_PREVIEW_BYTES = DESKTOP_BASH_MAX_BYTES - 2048; export type DesktopOutputLimitStrategy = 'prefix_lines' | 'head_tail_lines'; export type DesktopOutputOffsetUnit = 'line' | 'match' | 'file'; export interface DesktopOutputContinuationHint { - tool: 'read' | 'grep' | 'glob' | 'mavis' | 'bash'; + tool: 'read' | 'grep' | 'glob' | 'mavis' | 'bash' | 'task_output'; preserve_args: readonly string[]; instruction: string; } diff --git a/packages/agent-tools/src/desktop/types.ts b/packages/agent-tools/src/desktop/types.ts index 20f6f0b6..69f02edb 100644 --- a/packages/agent-tools/src/desktop/types.ts +++ b/packages/agent-tools/src/desktop/types.ts @@ -322,6 +322,7 @@ export interface LocalSandboxBashOperationsFactory { export interface LocalBackgroundBashExecutorResult { text: string; details?: Record; + isError?: boolean; } export interface LocalBackgroundBashExecutor { @@ -434,6 +435,8 @@ export interface LocalTaskOutputReadOptions extends TaskOutputReadOptions { export interface LocalTaskOutputReadResult extends TaskOutputReadResult { /** Status read from the same post-wait snapshot as the returned output. */ status?: BackgroundTaskStatus; + /** Task facts from the same post-wait snapshot as status. */ + task?: BackgroundTask; timedOut?: boolean; } diff --git a/packages/agent-tools/test/desktop/local-bash-output.test.ts b/packages/agent-tools/test/desktop/local-bash-output.test.ts new file mode 100644 index 00000000..45a12467 --- /dev/null +++ b/packages/agent-tools/test/desktop/local-bash-output.test.ts @@ -0,0 +1,197 @@ +/** + * Real foreground commands verify bounded head/tail previews and complete + * persisted output across the Pi executor and Desktop tool boundary. + */ + +import { mkdtemp, readFile, rm } from 'node:fs/promises'; +import { tmpdir } from 'node:os'; +import { join } from 'node:path'; + +import { afterEach, beforeEach, describe, expect, it, vi } from 'vitest'; + +import { LocalBashTool } from '../../src/desktop/local-pi-tools.js'; +import { DESKTOP_BASH_MAX_BYTES } from '../../src/desktop/output-limit.js'; +import { readPluginHookCompatibleToolResponse } from '../../src/plugin-hooks/vendor-tool-response.js'; +import type { LocalRuntimeToolContext } from '../../src/desktop/types.js'; + +const SESSION_CTX = { + sessionId: 'sess-test', + turnId: 'turn-test', +} satisfies LocalRuntimeToolContext; + +describe('LocalBashTool — foreground output truncation', () => { + let workspace: string; + + beforeEach(async () => { + workspace = await mkdtemp(join(tmpdir(), 'bash-output-')); + }); + afterEach(async () => { + await rm(workspace, { recursive: true, force: true }); + }); + + it('preserves exact stdout and stderr for Compatible PostToolUse', async () => { + const tool = new LocalBashTool(workspace, undefined, { mode: 'off' }); + const result = await tool.execute(SESSION_CTX, { + command: "printf 'from-stdout'; printf 'from-stderr' >&2", + }); + + expect(readPluginHookCompatibleToolResponse(result)).toEqual({ + stdout: 'from-stdout', + stderr: 'from-stderr', + interrupted: false, + }); + }); + + it('short output is returned verbatim, not truncated', async () => { + const tool = new LocalBashTool(workspace, undefined, { mode: 'off' }); + const result = await tool.execute(SESSION_CTX, { command: 'printf "a\\nb\\nc\\n"' }); + expect(result.tool_name).toBe('bash'); + expect(result.isError).toBeFalsy(); + expect(result.text).toContain('a'); + expect(result.text.toLowerCase()).not.toContain('full output:'); + }); + + it('caps 24-50 KiB foreground output with head and tail and persists the complete output', async () => { + const tool = new LocalBashTool(workspace, undefined, { mode: 'off' }); + const script = `for(let i=0;i<400;i++) console.log('row-'+i+'-'+'x'.repeat(90))`; + const command = `${JSON.stringify(process.execPath)} -e ${JSON.stringify(script)}`; + const result = await tool.execute(SESSION_CTX, { command }); + + expect(Buffer.byteLength(result.text, 'utf8')).toBeLessThanOrEqual( + DESKTOP_BASH_MAX_BYTES, + ); + expect(result.text).toContain('row-0-'); + expect(result.text).toContain('row-399-'); + expect(result.text).not.toContain('row-200-'); + expect(result.text.toLowerCase()).toContain('full output:'); + const fullOutput = await readFile(result.details?.fullOutputPath as string, 'utf8'); + expect(fullOutput.split('\n').filter(Boolean)).toHaveLength(400); + expect(fullOutput).toContain('row-200-'); + expect(result.details?.desktop_output_truncation).toMatchObject({ + truncated: true, + has_more: true, + strategy: 'head_tail_lines', + max_bytes: DESKTOP_BASH_MAX_BYTES, + continuation_hint: { + tool: 'read', + preserve_args: ['path'], + }, + }); + expect(result.details?.desktop_output_truncation).not.toHaveProperty('next_offset'); + expect(result.details?.desktop_output_truncation).not.toHaveProperty('offset_unit'); + }); + + it('caps a thrown timeout result while preserving the error ToolResult status', async () => { + const tool = new LocalBashTool(workspace, undefined, { mode: 'off' }); + const innerExecute = vi + .spyOn((tool as unknown as { tool: { execute: unknown } }).tool, 'execute') + .mockRejectedValue( + new Error( + `${'timeout-output\n'.repeat(2_000)}\n\nCommand timed out after 1 seconds`, + ), + ); + + const result = await tool.execute(SESSION_CTX, { command: 'ignored', timeout: 1 }); + + expect(innerExecute).toHaveBeenCalledOnce(); + expect(result.isError).toBe(true); + expect(Buffer.byteLength(result.text, 'utf8')).toBeLessThanOrEqual( + DESKTOP_BASH_MAX_BYTES, + ); + expect(result.text).toContain('Command timed out after 1 seconds'); + expect(result.text).toContain('desktop bash output truncated'); + expect(result.details?.desktop_output_truncation).toMatchObject({ + truncated: true, + has_more: true, + strategy: 'head_tail_lines', + max_bytes: DESKTOP_BASH_MAX_BYTES, + continuation_hint: { + tool: 'bash', + preserve_args: [], + }, + }); + }); + + it('still propagates an active AbortSignal instead of converting cancellation to a tool result', async () => { + const tool = new LocalBashTool(workspace, undefined, { mode: 'off' }); + const controller = new AbortController(); + vi.spyOn((tool as unknown as { tool: { execute: unknown } }).tool, 'execute').mockImplementation( + async () => { + controller.abort(); + throw new Error('Command aborted'); + }, + ); + + await expect( + tool.execute(SESSION_CTX, { command: 'ignored' }, controller.signal), + ).rejects.toThrow('Command aborted'); + }); + + it('caps a large nonzero-exit result while preserving the error ToolResult status', async () => { + const tool = new LocalBashTool(workspace, undefined, { mode: 'off' }); + // Let pending stderr writes drain before exiting with the intended error code. + const script = `for(let i=0;i<400;i++) console.error('error-row-'+i+'-'+'e'.repeat(90)); process.exitCode = 7`; + const command = `${JSON.stringify(process.execPath)} -e ${JSON.stringify(script)}`; + const result = await tool.execute(SESSION_CTX, { command }); + + expect(result.isError).toBe(true); + expect(Buffer.byteLength(result.text, 'utf8')).toBeLessThanOrEqual( + DESKTOP_BASH_MAX_BYTES, + ); + expect(result.text).toContain('error-row-0-'); + expect(result.text).toContain('error-row-399-'); + expect(result.text).not.toContain('error-row-200-'); + expect(result.text).toContain('Command exited with code 7'); + expect(result.text).toContain('desktop bash output truncated'); + const fullOutput = await readFile(result.details?.fullOutputPath as string, 'utf8'); + expect(fullOutput.split('\n').filter(Boolean)).toHaveLength(400); + expect(fullOutput).toContain('error-row-200-'); + expect(result.details?.desktop_output_truncation).toMatchObject({ + truncated: true, + has_more: true, + strategy: 'head_tail_lines', + max_bytes: DESKTOP_BASH_MAX_BYTES, + continuation_hint: { + tool: 'read', + preserve_args: ['path'], + }, + }); + }); + + it('output beyond 2000 lines preserves head and tail with a complete output file', async () => { + const tool = new LocalBashTool(workspace, undefined, { mode: 'off' }); + const script = `for(let i=0;i<3000;i++) console.log('spill-row-'+i+'-'+'z'.repeat(50))`; + const command = `${JSON.stringify(process.execPath)} -e ${JSON.stringify(script)}`; + const result = await tool.execute(SESSION_CTX, { command }); + + const lower = result.text.toLowerCase(); + expect(lower).toContain('full output:'); + expect(lower).toContain('original head+tail shown'); + expect(Buffer.byteLength(result.text, 'utf8')).toBeLessThanOrEqual( + DESKTOP_BASH_MAX_BYTES, + ); + + expect(result.text).toContain('spill-row-2999-'); + expect(result.text).toContain('spill-row-0-'); + expect(result.text).not.toContain('spill-row-1500-'); + expect(result.details?.desktop_output_truncation).toMatchObject({ + truncated: true, + has_more: true, + strategy: 'head_tail_lines', + max_bytes: DESKTOP_BASH_MAX_BYTES, + continuation_hint: { + tool: 'read', + preserve_args: ['path'], + }, + }); + + // Recover the spilled path from the footer and confirm it holds the FULL output. + const m = result.text.match(/full output:\s*(.+?)\]/i); + expect(m).not.toBeNull(); + const full = await readFile((m as RegExpMatchArray)[1].trim(), 'utf-8'); + const fullLines = full.split('\n').filter((l) => l.length > 0); + expect(fullLines).toHaveLength(3000); + expect(fullLines[0]).toContain('spill-row-0-'); + expect(fullLines[fullLines.length - 1]).toContain('spill-row-2999-'); + }); +}); diff --git a/packages/agent-tools/test/desktop/local-bash-timeout.test.ts b/packages/agent-tools/test/desktop/local-bash-timeout.test.ts new file mode 100644 index 00000000..cd87200f --- /dev/null +++ b/packages/agent-tools/test/desktop/local-bash-timeout.test.ts @@ -0,0 +1,63 @@ +import { describe, expect, it, vi } from 'vitest'; +import { LocalBashToolDef } from '../../src/desktop/builtin-defs.js'; +import { createLocalBashToolDefinition } from '../../src/desktop/local-bash-contract.js'; +import { LocalBashTool } from '../../src/desktop/local-pi-tools.js'; +import { resolveLocalBashTiming } from '../../src/desktop/local-bash-timing.js'; +import { + DEFAULT_BACKGROUND_BASH_MAX_RUN_MS, + resolveBackgroundBashMaxRunMs, +} from '../../../local-runtime/src/background-task/bash-runner-limits.js'; + +const context = { sessionId: 'test-session', turnId: 'test-turn', canConsumeBackgroundBashOutput: true }; + +describe('managed Bash timeout admission', () => { + it.each([ + [undefined, 3600], [7, 7], [3600, 3600], [7200, 3600], + [3_000_000, 3600], [0, 3600], [-5, 3600], [NaN, 3600], [Infinity, 3600], + ])('forwards the effective timeout for %s to the running command', async (timeout, effective) => { + const runManagedForeground = vi.fn(async () => ({ status: 'completed' as const, taskId: 'task', text: 'ok' })); + const tool = new LocalBashTool('/tmp/test-workspace', { startBackground: vi.fn(), runManagedForeground }, { mode: 'off' }); + const result = await tool.execute(context, { command: 'echo ok', timeout }); + expect(runManagedForeground).toHaveBeenCalledWith( + context, { command: 'echo ok', description: 'echo ok', timeout: effective }, 60_000, undefined, + ); + expect(result.details?.timing).toMatchObject({ commandTimeoutSeconds: effective }); + }); + + it('preserves the task identity after automatic promotion', async () => { + const startBackground = vi.fn(); + const tool = new LocalBashTool('/tmp/test-workspace', { + startBackground, + runManagedForeground: vi.fn(async () => ({ status: 'auto_promoted' as const, taskId: 'original-task' })), + }, { mode: 'off' }); + const result = await tool.execute(context, { command: 'sleep 65' }); + expect(result.details).toMatchObject({ task_id: 'original-task', timing: { commandTimeoutSeconds: 3600 } }); + expect(startBackground).not.toHaveBeenCalled(); + expect(result.text).toContain('without restarting it'); + }); + + it('keeps explicit background timeout input separate from the runtime watchdog', async () => { + const startBackground = vi.fn(async () => ({ status: 'started' as const, taskId: 'explicit-task' })); + const tool = new LocalBashTool('/tmp/test-workspace', { startBackground }, { mode: 'off' }); + await tool.execute(context, { command: 'echo ok', run_in_background: true }); + expect(startBackground.mock.calls[0]?.[1]).not.toHaveProperty('timeout'); + expect(DEFAULT_BACKGROUND_BASH_MAX_RUN_MS).toBe(3_600_000); + expect(resolveBackgroundBashMaxRunMs(undefined)).toBe(3_600_000); + expect(resolveBackgroundBashMaxRunMs(7)).toBe(3_600_000); + expect(resolveBackgroundBashMaxRunMs(7200)).toBe(7_200_000); + expect(resolveLocalBashTiming(7, 'explicit_background').commandTimeoutSeconds).toBe(7); + }); + + it('retains direct foreground limits and rejects invalid explicit background timeouts', () => { + expect(resolveLocalBashTiming(undefined, 'direct_foreground').commandTimeoutSeconds).toBe(120); + expect(resolveLocalBashTiming(7200, 'direct_foreground').commandTimeoutSeconds).toBe(300); + for (const mode of ['direct_foreground', 'explicit_background'] as const) { + for (const timeout of [0, -5, NaN, Infinity]) expect(() => resolveLocalBashTiming(timeout, mode)).toThrow(); + } + }); + + it('advertises the same limits to the model', () => { + expect(LocalBashToolDef.schema.properties.timeout.description).toContain('3600s'); + expect(createLocalBashToolDefinition({ background: true, shell: 'bash' }).schema.properties.timeout.description).toContain('1-hour'); + }); +}); diff --git a/packages/agent-tools/test/desktop/local-memory.test.ts b/packages/agent-tools/test/desktop/local-memory.test.ts new file mode 100644 index 00000000..a98f969d --- /dev/null +++ b/packages/agent-tools/test/desktop/local-memory.test.ts @@ -0,0 +1,63 @@ +import { describe, expect, it, vi } from 'vitest'; + +import { LocalMemoryTool } from '../../src/desktop/local-memory.js'; +import { MEMORY_TOOL_OUTPUT_MAX_BYTES } from '@mavis/shared/memory-limits'; +import type { LocalRuntimeToolContext } from '../../src/desktop/types.js'; + +const context: LocalRuntimeToolContext = { + sessionId: 'memory-session', + turnId: 'turn', + toolCallId: 'call', + agentName: 'mavis', +}; + +describe('LocalMemoryTool output limit', () => { + it.each(['user', 'main', 'topic'] as const)( + 'bounds oversized %s reads in both model-facing fields', + async (target) => { + const text = 'FIRST_FACT\n' + '记忆🙂'.repeat(20_000) + '\nLATEST_FACT'; + const execute = vi.fn(async () => ({ text, details: { ok: true, target } })); + const tool = new LocalMemoryTool({ execute }); + const input = { target, operation: 'read' as const, topicName: 'work' }; + const result = await tool.execute(context, input); + + expect(execute).toHaveBeenCalledWith(context, input, undefined); + expect(Buffer.byteLength(result.text!)).toBeLessThanOrEqual(MEMORY_TOOL_OUTPUT_MAX_BYTES); + expect(result.text).toContain('FIRST_FACT'); + expect(result.text).toContain('LATEST_FACT'); + expect(result.text).toContain('Memory output truncated'); + expect(Buffer.from(result.text!).toString('utf8')).toBe(result.text); + expect(result.content).toEqual([{ type: 'text', text: result.text }]); + expect(result.details).toMatchObject({ + kind: 'memory', + ok: true, + target, + desktop_output_truncation: { truncated: true, original_bytes: Buffer.byteLength(text) }, + }); + }, + ); + + it('bounds search results as well as direct reads', async () => { + const tool = new LocalMemoryTool({ + execute: async () => ({ text: JSON.stringify([{ line: '记'.repeat(80_000) }]) }), + }); + const result = await tool.execute(context, { + target: 'main', + operation: 'search', + query: '记', + }); + expect(Buffer.byteLength(result.text!)).toBeLessThanOrEqual(MEMORY_TOOL_OUTPUT_MAX_BYTES); + expect(result.details?.desktop_output_truncation).toMatchObject({ truncated: true }); + }); + + it.each(['short memory', 'a'.repeat(MEMORY_TOOL_OUTPUT_MAX_BYTES)])( + 'preserves an output within the budget', + async (text) => { + const tool = new LocalMemoryTool({ execute: async () => ({ text, details: { ok: true } }) }); + const result = await tool.execute(context, { target: 'main', operation: 'read' }); + expect(result.text).toBe(text); + expect(result.content).toEqual([{ type: 'text', text }]); + expect(result.details).toEqual({ kind: 'memory', ok: true }); + }, + ); +}); diff --git a/packages/config/src/config.ts b/packages/config/src/config.ts index 75bccea1..3501ec91 100644 --- a/packages/config/src/config.ts +++ b/packages/config/src/config.ts @@ -1544,6 +1544,28 @@ const MINIMAX_MODELS: Record = { }, capabilities: MINIMAX_M3_FILE_API_CAPABILITIES, }, + "MiniMax-M3.1-Flash-Preview": { + name: "M3.1-Flash-Preview", + attachment: true, + reasoning: true, + tool_call: true, + temperature: true, + modalities: { input: ["text", "image", "video"], output: ["text"] }, + limit: { context: 512000, output: 128000 }, + contextWindowOptions: [512000, 1000000], + contextWindowOptionHints: { "1000000": "higher_usage" }, + options: { reasoningSummary: "auto" }, + thinking: { + effortOptions: ["default", "low", "medium", "high", "xhigh", "max"], + defaultEffort: "default", + }, + thinking_config: { mode: "forced_on" }, + variants: { + "none-thinking": { thinking: { type: "disabled" } }, + thinking: { thinking: { type: "adaptive" } }, + }, + capabilities: MINIMAX_M3_FILE_API_CAPABILITIES, + }, "MiniMax-M2.7-highspeed": { name: "MiniMax-M2.7-highspeed", attachment: false, diff --git a/packages/config/src/tui-config.ts b/packages/config/src/tui-config.ts index 13838c51..b04ecd98 100644 --- a/packages/config/src/tui-config.ts +++ b/packages/config/src/tui-config.ts @@ -32,10 +32,14 @@ export interface TuiCustomStatusLineConfig { } export interface TuiConfig { + /** Ordered terminal title items. Null or an empty list disables title updates. */ + terminalTitle?: readonly string[] | null; /** Terminal notification policy. Unknown focus falls back to notifying. */ notifications?: { when?: 'unfocused' | 'always' | 'never'; method?: 'auto' | 'osc9' | 'osc777' | 'bel'; + /** Omit to enable all supported notification events; an empty list disables them. */ + events?: readonly string[]; }; /** Show contextual Tips in the idle composer header. Defaults to true. */ showTips?: boolean; @@ -75,12 +79,17 @@ export function parseTuiConfig(raw: Record): TuiConfig { typeof rawNotifications === 'object' && !Array.isArray(rawNotifications) ) { - const { when, method } = rawNotifications as Record; + const { when, method, events } = rawNotifications as Record; notifications = { ...(when === 'unfocused' || when === 'always' || when === 'never' ? { when } : {}), ...(method === 'auto' || method === 'osc9' || method === 'osc777' || method === 'bel' ? { method } : {}), + ...(Array.isArray(events) + ? { + events: events.filter((event): event is string => typeof event === 'string'), + } + : {}), }; } const rawStatusLine = tui.statusLine; @@ -92,6 +101,16 @@ export function parseTuiConfig(raw: Record): TuiConfig { : undefined; const customStatusLine = parseTuiCustomStatusLineConfig(tui.customStatusLine); return { + ...(tui.terminalTitle === null + ? { terminalTitle: null } + : Array.isArray(tui.terminalTitle) + ? { + terminalTitle: tui.terminalTitle + .filter((item): item is string => typeof item === 'string') + .map((item) => item.trim()) + .filter(Boolean), + } + : {}), ...(typeof tui.showTips === 'boolean' ? { showTips: tui.showTips } : {}), ...(notifications ? { notifications } : {}), ...(statusLine ? { statusLine } : {}), diff --git a/packages/config/test/builtin-model-fallback.test.ts b/packages/config/test/builtin-model-fallback.test.ts new file mode 100644 index 00000000..06ac508e --- /dev/null +++ b/packages/config/test/builtin-model-fallback.test.ts @@ -0,0 +1,100 @@ +import fs from "node:fs"; +import os from "node:os"; +import { join } from "node:path"; +import yaml from "js-yaml"; +import { afterEach, beforeEach, describe, expect, it, vi } from "vitest"; +import { + DEFAULT_MODEL_PRESETS, + getConfig, + resetConfig, +} from "../src/config.js"; +import { resolveModelAvailability } from "../src/model-availability.js"; + +const modelId = "MiniMax-M3.1-Flash-Preview"; +let dataDir: string; + +beforeEach(() => { + dataDir = fs.mkdtempSync(join(os.tmpdir(), "mcode-model-fallback-")); + vi.stubEnv("MINIMAX_DATA_DIR", dataDir); + vi.stubEnv("__MAVIS_RUNTIME_DATA_DIR", dataDir); + vi.stubEnv("__MAVIS_RUNTIME_MANAGED", "1"); + vi.stubEnv("__MAVIS_RUNTIME_DISABLE_GIT_AUTO_CONFIG", "1"); + vi.stubEnv("MAVIS_REGION", "en"); + vi.stubEnv("MAVIS_BUILD_ENV", "prod"); + resetConfig(); +}); + +afterEach(() => { + vi.unstubAllEnvs(); + resetConfig(); + fs.rmSync(dataDir, { recursive: true, force: true }); +}); + +describe("built-in model fallback", () => { + it("seeds M3.1 Flash Preview before a remote snapshot exists and preserves it on reload", () => { + const config = getConfig(); + const model = config.provider.minimax?.models?.[modelId]; + expect(model).toMatchObject({ + name: "M3.1-Flash-Preview", + attachment: true, + reasoning: true, + tool_call: true, + modalities: { input: ["text", "image", "video"], output: ["text"] }, + limit: { context: 512000, output: 128000 }, + contextWindowOptions: [512000, 1000000], + contextWindowOptionHints: { "1000000": "higher_usage" }, + thinking: { + effortOptions: ["default", "low", "medium", "high", "xhigh", "max"], + defaultEffort: "default", + }, + thinking_config: { mode: "forced_on" }, + variants: { thinking: { thinking: { type: "adaptive" } } }, + capabilities: { + support_files_api: true, + files_api_upload_endpoint: "/v1/files/upload", + max_image_bytes_inline: 10_485_760, + max_video_bytes_inline: 52_428_800, + max_request_body_bytes: 67_108_864, + max_attachments_count: 4, + }, + }); + expect(config.defaultModel).toBe("minimax/MiniMax-M3"); + resetConfig(); + expect(getConfig().provider.minimax?.models?.[modelId]).toEqual(model); + }); + + it("makes the fallback model available in every managed preset", () => { + for (const preset of Object.keys(DEFAULT_MODEL_PRESETS) as Array< + keyof typeof DEFAULT_MODEL_PRESETS + >) { + expect( + resolveModelAvailability({ + config: DEFAULT_MODEL_PRESETS[preset], + providerId: "minimax", + modelId, + preset, + source: "explicit_request", + }), + ).toEqual({ available: true, route: "managed_token_plan" }); + } + }); + + it("keeps an existing managed snapshot authoritative instead of adding fallback models", () => { + const models = { "MiniMax-M3": { name: "MiniMax-M3" } }; + fs.writeFileSync( + join(dataDir, "config.yaml"), + yaml.dump({ + provider: { + minimax: { + options: { + authMode: "managed-login", + baseURL: "https://agent.minimax.io/mavis/api/v1/llm/v1", + }, + models, + }, + }, + }), + ); + expect(getConfig().provider.minimax?.models).toEqual(models); + }); +}); diff --git a/packages/local-runtime-v2/assets/agents/_default/prompt-base-all.md.hbs b/packages/local-runtime-v2/assets/agents/_default/prompt-base-all.md.hbs index 37ed2a9e..b3adedcd 100644 --- a/packages/local-runtime-v2/assets/agents/_default/prompt-base-all.md.hbs +++ b/packages/local-runtime-v2/assets/agents/_default/prompt-base-all.md.hbs @@ -192,12 +192,7 @@ When the user's latest message contains the exact phrase `Prompt 灰度验证`, ### Non-interactive shell - Your shell is **non-interactive** — no TTY, no stdin, no prompt. Commands that wait for stdin - or require a terminal UI will hang forever. - -### Bash timeout - -- Bash `timeout` is measured in seconds. Foreground commands default to 120s, are capped at 300s, and yield to the same managed background process after 15s so the turn is not blocked. Start builds, tests, installs, downloads, servers, or diagnostics expected to take longer as background tasks up front; inspect their incremental output with the task tools. - Do not use excessively large timeouts to mask hung commands, and do not rerun a command after it yields—the returned task is the original process. + or require a terminal UI may fail or wait until the command timeout. Use non-interactive flags. ```bash # BAD — interactive commands hang diff --git a/packages/local-runtime-v2/assets/agents/_default/prompt-base-windows.md b/packages/local-runtime-v2/assets/agents/_default/prompt-base-windows.md index 79c46841..6e4bb0d8 100644 --- a/packages/local-runtime-v2/assets/agents/_default/prompt-base-windows.md +++ b/packages/local-runtime-v2/assets/agents/_default/prompt-base-windows.md @@ -1,45 +1,6 @@ ## Windows Shell Constraints -### PowerShell only - -- Use **PowerShell syntax only**. Do NOT use legacy DOS / `cmd.exe` commands (`cmd`, `cmd /c`, - `dir`, `type`, `copy`, `move`, `del`, `erase`, `rd`, `rmdir`, etc.). Use full PowerShell - cmdlets instead (`Get-ChildItem`, `Copy-Item`, `Move-Item`, `New-Item`). -- The bash tool already runs the command through the selected Windows shell. Do not wrap ordinary - bash-tool commands in `powershell -Command`. If you truly must launch a nested `pwsh`/PowerShell - process, prefer `pwsh -NoProfile -NonInteractive -Command '