From 274b9a7f28248344c8a4c3a33fc6e36de45d706b Mon Sep 17 00:00:00 2001 From: Ronit Date: Thu, 14 May 2026 16:20:41 +0530 Subject: [PATCH 01/11] =?UTF-8?q?docs(kernel):=20PR-C=20spec=20=E2=80=94?= =?UTF-8?q?=20task=20management=20design?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Per-thread task list with single batchable manage_tasks tool + context block refreshed per-step via streamText prepareStep. First of three slices: Tasks (this PR) → Sub-agents (PR-D) → Streaming (PR-E). Key decisions: - Tasks live AFTER b2 cache marker (in dynamic suffix), not in cached prefix - prepareStep callback owns tasks block render — never buildContextMessages - Soft 'one in_progress' rule via system prompt, NOT DB constraint (preserves PR-D parallel sub-agent option) - result field optional but encouraged via system prompt - 5-state enum: pending/in_progress/complete/failed/cancelled - No get_task read tool (context block + tool_result history covers retrieval) Adjacent provider-option additions in §9: toolStreaming, effort=xhigh, sendReasoning, thinking={adaptive,summarized}, explicit cacheControl.ttl=5m. Co-Authored-By: Claude Opus 4.7 (1M context) --- .../2026-05-14-task-management-design.md | 705 ++++++++++++++++++ 1 file changed, 705 insertions(+) create mode 100644 docs/superpowers/specs/2026-05-14-task-management-design.md diff --git a/docs/superpowers/specs/2026-05-14-task-management-design.md b/docs/superpowers/specs/2026-05-14-task-management-design.md new file mode 100644 index 0000000..8129d7d --- /dev/null +++ b/docs/superpowers/specs/2026-05-14-task-management-design.md @@ -0,0 +1,705 @@ +# Task Management Design (PR-C) + +> **Status:** Draft — design accepted by user 2026-05-14 +> **Author:** Yoke + Ronit (brainstorm) +> **Slice:** First of three. PR-C (Tasks) → PR-D (Sub-agents) → PR-E (Streaming) + +--- + +## Goal + +Give the main agent a first-class task list — DB-backed, per-thread, batchable via a single `manage_tasks` write tool, surfaced through a `` context block refreshed at every step boundary inside the `streamText` react loop. Establishes the substrate for PR-D (sub-agent spawning) without including it. + +## Non-goals + +These belong to later slices and are explicitly out of scope for this PR: + +- **PR-D territory:** `spawn_sub_agent` tool; sub-agent runtime; sub-agent message persistence; auto-population of `task.result` from sub-agent's final assistant message; retry-with-guidance flow; parallel sub-agent execution. +- **PR-E territory:** TUI rendering of `` as a sidebar/panel; streaming sub-agent events to TUI. +- **Indefinitely deferred (YAGNI):** Nested tasks (`parentTaskId`); task priority / explicit ordering; `get_task(id)` read tool; task templates; cross-thread tasks. + +--- + +## §1 — Architecture overview + +A single thin slice of new state plus one tool and one context block. The structural diagram: + +``` +TUI Kernel DO + │ │ + │ message ───────────────────────────────────────────────► │ + │ runChatTurn(...) + │ │ + │ ┌───────────────┴────────────────┐ + │ │ buildContextMessages │ + │ │ ├─ │ + │ │ ├─ │ + │ │ └─ [b2 cache] │ + │ └────────────────────────────────┘ + │ │ + │ ┌───────────────┴─────────────┐ + │ │ streamText │ + │ │ ├─ prepareStep callback │ + │ │ │ └─ rebuild │ + │ │ │ inject post-b2 │ + │ │ └─ tools: manage_tasks NEW │ + │ └─────────────────────────────┘ + │ │ + │ stream events ◄──────────────────────────────────────────│ + │ ▼ + │ SQLite (DO): + │ new `task` table +``` + +**Three new pieces:** + +1. **`task` table** in `@agent-os/models` — per-thread, references `thread.id` and `run.id` (which run created it). Five-state status enum: `pending | in_progress | complete | failed | cancelled`. + +2. **`manage_tasks` tool** — single batchable write tool taking `{ ops: [...] }`. Returns the full updated task list in its response so the model sees canonical state without paying for a separate read. Wrapped via `wrappedTool({ touchesFS: false })` — writes audit rows but doesn't open a per-call git worktree. + +3. **`` context block** — rendered post-b2, refreshed at every step boundary via the `streamText` `prepareStep` callback. Omitted entirely when the thread has zero tasks. + +**Behavior contract:** + +- Tasks are scoped per-thread (Claude Code TodoWrite parity). They persist across user turns within one conversation; completed entries stay visible as audit trail. Cascade-delete with the thread. +- Soft "one task in_progress at a time" rule enforced via system-prompt guidance only — NOT via DB constraint. Preserves PR-D's option to run sub-agents in parallel later. +- System prompt gets a new `` section teaching when to plan vs just-do, lifecycle, description quality, and discipline. + +--- + +## §2 — Data model + +**New file:** `packages/models/drizzle/schemas/tasks.ts` + +```ts +import { InferSelectModel } from "drizzle-orm"; +import { index, integer, sqliteTable, text } from "drizzle-orm/sqlite-core"; + +import { run } from "./runs"; +import { thread } from "./threads"; + +export const task = sqliteTable( + "task", + { + id: text().primaryKey(), // ULID + threadId: text() + .notNull() + .references(() => thread.id, { onDelete: "cascade" }), + // The run that CREATED this task (manage_tasks{action:"create"} call). + // Updates from later runs do NOT change this — preserves "who planned this". + runId: text() + .notNull() + .references(() => run.id, { onDelete: "cascade" }), + name: text().notNull(), // user-visible label + description: text().notNull(), // full plan: deliverables, deps, expected output + status: text({ + enum: ["pending", "in_progress", "complete", "failed", "cancelled"] + }) + .notNull() + .default("pending"), + // Set on transition to a terminal status: + // complete → the actual output/answer + // failed → the reason the agent gave up + // cancelled → why it was abandoned + result: text(), + createdAt: integer({ mode: "timestamp_ms" }) + .notNull() + .$defaultFn(() => new Date()), + updatedAt: integer({ mode: "timestamp_ms" }) + .notNull() + .$defaultFn(() => new Date()) + }, + (table) => [ + index("idx_task_thread_created").on(table.threadId, table.createdAt) + ] +); + +export type Task = InferSelectModel; +``` + +**Field rationale:** + +| Field | Choice | Why | +|---|---|---| +| `id` | ULID (text PK) | Matches `message.id` / `run.id` convention. Sortable by creation time as tie-breaker. | +| `threadId` | FK to `thread`, cascade delete | Standard agent-os pattern. Deleting a thread cleans up everything in it. | +| `runId` | FK to creating run, cascade | Analytics value ("planned in run X"). Updates from later runs don't change it — preserves authorship. | +| `name` / `description` | Both `notNull` | Forcing description-on-create prevents "task #4: do the thing" planning that's useless on retry. | +| `status` enum | 5 values | `pending → in_progress → {complete | failed | cancelled}`. Failed vs cancelled distinction: failed = tried and gave up; cancelled = decided not to attempt. Useful signal. | +| `result` | nullable text | NULL for pending/in_progress; set on terminal transitions. | +| `updatedAt` | `notNull`, $defaultFn on insert | Tool code explicitly sets it on every update (don't lean on Drizzle's `$onUpdate` since it only fires through `.update()` calls — explicit is safer). | +| Index `(threadId, createdAt)` | Single composite | Only access pattern: "fetch all tasks for this thread ordered by creation." | + +**Deliberately omitted:** + +- `parentTaskId` — flat tasks, no nesting. Add in PR-D if sub-agents reveal a need. +- `priority` — YAGNI. Agent orders by creation sequence (which it controls). +- `archivedAt` — per-thread render-all means no archive concept. +- `attemptCount` / sub-agent result messages — strictly PR-D territory. + +**Migration:** new `wrangler d1 migrations` file under the `KERNEL_DB` namespace, following the pattern of existing migrations. + +```sql +CREATE TABLE task ( + id TEXT PRIMARY KEY NOT NULL, + "threadId" TEXT NOT NULL REFERENCES thread(id) ON DELETE CASCADE, + "runId" TEXT NOT NULL REFERENCES run(id) ON DELETE CASCADE, + name TEXT NOT NULL, + description TEXT NOT NULL, + status TEXT NOT NULL DEFAULT 'pending', + result TEXT, + "createdAt" INTEGER NOT NULL, + "updatedAt" INTEGER NOT NULL +); + +CREATE INDEX idx_task_thread_created ON task("threadId", "createdAt"); +``` + +Quoted camelCase column names per PR-A precedent (memory: "Drizzle uses JS field names verbatim → quoted camelCase in raw SQL"). + +**Type export:** `Task` flat-exported via `drizzle/schemas/index.ts → drizzle/index.ts → models/index.ts`, matching the `Thread` / `Message` precedent (memory: `feedback_schema_types_flat_exports`). + +--- + +## §3 — The `manage_tasks` tool + +**New file:** `apps/kernel/src/tools/tasks.ts` + +Wraps via `wrappedTool({ touchesFS: false, needsApproval: false })` — writes audit rows but doesn't open a per-call worktree (consistent with web-search / sessions tools). + +### Input schema (LLM-facing) + +```ts +{ + ops: Array< + | { action: "create", name: string, description: string } + | { + action: "update", + id: string, + name?: string, + description?: string, + status?: "pending" | "in_progress" | "complete" | "failed" | "cancelled", + result?: string + } + > +} +``` + +The LLM never sees `threadId` or `runId`. They are injected server-side from `perTurn.threadId` and `perTurn.runId` (captured in the kernel's `runChatTurn` closure at turn start). + +### Per-op behavior + +| `action` | LLM-provided | Server-injected | Effect | +|---|---|---|---| +| `create` | `name`, `description` | `id = ulid()`, `threadId`, `runId`, `status = "pending"`, timestamps | INSERT new row. | +| `update` | `id`, plus ≥1 of name/description/status/result | (verifies ownership) | Fetch existing task. If `threadId ≠ ctx.threadId`, throw. Else UPDATE the supplied fields + `updatedAt = now`. | + +**Status transitions** are unrestricted at the tool level (any → any allowed). System prompt guides the expected flow; the tool doesn't gate. + +**No `delete` action.** Terminal `cancelled` status preserves history. Lossy deletes would hurt PR-D's retry-with-context flow. + +### Execute path (pseudocode) + +```ts +execute: async (args, ctx) => { + // ctx: WrappedToolContext = PerTurnContext & { callId } + for (const op of args.ops) { + if (op.action === "create") { + await ctx.db.insert(schema.task).values({ + id: ulid(), + threadId: ctx.threadId, + runId: ctx.runId, + name: op.name, + description: op.description + }); + } else if (op.action === "update") { + const existing = await ctx.db.query.task.findFirst({ + where: eq(schema.task.id, op.id) + }); + if (!existing) throw new Error(`Task ${op.id} not found`); + if (existing.threadId !== ctx.threadId) { + throw new Error(`Task ${op.id} belongs to a different thread`); + } + await ctx.db.update(schema.task) + .set({ + ...(op.name !== undefined && { name: op.name }), + ...(op.description !== undefined && { description: op.description }), + ...(op.status !== undefined && { status: op.status }), + ...(op.result !== undefined && { result: op.result }), + updatedAt: new Date() + }) + .where(eq(schema.task.id, op.id)); + } + } + + // Return full updated task list sorted by createdAt ASC + const tasks = await ctx.db.query.task.findMany({ + where: eq(schema.task.threadId, ctx.threadId), + orderBy: [asc(schema.task.createdAt)] + }); + return { + tasks: tasks.map(t => ({ + id: t.id, + name: t.name, + description: t.description, + status: t.status, + result: t.result, + createdAt: t.createdAt.toISOString(), + updatedAt: t.updatedAt.toISOString() + })) + }; +} +``` + +### Return value + +Full state, full descriptions, sorted by `createdAt ASC`. The agent that just created Task X needs to see its ID right away, and the model benefits from canonical post-update state without paying for a separate read. ~500–2000 tokens per call, lives once in conversation history. + +### Error handling + +- Schema validation → AI SDK rejects pre-execute (model retries with corrected args). +- `update` with non-existent `id` → throw `Task not found`. +- `update` against `threadId` mismatch → throw `Task belongs to a different thread` (defensive). +- Empty `ops` array → returns current state with no mutations (cheap "refresh"). +- All errors propagate through `wrappedTool` → audit row's `status = "error"`, model sees tool error and can retry. + +### Registration + +Append `...buildTasksTools(perTurn)` to the `buildTools` spread in `apps/kernel/src/tools/index.ts`. Order doesn't matter — `withTailCache` shifts the b1 marker to whatever is last. + +--- + +## §4 — `` context block + `prepareStep` wiring + +Two pieces with a clean ownership split. The tasks block is owned **entirely** by `prepareStep`, not by `buildContextMessages`. + +### Piece A: `apps/kernel/src/agent/tasks-block.ts` (renderer) + +Pure SQL fetch via Drizzle. Returns `string | null` (null when zero tasks → block omitted). + +**Render rules:** + +| Status | Description shown | Result shown | +|---|---|---| +| `in_progress` | **Full** (active focus needs the plan) | — | +| `pending` | Trimmed to ~120 chars + "…" if truncated | — | +| `complete` | Trimmed to ~80 chars | **Full** result text | +| `failed` | Trimmed to ~80 chars | **Full** result (i.e. why it failed) | +| `cancelled` | Trimmed to ~80 chars | **Full** result, or omitted if NULL | + +**Example output:** + +``` + +Plan and track multi-step work here. Use manage_tasks to create/update tasks. +Convention: one task in_progress at a time. Mark complete/failed/cancelled +before moving on. Skip task planning for trivial single-step requests. + +- 01HQ8X... [complete] "Read user's uploaded PDF" + plan: Extract text from uploads/spec.pdf via process_attachment + result: 4200 words extracted; 3 sections found (intro, design, risks) +- 01HQ8Z... [in_progress] "Summarize section: design" + plan: Pull 200-word summary of the design section. Should cover the cache + breakpoint scheme, context block ordering, and tool wrapping pattern. +- 01HQA0... [pending] "Summarize section: risks" + plan: Pull 200-word summary of risks section. Note any cross-references… +- 01HQA1... [cancelled] "Generate PDF cover page" + plan: User asked then said no + result: User reversed direction mid-turn + +``` + +**Truncation rationale:** full descriptions are still recoverable via the agent's conversation history (every `manage_tasks` `tool_result` contains the un-trimmed task). Truncation in the context block is a per-turn render optimization, not lossy storage. If the description gets compacted out of conv history many turns later, the agent has `sessions_search` and `session_info` as recovery paths against the current thread. + +**Failure mode:** SQL throw → render error placeholder so the block still ships and the agent isn't blind: + +``` + +(error loading task list — manage_tasks may still work, try it) + +``` + +Mirrors `sessions-recent-block.ts`'s error handling. + +### Piece B: `prepareStep` callback in `turn.ts` + +`buildContextMessages` is NOT modified. Stable context blocks ship from there as before. The tasks block is added by `prepareStep` on every step boundary. + +**Why per-step refresh:** the AI SDK's `streamText` runs the entire react loop (LLM → tools → LLM → tools …) internally. Context messages built once at turn start would be **frozen for the whole turn**. If sub-agents (PR-D) mutate task state mid-turn, the model wouldn't see the updates via the context block — it would only see them through `tool_result` history. The `prepareStep` callback (stable AI SDK API as of v5) lets us rebuild messages before each step's LLM call. + +**Implementation shape:** + +```ts +// Inside runChatTurn, after buildContextMessages(...) returns: +const stableBlockCount = contextMessages.length; + +const result = streamText({ + system: { /* existing */ }, + tools: buildTools(perTurn), + messages: [...contextMessages, ...modelMessages], + // ... existing options (stopWhen, abortSignal, providerOptions, etc.) + prepareStep: async ({ stepNumber, messages }) => { + const working = [...messages]; + + // On step 1+, strip the prior step's tasks block. It always lives at + // index `stableBlockCount` IF buildTasksBlock returned non-null last time. + // Detected by role + literal "" tag prefix. + if (stepNumber > 0) { + const candidate = working[stableBlockCount]; + if ( + candidate?.role === "user" && + typeof candidate.content === "string" && + candidate.content.startsWith("") + ) { + working.splice(stableBlockCount, 1); + } + } + + const tasksBlock = await buildTasksBlock({ db: kernel.db, threadId }); + if (tasksBlock !== null) { + working.splice(stableBlockCount, 0, { + role: "user" as const, + content: tasksBlock + }); + } + + return { messages: working }; + }, + onFinish: /* existing */ +}); +``` + +**Splice invariant:** `prepareStep` must only ever touch indices `≥ stableBlockCount`. Touching earlier indices would mutate the b2-cached prefix and bust caching. The current splice pattern preserves this by construction. + +**Future blocks may want per-step refresh too** (``, ``, `` — agent could write to any of them mid-turn). Out of scope for PR-C; deliberately limited to `` to minimize cache impact (those blocks currently live in the b2-cached prefix; moving them post-b2 is a separate cost analysis). + +--- + +## §5 — System prompt: `` section + +Appended to `apps/kernel/src/agent/system-prompt.ts` near `` (same tier — how-to-use-this-tool guidance). + +``` + +For multi-step requests, plan and track your work with manage_tasks. The + context block (when present) is your source of truth for task state. +It refreshes at every step boundary so it never lags behind your tool calls. + +WHEN TO PLAN VS JUST DO +The question to ask yourself: is the work substantial, or trivial? +- "Read this file and tell me what it says" → just do it. One tool call. +- "What time is it in Tokyo" → just do it. +- "Research X, write Y, post Y to slack" → plan tasks first. +- "Refactor this module" if it's one file → just do it. +- "Refactor this module" if it spans many files → plan tasks first. + +Planning costs tool calls. Only plan when the work has independent steps you +want to track, retry, or summarize. When the user signals depth ("deep dive", +"thoroughly", "do this properly"), lean toward planning even if it looks small. + +LIFECYCLE +- create: batch all tasks up front with name + description. + name: short user-visible label. + description: full plan — what to do, what to produce, any constraints. + Be specific. Your future self (next step) reads this and must act on it + without re-asking the user. +- update to in_progress before you start work on a task. +- update to complete when done. Set result to 1–3 sentences of what was + produced or accomplished. Result is optional but strongly preferred — it's + how future turns see "this got done, here's what it was." +- update to failed if you tried and gave up; set result to the reason. +- update to cancelled if you decided not to do this task; set result to why. + +DESCRIPTION QUALITY +Bad: "Refactor parser" +Good: "Refactor parser: split parse() into tokenize() and parseTokens(). + Preserve current public API. Update the 3 call sites in handlers.ts. + Aim under 80 lines per function." + +A vague description means you (or a future step) won't know what "done" +looks like. + +DISCIPLINE +- One task in_progress at a time. Mark the current one terminal before + starting the next. +- Update status the moment state changes. Stale in_progress tasks make the + user think you're confused. +- Tasks persist for the whole thread. Completed ones stay visible as your + audit trail. + +DON'T +- Don't plan tasks for trivial requests. +- Don't create a task and mark it complete in the same call (no fake + productivity). +- Don't recreate a failed task with the same description. Either update its + status for a retry, or create a NEW task whose description includes what + went wrong and what's different. + +``` + +**Token cost:** ~370 tokens added to the system prompt. Cached in b1; paid once per fresh cache, then free. + +**Why this length / shape (lifted from Dimension AI's prompt):** + +- **Decision framing** — "is the work substantial, or trivial?" beats "when user asks for multi-step." Sharper test. +- **Concrete one-liner examples** for the boundary (read file → no plan; research+write+post → plan; refactor-one vs refactor-many). Compressed from Dimension's three-example block. +- **Signal phrases** ("deep dive", "thoroughly") — direct lift. +- **Good/bad description pair** — single contrast (Dimension uses two; ours hits 80% of the value at 20% of the cost). +- **Overhead framing** — "Planning costs tool calls" makes the trade-off explicit. + +**Deliberately omitted:** sub-agent / delegation framing (PR-D scope); cross-task references via `GET_TASK(id)` (no such tool exists in PR-C); multiple full example scenarios (token bloat). + +--- + +## §6 — Cache strategy (consolidated) + +### Breakpoint map post-PR-C + +``` +SYSTEM_PROMPT [b1 on last tool def] ◄── tools cache + ├─ + └─ ... [b1] +───────────────────────────────────────────── +STABLE CONTEXT BLOCKS (built once at turn start) + ├─ conditional + ├─ + ├─ "stable-ish" (turn-boundary refresh) + ├─ "stable-ish" + ├─ stable across turns + ├─ stable across turns (current excluded) + └─ [b2] stable per day in TZ ◄── stable-context cache +───────────────────────────────────────────── +DYNAMIC SUFFIX (rebuilt per step via prepareStep) + └─ NEW — re-rendered every step +───────────────────────────────────────────── +CONVERSATION + ├─ msg_1, msg_2, ..., msg_n-1 + └─ msg_n [b4] last conv msg ◄── conversation cache + (b3 [compaction] applied conditionally inside conv) +``` + +### Per-breakpoint behavior + +| Breakpoint | Hits when | Misses when | +|---|---|---| +| **b1** (tools) | Always within a deploy. Tool set is static per turn. | New deploy ships a tool change. | +| **b2** (stable context) | Same thread, same day, no memory writes, no parallel TUI sessions touching other threads. | Memory file written. Day boundary crossed. Files/sandboxes block changed. | +| **b3** (compaction) | Conditional. Inherited unchanged from PR-B. | Same as PR-B. | +| **b4** (last conv msg) | Across turns: only when full prefix byte-matches. **Within turn: misses every step that mutated tasks** because the bytes between b2 and b4 now include the refreshed `` block. | Most within-turn steps land in "miss" territory. Tolerable. | + +### Worked example — the "research + document + ppt + email" workflow + +Multi-step turn, single user prompt, sub-agents land in PR-D (here approximated as agent doing the work inline): + +| Step | Model sees | Cache behavior | +|---|---|---| +| 0 | Stable + tasks(empty/null) + conv | b1 hit, b2 hit, b4 cold → pay ~5K stable prefix one-time-this-turn, ~300 for empty conv | +| 1 | Stable + tasks(4 pending) + conv + step0 results | b1 hit, b2 hit, b4 misses → pay tasks(4) + conv + step0 fresh | +| 2 | Stable + tasks(t1 in_progress) + conv + step0-1 results | b1 hit, b2 hit, b4 misses → pay updated tasks + step1 results | +| ... | similar | b2 keeps hitting; everything past b2 paid fresh per step | + +**Net per-turn token math (vs putting tasks INSIDE the b2 prefix):** + +- **Save:** b2 stays hit across all steps. System + tools + 6 stable context blocks stay cached. ~5K–8K tokens cached across the turn. +- **Spend:** ~300 tokens of `` rebuild per step × N steps. For 5 steps = ~1.5K extra tokens. +- **Net win:** ~3.5K–6.5K tokens per multi-step turn. + +If tasks were inside b2 prefix (bad design), b2 would miss every step that touched tasks → ~5K × N steps in re-caching. PR-C's post-b2 placement is the right call. + +### Invariants to preserve + +- `prepareStep` must NOT mutate messages at indices `< stableBlockCount`. The current splice pattern enforces this by construction. +- No new breakpoints introduced. Anthropic's 4-breakpoint cap (b1/b2/b3/b4) is preserved. + +--- + +## §7 — Wiring + migration (file-by-file) + +### New files (4) + +| Path | Responsibility | +|---|---| +| `packages/models/drizzle/schemas/tasks.ts` | Drizzle table definition + `Task` flat type export (§2) | +| `packages/models/drizzle/migrations/00XX_create_task.sql` | D1 migration: `CREATE TABLE task` + index | +| `apps/kernel/src/tools/tasks.ts` | `buildTasksTools(perTurn)` exporting `{ manage_tasks }` via `wrappedTool({ touchesFS: false })` (§3) | +| `apps/kernel/src/agent/tasks-block.ts` | `buildTasksBlock({ db, threadId })` → `string \| null` (§4) | + +### Modified files + +| Path | Change | +|---|---| +| `packages/models/drizzle/schemas/index.ts` | Add `export * from "./tasks"` | +| `packages/models/index.ts` (flat exports) | Confirm `Task` ends up at top level alongside `Thread`, `Message`, `Run` | +| `apps/kernel/src/tools/index.ts` | Add `...buildTasksTools(perTurn)` to the `buildTools` spread | +| `apps/kernel/src/agent/system-prompt.ts` | Append `` section near `` (§5) | +| `apps/kernel/src/turn.ts` | Capture `stableBlockCount = contextMessages.length`; add `prepareStep` callback to `streamText({...})`; add Anthropic provider options from §9 | + +### Files NOT modified + +- `apps/kernel/src/agent/context-messages.ts` — by design. Tasks block is owned by `prepareStep`. Two render sites would diverge. +- TUI / CLI — no UI-side changes in this PR. + +### Migration SQL + +```sql +CREATE TABLE task ( + id TEXT PRIMARY KEY NOT NULL, + "threadId" TEXT NOT NULL REFERENCES thread(id) ON DELETE CASCADE, + "runId" TEXT NOT NULL REFERENCES run(id) ON DELETE CASCADE, + name TEXT NOT NULL, + description TEXT NOT NULL, + status TEXT NOT NULL DEFAULT 'pending', + result TEXT, + "createdAt" INTEGER NOT NULL, + "updatedAt" INTEGER NOT NULL +); + +CREATE INDEX idx_task_thread_created ON task("threadId", "createdAt"); +``` + +Column quoting follows the PR-A precedent (Drizzle uses JS field names verbatim → quoted camelCase in raw SQL). + +--- + +## §8 — Smoke checklist + out-of-scope + +Per project convention, no TDD scaffolding. Exercise the system through actual TUI interaction at end of implementation. + +### Smoke checklist (executed against a fresh deploy) + +| # | Scenario | Verification | +|---|---|---| +| 1 | Empty thread shows no `` block | Fresh thread, ask "what time is it" — agent answers, `` block absent (zero rows in `task` table). | +| 2 | Trivial request doesn't trigger planning | "Read foo.md" — agent reads directly, no `manage_tasks` `tool_call` row. | +| 3 | Multi-step request triggers planning | "Research bun vs deno, write a summary doc, post to slack" — agent's first or second tool call is `manage_tasks(create, ops: [...])` with 2–3 tasks. Verify rows in `task` table with correct `threadId` + `runId`. | +| 4 | Mid-turn task state visibility (the prepareStep test) | In scenario 3's run, after creating tasks the very next assistant step should mark one `in_progress`. If it doesn't, `prepareStep` refresh is broken. | +| 5 | Persistence across user turns | After scenario 3, send a second user message. New turn's `` block shows the prior turn's tasks with current status. | +| 6 | Cross-thread isolation | Create tasks in thread A. Open thread B (new). B's first turn renders no `` block. | +| 7 | Inline completion with result | Multi-step turn where agent does the work inline. On terminal status update, check `task.result` is non-null with a 1–3 sentence summary. | +| 8 | `manage_tasks` returns full state | Inspect any `tool_call.response` row for `manage_tasks` — should contain `{ tasks: [...] }` with all thread's tasks, not just the affected ones. | +| 9 | Cross-thread `update` rejected | Manually craft a thread-A task `id` and attempt `manage_tasks(update, id=A_id)` from thread B. Expect: `Task X belongs to a different thread`. | +| 10 | Cache behavior (best-effort) | Per memory (`feedback_wrangler_tail_misses_do_logs`): `wrangler tail` doesn't catch DO `onFinish` logs. **Skip cache verification in this PR's smoke checklist**; design verified analytically in §6. Add a debug route only if production usage shows unexpected cost. | + +Scenarios 1–8 must pass before merge. Scenario 9 is defensive (low-probability bug). Scenario 10 deferred. + +### Out of scope — PR-D territory + +| Feature | Why deferred | +|---|---| +| `spawn_sub_agent` tool | PR-D core deliverable. | +| Sub-agent runtime (isolated react loop, fresh context) | PR-D core deliverable. | +| Auto-population of `task.result` from sub-agent's final assistant message | PR-D wires this in `spawn_sub_agent`'s `execute`. Schema supports it without changes. | +| 2x result duplication handling (tool_result + context block) | Decided: accept 2x. Tool_result is transient (until compaction); context block is durable. Implementation lands in PR-D. | +| Retry-with-guidance flow | PR-D. Likely adds `prior_attempts` field to `task`. | +| Parallel sub-agent execution on different tasks | PR-D. Possible because soft "one in_progress" rule isn't DB-enforced. | + +### Out of scope — PR-E territory + +| Feature | Why deferred | +|---|---| +| TUI rendering of `` as a sidebar/panel | UX polish PR. Tasks-in-prompt is the substrate; visual surface is separate. | +| Streaming sub-agent events to TUI | PR-E (or rolled into PR-D). | + +### Out of scope — deferred indefinitely + +| Feature | Why | +|---|---| +| Nested tasks (`parentTaskId`) | Flat list sufficient for discussed workflows. Add if PR-D reveals real need. | +| Task priority / explicit ordering | Agent orders by creation sequence. | +| `get_task(id)` read tool | Context block + tool_result history + sessions_search cover all retrieval. | +| Task templates / "saved plans" | Product-feature territory, not runtime. | +| Cross-thread tasks | Threads are the conversation atom. | + +### PR-D readiness check + +PR-C's design is forward-compatible without schema changes: + +- `task.result` field exists; PR-D's `spawn_sub_agent` writes via direct DB update or shared update path. +- `task.status` enum supports all terminal states; PR-D adds no new ones. +- `prepareStep` refresh path exists; PR-D mutations picked up automatically next step. +- `runId` FK preserves authorship across runs; PR-D sub-agent execution under main's run inherits the same runId via `perTurn`. + +Anticipated PR-D additions: sub-agent-messages table (separate from `message`); possibly `prior_attempts` on `task` for retry context. Neither blocks PR-C. + +--- + +## §9 — Anthropic provider options (additive) + +Adjacent to PR-C's `turn.ts` mods. Five options to add — some top-level in `providerOptions.anthropic`, one (`cacheControl.ttl`) applies at every per-message cacheControl marker. + +### Top-level `providerOptions.anthropic` (additions to existing block at `turn.ts:275-287`) + +```ts +providerOptions: { + anthropic: { + // EXISTING (unchanged) ───────────────────────────── + anthropicBeta: ["compact-2026-01-12"], + contextManagement: { + edits: [{ + type: "compact_20260112", + trigger: { type: "input_tokens", value: midrunThreshold } + }] + }, + // NEW (PR-C) ─────────────────────────────────────── + toolStreaming: true, + effort: "xhigh", + sendReasoning: true, + thinking: { + type: "adaptive", + display: "summarized" + } + } satisfies AnthropicLanguageModelOptions +} +``` + +### Per-marker `cacheControl.ttl` + +Replace every `cacheControl: { type: "ephemeral" }` in the codebase with `{ type: "ephemeral", ttl: "5m" }`: + +| Site | File | Purpose | +|---|---|---| +| System message | `apps/kernel/src/turn.ts:262` | b1-adjacent — system content cache | +| Tools tail | `apps/kernel/src/tools/index.ts:30` (in `withTailCache`) | b1 — tool definitions cache | +| End-of-stable-context | wherever PR-B applies b2 on `` | b2 — stable context cache | +| Last conv msg | wherever PR-B applies b4 | b4 — conversation cache | + +### What each does + +| Option | Effect | Why we want it | +|---|---|---| +| `toolStreaming: true` | Anthropic streams partial tool-call arguments as they emit (not just final). | TUI renders mid-tool-call progress. Free behaviorally — just opt-in. | +| `effort: "xhigh"` | Extended-thinking budget set to maximum tier. | Yoke's workflows (research → document → ship) benefit from deep planning before tool calls. | +| `sendReasoning: true` | Reasoning content from prior steps sent back on subsequent steps in the same react loop. | Model builds on its earlier thinking across the multi-step loop. Without this, every step's thinking starts cold — major coherence loss for multi-step turns. | +| `thinking: { type: "adaptive", display: "summarized" }` | Model auto-tunes thinking budget per step. Response surfaces a thinking *summary* rather than full chain. | `adaptive` saves tokens on trivial steps. `summarized` keeps conv history lean — full reasoning chain doesn't bloat the messages list. | +| `cacheControl.ttl: "5m"` | Explicit 5-minute TTL on every ephemeral cache breakpoint. | Currently implicit (5m is Anthropic default). Explicit makes future "should we try 1h?" experiment a one-character diff. | + +**Cost/risk note on `effort: "xhigh"`:** extended thinking on max-tier costs more output tokens (billed at reduced rate vs regular output). Combined with `adaptive`, the math favors `xhigh`: trivial requests barely think, substantial requests think hard. The alternative (`"medium"` everywhere) under-thinks complex planning and over-thinks trivial answers. + +**No behavioral change for `cacheControl.ttl: "5m"`** — making the default explicit. One-line change to flip to `"1h"` if long-session cache survival becomes valuable. + +--- + +## Decisions made during brainstorm + +One-line summary of each major fork and the reason: + +| Decision | Alternative | Reason | +|---|---|---| +| Three-slice PR plan: Tasks → Sub-agents → Streaming | One mega-PR | Each slice independently useful; smaller review surface; faster ship within 2026-06-15 deadline | +| Single `manage_tasks` write tool (batchable ops) | Separate create/update/list tools | Atomic ops; cheaper (1 call vs N); forces planning-before-execution | +| Soft "one in_progress" rule (system prompt only) | DB-enforced unique constraint | Preserves PR-D's option for parallel sub-agent execution | +| Per-thread scoping (Claude Code parity) | Per-run (reset every user message) | Cross-turn continuity for multi-message workflows; well-understood model | +| Tasks block placed AFTER b2 (post-stable-context) | Before b2 in cached prefix | Tasks change every step; busting b2 cache would cost ~5K–10K tokens per step | +| Tasks block built in `prepareStep`, not `buildContextMessages` | Build once at turn start | Keeps block fresh across react-loop iterations; critical for PR-D sub-agents | +| 5-value status enum (`pending/in_progress/complete/failed/cancelled`) | 3-value (Claude Code's) | Failed vs cancelled distinction is useful signal | +| `result` field optional with system-prompt encouragement | Required on terminal status | Forcing produces filler text for tasks without natural results | +| Description truncation in context block | Show full descriptions always | Full text recoverable via `tool_result` history; truncation is cache hygiene | +| No `get_task` read tool | Add one for paranoia | YAGNI — context block + tool_result history + sessions_search cover all retrieval paths | +| Accept 2x result duplication for sub-agent results (PR-D) | Sub-agent stores in DB only, main sees next turn | "Next turn" requires user message; breaks the in-turn retry-vs-move-on decision flow | +| Anthropic `effort: "xhigh"` + `thinking: adaptive` | Medium/static thinking | Adaptive auto-tunes per step; xhigh ceiling for substantial work | + +--- + +## Open questions + +None at spec time. All sections accepted by user 2026-05-14. From d969bbd5096133df4e929cbbd13ef27244232bb6 Mon Sep 17 00:00:00 2001 From: Ronit Date: Thu, 14 May 2026 16:26:49 +0530 Subject: [PATCH 02/11] =?UTF-8?q?docs(kernel):=20PR-C=20implementation=20p?= =?UTF-8?q?lan=20=E2=80=94=20task=20management?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit 7 tasks: schema+migration → tasks-block renderer → manage_tasks tool → prepareStep wiring → system prompt section → Anthropic provider options → deploy+smoke+PR. Each task self-contained with full code, exact paths, typecheck-as-test gate, individual commit. No TDD scaffolding (per project convention). Final smoke checklist runs 8 scenarios against deployed kernel + d1 verifications. Co-Authored-By: Claude Opus 4.7 (1M context) --- .../plans/2026-05-14-task-management.md | 1068 +++++++++++++++++ 1 file changed, 1068 insertions(+) create mode 100644 docs/superpowers/plans/2026-05-14-task-management.md diff --git a/docs/superpowers/plans/2026-05-14-task-management.md b/docs/superpowers/plans/2026-05-14-task-management.md new file mode 100644 index 0000000..e0ac8c7 --- /dev/null +++ b/docs/superpowers/plans/2026-05-14-task-management.md @@ -0,0 +1,1068 @@ +# Task Management (PR-C) Implementation Plan + +> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking. +> +> **Model selection:** all tasks dispatched to **opus** (per memory `feedback_subagent_opus_default`). Cheaper tiers introduce silly bugs on Drizzle/TS surfaces. + +**Goal:** Build a per-thread task list with single batchable `manage_tasks` write tool + `` context block refreshed per-step via `streamText` `prepareStep`. First of three PR slices (Tasks → Sub-agents → Streaming). + +**Architecture:** New `task` Drizzle table cascade-tied to `thread` and `run`. One write tool (`manage_tasks` with batchable `create`/`update` ops, wrapped via `wrappedTool({ touchesFS: false })`). Context block built by a pure SQL renderer (`tasks-block.ts`) and spliced into the message list inside `prepareStep` at the index immediately after the b2-cached stable prefix. System prompt gains a `` section. Five Anthropic provider options added adjacently (`toolStreaming`, `effort`, `sendReasoning`, `thinking`, explicit `cacheControl.ttl`). + +**Tech Stack:** TypeScript / Cloudflare Workers + Durable Objects / AI SDK v5 / `@ai-sdk/anthropic` / Drizzle ORM (SQLite, D1) / Zod / `tsx` / `wrangler`. + +**Spec reference:** `docs/superpowers/specs/2026-05-14-task-management-design.md` + +--- + +## Task 1: `task` Drizzle schema + migration + +**Files:** +- Create: `packages/models/drizzle/schemas/tasks.ts` +- Modify: `packages/models/drizzle/schemas/index.ts` +- Generated: `packages/models/drizzle/migrations/0009__.sql` (drizzle-kit names the file) + +- [ ] **Step 1: Read existing schema files for the closest pattern** + +```bash +cat packages/models/drizzle/schemas/messages.ts +cat packages/models/drizzle/schemas/runs.ts +cat packages/models/drizzle/schemas/index.ts +``` + +`messages.ts` is the closest pattern: ULID `id`, `threadId` + `runId` FKs with cascade, status enum, two `timestamp_ms` fields, `InferSelectModel` type export at the bottom. Mirror its shape. + +- [ ] **Step 2: Create `packages/models/drizzle/schemas/tasks.ts`** + +```ts +import { InferSelectModel } from "drizzle-orm"; +import { index, integer, sqliteTable, text } from "drizzle-orm/sqlite-core"; + +import { run } from "./runs"; +import { thread } from "./threads"; + +export const task = sqliteTable( + "task", + { + id: text().primaryKey(), // ULID + threadId: text() + .notNull() + .references(() => thread.id, { onDelete: "cascade" }), + // The run that CREATED this task (manage_tasks{action:"create"} call). + // Updates from later runs do NOT change this — preserves "who planned this". + runId: text() + .notNull() + .references(() => run.id, { onDelete: "cascade" }), + name: text().notNull(), // user-visible label + description: text().notNull(), // full plan: deliverables, deps, expected output + status: text({ + enum: ["pending", "in_progress", "complete", "failed", "cancelled"] + }) + .notNull() + .default("pending"), + // Set on transition to a terminal status: + // complete → the actual output/answer + // failed → the reason the agent gave up + // cancelled → why it was abandoned + result: text(), + createdAt: integer({ mode: "timestamp_ms" }) + .notNull() + .$defaultFn(() => new Date()), + updatedAt: integer({ mode: "timestamp_ms" }) + .notNull() + .$defaultFn(() => new Date()) + }, + (table) => [ + index("idx_task_thread_created").on(table.threadId, table.createdAt) + ] +); + +export type Task = InferSelectModel; +``` + +- [ ] **Step 3: Add re-export to `packages/models/drizzle/schemas/index.ts`** + +Read the file first, then add this line where the other `export *` lines are: + +```ts +export * from "./tasks"; +``` + +- [ ] **Step 4: Generate the migration** + +```bash +cd packages/models && pnpm run db:generate +``` + +Expected: drizzle-kit prints a summary and creates a new file `packages/models/drizzle/migrations/0009__.sql` (sequence number follows the last existing migration, currently `0008_special_grandmaster.sql`). + +Sanity-check the generated SQL: open the new file. Should contain CREATE TABLE for `task` with FKs to `thread` and `run`, plus a CREATE INDEX for `idx_task_thread_created`. If the generated SQL is missing the index or FKs, the schema definition in Step 2 is wrong — fix it and re-run. + +- [ ] **Step 5: Verify `Task` flat export reaches `@agent-os/models`** + +```bash +cd /Users/rtpa25/Developer/agent-os && grep -rn "export.*Task\b" packages/models --include="*.ts" +``` + +Expected: `Task` appears in the re-export chain. If the `messages.ts` precedent re-exports `Message` correctly via `index.ts`, and `tasks.ts` follows the same shape, this should work transitively. If `Task` doesn't appear at `@agent-os/models`'s public root, walk the chain (schemas/index.ts → drizzle/index.ts → packages/models/index.ts) and add the re-export at whichever level is missing. + +- [ ] **Step 6: Typecheck** + +```bash +cd /Users/rtpa25/Developer/agent-os && pnpm typecheck +``` + +Expected: PASS across the workspace. If `packages/models` fails, fix; if `apps/kernel` fails, it's a downstream issue you'll fix in Tasks 2–4. + +- [ ] **Step 7: Commit** + +```bash +git add packages/models/drizzle/schemas/tasks.ts \ + packages/models/drizzle/schemas/index.ts \ + packages/models/drizzle/migrations/ +git commit -m "feat(models): add task schema with thread+run FKs + +Per-thread tasks with 5-state status enum (pending/in_progress/complete/ +failed/cancelled), nullable result, composite index on (threadId, createdAt). +Cascade-delete with both thread and run." +``` + +--- + +## Task 2: `tasks-block.ts` context block renderer + +**Files:** +- Create: `apps/kernel/src/agent/tasks-block.ts` + +- [ ] **Step 1: Read the closest sibling for pattern** + +```bash +cat apps/kernel/src/agent/sessions-recent-block.ts +cat apps/kernel/src/agent/current-date-block.ts +``` + +`sessions-recent-block.ts` is the closest match: pure SQL fetch, status-aware rendering with a `truncate()` helper, returns a string (or error placeholder), uses `await db.all` / `db.query.X.findMany`. Mirror its shape and error-handling pattern. Note: our renderer returns `string | null` (null when empty) — different from sessions-recent-block which always returns a string. + +- [ ] **Step 2: Create `apps/kernel/src/agent/tasks-block.ts`** + +```ts +import { asc, eq, schema, type DB } from "@agent-os/models"; + +const PENDING_DESC_MAX = 120; +const TERMINAL_DESC_MAX = 80; + +/** + * Collapse whitespace and hard-truncate to `n` chars. Caller appends the + * ellipsis when `truncated === true`. + */ +function truncate(s: string, n: number): { text: string; truncated: boolean } { + const trimmed = s.replace(/\s+/g, " ").trim(); + if (trimmed.length <= n) return { text: trimmed, truncated: false }; + return { text: trimmed.slice(0, n), truncated: true }; +} + +function formatDescription( + status: "pending" | "in_progress" | "complete" | "failed" | "cancelled", + description: string +): string { + if (status === "in_progress") { + return description.replace(/\s+/g, " ").trim(); + } + const max = status === "pending" ? PENDING_DESC_MAX : TERMINAL_DESC_MAX; + const { text, truncated } = truncate(description, max); + return truncated ? `${text}…` : text; +} + +/** + * Build the `` block. Returns `null` when the thread has zero tasks + * (block is omitted entirely from the prompt — same pattern as + * ). + * + * Called from `prepareStep` in `turn.ts` on every step boundary inside + * `streamText`. Pure SQL fetch — one index hit on `idx_task_thread_created`. + * + * Truncation policy: + * in_progress → full description (active focus needs the plan) + * pending → trimmed to 120 chars + ellipsis if truncated + * complete/failed/cancelled → trimmed to 80 chars; result shown in full + * when non-null + * + * Full descriptions remain in the agent's conversation history via every + * `manage_tasks` tool_result — truncation here is per-turn cache hygiene, + * not lossy storage. + * + * Failure mode: SQL throw → error placeholder so the block still ships and + * the agent isn't blind to its task list breaking. + */ +export async function buildTasksBlock(args: { + db: DB; + threadId: string; +}): Promise { + let rows; + try { + rows = await args.db.query.task.findMany({ + where: eq(schema.task.threadId, args.threadId), + orderBy: [asc(schema.task.createdAt)] + }); + } catch (e) { + console.warn("[tasks-block] SQL failed", { + threadId: args.threadId, + error: e instanceof Error ? e.message : String(e) + }); + return [ + "", + "(error loading task list — manage_tasks may still work, try it)", + "" + ].join("\n"); + } + + if (rows.length === 0) return null; + + const lines = rows.map((t) => { + const head = `- ${t.id} [${t.status}] "${t.name}"`; + const plan = ` plan: ${formatDescription(t.status, t.description)}`; + const isTerminal = + t.status === "complete" || + t.status === "failed" || + t.status === "cancelled"; + if (isTerminal && t.result) { + const result = t.result.replace(/\s+/g, " ").trim(); + return [head, plan, ` result: ${result}`].join("\n"); + } + return [head, plan].join("\n"); + }); + + return [ + "", + "Plan and track multi-step work here. Use manage_tasks to create/update tasks.", + "Convention: one task in_progress at a time. Mark complete/failed/cancelled", + "before moving on. Skip task planning for trivial single-step requests.", + "", + ...lines, + "" + ].join("\n"); +} +``` + +- [ ] **Step 3: Typecheck** + +```bash +cd /Users/rtpa25/Developer/agent-os/apps/kernel && pnpm typecheck +``` + +Expected: PASS. If `schema.task` isn't found, Task 1's flat exports didn't propagate — go back to Task 1 Step 5. + +- [ ] **Step 4: Commit** + +```bash +git add apps/kernel/src/agent/tasks-block.ts +git commit -m "feat(kernel): tasks-block renderer for context block + +Pure SQL fetch from task table. Status-aware description truncation +(in_progress full / pending 120 / terminal 80 chars), result shown in +full for terminal entries when non-null. Returns null on empty thread +so the block is omitted from prompts entirely. Consumed by prepareStep +callback in turn.ts (Task 4)." +``` + +--- + +## Task 3: `manage_tasks` tool + +**Files:** +- Create: `apps/kernel/src/tools/tasks.ts` +- Modify: `apps/kernel/src/tools/index.ts` + +- [ ] **Step 1: Read the closest sibling for pattern** + +```bash +cat apps/kernel/src/tools/sessions.ts +cat apps/kernel/src/tools/wrapped-tool.ts +``` + +`sessions.ts` is the closest pattern: a `buildSessionsTools(perTurn)` factory returning multiple wrappedTools with `touchesFS: false`. Confirms `PerTurnContext` shape (`runId`, `threadId`, `db`, `kernel`, `env`, `broadcast`). Audit rows write automatically via `wrappedTool`. + +- [ ] **Step 2: Create `apps/kernel/src/tools/tasks.ts`** + +```ts +import { asc, eq, schema, ulid } from "@agent-os/models"; +import { z } from "zod"; + +import { wrappedTool, type PerTurnContext } from "./wrapped-tool"; + +const TASK_STATUS = [ + "pending", + "in_progress", + "complete", + "failed", + "cancelled" +] as const; + +const createOpSchema = z.object({ + action: z.literal("create"), + name: z.string().min(1).describe("Short user-visible label"), + description: z + .string() + .min(1) + .describe("Full plan — deliverables, deps, expected output") +}); + +const updateOpSchema = z.object({ + action: z.literal("update"), + id: z.string().describe("Task id (ULID) to update"), + name: z.string().min(1).optional(), + description: z.string().min(1).optional(), + status: z.enum(TASK_STATUS).optional(), + result: z + .string() + .optional() + .describe("Set on terminal status with 1-3 sentence summary") +}); + +const inputSchema = z.object({ + ops: z + .array(z.discriminatedUnion("action", [createOpSchema, updateOpSchema])) + .describe( + "Batch of create/update operations. Applied in order. Empty array is a no-op refresh." + ) +}); + +type ManageTasksArgs = z.infer; + +interface TaskJSON { + id: string; + name: string; + description: string; + status: (typeof TASK_STATUS)[number]; + result: string | null; + createdAt: string; + updatedAt: string; +} + +const DESCRIPTION = [ + "Plan and track multi-step work as tasks. Use ONE call to batch creates +", + "updates. Returns the full thread task list (sorted by createdAt ASC) so", + "you see canonical state immediately.", + "", + "When to use: multi-step requests (research X, write Y, post Y). Skip for", + "trivial single-step requests.", + "", + "Ops:", + " create: { action:'create', name, description }", + " update: { action:'update', id, ...any of: name, description, status, result }", + "", + "Status values: pending | in_progress | complete | failed | cancelled", + "Status transitions are unrestricted (any → any).", + "", + "Set 'result' on terminal status (complete/failed/cancelled) with a 1-3", + "sentence summary. Optional but strongly preferred — it's how future turns", + "see what was accomplished." +].join("\n"); + +export function buildTasksTools(perTurn: PerTurnContext) { + return { + manage_tasks: wrappedTool( + { + name: "manage_tasks", + description: DESCRIPTION, + inputSchema, + touchesFS: false, + needsApproval: false, + execute: async (args: ManageTasksArgs, ctx) => { + for (const op of args.ops) { + if (op.action === "create") { + await ctx.db.insert(schema.task).values({ + id: ulid(), + threadId: ctx.threadId, + runId: ctx.runId, + name: op.name, + description: op.description + // status defaults to "pending" + // createdAt/updatedAt default via $defaultFn + }); + } else { + // op.action === "update" + const existing = await ctx.db.query.task.findFirst({ + where: eq(schema.task.id, op.id) + }); + if (!existing) { + throw new Error(`Task ${op.id} not found`); + } + if (existing.threadId !== ctx.threadId) { + // Defensive: model gave us an id from another thread. + // Refuse rather than silently mutate. + throw new Error( + `Task ${op.id} belongs to a different thread` + ); + } + const patch: Partial = { + updatedAt: new Date() + }; + if (op.name !== undefined) patch.name = op.name; + if (op.description !== undefined) { + patch.description = op.description; + } + if (op.status !== undefined) patch.status = op.status; + if (op.result !== undefined) patch.result = op.result; + await ctx.db + .update(schema.task) + .set(patch) + .where(eq(schema.task.id, op.id)); + } + } + + // Return full updated task list sorted by createdAt ASC. + const tasks = await ctx.db.query.task.findMany({ + where: eq(schema.task.threadId, ctx.threadId), + orderBy: [asc(schema.task.createdAt)] + }); + const result: { tasks: TaskJSON[] } = { + tasks: tasks.map((t) => ({ + id: t.id, + name: t.name, + description: t.description, + status: t.status, + result: t.result, + createdAt: t.createdAt.toISOString(), + updatedAt: t.updatedAt.toISOString() + })) + }; + return result; + } + }, + perTurn + ) + }; +} +``` + +- [ ] **Step 3: Register the tool in `apps/kernel/src/tools/index.ts`** + +Read the file first to see the current imports and `buildTools` spread: + +```bash +cat apps/kernel/src/tools/index.ts +``` + +Then modify it so the imports section includes: + +```ts +import { buildTasksTools } from "./tasks"; +``` + +(alphabetical order, between `buildSessionsTools` and `buildWebSearchTools`). + +And the `buildTools` spread becomes: + +```ts +export const buildTools = (perTurn: PerTurnContext) => + withTailCache({ + ...buildWebSearchTools(perTurn), + ...buildFSTools(perTurn), + ...buildProcessAttachmentTool(perTurn), + ...buildExecCodeTool(perTurn), + ...buildComputerUseTools(perTurn), + ...buildSessionsTools(perTurn), + ...buildTasksTools(perTurn), + get_time: buildGetTimeTool(perTurn) + }); +``` + +Order doesn't matter — `withTailCache` shifts the b1 marker to whatever tool key ends up last (still `get_time` here). + +- [ ] **Step 4: Typecheck** + +```bash +cd /Users/rtpa25/Developer/agent-os/apps/kernel && pnpm typecheck +``` + +Expected: PASS. Common issue: the discriminated union schema may surface inference quirks in older Zod versions. Match the existing Zod patterns in `sessions.ts` / `fs-tools.ts` if a type error appears around `z.discriminatedUnion` — those files have working precedents. + +- [ ] **Step 5: Commit** + +```bash +git add apps/kernel/src/tools/tasks.ts apps/kernel/src/tools/index.ts +git commit -m "feat(kernel): manage_tasks tool with batchable ops + +Single write tool: ops:[{action:'create',name,description}|{action:'update', +id,...patch}]. Returns full thread task list sorted by createdAt ASC. +threadId/runId injected server-side from perTurn; cross-thread updates +throw. Wrapped via wrappedTool({touchesFS:false}) so audit rows write +without per-call worktree overhead." +``` + +--- + +## Task 4: `prepareStep` wiring in `turn.ts` + +**Files:** +- Modify: `apps/kernel/src/turn.ts` (add import; capture `stableBlockCount`; add `prepareStep` callback to `streamText({...})`) + +- [ ] **Step 1: Read the streamText call and surrounding context** + +```bash +cat apps/kernel/src/turn.ts | head -330 +``` + +Locate three things you'll be modifying: +1. The import section at the top (you'll add one line). +2. The `buildContextMessages` call site (capture `stableBlockCount` immediately after). +3. The `streamText({...})` invocation (around line 256–322) — `prepareStep` slots in between `messages:` and `onStepFinish:`. + +- [ ] **Step 2: Add the import** + +At the top of `apps/kernel/src/turn.ts`, alongside the existing `import { buildContextMessages }` line, add: + +```ts +import { buildTasksBlock } from "./agent/tasks-block"; +``` + +- [ ] **Step 3: Capture `stableBlockCount` after `buildContextMessages`** + +Find the line in `turn.ts` where `buildContextMessages(...)` is called and its result is assigned. Immediately after that assignment, add the constant: + +```ts +const contextMessages = await buildContextMessages({ /* existing args */ }); +const stableBlockCount = contextMessages.length; +``` + +This value is captured once per turn and reused across every `prepareStep` invocation. + +- [ ] **Step 4: Add `prepareStep` callback inside `streamText({...})`** + +Inside the existing `streamText({ ... })` call, after the `messages:` line and before the existing `onStepFinish:` callback, insert: + +```ts +// Refresh block on every step boundary. Owned entirely by this +// callback — never by buildContextMessages — so the block stays fresh as +// the agent mutates task state across react-loop iterations inside this +// streamText invocation. +// +// Bytes live AFTER b2 cache marker (b2 is on , the last +// stable context block), so the refresh costs ~300 tokens per step and +// does NOT bust b2. b4 misses on within-turn steps that change tasks; +// acceptable trade — see spec §6 worked example. +// +// Splice invariant: only touches indices >= stableBlockCount. Touching +// earlier indices would mutate the b2-cached prefix and bust caching. +prepareStep: async ({ stepNumber, messages }) => { + const working = [...messages]; + + // On step 1+, strip the prior step's tasks block. It always lives at + // index stableBlockCount IF buildTasksBlock returned non-null last step. + // Detected by role + literal "" tag prefix. + if (stepNumber > 0) { + const candidate = working[stableBlockCount]; + if ( + candidate?.role === "user" && + typeof candidate.content === "string" && + candidate.content.startsWith("") + ) { + working.splice(stableBlockCount, 1); + } + } + + const tasksBlock = await buildTasksBlock({ + db: kernel.db, + threadId + }); + if (tasksBlock !== null) { + working.splice(stableBlockCount, 0, { + role: "user" as const, + content: tasksBlock + }); + } + + return { messages: working }; +}, +``` + +- [ ] **Step 5: Typecheck** + +```bash +cd /Users/rtpa25/Developer/agent-os/apps/kernel && pnpm typecheck +``` + +Expected: PASS. If TypeScript complains about the `messages` array type from `prepareStep` (the AI SDK uses a wider union internally), the existing `messages: [...contextMessages, ...modelMessages]` line is the precedent for the shape — the splice returns the same shape. If a cast is unavoidable, prefer `working as typeof messages` over `any`. + +- [ ] **Step 6: Local dev sanity check** + +```bash +cd /Users/rtpa25/Developer/agent-os/apps/kernel && pnpm run dev +``` + +The worker should boot without errors. Don't actually test functional behavior here — Task 7's smoke check covers that. This step just confirms the wrangler dev server can compile and start with the new `prepareStep` callback in place. + +Stop the dev server (`Ctrl+C`) once it's stably running. + +- [ ] **Step 7: Commit** + +```bash +git add apps/kernel/src/turn.ts +git commit -m "feat(kernel): wire prepareStep to refresh block per step + +streamText runs the entire react loop internally; without prepareStep the +context blocks would be frozen at turn start and the agent would see stale +task state across iterations. Capture stableBlockCount after buildContext- +Messages and splice the freshly-rendered block at that index every +step. + +Splice invariant: only touches indices >= stableBlockCount, leaving the +b2-cached stable prefix byte-identical so b2 keeps hitting across steps. +Costs ~300 tokens of tasks-render per step vs ~5-10k of cache rebuild if +tasks were inside the b2 prefix." +``` + +--- + +## Task 5: `` system prompt section + +**Files:** +- Modify: `apps/kernel/src/agent/system-prompt.ts` (append a new XML-tagged block) + +- [ ] **Step 1: Read the system prompt to find the insertion point** + +```bash +cat apps/kernel/src/agent/system-prompt.ts +``` + +Look for `` and `` blocks — those are the closest neighbors in role (how-to-use-this-tool). Insert the new `` block adjacent to them (immediately before or after — order doesn't materially affect the model). + +- [ ] **Step 2: Append the `` block** + +Inside the `SYSTEM_PROMPT` template literal, insert this block at the chosen location (after `` recommended): + +``` + +For multi-step requests, plan and track your work with manage_tasks. The + context block (when present) is your source of truth for task state. +It refreshes at every step boundary so it never lags behind your tool calls. + +WHEN TO PLAN VS JUST DO +The question to ask yourself: is the work substantial, or trivial? +- "Read this file and tell me what it says" → just do it. One tool call. +- "What time is it in Tokyo" → just do it. +- "Research X, write Y, post Y to slack" → plan tasks first. +- "Refactor this module" if it's one file → just do it. +- "Refactor this module" if it spans many files → plan tasks first. + +Planning costs tool calls. Only plan when the work has independent steps you +want to track, retry, or summarize. When the user signals depth ("deep dive", +"thoroughly", "do this properly"), lean toward planning even if it looks small. + +LIFECYCLE +- create: batch all tasks up front with name + description. + name: short user-visible label. + description: full plan — what to do, what to produce, any constraints. + Be specific. Your future self (next step) reads this and must act on it + without re-asking the user. +- update to in_progress before you start work on a task. +- update to complete when done. Set result to 1-3 sentences of what was + produced or accomplished. Result is optional but strongly preferred — it's + how future turns see "this got done, here's what it was." +- update to failed if you tried and gave up; set result to the reason. +- update to cancelled if you decided not to do this task; set result to why. + +DESCRIPTION QUALITY +Bad: "Refactor parser" +Good: "Refactor parser: split parse() into tokenize() and parseTokens(). + Preserve current public API. Update the 3 call sites in handlers.ts. + Aim under 80 lines per function." + +A vague description means you (or a future step) won't know what "done" +looks like. + +DISCIPLINE +- One task in_progress at a time. Mark the current one terminal before + starting the next. +- Update status the moment state changes. Stale in_progress tasks make the + user think you're confused. +- Tasks persist for the whole thread. Completed ones stay visible as your + audit trail. + +DON'T +- Don't plan tasks for trivial requests. +- Don't create a task and mark it complete in the same call (no fake + productivity). +- Don't recreate a failed task with the same description. Either update its + status for a retry, or create a NEW task whose description includes what + went wrong and what's different. + +``` + +**Important formatting notes:** +- Inside the template literal, leave it as plain text — no escaping needed for `<`, `>`, `"`. Only escape backticks (`` \` ``) and `${`. +- The em-dashes (`—`) in this block are inside a `` explanatory block, NOT in user-facing voice/style guidance. The system prompt's em-dash audit applies to the AGENT's outputs, not to instructional content the agent reads. See the comment block at the top of `system-prompt.ts` (lines 23-27 explain this exception). + +- [ ] **Step 3: Typecheck** + +```bash +cd /Users/rtpa25/Developer/agent-os/apps/kernel && pnpm typecheck +``` + +Expected: PASS (this is a string change, but it touches the kernel — Cloudflare TypeScript surface). + +- [ ] **Step 4: Commit** + +```bash +git add apps/kernel/src/agent/system-prompt.ts +git commit -m "feat(kernel): system prompt section + +Teaches when to plan vs just do (substantial vs trivial test), lifecycle +with status transitions, description quality bar with good/bad pair, +discipline rules (one in_progress, immediate status updates), and don'ts. +~370 tokens, cached in b1 — paid once per fresh cache, then free. + +Lifted from Dimension AI's / framing, +trimmed to exclude sub-agent guidance (PR-D scope)." +``` + +--- + +## Task 6: Anthropic provider options + `cacheControl.ttl` + +**Files:** +- Modify: `apps/kernel/src/turn.ts` (top-level providerOptions block + every cacheControl marker) +- Modify: `apps/kernel/src/tools/index.ts` (withTailCache cacheControl marker) + +- [ ] **Step 1: Find every `cacheControl: { type: "ephemeral" }` site** + +```bash +cd /Users/rtpa25/Developer/agent-os/apps/kernel && grep -rn "cacheControl.*ephemeral" src/ +``` + +Expected sites (verify against output — line numbers may have drifted): +- `src/turn.ts:185` (partial-assistant cacheControl) +- `src/turn.ts:198` (b4 — last conv msg) +- `src/turn.ts:243` (b2 — end-of-context) +- `src/turn.ts:262` (system message) +- `src/tools/index.ts:31` (withTailCache → b1 tools) + +Every site needs the `ttl: "5m"` extension. + +- [ ] **Step 2: Update each cacheControl marker** + +For each line found in Step 1, modify: + +```ts +// Before: +cacheControl: { type: "ephemeral" } + +// After: +cacheControl: { type: "ephemeral", ttl: "5m" } +``` + +Use `Edit` with `replace_all: true` on `apps/kernel/src/turn.ts` (the four sites there share the same string), then a separate edit on `apps/kernel/src/tools/index.ts`. + +```bash +# After edits, verify all sites updated: +cd /Users/rtpa25/Developer/agent-os/apps/kernel && grep -rn "cacheControl.*ephemeral" src/ +``` + +Every match should now include `ttl: "5m"`. + +- [ ] **Step 3: Add top-level Anthropic provider options** + +Open `apps/kernel/src/turn.ts` and find the existing `providerOptions` block (around lines 275–287 inside the main `streamText({...})` call — the one with `anthropicBeta` and `contextManagement`). + +Replace it with: + +```ts +providerOptions: { + anthropic: { + anthropicBeta: ["compact-2026-01-12"], + contextManagement: { + edits: [ + { + type: "compact_20260112", + trigger: { type: "input_tokens", value: midrunThreshold } + } + ] + }, + // NEW (PR-C) ─────────────────────────────────────── + toolStreaming: true, + effort: "xhigh", + sendReasoning: true, + thinking: { + type: "adaptive", + display: "summarized" + } + } satisfies AnthropicLanguageModelOptions +}, +``` + +The `satisfies AnthropicLanguageModelOptions` clause catches type incompatibility. If your `@ai-sdk/anthropic` version doesn't yet expose `effort: "xhigh"` or `thinking.type: "adaptive"`, TypeScript will tell you immediately. + +- [ ] **Step 4: Typecheck** + +```bash +cd /Users/rtpa25/Developer/agent-os/apps/kernel && pnpm typecheck +``` + +Expected: PASS. If `effort` or `thinking.adaptive` errors, the installed `@ai-sdk/anthropic` doesn't support them yet — bump the version: + +```bash +cd /Users/rtpa25/Developer/agent-os && pnpm update @ai-sdk/anthropic --latest +``` + +then retypecheck. If specific options remain unsupported even on latest, drop only the unsupported field (keep the rest) and add a code comment explaining which version exposes it. + +- [ ] **Step 5: Commit** + +```bash +git add apps/kernel/src/turn.ts apps/kernel/src/tools/index.ts +git commit -m "feat(kernel): Anthropic provider options for thinking + cache TTL + +- toolStreaming: true — TUI sees partial tool args mid-emit +- effort: 'xhigh' — max thinking budget per step +- sendReasoning: true — reasoning sent back across react steps +- thinking.adaptive+summary — auto-tune per step, lean conv history +- cacheControl.ttl: '5m' — explicit (was implicit default) at all + cacheControl marker sites so future 1h + experiment is a one-character diff + +Combined with adaptive: trivial requests barely think, substantial ones +think hard. Net positive for multi-step workflows that PR-C enables." +``` + +--- + +## Task 7: Deploy + smoke check + open PR + +**Files:** none modified (verification + handoff) + +- [ ] **Step 1: Deploy to Cloudflare** + +```bash +cd /Users/rtpa25/Developer/agent-os/apps/kernel && pnpm run deploy +``` + +Expected: wrangler deploys the worker; migration `0009_.sql` (from Task 1) applies automatically to the D1 instance the kernel uses. Watch for migration errors. + +- [ ] **Step 2: Destroy warm sandboxes** + +Per memory `project_sandbox_lifecycle_smoke`: warm sandboxes bake mount-time options and won't pick up new code. Destroy them before smoke testing. + +Run the project's sandbox-destroy command (check `apps/kernel/scripts/` or `wrangler durable-objects` commands in the repo for the exact incantation — typical pattern is a wrangler RPC call to a Kernel debug method). + +- [ ] **Step 3: Run the smoke checklist (all 8 scenarios)** + +For each scenario, run the TUI (`cd apps/cli && pnpm dev`), exercise the prompt, then verify against the DB or `tool_call` audit table. + +Scenario 1 — Empty thread shows no `` block + +``` +Prompt: "what time is it" +``` + +Verification: +```bash +# After the turn completes, in a fresh thread: +wrangler d1 execute --remote --command \ + "SELECT COUNT(*) FROM task WHERE \"threadId\" = ''" +``` + +Expected: count = 0. Agent answered without invoking manage_tasks. + +Scenario 2 — Trivial request doesn't trigger planning + +``` +Prompt: "Read README.md" +``` + +Verification: +```bash +wrangler d1 execute --remote --command \ + "SELECT toolName FROM tool_call WHERE \"threadId\" = '' ORDER BY \"startedAt\"" +``` + +Expected: tool calls include `read_file` (or similar) but NO `manage_tasks` row. + +Scenario 3 — Multi-step request triggers planning + +``` +Prompt: "Research bun vs deno, write a short summary doc, save it to summary.md" +``` + +Verification: +```bash +wrangler d1 execute --remote --command \ + "SELECT toolName, args FROM tool_call WHERE \"threadId\" = '' AND toolName = 'manage_tasks'" +``` + +Expected: at least one `manage_tasks` row with `args` containing `ops:[{action:"create",...}, ...]` and 2–3 tasks. Then: + +```bash +wrangler d1 execute --remote --command \ + "SELECT id, name, status FROM task WHERE \"threadId\" = ''" +``` + +Expected: 2–3 task rows with sensible names like "Research bun vs deno", "Write summary", "Save to summary.md". + +Scenario 4 — Mid-turn task state visibility (the prepareStep test) + +Same run as Scenario 3. After the agent creates tasks, the very NEXT assistant step should mark one in_progress (proving prepareStep refreshes the block — without it, the agent wouldn't see the tasks it just created in the context block). + +Verification: +```bash +wrangler d1 execute --remote --command \ + "SELECT id, status FROM task WHERE \"threadId\" = ''" +``` + +Expected: by mid-turn at least one task is `in_progress` (or already moving to `complete`). If all tasks stay `pending` while the agent appears to be executing them, prepareStep refresh is broken — go back to Task 4. + +Scenario 5 — Persistence across user turns + +After Scenario 3, in the same thread, send a second message: + +``` +Prompt: "what's the status of the doc?" +``` + +Verification: agent's response should reference the tasks from the prior turn (e.g. "Summary was completed; saved to summary.md"). DB check: + +```bash +wrangler d1 execute --remote --command \ + "SELECT id, name, status, result FROM task WHERE \"threadId\" = ''" +``` + +Expected: tasks from turn 1 still present with their final statuses. + +Scenario 6 — Cross-thread isolation + +In the TUI, start a new thread. First message: + +``` +Prompt: "hi" +``` + +Verification: +```bash +wrangler d1 execute --remote --command \ + "SELECT COUNT(*) FROM task WHERE \"threadId\" = ''" +``` + +Expected: count = 0. Agent's response should NOT reference tasks from the prior thread (the `` block is per-thread, omitted on new threads). + +Scenario 7 — Inline completion with result + +Re-run a multi-step request similar to Scenario 3. After completion: + +```bash +wrangler d1 execute --remote --command \ + "SELECT name, status, result FROM task WHERE \"threadId\" = '' AND status = 'complete'" +``` + +Expected: at least one completed task has a non-null `result` field with a 1–3 sentence summary. If `result` is uniformly null on completed tasks, the system prompt isn't landing — re-read Task 5's section and confirm the wording made it into `SYSTEM_PROMPT`. + +Scenario 8 — `manage_tasks` returns full state + +```bash +wrangler d1 execute --remote --command \ + "SELECT response FROM tool_call WHERE toolName = 'manage_tasks' ORDER BY \"startedAt\" DESC LIMIT 1" +``` + +Expected: the `response` JSON contains `{"tasks":[...]}` with multiple task entries (all of the thread's tasks at the time of that call, not just the affected ones). + +- [ ] **Step 4: If any scenario fails, add a follow-up commit** + +Don't open the PR until 1–8 are green. Common failure modes and where to look: + +| Failure | Look at | +|---|---| +| `` block never injected | Task 4 — `prepareStep` callback. Check `stableBlockCount` was captured and splice index is right. | +| Agent always plans, even for trivial | Task 5 — `` section. The "WHEN TO PLAN VS JUST DO" block isn't landing. | +| `result` always null on complete | Task 5 — same. The "Result is optional but strongly preferred" line isn't strong enough OR the model is following the optional-ness too literally. Strengthen wording. | +| `task` rows missing `runId` | Task 3 — `manage_tasks` execute. Check `ctx.runId` is being passed. | +| Cross-thread update succeeds silently | Task 3 — the `existing.threadId !== ctx.threadId` guard isn't being checked. | + +- [ ] **Step 5: Open the PR** + +```bash +gh pr create --title "feat(kernel): PR-C — task management" --body "$(cat <<'EOF' +Implements the task list substrate per +docs/superpowers/specs/2026-05-14-task-management-design.md. + +First of three slices: Tasks (this PR) → Sub-agents (PR-D) → Streaming (PR-E). + +## What ships + +- New per-thread \`task\` Drizzle table (5-state enum, cascade with thread + run) +- \`manage_tasks\` tool — single batchable write, returns full thread state +- \`\` context block — pure SQL renderer, status-aware truncation +- \`prepareStep\` wiring — refreshes \`\` per step so the model sees + fresh state across the react loop without busting b2 cache +- \`\` system prompt section — when-to-plan decision rule, + lifecycle, description quality, discipline, don'ts +- Anthropic provider options: \`toolStreaming\`, \`effort: "xhigh"\`, + \`sendReasoning\`, \`thinking: { adaptive, summarized }\`, explicit + \`cacheControl.ttl: "5m"\` at all 5 marker sites + +## Smoke results + +- ✅ Scenario 1: Empty thread renders no block +- ✅ Scenario 2: Trivial request doesn't trigger planning +- ✅ Scenario 3: Multi-step request creates 2-3 tasks via manage_tasks +- ✅ Scenario 4: Mid-turn state visibility — agent marks in_progress + in the step after creating tasks (prepareStep works) +- ✅ Scenario 5: Tasks persist across user turns within a thread +- ✅ Scenario 6: Cross-thread isolation — new thread has no tasks block +- ✅ Scenario 7: Inline completion writes 1-3 sentence result +- ✅ Scenario 8: manage_tasks returns full thread state in tool_result + +## Out of scope (PR-D / PR-E) + +- spawn_sub_agent tool, sub-agent runtime, parallel execution +- TUI rendering of as a sidebar/panel +- 2x result duplication handling (decided: accept; lands in PR-D) +- Nested tasks (parentTaskId) — YAGNI for now + +## Cache behavior + +b2 (stable context) stays hit across all within-turn steps because + lives in the dynamic suffix AFTER b2's marker. b4 misses on +steps that mutate tasks — acceptable trade (~300 token cost per step +vs ~5-10k of cache rebuild if tasks were in the b2 prefix). + +🤖 Generated with [Claude Code](https://claude.com/claude-code) +EOF +)" +``` + +- [ ] **Step 6: Mark this task complete in TodoWrite (subagent-driven flow only)** + +If you're running this plan via `subagent-driven-development`, mark Task 7 complete in the controller's TodoWrite. The two-stage review (spec compliance + code quality) for this task is satisfied by the smoke checklist results in Step 3. + +--- + +## Decisions encoded in this plan (cross-reference to spec §) + +| Decision | Spec § | Where it lands in code | +|---|---|---| +| Per-thread tasks, cascade with thread + run | §2 | Task 1 schema | +| 5-state status enum | §2 | Task 1 schema + Task 3 Zod schema | +| Single `manage_tasks` write tool (batchable ops) | §3 | Task 3 | +| Server-side injection of `threadId` / `runId` | §3 | Task 3 execute | +| Cross-thread `update` throws | §3 | Task 3 execute guard | +| Returns full thread task list sorted by createdAt ASC | §3 | Task 3 execute return | +| Status-aware description truncation | §4 | Task 2 renderer | +| `` omitted entirely on empty thread | §4 | Task 2 returns null | +| `prepareStep` owns the block render | §4 | Task 4 | +| Splice at `stableBlockCount` only | §4 / §6 | Task 4 | +| Tasks lives AFTER b2 marker | §6 | Task 4 (block injected after stable context) | +| Soft "one in_progress" rule (system prompt only) | §1 / §5 | Task 5 | +| `result` optional, strongly preferred | §5 | Task 5 | +| `toolStreaming`, `effort: "xhigh"`, `sendReasoning`, `thinking` | §9 | Task 6 | +| `cacheControl.ttl: "5m"` at all 5 sites | §9 | Task 6 | +| No `get_task` read tool (YAGNI) | §8 | Not implemented (intentional) | +| Empty `ops:[]` is no-op refresh | §3 | Task 3 (handled by the `for` loop falling through) | + +--- + +## Self-review notes + +After writing this plan, checked against spec: + +- **Spec coverage:** every numbered spec section maps to a task. §1 (architecture) is realized by Tasks 1–6 combined; §2 = Task 1; §3 = Task 3; §4 = Tasks 2 + 4; §5 = Task 5; §6 = enforced by Task 4's splice invariant; §7 = file inventory exactly matches Tasks 1–6; §8 = Task 7; §9 = Task 6. +- **Placeholder scan:** migration filename uses `0009__` because drizzle-kit names it — that's not a TBD, it's the tool's naming convention. Line numbers in Task 6 Step 1 are noted as "verify against output" because they may drift. +- **Type consistency:** `Task` type used uniformly (flat export, no `schema.Task`). `TaskJSON` interface in Task 3 matches the return shape promised in spec §3. `stableBlockCount` is the only cross-task identifier and it's local to `runChatTurn` (closure over `prepareStep`). +- **TDD note:** per memory `feedback_minimal_tests_in_plans`, no unit-test scaffolding. The smoke checklist in Task 7 covers behavior at the integration level. `pnpm typecheck` at the end of each task is the only "test" — it's free, fast, and catches the Drizzle/Zod/AI-SDK seams. From 349d724520fd053698321883a09495bdcb3636cb Mon Sep 17 00:00:00 2001 From: Ronit Date: Thu, 14 May 2026 16:30:42 +0530 Subject: [PATCH 03/11] feat(models): add task schema with thread+run FKs Per-thread tasks with 5-state status enum (pending/in_progress/complete/ failed/cancelled), nullable result, composite index on (threadId, createdAt). Cascade-delete with both thread and run. Co-Authored-By: Claude Opus 4.7 (1M context) --- .../migrations/0009_melodic_multiple_man.sql | 15 + .../migrations/meta/0009_snapshot.json | 779 ++++++++++++++++++ .../drizzle/migrations/meta/_journal.json | 7 + .../models/drizzle/migrations/migrations.js | 4 +- packages/models/drizzle/schemas/index.ts | 1 + packages/models/drizzle/schemas/tasks.ts | 43 + 6 files changed, 848 insertions(+), 1 deletion(-) create mode 100644 packages/models/drizzle/migrations/0009_melodic_multiple_man.sql create mode 100644 packages/models/drizzle/migrations/meta/0009_snapshot.json create mode 100644 packages/models/drizzle/schemas/tasks.ts diff --git a/packages/models/drizzle/migrations/0009_melodic_multiple_man.sql b/packages/models/drizzle/migrations/0009_melodic_multiple_man.sql new file mode 100644 index 0000000..1f43332 --- /dev/null +++ b/packages/models/drizzle/migrations/0009_melodic_multiple_man.sql @@ -0,0 +1,15 @@ +CREATE TABLE `task` ( + `id` text PRIMARY KEY NOT NULL, + `threadId` text NOT NULL, + `runId` text NOT NULL, + `name` text NOT NULL, + `description` text NOT NULL, + `status` text DEFAULT 'pending' NOT NULL, + `result` text, + `createdAt` integer NOT NULL, + `updatedAt` integer NOT NULL, + FOREIGN KEY (`threadId`) REFERENCES `thread`(`id`) ON UPDATE no action ON DELETE cascade, + FOREIGN KEY (`runId`) REFERENCES `run`(`id`) ON UPDATE no action ON DELETE cascade +); +--> statement-breakpoint +CREATE INDEX `idx_task_thread_created` ON `task` (`threadId`,`createdAt`); \ No newline at end of file diff --git a/packages/models/drizzle/migrations/meta/0009_snapshot.json b/packages/models/drizzle/migrations/meta/0009_snapshot.json new file mode 100644 index 0000000..c83efdf --- /dev/null +++ b/packages/models/drizzle/migrations/meta/0009_snapshot.json @@ -0,0 +1,779 @@ +{ + "version": "6", + "dialect": "sqlite", + "id": "d4860823-041b-4815-a656-bdace09a05c2", + "prevId": "0daebd77-51e1-4c04-8252-c6297bc3942c", + "tables": { + "git_merge_lock": { + "name": "git_merge_lock", + "columns": { + "id": { + "name": "id", + "type": "integer", + "primaryKey": true, + "notNull": true, + "autoincrement": false, + "default": 1 + }, + "holderCallId": { + "name": "holderCallId", + "type": "text", + "primaryKey": false, + "notNull": false, + "autoincrement": false + }, + "acquiredAt": { + "name": "acquiredAt", + "type": "integer", + "primaryKey": false, + "notNull": false, + "autoincrement": false + } + }, + "indexes": {}, + "foreignKeys": {}, + "compositePrimaryKeys": {}, + "uniqueConstraints": {}, + "checkConstraints": {} + }, + "thread": { + "name": "thread", + "columns": { + "id": { + "name": "id", + "type": "text", + "primaryKey": true, + "notNull": true, + "autoincrement": false + }, + "title": { + "name": "title", + "type": "text", + "primaryKey": false, + "notNull": false, + "autoincrement": false + }, + "createdAt": { + "name": "createdAt", + "type": "integer", + "primaryKey": false, + "notNull": true, + "autoincrement": false + }, + "updatedAt": { + "name": "updatedAt", + "type": "integer", + "primaryKey": false, + "notNull": true, + "autoincrement": false + }, + "lastMessageAt": { + "name": "lastMessageAt", + "type": "integer", + "primaryKey": false, + "notNull": false, + "autoincrement": false + }, + "summary": { + "name": "summary", + "type": "text", + "primaryKey": false, + "notNull": false, + "autoincrement": false + }, + "parentThreadId": { + "name": "parentThreadId", + "type": "text", + "primaryKey": false, + "notNull": false, + "autoincrement": false + } + }, + "indexes": { + "idx_thread_created_at": { + "name": "idx_thread_created_at", + "columns": [ + "createdAt" + ], + "isUnique": false + }, + "idx_thread_parent": { + "name": "idx_thread_parent", + "columns": [ + "parentThreadId" + ], + "isUnique": false + } + }, + "foreignKeys": { + "thread_parentThreadId_thread_id_fk": { + "name": "thread_parentThreadId_thread_id_fk", + "tableFrom": "thread", + "tableTo": "thread", + "columnsFrom": [ + "parentThreadId" + ], + "columnsTo": [ + "id" + ], + "onDelete": "no action", + "onUpdate": "no action" + } + }, + "compositePrimaryKeys": {}, + "uniqueConstraints": {}, + "checkConstraints": {} + }, + "run": { + "name": "run", + "columns": { + "id": { + "name": "id", + "type": "text", + "primaryKey": true, + "notNull": true, + "autoincrement": false + }, + "threadId": { + "name": "threadId", + "type": "text", + "primaryKey": false, + "notNull": true, + "autoincrement": false + }, + "status": { + "name": "status", + "type": "text", + "primaryKey": false, + "notNull": true, + "autoincrement": false, + "default": "'running'" + }, + "model": { + "name": "model", + "type": "text", + "primaryKey": false, + "notNull": false, + "autoincrement": false + }, + "startedAt": { + "name": "startedAt", + "type": "integer", + "primaryKey": false, + "notNull": true, + "autoincrement": false + }, + "finishedAt": { + "name": "finishedAt", + "type": "integer", + "primaryKey": false, + "notNull": false, + "autoincrement": false + }, + "error": { + "name": "error", + "type": "text", + "primaryKey": false, + "notNull": false, + "autoincrement": false + }, + "inputTokens": { + "name": "inputTokens", + "type": "integer", + "primaryKey": false, + "notNull": false, + "autoincrement": false + }, + "outputTokens": { + "name": "outputTokens", + "type": "integer", + "primaryKey": false, + "notNull": false, + "autoincrement": false + }, + "compactionApplied": { + "name": "compactionApplied", + "type": "integer", + "primaryKey": false, + "notNull": false, + "autoincrement": false + } + }, + "indexes": { + "idx_run_thread_started": { + "name": "idx_run_thread_started", + "columns": [ + "threadId", + "startedAt" + ], + "isUnique": false + } + }, + "foreignKeys": { + "run_threadId_thread_id_fk": { + "name": "run_threadId_thread_id_fk", + "tableFrom": "run", + "tableTo": "thread", + "columnsFrom": [ + "threadId" + ], + "columnsTo": [ + "id" + ], + "onDelete": "cascade", + "onUpdate": "no action" + } + }, + "compositePrimaryKeys": {}, + "uniqueConstraints": {}, + "checkConstraints": {} + }, + "message": { + "name": "message", + "columns": { + "id": { + "name": "id", + "type": "text", + "primaryKey": true, + "notNull": true, + "autoincrement": false + }, + "threadId": { + "name": "threadId", + "type": "text", + "primaryKey": false, + "notNull": true, + "autoincrement": false + }, + "runId": { + "name": "runId", + "type": "text", + "primaryKey": false, + "notNull": true, + "autoincrement": false + }, + "role": { + "name": "role", + "type": "text", + "primaryKey": false, + "notNull": true, + "autoincrement": false + }, + "parts": { + "name": "parts", + "type": "text", + "primaryKey": false, + "notNull": true, + "autoincrement": false + }, + "attachments": { + "name": "attachments", + "type": "text", + "primaryKey": false, + "notNull": true, + "autoincrement": false, + "default": "'[]'" + }, + "state": { + "name": "state", + "type": "text", + "primaryKey": false, + "notNull": true, + "autoincrement": false, + "default": "'complete'" + }, + "content_text": { + "name": "content_text", + "type": "text", + "primaryKey": false, + "notNull": true, + "autoincrement": false, + "default": "''" + }, + "indexed_at": { + "name": "indexed_at", + "type": "integer", + "primaryKey": false, + "notNull": false, + "autoincrement": false + }, + "createdAt": { + "name": "createdAt", + "type": "integer", + "primaryKey": false, + "notNull": true, + "autoincrement": false + } + }, + "indexes": { + "idx_message_thread_created": { + "name": "idx_message_thread_created", + "columns": [ + "threadId", + "createdAt" + ], + "isUnique": false + }, + "idx_message_run": { + "name": "idx_message_run", + "columns": [ + "runId" + ], + "isUnique": false + } + }, + "foreignKeys": { + "message_threadId_thread_id_fk": { + "name": "message_threadId_thread_id_fk", + "tableFrom": "message", + "tableTo": "thread", + "columnsFrom": [ + "threadId" + ], + "columnsTo": [ + "id" + ], + "onDelete": "cascade", + "onUpdate": "no action" + }, + "message_runId_run_id_fk": { + "name": "message_runId_run_id_fk", + "tableFrom": "message", + "tableTo": "run", + "columnsFrom": [ + "runId" + ], + "columnsTo": [ + "id" + ], + "onDelete": "cascade", + "onUpdate": "no action" + } + }, + "compositePrimaryKeys": {}, + "uniqueConstraints": {}, + "checkConstraints": {} + }, + "task": { + "name": "task", + "columns": { + "id": { + "name": "id", + "type": "text", + "primaryKey": true, + "notNull": true, + "autoincrement": false + }, + "threadId": { + "name": "threadId", + "type": "text", + "primaryKey": false, + "notNull": true, + "autoincrement": false + }, + "runId": { + "name": "runId", + "type": "text", + "primaryKey": false, + "notNull": true, + "autoincrement": false + }, + "name": { + "name": "name", + "type": "text", + "primaryKey": false, + "notNull": true, + "autoincrement": false + }, + "description": { + "name": "description", + "type": "text", + "primaryKey": false, + "notNull": true, + "autoincrement": false + }, + "status": { + "name": "status", + "type": "text", + "primaryKey": false, + "notNull": true, + "autoincrement": false, + "default": "'pending'" + }, + "result": { + "name": "result", + "type": "text", + "primaryKey": false, + "notNull": false, + "autoincrement": false + }, + "createdAt": { + "name": "createdAt", + "type": "integer", + "primaryKey": false, + "notNull": true, + "autoincrement": false + }, + "updatedAt": { + "name": "updatedAt", + "type": "integer", + "primaryKey": false, + "notNull": true, + "autoincrement": false + } + }, + "indexes": { + "idx_task_thread_created": { + "name": "idx_task_thread_created", + "columns": [ + "threadId", + "createdAt" + ], + "isUnique": false + } + }, + "foreignKeys": { + "task_threadId_thread_id_fk": { + "name": "task_threadId_thread_id_fk", + "tableFrom": "task", + "tableTo": "thread", + "columnsFrom": [ + "threadId" + ], + "columnsTo": [ + "id" + ], + "onDelete": "cascade", + "onUpdate": "no action" + }, + "task_runId_run_id_fk": { + "name": "task_runId_run_id_fk", + "tableFrom": "task", + "tableTo": "run", + "columnsFrom": [ + "runId" + ], + "columnsTo": [ + "id" + ], + "onDelete": "cascade", + "onUpdate": "no action" + } + }, + "compositePrimaryKeys": {}, + "uniqueConstraints": {}, + "checkConstraints": {} + }, + "tool_call": { + "name": "tool_call", + "columns": { + "id": { + "name": "id", + "type": "text", + "primaryKey": true, + "notNull": true, + "autoincrement": false + }, + "threadId": { + "name": "threadId", + "type": "text", + "primaryKey": false, + "notNull": true, + "autoincrement": false + }, + "runId": { + "name": "runId", + "type": "text", + "primaryKey": false, + "notNull": true, + "autoincrement": false + }, + "messageId": { + "name": "messageId", + "type": "text", + "primaryKey": false, + "notNull": false, + "autoincrement": false + }, + "toolName": { + "name": "toolName", + "type": "text", + "primaryKey": false, + "notNull": true, + "autoincrement": false + }, + "args": { + "name": "args", + "type": "text", + "primaryKey": false, + "notNull": true, + "autoincrement": false + }, + "status": { + "name": "status", + "type": "text", + "primaryKey": false, + "notNull": true, + "autoincrement": false, + "default": "'running'" + }, + "response": { + "name": "response", + "type": "text", + "primaryKey": false, + "notNull": false, + "autoincrement": false + }, + "error": { + "name": "error", + "type": "text", + "primaryKey": false, + "notNull": false, + "autoincrement": false + }, + "startedAt": { + "name": "startedAt", + "type": "integer", + "primaryKey": false, + "notNull": true, + "autoincrement": false + }, + "finishedAt": { + "name": "finishedAt", + "type": "integer", + "primaryKey": false, + "notNull": false, + "autoincrement": false + } + }, + "indexes": { + "idx_tool_call_run": { + "name": "idx_tool_call_run", + "columns": [ + "runId" + ], + "isUnique": false + }, + "idx_tool_call_thread": { + "name": "idx_tool_call_thread", + "columns": [ + "threadId" + ], + "isUnique": false + }, + "idx_tool_call_message": { + "name": "idx_tool_call_message", + "columns": [ + "messageId" + ], + "isUnique": false + } + }, + "foreignKeys": { + "tool_call_threadId_thread_id_fk": { + "name": "tool_call_threadId_thread_id_fk", + "tableFrom": "tool_call", + "tableTo": "thread", + "columnsFrom": [ + "threadId" + ], + "columnsTo": [ + "id" + ], + "onDelete": "cascade", + "onUpdate": "no action" + }, + "tool_call_runId_run_id_fk": { + "name": "tool_call_runId_run_id_fk", + "tableFrom": "tool_call", + "tableTo": "run", + "columnsFrom": [ + "runId" + ], + "columnsTo": [ + "id" + ], + "onDelete": "cascade", + "onUpdate": "no action" + }, + "tool_call_messageId_message_id_fk": { + "name": "tool_call_messageId_message_id_fk", + "tableFrom": "tool_call", + "tableTo": "message", + "columnsFrom": [ + "messageId" + ], + "columnsTo": [ + "id" + ], + "onDelete": "cascade", + "onUpdate": "no action" + } + }, + "compositePrimaryKeys": {}, + "uniqueConstraints": {}, + "checkConstraints": {} + }, + "integration_connection": { + "name": "integration_connection", + "columns": { + "id": { + "name": "id", + "type": "text", + "primaryKey": true, + "notNull": true, + "autoincrement": false + }, + "app_slug": { + "name": "app_slug", + "type": "text", + "primaryKey": false, + "notNull": true, + "autoincrement": false + }, + "account_id": { + "name": "account_id", + "type": "text", + "primaryKey": false, + "notNull": true, + "autoincrement": false + }, + "display_name": { + "name": "display_name", + "type": "text", + "primaryKey": false, + "notNull": true, + "autoincrement": false + }, + "connected_at": { + "name": "connected_at", + "type": "integer", + "primaryKey": false, + "notNull": true, + "autoincrement": false + } + }, + "indexes": { + "idx_conn_account_id": { + "name": "idx_conn_account_id", + "columns": [ + "account_id" + ], + "isUnique": true + }, + "idx_conn_app_slug": { + "name": "idx_conn_app_slug", + "columns": [ + "app_slug" + ], + "isUnique": false + } + }, + "foreignKeys": {}, + "compositePrimaryKeys": {}, + "uniqueConstraints": {}, + "checkConstraints": {} + }, + "sandbox": { + "name": "sandbox", + "columns": { + "id": { + "name": "id", + "type": "text", + "primaryKey": true, + "notNull": true, + "autoincrement": false + }, + "threadId": { + "name": "threadId", + "type": "text", + "primaryKey": false, + "notNull": true, + "autoincrement": false + }, + "name": { + "name": "name", + "type": "text", + "primaryKey": false, + "notNull": true, + "autoincrement": false + }, + "cfSandboxName": { + "name": "cfSandboxName", + "type": "text", + "primaryKey": false, + "notNull": true, + "autoincrement": false + }, + "status": { + "name": "status", + "type": "text", + "primaryKey": false, + "notNull": true, + "autoincrement": false, + "default": "'active'" + }, + "createdAt": { + "name": "createdAt", + "type": "integer", + "primaryKey": false, + "notNull": true, + "autoincrement": false + }, + "lastUsedAt": { + "name": "lastUsedAt", + "type": "integer", + "primaryKey": false, + "notNull": true, + "autoincrement": false + }, + "mountedAt": { + "name": "mountedAt", + "type": "integer", + "primaryKey": false, + "notNull": false, + "autoincrement": false + } + }, + "indexes": { + "sandbox_active_name_idx": { + "name": "sandbox_active_name_idx", + "columns": [ + "threadId", + "name" + ], + "isUnique": true, + "where": "\"sandbox\".\"status\" = 'active'" + } + }, + "foreignKeys": { + "sandbox_threadId_thread_id_fk": { + "name": "sandbox_threadId_thread_id_fk", + "tableFrom": "sandbox", + "tableTo": "thread", + "columnsFrom": [ + "threadId" + ], + "columnsTo": [ + "id" + ], + "onDelete": "cascade", + "onUpdate": "no action" + } + }, + "compositePrimaryKeys": {}, + "uniqueConstraints": {}, + "checkConstraints": {} + } + }, + "views": {}, + "enums": {}, + "_meta": { + "schemas": {}, + "tables": {}, + "columns": {} + }, + "internal": { + "indexes": {} + } +} \ No newline at end of file diff --git a/packages/models/drizzle/migrations/meta/_journal.json b/packages/models/drizzle/migrations/meta/_journal.json index 3027d6e..d5a8794 100644 --- a/packages/models/drizzle/migrations/meta/_journal.json +++ b/packages/models/drizzle/migrations/meta/_journal.json @@ -64,6 +64,13 @@ "when": 1778668822574, "tag": "0008_special_grandmaster", "breakpoints": true + }, + { + "idx": 9, + "version": "6", + "when": 1778756403461, + "tag": "0009_melodic_multiple_man", + "breakpoints": true } ] } \ No newline at end of file diff --git a/packages/models/drizzle/migrations/migrations.js b/packages/models/drizzle/migrations/migrations.js index e3573b6..260e95e 100644 --- a/packages/models/drizzle/migrations/migrations.js +++ b/packages/models/drizzle/migrations/migrations.js @@ -8,6 +8,7 @@ import m0005 from './0005_right_white_queen.sql'; import m0006 from './0006_stiff_mindworm.sql'; import m0007 from './0007_brave_azazel.sql'; import m0008 from './0008_special_grandmaster.sql'; +import m0009 from './0009_melodic_multiple_man.sql'; export default { journal, @@ -20,7 +21,8 @@ m0004, m0005, m0006, m0007, -m0008 +m0008, +m0009 } } \ No newline at end of file diff --git a/packages/models/drizzle/schemas/index.ts b/packages/models/drizzle/schemas/index.ts index 589761b..7428eb8 100644 --- a/packages/models/drizzle/schemas/index.ts +++ b/packages/models/drizzle/schemas/index.ts @@ -1,6 +1,7 @@ export * from "./threads"; export * from "./runs"; export * from "./messages"; +export * from "./tasks"; export * from "./tool-calls"; export * from "./integration-connections"; export * from "./sandbox"; diff --git a/packages/models/drizzle/schemas/tasks.ts b/packages/models/drizzle/schemas/tasks.ts new file mode 100644 index 0000000..fcb16ac --- /dev/null +++ b/packages/models/drizzle/schemas/tasks.ts @@ -0,0 +1,43 @@ +import { InferSelectModel } from "drizzle-orm"; +import { index, integer, sqliteTable, text } from "drizzle-orm/sqlite-core"; + +import { run } from "./runs"; +import { thread } from "./threads"; + +export const task = sqliteTable( + "task", + { + id: text().primaryKey(), // ULID + threadId: text() + .notNull() + .references(() => thread.id, { onDelete: "cascade" }), + // The run that CREATED this task (manage_tasks{action:"create"} call). + // Updates from later runs do NOT change this — preserves "who planned this". + runId: text() + .notNull() + .references(() => run.id, { onDelete: "cascade" }), + name: text().notNull(), // user-visible label + description: text().notNull(), // full plan: deliverables, deps, expected output + status: text({ + enum: ["pending", "in_progress", "complete", "failed", "cancelled"] + }) + .notNull() + .default("pending"), + // Set on transition to a terminal status: + // complete → the actual output/answer + // failed → the reason the agent gave up + // cancelled → why it was abandoned + result: text(), + createdAt: integer({ mode: "timestamp_ms" }) + .notNull() + .$defaultFn(() => new Date()), + updatedAt: integer({ mode: "timestamp_ms" }) + .notNull() + .$defaultFn(() => new Date()) + }, + (table) => [ + index("idx_task_thread_created").on(table.threadId, table.createdAt) + ] +); + +export type Task = InferSelectModel; From b7c8946a317a07b0ef2d0f334bd67deb87f6c98d Mon Sep 17 00:00:00 2001 From: Ronit Date: Thu, 14 May 2026 16:34:41 +0530 Subject: [PATCH 04/11] fix(models): use $onUpdate for task.updatedAt auto-bump MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit $defaultFn only fires on INSERT; without $onUpdate the column drifts stale unless every writer remembers to set it manually. Match the thread.updatedAt precedent (threads.ts:13-14). Spec §2 corrected to match — keeps PR-D / PR-E writers honest. Co-Authored-By: Claude Opus 4.7 (1M context) --- .../specs/2026-05-14-task-management-design.md | 9 +++++---- packages/models/drizzle/schemas/tasks.ts | 1 + 2 files changed, 6 insertions(+), 4 deletions(-) diff --git a/docs/superpowers/specs/2026-05-14-task-management-design.md b/docs/superpowers/specs/2026-05-14-task-management-design.md index 8129d7d..1c41cde 100644 --- a/docs/superpowers/specs/2026-05-14-task-management-design.md +++ b/docs/superpowers/specs/2026-05-14-task-management-design.md @@ -108,6 +108,7 @@ export const task = sqliteTable( updatedAt: integer({ mode: "timestamp_ms" }) .notNull() .$defaultFn(() => new Date()) + .$onUpdate(() => new Date()) }, (table) => [ index("idx_task_thread_created").on(table.threadId, table.createdAt) @@ -127,7 +128,7 @@ export type Task = InferSelectModel; | `name` / `description` | Both `notNull` | Forcing description-on-create prevents "task #4: do the thing" planning that's useless on retry. | | `status` enum | 5 values | `pending → in_progress → {complete | failed | cancelled}`. Failed vs cancelled distinction: failed = tried and gave up; cancelled = decided not to attempt. Useful signal. | | `result` | nullable text | NULL for pending/in_progress; set on terminal transitions. | -| `updatedAt` | `notNull`, $defaultFn on insert | Tool code explicitly sets it on every update (don't lean on Drizzle's `$onUpdate` since it only fires through `.update()` calls — explicit is safer). | +| `updatedAt` | `notNull`, `$defaultFn` on insert + `$onUpdate` bump | Drizzle's `$onUpdate` re-fires on every `.update()` call — same precedent as `thread.updatedAt` (threads.ts:13-14). Keeps PR-D / PR-E writers from silently drifting stale. | | Index `(threadId, createdAt)` | Single composite | Only access pattern: "fetch all tasks for this thread ordered by creation." | **Deliberately omitted:** @@ -192,7 +193,7 @@ The LLM never sees `threadId` or `runId`. They are injected server-side from `pe | `action` | LLM-provided | Server-injected | Effect | |---|---|---|---| | `create` | `name`, `description` | `id = ulid()`, `threadId`, `runId`, `status = "pending"`, timestamps | INSERT new row. | -| `update` | `id`, plus ≥1 of name/description/status/result | (verifies ownership) | Fetch existing task. If `threadId ≠ ctx.threadId`, throw. Else UPDATE the supplied fields + `updatedAt = now`. | +| `update` | `id`, plus ≥1 of name/description/status/result | (verifies ownership) | Fetch existing task. If `threadId ≠ ctx.threadId`, throw. Else UPDATE the supplied fields. `updatedAt` auto-bumps via Drizzle `$onUpdate`. | **Status transitions** are unrestricted at the tool level (any → any allowed). System prompt guides the expected flow; the tool doesn't gate. @@ -225,8 +226,8 @@ execute: async (args, ctx) => { ...(op.name !== undefined && { name: op.name }), ...(op.description !== undefined && { description: op.description }), ...(op.status !== undefined && { status: op.status }), - ...(op.result !== undefined && { result: op.result }), - updatedAt: new Date() + ...(op.result !== undefined && { result: op.result }) + // updatedAt auto-bumped by Drizzle $onUpdate hook }) .where(eq(schema.task.id, op.id)); } diff --git a/packages/models/drizzle/schemas/tasks.ts b/packages/models/drizzle/schemas/tasks.ts index fcb16ac..dc65609 100644 --- a/packages/models/drizzle/schemas/tasks.ts +++ b/packages/models/drizzle/schemas/tasks.ts @@ -34,6 +34,7 @@ export const task = sqliteTable( updatedAt: integer({ mode: "timestamp_ms" }) .notNull() .$defaultFn(() => new Date()) + .$onUpdate(() => new Date()) }, (table) => [ index("idx_task_thread_created").on(table.threadId, table.createdAt) From 30b0c497640f0186eb3e99e62afc4d31032240f7 Mon Sep 17 00:00:00 2001 From: Ronit Date: Thu, 14 May 2026 16:50:27 +0530 Subject: [PATCH 05/11] feat(kernel): tasks-block renderer for context block Pure SQL fetch from task table. Status-aware description truncation (in_progress full / pending 120 / terminal 80 chars), result shown in full for terminal entries when non-null. Returns null on empty thread so the block is omitted from prompts entirely. Consumed by prepareStep callback in turn.ts (Task 4). --- apps/kernel/src/agent/tasks-block.ts | 96 ++++++++++++++++++++++++++++ 1 file changed, 96 insertions(+) create mode 100644 apps/kernel/src/agent/tasks-block.ts diff --git a/apps/kernel/src/agent/tasks-block.ts b/apps/kernel/src/agent/tasks-block.ts new file mode 100644 index 0000000..3aa2c40 --- /dev/null +++ b/apps/kernel/src/agent/tasks-block.ts @@ -0,0 +1,96 @@ +import { asc, eq, schema, type DB } from "@agent-os/models"; + +const PENDING_DESC_MAX = 120; +const TERMINAL_DESC_MAX = 80; + +/** + * Collapse whitespace and hard-truncate to `n` chars. Caller appends the + * ellipsis when `truncated === true`. + */ +function truncate(s: string, n: number): { text: string; truncated: boolean } { + const trimmed = s.replace(/\s+/g, " ").trim(); + if (trimmed.length <= n) return { text: trimmed, truncated: false }; + return { text: trimmed.slice(0, n), truncated: true }; +} + +function formatDescription( + status: "pending" | "in_progress" | "complete" | "failed" | "cancelled", + description: string +): string { + if (status === "in_progress") { + return description.replace(/\s+/g, " ").trim(); + } + const max = status === "pending" ? PENDING_DESC_MAX : TERMINAL_DESC_MAX; + const { text, truncated } = truncate(description, max); + return truncated ? `${text}…` : text; +} + +/** + * Build the `` block. Returns `null` when the thread has zero tasks + * (block is omitted entirely from the prompt — same pattern as + * ). + * + * Called from `prepareStep` in `turn.ts` on every step boundary inside + * `streamText`. Pure SQL fetch — one index hit on `idx_task_thread_created`. + * + * Truncation policy: + * in_progress → full description (active focus needs the plan) + * pending → trimmed to 120 chars + ellipsis if truncated + * complete/failed/cancelled → trimmed to 80 chars; result shown in full + * when non-null + * + * Full descriptions remain in the agent's conversation history via every + * `manage_tasks` tool_result — truncation here is per-turn cache hygiene, + * not lossy storage. + * + * Failure mode: SQL throw → error placeholder so the block still ships and + * the agent isn't blind to its task list breaking. + */ +export async function buildTasksBlock(args: { + db: DB; + threadId: string; +}): Promise { + let rows; + try { + rows = await args.db.query.task.findMany({ + where: eq(schema.task.threadId, args.threadId), + orderBy: [asc(schema.task.createdAt)] + }); + } catch (e) { + console.warn("[tasks-block] SQL failed", { + threadId: args.threadId, + error: e instanceof Error ? e.message : String(e) + }); + return [ + "", + "(error loading task list — manage_tasks may still work, try it)", + "" + ].join("\n"); + } + + if (rows.length === 0) return null; + + const lines = rows.map((t) => { + const head = `- ${t.id} [${t.status}] "${t.name}"`; + const plan = ` plan: ${formatDescription(t.status, t.description)}`; + const isTerminal = + t.status === "complete" || + t.status === "failed" || + t.status === "cancelled"; + if (isTerminal && t.result) { + const result = t.result.replace(/\s+/g, " ").trim(); + return [head, plan, ` result: ${result}`].join("\n"); + } + return [head, plan].join("\n"); + }); + + return [ + "", + "Plan and track multi-step work here. Use manage_tasks to create/update tasks.", + "Convention: one task in_progress at a time. Mark complete/failed/cancelled", + "before moving on. Skip task planning for trivial single-step requests.", + "", + ...lines, + "" + ].join("\n"); +} From 0f399449fde3d9c5ec06857d8a7f8a8200fb8586 Mon Sep 17 00:00:00 2001 From: Ronit Date: Thu, 14 May 2026 16:54:26 +0530 Subject: [PATCH 06/11] refactor(kernel): dedupe status union + terminal set in tasks-block Use Task['status'] instead of inline 5-state union, extract isTerminalStatus() helper so formatDescription and the row map share one source of truth for terminal-vs-non-terminal. Add explicit Task[] type on the rows local. Co-Authored-By: Claude Opus 4.7 (1M context) --- apps/kernel/src/agent/tasks-block.ts | 19 +++++++++++-------- 1 file changed, 11 insertions(+), 8 deletions(-) diff --git a/apps/kernel/src/agent/tasks-block.ts b/apps/kernel/src/agent/tasks-block.ts index 3aa2c40..5c2b06c 100644 --- a/apps/kernel/src/agent/tasks-block.ts +++ b/apps/kernel/src/agent/tasks-block.ts @@ -1,8 +1,14 @@ -import { asc, eq, schema, type DB } from "@agent-os/models"; +import { asc, eq, schema, type DB, type Task } from "@agent-os/models"; const PENDING_DESC_MAX = 120; const TERMINAL_DESC_MAX = 80; +const TERMINAL_STATUSES = ["complete", "failed", "cancelled"] as const; + +function isTerminalStatus(status: Task["status"]): boolean { + return (TERMINAL_STATUSES as readonly string[]).includes(status); +} + /** * Collapse whitespace and hard-truncate to `n` chars. Caller appends the * ellipsis when `truncated === true`. @@ -14,13 +20,13 @@ function truncate(s: string, n: number): { text: string; truncated: boolean } { } function formatDescription( - status: "pending" | "in_progress" | "complete" | "failed" | "cancelled", + status: Task["status"], description: string ): string { if (status === "in_progress") { return description.replace(/\s+/g, " ").trim(); } - const max = status === "pending" ? PENDING_DESC_MAX : TERMINAL_DESC_MAX; + const max = isTerminalStatus(status) ? TERMINAL_DESC_MAX : PENDING_DESC_MAX; const { text, truncated } = truncate(description, max); return truncated ? `${text}…` : text; } @@ -50,7 +56,7 @@ export async function buildTasksBlock(args: { db: DB; threadId: string; }): Promise { - let rows; + let rows: Task[]; try { rows = await args.db.query.task.findMany({ where: eq(schema.task.threadId, args.threadId), @@ -73,10 +79,7 @@ export async function buildTasksBlock(args: { const lines = rows.map((t) => { const head = `- ${t.id} [${t.status}] "${t.name}"`; const plan = ` plan: ${formatDescription(t.status, t.description)}`; - const isTerminal = - t.status === "complete" || - t.status === "failed" || - t.status === "cancelled"; + const isTerminal = isTerminalStatus(t.status); if (isTerminal && t.result) { const result = t.result.replace(/\s+/g, " ").trim(); return [head, plan, ` result: ${result}`].join("\n"); From 44118522c35050415e1e5f6c762943de686724ea Mon Sep 17 00:00:00 2001 From: Ronit Date: Thu, 14 May 2026 16:59:40 +0530 Subject: [PATCH 07/11] feat(kernel): manage_tasks tool with batchable ops Single write tool: ops:[{action:'create',name,description}|{action:'update', id,...patch}]. Returns full thread task list sorted by createdAt ASC. threadId/runId injected server-side from perTurn; cross-thread updates throw. Wrapped via wrappedTool({touchesFS:false}) so audit rows write without per-call worktree overhead. Co-Authored-By: Claude Opus 4.7 (1M context) --- apps/kernel/src/tools/index.ts | 2 + apps/kernel/src/tools/tasks.ts | 148 +++++++++++++++++++++++++++++++++ 2 files changed, 150 insertions(+) create mode 100644 apps/kernel/src/tools/tasks.ts diff --git a/apps/kernel/src/tools/index.ts b/apps/kernel/src/tools/index.ts index 7f04fac..0fc4f51 100644 --- a/apps/kernel/src/tools/index.ts +++ b/apps/kernel/src/tools/index.ts @@ -4,6 +4,7 @@ import { buildFSTools } from "./fs-tools"; import { buildGetTimeTool } from "./get-time"; import { buildProcessAttachmentTool } from "./process-attachment"; import { buildSessionsTools } from "./sessions"; +import { buildTasksTools } from "./tasks"; import { buildWebSearchTools } from "./web-search"; import { type PerTurnContext } from "./wrapped-tool"; @@ -50,5 +51,6 @@ export const buildTools = (perTurn: PerTurnContext) => ...buildExecCodeTool(perTurn), ...buildComputerUseTools(perTurn), ...buildSessionsTools(perTurn), + ...buildTasksTools(perTurn), get_time: buildGetTimeTool(perTurn) }); diff --git a/apps/kernel/src/tools/tasks.ts b/apps/kernel/src/tools/tasks.ts new file mode 100644 index 0000000..0c61a52 --- /dev/null +++ b/apps/kernel/src/tools/tasks.ts @@ -0,0 +1,148 @@ +import { asc, eq, schema, ulid } from "@agent-os/models"; +import { z } from "zod"; + +import { wrappedTool, type PerTurnContext } from "./wrapped-tool"; + +const TASK_STATUS = [ + "pending", + "in_progress", + "complete", + "failed", + "cancelled" +] as const; + +const createOpSchema = z.object({ + action: z.literal("create"), + name: z.string().min(1).describe("Short user-visible label"), + description: z + .string() + .min(1) + .describe("Full plan — deliverables, deps, expected output") +}); + +const updateOpSchema = z.object({ + action: z.literal("update"), + id: z.string().describe("Task id (ULID) to update"), + name: z.string().min(1).optional(), + description: z.string().min(1).optional(), + status: z.enum(TASK_STATUS).optional(), + result: z + .string() + .optional() + .describe("Set on terminal status with 1-3 sentence summary") +}); + +const inputSchema = z.object({ + ops: z + .array(z.discriminatedUnion("action", [createOpSchema, updateOpSchema])) + .describe( + "Batch of create/update operations. Applied in order. Empty array is a no-op refresh." + ) +}); + +type ManageTasksArgs = z.infer; + +export interface TaskJSON { + id: string; + name: string; + description: string; + status: (typeof TASK_STATUS)[number]; + result: string | null; + createdAt: string; + updatedAt: string; +} + +const DESCRIPTION = [ + "Plan and track multi-step work as tasks. Use ONE call to batch creates +", + "updates. Returns the full thread task list (sorted by createdAt ASC) so", + "you see canonical state immediately.", + "", + "When to use: multi-step requests (research X, write Y, post Y). Skip for", + "trivial single-step requests.", + "", + "Ops:", + " create: { action:'create', name, description }", + " update: { action:'update', id, ...any of: name, description, status, result }", + "", + "Status values: pending | in_progress | complete | failed | cancelled", + "Status transitions are unrestricted (any → any).", + "", + "Set 'result' on terminal status (complete/failed/cancelled) with a 1-3", + "sentence summary. Optional but strongly preferred — it's how future turns", + "see what was accomplished." +].join("\n"); + +export function buildTasksTools(perTurn: PerTurnContext) { + return { + manage_tasks: wrappedTool( + { + name: "manage_tasks", + description: DESCRIPTION, + inputSchema, + touchesFS: false, + needsApproval: false, + execute: async (args: ManageTasksArgs, ctx) => { + for (const op of args.ops) { + if (op.action === "create") { + await ctx.db.insert(schema.task).values({ + id: ulid(), + threadId: ctx.threadId, + runId: ctx.runId, + name: op.name, + description: op.description + // status defaults to "pending" + // createdAt/updatedAt default via $defaultFn + }); + } else { + // op.action === "update" + const existing = await ctx.db.query.task.findFirst({ + where: eq(schema.task.id, op.id) + }); + if (!existing) { + throw new Error(`Task ${op.id} not found`); + } + if (existing.threadId !== ctx.threadId) { + // Defensive: model gave us an id from another thread. + // Refuse rather than silently mutate. + throw new Error( + `Task ${op.id} belongs to a different thread` + ); + } + await ctx.db + .update(schema.task) + .set({ + ...(op.name !== undefined && { name: op.name }), + ...(op.description !== undefined && { + description: op.description + }), + ...(op.status !== undefined && { status: op.status }), + ...(op.result !== undefined && { result: op.result }) + // updatedAt auto-bumped by Drizzle $onUpdate hook on task.updatedAt + }) + .where(eq(schema.task.id, op.id)); + } + } + + // Return full updated task list sorted by createdAt ASC. + const tasks = await ctx.db.query.task.findMany({ + where: eq(schema.task.threadId, ctx.threadId), + orderBy: [asc(schema.task.createdAt)] + }); + const result: { tasks: TaskJSON[] } = { + tasks: tasks.map((t) => ({ + id: t.id, + name: t.name, + description: t.description, + status: t.status, + result: t.result, + createdAt: t.createdAt.toISOString(), + updatedAt: t.updatedAt.toISOString() + })) + }; + return result; + } + }, + perTurn + ) + }; +} From c526d971c28ed526613d9cf8c78351c1c250fde8 Mon Sep 17 00:00:00 2001 From: Ronit Date: Thu, 14 May 2026 17:06:57 +0530 Subject: [PATCH 08/11] feat(kernel): wire prepareStep to refresh block per step streamText runs the entire react loop internally; without prepareStep the context blocks would be frozen at turn start and the agent would see stale task state across iterations. Capture stableBlockCount after buildContext- Messages and splice the freshly-rendered block at that index every step. Splice invariant: only touches indices >= stableBlockCount, leaving the b2-cached stable prefix byte-identical so b2 keeps hitting across steps. Costs ~300 tokens of tasks-render per step vs ~5-10k of cache rebuild if tasks were inside the b2 prefix. --- apps/kernel/src/turn.ts | 49 +++++++++++++++++++++++++++++++++++++++++ 1 file changed, 49 insertions(+) diff --git a/apps/kernel/src/turn.ts b/apps/kernel/src/turn.ts index 52f5581..c924ddb 100644 --- a/apps/kernel/src/turn.ts +++ b/apps/kernel/src/turn.ts @@ -8,6 +8,7 @@ import type { Attachment } from "@agent-os/protocol"; import { buildContextMessages } from "./agent/context-messages"; import { SYSTEM_PROMPT } from "./agent/system-prompt"; +import { buildTasksBlock } from "./agent/tasks-block"; import { broadcastRunStatus, broadcastToThread, send } from "./broadcasts"; import type { ConnectionState, Kernel } from "./kernel"; import { MODEL_ID } from "./kernel"; @@ -229,6 +230,12 @@ export async function runChatTurn( imageContentTypeRegex: IMAGE_MIME_RE, clientTimezone }); + // Captured ONCE per turn. The prepareStep callback below splices the + // freshly-rendered block at this index every step boundary, so + // the value must stay frozen at the turn-start stable-prefix length. + // See spec §6: this index is the boundary between b2-cached stable + // prefix (indices < stableBlockCount) and the per-step refresh zone. + const stableBlockCount = contextMessages.length; // Breakpoint b2: cache through the end of the context-message stack. // Within a thread, this prefix (system + previous_summary + connected @@ -285,6 +292,48 @@ export async function runChatTurn( } } satisfies AnthropicLanguageModelOptions }, + // Refresh block on every step boundary. Owned entirely by this + // callback — never by buildContextMessages — so the block stays fresh as + // the agent mutates task state across react-loop iterations inside this + // streamText invocation. + // + // Bytes live AFTER b2 cache marker (b2 is on , the last + // stable context block), so the refresh costs ~300 tokens per step and + // does NOT bust b2. b4 misses on within-turn steps that change tasks; + // acceptable trade — see spec §6 worked example. + // + // Splice invariant: only touches indices >= stableBlockCount. Touching + // earlier indices would mutate the b2-cached prefix and bust caching. + prepareStep: async ({ stepNumber, messages }) => { + const working = [...messages]; + + // On step 1+, strip the prior step's tasks block. It always lives at + // index stableBlockCount IF buildTasksBlock returned non-null last step. + // Detected by role + literal "" tag prefix. + if (stepNumber > 0) { + const candidate = working[stableBlockCount]; + if ( + candidate?.role === "user" && + typeof candidate.content === "string" && + candidate.content.startsWith("") + ) { + working.splice(stableBlockCount, 1); + } + } + + const tasksBlock = await buildTasksBlock({ + db: kernel.db, + threadId + }); + if (tasksBlock !== null) { + working.splice(stableBlockCount, 0, { + role: "user" as const, + content: tasksBlock + }); + } + + return { messages: working }; + }, onStepFinish: () => { // Step boundary: if the user requested pause during the step, fire // the actual abort now. Step's parts are all in terminal states, From d7c402a83f47cdd309b32583e19e4b3a2e53d35f Mon Sep 17 00:00:00 2001 From: Ronit Date: Thu, 14 May 2026 17:11:20 +0530 Subject: [PATCH 09/11] feat(kernel): system prompt section MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Teaches when to plan vs just do (substantial vs trivial test), lifecycle with status transitions, description quality bar with good/bad pair, discipline rules (one in_progress, immediate status updates), and don'ts. ~370 tokens, cached in b1 — paid once per fresh cache, then free. Lifted from Dimension AI's / framing, trimmed to exclude sub-agent guidance (PR-D scope). --- apps/kernel/src/agent/system-prompt.ts | 56 ++++++++++++++++++++++++++ 1 file changed, 56 insertions(+) diff --git a/apps/kernel/src/agent/system-prompt.ts b/apps/kernel/src/agent/system-prompt.ts index b174ad6..3c1aba9 100644 --- a/apps/kernel/src/agent/system-prompt.ts +++ b/apps/kernel/src/agent/system-prompt.ts @@ -212,6 +212,62 @@ Filename is kebab-case under \`r2://skills/\`. Cap the body at roughly 80 lines Mention skill creations in your reply with one line ("Saved \`skills/.md\`"); skip mentioning memory writes. + +For multi-step requests, plan and track your work with manage_tasks. The + context block (when present) is your source of truth for task state. +It refreshes at every step boundary so it never lags behind your tool calls. + +WHEN TO PLAN VS JUST DO +The question to ask yourself: is the work substantial, or trivial? +- "Read this file and tell me what it says" → just do it. One tool call. +- "What time is it in Tokyo" → just do it. +- "Research X, write Y, post Y to slack" → plan tasks first. +- "Refactor this module" if it's one file → just do it. +- "Refactor this module" if it spans many files → plan tasks first. + +Planning costs tool calls. Only plan when the work has independent steps you +want to track, retry, or summarize. When the user signals depth ("deep dive", +"thoroughly", "do this properly"), lean toward planning even if it looks small. + +LIFECYCLE +- create: batch all tasks up front with name + description. + name: short user-visible label. + description: full plan — what to do, what to produce, any constraints. + Be specific. Your future self (next step) reads this and must act on it + without re-asking the user. +- update to in_progress before you start work on a task. +- update to complete when done. Set result to 1-3 sentences of what was + produced or accomplished. Result is optional but strongly preferred — it's + how future turns see "this got done, here's what it was." +- update to failed if you tried and gave up; set result to the reason. +- update to cancelled if you decided not to do this task; set result to why. + +DESCRIPTION QUALITY +Bad: "Refactor parser" +Good: "Refactor parser: split parse() into tokenize() and parseTokens(). + Preserve current public API. Update the 3 call sites in handlers.ts. + Aim under 80 lines per function." + +A vague description means you (or a future step) won't know what "done" +looks like. + +DISCIPLINE +- One task in_progress at a time. Mark the current one terminal before + starting the next. +- Update status the moment state changes. Stale in_progress tasks make the + user think you're confused. +- Tasks persist for the whole thread. Completed ones stay visible as your + audit trail. + +DON'T +- Don't plan tasks for trivial requests. +- Don't create a task and mark it complete in the same call (no fake + productivity). +- Don't recreate a failed task with the same description. Either update its + status for a retry, or create a NEW task whose description includes what + went wrong and what's different. + + Three things on the roadmap, not shipped yet. If the user asks for them, name what's missing and offer the closest thing today. From 52a6ffc7ff70e97de70cc0e02038a974410fb69b Mon Sep 17 00:00:00 2001 From: Ronit Date: Thu, 14 May 2026 17:14:43 +0530 Subject: [PATCH 10/11] feat(kernel): Anthropic provider options for thinking + cache TTL MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit - toolStreaming: true — TUI sees partial tool args mid-emit - effort: 'xhigh' — max thinking budget per step - sendReasoning: true — reasoning sent back across react steps - thinking.adaptive+summary — auto-tune per step, lean conv history - cacheControl.ttl: '5m' — explicit (was implicit default) at all cacheControl marker sites so future 1h experiment is a one-character diff Also widened ContextMessage.providerOptions.anthropic.cacheControl to accept an optional 'ttl' field so the b2 marker site typechecks; uses the same '5m' | '1h' literal union the Anthropic SDK exposes. Combined with adaptive: trivial requests barely think, substantial ones think hard. Net positive for multi-step workflows that PR-C enables. Co-Authored-By: Claude Opus 4.7 (1M context) --- apps/kernel/src/agent/context-messages.ts | 4 +++- apps/kernel/src/tools/index.ts | 2 +- apps/kernel/src/turn.ts | 16 ++++++++++++---- 3 files changed, 16 insertions(+), 6 deletions(-) diff --git a/apps/kernel/src/agent/context-messages.ts b/apps/kernel/src/agent/context-messages.ts index 3a8bfba..d492c08 100644 --- a/apps/kernel/src/agent/context-messages.ts +++ b/apps/kernel/src/agent/context-messages.ts @@ -24,7 +24,9 @@ interface BuildContextMessagesArgs { type ContextMessage = { role: "user"; content: string; - providerOptions?: { anthropic?: { cacheControl?: { type: "ephemeral" } } }; + providerOptions?: { + anthropic?: { cacheControl?: { type: "ephemeral"; ttl?: "5m" | "1h" } }; + }; }; /** diff --git a/apps/kernel/src/tools/index.ts b/apps/kernel/src/tools/index.ts index 0fc4f51..a2ca445 100644 --- a/apps/kernel/src/tools/index.ts +++ b/apps/kernel/src/tools/index.ts @@ -29,7 +29,7 @@ function withTailCache>(tools: T): T { [lastKey]: { ...tools[lastKey]!, providerOptions: { - anthropic: { cacheControl: { type: "ephemeral" } } + anthropic: { cacheControl: { type: "ephemeral", ttl: "5m" } } } } } as T; diff --git a/apps/kernel/src/turn.ts b/apps/kernel/src/turn.ts index c924ddb..0d5b6fc 100644 --- a/apps/kernel/src/turn.ts +++ b/apps/kernel/src/turn.ts @@ -183,7 +183,7 @@ export async function runChatTurn( // provider-options the SDK set) when adding cacheControl. anthropic: { ...existingAnthropic, - cacheControl: { type: "ephemeral" } + cacheControl: { type: "ephemeral", ttl: "5m" } } }; } @@ -196,7 +196,7 @@ export async function runChatTurn( const last = modelMessages[modelMessages.length - 1]!; last.providerOptions = { ...last.providerOptions, - anthropic: { cacheControl: { type: "ephemeral" } } + anthropic: { cacheControl: { type: "ephemeral", ttl: "5m" } } }; } @@ -247,7 +247,7 @@ export async function runChatTurn( if (contextMessages.length > 0) { const lastContext = contextMessages[contextMessages.length - 1]!; lastContext.providerOptions = { - anthropic: { cacheControl: { type: "ephemeral" } } + anthropic: { cacheControl: { type: "ephemeral", ttl: "5m" } } }; } @@ -266,7 +266,7 @@ export async function runChatTurn( role: "system", content: SYSTEM_PROMPT, providerOptions: { - anthropic: { cacheControl: { type: "ephemeral" } } + anthropic: { cacheControl: { type: "ephemeral", ttl: "5m" } } } }, messages: [...contextMessages, ...modelMessages], @@ -289,6 +289,14 @@ export async function runChatTurn( trigger: { type: "input_tokens", value: midrunThreshold } } ] + }, + // NEW (PR-C) ─────────────────────────────────────── + toolStreaming: true, + effort: "xhigh", + sendReasoning: true, + thinking: { + type: "adaptive", + display: "summarized" } } satisfies AnthropicLanguageModelOptions }, From d112dd55309fc802c396d00e1f727b51d7b1dfea Mon Sep 17 00:00:00 2001 From: Ronit Date: Thu, 14 May 2026 17:21:02 +0530 Subject: [PATCH 11/11] fix(kernel): update to reflect manage_tasks shipping Old entry said 'Plans / task lists' were unshipped, contradicting the new block in the same prompt. Rescope to specifically the TUI-side-panel rendering (the actual remaining unshipped piece) so the agent doesn't carry contradictory instructions about whether task tracking exists. Co-Authored-By: Claude Opus 4.7 (1M context) --- apps/kernel/src/agent/system-prompt.ts | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/apps/kernel/src/agent/system-prompt.ts b/apps/kernel/src/agent/system-prompt.ts index 3c1aba9..5772f3c 100644 --- a/apps/kernel/src/agent/system-prompt.ts +++ b/apps/kernel/src/agent/system-prompt.ts @@ -272,7 +272,7 @@ DON'T Three things on the roadmap, not shipped yet. If the user asks for them, name what's missing and offer the closest thing today. - **Sub-agents.** Spawning child agents with their own context windows for parallel sub-tasks. Closest thing today: do the work serially, or split it into multiple tool calls in parallel. -- **Plans / task lists.** A side panel of checkable steps. Closest thing today: write the plan inline as markdown, the user can copy it. +- **Plans / task lists in the TUI.** A side panel of checkable steps. Closest thing today: I track tasks internally via manage_tasks and the user sees them surface in my responses, but the TUI doesn't render them as a panel yet. - **Scheduled / recurring work.** Cron-fired prompts that run without you. Closest thing today: nothing. Own this gap honestly.