Skip to content

chore(deps): update pi-agent-core and pi-ai to 1.1.0 - #385

Merged
nickruigrok merged 1 commit into
mainfrom
chore/pi-agent-core-1-1
Oct 11, 2026
Merged

nickruigrok merged 1 commit into
mainfrom
chore/pi-agent-core-1-1

Conversation

@nickruigrok

@nickruigrok nickruigrok commented Oct 11, 2026 •

Copy link
Copy Markdown
Contributor

First PR of the GRA-349 stack (Pi agent core with one bounded Code Mode tool).

Pins @earendil-works/pi-agent-core and @earendil-works/pi-ai to exactly 1.1.0. It is the newest release past the three-day wait: released 2026-10-07 at 22:12 UTC, eligible since 2026-10-10 at 22:12 UTC. Both packages stay on the same tested version.

What changed upstream (0.87.1 to 1.1.0)

  • pi-agent-core 1.0.0 removed the harness, sessions, compaction and skills. Core never used any of them; it uses only runAgentLoop, AgentTool and AgentMessage.
  • pi-ai 1.1.0 now requires a stream function to return an AssistantMessageEventStream. Core's gateway already builds one with createAssistantMessageEventStream, so type-checking still passes.
  • Messages carry a new durationMs field. It is optional, and the stored transcript takes it as it comes.
  • pi-ai 1.1.0 retries server_busy and "servers are currently busy" provider errors. Those retries happen inside core's maxRetries: 2, so a request still makes at most three provider attempts, each admitted by the gateway.
  • pi-ai now estimates input at 3.5 characters per token, instead of 4, when it works out output limits. Core sizes requests itself (requestChars, 3 characters per token), so only pi's own output-limit estimate changes.
  • The tool onUpdate callback and tool_execution_update events are available. PR 4 of this stack uses them for progress.

pi-telemetry

pi-ai keeps its dependency on @earendil-works/pi-telemetry (0.87.1 already had it). I read the 1.1.0 code:

  • It holds a TelemetryContext contract, a no-op context and an in-memory reference context.
  • It has no exporter, no network calls and no globals.
  • pi-ai only imports its types. Its one runtime use passes on a caller-supplied telemetryContext (api/simple-options.js).
  • Core passes no telemetry context, so nothing is recorded or sent.

Every model request still goes through core's gateway (models(env).agent, admitted by the ModelLedger). The provider SDKs are still only reached through the gateway's bound stream functions.

Evidence

  • vp check: passes (format, lint, types).
  • vp test run test/agent.test.ts test/models.test.ts test/agent-instructions.test.ts: 97 passed.
  • vp test run test/model-ledger.test.ts test/model-requests.test.ts test/model-rules.test.ts test/model-budgets.test.ts test/model-prices.test.ts test/chats.test.ts: 95 passed.

The newest Pi release past the three-day wait. The agent loop and the
gateway provider streams use the same APIs, so core does not change.
pi-ai keeps its pi-telemetry dependency: types plus a no-op and an
in-memory context, with no exporter. Core passes no telemetry context,
so nothing is recorded or sent, and every model request still goes
through the gateway.
@nickruigrok
nickruigrok force-pushed the chore/pi-agent-core-1-1 branch from b3dd385 to df18454 Compare October 11, 2026 03:06
@nickruigrok
nickruigrok marked this pull request as ready for review October 11, 2026 03:08
@nickruigrok

Copy link
Copy Markdown
Contributor Author

@greptileai review

1 similar comment
@nickruigrok

Copy link
Copy Markdown
Contributor Author

@greptileai review

@greptile-apps

greptile-apps Bot commented Oct 11, 2026 •

Copy link
Copy Markdown

RetriggerConfidence Score: 5/5

[High impact] The PR appears safe to merge; no actionable failure caused by the dependency update was established.

Summary

Pins @earendil-works/pi-agent-core and @earendil-works/pi-ai to 1.1.0 and updates the matching lockfile.

  • Both Pi packages now use version 1.1.0.

Reviews (1) · Last reviewed commit: "chore(deps): update pi-agent-core and pi..." · Reviewed by Greptile

@nickruigrok
nickruigrok merged commit f0337af into main Oct 11, 2026
9 checks passed
@nickruigrok
nickruigrok deleted the chore/pi-agent-core-1-1 branch October 11, 2026 03:53
nickruigrok added a commit that referenced this pull request Oct 11, 2026
… outcomes (#387)

Second PR of the GRA-349 stack, on top of #385. It brings the existing
`executeCode` and `runCode` up to the module contract in section 21.1.
The executor and the tool are the same ones as before; nothing new runs
beside them.

## What changes
- **Input.** The tool takes `{ code }` and nothing else. The declaration
sets `additionalProperties: false`, and `prepareArguments` refuses
anything else with a short message. Pi's own validation error would
repeat the whole input, which can be 200 KB, back to the model and into
the transcript.
- **Source limit.** A module over 200,000 characters is refused before
it loads. It ends with status `refused` and a bounded diagnostic.
- **Signature.** The default export is called as `(self, env, ctx)`:
  - `self` is `null`;
  - `env` holds the same seven scoped stubs as before (decision D1);
- `ctx` is a frozen object with only the run's own `waitUntil`, so
nothing of core's and no `exports`. `waitUntil` extends no lease,
because the host closes the run's stubs when the run ends.
- The agent instructions and every test module now use the new
signature. This is a hard cut.
- **Verify, then run.** The harness has two entry points:
- `verify` imports the module and checks its default export. Whatever
the module does while it loads runs inside the isolate's limits. A
failure there is `load_failed`, and what the module logged before
failing is kept.
  - `run` then calls the function.
- **Wall clock.** The limit is set before the isolate is asked to load
the module, and covers loading, verify and run. Workerd gives no way to
observe when the load itself starts, so this ordering is in the code,
not under test.
- **Statuses.** The host sets every outcome; the code can't. They are
`completed`, `threw`, `refused`, `load_failed`, `cpu_exceeded`,
`crashed`, `wall_expired` and `cancelled`.
- A failed run is a Pi `isError` result. The status is kept in the tool
result's `details`, and the model reads it as `Error (<status>):`.
- When workerd sees code waiting on something that can never settle, it
cancels the request as hung. That counts as `wall_expired`.
- **Test-only limit.** `CODE_RUN_TIMEOUT_MS` shortens the wall clock in
tests only, like `APP_CALL_TIMEOUT_MS`. It can only lower the limit, and
the model can't set it.

## Not here
- **CPU limit (D2).** `cpu_exceeded` maps the runtime's message, but
plain workerd doesn't enforce `limits.cpuMs`. The CPU proof is
GRA-382's; PR 5 of this stack records `cpu_limit_unsupported` on plain
workerd.
- **Lease and revocation.** The persisted record, revocation and closing
the lease are PR 3. Output budget and streaming are PR 4.

## Evidence
New tests in `apps/core/test/agent.test.ts`, all through the real
Workspace DO, Worker Loader and Pi loop, with only the model provider
faked:
- A full module: helpers, `self === null`, `ctx` keys `["waitUntil"]`,
and two parallel native calls.
- An oversized module is `refused` before loading. It would log while
loading, and it doesn't.
- No default export, a non-function export, a module that throws while
loading (its log is kept) and a syntax error are all `load_failed`.
- A function that throws is `threw`, after the module's load-time log.
- Extra input fields are refused, the module doesn't run, and the input
isn't repeated back.
- A promise that can never settle is `wall_expired`.
- A 60 s wait under a 1 s test limit is `wall_expired` within 10 s.
- Cancellation is now reported as `Error (cancelled)`.

Commands:
- `vp check`: passes.
- `vp test run test/agent.test.ts test/run-fixes.test.ts`: 54 passed.
- `vp test run test/chat-connections.test.ts
test/agent-knowledge.test.ts test/agent-instructions.test.ts
test/agent-connections.test.ts test/run-fixes.test.ts`, and the same
with `test/packages.test.ts test/dependencies.test.ts
test/chats.test.ts`: all pass.
- `vp test run -c vite.screens.config.ts test/agent-apps.test.ts
test/agent-builds.test.ts test/preview-repairs.test.ts`: 38 passed.



## Review fixes (711c4a5)
**Must-fix: the code could forge the host-set statuses.** It could make
`cpu_exceeded`, `wall_expired` or `crashed` appear by making a harness
call reject with a message it chose. Three routes did it: a thrown
Proxy, built-ins patched after load, and replacing
`Harness.prototype.run` through an import of the harness.
- The harness takes every built-in it uses before the code loads:
`Object.keys`, `Proxy`, `Reflect`, `Error`, the array and string
helpers, `JSON.stringify` and `String`.
- The entrypoint class and its prototype are frozen before load.
- `render`, `show` and `describe` are total. They catch everything, fall
back to fixed text, and read a thrown value's `stack` with `Reflect.get`
instead of `instanceof`.
- `verify` and `run` return every JS outcome and never reject from JS. A
rejection is therefore always the runtime's own.
- The host maps `cpu_exceeded` only from the runtime's exact message,
matched whole, and `wall_expired` only from its hang message (prefix).

**Nits:**
- `run()` with no function to call ends as `load_failed` (the outcome
carries `loaded`).
- The `crashed` doc now says a rejection while loading ends as
`load_failed` and one while running as `crashed`.
- An outcome that fails the schema ends as `load_failed` from verify and
`threw` from run; this is documented.

**New tests:**
- A thrown Proxy whose traps throw the runtime's message ends as
`threw`.
- An error carrying the runtime's message ends as `threw`.
- Built-ins replaced as the module loads end as `threw`.
- Harness methods replaced through an import end as `completed`.
- A Proxy thrown while loading ends as `load_failed`.
- The hang message thrown while loading ends as `load_failed`.
- A module of exactly 200,000 characters runs, and one character more is
refused.
- `ctx.waitUntil` work calling env after the run returned is never
answered. On workerd, that work either reaches core and is refused at
the lease, or never reaches core at all; the test accepts both, so it
doesn't prove which.

Evidence: `vp test run test/agent.test.ts` (57 passed at this commit),
plus the connection, knowledge, chat and run-fix test files (42 passed).


## Delta review fix (b4ae7c2)
**Must-fix: forgery through structured clone.** A setter on
`Array.prototype[0]` fired when the harness pushed to its logs array,
installed a throwing getter, and made the RPC serialisation of the
outcome reject with the code's own message.
- `verify` and `run` now return **one string**: the outcome's JSON, put
together by hand from primitives (two booleans, the logs, and one result
or error string), with a fixed fallback.
- Logs are kept as the JSON of their lines (`logsJson`), built by string
concatenation. No array or object of the harness's is ever filled
through [[Set]] or sent.
- The host checks it is a string and bounded, parses it, then applies
the schema.
- New entries in the forgery table: a setter on `Array.prototype[0]` set
up while the module loads and inside the function, each throwing the CPU
and the hang message. They end as `completed` or `threw`.
- Mutation check: against the previous harness (711c4a5), all four new
entries fail.

Evidence: `vp test run test/agent.test.ts` passes (61), and so do the
connection, knowledge, chat and run-fix test files (42).
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant