Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -115,7 +115,7 @@ kcode

`--use` tests the first listed model before saving and selecting it. A failed connection test saves nothing. Omit `--use` to save without testing or changing the default model. For custom/local models, add `--context-limit 32768 --output-limit 4096` (use your server's actual limits). Each value must be a positive safe integer and applies to every repeated `--model`. Inspect configured limits with `kcode provider list --json`. Omitting these flags preserves the existing model-limit defaults.

Supported API formats: `openai-completions`, `openai-responses`, and `anthropic-messages`. See the [model examples](docs/examples.md#2-choose-your-own-model) for environment variable setup, connection checks, and model overrides for a single run.
Supported API formats: `openai-completions`, `openai-responses`, and `anthropic-messages`. An endpoint that needs no authentication — a server you run locally, typically — is a provider without a key: omit `--api-key-env`, leave `MCODE_PROVIDER_API_KEY` unset, and the connection is stored with no credential, so no request to it carries one. See the [model examples](docs/examples.md#2-choose-your-own-model) for environment variable setup, connection checks, and model overrides for a single run.

</details>

Expand Down
5 changes: 3 additions & 2 deletions docs/examples.md
Original file line number Diff line number Diff line change
Expand Up @@ -56,16 +56,17 @@ pnpm kcode exec "Explain this project's test entry points" --model <provider-id>

Replace the example URL, model name, and IDs with your configuration and the IDs returned by the list command. `--use` tests the first listed model, then saves the provider and selects that model as the default. A failed connection test exits nonzero without saving or changing the default; correct the URL, key, or first model ID and retry. Omit `--use` to save without a connection test or default-model change. `exec --model` overrides only the current run. Backslash line continuations are for POSIX shells; use a single line in PowerShell.

For a local server, configure its actual token limits explicitly:
A local server that checks no credential needs no key: omit `--api-key-env` and leave `MCODE_PROVIDER_API_KEY` unset, since that variable is the default when the flag is absent — then configure the server's actual token limits explicitly:

```bash
pnpm kcode provider add --name local-models --base-url http://localhost:8080/v1 \
--api-format openai-completions --model local-model --model another-model \
--api-key-env MCODE_PROVIDER_API_KEY \
--context-limit 32768 --output-limit 4096 --use
pnpm kcode provider list --json
```

The provider is stored with no `apiKey` and no `authMode`, and every request to it — the connection test, model discovery, and a turn — is sent with no `Authorization` and no `x-api-key` header. Pass `--api-key-env <name>` when the server behind the URL does check a credential; an explicitly named variable that holds nothing is an error, because the key was asked for and not found. The same choice exists in the TUI: the catalogue's **Local model** entry pre-fills an OpenAI-compatible endpoint, skips the protocol step, and leaves the API Key field empty, while the `/provider` editor shows an entry without a key as `Not set · a request carries no credential`, and an empty API Key field there removes a saved key.

`--context-limit` and `--output-limit` each accept a positive safe integer (at most `9007199254740991`). Either flag can be used independently. The same limits apply to every repeated `--model`; only the first model is tested and selected by `--use`. The JSON list shows the configured values as `contextLimit` and `maxOutputTokens`. Without these flags, the existing defaults remain unchanged (unknown custom models currently fall back to 200,000 context tokens and 16,384 output tokens). Model discovery does not infer your local server's context size.

`--api-key-env` reads the current environment variable value and stores that value in the active profile's `config.yaml`; it does not save an environment-variable reference. The file still contains plaintext credentials. On POSIX systems, config writes and temporary copies use `0600`. When loading existing files, KCode removes group/other access while preserving the owner's permissions; already-private files such as `0400` or `0600` do not require a permission change. Loading fails if an unsafe main config cannot be restricted. Older migration backups are also checked, but inspection or repair failures produce a warning identifying the directory or backup that needs manual attention rather than preventing the main config from loading. Windows file modes do not provide equivalent ACL protection; restrict access to the profile directory using Windows permissions.
Expand Down
2 changes: 1 addition & 1 deletion docs/tui-capabilities.md
Original file line number Diff line number Diff line change
Expand Up @@ -7,7 +7,7 @@ The evidence column summarizes the historical TUI 0.3.11 restoration record from
| Capability | Implementation | Evidence |
| --- | --- | --- |
| MiniMax login, logout, Token Plan, quota, check-in | OAuth Core, account clients, and TUI / CLI entry points restored; `/login <provider>` names the sign-in it starts, and every other provider is connected in `/provider` | Login-state, auth-command, check-in, provider, and ACP tests; production Token Plan session and resume passed |
| BYOK / custom models | Retained; headless model overrides no longer inherit the default Token Plan login requirement. `/provider` lists every connection the runtime resolves — entries the `provider` tree of `config.yaml` declares, and the user's `custom_provider` connections — tests each against its own endpoint and key, and connects a new provider from the same catalogue the model picker offers | Local protocol server drives runtime, resume, and file tools; live BYOK session passed; managed models still reject unauthenticated requests. The panel's list, test and connect paths are covered by unit tests and by a built CLI against an isolated config; the endpoint a builtin entry is tested against is a local stand-in, not a live service |
| BYOK / custom models | Retained; headless model overrides no longer inherit the default Token Plan login requirement. `/provider` lists every connection the runtime resolves — entries the `provider` tree of `config.yaml` declares, and the user's `custom_provider` connections — tests each against its own endpoint and key, and connects a new provider from the same catalogue the model picker offers, including a **Local model** entry for an OpenAI-compatible server. A key is optional everywhere: an entry saved without one declares no credential scheme, and no request to it carries `Authorization` or `x-api-key`; the `/provider` editor marks such an entry `Not set · a request carries no credential` and removes a saved key when its API Key field is submitted empty | Local protocol server drives runtime, resume, and file tools; live BYOK session passed; managed models still reject unauthenticated requests. The panel's list, test and connect paths are covered by unit tests and by a built CLI against an isolated config; the endpoint a builtin entry is tested against is a local stand-in, not a live service. A keyless connection is covered end to end by `test:byok` against a loopback stand-in: add, list, connection test and a turn, with every request asserting the absence of both credential headers |
| GitHub Copilot provider (fork addition) | Connector signs in to a Copilot subscription and reads the account's live `/models` catalog: per-model context and output limits, the reasoning-effort levels the account reports, and the wire protocol each model advertises. Sign-in is the GitHub device flow through pi's Copilot OAuth provider, or a token supplied out of band; the credential stays in the provider credential store | Discovery and connector tests run against a captured catalog; live completions passed on `openai-responses`, `anthropic-messages` and `openai-completions`; `kcode exec` passed end to end on all three protocols and in all four egress modes. Both TUI entry points carry the connect action and run the device flow: the model picker row, and a `/provider` row that Runtime synthesizes from the sign-in status before any entry exists. Once the connector has written its entry, that entry is the only Copilot row — Codex instead keeps its synthetic row beside the entry. The enterprise API host is unit-tested only |
| mcode-tools | Public package's unchanged 0.0.4 artifact, launcher, and short-lived token leases restored | Archive SHA-512, CLI SHA-256, real CLI startup, lease and host integration tests; production shared-broker authentication passed |
| Search and Matrix image / audio / video tools | Matrix MCP and tool assembly restored; search is independent of mcode-tools | MCP configuration and auth-isolation tests; actual search passed; generation requests not run |
Expand Down
16 changes: 16 additions & 0 deletions docs/verification.md
Original file line number Diff line number Diff line change
Expand Up @@ -156,6 +156,22 @@ A real PTY session of the built CLI with the same isolated config showed the pan

NOT RUN: a real OpenRouter call (the endpoint above is a local stand-in, so the resolution path is proven, not the service), and Windows or macOS execution.

### Local endpoint without authentication, 2026-09-23

Verification results for the keyless local-endpoint support at code revision `e2a4dbf` (the documentation commit that follows changes no code).

A model served on an OpenAI-compatible endpoint that checks no credential — Ollama, llama.cpp, vLLM, LM Studio — is now a provider without a key. Nothing infers "no authentication" from a URL and no new setting was introduced: an entry that carries no key declares no credential scheme, and no request to it is sent with one. `packages/shared/src/credential-headers.ts` names the credential headers the three protocols use, clears them on the request's own options (the last header source `pi-ai` merges, and the one its SDKs read a `null` from as "remove this default header"), and owns the placeholder key the vendored transport requires — `openai-completions` throws on a falsy key, so the resolver hands such a model one that never reaches the wire.

`pnpm verify` passed all 14 gates on Linux arm64 with Node.js 26.5.1 and pnpm 9.12.0: source inventory (4,239 files), generated paths (127 package exports), source export, release tooling (44 tests), typecheck, build, standalone boundary, egress boundary, built artifacts (4 tests), capabilities (4,649 tests in 175 files), status contract (9 tests), CLI/ACP smoke (21 tests), offline BYOK (4 tests), and permission policy (142 tests). `test:windows`, `test:sandbox` and `test:release-package` are skipped on this platform.

The change adds 16 tests — 15 to the capability suite (one of them in a new file, which moves its file count from 174 to 175) and 1 to the offline BYOK suite, whose test count moves from 3 to 4. In `packages/agent-core/test/unit/pi-turn-runner/unauthenticated-endpoint.test.ts`, 7 cases pin the header clearing: both credential headers are cleared while the transport still receives the key it requires, other headers survive, a header the connection declares itself is preserved, and a model without the flag is left untouched. Two cases in the model-system service suite pin the saved shape (a candidate without a key stores no `apiKey` and no `authMode`, and its gated connection test sends no credential) and the empty-key convention (an empty key is stored as no credential, not rejected as invalid). Two resolver cases pin the resolved config on the custom-provider route and on the builtin route: the placeholder key plus `unauthenticatedEndpoint`, and no credential header. The reworked plan case states the new rule (a keyless plan is returned instead of a refusal) and keeps the incomplete-endpoint refusal covered. One case in the shared token-counter suite pins that a keyless endpoint is counted locally rather than probed. One onboarding case pins the catalogue's **Local model** entry and its keyless save; two editor cases pin a connection that carries no key being tested and saved, and an empty API Key field clearing a saved key.

The end-to-end proof is the offline BYOK suite's new case, which runs the built CLI against a loopback stand-in and an isolated data directory with no provider key in the environment: `provider add` saves the connection with `options` exactly `{ baseURL }`, `provider list` reports `hasApiKey: false`, the connection test answers, and `exec` completes a real turn against the stand-in. Every request the stand-in saw is asserted to carry neither `Authorization` nor `x-api-key` — including the auxiliary `POST /v1/responses/input_tokens` counter probe that the first run of this case caught, which is why the remote token counter now stays local for such an endpoint.

A real PTY session of the built CLI, driven against the same kind of stand-in on `127.0.0.1:11434`, showed the new surface: the catalogue renders `› Local model OpenAI-compatible server · API key optional`, the flow reaches the step labelled `API Key (optional)` under `Optional: a server that needs no authentication is connected with the field left empty`, submitting it empty connects (`Provider added: Local model`), and the panel then lists that connection as `Enabled · 1 model` beside an active `✦ local-model` status line. The stand-in logged one request from that session — `POST /v1/chat/completions`, `stream: false` — with `authorization: null` and `x-api-key: null`.

NOT RUN: any live provider or service call (every endpoint above is a local stand-in, so the resolution and request paths are proven, not a third-party service); the legacy v1 host's own model resolver, which still requires a key and is not the runtime the product drives; and Windows or macOS execution.

### Release-preparation verification, 2026-09-12

`pnpm verify` was run on `3de31f0e365e635c52d661c0a14c9ee69e65f099` in an isolated worktree after a frozen-lockfile install, on macOS arm64 with Node.js 26.4.0 and pnpm 9.12.0.
Expand Down
33 changes: 31 additions & 2 deletions packages/agent-core/src/pi-turn-runner/llm.ts
Original file line number Diff line number Diff line change
@@ -1,5 +1,6 @@
import type { Agent, AgentMessage, StreamFn } from '@earendil-works/pi-agent-core';
import { streamSimple, type CacheRetention, type SimpleStreamOptions } from '@earendil-works/pi-ai';
import { withClearedCredentialHeaders } from '@mavis/shared';
import { LLM_REQUEST_TIMEOUT_MS } from './defaults.js';
import type {
PiAfterLlmReplacementCommit,
Expand All @@ -14,10 +15,15 @@ export function composeStreamFn(resolved: LLMModelConfig): StreamFn {
// Order matters: each wrapper only fills the option it owns, while
// explicit per-call options from upstream wrappers still win. The host
// ceiling is the exception and therefore the innermost link: it has to see
// the fully accumulated options in order to shrink them.
// the fully accumulated options in order to shrink them. Clearing a
// credential header of an endpoint that needs no authentication sits below
// even that, because it must see every header the request would carry.
const callerHeaders =
resolved.headers && Object.keys(resolved.headers).length > 0 ? resolved.headers : undefined;
const hostClamped = withHostMaxOutputTokens(resolved.streamFn, resolved.hostMaxOutputTokens);
const baseStream = resolved.unauthenticatedEndpoint
? withUnauthenticatedEndpointHeaders(resolved.streamFn)
: resolved.streamFn;
const hostClamped = withHostMaxOutputTokens(baseStream, resolved.hostMaxOutputTokens);
const cacheWrapped = withCacheRetention(hostClamped, resolved.cacheRetention);
const maxTokensWrapped =
typeof resolved.maxTokens === 'number' && resolved.maxTokens > 0
Expand Down Expand Up @@ -539,6 +545,29 @@ function withHeaders(inner: StreamFn | undefined, headers: Record<string, string
}) as StreamFn;
}

/**
* Keeps the placeholder key of an endpoint that needs no authentication off the
* wire.
*
* The resolver hands such a model a key because the provider SDKs will not build
* a client without one; the header that would carry it is cleared here, on the
* request's own options, which is the last header source pi-ai merges — a `null`
* value is what its SDKs read as "remove this default header".
*/
function withUnauthenticatedEndpointHeaders(inner: StreamFn | undefined): StreamFn {
const base = inner ?? streamSimple;
return ((model, context, options) => {
const headers = withClearedCredentialHeaders(
options?.headers as Readonly<Record<string, string>> | undefined,
);
return base(model, context, {
...(options ?? {}),
// pi-ai types these as strings; a null is forwarded to the SDK untouched.
headers: headers as Record<string, string>,
});
}) as StreamFn;
}

function withFetch(inner: StreamFn | undefined, fetch: SimpleStreamOptions['fetch']): StreamFn {
const base = inner ?? streamSimple;
return ((model, context, options) => {
Expand Down
9 changes: 9 additions & 0 deletions packages/agent-core/src/pi-turn-runner/types.ts
Original file line number Diff line number Diff line change
Expand Up @@ -52,6 +52,15 @@ export interface LLMModelConfig {
*/
hostMaxOutputTokens?: number;
headers?: Record<string, string>;
/**
* The endpoint this model runs on needs no credential.
*
* The provider SDKs refuse to build a client without a key, so the resolver
* hands them one that is never meant to be sent; this flag is what keeps that
* placeholder off the wire, by clearing the credential headers each request
* would otherwise carry.
*/
unauthenticatedEndpoint?: true;
fetch?: SimpleStreamOptions['fetch'];
/** Provider payload transform for the main assistant response. */
payloadTransform?: SimpleStreamOptions['onPayload'];
Expand Down
Loading
Loading