fix(copilot): route responses-only models (grok-4.x, gpt-5.x) to /responses API - #1
Closed
MaxMoldmann wants to merge 2437 commits into
Closed
MaxMoldmann wants to merge 2437 commits into
MaxMoldmann wants to merge 2437 commits into
Conversation
…afe config Preserve contributor ancestry while integrating only 57d5878..bffc2d3 on the current approved-PR candidate. Do not import the old branch tree or unrelated divergent changes. Keep thought signatures on persisted tool calls and provider resume handles separate from jcode session IDs. Retain Gemini/Antigravity compaction and Gemini schema/client compatibility changes. Review corrections: do not enable Cursor summary compaction until its prompt truncation is safe. Resolve Gemini config without sticky process-environment exports, preserve the project alias, and invalidate cached runtime state on route/project changes. Apply compaction caps on construction and refresh them before requests while retaining the uncapped model window. Restrict HTTP retries by method/status/quota, honor bounded Retry-After hints, and avoid replaying onboarding. Validation: provider/schema/message suites, config reload and compaction suites, local HTTP retry fixture, actual runtime cache replacement, and two-turn agent signature/session-identity regression all pass. Shared-target artifacts were rebuilt under the host Cargo gate after a stale dependency caused a false missing-method failure. Live Gemini acceptance remains untested because credentials are unavailable.
* ci: classify PR labels with Jev semantic scope * ci: pin validated Jev semantic labeler revision * ci: allow manually labeling an existing PR
Configure root reasoning effort independently for light and deep swarm modes.
* ci: label PRs after current-head Greptile reviews * ci: isolate trusted review concurrency and pin verified action
Preserve current master and original PR ancestry. The two formatting fixes are already present; apply only the three missing workflow fixes. Duplicate-rejecting YAML parsing passes for every workflow, and negative controls reproduce each old duplicate env failure. Other workflow semantics are unchanged.
Apply only the two scoped contribution commits while preserving current master and original PR ancestry. Keep fallback branch and automatic pull policies unchanged.
Preserve current master and original author history. Resolve terminal test overlap by retaining tmux extended-key coverage and checking the Kitty push sequence specifically rather than rejecting modifyOtherKeys. Candidate is pending combined validation.
… fixes The exact scoped three-fix patch is already present on current master. Preserve original contribution ancestry without replacing newer code. Current candidate tests and public-interface validation follow.
…ixes The scoped three commits are patch-equivalent and all changed blobs matched current master before the other approved integrations. Preserve original PR ancestry without replacing newer code.
…GUI clients - jcode-base: Claude /limit-reset contract (read-only at-wall offer lookup, profile-pinned organization, confirmed claim, cache and cooldown invalidation) - jcode-base: account-scoped OpenAI reset preparation and GUI review details - protocol/daemon: invalidate_anthropic_usage, and both invalidations are now lightweight one-shot control requests - harness API/SDKs: invalidate_usage request for Rust and TypeScript - CLI: jcode usage --json reports redeemable banked resets per login
paused_jcode_shell_command is only used by tests, and escape_shell_single_quotes is always called fully qualified, so both imports warned on macOS and tripped the zero-warning budget in the macOS Build & Test job.
After the CLI launcher bundle was dropped, only tests call it, so macOS builds reported it as dead code and exceeded the zero-warning budget.
…on id Reload handoff re-execs clients with --resume <id>. When the id existed only server-side, find_session_by_name_or_id fell through to a title scan that parsed every session file (~2600 files, 1.1GB), stalling client startup ~15s. Generated ids can never match a short name or title, so bail early.
Jcode talks to provider APIs directly (Claude via the native Anthropic OAuth/API runtime) rather than shelling out to vendor CLIs. Drop the deprecated jcode-provider-claude-cli-runtime crate, MultiProvider's claude slot and use_claude_cli flag, and the ClaudeSubprocess provider choice. --provider claude-subprocess remains a hidden alias for claude, and JCODE_USE_CLAUDE_CLI is now ignored with a warning. test_api and the real provider smoke script now exercise the direct Anthropic runtime.
…the aws CLI Enable aws-config's credentials-login feature so ProfileFileCredentialsProvider reads `aws login` (login_session) profiles natively, and drop the `aws configure export-credentials` subprocess.
Voice sends wait on classification. Batches share identical state and are independent, so run them with try_join_all instead of sequential round trips.
Streaming events went through block_in_place, handing the worker core to a fresh blocking thread per delta. With Tokio's default 512-thread, 10 s keep-alive pool and glibc's per-thread arenas, one desktop connection grew the bridge to 69 threads and ~123 MB of anonymous memory. - Only enter block_in_place for the event kinds that read session files. - Bound the blocking pool (8) with a 2 s keep-alive. - Cap glibc arenas at 2 and trim after heavy requests, large frames, and client disconnect. - Release oversized frame buffers instead of pinning their capacity. - Cap session-scan stat workers at 4.
…ead of ACP Grok Build no longer spawns `grok agent stdio`. It calls the Grok CLI chat proxy (https://cli-chat-proxy.grok.com/v1, OpenAI-compatible) directly through the OpenAI-compatible runtime, so Jcode owns tool execution. - Auth: the xAI OIDC session in $GROK_HOME/auth.json (default ~/.grok/auth.json). Only the https://auth.x.ai::<grok-cli client id> entry is used. expires_at (or the JWT exp claim) is honoured, the token is refreshed via refresh_token with a 60s skew, and a 401 forces one refresh-and-replay. The token is re-read for each request. - Identity: requests present the official Grok CLI headers (User-Agent grok-cli/<ver>, X-XAI-Token-Auth: xai-grok-cli, x-grok-client-version, x-grok-client-identifier: grok-shell, x-grok-client-surface, and per-turn x-grok-model-override / x-grok-conv-id / x-grok-req-id). These were taken from the Grok CLI 1.0.41 binary. JCODE_GROK_CLI_VERSION overrides the version. - Login: native xAI OAuth device flow (Grok CLI client id) in both the CLI and the TUI. It writes to the shared Grok credential store. The managed Grok binary download and the ACP bin/tests/dependency have been removed. - Route ids (grok-build:<model>, grok-build-acp) and /model selection are unchanged.
…afe batch Typesafe latency is bimodal and sticky per connection (~150ms or 2-12s). Voice insert/send waits on Jev, so race duplicates on separate clients after 300/700/1500ms and take the first valid answer. Primary failures still return immediately and are never duplicated. Typesafe direct accepts all 26 voice questions in one request, and batches now run concurrently. Measured live: unhedged median ~5s, hedged median ~0.7-1.1s.
Managed Jcode Cloud hosts authorize a fresh key for each connection and publish their host keys through the control plane. SshConnectOptions can now supply that identity exclusively, pin known_hosts without consulting user or system files, and ignore user SSH config. Adds an ssh_prompt example that creates a session and runs one prompt over SSH.
Agents restarted the bridge by hand from bash (kill, then setsid nohup). When the agent's own session went through that bridge, the kill cut its connection, the turn was interrupted before the relaunch ran, and every Desktop panel was stranded on a dead socket. reload-bridge (in selfdev and desktop_selfdev) runs the restart as a daemon-owned background task that survives the caller disconnecting. It preflights the new binary, records the old bridge's command line from its peer PID, stops it, starts the new one, verifies the socket accepts, and relaunches the old command line if the new bridge exits or never accepts. Prompts now steer agents to it instead of manual restarts.
MaxMoldmann
force-pushed
the
fix/copilot-responses-api-routing
branch
from
September 24, 2026 22:35
5592d5c to
5d987c2
Compare
MaxMoldmann
marked this pull request as ready for review
September 24, 2026 22:42
MaxMoldmann
force-pushed
the
fix/copilot-responses-api-routing
branch
from
September 24, 2026 22:47
5d987c2 to
f98937a
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
Models like
grok-4.5,grok-4.6,grok-4.7,gpt-5.3-codex,gpt-5.5, and other newer Copilot models fail with HTTP 400unsupported_api_for_modelwhen selected in jcode. The root cause is that Copilot only exposes these models via the OpenAI Responses API (/responsesendpoint), not/chat/completions.From the
/modelscatalog for these models:{ "supported_endpoints": ["/responses"] }Fix
Read
supported_endpointsfrom the/modelscatalog at startup. For any model that lists/responsesbut not/chat/completions, route requests to/responsesinstead.Changed Files
crates/jcode-base/src/auth/copilot.rs- Addsupported_endpoints: Vec<String>toCopilotModelInfo; addneeds_responses_api()method.crates/jcode-provider-copilot-runtime/src/lib.rs- Addresponses_model_idsset populated from/modelson startup. Instream_request, branch onmodel_needs_responses_api(): convert messages to Responses API input items and POST to/responses. Addprocess_responses_sse_streamto parse SSE events.crates/jcode-provider-copilot-runtime/Cargo.toml- Addjcode-provider-openaidependency.crates/jcode-tui/src/tui/app/commands.rs- Addunsupported_api_for_modeltois_fatal_model_endpoint_error.Tests
34/34 copilot-runtime unit tests pass. 7 new
needs_responses_apiauth tests, 3 new routing tests added.Live integration test against
api.githubcopilot.com:/modelslists grok-4.5/4.6/4.7 as responses-only/chat/completionsfor gpt-5.3-codexunsupported_api_for_model(root cause)/responsesstreaming for gpt-5.3-codex/responsesstreaming for grok-4.5/responsesstreaming for grok-4.6/responsesstreaming for grok-4.7Affected Models
gpt-5.3-codex,gpt-5.4-mini,gpt-5.5,gpt-5.6-luna,gpt-5.6-sol,gpt-5.6-terra,gpt-6-astra,gpt-6-luna,gpt-6-sol,grok-4.5,grok-4.6,grok-4.7,mai-code-1.1-flashLinked Issue
Closes 1jehuang#1128