Skip to content

[WRONG BRANCH] release: 2.83.0 - #6910

Merged
lidge-jun merged 26 commits into
mainfrom
codex/promote-main-2.83.0-r2
Oct 10, 2026
Merged

lidge-jun merged 26 commits into
mainfrom
codex/promote-main-2.83.0-r2

Conversation

@lidge-jun

Copy link
Copy Markdown
Owner

Summary

Promote the frozen dev candidate c61c3f62e6069ded1705f91846f36e9a4d72dbae to main as stable 2.83.0. The promotion head eb06a2c418bd1b992f6c8a13a5fdd81228dc09a2 merges in main ancestry and has exactly the candidate tree (42b7180671). The version is 2.83.0 in package.json, desktop tauri.conf.json, Cargo.toml and the opencodex-desktop entry of Cargo.lock. The dev pre-move has merged (#6902), so dev now carries 2.84.0.

2.83.0 contains 25 commits since 2.82.0. Most are bug fixes from today's lanes: #6850, #6859, #6858, #6874, #6869, #6870, #6828, #6879, #6881, #6882, #6883, #6887, #6888, #6885, #6889, #6864, #6880, #6886, #6890 and #6893. Two are fixes found by the release regression audit: #6903 and #6904. The rest are devlog-only moves.

Release regression audit. This audit is also the security review of the release diff. Three independent reviewers covered server/logs, providers/auth/config/coupons, and desktop/CLI/infra. They looked for regressions against 2.82.0 and gave a disposition for each changed security boundary. The first pass blocked the release on two P1 regressions, both reproduced by the coordinator:

Both fixes passed independent re-review, and the candidate was re-frozen on top of them. No P0/P1 remains.

Packed-tarball smoke at the candidate. npm pack produced 1984 files, including gui/dist and bin/ocx.mjs. The tarball was installed into an isolated prefix and home. ocx --version printed 2.83.0, and /healthz returned ok with version 2.83.0. The process then shut down cleanly.

Main-review disposition. This is a maintainer-controlled release promotion (MAINTAINERS.md: "Promotion from dev to main and npm releases is maintainer-controlled"). The dev maintainer-integration exception does not cover main. Project owner @lidge-jun decided in-session (2026-10-11): "네, 오너 권한으로 머지 (권장)". That decision authorizes merging this promotion with owner authority (--admin) without a second-maintainer approval. It was given separately from the release instruction "배포 회귀 검증까지 하고 배포까지 완료해줘". The PR targets main deliberately, so the enforce-target wrong_base failure is expected (as in #6771 and #6878).

Verification

Grok reset-coupon dialog included in this promotion (from merged #6870)

Checklist

  • Scope stays focused and avoids unrelated cleanup.
  • Docs or release notes were updated when needed.
  • Security-sensitive changes were reviewed for secrets, auth, and unsafe defaults.

github-actions Bot and others added 26 commits October 10, 2026 17:31
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
* fix(config): serialize Grok and link writers under existing locks

Hold the canonical config lock through Grok injection and cleanup, validate publication destinations, and preserve refresh-only admission. Coordinate link store mutations and revalidate recorded links before activating a bound listener.

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(grok): retain refresh skip on destination drift

---------

Co-authored-by: Epinephrine <27862058+luvs01@users.noreply.github.com>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Claude Code 2.1.294+ resolves the `sonnet` alias to claude-sonnet-5-5, but
the native tier table still pinned an unset Sonnet slot to claude-sonnet-5,
so subscription launches through `ocx claude` hid Sonnet 5.5 from the picker.

Refs #6873
…el (#6859)

* fix(codex): never discard a rotated native-main refresh token on cancel

A successful refresh_token exchange retires the token it sent. Two paths in
resolveMainAccountToken still lost the replacement when the caller went
away:

- the caller's signal was passed into the exchange, so a disconnect after
  the POST was sent aborted the request or its body read;
- after the exchange resolved, a caller-abort check threw before
  persistRefreshedMainAuthJson ran (added in #5268).

Either way auth.json kept the retired token and the account needed a
manual re-login. Caller cancellation now stops only the waits before the
exchange. The exchange observes the refresh's own 30-second timeout, the
result is published through the unchanged snapshot/identity guards, and the
caller's abort reason is thrown after the commit. An external writer that
replaced auth.json during the exchange still wins.

The claim and refresh-lock helpers await the callback through completion,
so a cancelled wait cannot release an acquired lock mid-exchange.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(codex): assert the persisted account identity after cancelled refresh publication

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: JUN <jun@lidgeai.com>
* fix: release turn admission on client cancellation

* fix: bind internal decision cancellation at admission

* fix: release pending Live upgrade admission on client abort

* test(layout): register active-turn-disconnect on an existing line to stay under the size threshold

---------

Co-authored-by: JUN <jun@lidgeai.com>
…#6858)

Ollama Cloud's mistral-large-4 streams several distinct tool calls in one
frame, each with its own id but all with function.index 0. The parser keyed
calls by index, so the second call collided with the first and the stream
failed with "changed a tool-call id for an existing index" (or "reused a
tool-call index for another function" when the names differ).

A valid native id is now the call's identity. Index, then position, is the
fallback only for entries without an id, and a later id adopts the call that
was first seen without one at the same index. Two entries in one frame share
a call only when they share an id. Budget keys follow first-seen order
instead of the index, so distinct calls at one index keep separate argument
and metadata accounting. The parallelToolCalls:false guard is unchanged: the
captured shape now reaches it and fails closed with the parallel-call error.

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
…_config.json store (#6828)

* fix(zcode): carry reasoning ladders and image input into the provider_config.json store

ZCode 3.14 moved custom providers to v2/provider_config.json (#5348), and the
store export carried over only each model's contextWindow — the reasoning
ladder #4147 added to the legacy v2/config.json block and the catalog's image
modality were both dropped, so every opencodex model showed no thought-level
picker and no attachment support in current ZCode builds.

Write the capabilities the client's own store schema accepts, extracted from
the ZCode CLI bundle's strict zod definitions: config.optionSpecs.reasoningLevel.
values (the same Codex-vocabulary ladder the legacy block emits, minus the none
sentinel; the chosen level is validated against values and forwarded verbatim
to reasoning.effort) and config.properties.inputFormat.supportsImage when the
catalog row declares image input. A model rule now ships when the row has any
of contextWindow, image input, or a selectable ladder; rows with none still
ship none, per the no-guessed-capability rule. Effort filtering is shared with
the legacy exporter via zcodeSelectableEfforts in model-metadata.

* chore: re-trigger review readiness gate after checklist attestation

---------

Co-authored-by: Blazetes <Blazetes@users.noreply.github.com>
…text reasoning (#6882)

* fix(responses): warn once when a custom provider drops replayed plaintext reasoning (#6675)

* fix(responses): treat Azure per-resource hosts as curated destinations in the reasoning notice
…x malformed call (#6881)

* fix(google): let the empty-completion retry replay a pre-output Vertex malformed call (#6876)

* fix(google): carry Vertex usage on truncation errors so a replay meters both attempts
…ts (#6883)

A raised inbound body reservation was released only on response EOF/error, a settled explicit cancel or a thrown handler error. If the client aborted and the response was never consumed, the reservation stayed held and later covered POSTs got 503 capacity exhausted until restart. Observe the request abort signal once work returns, cancel the producer through the shared idempotent path, and release once after cancellation settles. Closes #6793.
)

A synthetic 429 upstream_reset_replay_refused was attributed like a provider throttle (rate-limit/permitted) because attribution treated every 429 as a rate limit. Read the in-process replay-refusal marker in the deferred logger and carry a synthetic transport-ambiguous cause hint into final request and attempt attribution. Client-visible refusal behavior is unchanged. Closes #6658.
)

Read a fresh user environment block after the registry transaction and evaluate searchable PATH entries in order until the Desktop directory. Preserve the legacy registry probe only as a conservative fallback when the fresh PATH is unavailable. Cover UTF-16 parsing and PATH precedence with portable tests.
…#6472) (#6885)

* fix(copilot): support Auto-only accounts and plan-aware model selection

* fix(copilot): keep Auto visible with existing new-model Off policy

* fix(copilot): harden Auto renewal and refusal handling

Keep ephemeral session credentials redacted, preserve account-owned
snapshots through bounded OAuth renewal, and renegotiate expired queued
sessions without a nested concurrency lease. Add focused regressions,
pin public protocol evidence, and expose selection near the settings top.

* docs(copilot): document callback contracts and isolate GUI test globals

* docs(copilot): describe test fixtures and clarify Russian Auto hint

* Fix Copilot Auto OAuth refresh wire handoff and refusal mapping

* fix(copilot): preserve unknown permissions and shared selection policy

* refactor(copilot): limit Student fix to automatic negotiation

Defer configuration, GUI and CLI selection modes; preserve core recovery and operator policies. Scope API version to Auto requests and fix capacity cache eviction.

* fix(copilot): hand an Auto session renewed onto Chat back to the native owner

When a queued native Responses session for Copilot Auto expired and dispatch-time
renegotiation selected the Chat wire, the passthrough path rebuilt and sent
/chat/completions but still delivered the reply natively, so a Responses client
received Chat choices (JSON) or Chat chunks ending in adapter_eof (SSE). The
dispatch fence now signals the passthrough owner before the inference send and
reuses the existing 401/429 Chat handoff.

Co-authored-by: Zehua Kcriss Li <65699115+ZehuaKcrissLi@users.noreply.github.com>

* test(copilot): keep the renewed bearer assertion out of the privacy scan

---------

Co-authored-by: Zehua Kcriss Li <65699115+ZehuaKcrissLi@users.noreply.github.com>
Co-authored-by: Kcriss <jiafujun123@gmail.com>
… terminals (#6889)

* fix(logs): keep recognized bare upstream error classes in synthesized terminals

A bare upstream error that ended the WebSocket exchange was turned into a synthesized response.failed carrying upstream_server_error, and captureTerminalHttpStatus ignored bare errors, so structured rate-limit and overload codes were recorded as a generic 502. Preserve recognized rate-limit and overload classes in both relay modes, record them as 429/503 provisionally, let a genuine terminal supersede them, and record a bounded upstreamErrorType. Terminal-status diagnostics move to request-log-terminal-status.ts to keep request-log.ts under the size gate. No replay is added. Refs #6740.

* fix(logs): record status only for recognized bare upstream errors

Unrecognized bare errors recorded a provisional 502, which overrode explicit status fields that combo stream preflight reads (status/http_status on the event), and a completed terminal recorded 200. Record a provisional status only for recognized rate-limit and overload classes, and let response.completed only retire provisional evidence.
* test: isolate Windows host state and bound native enumeration

* test: scope native CIM ownership checks to the fixture PID

* test: key the trust fixture by the platform project path

The positive control added for the untrusted-project case wrote the trust key with a hard-coded Windows separator, so on Linux and macOS the trusted project never matched and the expected warning was absent (Cross-platform CI run 38007605487, test 1/4). A TOML literal string keeps the platform path verbatim, as the neighbouring uncached-collection test already does.

---------

Co-authored-by: JUN <jun@lidgeai.com>
…6880)

* fix(chat): promote caller x-session-id on Chat Completions ingress

Chat Completions passed the original request to admission and the ChatGPT bridge, which forwards only canonical session headers, so a caller x-session-id never reached ChatGPT-backed models and prompt-cache reuse was lost. Apply the same withCallerSessionIdentity promotion used by Responses and Messages before admission. Closes #6520.

* docs(structure): record Chat Completions x-session-id promotion

* test(server): pin the Chat Completions caller-session wiring

* test(server): pin the promoted Chat Completions request in the loopback CORS oracle
…ble observation (#6886)

* fix(codex): keep rediscovering main-account credits after an unavailable observation

The recovery sweep only renewed already spendable credit evidence, so one zero, null, restricted or retracted WHAM read stopped discovery and the main account stayed refused locally even after credits were usable again. Keep consented, unpaused accounts with an exhausted window eligible for paced lookups, keep admission gated on fresh positive evidence, and bound the consented refusal's Retry-After to the base metadata interval. Closes #6845.

* test(codex): pin recovery pacing when credit evidence is missing and WHAM refuses auth
…6890)

* fix(responses): report spend-ledger storage refusals as local 503s

A SpendLedgerFileRefusedError (for example an extra hard link from a sync daemon) was caught as a transport failure, returned as 502 Provider unreachable, attributed upstream, and charged to provider and account health. Map it, including nested retry causes, to 503 server_error spend_ledger_storage_unavailable before transport classification on passthrough, translated, compact and runTurn paths; log it as a local refusal and skip health penalties. The guard is unchanged. Refs #6314.

* test(responses): register the spend-storage-error owner in the core module inventory
The connected-sibling recycle case spawns a Bun child with piped stdout and stderr and, until the replacement wait, never looked at either. On Windows CI the child once wrote no readiness marker for 45.9 s and the test failed with a bare timeout, leaving no trace of what the child was doing (#4956 family). Drain both pipes from spawn into bounded buffers and append the child's pid, exit code and recent output to every wait's timeout. The replacement wait no longer opens a second reader on an already drained stream.

Refs #4956
…ertain outcomes (#6870)

* fix(grok): atomically claim coupon attempts and reconcile without respending

Persist one winning attempt before transport. Keep uncertain outcomes replay-only, refuse unreadable ledgers, replay existing records at capacity, and synchronize the route contracts and regression fixtures.

Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(grok): guard coupon settlement and preserve unknown outcomes

* fix(grok): require explicit coupon RPC status confirmation

* Retry confirmed Grok coupon settlement under transient contention

Retry only local settlement for ConfigMutationLockError caused by SQLITE_BUSY, with five attempts and four short waits. Preserve changed-claim and terminal guards, stop on other errors, and never repeat redemption. Cover real process contention, exhaustion, refusal classes and identity drift; clarify structure contracts and repair the garbled separator.

* fix(grok): preserve coupon claims and unknown operation identity

Flush coupon claims before dispatch, quarantine ambiguous version-one records,
and retain unknown attempts across CLI errors and dashboard dialog reopen.
Keep remaining-coupon availability separate from redemption confirmation.

* fix(grok): preserve coupon identity and release definitive failures

* fix(grok): keep coupon recovery holds on operation_state_changed

A settlement race that returns operation_state_changed left the GUI and CLI
treating the redemption as definitively failed, dropping the hold and the
operation id the user needs to inspect the attempt. Treat it like the other
unresolved codes. The coupon modal now reads uncertainty only from the
controller, so a reopened modal clears once the original request is confirmed.

Adds GUI and CLI regressions for both paths.

* fix(grok): replay a durable terminal winner to concurrent coupon refusals

A preflight refusal that lost a race to a concurrent definitive outcome
returned operation_state_changed, which the GUI and CLI now hold as an
unresolved attempt; the hold could never clear because the operation was
already final. The refusal now settles or inspects the ledger inside one
config mutation transaction and replays the definitive winner when account
and token match; genuinely unresolved states still return
operation_state_changed. The GUI retires the active attempt as soon as an
outcome is definitive, before awaiting the refresh, so a reopened modal cannot
recreate a hold.

* fix(grok): replay a definitive winner to a lost coupon claim

A request that lost the atomic open -> attempted claim always answered
attempt_in_progress, even when the winning request had already settled, which
left the dashboard holding a finished operation. A lost claim now reads the
durable winner and replays a matching definitive outcome with the same
response as a late refusal; unresolved or mismatched operations, and a failed
read of the winner, still return attempt_in_progress so no caller is told it
may open a replacement operation.

---------

Co-authored-by: Epinephrine <27862058+luvs01@users.noreply.github.com>
Co-authored-by: JUN <jun@lidgeai.com>
…ach log rows (#6903)

* fix(logs): drop credential-shaped upstream diagnostics before they reach log rows

* fix(logs): record only known upstream error classes

* fix(logs): keep only known bare top-level error codes
@lidge-jun
lidge-jun requested a review from Ingwannu as a code owner October 10, 2026 18:08
@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Oct 10, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-10-10T18:43:41.536998Z eb06a2c Draft marked ready
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@github-actions

Copy link
Copy Markdown
Contributor

✅ Deterministic PR hygiene checks passed.

@coderabbitai

coderabbitai Bot commented Oct 10, 2026

Copy link
Copy Markdown
Contributor

Walkthrough

Warning

Review details and warnings were omitted to fit the comment limit.

@github-actions github-actions Bot changed the title release: 2.83.0 [WRONG BRANCH] release: 2.83.0 Oct 10, 2026
@github-actions

github-actions Bot commented Oct 10, 2026 •

Copy link
Copy Markdown
Contributor

⏳ DRAFT

  • wrong target branch (main); retarget to dev.

What to do

  • Retarget this PR to dev — all contributions go to dev.

Its title has been prefixed with [WRONG BRANCH].
Automatic draft conversion failed (token cannot change draft status). Please convert this pull request to a draft manually. The required enforce-target check will keep failing until every issue above is resolved.

@github-actions
github-actions Bot marked this pull request as draft October 10, 2026 18:08

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: eb06a2c418

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

} }, 409, req, config);
let opRecord: GrokResetCouponOperationRecord;
try {
opRecord = openGrokResetCouponOperation({ accountId, tokenId: requestedTokenId, operationId: effectiveOpId });

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Replay settled coupon operations before OAuth resolution

Replay lookup occurs only after getValidAccessSnapshotForAccount, so if a redemption succeeded but its response was lost and the xAI credential is then expired, revoked, or removed, retrying the preserved operation ID returns auth_failed instead of the durable redeemed result. Both the CLI and dashboard treat that code as definitive, allowing the user to create another operation and potentially consume a second coupon. Read and return terminal ledger outcomes before requiring credentials; authenticate only operations that still need upstream inspection or execution.

AGENTS.md reference: AGENTS.md:L180-L182

Useful? React with 👍 / 👎.

Comment on lines +192 to +193
const held = uncertainRef.current.get(accountId) ?? activeRedeems.current.get(accountId);
if (held) return hold(held);

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Let held coupon attempts query their idempotent operation

After a timeout or an uncertain response, every later redeem call returns the local hold without contacting the server, while the dialog's “Re-read account” action performs only the coupon-list GET. If the original request subsequently settles successfully, the management endpoint can now replay that definitive result for the same operation ID, but the dashboard can never obtain it: the account remains blocked until reload, which then discards the safety hold entirely. Recheck the consume endpoint with the stored token and operation ID so its idempotent replay can clear a definitive outcome.

AGENTS.md reference: gui/AGENTS.md:L7-L10

Useful? React with 👍 / 👎.

}
/** Isolate discovery caches by endpoint and credential without storing plaintext cache keys. */
function authority(provider: OcxProviderConfig): string {
return createHash("sha256").update(JSON.stringify([provider.baseUrl, provider.apiKey])).digest("hex");
@lidge-jun

Copy link
Copy Markdown
Owner Author

Gate dispositions for this promotion

enforce-target: expected failure. This is a dev→main promotion, and enforce-target accepts only dev and open stacked bases. The same failure occurred on #6771 and #6878.

CodeQL: 34 "new" alerts, each mapped against the release diff 55557dc64e..c61c3f62e6.

  • 33 are not on any line this release changes. 23 are in files the release does not touch. 10 are on unchanged lines of files it does touch: provider-models.ts (5 hash and 1 URL), cursor/catalog.ts (2 regex), google.ts (escaping and randomness), openai-responses/passthrough.ts (regex) and bridge/errors.ts (stack trace). CodeQL reports them because a data flow passes through changed code, or because the diff is large. They are pre-existing, and this promotion leaves them unchanged.
  • 1 is on a changed line: src/providers/github-copilot-auto.ts:38, "password hash with insufficient computational effort". This is a false positive. authority() computes SHA-256 over [baseUrl, apiKey] only as an in-memory key that isolates the Copilot discovery cache per credential. It is not used for authentication, it never stores a password, and it is never persisted. API keys are high-entropy random tokens, so a slow KDF adds nothing. provider-models.ts:346 already uses the same pattern.

Security review of the release diff is the independent regression audit described in the PR body. It ended with no P0/P1.

@lidge-jun
lidge-jun marked this pull request as ready for review October 10, 2026 18:39
@lidge-jun
lidge-jun merged commit 33b3218 into main Oct 10, 2026
146 of 152 checks passed
@lidge-jun
lidge-jun deleted the codex/promote-main-2.83.0-r2 branch October 10, 2026 18:39

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

await redeemGrokResetCoupon({
accessToken: tokenSnapshot.accessToken,
tokenId: resolvedTokenId,
});

P2 Badge Bound the coupon redemption request

When Grok accepts the connection but never completes the response, this call has neither an independent deadline nor a signal, so the promise can remain pending after the dashboard's 30-second fetch timeout. The finally block never clears grokCouponInFlightAttempts, every retry of the preserved operation returns attempt_in_progress, and a reload can discard the UI hold and allow a second operation to spend another coupon. Pass a server-owned bounded signal to redeemGrokResetCoupon so a timeout leaves the durable operation in the existing uncertain state without tying completion to the client disconnect.

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

8 participants