Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 2 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,6 +2,8 @@

## v0.14.17 (unreleased)

- Fix intermittent upstream 400 on Chat passthrough when the client sends only `max_completion_tokens` (`fix(chat)`, issue #35): the `max_tokens_cap` / `max_tokens_cap_per_model` budget injection unconditionally filled `max_tokens` whenever the client omitted it, and the resulting body then carried **both** `max_tokens` (=cap) and the client's `max_completion_tokens`. OpenAI deprecated `max_tokens` in favor of `max_completion_tokens` and requires them mutually exclusive; opencode zen 的部分后端实例严格校验并拒绝(`max_tokens and max_completion_tokens cannot both be set`),而另一些实例放行 —— 同一请求重放经常成功,表现为偶发失败且对 400 不做重试的客户端(如 ZCode)硬失败。`convertRequest` 现在在客户端已带 `max_completion_tokens` 时跳过 `max_tokens` 注入(双字段均未设置时仍注入 cap;显式 `max_tokens` 的收敛行为不变)。回归测试 `TestConvertRequest_NoMaxTokensInjectionWhenMaxCompletionTokensSet`(chat_bridge_test.go)。

- Fix streamed `tool_calls` landing outside `delta` on reasoning-heavy models (`fix(chat)`, issue #34): `muse-spark-*-contributor` intermittently emits stream `tool_calls` as a sibling of `delta` (`choices[0].tool_calls`) instead of inside it, so standard clients reading `delta.tool_calls` miss the call while `finish_reason` is still `tool_calls` and loop retries. New `hoistChoiceSiblingToolCalls` normalizes the shape at ingress (append after existing `delta.tool_calls`, idempotent) and is wired into every chat stream consumer: direct passthrough (`convertStreamChunkWithUsage`, before case-restore) + stats, `rawSSEReader` native-detect (reserialized), non-stream aggregator (`aggregateOpenAIStream`), `claudeStreamHandler`, `responsesStreamHandler`. Regression tests `TestChatStreamSiblingToolCallsHoistedIntoDelta` + end-to-end `TestChatStreamSiblingToolCalls_EndToEnd`. Live-verified on `muse-spark-1.3-contributor` (`stream:true` 7/7 clean, `delta.tool_calls` × 2 + `finish_reason=tool_calls`); note `stream:false` on the same model separately returns empty `content` with no `tool_calls` (independent issue, not covered).

- Protocol parity vs sub2api `apicompat` (`fix(chat)`, `fix(claude)`):对照 Wei-Shaw/sub2api `backend/internal/pkg/apicompat` 全量审计三条转换链路并补齐 9 项差异,官方 Responses 流事件文档确认事件语义,真实流量三协议 9/9 验证通过。
Expand Down
1 change: 1 addition & 0 deletions docs/API.md
Original file line number Diff line number Diff line change
Expand Up @@ -71,6 +71,7 @@
- `stream`
- `temperature`(闭区间 `0..2`)
- `max_tokens`
- `max_completion_tokens`(clamped 到 `[128, cap]`;已带该字段时不再注入 `max_tokens`,遵循 OpenAI 二者互斥约束,避免上游 400,issue #35)
- `top_p`
- `thinking`
- `reasoning_effort`
Expand Down
2 changes: 1 addition & 1 deletion docs/CONFIGURATION.md
Original file line number Diff line number Diff line change
Expand Up @@ -132,7 +132,7 @@ cp config.example.json config.json

本批次行为补充(仅本版起):

- `max_tokens`(Anthropic/Chat)/ `max_output_tokens`(Responses)全局与按模型上限取 `max_tokens_cap` / `max_tokens_cap_per_model`。配置了 cap 时,**所有**发给上游的请求(Responses/Chat/Anthropic 的直通与翻译路径)缺省该字段都会自动注入 = cap;已设置的收敛到 `[128, cap]`。未配置 cap 时 Chat 直通保持缺省,Anthropic / Chat→Anthropic / Chat→Responses 因 max_tokens 必填兜底 8192。count_tokens 直通只降不补不注入。
- `max_tokens`(Anthropic/Chat)/ `max_output_tokens`(Responses)全局与按模型上限取 `max_tokens_cap` / `max_tokens_cap_per_model`。配置了 cap 时,**所有**发给上游的请求(Responses/Chat/Anthropic 的直通与翻译路径)缺省该字段都会自动注入 = cap;已设置的收敛到 `[128, cap]`。未配置 cap 时 Chat 直通保持缺省,Anthropic / Chat→Anthropic / Chat→Responses 因 max_tokens 必填兜底 8192。count_tokens 直通只降不补不注入。**例外**:Chat 直通路径客户端已带 `max_completion_tokens` 时不补 `max_tokens`(OpenAI 要求二者互斥,部分上游对并存严格 400,issue #35)。
- chat 入站的 `max_completion_tokens` 优先于 `max_tokens` 指导预算(按 Worker A/B 的 OpenAI 现代字段语义),响应走的 `store:false` 且上游为 reasoning 时带 `include:["reasoning.encrypted_content"]` 由 A/B 补齐。
- thinking 模式与 `temperature`/`top_p`/`top_k` 互斥:开启 thinking 时剥离这些采样参数(避免上游 400),同时保留 `output_config.effort` → `reasoning_effort` 映射。
- tool_use/tool_result 配对归一:Claude 历史的 orphan tool_use(无 matching tool_result)在翻译为 chat/responses 前补占位 tool_result 或 drop,保证上游不再因序列非法 400。
Expand Down
8 changes: 7 additions & 1 deletion internal/app/chat.go
Original file line number Diff line number Diff line change
Expand Up @@ -1461,11 +1461,17 @@ func convertRequest(req *OpenAIRequest) map[string]any {
if req.Temperature != nil {
converted["temperature"] = *req.Temperature
}
// 客户端已用现代的 max_completion_tokens(顶层 typed 字段或 extra_body,
// SDK extra_body 顶层合并的惯例)时不得注入 max_tokens —— OpenAI 已废弃
// max_tokens 且要求二者互斥,zen 部分后端对并存严格 400 (cannot both be
// set),注入会把可成功的请求打成偶发硬失败(issue #35)。
clientHasMaxCompletionTokens := req.MaxCompletionTokens != nil ||
extraBodyValue(req, "max_completion_tokens") != nil
if req.MaxTokens != nil {
// clampMaxTokens 复用 anthropic_protocol.go 的 cap 收敛;
// chat 入站的 max_tokens 是客户端可选字段,下限收敛无害。
converted["max_tokens"] = clampMaxTokens(*req.MaxTokens, config.MaxTokensCapFor(req.Model))
} else if cap := config.MaxTokensCapFor(req.Model); cap > 0 {
} else if cap := config.MaxTokensCapFor(req.Model); cap > 0 && !clientHasMaxCompletionTokens {
// 未显式设置时注入 cap:与 responses 直通口径一致(cap 即上游默认
// 预算,避免上游按自身小默认截断)。
converted["max_tokens"] = cap
Expand Down
71 changes: 71 additions & 0 deletions internal/app/chat_bridge_test.go
Original file line number Diff line number Diff line change
Expand Up @@ -630,3 +630,74 @@ func TestChatToResponsesBody_ReasoningSummaryAuto(t *testing.T) {
t.Fatalf("reasoning = %#v, want {effort:ultra summary:auto}", r)
}
}

// Issue #35:客户端只带 max_completion_tokens(不带 max_tokens)时,cap 注入
// 不得补 max_tokens —— OpenAI 已废弃 max_tokens 并要求二者互斥,zen 部分
// 后端对并存严格 400 (cannot both be set),注入造成偶发硬失败。
func TestConvertRequest_NoMaxTokensInjectionWhenMaxCompletionTokensSet(t *testing.T) {
old := config.Get()
config.Update(func(s *config.Snapshot) {
s.MaxTokensCap = 200000
s.MaxTokensCapPerModel = map[string]int{"big-pickle": 200000}
})
t.Cleanup(func() { config.Update(func(s *config.Snapshot) { *s = old }) })

// 仅 max_completion_tokens(按 issue #35 原始线上 JSON 解析):max_tokens
// 不出现,max_completion_tokens 原样转发。
req1 := &OpenAIRequest{}
if err := json.Unmarshal([]byte(`{
"model": "big-pickle",
"stream": true,
"max_completion_tokens": 32000,
"messages": [{"role": "user", "content": "hi"}]
}`), req1); err != nil {
t.Fatal(err)
}
out := convertRequest(req1)
if v, ok := out["max_tokens"]; ok {
t.Fatalf("max_tokens 不应被注入, got %#v", v)
}
if out["max_completion_tokens"] != 32000 {
t.Fatalf("max_completion_tokens = %#v, want 32000", out["max_completion_tokens"])
}

// 双字段均未设置:cap 注入照旧(无 conflict 风险)。
out2 := convertRequest(&OpenAIRequest{
Model: "big-pickle",
Messages: []Message{{Role: "user", Content: "hi"}},
})
if out2["max_tokens"] != 200000 {
t.Fatalf("max_tokens = %#v, want 200000 (cap 注入)", out2["max_tokens"])
}

// 仅 max_tokens:注入路径不介入,显式值照常收敛。
out3 := convertRequest(&OpenAIRequest{
Model: "big-pickle",
Messages: []Message{{Role: "user", Content: "hi"}},
MaxTokens: ptr(5000),
})
if out3["max_tokens"] != 5000 {
t.Fatalf("max_tokens = %#v, want 5000", out3["max_tokens"])
}
if _, ok := out3["max_completion_tokens"]; ok {
t.Fatalf("max_completion_tokens 不应出现")
}

// extra_body 携带 max_completion_tokens(SDK extra_body 顶层合并):
// 同样不得注入 max_tokens,且客户端值原样上行。
req4 := &OpenAIRequest{}
if err := json.Unmarshal([]byte(`{
"model": "big-pickle",
"messages": [{"role": "user", "content": "hi"}],
"extra_body": {"max_completion_tokens": 4096}
}`), req4); err != nil {
t.Fatal(err)
}
out4 := convertRequest(req4)
if v, ok := out4["max_tokens"]; ok {
t.Fatalf("extra_body 场景 max_tokens 不应被注入, got %#v", v)
}
if out4["max_completion_tokens"] != float64(4096) {
t.Fatalf("max_completion_tokens = %#v (%T), want 4096 (extra_body 透传)", out4["max_completion_tokens"], out4["max_completion_tokens"])
}
}
Loading