Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
8 changes: 8 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,6 +2,14 @@

## v0.14.12 (unreleased)

- Protocol parity vs sub2api `apicompat` (`fix(chat)`, `fix(claude)`):对照 Wei-Shaw/sub2api `backend/internal/pkg/apicompat` 全量审计三条转换链路并补齐 9 项差异,官方 Responses 流事件文档确认事件语义,真实流量三协议 9/9 验证通过。
- chat→responses 请求体:`parallel_tool_calls` / `service_tier` 透传上游(顶层键经 `preserveChatPassthroughKeys` 从原始 body 回填 ExtraBody,类型化结构装不下;`extra_body` 显式键优先);`response_format`(`json_schema` 展平 / `json_object` 透传)映射为 `text.format`(此前直接丢弃,structured-output 约束在该路径失效)。
- responses→chat 流/聚合/非流式:`custom_tool_call`(custom/freeform 工具)纳入工具槽位,`custom_tool_call_input.delta/done` 与 function_call 同形累积(done 读 `input` 键);`incomplete` 按 `reason` 细分(`content_filter`→`content_filter`,其余→`length`,此前一律 `length`);上游 `service_tier` 回写 chat 顶层(流 chunk + 聚合 + 非流式);`response.done`(Realtime/WS 别名)与 `completed` 同等终结(此前聚合路径直接原样回吐上游体)。
- chat→anthropic 请求体:thinking 生效即剥离 `temperature`/`top_p`(此前无条件透传,上游 400;与本项目 `convertClaudeRequest` 同口径)。
- claude 响应:`content_filter` 的 `stop_reason` 由非法枚举 `refusal` 改为 `end_turn`(拒绝文本已进内容);`stop` 分支加 tool_use block 存在性回退(防止客户端不回传工具结果、对话卡死)。
- usage/finish 闭集合:`anthropicUsageToChat` 补 `cache_creation_input_tokens`→`prompt_tokens_details.cache_creation_tokens` 归位;`normalizeFinishReason` 未知原因闭集合回退 `stop`(此前透传污染)。
- 回归测试:`protocol_parity_sub2api_test.go` 新增 11 用例(custom 工具同 index/`input` 去重、`content_filter` 双路径、`response.done` 哨兵、`service_tier` 双路径、thinking 剥参、`cache_creation` 归位、未知 finish 闭集合、ExtraBody 回填等)。`docs/API.md` 同步。
- 真实流量(新二进制,18358 direct + 18359 muse-spark responses-rule):direct chat 工具单 index、claude/responses 工具、chat 流 `[DONE]`、responses 流 `completed`、messages 流 `message_stop` 9/9 通过;thinking+temperature 同传 200。
- Fix muse-spark via Claude protocol never hitting upstream prefix-cache (`fix(cache)`): `claudeToResponsesBody` never injected `prompt_cache_key` / `prompt_cache_retention` — the chat→responses, native passthrough, and claude→anthropic paths all had the injection, only the claude→responses path missed it, so `/v1/messages` traffic on muse-spark models never warmed the upstream cache (`stats.json` showed ~3M prompt tokens with zero `cache_read_tokens` while chat-path `big-pickle` cached normally). The body now goes through `applyResponsesCacheHintsToRawBody` (responses-scoped: retention + session-derived key, no top-level `cache_control`), signature gains a `ctx` param at all 4 call sites. Also drops the non-spec top-level `stop` field the converter used to emit (chat path already dropped it). Regression tests in `claude_responses_cache_test.go`.
- Fix non-stream upstream error-in-200 swallowed as empty message (`fix(claude)`): some upstreams return HTTP 200 with an error-shaped body (`{"type":"error","error":{...}}`, e.g. model-side generation failure); `convertResponsesToClaude` failed to parse it as a success response and fell back to an empty text message with no usage, hiding the cause. The non-stream 2xx branches (forward first-try + same-protocol retry + probe) now detect the error shape via `isResponsesErrorBody` and relay it through the existing `convertResponsesErrorToClaude` as HTTP 502 with the upstream message preserved. Stream path already relayed `response.failed`/`error` events.

Expand Down
2 changes: 2 additions & 0 deletions docs/API.md
Original file line number Diff line number Diff line change
Expand Up @@ -100,6 +100,8 @@
- 显式零值的 `temperature`(闭区间 `0..2`)、`top_p`、`frequency_penalty`、`presence_penalty`
- `max_output_tokens`、`stop`、`user`、`parallel_tool_calls`、`stream_options`、`store`
- 函数工具、项目已有的内置工具、`tool_choice`、`reasoning`、`metadata`
- Chat 经 `protocol_rules` 走原生 Responses 上游时:`parallel_tool_calls` / `service_tier` 透传上游;`response_format`(`json_schema` 展平 / `json_object` 透传 `type`)映射为 `text.format`;`custom_tool_call`(custom/freeform 工具)的 `input` 增量与 `done` 按同一 `tool_calls` index 累积;`incomplete` 按 `reason` 细分 `finish_reason`(`max_output_tokens`→`length`,`content_filter`→`content_filter`);上游 `service_tier` 回写 chat 顶层;`response.done`(Realtime/WS 别名)与 `completed` 同等终结流
- Chat 经 `protocol_rules` 走原生 Anthropic 上游时:thinking 生效即剥离 `temperature`/`top_p`(与 Claude 入站同口径,避免上游 400)
- Anthropic-style `tool_result`(`call_id`,缺省时用 `tool_use_id`;`content` 支持 string、字符串数组、`{type:"text"|"input_text"|"output_text",text}` blocks;`is_error:true` 加 `Error: ` 前缀)
- 正常终态 `response.completed`;长度截断终态 `response.incomplete`,reason 为 `max_output_tokens`

Expand Down
9 changes: 7 additions & 2 deletions docs/keypool-design.md
Original file line number Diff line number Diff line change
Expand Up @@ -36,13 +36,18 @@
## 3. 选择器(`internal/app/keypool.go`,内存状态,仿 `socks5RRIndex` 原子模式)

* `keyRRIndex atomic.Uint64`:`round_robin` 取 `Add % len`;`weighted` 取 `Add % totalWeight` 走权重区间;
`sticky` 取 `fnv32a(stickyKeyForRequest 同源串) % len`(与 egress sticky 同键,保证 prompt cache 亲和)。
`sticky` 取 `fnv32a(stickySessionBase) % len`。身份先按下游凭证、再按客户端会话
划分:同一下游 token 的不同客户端会话(`tok:<token>|cli:<会话哈希>`,与 egress
完全同源)散开——会话级负载均衡;同一会话粘定同一 key——prompt cache 亲和。
无 token 的 admin 请求以 `sess:<scope>`(`|` 去掉的会话后缀)散列,会话后缀
为空时所有 token-less 流量仍共用一槽。`attempt>0` 的池 failover 重试混入常量
后缀 `|pool-retry`,跳离刚失败的 key(重试之间仍粘同一备选 key,保持亲和)。
* 候选过滤:`enabled && group匹配 && now > cooldownUntil`;全冷却 → 放行最早过期的那把(不断服)。
* 状态表(`keypoolMu` 守卫):`{cooldownUntil, consecutiveFails}`;成功清零。

## 4. Failover(改 `opencode.go:582` 重试循环内两点,不建新子系统)

* 每 attempt:`attemptAuth, keyID := selectPoolKey(auth, modelID, attempt)`,
* 每 attempt:`attemptAuth, keyID := selectPoolKey(auth, modelID, bodyMap, headers, scope, attempt)`,
以值拷贝传入 `selectUpstreamTarget` + `buildOCRequestWithSubpath` + `invalidateUpstreamTarget`
(三处都收 `auth` 值类型,race-free;sticky egress 按池 key 绑定)。
* 记账(`ReportKeyResult`):
Expand Down
2 changes: 1 addition & 1 deletion internal/app/auth_test.go
Original file line number Diff line number Diff line change
Expand Up @@ -72,7 +72,7 @@ func TestSelectPoolKey_AdminAutoPicked(t *testing.T) {
keypoolMu.Unlock()
}()

auth, keyID, pooled := selectPoolKey(UpstreamAuth{Mode: AuthRouteAdmin, Source: "admin"}, "big-pickle", 0)
auth, keyID, pooled := selectPoolKey(UpstreamAuth{Mode: AuthRouteAdmin, Source: "admin"}, "big-pickle", nil, nil, "", 0)
if !pooled || keyID != "k1" {
t.Fatalf("admin auth should be picked from pool: pooled=%v keyID=%q auth=%+v", pooled, keyID, auth)
}
Expand Down
30 changes: 30 additions & 0 deletions internal/app/chat.go
Original file line number Diff line number Diff line change
Expand Up @@ -356,6 +356,31 @@ func clientStreamUsageWanted(body []byte) bool {
return false
}

// preserveChatPassthroughKeys 把 OpenAIRequest 类型化结构装不下的顶层直通
// 字段从原始 body 回填进 ExtraBody(只补缺,不覆盖 extra_body 内显式键),供
// chat→responses 上游体透传(对齐 sub2api ChatCompletionsToResponses 的
// ParallelToolCalls/ServiceTier)。没有这些键时不建 ExtraBody。
func preserveChatPassthroughKeys(body []byte, req *OpenAIRequest) {
var raw map[string]any
if json.Unmarshal(body, &raw) != nil || req == nil {
return
}
for _, key := range []string{"parallel_tool_calls", "service_tier"} {
v, ok := raw[key]
if !ok || v == nil {
continue
}
if req.ExtraBody != nil {
if _, exists := req.ExtraBody[key]; exists {
continue
}
} else {
req.ExtraBody = map[string]any{}
}
req.ExtraBody[key] = v
}
}

// convertStreamChunkWithUsage 转换流式 chunk,并在同一次解析中顺带返回 usage。
// 注意:流循环(chat.go 的 stream 处理)仍会为流统计单独解析一次 chunk;
// 这里的 "顺带提取" 只是免去了 usage 的第三次解析。
Expand Down Expand Up @@ -492,6 +517,11 @@ func chatCompletionsHandler(w http.ResponseWriter, r *http.Request) {
http.Error(w, "Invalid JSON", http.StatusBadRequest)
return
}
// OpenAIRequest 类型化结构装不下的顶层直通字段(parallel_tool_calls /
// service_tier)从原始 body 回填进 ExtraBody,供 chat→responses 上游体
// 透传(对齐 sub2api ChatCompletionsToResponses)。只补缺,不覆盖客户端
// 已显式放在 extra_body 内的同名键。
preserveChatPassthroughKeys(body, &req)
modelIn := req.Model
req.Model = resolveModelForAuth(auth, req.Model)
if req.Model == "" {
Expand Down
22 changes: 18 additions & 4 deletions internal/app/chat_protocol.go
Original file line number Diff line number Diff line change
Expand Up @@ -10,7 +10,9 @@ import (
)

// normalizeFinishReason maps Anthropic stop reasons onto the closed set used
// by Chat Completions.
// by Chat Completions. Unknown reasons fall back to "stop":Chat
// finish_reason 是闭集合,透传上游新枚举(如 pause_turn)会污染下游(对齐
// sub2api 各映射器的闭集合输出,无透传分支)。
func normalizeFinishReason(reason string) string {
switch reason {
case "end_turn", "stop_sequence", "stop":
Expand All @@ -22,7 +24,7 @@ func normalizeFinishReason(reason string) string {
case "refusal", "content_filter":
return "content_filter"
default:
return reason
return "stop"
}
}

Expand Down Expand Up @@ -51,8 +53,8 @@ func anthropicUsageToChat(usage map[string]any) map[string]any {
}
}
// Anthropic 缓存读/写 token 顶层键透传,并同时归入 chat 约定位置
// prompt_tokens_details.cached_tokens(对齐 sub2api 对 Chat Completions
// usage 的形状;原有顶层键透传保留,不改已有调用方行为)。
// prompt_tokens_details.cached_tokens / .cache_creation_tokens(对齐 sub2api
// 对 Chat Completions usage 的形状;原有顶层键透传保留,不改已有调用方行为)。
if v, ok := numberAsFloat(usage["cache_read_input_tokens"]); ok {
details, _ := out["prompt_tokens_details"].(map[string]any)
if details == nil {
Expand All @@ -63,6 +65,18 @@ func anthropicUsageToChat(usage map[string]any) map[string]any {
}
out["prompt_tokens_details"] = details
}
// cache_creation_input_tokens 同样归位到 details(对齐 sub2api
// promptDetailsFromResponses),否则按 OpenAI 形状计费/统计时写缓存 token 丢失。
if v, ok := numberAsFloat(usage["cache_creation_input_tokens"]); ok {
details, _ := out["prompt_tokens_details"].(map[string]any)
if details == nil {
details = map[string]any{}
}
if existing, eok := numberAsFloat(details["cache_creation_tokens"]); !eok || existing == 0 {
details["cache_creation_tokens"] = v
}
out["prompt_tokens_details"] = details
}
// 上游发 output_tokens_details.thinking_tokens 时归位到 chat 的
// completion_tokens_details.reasoning_tokens。
if outDetails, ok := usage["output_tokens_details"].(map[string]any); ok {
Expand Down
39 changes: 23 additions & 16 deletions internal/app/chat_to_anthropic.go
Original file line number Diff line number Diff line change
Expand Up @@ -373,11 +373,26 @@ func chatToAnthropicBodyWithRaw(req *OpenAIRequest, modelID string, rawBody map[
}
maxTokens := resolveMaxTokens(rawBody, req, modelID)
body["max_tokens"] = maxTokens
if req.Temperature != nil {
body["temperature"] = *req.Temperature
// thinking 与采样参数互斥(对齐 sub2api 与本项目 convertClaudeRequest
// anthropic_protocol.go:119-144):thinking 生效(budget>0 将写入)时剥离
// temperature/top_p,避免上游 Anthropic 400。判定先行,写入分支据此跳过。
thinkingBudget := 0
if !config.ForceDisableThinking() && !isThinkingDisabled(req.Thinking) {
effort := req.ReasoningEffort
if effort == "" {
effort = reasoningEffortFromThinking(req.Thinking)
}
if effort != "" && effort != "none" {
thinkingBudget = effortToThinkingBudget(mappedReasoningEffort(effort))
}
}
if req.TopP != nil {
body["top_p"] = *req.TopP
if thinkingBudget <= 0 {
if req.Temperature != nil {
body["temperature"] = *req.Temperature
}
if req.TopP != nil {
body["top_p"] = *req.TopP
}
}
if stop := extraBodyValue(req, "stop"); stop != nil {
if arr, ok := stop.([]any); ok {
Expand Down Expand Up @@ -417,18 +432,10 @@ func chatToAnthropicBodyWithRaw(req *OpenAIRequest, modelID string, rawBody map[
body["tool_choice"] = choice
}
}
// thinking:effort 映射为预算;ForceDisableThinking 或客户端显式禁用则省略。
if !config.ForceDisableThinking() && !isThinkingDisabled(req.Thinking) {
effort := req.ReasoningEffort
if effort == "" {
effort = reasoningEffortFromThinking(req.Thinking)
}
if effort != "" && effort != "none" {
effort = mappedReasoningEffort(effort)
if budget := effortToThinkingBudget(effort); budget > 0 {
body["thinking"] = map[string]any{"type": "enabled", "budget_tokens": budget}
}
}
// thinking 写入:预算已在上方预计算(thinkingBudget),与采样参数剥离
// 用同一判定,避免两处推导不一致。
if thinkingBudget > 0 {
body["thinking"] = map[string]any{"type": "enabled", "budget_tokens": thinkingBudget}
}
b, err := json.Marshal(body)
if err != nil {
Expand Down
Loading
Loading