Release/v6.8.4 - #128
Release/v6.8.4#128
Conversation
Replace the Axon context-window model variants with an OSS catalog served through the MatterAI gateway: meta/muse-spark-1.2-contributor, deepseek/deepseek-v4-flash-0731, zai/glm-5.3, zai/glm-5.3-flash, gpt-5.6-sol, gpt-5.6-luna, and gemini-3.7-flash. - Extension catalog (kilocode-models.ts) and the webview copy (useOpenRouterModelProviders) now share a single OSS_MODEL_BASE variant with per-model published pricing (USD per token, OpenRouter format) and 1:1 API model ids; getKilocodeApiModelId passes unknown ids through. - Default model is deepseek/deepseek-v4-flash-0731: openRouterDefaultModelId in @roo-code/types, the CLI kilocodeModel defaults (config defaults, provider settings, browser auth), and web-evals MODEL_DEFAULT. - Model selector drops the 232k/400k context toggle and its sort-priority helper; stored Axon selections reset to the default on next launch via the existing isValidKilocodeModel stale-model check. - Update kilocode-models, kilocode-openrouter, and CLI provider-merge tests for the new catalog and default.
…ng blocks after a read-only batch - AssistantMessageParser: synthesize a stable `native-tool-call-<index>` id when a delta carries an index but no id instead of dropping the call. Some OpenAI-compatible providers never send an id (or only send it in a later fragment); dropping the first delta silently discarded the whole tool call, and every later argument fragment for that index then hit "arguments for unknown tool call" and was dropped too. `name` is now optional in NativeToolCall since only the first delta carries it. - presentAssistantMessage: after a parallel read-only batch stops at the first non-parallelizable block (e.g. execute_command following a run of searches), continue the presentation chain so trailing blocks still execute; the batch only marks userMessageContentReady when it consumed every block, instead of firing the next API request early. - The tool-failure fallback only signals readiness for the LAST content block, so a mid-message failure no longer tells the task loop the message is complete while later blocks still have to run. - Regression tests for id-less tool call accumulation, the trailing execute_command after a read-only batch, and mid-message failure readiness.
…ff prefix Port codex-style context window management so the model adapts before and after auto-compaction instead of redoing work: - Task: when usage enters a 10-point band below the condense threshold (CONTEXT_WARNING_BAND_PERCENT), push a one-time <context_window_warning> user message telling the model to finish in-flight edits, stop broad searches and large reads, and keep the todo list current before compaction replaces raw tool outputs with a summary. The warning re-arms when usage drops back below the band (e.g. after a successful condensation), so each context window gets at most one warning. - Condense: the summary prompt now records exploration already performed (searches, reads, investigations and their conclusions) and demands enough detail (paths, signatures, line numbers, error text) to continue without re-reading files. The inserted summary message is prefixed with a [CONTEXT COMPACTION] handoff note (SUMMARY_PREFIX) instructing the model to treat the summary as the authoritative record and continue from NEXT STEPS instead of re-searching, re-reading, or re-deriving. - Update condense tests for the prefix and new prompt wording.
…pace rules - file_edit/multi_file_edit guidance now keys on edits that are confirmed and ready now rather than a blanket "always batch 2+ edits": make a ready edit immediately instead of holding it back, keep batches small and cohesive (the edits belonging to the current step), and never accumulate a whole task into one giant multi-file edit. - Replace "Plan before editing" with "Edit early, iterate in small steps": make the first edit as soon as one file's change is confirmed, alternate editing and checking (typecheck, test, targeted read) as the intended workflow, and track remaining work with update_todo_list instead of holding a full multi-file plan in context. Do not re-read a file just to confirm an edit succeeded. - Add a Multi-repo workspaces section: work inside the repo that owns the code being changed, cross into another repo only when the task requires it, and never interleave reads across repos. - Refresh prompt snapshots.
- Bump the extension version from 6.8.2 to 6.8.4. - Add the v6.8.3 changelog entry documenting the OSS model catalog that replaces the Axon models, including per-model pricing and the new deepseek/deepseek-v4-flash-0731 default. - Consume the applied changesets (axon default model, pasted image paths, Fireworks provider removal).
…ility - Fetch the model catalog from the MatterAI backend and register it via registerDynamicKilocodeModels, replacing static hardcoded lists; refresh on window focus, manual refresh, and a 10-minute poller with forceRefresh cache bypass - Show per-model usage (weekly/monthly share of the shared plan pool, cost-multiplier badges) in the usage dialog and a new read-only Settings Model Usage section; /axoncode/profile now returns modelUsage - Fix refreshKilocodeModels wiping OpenRouter models by fetching openrouter and kilocode-openrouter in parallel and merging results - Fetch the catalog at provider startup so stale axon model selections are reset via isValidKilocodeModel against a populated catalog
- Add iconUrl to ModelInfo and forward it from the MatterAI /v1/web/models catalog through parseOpenRouterModel, registerDynamicKilocodeModels, and the static KILO_CODE_MODELS loop - New ProviderLogo component renders the SVG on a white circular background (visible on dark themes) - Show provider logos in the model selector (dropdown items + trigger), usage dialog, and settings model usage section - Add iconUrl to AxonCodeModelUsage for /axoncode/profile model usage entries
Replace the retired axon-mini/axon-code/axon-code-2 table with the seven OSS models served dynamically from the MatterAI backend catalog, with providers, credit multipliers, and use cases.
There was a problem hiding this comment.
🧪 PR Review is completed: Release v6.8.4 introduces dynamic model catalogs, context budget warnings, and model usage UI. Found a pricing/description mismatch on the new default model info, a redundant double-fetch on window focus in the webview router-models hook, debug console.log in Task.ts, and an as any bypassing the new response payload type. Reviewed src/api/providers/fetchers/modelCache.ts: no issues found (auth headers composed correctly, forceRefresh plumbed through). Reviewed src/api/providers/fetchers/openrouter.ts: no issues found (dynamic fetch is try/catch guarded; the ~55-line duplication between the two fetch paths is acceptable for now). Reviewed src/api/providers/kilocode-models.ts: no issues found (scientific-notation price coercion is documented as lossless via parseFloat). Reviewed src/core/assistant-message/AssistantMessageParser.ts: no issues found (synthesized tool-call id stays stable via the index→id map). Reviewed src/core/assistant-message/presentAssistantMessage.ts: no issues found (readiness gating and batch continuation are correct; cleanup of lock before recursion verified). Reviewed src/core/webview/ClineProvider.ts: no issues found (modelsRefreshInterval is unref'd and cleared on dispose; focus refresh is throttled). Reviewed src/core/webview/webviewMessageHandler.ts: no issues found. Reviewed src/core/condense/index.ts: no issues found (SUMMARY_PREFIX is exported and consumed consistently). Reviewed src/core/prompts/system.ts and snapshot files: intentional prompt updates, no issues. Reviewed webview-ui/src/components/kilocode/chat/ModelSelector.tsx: no issues found. Reviewed webview-ui/src/components/ui/hooks/useOpenRouterModelProviders.ts: no issues found (static catalog intentionally emptied; fallback provider is display-only). Reviewed test files (kilocode-models.spec.ts, kilocode-openrouter.spec.ts, AssistantMessageParser.spec.ts, presentAssistantMessage.spec.ts, ClineProvider.spec.ts, condense specs): no issues found.
Skipped files
.changeset/axon-auto-default-model.md: Skipped file pattern.changeset/pasted-image-paths.md: Skipped file pattern.changeset/remove-fireworks-provider.md: Skipped file patternCHANGELOG.md: Skipped file patternREADME.md: Skipped file patternwebview-ui/src/i18n/locales/en/settings.json: Skipped file pattern
⬇️ Low Priority Suggestions (4)
packages/types/src/providers/openrouter.ts (1 suggestion)
Location:
packages/types/src/providers/openrouter.ts(Lines 12-16)🟡 Data Integrity / Business Logic
Issue:
openRouterDefaultModelIdis now"deepseek/deepseek-v4-flash-0731"(L4), but the accompanyingopenRouterDefaultModelInfodescribes "GLM 5.3 Flash" and carries GLM 5.3 Flash pricing (input $0.15/M, output $0.50/M, cache read $0.03/M). Per the catalog insrc/api/providers/kilocode-models.ts, DeepSeek V4 Flash is priced at $0.14/M input, $0.28/M output, $0.028/M cache read. Any cost/usage calculation or UI display that leans on this default ModelInfo will report the wrong model name and wrong prices for the default model.Fix: Align the description and pricing fields with DeepSeek V4 Flash's published rates.
Impact: Correct cost estimation and model metadata for the default model.
- inputPrice: 0.00000015, - outputPrice: 0.0000005, - cacheWritesPrice: 0, - cacheReadsPrice: 0.00000003, - description: "GLM 5.3 Flash is a fast, low cost open model for everyday coding tasks.", + inputPrice: 0.00000014, + outputPrice: 0.00000028, + cacheWritesPrice: 0, + cacheReadsPrice: 0.000000028, + description: "DeepSeek V4 Flash is a fast, low cost open model for low-effort day-to-day coding tasks.",
webview-ui/src/components/ui/hooks/useRouterModels.ts (1 suggestion)
Location:
webview-ui/src/components/ui/hooks/useRouterModels.ts(Lines 83-88)🟠 Performance
Issue: The new
useQueryconfig setsrefetchOnWindowFocus: truewhile the new customfocuslistener in the same hook already callsqueryClient.invalidateQueries({ queryKey: ["routerModels"] })on the same window-focus event. Both fire per focus, so each focus triggers tworequestRouterModelspostMessage round-trips — and the extension host also force-refreshes models on window focus (ClineProvider.onDidChangeWindowState), making it up to three fetch cycles per focus. Note: in TanStack Query v5refetchOnWindowFocusdefaults to true, so simply deleting the line may not fix it.Fix: Disable react-query's built-in focus refetch and keep the throttled custom listener (which also covers the
didBecomeVisiblemessage) as the single focus trigger.Impact: Halves the webview-side fetch chatter on every window focus; reduces load on the OpenRouter/MatterAI model endpoints.
- return useQuery({ - queryKey: ["routerModels", queryKey], - queryFn: () => getRouterModels(), - refetchInterval: 10 * 60 * 1000, - refetchOnWindowFocus: true, - }) + return useQuery({ + queryKey: ["routerModels", queryKey], + queryFn: () => getRouterModels(), + refetchInterval: 10 * 60 * 1000, + refetchOnWindowFocus: false, + })
src/core/task/Task.ts (1 suggestion)
Location:
src/core/task/Task.ts(Lines 3970-3973)🔵 Logging
Issue: Debug
console.logstatement left in the new context-budget warning path. Task.ts routes diagnostics through its logging facilities; a raw console.log fires on every warning event in the extension host output.Fix: Remove the console.log block; the
<context_window_warning>text pushed touserMessageContentalready carries the same information to the model/UI.Impact: Cleaner extension host output; no behavioral change.
- console.log( - `[Task#${this.taskId}] Context budget warning: ${contextPercent.toFixed(1)}% used ` + - `(~${tokensRemaining} of ${contextWindow} tokens remain), condense threshold ${effectiveThreshold}%.`, - ) +
webview-ui/src/components/settings/ModelUsageSettings.tsx (1 suggestion)
Location:
webview-ui/src/components/settings/ModelUsageSettings.tsx(Lines 111-111)🔵 Type Safety
Issue: The
profileDataResponsehandler castsmessage.payloadwithas any, bypassing theProfileDataResponsePayloadinterface that already exists in@roo/WebviewMessage(and is used by the backend response shape). This discards type checking onpayload.success/payload.datain a brand-new file where the types are readily available.Fix: Cast to the exported payload type instead. Also add
ProfileDataResponsePayloadto the existing type import on line 5.Impact: Compile-time safety for the profile response contract; no runtime change.
- const payload = message.payload as any + const payload = message.payload as ProfileDataResponsePayload | undefined
No description provided.