Skip to content

Release/v6.8.4 - #128

Merged
code-crusher merged 9 commits into
mainfrom
release/v6.8.4
Sep 3, 2026
Merged

code-crusher merged 9 commits into
mainfrom
release/v6.8.4

Conversation

@code-crusher

Copy link
Copy Markdown
Member

No description provided.

matterai-app Bot and others added 9 commits September 2, 2026 15:36
Replace the Axon context-window model variants with an OSS catalog served
through the MatterAI gateway: meta/muse-spark-1.2-contributor,
deepseek/deepseek-v4-flash-0731, zai/glm-5.3, zai/glm-5.3-flash,
gpt-5.6-sol, gpt-5.6-luna, and gemini-3.7-flash.

- Extension catalog (kilocode-models.ts) and the webview copy
  (useOpenRouterModelProviders) now share a single OSS_MODEL_BASE variant
  with per-model published pricing (USD per token, OpenRouter format) and
  1:1 API model ids; getKilocodeApiModelId passes unknown ids through.
- Default model is deepseek/deepseek-v4-flash-0731: openRouterDefaultModelId
  in @roo-code/types, the CLI kilocodeModel defaults (config defaults,
  provider settings, browser auth), and web-evals MODEL_DEFAULT.
- Model selector drops the 232k/400k context toggle and its sort-priority
  helper; stored Axon selections reset to the default on next launch via
  the existing isValidKilocodeModel stale-model check.
- Update kilocode-models, kilocode-openrouter, and CLI provider-merge
  tests for the new catalog and default.
…ng blocks after a read-only batch

- AssistantMessageParser: synthesize a stable `native-tool-call-<index>` id
  when a delta carries an index but no id instead of dropping the call.
  Some OpenAI-compatible providers never send an id (or only send it in a
  later fragment); dropping the first delta silently discarded the whole
  tool call, and every later argument fragment for that index then hit
  "arguments for unknown tool call" and was dropped too. `name` is now
  optional in NativeToolCall since only the first delta carries it.
- presentAssistantMessage: after a parallel read-only batch stops at the
  first non-parallelizable block (e.g. execute_command following a run of
  searches), continue the presentation chain so trailing blocks still
  execute; the batch only marks userMessageContentReady when it consumed
  every block, instead of firing the next API request early.
- The tool-failure fallback only signals readiness for the LAST content
  block, so a mid-message failure no longer tells the task loop the
  message is complete while later blocks still have to run.
- Regression tests for id-less tool call accumulation, the trailing
  execute_command after a read-only batch, and mid-message failure
  readiness.
…ff prefix

Port codex-style context window management so the model adapts before and
after auto-compaction instead of redoing work:

- Task: when usage enters a 10-point band below the condense threshold
  (CONTEXT_WARNING_BAND_PERCENT), push a one-time <context_window_warning>
  user message telling the model to finish in-flight edits, stop broad
  searches and large reads, and keep the todo list current before
  compaction replaces raw tool outputs with a summary. The warning re-arms
  when usage drops back below the band (e.g. after a successful
  condensation), so each context window gets at most one warning.
- Condense: the summary prompt now records exploration already performed
  (searches, reads, investigations and their conclusions) and demands
  enough detail (paths, signatures, line numbers, error text) to continue
  without re-reading files. The inserted summary message is prefixed with
  a [CONTEXT COMPACTION] handoff note (SUMMARY_PREFIX) instructing the
  model to treat the summary as the authoritative record and continue
  from NEXT STEPS instead of re-searching, re-reading, or re-deriving.
- Update condense tests for the prefix and new prompt wording.
…pace rules

- file_edit/multi_file_edit guidance now keys on edits that are confirmed
  and ready now rather than a blanket "always batch 2+ edits": make a
  ready edit immediately instead of holding it back, keep batches small
  and cohesive (the edits belonging to the current step), and never
  accumulate a whole task into one giant multi-file edit.
- Replace "Plan before editing" with "Edit early, iterate in small
  steps": make the first edit as soon as one file's change is confirmed,
  alternate editing and checking (typecheck, test, targeted read) as the
  intended workflow, and track remaining work with update_todo_list
  instead of holding a full multi-file plan in context. Do not re-read a
  file just to confirm an edit succeeded.
- Add a Multi-repo workspaces section: work inside the repo that owns the
  code being changed, cross into another repo only when the task requires
  it, and never interleave reads across repos.
- Refresh prompt snapshots.
- Bump the extension version from 6.8.2 to 6.8.4.
- Add the v6.8.3 changelog entry documenting the OSS model catalog that
  replaces the Axon models, including per-model pricing and the new
  deepseek/deepseek-v4-flash-0731 default.
- Consume the applied changesets (axon default model, pasted image paths,
  Fireworks provider removal).
…ility

- Fetch the model catalog from the MatterAI backend and register it via registerDynamicKilocodeModels, replacing static hardcoded lists; refresh on window focus, manual refresh, and a 10-minute poller with forceRefresh cache bypass
- Show per-model usage (weekly/monthly share of the shared plan pool, cost-multiplier badges) in the usage dialog and a new read-only Settings Model Usage section; /axoncode/profile now returns modelUsage
- Fix refreshKilocodeModels wiping OpenRouter models by fetching openrouter and kilocode-openrouter in parallel and merging results
- Fetch the catalog at provider startup so stale axon model selections are reset via isValidKilocodeModel against a populated catalog
- Add iconUrl to ModelInfo and forward it from the MatterAI /v1/web/models catalog through parseOpenRouterModel, registerDynamicKilocodeModels, and the static KILO_CODE_MODELS loop
- New ProviderLogo component renders the SVG on a white circular background (visible on dark themes)
- Show provider logos in the model selector (dropdown items + trigger), usage dialog, and settings model usage section
- Add iconUrl to AxonCodeModelUsage for /axoncode/profile model usage entries
Replace the retired axon-mini/axon-code/axon-code-2 table with the seven
OSS models served dynamically from the MatterAI backend catalog, with
providers, credit multipliers, and use cases.
@code-crusher
code-crusher merged commit 518c92b into main Sep 3, 2026
1 of 8 checks passed
@code-crusher
code-crusher deleted the release/v6.8.4 branch September 3, 2026 09:57
@matterai-app

matterai-app Bot commented Sep 3, 2026

Copy link
Copy Markdown
Contributor

Summary By MatterAI MatterAI logo ### 🔄 What Changed This pull request (Release/v6.8.4) updates the default OSS model to deepseek/deepseek-v4-flash-0731, adds dynamic catalog registration tests, improves assistant message streaming parser robustness for missing tool IDs, fixes parallel read-only batch scheduling issues, and introduces new model usage settings and UI components. ### 🔍 Impact of the Change Enhances model selection reliability, prevents task loops by ensuring proper execution of trailing non-parallelizable tools, and provides users with granular visibility into plan and model usage windows. ### 📁 Total Files Changed
Click to Expand

File ChangeLog
Schemas Config (apps/web-evals/src/lib/schemas.ts) Updated MODEL_DEFAULT to DeepSeek v4 Flash.
Persistence Test (cli/src/config/__tests__/persistence-provider-merge.test.ts) Added test assertion for kilocodeModel persistence.
CLI Defaults (cli/src/config/defaults.ts) Updated default kilocodeModel configuration.
Model Constants (cli/src/constants/providers/models.ts) Commented out alphabetized sorting on rest model IDs.
Provider Settings (cli/src/constants/providers/settings.ts) Updated default provider settings and PROVIDER_DEFAULT_MODELS.
Browser Auth (cli/src/utils/browserAuth.ts) Configured default kilocodeModel in OAuth persistence.
Model Types (packages/types/src/model.ts) Added optional iconUrl property to model schema.
OpenRouter Types (packages/types/src/providers/openrouter.ts) Updated default OpenRouter model ID and model info.
Kilocode Models Test (src/api/providers/__tests__/kilocode-models.spec.ts) Added comprehensive test suite for dynamic model registration and pricing.
Kilocode OpenRouter Test (src/api/providers/__tests__/kilocode-openrouter.spec.ts) Updated test cases for GLM 5.3 Flash model handler options.
Model Cache (src/api/providers/fetchers/modelCache.ts) Enhanced fetch headers to include Authorization and Cache-Control.
OpenRouter Fetcher (src/api/providers/fetchers/openrouter.ts) Updated dynamic model fetching and iconUrl parsing logic.
Kilocode Models (src/api/providers/kilocode-models.ts) Refactored static catalog to support dynamic backend registration.
Assistant Parser (src/core/assistant-message/AssistantMessageParser.ts) Synthesizes stable IDs for tool calls lacking IDs during streaming.
Parser Tests (src/core/assistant-message/__tests__/AssistantMessageParser.spec.ts) Added test for tool calls missing ID deltas.
Presentation Tests (src/core/assistant-message/__tests__/presentAssistantMessage.spec.ts) Added regression tests for trailing tool execution and failure recovery.
Native Tool Call (src/core/assistant-message/kilocode/native-tool-call.ts) Added optional name property to function delta type.
Present Assistant Message (src/core/assistant-message/presentAssistantMessage.ts) Fixed batch scheduler readiness flag to ensure trailing non-parallel tools execute.
Condense Tests (src/core/condense/__tests__/condense.spec.ts) Updated summary prefix assertions.
Condense Index Tests (src/core/condense/__tests__/index.spec.ts) Updated conversation summary expectations.
Condense Module (src/core/condense/index.ts) Added SUMMARY_PREFIX context handoff constant.
Prompt Snapshots (src/core/prompts/__tests__/__snapshots__/) Updated system and mode prompt snapshots with editing discipline guidelines.
System Prompts (src/core/prompts/system.ts) Added prompt guidelines for early editing and iterative steps.
Task Manager (src/core/task/Task.ts) Implemented context budget warning logic near the condense threshold.
Cline Provider (src/core/webview/ClineProvider.ts) Added background model refresh intervals and stale model validation on startup.
Provider Tests (src/core/webview/__tests__/ClineProvider.spec.ts) Seeded dynamic catalog in provider tests.
Webview Message Handler (src/core/webview/webviewMessageHandler.ts) Added forceRefresh handling for model requests.
Package JSON (src/package.json) Bumped version to 6.8.4.
Webview Messages (src/shared/WebviewMessage.ts) Added AxonCodeModelUsage interface definition.
Shared API (src/shared/api.ts) Added forceRefresh to common fetch parameters.
Chat Row (webview-ui/src/components/chat/ChatRow.tsx) Adjusted border radius and overflow classes for code blocks.
Chat Tabs (webview-ui/src/components/chat/ChatTabs.tsx) Refactored tab styling and active background colors.
Usage Dialog (webview-ui/src/components/chat/UsageDialog.tsx) Added Model Usage section with progress bars and provider logos.
Chat Layout (webview-ui/src/components/chat/chatLayout.ts) Added horizontal padding constant.
Sticky User Message (webview-ui/src/components/kilocode/StickyUserMessage.tsx) Adjusted padding and margin styles.
Model Selector (webview-ui/src/components/kilocode/chat/ModelSelector.tsx) Integrated provider logos and updated model refresh handling.
Provider Models Hook (webview-ui/src/components/kilocode/hooks/useProviderModels.ts) Exposed refetchRouterModels.
Model Usage Settings (webview-ui/src/components/settings/ModelUsageSettings.tsx) Created settings component for plan limits and per-model quotas.
Settings View (webview-ui/src/components/settings/SettingsView.tsx) Added Model Usage settings tab.
Preferred Models Hook (webview-ui/src/components/ui/hooks/usePreferredModels.ts) Commented out model ID sorting.
OpenRouter Providers Hook (webview-ui/src/components/ui/hooks/useOpenRouterModelProviders.ts) Updated static fallback provider mapping.
Router Models Hook (webview-ui/src/components/ui/hooks/useRouterModels.ts) Added window focus and visibility change listeners for query invalidation.
UI Index (webview-ui/src/components/ui/index.ts) Exported ProviderLogo.
Provider Logo (webview-ui/src/components/ui/provider-logo.tsx) Created wrapper component for provider SVG logos on white background.
### 🧪 Test Added/Recommended #### Added - Unit tests for dynamic KiloCode model catalog registration, pricing coercion, and sparse entry fallbacks in `src/api/providers/__tests__/kilocode-models.spec.ts`. - Regression tests for trailing non-parallelizable tool execution and failure recovery in `src/core/assistant-message/__tests__/presentAssistantMessage.spec.ts`. - Streaming parser test for tool calls lacking an explicit ID in `src/core/assistant-message/__tests__/AssistantMessageParser.spec.ts`. #### Recommended - Add integration tests for background model refresh intervals and focus-based cache invalidation in `ClineProvider`. ### 🔒Security Vulnerabilities N/A

@matterai-app matterai-app Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🧪 PR Review is completed: Release v6.8.4 introduces dynamic model catalogs, context budget warnings, and model usage UI. Found a pricing/description mismatch on the new default model info, a redundant double-fetch on window focus in the webview router-models hook, debug console.log in Task.ts, and an as any bypassing the new response payload type. Reviewed src/api/providers/fetchers/modelCache.ts: no issues found (auth headers composed correctly, forceRefresh plumbed through). Reviewed src/api/providers/fetchers/openrouter.ts: no issues found (dynamic fetch is try/catch guarded; the ~55-line duplication between the two fetch paths is acceptable for now). Reviewed src/api/providers/kilocode-models.ts: no issues found (scientific-notation price coercion is documented as lossless via parseFloat). Reviewed src/core/assistant-message/AssistantMessageParser.ts: no issues found (synthesized tool-call id stays stable via the index→id map). Reviewed src/core/assistant-message/presentAssistantMessage.ts: no issues found (readiness gating and batch continuation are correct; cleanup of lock before recursion verified). Reviewed src/core/webview/ClineProvider.ts: no issues found (modelsRefreshInterval is unref'd and cleared on dispose; focus refresh is throttled). Reviewed src/core/webview/webviewMessageHandler.ts: no issues found. Reviewed src/core/condense/index.ts: no issues found (SUMMARY_PREFIX is exported and consumed consistently). Reviewed src/core/prompts/system.ts and snapshot files: intentional prompt updates, no issues. Reviewed webview-ui/src/components/kilocode/chat/ModelSelector.tsx: no issues found. Reviewed webview-ui/src/components/ui/hooks/useOpenRouterModelProviders.ts: no issues found (static catalog intentionally emptied; fallback provider is display-only). Reviewed test files (kilocode-models.spec.ts, kilocode-openrouter.spec.ts, AssistantMessageParser.spec.ts, presentAssistantMessage.spec.ts, ClineProvider.spec.ts, condense specs): no issues found.

Skipped files
  • .changeset/axon-auto-default-model.md: Skipped file pattern
  • .changeset/pasted-image-paths.md: Skipped file pattern
  • .changeset/remove-fireworks-provider.md: Skipped file pattern
  • CHANGELOG.md: Skipped file pattern
  • README.md: Skipped file pattern
  • webview-ui/src/i18n/locales/en/settings.json: Skipped file pattern
⬇️ Low Priority Suggestions (4)
packages/types/src/providers/openrouter.ts (1 suggestion)

Location: packages/types/src/providers/openrouter.ts (Lines 12-16)

🟡 Data Integrity / Business Logic

Issue: openRouterDefaultModelId is now "deepseek/deepseek-v4-flash-0731" (L4), but the accompanying openRouterDefaultModelInfo describes "GLM 5.3 Flash" and carries GLM 5.3 Flash pricing (input $0.15/M, output $0.50/M, cache read $0.03/M). Per the catalog in src/api/providers/kilocode-models.ts, DeepSeek V4 Flash is priced at $0.14/M input, $0.28/M output, $0.028/M cache read. Any cost/usage calculation or UI display that leans on this default ModelInfo will report the wrong model name and wrong prices for the default model.

Fix: Align the description and pricing fields with DeepSeek V4 Flash's published rates.

Impact: Correct cost estimation and model metadata for the default model.

-  	inputPrice: 0.00000015,
-  	outputPrice: 0.0000005,
-  	cacheWritesPrice: 0,
-  	cacheReadsPrice: 0.00000003,
-  	description: "GLM 5.3 Flash is a fast, low cost open model for everyday coding tasks.",
+  	inputPrice: 0.00000014,
+  	outputPrice: 0.00000028,
+  	cacheWritesPrice: 0,
+  	cacheReadsPrice: 0.000000028,
+  	description: "DeepSeek V4 Flash is a fast, low cost open model for low-effort day-to-day coding tasks.",
webview-ui/src/components/ui/hooks/useRouterModels.ts (1 suggestion)

Location: webview-ui/src/components/ui/hooks/useRouterModels.ts (Lines 83-88)

🟠 Performance

Issue: The new useQuery config sets refetchOnWindowFocus: true while the new custom focus listener in the same hook already calls queryClient.invalidateQueries({ queryKey: ["routerModels"] }) on the same window-focus event. Both fire per focus, so each focus triggers two requestRouterModels postMessage round-trips — and the extension host also force-refreshes models on window focus (ClineProvider.onDidChangeWindowState), making it up to three fetch cycles per focus. Note: in TanStack Query v5 refetchOnWindowFocus defaults to true, so simply deleting the line may not fix it.

Fix: Disable react-query's built-in focus refetch and keep the throttled custom listener (which also covers the didBecomeVisible message) as the single focus trigger.

Impact: Halves the webview-side fetch chatter on every window focus; reduces load on the OpenRouter/MatterAI model endpoints.

-  	return useQuery({
-  		queryKey: ["routerModels", queryKey],
-  		queryFn: () => getRouterModels(),
-  		refetchInterval: 10 * 60 * 1000,
-  		refetchOnWindowFocus: true,
-  	})
+  	return useQuery({
+  		queryKey: ["routerModels", queryKey],
+  		queryFn: () => getRouterModels(),
+  		refetchInterval: 10 * 60 * 1000,
+  		refetchOnWindowFocus: false,
+  	})
src/core/task/Task.ts (1 suggestion)

Location: src/core/task/Task.ts (Lines 3970-3973)

🔵 Logging

Issue: Debug console.log statement left in the new context-budget warning path. Task.ts routes diagnostics through its logging facilities; a raw console.log fires on every warning event in the extension host output.

Fix: Remove the console.log block; the <context_window_warning> text pushed to userMessageContent already carries the same information to the model/UI.

Impact: Cleaner extension host output; no behavioral change.

-  					console.log(
-  						`[Task#${this.taskId}] Context budget warning: ${contextPercent.toFixed(1)}% used ` +
-  							`(~${tokensRemaining} of ${contextWindow} tokens remain), condense threshold ${effectiveThreshold}%.`,
-  					)
+  
webview-ui/src/components/settings/ModelUsageSettings.tsx (1 suggestion)

Location: webview-ui/src/components/settings/ModelUsageSettings.tsx (Lines 111-111)

🔵 Type Safety

Issue: The profileDataResponse handler casts message.payload with as any, bypassing the ProfileDataResponsePayload interface that already exists in @roo/WebviewMessage (and is used by the backend response shape). This discards type checking on payload.success / payload.data in a brand-new file where the types are readily available.

Fix: Cast to the exported payload type instead. Also add ProfileDataResponsePayload to the existing type import on line 5.

Impact: Compile-time safety for the profile response contract; no runtime change.

-  				const payload = message.payload as any
+  			const payload = message.payload as ProfileDataResponsePayload | undefined

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant