Skip to content

Fix selected configured model reporting in headless output and session metadata - #364

Merged
Brian Krabach (bkrabach) merged 1 commit into
mainfrom
fix/headless-selected-model-metadata
Oct 1, 2026
Merged

Brian Krabach (bkrabach) merged 1 commit into
mainfrom
fix/headless-selected-model-metadata

Conversation

@bkrabach

@bkrabach Brian Krabach (bkrabach) commented Oct 1, 2026 •

Copy link
Copy Markdown
Collaborator

Problem

Actual headless CLI qualification with OpenAI SDK 3.22.1 found that selecting a non-first mounted provider sent the correct request but reported the first provider's model in JSON and saved session metadata. The initial --provider openai --model gpt-4.1 fixture reached gpt-4.1 while reporting vllm/meta-llama/Llama-3-8B. Four new Sol/Astra JSON/json-trace tests reproduced the same defect in both output and real SessionStore reloads before repair.

Bounded repair

  • Touch only amplifier_app_cli/main.py and tests/test_headless_session_persistence.py.
  • Display the configured conversation model using the existing conversation pin, then streaming-loop priority ordering with stable insertion-order ties. Read get_info().defaults['model'], with legacy attribute fallback. Never borrow a model from an unselected mount.
  • Freeze one label for JSON and persistence before execution, inside the cleanup boundary.
  • Add model_source='configured_default' or 'unknown'. This is configured-default reporting, not a claim of observed per-call routing, account identity, or the server's underlying model deployment.
  • Best-effort reporting failure returns unknown; it must not discard a completed response or bypass cleanup.
  • Preserve provider mount order, settings, routing, session ownership and canonical transcript content. No dependency change or new observer in production.

Trade-off: a scalar configured-default label does not summarize multi-model role routing. This PR makes its provenance explicit rather than adding unscoped event inference or changing selection.

Verification

Exact candidate 2b64f76; existing task-owned Linux-aarch64 DTU; Python 3.13.15, Core 2.0.1, Foundation 89575c3482e3e8afe5a03df72e723cf815fa1f6c, SDK 3.22.1. Normal package resolution and non-editable CLI wheel.

  • Four baseline regression failures: Sol/Astra x JSON/json-trace reported the first mounted vLLM model, despite a selected OpenAI provider.
  • Final default suite: 2478 passed, 4 skipped, 13 deselected, 1 xfailed, exit 0, 26.94 seconds. Skips: real macOS kqueue, two missing hooks-logging production-path checks, real tmux. Deselected tests are existing PTY integration tests.
  • New priority/config/default/stable-tie, missing model, pin, and broken pin/model-property tests. Broken lookup cases prove successful response persistence and cleanup awaited once.
  • Independent bounded source review's cleanup finding repaired and re-reviewed: no remaining critical source issue. Review is not GitHub approval.
  • amplifier-tester's frozen CLI qualification verified 454 archive files and all 101 installed Python files against the candidate; installed main.py SHA256 6e5f3f0fba88f3ca3e455b2fa78a0075446e0c5dd15125990a1b457708e1d32b.
  • Four actual CLI controlled-TCP cases: gpt-6.1-sol SSE, gpt-6-astra SSE, openai/gpt-oss-120b JSON, openai/gpt-oss-120b SSE, all with documented reasoning effort high. All execute the real read_file tool, return unpredictable fixture contents with invocation call_id distinct from output item id on the second HTTP request, and then receive the final success marker.
  • Correct selected instance/model appears in both JSON and actual saved metadata for every case, with model_source='configured_default'. Sol/Astra omit unsupported sampling fields. SDK 3.22.1 remains installed before/after every CLI preparation.
  • GPT-OSS uses production Harmony accounting, no accounting mocks: zero server usage is replaced with validated nonzero computed input/output usage, consistent with canonical provider events.
  • 56 controlled-server assertions pass; 49 installed packages compatible. No shared-venv guard bypass; original DTU settings restored and test servers stopped.
  • Root independently reread all case receipts, SDK snapshots and server assertions, then reran Sol SSE and GPT-OSS SSE against the same exact frozen CLI/provider bytes. Both exited0 with real tool round-trips, correct JSON/saved metadata and validated production Harmony accounting. Independent receipt phase3-20261001-010143-43a1f3 retained separately from the specialist's original result.
  • GitHub CI36799174201 completed success on exact2b64f760: all eight jobs pass, including Python3.11/3.12 on Linux/macOS/Windows and real PTY integration on Linux/macOS. This supplements, rather than relabels, the DTU's explicitly deselected PTY tests.

The coupled vLLM SDK repair is microsoft/amplifier-module-provider-vllm#46 at 86d7d8ebcff4ba95427aa60bc0880d4632962b0e. Its own 333-test four-SDK matrix and all eight CI jobs are green. Neither PR is merged or installed on the live host by this qualification.

Limits: controlled local server, not real hosted model inference or production vLLM deployment; headless single-shot paths, not full interactive UX. An archive install's version banner falls back to 0.1.0 without .git; package metadata remains 0.1.1 and exact source identity is established by hashes, not the banner. Preserve initial failing receipts rather than relabeling them.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants