feat(inference): add llmisvc_model_provider_resolver - #699
Conversation
|
Missing Signed-off-by: 5113bb3. All commits require sign-off (via |
5113bb3 to
d11ef9e
Compare
|
AI tool authorship detected:
Sorry, this project does not accept commits authored by tools as valid. |
d11ef9e to
b0d0e9a
Compare
b53a455 to
d03e60a
Compare
|
Unsigned commits: d03e60a. Please sign your commits. |
| return Some(from_header); | ||
| } | ||
|
|
||
| obj.get("model") |
There was a problem hiding this comment.
Why this if model_to_header ?
There was a problem hiding this comment.
model_to_header puts the model in the header.
But the model serving needs the header to be a "canonical id" publishers//models/<MODEL_NAME>. So the way BBR works in 3.5 is that we have the user put this value in the body "models" value and we move it into the header and replace it with the real <MODEL_NAME>
This resolver is that second part.
External Models do something similar but replacement will happen based on the ExternalModel CRD.
There was a problem hiding this comment.
I mean, this seems to "if it's not in the header as provided by model_to_header", it will fall back to getting the value from where model_to_header would/should have gotten the value.
There was a problem hiding this comment.
Hmm maybe we can sync on this so I can better understand but I would think that is ok right?
It is faster to get this from the header and we fallback to the model if it is not there seems reasonable. But I can just remove the fallback if we think that makes more sense.
There was a problem hiding this comment.
Agreed after sync — dropped the body fallback. This filter now assumes model_to_header (or an equivalent) already set the routing header. If the header is absent/empty, we no-op and leave the body alone.
| return Ok(FilterAction::Continue); | ||
| } | ||
|
|
||
| obj.insert("model".to_owned(), serde_json::Value::String(short_name.to_owned())); |
There was a problem hiding this comment.
Is there ever a chance that "model" won't be in the request body on a valid request?
There was a problem hiding this comment.
Not if it is following the OpenAI spec I don't believe, and it has to follow that spec if it is a vLLM model which should be the only thing that this filter would catch.
There was a problem hiding this comment.
Ok. I'm hinting at we could pretty easy do a StringBuffer splice to inject the new model value to avoid the full DOM deserialize and serialize of the complete body. (handling missing model fields is a little bit more tricky, but a replace in a buffer is easy)
There was a problem hiding this comment.
What you are pointing at: today we serde_json::from_slice the whole request body into a DOM, mutate the "model" field, then re-serialize with replace_json_body. That is correct but pays a full parse/serialize of the JSON for a single string field change.
A StringBuffer-style splice would instead find the existing "model":"..." bytes in the buffered body and overwrite/replace just that string in place (or with a small rebuild of the surrounding bytes), avoiding a full DOM round-trip. That is attractive for hot paths, especially once we know "model" is always present on OpenAI/vLLM chat completions which is all this should be targeting currently.
For this PR I am keeping the DOM path: we already share replace_json_body with other filters, and with the header-required + no-invent-model guards the edge cases are simpler. Happy to follow up with a targeted byte/string splice later if profiling shows the parse/serialize cost matters here.
d03e60a to
8f9c832
Compare
praxis-bot
left a comment
There was a problem hiding this comment.
praxis-bot review
Clean filter implementation with good test coverage, proper config validation, and correct use of replace_json_body. One edge case around body mutation when the "model" field is absent.
Findings: 1 medium
| return Ok(FilterAction::Continue); | ||
| } | ||
|
|
||
| obj.insert("model".to_owned(), serde_json::Value::String(short_name.to_owned())); |
There was a problem hiding this comment.
[Medium] When resolve_model_name obtains the model from the header (not the body), the body might not contain a "model" key at all. In that case obj.insert(...) silently adds a "model" field the caller never sent, which could surprise backends that do not expect one (e.g. non-completions endpoints that happen to share a pipeline).
Guard the insert so it only rewrites an existing field:
if !obj.contains_key("model") {
return Ok(FilterAction::Continue);
}
obj.insert("model".to_owned(), serde_json::Value::String(short_name.to_owned()));Add a unit test: header carries a publisher ID, body is {"messages":[]} (no "model"), assert the body is unchanged after the filter runs.
There was a problem hiding this comment.
Fixed in the latest push: we only rewrite when the body already has a "model" key; otherwise no-op (no invented field). Also added a unit test for header publisher ID + body without "model".
praxis-bot
left a comment
There was a problem hiding this comment.
praxis-bot review (round 2)
Fix commit is clean: backtick formatting, CI metrics_route field, and test loop refactor all look correct. One nit below. The prior Medium (body mutation when "model" absent) remains open.
Findings: 1 nit
| assert_eq!( | ||
| llmisvc_short_model_name("publishers/ns/models/a/b"), | ||
| Some("a/b"), | ||
| "SplitN keeps remainder after first /models/" |
There was a problem hiding this comment.
[Nit] Assertion message says SplitN but the implementation (line 279) uses split_once. They are semantically similar, but the message should match the actual method for accuracy.
There was a problem hiding this comment.
Fixed — assertion message now says split_once.
praxis-bot
left a comment
There was a problem hiding this comment.
praxis-bot review (round 3)
Clean implementation with good separation of concerns. Config validation (deny_unknown_fields, empty-header rejection, validate_max_body_bytes), the header-vs-body resolution chain, and the split_once-based publisher-ID parsing are all correct. Integration tests exercise the core rewrite, routing-header preservation, and non-publisher passthrough end-to-end. The register.rs test refactor to a loop is a nice cleanup.
The prior Medium from round 1 (body mutation when "model" absent -- obj.insert(...) adds a field the caller never sent) remains the only actionable concern and is not yet addressed.
Findings: 0 new (prior Medium still open)
a1d79f9 to
b6c729b
Compare
praxis-bot
left a comment
There was a problem hiding this comment.
PR Re-Review
Summary: Previous findings (Medium: body mutation when model absent, Nit: SplitN assertion message) are both resolved. Commit 8656cdd5 added the guard to prevent inventing a body "model" field, and the test leaves_body_unchanged_when_header_publisher_id_but_no_body_model exercises this path. The assertion message now correctly references split_once. One new medium finding on missing test coverage.
| Severity | Count |
|---|---|
| Medium | 1 |
Signed-off-by: Sébastien Han <seb@redhat.com>
|
@jland-redhat please rebase |
Signed-off-by: jland <jland@redhat.com>
Backtick CamelCase identifiers for doc_markdown, shrink tests under too_many_lines, and set metrics_route when building against praxis main. Signed-off-by: jland <jland@redhat.com>
Assume model_to_header already set the routing header; no-op when it is missing. Do not invent a body model field, and rustfmt the llmisvc tests. Signed-off-by: jland <jland@redhat.com>
Signed-off-by: jland <jland@redhat.com>
b6c729b to
76cb977
Compare
Sync with ai-gateway-controller hack/tekton: default PRAXIS_EXTPROC_IMAGE to odh-praxis-extproc:pr699-76cb977 (praxis-proxy/ai#699 @ 76cb977) with TODO(before-merge) to revert once published. Drop must-gather step. Co-authored-by: Cursor <cursoragent@cursor.com>
Default PRAXIS_EXTPROC_IMAGE to the praxis-proxy/ai#699 build (llmisvc_model_provider_resolver) in prow/Tekton and enable the filter in the vendored praxis configmap. Mark all overrides with TODO(before-merge) and keep external-model pytest excluded until the reconciler lands. Co-authored-by: Cursor <cursoragent@cursor.com>
Default PRAXIS_EXTPROC_IMAGE to the praxis-proxy/ai#699 build (llmisvc_model_provider_resolver) in prow/Tekton and enable the filter in the vendored praxis configmap. Mark all overrides with TODO(before-merge) and keep external-model pytest excluded until the reconciler lands. Co-authored-by: Cursor <cursoragent@cursor.com> Signed-off-by: jland <jland@redhat.com>
f907e95 to
36d2af1
Compare
Defer model_to_header promotion to end-of-stream so llmisvc can observe the pending routing header during StreamBuffer pre-read, and add example config, clippy, and coverage test fixes. Signed-off-by: jland <jland@redhat.com>
36d2af1 to
d427d18
Compare
praxis-bot
left a comment
There was a problem hiding this comment.
praxis-bot review (round 5)
All three prior findings are resolved: the body-mutation guard is in place, the assertion message matches split_once, and the deny_unknown_fields test is present. One new finding on license headers.
| Severity | Count |
|---|---|
| Medium | 1 |
Upgrade rustls to 0.23.45 for RUSTSEC-2026-0285 and use Apache-2.0 SPDX headers on llmisvc filter and integration test files per repository convention. Signed-off-by: jland <jland@redhat.com>
shaneutt
left a comment
There was a problem hiding this comment.
I have comments to resolve, but most of them are about language as opposed to structure. Thanks @jland-redhat 👍
Signed-off-by: jland <jland@redhat.com>
0846114 to
f740c80
Compare
Pin praxis-ai-apis and praxis-ai-filters to praxis-proxy/ai#699 and align praxis-core/filter on crates.io v0.5.4 so PipelineExtension stays unified. Initialize new HttpFilterContext fields and add an llmisvc example config. Signed-off-by: jland <jland@redhat.com> Co-authored-by: Cursor <cursoragent@cursor.com>
* ci: add MaaS e2e tests and Konflux group-test pipeline Vendor MaaS e2e infrastructure via remote sync, add ai-gateway-controller prow runner (excluding external-model tests), and wire Konflux group testing with enable-group-testing on PR builds. Signed-off-by: jland <jland@redhat.com> Co-authored-by: Cursor <cursoragent@cursor.com> Signed-off-by: jland <jland@redhat.com> * fix(ci): use operator mode and latest MaaS images for e2e Deploy without local maas-controller/ tree: operator mode with maas-api:latest and maas-controller:latest from MaaS main pushes, PR ai-gateway-controller image unchanged from Konflux snapshot. Signed-off-by: jland <jland@redhat.com> Co-authored-by: Cursor <cursoragent@cursor.com> Signed-off-by: jland <jland@redhat.com> * fix(ci): default e2e to kustomize deploy with synced MaaS manifests Avoid operator-mode prerequisite races by syncing deployment/ from MaaS main, keeping a custom deploy-platform.sh, and hardening platform install (idempotent operators, scoped AuthPolicy waits, ocproute gateway defaults). Co-authored-by: Cursor <cursoragent@cursor.com> Signed-off-by: jland <jland@redhat.com> * fix(ci): hand off ext_proc from maas IPP to praxis-extproc Delete maas-controller legacy payload-processing before ai-gateway-controller installs praxis, protect resources with managed=false, and pin the parameters ConfigMap to opendatahub so RELATED_IMAGE_ODH_PRAXIS_EXTPROC_IMAGE resolves. Co-authored-by: Cursor <cursoragent@cursor.com> Signed-off-by: jland <jland@redhat.com> * fix(ci): exclude synced MaaS deployment from typos deployment/ is copied from models-as-a-service for kustomize e2e and uses Envoy terms (INSERT_BEFORE, otel_als_cluster) that trigger false positives. Co-authored-by: Cursor <cursoragent@cursor.com> Signed-off-by: jland <jland@redhat.com> * fix(ci): tee prow output to prow.log in group test pipeline Capture full e2e step stdout/stderr in CI artifacts so Konflux failures show the actual error instead of only post-mortem auth-debug dumps. Signed-off-by: jland <jland@redhat.com> Co-authored-by: Cursor <cursoragent@cursor.com> Signed-off-by: jland <jland@redhat.com> * fix(ci): use path-based inference URL when models API returns BBR gateway root BBR clusters return model URLs with no path (gateway root only). validate-deployment was rewriting those to HOST/v1/chat/completions which has no HTTPRoute. Export E2E_MODEL_PATH from the prow runner and teach validate-deployment to use it. Co-authored-by: Cursor <cursoragent@cursor.com> Signed-off-by: jland <jland@redhat.com> * fix(ci): improve routing artifacts, drop must-gather, align praxis PR tag - Collect full HTTPRoute/LLMIS/MaaSModelRef YAML under llm-routing/ - Expand cluster-state.log with llm namespace routes and ext_proc images - Remove oc adm must-gather from group-test pipeline (slow, low signal) - Set PRAXIS_EXTPROC_IMAGE from snapshot or matching odh-pr tag in CI Co-authored-by: Cursor <cursoragent@cursor.com> Signed-off-by: jland <jland@redhat.com> * ci(e2e): pin praxis-extproc to pr699-76cb977 for Konflux group-test Default PRAXIS_EXTPROC_IMAGE to the praxis-proxy/ai#699 build (llmisvc_model_provider_resolver) in prow/Tekton and enable the filter in the vendored praxis configmap. Mark all overrides with TODO(before-merge) and keep external-model pytest excluded until the reconciler lands. Co-authored-by: Cursor <cursoragent@cursor.com> Signed-off-by: jland <jland@redhat.com> * fix(ci): annotate AITenant before pausing maas-controller webhook AITenant validation is served by maas-controller-webhook-service. Scaling maas-controller to 0 for the IPP handoff left no endpoints and blocked the praxis opt-in annotation in Konflux e2e. Signed-off-by: jland <jland@redhat.com> Co-authored-by: Cursor <cursoragent@cursor.com> * ci(e2e): point praxis PR699 image at quay.io/maas Publish path for llmisvc_model_provider_resolver CI pin is quay.io/maas/odh-praxis-extproc:pr699-76cb977 (maas org on Quay). Signed-off-by: jland <jland@redhat.com> Co-authored-by: Cursor <cursoragent@cursor.com> * fix(ci): resume maas-controller webhook during praxis handoff ai-gateway-controller must patch AITenant finalizers through maas-controller's validating webhook. Pause maas-controller only for the legacy IPP delete, then resume before installing aigc. Also wait for aigc reconcile, drop stale IPP if wrong image, and require exact PRAXIS_EXTPROC_IMAGE (300s default unchanged). Signed-off-by: jland <jland@redhat.com> Co-authored-by: Cursor <cursoragent@cursor.com> * fix(e2e): praxis IPP log check and restore must-gather artifacts Skip legacy IPP log markers for default praxis-extproc dataplane (praxis is quiet at INFO; routing already verified via HTTP 200). Restore Tekton must-gather step with output redirected to must-gather.log. Signed-off-by: jland <jland@redhat.com> Co-authored-by: Cursor <cursoragent@cursor.com> * docs(e2e): note praxis lacks per-request log markers praxis-extproc does not emit legacy IPP log lines at INFO and :9090 metrics were empty in e2e probes; per-tenant tests skip log checks for praxis default dataplane and rely on HTTP 200 instead. Signed-off-by: jland <jland@redhat.com> Co-authored-by: Cursor <cursoragent@cursor.com> * fix(e2e): harden duplicate-subscription warmup under parallel load Warmup failed with 403 before duplicate-header logic ran. Add API-key propagation delay, extend poll/retry windows, and teach _poll_status to re-check gateway AuthPolicy on transient empty 401/403. Signed-off-by: jland <jland@redhat.com> Co-authored-by: Cursor <cursoragent@cursor.com> * chore(test): gitignore local e2e reports and Python cache Ignore test/e2e/reports/* (CI/debug artifacts) while keeping the directory via .gitkeep; also ignore __pycache__, .pytest_cache, and .venv. Signed-off-by: jland <jland@redhat.com> Co-authored-by: Cursor <cursoragent@cursor.com> * chore(test): fix e2e reports gitignore for IDE visibility Signed-off-by: jland <jland@redhat.com> * ci(e2e): dynamic MaaS checkout instead of vendored tests Fetch models-as-a-service at prow runtime into test/maas-e2e/ (pinned by test/maas-e2e.lock). Keep ai-gateway-controller orchestration in test/e2e/scripts/; remove sync-maas-e2e-tests.sh and vendored tests/deployment/scripts. Add test/e2e/TODO.md with excluded test allowlist and upstream MaaS gaps. Signed-off-by: jland <jland@redhat.com> Co-authored-by: Cursor <cursoragent@cursor.com> * docs(e2e): document fixed MaaS pin at main tip 53fdb8a1 Clarify that prow fetches maas-e2e.lock commit, not rolling main, and document post-fetch patch vs upstream MaaS PR trade-off. Signed-off-by: jland <jland@redhat.com> Co-authored-by: Cursor <cursoragent@cursor.com> * ci(e2e): pin MaaS fork with praxis and poll fixes Point runtime fetch at jland-redhat/models-as-a-service@630e7b5 until upstream merges aigc-e2e-praxis-and-poll-fixes. Co-authored-by: Cursor <cursoragent@cursor.com> Signed-off-by: jland <jland@redhat.com> Co-authored-by: Cursor <cursoragent@cursor.com> Signed-off-by: jland <jland@redhat.com> * fix(ci): preserve AIGC_PROJECT_ROOT after MaaS auth_utils MaaS auth_utils.sh overwrites PROJECT_ROOT with the fetched checkout root, so maas-image-defaults.sh was resolved under test/maas-e2e/. Co-authored-by: Cursor <cursoragent@cursor.com> Signed-off-by: jland <jland@redhat.com> * fix(ci): restore execute bit on patched MaaS scripts Python patches rewrite deploy.sh without +x after fetch; chmod scripts after patching and invoke deploy.sh via bash in deploy-platform. Co-authored-by: Cursor <cursoragent@cursor.com> Signed-off-by: jland <jland@redhat.com> * fix(ci): point deploy-models at MaaS checkout fixtures Dynamic e2e left PROJECT_ROOT on ai-gateway-controller, so deploy-models built test/e2e/fixtures from the wrong repo. Patch deploy-models to use MAAS_CHECKOUT_ROOT and stop patch-maas-deploy from exiting early. Co-authored-by: Cursor <cursoragent@cursor.com> Signed-off-by: jland <jland@redhat.com> * ci(e2e): default praxis-extproc to odh-stable praxis-extproc#79 (9872fc9) merged llmisvc_model_provider_resolver and Quay odh-stable is built from that commit; drop the manual pr699 image pin. Co-authored-by: Cursor <cursoragent@cursor.com> Signed-off-by: jland <jland@redhat.com> * fix(ci): enable tenant namespace discovery during deploy Patch maas-controller before validation and pytest so the rollout finishes while models deploy, not seconds before parallel e2e starts. Signed-off-by: jland <jland@redhat.com> Co-authored-by: Cursor <cursoragent@cursor.com> * fix(ci): add MaaS must-gather and defer IPP migration label Collect maas.opendatahub.io CRs, tenant readiness, and controller logs into gather-maas/ for Konflux artifacts. Document the SkipIPP migration blocker and keep the managed-by pod label workaround commented out until maas-controller handles praxis payload-processing correctly. Co-authored-by: Cursor <cursoragent@cursor.com> Signed-off-by: jland <jland@redhat.com> * fix(tenant): re-enable managed-by label for IPP migration Stamp praxis payload-processing Deployments with app.kubernetes.io/managed-by=ai-gateway-controller so maas-controller SkipIPP cleanup unblocks MaasTenantConfig Ready. Re-enable deploy-script patch and default-tenant wait as a temporary workaround until upstream maas-controller fixes the SkipIPP path. Signed-off-by: jland <jland@redhat.com> Co-authored-by: Cursor <cursoragent@cursor.com> * fix(ci): label praxis writer pods for IPP migration maas-controller ensureIPPWritersStopped inspects live Pod labels, not Deployment spec. SSA-owned Deployments often reject oc patch, so stamp running payload-processing pods directly and re-label when the tenant config reports IPP writer pod still present. Signed-off-by: jland <jland@redhat.com> Co-authored-by: Cursor <cursoragent@cursor.com> * fix(ci): pause maas-controller during praxis pod labeling SkipIPP cleanup runs as soon as praxis writer pods appear, before the managed-by label workaround can run. Pause maas-controller after aigc engages, label every payload-processing pod, then resume and restart maas-controller to retry MaasTenantConfig reconcile. Signed-off-by: jland <jland@redhat.com> Co-authored-by: Cursor <cursoragent@cursor.com> * fix(ci): expand MaaS must-gather for CRs and HTTPRoutes Auto-discover all maas.opendatahub.io and inference.opendatahub.io resources, collect Gateway API HTTPRoutes/Gateways with short-name fallbacks and per-namespace dumps, and add Kuadrant plus Istio gateway networking artifacts that oc adm must-gather omits. Signed-off-by: jland <jland@redhat.com> Co-authored-by: Cursor <cursoragent@cursor.com> * chore(e2e): pin MaaS checkout to upstream main 5ece7d3 Switch maas-e2e.lock from the fork to opendatahub-io/models-as-a-service at 5ece7d3 (PR #1493: praxis IPP log checks and poll flake hardening). fetch-maas-e2e.sh now reads maas_repo from the lock file. Signed-off-by: jland <jland@redhat.com> Co-authored-by: Cursor <cursoragent@cursor.com> * chore(manifests): vendor praxis-extproc d030ea0 for BBR (#82) Bump get-manifests pin to praxis-extproc main after PR #82 enables body-based routing: llmisvc_model_provider_resolver moves to the pre-extproc (BBR) chain where model_to_header runs, not post-auth extproc. Signed-off-by: jland <jland@redhat.com> Co-authored-by: Cursor <cursoragent@cursor.com> * ci(e2e): test MaaS #1505 without IPP migration workarounds Drop aigc-side IPP migration workarounds (pod labeling, maas-controller pause/restart loop, stampAIGCManagedByOnDeployment) now that maas-controller fix/praxis-ipp-writer-pod-skip (PR #1505) skips praxis-owned writer pods. Pin MaaS checkout to 3255890 and maas-controller image to odh-pr-1505. Co-authored-by: Cursor <cursoragent@cursor.com> Signed-off-by: jland <jland@redhat.com> * fix(ci): patch gateway for llm routes and pass E2E_MODEL_PATH to validate Ensure maas-default-gateway allows HTTPRoutes from MODEL_NAMESPACE (llm) and export MAAS_GATEWAY_HOST when running validate-deployment.sh. Co-authored-by: Cursor <cursoragent@cursor.com> Signed-off-by: jland <jland@redhat.com> * ci(e2e): pin MaaS to PR #1508 (ipp-migration-cleanup-complete) Point lock and maas-controller image at ryancham715/rq-95503 (5c36d49 / odh-pr-1508). Cleanup workarounds remain absent; validate against Ryan's IPP backend-swap handoff fix. Co-authored-by: Cursor <cursoragent@cursor.com> Signed-off-by: jland <jland@redhat.com> * Remove standalone praxis-ai hop from the dataplane. ExtProc is the only required dataplane image; ExternalModel routes now backend to provider Services so packaging no longer needs --praxis-image. Co-authored-by: Cursor <cursoragent@cursor.com> Signed-off-by: jland <jland@redhat.com> * fix(controller): align const block for gofmt/gci Signed-off-by: jland <jland@redhat.com> Co-authored-by: Cursor <cursoragent@cursor.com> * e2e: align artifact collection with MaaS patterns Dump DSC, HTTPRoutes, and Gateways into maas-crs/ using the same kubectl get CRD -A -o yaml flow as MaaS collect_maas_crs. Fix must-gather bash -c subshells that could not call _k. Prefer odh-pr-<N> controller image in Tekton. Pin maas-e2e.lock to ryancham715/models-as-a-service. Signed-off-by: jland <jland@redhat.com> Co-authored-by: Cursor <cursoragent@cursor.com> * fix(e2e): opt praxis handoff in on MaasTenantConfig The deploy script annotated AITenant and waited for the praxis finalizer there, but the controller only reads payload-processing-type and writes praxis-cleanup on MaasTenantConfig. Leave maas-controller running so it can signal cleanup-complete before extproc is created. Signed-off-by: jland <jland@redhat.com> Co-authored-by: Cursor <cursoragent@cursor.com> * fix(ci): deploy odh-pr tag and collect billing-style artifacts Group-test was still installing the Konflux snapshot digest, which lags the floating odh-pr-N tag. Artifact collection now follows maas-billing auth_utils (maas-crs, cluster-state, pod logs) and also dumps DataScienceCluster and HTTPRoutes. Signed-off-by: jland <jland@redhat.com> Co-authored-by: Cursor <cursoragent@cursor.com> * fix(e2e): set MaasTenantConfig praxis handoff annotations Opt into praxis on MaasTenantConfig with payload-processing-type=praxis and payload-processing-status=cleanup-complete, verify they stick, and remove legacy IPP so aigc can deploy ExtProc. Signed-off-by: jland <jland@redhat.com> Co-authored-by: Cursor <cursoragent@cursor.com> * fix(ci): share artifacts volume and collect in e2e step Mirror the MaaS group-test pattern: shared emptyDir for maas-crs/pod-logs across e2e, collect-maas-artifacts, must-gather, and git-push. Signed-off-by: jland <jland@redhat.com> Co-authored-by: Cursor <cursoragent@cursor.com> Signed-off-by: jland <jland@redhat.com> * fix(ci): sync group-test pin — unset MAAS_*:latest Mirror odh-konflux-central: leave MAAS images unset so prow maas-image-defaults.sh pins odh-pr-1508 (keeps payload-processing-type). Also use stepTemplate for artifacts volumeMounts (ref steps). Signed-off-by: jland <jland@redhat.com> Co-authored-by: Cursor <cursoragent@cursor.com> * ci: retrigger group-test after maas image pin Signed-off-by: jland <jland@redhat.com> Co-authored-by: Cursor <cursoragent@cursor.com> --------- Signed-off-by: jland <jland@redhat.com> Co-authored-by: Cursor <cursoragent@cursor.com>
Signed-off-by: Sébastien Han <seb@redhat.com>
Summary
llmisvc_model_provider_resolver, porting only the LLMISvc / KServe BBR body-rewrite path from IPP’smodel-provider-resolver.X-Model, aligned withmodel_to_header), fall back to body"model", and when the value is a publisher ID (publishers/.../models/<name>) rewrite the body"model"to<name>only.llmisvc_model_provider_resolver.publisher_idfor metering.Does not port ExternalModel / ExternalProvider resolution, weighted provider selection, Host rewrite, api-format detection, or credential handling.
Sister PR (merge after this)
praxis-extproc: fix ExtProc body-mutation path to setcontent-lengthto the mutated body length (avoids Envoymismatch_between_content_length_and_the_length_of_the_mutated_body/ 500).Without the ExtProc follow-up, body rewrites that change length will fail in Envoy
BUFFERED+ headerSENDmode even though this filter’s rewrite is correct.Test plan
{ "id": "chatcmpl-efb19481-6952-5be9-9572-ebfb8aa9070b", "model": "demo/sim-stream", "object": "chat.completion", "choices": [ { "finish_reason": "stop", "message": { "role": "assistant", "content": "I am fine, how are you today? ..." } } ] }