Skip to content

feat(inference): add llmisvc_model_provider_resolver - #699

Merged
shaneutt merged 9 commits into
praxis-proxy:mainfrom
jland-redhat:llmisvc_model_provider_resolver
Sep 14, 2026
Merged

shaneutt merged 9 commits into
praxis-proxy:mainfrom
jland-redhat:llmisvc_model_provider_resolver

Conversation

@jland-redhat

@jland-redhat jland-redhat commented Aug 10, 2026 •

Copy link
Copy Markdown
Contributor

Summary

  • Add llmisvc_model_provider_resolver, porting only the LLMISvc / KServe BBR body-rewrite path from IPP’s model-provider-resolver.
  • Prefer a configurable model header (default X-Model, aligned with model_to_header), fall back to body "model", and when the value is a publisher ID (publishers/.../models/<name>) rewrite the body "model" to <name> only.
  • Leave the routing header untouched so KServe can still route on the publisher ID; stash the original ID in llmisvc_model_provider_resolver.publisher_id for metering.
  • Includes unit tests, example config, integration coverage, and generated filter docs.

Does not port ExternalModel / ExternalProvider resolution, weighted provider selection, Host rewrite, api-format detection, or credential handling.

Sister PR (merge after this)

Without the ExtProc follow-up, body rewrites that change length will fail in Envoy BUFFERED + header SEND mode even though this filter’s rewrite is correct.

Test plan

  • Unit tests for rewrite / header preference / body fallback / non-publisher passthrough
  • Example config + integration tests
  • Validated on local cluster with publisher-ID model request; upstream returned a completion:
{
  "id": "chatcmpl-efb19481-6952-5be9-9572-ebfb8aa9070b",
  "model": "demo/sim-stream",
  "object": "chat.completion",
  "choices": [
    {
      "finish_reason": "stop",
      "message": {
        "role": "assistant",
        "content": "I am fine, how are you today? ..."
      }
    }
  ]
}

@jland-redhat
jland-redhat requested review from a team and leseb August 10, 2026 20:38
@praxis-bot-app

Copy link
Copy Markdown

Missing Signed-off-by: 5113bb3. All commits require sign-off (via git commit --signoff).

@praxis-bot-app

Copy link
Copy Markdown

AI tool authorship detected:

  • d11ef9e: Co-authored-by: Cursor <cursoragent@cursor.com>

Sorry, this project does not accept commits authored by tools as valid.
Commits need to be authored by and signed-off by the human(s) responsible for the PR, with their name and contact.

@Jaland
Jaland force-pushed the llmisvc_model_provider_resolver branch from d11ef9e to b0d0e9a Compare August 10, 2026 20:46
@jland-redhat jland-redhat changed the title Adding nnew llmisvc_model_provider_resolver feat(inference): add llmisvc_model_provider_resolver Aug 10, 2026
@Jaland
Jaland force-pushed the llmisvc_model_provider_resolver branch from b53a455 to d03e60a Compare August 10, 2026 21:14
@praxis-bot-app

Copy link
Copy Markdown

Unsigned commits: d03e60a. Please sign your commits.

return Some(from_header);
}

obj.get("model")

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Why this if model_to_header ?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

model_to_header puts the model in the header.

But the model serving needs the header to be a "canonical id" publishers//models/<MODEL_NAME>. So the way BBR works in 3.5 is that we have the user put this value in the body "models" value and we move it into the header and replace it with the real <MODEL_NAME>

This resolver is that second part.

External Models do something similar but replacement will happen based on the ExternalModel CRD.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I mean, this seems to "if it's not in the header as provided by model_to_header", it will fall back to getting the value from where model_to_header would/should have gotten the value.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hmm maybe we can sync on this so I can better understand but I would think that is ok right?

It is faster to get this from the header and we fallback to the model if it is not there seems reasonable. But I can just remove the fallback if we think that makes more sense.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Agreed after sync — dropped the body fallback. This filter now assumes model_to_header (or an equivalent) already set the routing header. If the header is absent/empty, we no-op and leave the body alone.

return Ok(FilterAction::Continue);
}

obj.insert("model".to_owned(), serde_json::Value::String(short_name.to_owned()));

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is there ever a chance that "model" won't be in the request body on a valid request?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not if it is following the OpenAI spec I don't believe, and it has to follow that spec if it is a vLLM model which should be the only thing that this filter would catch.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ok. I'm hinting at we could pretty easy do a StringBuffer splice to inject the new model value to avoid the full DOM deserialize and serialize of the complete body. (handling missing model fields is a little bit more tricky, but a replace in a buffer is easy)

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

What you are pointing at: today we serde_json::from_slice the whole request body into a DOM, mutate the "model" field, then re-serialize with replace_json_body. That is correct but pays a full parse/serialize of the JSON for a single string field change.

A StringBuffer-style splice would instead find the existing "model":"..." bytes in the buffered body and overwrite/replace just that string in place (or with a small rebuild of the surrounding bytes), avoiding a full DOM round-trip. That is attractive for hot paths, especially once we know "model" is always present on OpenAI/vLLM chat completions which is all this should be targeting currently.

For this PR I am keeping the DOM path: we already share replace_json_body with other filters, and with the header-required + no-invent-model guards the edge cases are simpler. Happy to follow up with a targeted byte/string splice later if profiling shows the parse/serialize cost matters here.

@Jaland
Jaland force-pushed the llmisvc_model_provider_resolver branch from d03e60a to 8f9c832 Compare August 11, 2026 00:45

@praxis-bot praxis-bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

praxis-bot review

Clean filter implementation with good test coverage, proper config validation, and correct use of replace_json_body. One edge case around body mutation when the "model" field is absent.

Findings: 1 medium

return Ok(FilterAction::Continue);
}

obj.insert("model".to_owned(), serde_json::Value::String(short_name.to_owned()));

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[Medium] When resolve_model_name obtains the model from the header (not the body), the body might not contain a "model" key at all. In that case obj.insert(...) silently adds a "model" field the caller never sent, which could surprise backends that do not expect one (e.g. non-completions endpoints that happen to share a pipeline).

Guard the insert so it only rewrites an existing field:

if !obj.contains_key("model") {
    return Ok(FilterAction::Continue);
}

obj.insert("model".to_owned(), serde_json::Value::String(short_name.to_owned()));

Add a unit test: header carries a publisher ID, body is {"messages":[]} (no "model"), assert the body is unchanged after the filter runs.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in the latest push: we only rewrite when the body already has a "model" key; otherwise no-op (no invented field). Also added a unit test for header publisher ID + body without "model".

@praxis-bot praxis-bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

praxis-bot review (round 2)

Fix commit is clean: backtick formatting, CI metrics_route field, and test loop refactor all look correct. One nit below. The prior Medium (body mutation when "model" absent) remains open.

Findings: 1 nit

assert_eq!(
llmisvc_short_model_name("publishers/ns/models/a/b"),
Some("a/b"),
"SplitN keeps remainder after first /models/"

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[Nit] Assertion message says SplitN but the implementation (line 279) uses split_once. They are semantically similar, but the message should match the actual method for accuracy.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed — assertion message now says split_once.

@praxis-bot praxis-bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

praxis-bot review (round 3)

Clean implementation with good separation of concerns. Config validation (deny_unknown_fields, empty-header rejection, validate_max_body_bytes), the header-vs-body resolution chain, and the split_once-based publisher-ID parsing are all correct. Integration tests exercise the core rewrite, routing-header preservation, and non-publisher passthrough end-to-end. The register.rs test refactor to a loop is a nice cleanup.

The prior Medium from round 1 (body mutation when "model" absent -- obj.insert(...) adds a field the caller never sent) remains the only actionable concern and is not yet addressed.

Findings: 0 new (prior Medium still open)

@Jaland
Jaland force-pushed the llmisvc_model_provider_resolver branch from a1d79f9 to b6c729b Compare August 14, 2026 16:50

@praxis-bot praxis-bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

PR Re-Review

Summary: Previous findings (Medium: body mutation when model absent, Nit: SplitN assertion message) are both resolved. Commit 8656cdd5 added the guard to prevent inventing a body "model" field, and the test leaves_body_unchanged_when_header_publisher_id_but_no_body_model exercises this path. The assertion message now correctly references split_once. One new medium finding on missing test coverage.

Severity Count
Medium 1

Comment thread filters/src/inference/llmisvc_model_provider_resolver.rs
@shaneutt
shaneutt self-requested a review as a code owner August 28, 2026 17:11
shaneutt pushed a commit that referenced this pull request Aug 28, 2026
Signed-off-by: Sébastien Han <seb@redhat.com>
@leseb

leseb commented Sep 3, 2026

Copy link
Copy Markdown
Collaborator

@jland-redhat please rebase

Signed-off-by: jland <jland@redhat.com>
Backtick CamelCase identifiers for doc_markdown, shrink
tests under too_many_lines, and set metrics_route when
building against praxis main.

Signed-off-by: jland <jland@redhat.com>
Assume model_to_header already set the routing header;
no-op when it is missing. Do not invent a body model
field, and rustfmt the llmisvc tests.

Signed-off-by: jland <jland@redhat.com>
Signed-off-by: jland <jland@redhat.com>
@Jaland
Jaland force-pushed the llmisvc_model_provider_resolver branch from b6c729b to 76cb977 Compare September 13, 2026 22:08
@Jaland
Jaland requested a review from a team as a code owner September 13, 2026 22:08
Jaland pushed a commit to jland-redhat/odh-konflux-central that referenced this pull request Sep 14, 2026
Sync with ai-gateway-controller hack/tekton: default PRAXIS_EXTPROC_IMAGE
to odh-praxis-extproc:pr699-76cb977 (praxis-proxy/ai#699 @ 76cb977) with
TODO(before-merge) to revert once published. Drop must-gather step.

Co-authored-by: Cursor <cursoragent@cursor.com>
jland-redhat added a commit to jland-redhat/ai-gateway-controller that referenced this pull request Sep 14, 2026
Default PRAXIS_EXTPROC_IMAGE to the praxis-proxy/ai#699 build
(llmisvc_model_provider_resolver) in prow/Tekton and enable the filter in
the vendored praxis configmap. Mark all overrides with TODO(before-merge)
and keep external-model pytest excluded until the reconciler lands.

Co-authored-by: Cursor <cursoragent@cursor.com>
jland-redhat added a commit to jland-redhat/ai-gateway-controller that referenced this pull request Sep 14, 2026
Default PRAXIS_EXTPROC_IMAGE to the praxis-proxy/ai#699 build
(llmisvc_model_provider_resolver) in prow/Tekton and enable the filter in
the vendored praxis configmap. Mark all overrides with TODO(before-merge)
and keep external-model pytest excluded until the reconciler lands.

Co-authored-by: Cursor <cursoragent@cursor.com>
Signed-off-by: jland <jland@redhat.com>
@Jaland
Jaland force-pushed the llmisvc_model_provider_resolver branch 2 times, most recently from f907e95 to 36d2af1 Compare September 14, 2026 16:01
Defer model_to_header promotion to end-of-stream so llmisvc can
observe the pending routing header during StreamBuffer pre-read,
and add example config, clippy, and coverage test fixes.

Signed-off-by: jland <jland@redhat.com>
@Jaland
Jaland force-pushed the llmisvc_model_provider_resolver branch from 36d2af1 to d427d18 Compare September 14, 2026 16:03

@praxis-bot praxis-bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

praxis-bot review (round 5)

All three prior findings are resolved: the body-mutation guard is in place, the assertion message matches split_once, and the deny_unknown_fields test is present. One new finding on license headers.

Severity Count
Medium 1

Comment thread filters/src/inference/llmisvc_model_provider_resolver.rs
Upgrade rustls to 0.23.45 for RUSTSEC-2026-0285 and use
Apache-2.0 SPDX headers on llmisvc filter and integration
test files per repository convention.

Signed-off-by: jland <jland@redhat.com>

@shaneutt shaneutt left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I have comments to resolve, but most of them are about language as opposed to structure. Thanks @jland-redhat 👍

Comment thread docs/filters/reference.md Outdated
Comment thread filters/src/inference/llmisvc_model_provider_resolver.rs
Comment thread filters/src/inference/llmisvc_model_provider_resolver.rs Outdated
@Jaland
Jaland force-pushed the llmisvc_model_provider_resolver branch from 0846114 to f740c80 Compare September 14, 2026 20:32
@shaneutt
shaneutt merged commit da29eb4 into praxis-proxy:main Sep 14, 2026
34 checks passed
alexsnaps pushed a commit to opendatahub-io/praxis-extproc that referenced this pull request Sep 15, 2026
Pin praxis-ai-apis and praxis-ai-filters to praxis-proxy/ai#699 and align
praxis-core/filter on crates.io v0.5.4 so PipelineExtension stays unified.
Initialize new HttpFilterContext fields and add an llmisvc example config.

Signed-off-by: jland <jland@redhat.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
openshift-merge-bot Bot pushed a commit to opendatahub-io/ai-gateway-controller that referenced this pull request Sep 19, 2026
* ci: add MaaS e2e tests and Konflux group-test pipeline

Vendor MaaS e2e infrastructure via remote sync, add ai-gateway-controller
prow runner (excluding external-model tests), and wire Konflux group testing
with enable-group-testing on PR builds.

Signed-off-by: jland <jland@redhat.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Signed-off-by: jland <jland@redhat.com>

* fix(ci): use operator mode and latest MaaS images for e2e

Deploy without local maas-controller/ tree: operator mode with
maas-api:latest and maas-controller:latest from MaaS main pushes,
PR ai-gateway-controller image unchanged from Konflux snapshot.

Signed-off-by: jland <jland@redhat.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Signed-off-by: jland <jland@redhat.com>

* fix(ci): default e2e to kustomize deploy with synced MaaS manifests

Avoid operator-mode prerequisite races by syncing deployment/ from MaaS main,
keeping a custom deploy-platform.sh, and hardening platform install (idempotent
operators, scoped AuthPolicy waits, ocproute gateway defaults).

Co-authored-by: Cursor <cursoragent@cursor.com>
Signed-off-by: jland <jland@redhat.com>

* fix(ci): hand off ext_proc from maas IPP to praxis-extproc

Delete maas-controller legacy payload-processing before ai-gateway-controller
installs praxis, protect resources with managed=false, and pin the parameters
ConfigMap to opendatahub so RELATED_IMAGE_ODH_PRAXIS_EXTPROC_IMAGE resolves.

Co-authored-by: Cursor <cursoragent@cursor.com>
Signed-off-by: jland <jland@redhat.com>

* fix(ci): exclude synced MaaS deployment from typos

deployment/ is copied from models-as-a-service for kustomize e2e and uses
Envoy terms (INSERT_BEFORE, otel_als_cluster) that trigger false positives.

Co-authored-by: Cursor <cursoragent@cursor.com>
Signed-off-by: jland <jland@redhat.com>

* fix(ci): tee prow output to prow.log in group test pipeline

Capture full e2e step stdout/stderr in CI artifacts so Konflux failures
show the actual error instead of only post-mortem auth-debug dumps.

Signed-off-by: jland <jland@redhat.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Signed-off-by: jland <jland@redhat.com>

* fix(ci): use path-based inference URL when models API returns BBR gateway root

BBR clusters return model URLs with no path (gateway root only). validate-deployment
was rewriting those to HOST/v1/chat/completions which has no HTTPRoute. Export
E2E_MODEL_PATH from the prow runner and teach validate-deployment to use it.

Co-authored-by: Cursor <cursoragent@cursor.com>
Signed-off-by: jland <jland@redhat.com>

* fix(ci): improve routing artifacts, drop must-gather, align praxis PR tag

- Collect full HTTPRoute/LLMIS/MaaSModelRef YAML under llm-routing/
- Expand cluster-state.log with llm namespace routes and ext_proc images
- Remove oc adm must-gather from group-test pipeline (slow, low signal)
- Set PRAXIS_EXTPROC_IMAGE from snapshot or matching odh-pr tag in CI

Co-authored-by: Cursor <cursoragent@cursor.com>
Signed-off-by: jland <jland@redhat.com>

* ci(e2e): pin praxis-extproc to pr699-76cb977 for Konflux group-test

Default PRAXIS_EXTPROC_IMAGE to the praxis-proxy/ai#699 build
(llmisvc_model_provider_resolver) in prow/Tekton and enable the filter in
the vendored praxis configmap. Mark all overrides with TODO(before-merge)
and keep external-model pytest excluded until the reconciler lands.

Co-authored-by: Cursor <cursoragent@cursor.com>
Signed-off-by: jland <jland@redhat.com>

* fix(ci): annotate AITenant before pausing maas-controller webhook

AITenant validation is served by maas-controller-webhook-service. Scaling
maas-controller to 0 for the IPP handoff left no endpoints and blocked the
praxis opt-in annotation in Konflux e2e.

Signed-off-by: jland <jland@redhat.com>
Co-authored-by: Cursor <cursoragent@cursor.com>

* ci(e2e): point praxis PR699 image at quay.io/maas

Publish path for llmisvc_model_provider_resolver CI pin is
quay.io/maas/odh-praxis-extproc:pr699-76cb977 (maas org on Quay).

Signed-off-by: jland <jland@redhat.com>
Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(ci): resume maas-controller webhook during praxis handoff

ai-gateway-controller must patch AITenant finalizers through
maas-controller's validating webhook. Pause maas-controller only for
the legacy IPP delete, then resume before installing aigc.

Also wait for aigc reconcile, drop stale IPP if wrong image, and require
exact PRAXIS_EXTPROC_IMAGE (300s default unchanged).

Signed-off-by: jland <jland@redhat.com>
Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(e2e): praxis IPP log check and restore must-gather artifacts

Skip legacy IPP log markers for default praxis-extproc dataplane (praxis
is quiet at INFO; routing already verified via HTTP 200). Restore Tekton
must-gather step with output redirected to must-gather.log.

Signed-off-by: jland <jland@redhat.com>
Co-authored-by: Cursor <cursoragent@cursor.com>

* docs(e2e): note praxis lacks per-request log markers

praxis-extproc does not emit legacy IPP log lines at INFO and :9090
metrics were empty in e2e probes; per-tenant tests skip log checks for
praxis default dataplane and rely on HTTP 200 instead.

Signed-off-by: jland <jland@redhat.com>
Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(e2e): harden duplicate-subscription warmup under parallel load

Warmup failed with 403 before duplicate-header logic ran. Add API-key
propagation delay, extend poll/retry windows, and teach _poll_status to
re-check gateway AuthPolicy on transient empty 401/403.

Signed-off-by: jland <jland@redhat.com>
Co-authored-by: Cursor <cursoragent@cursor.com>

* chore(test): gitignore local e2e reports and Python cache

Ignore test/e2e/reports/* (CI/debug artifacts) while keeping the
directory via .gitkeep; also ignore __pycache__, .pytest_cache, and .venv.

Signed-off-by: jland <jland@redhat.com>
Co-authored-by: Cursor <cursoragent@cursor.com>

* chore(test): fix e2e reports gitignore for IDE visibility

Signed-off-by: jland <jland@redhat.com>

* ci(e2e): dynamic MaaS checkout instead of vendored tests

Fetch models-as-a-service at prow runtime into test/maas-e2e/ (pinned by
test/maas-e2e.lock). Keep ai-gateway-controller orchestration in test/e2e/scripts/;
remove sync-maas-e2e-tests.sh and vendored tests/deployment/scripts.

Add test/e2e/TODO.md with excluded test allowlist and upstream MaaS gaps.

Signed-off-by: jland <jland@redhat.com>
Co-authored-by: Cursor <cursoragent@cursor.com>

* docs(e2e): document fixed MaaS pin at main tip 53fdb8a1

Clarify that prow fetches maas-e2e.lock commit, not rolling main, and
document post-fetch patch vs upstream MaaS PR trade-off.

Signed-off-by: jland <jland@redhat.com>
Co-authored-by: Cursor <cursoragent@cursor.com>

* ci(e2e): pin MaaS fork with praxis and poll fixes

Point runtime fetch at jland-redhat/models-as-a-service@630e7b5 until
upstream merges aigc-e2e-praxis-and-poll-fixes.

Co-authored-by: Cursor <cursoragent@cursor.com>
Signed-off-by: jland <jland@redhat.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Signed-off-by: jland <jland@redhat.com>

* fix(ci): preserve AIGC_PROJECT_ROOT after MaaS auth_utils

MaaS auth_utils.sh overwrites PROJECT_ROOT with the fetched checkout
root, so maas-image-defaults.sh was resolved under test/maas-e2e/.

Co-authored-by: Cursor <cursoragent@cursor.com>
Signed-off-by: jland <jland@redhat.com>

* fix(ci): restore execute bit on patched MaaS scripts

Python patches rewrite deploy.sh without +x after fetch; chmod scripts
after patching and invoke deploy.sh via bash in deploy-platform.

Co-authored-by: Cursor <cursoragent@cursor.com>
Signed-off-by: jland <jland@redhat.com>

* fix(ci): point deploy-models at MaaS checkout fixtures

Dynamic e2e left PROJECT_ROOT on ai-gateway-controller, so deploy-models
built test/e2e/fixtures from the wrong repo. Patch deploy-models to use
MAAS_CHECKOUT_ROOT and stop patch-maas-deploy from exiting early.

Co-authored-by: Cursor <cursoragent@cursor.com>
Signed-off-by: jland <jland@redhat.com>

* ci(e2e): default praxis-extproc to odh-stable

praxis-extproc#79 (9872fc9) merged llmisvc_model_provider_resolver and
Quay odh-stable is built from that commit; drop the manual pr699 image pin.

Co-authored-by: Cursor <cursoragent@cursor.com>
Signed-off-by: jland <jland@redhat.com>

* fix(ci): enable tenant namespace discovery during deploy

Patch maas-controller before validation and pytest so the rollout
finishes while models deploy, not seconds before parallel e2e starts.

Signed-off-by: jland <jland@redhat.com>
Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(ci): add MaaS must-gather and defer IPP migration label

Collect maas.opendatahub.io CRs, tenant readiness, and controller logs
into gather-maas/ for Konflux artifacts. Document the SkipIPP migration
blocker and keep the managed-by pod label workaround commented out
until maas-controller handles praxis payload-processing correctly.

Co-authored-by: Cursor <cursoragent@cursor.com>
Signed-off-by: jland <jland@redhat.com>

* fix(tenant): re-enable managed-by label for IPP migration

Stamp praxis payload-processing Deployments with
app.kubernetes.io/managed-by=ai-gateway-controller so maas-controller
SkipIPP cleanup unblocks MaasTenantConfig Ready. Re-enable deploy-script
patch and default-tenant wait as a temporary workaround until upstream
maas-controller fixes the SkipIPP path.

Signed-off-by: jland <jland@redhat.com>
Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(ci): label praxis writer pods for IPP migration

maas-controller ensureIPPWritersStopped inspects live Pod labels, not
Deployment spec. SSA-owned Deployments often reject oc patch, so stamp
running payload-processing pods directly and re-label when the tenant
config reports IPP writer pod still present.

Signed-off-by: jland <jland@redhat.com>
Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(ci): pause maas-controller during praxis pod labeling

SkipIPP cleanup runs as soon as praxis writer pods appear, before the
managed-by label workaround can run. Pause maas-controller after aigc
engages, label every payload-processing pod, then resume and restart
maas-controller to retry MaasTenantConfig reconcile.

Signed-off-by: jland <jland@redhat.com>
Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(ci): expand MaaS must-gather for CRs and HTTPRoutes

Auto-discover all maas.opendatahub.io and inference.opendatahub.io
resources, collect Gateway API HTTPRoutes/Gateways with short-name
fallbacks and per-namespace dumps, and add Kuadrant plus Istio gateway
networking artifacts that oc adm must-gather omits.

Signed-off-by: jland <jland@redhat.com>
Co-authored-by: Cursor <cursoragent@cursor.com>

* chore(e2e): pin MaaS checkout to upstream main 5ece7d3

Switch maas-e2e.lock from the fork to opendatahub-io/models-as-a-service
at 5ece7d3 (PR #1493: praxis IPP log checks and poll flake hardening).
fetch-maas-e2e.sh now reads maas_repo from the lock file.

Signed-off-by: jland <jland@redhat.com>
Co-authored-by: Cursor <cursoragent@cursor.com>

* chore(manifests): vendor praxis-extproc d030ea0 for BBR (#82)

Bump get-manifests pin to praxis-extproc main after PR #82 enables
body-based routing: llmisvc_model_provider_resolver moves to the
pre-extproc (BBR) chain where model_to_header runs, not post-auth extproc.

Signed-off-by: jland <jland@redhat.com>
Co-authored-by: Cursor <cursoragent@cursor.com>

* ci(e2e): test MaaS #1505 without IPP migration workarounds

Drop aigc-side IPP migration workarounds (pod labeling, maas-controller
pause/restart loop, stampAIGCManagedByOnDeployment) now that maas-controller
fix/praxis-ipp-writer-pod-skip (PR #1505) skips praxis-owned writer pods.

Pin MaaS checkout to 3255890 and maas-controller image to odh-pr-1505.

Co-authored-by: Cursor <cursoragent@cursor.com>

Signed-off-by: jland <jland@redhat.com>

* fix(ci): patch gateway for llm routes and pass E2E_MODEL_PATH to validate

Ensure maas-default-gateway allows HTTPRoutes from MODEL_NAMESPACE (llm)
and export MAAS_GATEWAY_HOST when running validate-deployment.sh.

Co-authored-by: Cursor <cursoragent@cursor.com>

Signed-off-by: jland <jland@redhat.com>

* ci(e2e): pin MaaS to PR #1508 (ipp-migration-cleanup-complete)

Point lock and maas-controller image at ryancham715/rq-95503 (5c36d49 /
odh-pr-1508). Cleanup workarounds remain absent; validate against Ryan's
IPP backend-swap handoff fix.

Co-authored-by: Cursor <cursoragent@cursor.com>

Signed-off-by: jland <jland@redhat.com>

* Remove standalone praxis-ai hop from the dataplane.

ExtProc is the only required dataplane image; ExternalModel routes now backend to provider Services so packaging no longer needs --praxis-image.

Co-authored-by: Cursor <cursoragent@cursor.com>

Signed-off-by: jland <jland@redhat.com>

* fix(controller): align const block for gofmt/gci

Signed-off-by: jland <jland@redhat.com>
Co-authored-by: Cursor <cursoragent@cursor.com>

* e2e: align artifact collection with MaaS patterns

Dump DSC, HTTPRoutes, and Gateways into maas-crs/ using the same
kubectl get CRD -A -o yaml flow as MaaS collect_maas_crs. Fix must-gather
bash -c subshells that could not call _k. Prefer odh-pr-<N> controller
image in Tekton. Pin maas-e2e.lock to ryancham715/models-as-a-service.

Signed-off-by: jland <jland@redhat.com>
Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(e2e): opt praxis handoff in on MaasTenantConfig

The deploy script annotated AITenant and waited for the praxis finalizer there, but the controller only reads payload-processing-type and writes praxis-cleanup on MaasTenantConfig. Leave maas-controller running so it can signal cleanup-complete before extproc is created.

Signed-off-by: jland <jland@redhat.com>
Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(ci): deploy odh-pr tag and collect billing-style artifacts

Group-test was still installing the Konflux snapshot digest, which lags the floating odh-pr-N tag. Artifact collection now follows maas-billing auth_utils (maas-crs, cluster-state, pod logs) and also dumps DataScienceCluster and HTTPRoutes.

Signed-off-by: jland <jland@redhat.com>
Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(e2e): set MaasTenantConfig praxis handoff annotations

Opt into praxis on MaasTenantConfig with payload-processing-type=praxis and payload-processing-status=cleanup-complete, verify they stick, and remove legacy IPP so aigc can deploy ExtProc.

Signed-off-by: jland <jland@redhat.com>
Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(ci): share artifacts volume and collect in e2e step

Mirror the MaaS group-test pattern: shared emptyDir for maas-crs/pod-logs
across e2e, collect-maas-artifacts, must-gather, and git-push.

Signed-off-by: jland <jland@redhat.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Signed-off-by: jland <jland@redhat.com>

* fix(ci): sync group-test pin — unset MAAS_*:latest

Mirror odh-konflux-central: leave MAAS images unset so prow
maas-image-defaults.sh pins odh-pr-1508 (keeps payload-processing-type).
Also use stepTemplate for artifacts volumeMounts (ref steps).

Signed-off-by: jland <jland@redhat.com>
Co-authored-by: Cursor <cursoragent@cursor.com>

* ci: retrigger group-test after maas image pin

Signed-off-by: jland <jland@redhat.com>
Co-authored-by: Cursor <cursoragent@cursor.com>

---------

Signed-off-by: jland <jland@redhat.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
yossiovadia pushed a commit to yossiovadia/ai that referenced this pull request Sep 23, 2026
Signed-off-by: Sébastien Han <seb@redhat.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

blocker This is blocking other work significantly

Projects

Development

Successfully merging this pull request may close these issues.

5 participants