Declaratively deploy and manage VMware Private AI Services (PAIS) resources — from Kubernetes ModelEndpoint deployments and InferenceGatewayRoutes down to Data Sources, Knowledge Bases, MCP Tool integrations, and AI Agents — from a single unified YAML configuration file, reconciled automatically via GitHub Actions.
- Kubernetes CRD API Reference: https://developer.broadcom.com/xapis/vmware-private-ai-service-kubernetes-api/latest/api-docs.html
- REST Data & Agent Plane API Reference: https://developer.broadcom.com/xapis/vmware-private-ai-service-api/latest/
- Architectural Overview
- Authentication & Login Architecture
- Key Capabilities
- Repository Layout
- Prerequisites
- The Unified Configuration File (
config.yaml) - Model Endpoint Operations & Lifecycle Management
- Secrets and Environment Variables
- Local Execution & Dry Runs
- GitHub Secrets & GitOps Pipeline Setup
- End-to-End Walkthrough Example
- Reconciliation & Removal Logic
- Troubleshooting
- CRD & REST API Reference Matrix
VMware Private AI Services (PAIS) operates across two primary planes:
-
Control & Compute Plane (Kubernetes CRDs -
pais.vmware.com/v1alpha1):ModelEndpoint: Serves AI models (LLMs/Embeddings) on vSphere / VKS node pools using engines like vLLM, Infinity, or LlamaCPP. Customizes GPU classes (virtualMachineClassName), storage classes, OCI registry model references (ociRef), replicas, CLI flags, and shared memory sizes.InferenceGatewayRoute: Defines routing rules mapping client requests (byroutingName) to backend ModelEndpoints, cross-namespace PAIS ingress services, or external cloud LLMs (OpenAI, Anthropic).
-
Data & Agent Plane (REST APIs):
- Data Sources & Knowledge Bases: Ingests document stores (Google Drive, S3, SharePoint) and splits/embeds them into vector indexes.
- MCP Tool Integration: Connects external Model Context Protocol (MCP) servers and approves specific tools.
- Agents: RAG-enabled agents combining Knowledge Base search (REX tools) and external MCP tools.
┌─────────────────────────────────────────────────────────────────────────────┐
│ Unified GitOps (config.yaml) │
└──────────────────────────────────────┬──────────────────────────────────────┘
│ Push / Reconcile
▼
┌─────────────────────────────────────────────────────────────────────────────┐
│ GitHub Actions Pipeline (pais-gitops.yml) │
│ │
│ Step 0: k8s_manager.py Step 1-6: setup_pais.py / cleanup_pais.py │
└──────────────────┬──────────────────────────────────┬───────────────────────┘
│ │
│ kubectl apply / CRD Manifests │ OIDC Bearer REST API
▼ ▼
┌──────────────────────────────────────┐ ┌───────────────────────────────────┐
│ Kubernetes Control Plane │ │ PAIS Data Plane REST API │
│ (pais.vmware.com/v1alpha1) │ │ │
│ ├── ModelEndpoints (vLLM/Infinity) │ │ ├── Data Sources & KBs │
│ └── InferenceGatewayRoutes │ │ ├── MCP Servers & Tool Approvals │
└──────────────────────────────────────┘ │ └── RAG Agents & REX Tools │
└───────────────────────────────────┘
A common question when deploying ModelEndpoints is: How does the pipeline log in to deploy Kubernetes CRDs and pull model weights?
The tooling uses a three-part authentication model:
┌─────────────────────────────────────────────────────────────────────────────┐
│ Authentication & Login Flow │
├─────────────────────────────────────────────────────────────────────────────┤
│ 1. Kubernetes Cluster Login: │
│ - VCF CLI Context (`vcf context create` & `vcf context use` via VCF_*) │
│ - OR Kubeconfig (`KUBECONFIG_DATA` base64) / Bearer Token (`KUBE_TOKEN`) │
│ │
│ 2. OCI Model Registry Pull Authentication: │
│ - Automated `kubernetes.io/dockerconfigjson` Secret creation │
│ - Generated from `HARBOR_REGISTRY`, `HARBOR_USERNAME`, `HARBOR_PASSWORD` │
│ - Referenced by `ModelEndpoint.spec.model.pullSecrets` │
│ │
│ 3. PAIS REST API Data Plane Authentication: │
│ - OIDC Resource Owner Password Flow via `PAIS_TOKEN_URL`, `PAIS_USERNAME`│
└─────────────────────────────────────────────────────────────────────────────┘
k8s_manager.py authenticates to the VCF Supervisor Cluster / VKS cluster using one of three supported methods:
-
Method A: VCF CLI Context Login (
vcf context create&vcf context use)
Replaces the deprecatedkubectl vsphere login. Generates and uses the context using your VCF API token or credentials, CCI type, tenant name, and namespace/project structure:vcf context create vcf05paif \ --endpoint=${VCF_ENDPOINT} \ --api-token=${VCF_API_TOKEN} \ --type cci \ --insecure-skip-tls-verify \ --auth-type basic \ --tenant-name all-apps vcf context use vcf05paif:${VCF_NAMESPACE}:${VCF_PROJECT} # e.g., vcf context use vcf05paif:paif-4hqfz:paif-project
-
Method B: ServiceAccount / Bearer Token Login
WhenKUBE_SERVERandKUBE_TOKENare provided, the script sets up a context:kubectl config set-cluster pais-cluster --server=${KUBE_SERVER} --insecure-skip-tls-verify=true kubectl config set-credentials pais-sa --token=${KUBE_TOKEN} kubectl config set-context pais-context --cluster=pais-cluster --user=pais-sa --namespace=${KUBE_NAMESPACE} kubectl config use-context pais-context
-
Method C: GitHub Actions Kubeconfig Secret (
KUBECONFIG_DATA)
The workflow decodesKUBECONFIG_DATAto~/.kube/configbefore executing python scripts.
Model weights are packaged as OCI artifacts (ociRef) in an internal registry like Harbor. For VKS worker nodes to pull these artifacts, a Kubernetes secret of type kubernetes.io/dockerconfigjson must exist in the target namespace.
If your Harbor registry uses self-signed SSL/TLS certificates, the framework supports two options:
-
Option A: Insecure Registry Mode (
insecure: true/HARBOR_INSECURE=true)- Automatically attaches
pais.vmware.com/insecure-registry: "true"labels and annotations to theharbor-registry-secret. - Tells node pools and container engines to skip TLS certificate verification when pulling model weights.
- Automatically attaches
-
Option B: Self-Signed CA Certificate Injection (
ca_cert/HARBOR_CA_CERT)- Injects the PEM-encoded CA certificate string into
ca.crtinside theharbor-registry-secret. - Nodes and containers mount this CA certificate to establish a trusted TLS connection without disabling SSL verification.
- Injects the PEM-encoded CA certificate string into
k8s_manager.py automatically generates and applies this secret manifest from registry configuration:
apiVersion: v1
kind: Secret
metadata:
name: harbor-registry-secret
namespace: default
labels:
pais.vmware.com/insecure-registry: "true"
annotations:
pais.vmware.com/insecure-registry: "true"
type: kubernetes.io/dockerconfigjson
data:
.dockerconfigjson: <base64-encoded-docker-auth>
ca.crt: <base64-encoded-ca-cert-pem> # (Optional, included if HARBOR_CA_CERT is provided)ModelEndpoint resources reference this secret via spec.model.pullSecrets: [{ name: "harbor-registry-secret" }].
For Data Sources, Knowledge Bases, MCP Servers, and Agents, pais_client.py uses OpenID Connect (OIDC) Resource Owner Password Flow against the IdP endpoint configured in PAIS_TOKEN_URL.
- Automated Model Endpoint Deployment: Deploy vLLM or Infinity model servers with dedicated vGPU classes (
nvidia-a10g-gpu-class), vSAN storage, and custom engine parameters (--gpu-memory-utilization,--max-model-len). - Inference Gateway Route Management: Automatically expose models via Gateway routes or connect to external cloud models smoothly.
- Automated Login & Pull Secret Management: Handles cluster authentication (
kubectl vsphere loginor Kubeconfig) and OCI model pull secrets (harbor-registry-secret) automatically. - RAG & Agent Builder: Pair deployed embedding models (
Infinity) with Knowledge Bases and pair deployed LLMs (vLLM) with Agents. - Idempotent Reconciler: Re-running pipelines against unchanged configs performs no duplicate creations.
- Diff-Based Cleanup: Deleting a
model_endpoint,inference_route,data_source,knowledge_base,mcp_server, oragentfromconfig.yamlautomatically deletes the corresponding resource in the cluster/API in safe dependency order. - Artifact Generation: The pipeline outputs standalone
k8s-manifests/pais-resources.yamlmulti-doc manifests for GitOps review or ArgoCD/Flux sync.
pais-gitops/ # Repository Root
├── .github/
│ └── workflows/
│ └── pais-gitops.yml # CI/CD Multi-Tenant Reconcile Workflow
├── tenants/ # Directory-Based Multi-Tenancy (Approach A)
│ ├── finance-org/
│ │ └── config.yaml # Finance VCFA namespace & AI workloads
│ ├── hr-org/
│ │ └── config.yaml # HR VCFA namespace & AI workloads
│ ├── shared-services/
│ │ └── config.yaml # Shared embeddings, gateway routes, & vector KBs
│ └── README.md # Multi-Tenancy Documentation
├── reconcile_all.py # Multi-Tenant Orchestrator (scans tenants/*/config.yaml)
├── k8s_manager.py # K8s Auth, Pull Secrets, CRD Generator & kubectl runner
├── pais_client.py # PAIS REST API Client, Auth & Helpers
├── setup_pais.py # GitOps Apply Script (CRDs + REST API)
├── cleanup_pais.py # GitOps Cleanup Script (Diff-based deletions)
├── config.yaml # Default / Primary Desired State Config
├── config_template.yaml # Reference Template
├── requirements.txt # Python Dependencies (httpx, httpx-auth, pyyaml)
├── .gitignore
└── README.md # Documentation
- A PAIS Kubernetes Cluster: Access to a VKS / vSphere cluster with PAIS installed (
pais.vmware.com/v1alpha1CRDs registered). - PAIS OIDC Credentials: Client ID, Username, Password, Token URL from your PAIS IdP.
- vSphere / Cluster Credentials or Kubeconfig: Credentials to authenticate to the K8s API (
VSPHERE_USER/PASSorKUBECONFIG_DATA). - OCI Model Registry Credentials: Harbor or Docker registry username/password storing model OCI artifacts (
HARBOR_USERNAME/PASSWORD). - Python 3.11+ for local dry runs or manual script execution.
kubernetes:
namespace: "default"
# Cluster Login (Optional: if using vSphere SSO login)
vsphere:
server: "${VSPHERE_SERVER}"
username: "${VSPHERE_USER}"
password: "${VSPHERE_PASSWORD}"
namespace: "${VSPHERE_NAMESPACE}"
# Harbor / OCI Registry Secret Provisioning
registry:
server: "${HARBOR_REGISTRY}"
username: "${HARBOR_USERNAME}"
password: "${HARBOR_PASSWORD}"
secret_name: "harbor-registry-secret"ModelEndpoint defines how an AI model is served on Kubernetes node pools:
model_endpoints:
- name: "llama3-8b-endpoint"
namespace: "default"
type: "Completions" # Completions | Embeddings
engine: "vLLM" # vLLM | Infinity | LlamaCPP
routing_name: "meta-llama/Meta-Llama-3.1-8B-Instruct"
replicas: 1
virtual_machine_class_name: "nvidia-a10g-gpu-class"
storage_class_name: "vsan-default-storage-class"
failure_domain: "zone-1" # vSphere Zone (optional)
model:
oci_ref: "harbor.internal.example.com/pais/models/meta-llama-3.1-8b-instruct:v1"
pull_secrets:
- name: "harbor-registry-secret" # Auto-created by k8s_manager
inference_server_customization:
cli_args:
- "--max-model-len=8192"
- "--gpu-memory-utilization=0.90"
env_vars:
- name: "PAIH_MODEL_ID"
value: "meta-llama/Meta-Llama-3.1-8B-Instruct"
shared_memory_mount_size: "2Gi"InferenceGatewayRoute maps API client requests to local ModelEndpoint services or external providers (e.g., Google Gemini, OpenAI):
# Provision Kubernetes Secret for backend API token authentication
pais:
api_tokens:
- name: "shared-models-token"
api_token: "${SHARED_MODELS_API_TOKEN}"
- name: "gemini-token"
api_token: "${GEMINI_TOKEN}" # Injected via GEMINI_TOKEN secret/env var
inference_routes:
# Route 1: Local Model Endpoint
- name: "route-llama3-8b"
namespace: "default"
type: "Completions" # Completions | Embeddings
engine: "vLLM" # vLLM | Infinity | LlamaCPP | OpenAI
matches:
routing_name: "meta-llama/Meta-Llama-3.1-8B-Instruct"
backend:
http_base_url: "http://llama3-8b-endpoint.default.svc.cluster.local"
model_id: "meta-llama/Meta-Llama-3.1-8B-Instruct"
tls:
verification: "strict" # strict | caOnly | none | mutual
auth:
api_token_ref:
name: "shared-models-token"
# Route 2: External Google Gemini Endpoint
- name: "gemini-flash-latest-service"
type: "Completions"
engine: "GoogleOpenAI" # Or OpenAI
matches:
routing_name: "gemini-flash-latest"
backend:
http_base_url: "https://generativelanguage.googleapis.com/v1beta/openai"
model_id: "models/gemini-flash-latest"
tls:
verification: "strict"
auth:
api_token_ref:
name: "gemini-token" # References gemini-token secret abovedata_sources:
- name: "gdrive-product-docs"
type: "GOOGLE_DRIVE"
origin_url: "https://drive.google.com/drive/u/0/folders/abc123"
credentials: "${GDRIVE_CREDENTIALS}"
test_connection: true
knowledge_bases:
- name: "product-docs-kb"
data_origin_type: "DATA_SOURCES"
index_refresh_policy: { policy_type: "MANUAL" }
data_sources: [ "gdrive-product-docs" ]
index:
name: "product-docs-index"
embeddings_model_endpoint: "BAAI/bge-small-en-v1.5" # References deployed embedding route
text_splitting: "SENTENCE"
chunk_size: 100
chunk_overlap: 0
trigger_indexing: true
wait_for_indexing: truemcp_servers:
- name: "weather-service"
url: "https://weather-mcp-server.example.com"
transport: "STREAMABLE_HTTP"
approve_tools:
- "get_current_weather"
- "get_forecast"agents:
- name: "product-support-agent"
model: "meta-llama/Meta-Llama-3.1-8B-Instruct" # References deployed completions route
instructions: "You are a helpful product support assistant."
knowledge_bases:
- name: "product-docs-kb"
top_n: 5
similarity_cutoff: 0.65
mcp_tools:
- server: "weather-service"
tool_name: "get_current_weather"Common Day-2 operations managed via GitOps:
- Scaling Replicas: Edit
replicas: 1toreplicas: 3inconfig.yamland push. - Upgrading Model Versions: Update
oci_reftag (e.g.:v1->:v2) inconfig.yamland push. - GPU Sizing Tuning: Modify
virtual_machine_class_nameorcli_args(--gpu-memory-utilization=0.95). - Decommissioning a Model: Remove the endpoint and route entries from
config.yaml. Thecleanup_pais.pyscript automatically removes the CRDs from Kubernetes.
Secret interpolation uses ${ENV_VAR_NAME} syntax:
| Environment Variable | Description |
|---|---|
PAIS_BASE_URL |
Base URL of PAIS REST Data Plane API |
PAIS_TOKEN_URL |
OIDC Token URL |
PAIS_CLIENT_ID |
OIDC Client ID |
PAIS_CLIENT_SECRET |
(Optional) OIDC Client Secret (for confidential clients) |
PAIS_SCOPE |
(Optional) OIDC Scope (leave blank if default scope used) |
PAIS_USERNAME |
OIDC Admin / User Username |
PAIS_PASSWORD |
OIDC Password |
PAIS_TOKEN |
(Optional) Pre-created Authentik / OIDC Bearer Token (bypasses password login) |
VCF_ENDPOINT / VSPHERE_SERVER |
VCF Supervisor Cluster FQDN or IP |
VCF_USER / VSPHERE_USER |
VCF / vSphere Username |
VCF_PASSWORD / VSPHERE_PASSWORD |
VCF / vSphere Password |
VCF_NAMESPACE / VSPHERE_NAMESPACE |
VCF / vSphere Namespace |
VCF_PROJECT / PROJECT_NAME |
VCF Automation Project Name |
HARBOR_REGISTRY |
Harbor / OCI Registry FQDN (e.g. harbor.internal.example.com) |
HARBOR_USERNAME |
Harbor Registry Username |
HARBOR_PASSWORD |
Harbor Registry Password |
HARBOR_INSECURE |
Set to true to allow self-signed certificates or HTTP |
HARBOR_CA_CERT |
(Optional) PEM-encoded self-signed CA certificate for Harbor |
GDRIVE_CREDENTIALS |
Service Account JSON string for Google Drive |
S3_CREDENTIALS |
S3 Access Key / Secret JSON string |
GEMINI_TOKEN |
(Optional) API Key / token for Google Gemini endpoint (e.g. Google Generative Language API) |
KUBECONFIG_DATA |
(Optional) Base64-encoded Kubeconfig for direct kubectl apply |
Run local dry runs to preview CRD generation, pull secrets, and REST API execution without making live changes:
# 1. Install dependencies
pip install -r requirements.txt
# 2. Preview additions & updates (Generates k8s-manifests/pais-resources.yaml)
python setup_pais.py --config config.yaml --dry-run --verbose
# 3. Inspect generated Kubernetes manifests
cat k8s-manifests/pais-resources.yaml
# 4. Reconcile all tenant configurations in dry-run mode (Approach A)
python reconcile_all.py --dry-runThe framework scales to multiple VCFA organizations and Kubernetes namespaces using Directory-Based Isolation (Approach A).
tenants/pais-shared-org/config.yaml: Shared services tenant hosting shared embedding models (bge-small), shared gateway routes, shared vector KBs, and MCP tools.tenants/pais-all-apps/config.yaml: Application workload tenant hosting LLM model endpoints (gemma-2-9b-v6), gateway routes, application KBs, and AI agents (product-support-agent).
- PAIS REST API:
PAIS_SHARED_BASE_URL,PAIS_SHARED_TOKEN_URL,PAIS_SHARED_CLIENT_ID,PAIS_SHARED_CLIENT_SECRET,PAIS_SHARED_USERNAME,PAIS_SHARED_PASSWORD - vSphere / VCF CLI:
VCF_SHARED_ENDPOINT,VCF_SHARED_API_TOKEN,VCF_SHARED_USER,VCF_SHARED_PASSWORD,VCF_SHARED_NAMESPACE,VCF_SHARED_PROJECT - Inference Route Tokens:
SHARED_MODELS_API_TOKEN,GEMINI_TOKEN
- PAIS REST API:
PAIS_ALL_APPS_BASE_URL,PAIS_ALL_APPS_TOKEN_URL,PAIS_ALL_APPS_CLIENT_ID,PAIS_ALL_APPS_CLIENT_SECRET,PAIS_ALL_APPS_USERNAME,PAIS_ALL_APPS_PASSWORD - vSphere / VCF CLI:
VCF_ALL_APPS_ENDPOINT,VCF_ALL_APPS_API_TOKEN,VCF_ALL_APPS_USER,VCF_ALL_APPS_PASSWORD,VCF_ALL_APPS_NAMESPACE,VCF_ALL_APPS_PROJECT - Inference Route Token:
ALL_APPS_BACKEND_API_TOKEN
- Folder per Tenant (
tenants/<tenant-name>/config.yaml): Each tenant maintains its ownconfig.yamlintenants/<tenant-name>/. - VCFA Namespace Context Switching:
Each tenant's config specifies its own
vsphereparameters:Whenvsphere: namespace: "${VCF_ALL_APPS_NAMESPACE}" project_name: "${VCF_ALL_APPS_PROJECT}" context_name: "vcf-pais-all-apps"
setup_pais.pyprocesses a tenant,k8s_manager.pyswitches to that tenant's VCF context (vcf context use), preventing cross-tenant resource contamination. - Multi-Tenant Orchestration:
reconcile_all.pyautomatically discovers all tenant config files and executes setup and diff-based cleanup sequentially for every tenant. - CI/CD Pipeline Integration:
The GitHub Actions workflow automatically detects changes across
tenants/**/config.yamland reconciles all active tenant states idempotently.
Add the following Repository Secrets under Settings ▸ Secrets and variables ▸ Actions:
# PAIS Shared Org Secrets
gh secret set PAIS_SHARED_BASE_URL --body "https://pais-shared.example.com"
gh secret set PAIS_SHARED_TOKEN_URL --body "https://idp.example.com/realms/shared/protocol/openid-connect/token"
gh secret set PAIS_SHARED_CLIENT_ID --body "pais-shared-client"
gh secret set VCF_SHARED_ENDPOINT --body "https://vc-shared.domain.local"
gh secret set VCF_SHARED_API_TOKEN --body "shared-org-vcf-api-token"
gh secret set VCF_SHARED_NAMESPACE --body "pais-shared-ns"
gh secret set VCF_SHARED_PROJECT --body "pais-shared-project"
gh secret set SHARED_MODELS_API_TOKEN --body "shared-backend-token-value"
# PAIS All-Apps Org Secrets
gh secret set PAIS_ALL_APPS_BASE_URL --body "https://pais-apps.example.com"
gh secret set PAIS_ALL_APPS_TOKEN_URL --body "https://idp.example.com/realms/apps/protocol/openid-connect/token"
gh secret set PAIS_ALL_APPS_CLIENT_ID --body "pais-apps-client"
gh secret set VCF_ALL_APPS_ENDPOINT --body "https://vc-apps.domain.local"
gh secret set VCF_ALL_APPS_API_TOKEN --body "all-apps-vcf-api-token"
gh secret set VCF_ALL_APPS_NAMESPACE --body "pais-all-apps-ns"
gh secret set VCF_ALL_APPS_PROJECT --body "pais-all-apps-project"
gh secret set ALL_APPS_BACKEND_API_TOKEN --body "apps-backend-token-value"
# External Provider Secrets (for Gemini endpoints)
gh secret set GEMINI_TOKEN --body "your-gemini-api-key"
# OCI Registry Secrets (for pulling model artifacts)
gh secret set HARBOR_REGISTRY --body "harbor.internal.example.com"
gh secret set HARBOR_USERNAME --body "robot$pais-puller"
gh secret set HARBOR_PASSWORD --body "your-harbor-secret"-
Branch Checkout:
git checkout -b feature/model-updates
-
Update
config.yaml: Define new model endpoints, routes, cluster login parameters, data sources, and agents. -
Commit, Push, and Create PR:
git add . git commit -m "Deploy Llama3 8B vLLM endpoint and support agent" git push -u origin feature/model-updates # Open Pull Request targeting `main`
-
GitHub Actions Execution:
- Step 0: Authenticates to cluster via
vcf context create&vcf context use(or Kubeconfig), generatesharbor-registry-secret, buildspais.vmware.com/v1alpha1ModelEndpointandInferenceGatewayRoutemanifests, and applies viakubectl. - Step 1: Provisions S3 / Google Drive Data Sources.
- Step 2: Provisions Knowledge Bases and triggers indexing.
- Step 3-5: Registers MCP Servers and approves tools.
- Step 6: Provisions Agent linked to Knowledge Base REX search tools and MCP tools.
- Step 7: Uploads generated
k8s-manifests/pais-resources.yamlartifact to GitHub Actions summary.
- Step 0: Authenticates to cluster via
- Ordering:
- Apply Phase: K8s Cluster Login ➔ Registry Secret Provisioning ➔ K8s CRDs (ModelEndpoints & GatewayRoutes) ➔ Data Sources ➔ Knowledge Bases & Indexes ➔ MCP Servers ➔ Tool Approvals ➔ Agents.
- Cleanup Phase: Agents ➔ Tool Approval Revocation ➔ Knowledge Base Links ➔ Knowledge Bases ➔ MCP Servers ➔ Data Sources ➔ K8s CRDs (GatewayRoutes & ModelEndpoints).
- CRD Diffing: Objects are matched by
metadata.name. Deleting an item fromconfig.yamltriggers a targetedkubectl deletecommand.
| Issue | Resolution |
|---|---|
vcf context create fails |
Ensure VCF_ENDPOINT (or VSPHERE_SERVER), VCF_USER, and VCF_PASSWORD secrets are correct, and the vcf CLI is installed on the runner. |
ModelEndpoint status ImagePullBackOff |
Verify harbor-registry-secret creation. Ensure HARBOR_REGISTRY, HARBOR_USERNAME, and HARBOR_PASSWORD are valid and the user has pull permissions on the OCI repository. |
ModelEndpoint status Pending |
Check node pool vGPU availability (virtualMachineClassName) or vSphere Zone (failureDomain). |
| Agent return code 404 on Model | Verify that the routing_name in ModelEndpoint matches the matches.routing_name in InferenceGatewayRoute. |
| Capability | Resource Kind / API Path | API Group / Endpoint |
|---|---|---|
| Cluster Login | kubectl vsphere login |
vSphere Supervisor SSO |
| Registry Pull Secret | Secret (dockerconfigjson) |
core/v1 |
| Model Endpoint | ModelEndpoint |
pais.vmware.com/v1alpha1 |
| Gateway Routing | InferenceGatewayRoute |
pais.vmware.com/v1alpha1 |
| Data Source | REST Data Source | /api/v1/control/data-sources |
| Knowledge Base | REST Knowledge Base | /api/v1/control/knowledge-bases |
| Index & Search | REST Index & REX Tool | /api/v1/control/knowledge-bases/{id}/indexes |
| MCP Server | REST MCP Server | /api/v1/control/mcp-servers |
| Agent Builder | REST Agent | /api/v1/compatibility/openai/v1/agents |