From 8c4bb7068f19666268b9eccec0b7fbc415b816f0 Mon Sep 17 00:00:00 2001 From: copilot-review Date: Sat, 5 Sep 2026 04:42:44 +0000 Subject: [PATCH] Reviewed range: 4a10607..a8344d3 a8344d3 Delete zeronym-22aa9851caf68-high-medium directory --- ...hing-binds-the-attested-code-to-zeronym.md | 497 ------------- ...lished-transaction-is-self-timestamping.md | 604 ---------------- ...-debug-mode-which-turns-attestation-off.md | 361 ---------- ...ction-flood-starves-migration-diversion.md | 551 -------------- ...tp-lookup-path-has-no-concurrency-bound.md | 448 ------------ ...cated-fill-silently-destroys-migrations.md | 450 ------------ ...uffer-without-bound-and-oom-the-enclave.md | 605 ---------------- ...cknowledged-migrations-into-silent-loss.md | 643 ----------------- ...ct-logs-migration-value-balance-at-info.md | 458 ------------ ...-a-plaintext-copy-with-nothing-breaking.md | 535 -------------- ...-every-proposed-tip-rate-defence-misses.md | 413 ----------- ...ch-flag-and-ssh-keys-is-not-gated-on-it.md | 471 ------------ ...tion-has-an-unbounded-undetected-window.md | 456 ------------ ...a-defence-the-platform-does-not-rely-on.md | 509 ------------- ...-branch-id-diverts-all-shielded-traffic.md | 419 ----------- ...ect-pcr01-failure-and-accept-pcr2-alone.md | 492 ------------- ...oth-enclaves-have-unrestricted-outbound.md | 405 ----------- ...s-the-running-binary-to-expected-sha256.md | 410 ----------- ...b-chain-unbounded-indexer-response-body.md | 353 --------- ...aino-node-rejections-are-never-verdicts.md | 396 ---------- ...igration-on-single-unverifiable-verdict.md | 390 ---------- ...y-document-and-unobservable-when-absent.md | 591 --------------- ...can-drive-the-fresh-identity-fleet-kill.md | 681 ------------------ ...h-the-shipped-config-makes-the-operator.md | 300 -------- ...tity-permanently-invalidates-every-shim.md | 458 ------------ ...ect-already-owns-is-never-applied-to-it.md | 384 ---------- ...flood-starves-gettransaction-fleet-wide.md | 555 -------------- ...nch-resets-last-advance-masks-stale-tip.md | 362 ---------- .../hub-tip-advance-unbounded-flush-clock.md | 348 --------- ...overshoot-latches-hub-permanently-stale.md | 374 ---------- ...-pre-publication-transaction-disclosure.md | 367 ---------- ...e-is-adversary-selected-regardless-of-k.md | 364 ---------- ...-read-so-every-hub-refusal-is-invisible.md | 446 ------------ ...ery-documented-verification-step-passes.md | 395 ---------- ...s-not-judged-into-permanent-destruction.md | 640 ---------------- ...run-on-any-commit-that-published-a-hash.md | 530 -------------- ...pty-hub-nym-silently-selects-no-privacy.md | 439 ----------- ...so-an-unacked-frame-retransmits-forever.md | 437 ----------- ...-silently-destroys-acknowledged-submits.md | 414 ----------- ...d-concurrency-enclave-memory-exhaustion.md | 509 ------------- ...address-routing-readme-amount-overclaim.md | 381 ---------- ...-and-achieved-batch-size-over-counts-it.md | 484 ------------- 42 files changed, 19325 deletions(-) delete mode 100644 zeronym-22aa9851caf68-high-medium/high/caution-verify-reproduces-from-a-repository-the-operator-nominates-so-nothing-binds-the-attested-code-to-zeronym.md delete mode 100644 zeronym-22aa9851caf68-high-medium/high/core-linkage-survives-in-the-attested-deployment-because-the-wallet-leg-is-unpadded-and-the-published-transaction-is-self-timestamping.md delete mode 100644 zeronym-22aa9851caf68-high-medium/high/deploy-script-defaults-to-debug-mode-which-turns-attestation-off.md delete mode 100644 zeronym-22aa9851caf68-high-medium/high/gettransaction-flood-starves-migration-diversion.md delete mode 100644 zeronym-22aa9851caf68-high-medium/high/hub-http-lookup-path-has-no-concurrency-bound.md delete mode 100644 zeronym-22aa9851caf68-high-medium/high/hub-queue-unauthenticated-fill-silently-destroys-migrations.md delete mode 100644 zeronym-22aa9851caf68-high-medium/high/hub-surb-starved-lookup-replies-grow-the-sdk-pending-buffer-without-bound-and-oom-the-enclave.md delete mode 100644 zeronym-22aa9851caf68-high-medium/high/junk-sendtransaction-flood-consumes-the-shims-whole-mixnet-egress-and-converts-acknowledged-migrations-into-silent-loss.md delete mode 100644 zeronym-22aa9851caf68-high-medium/high/log-verdict-logs-migration-value-balance-at-info.md delete mode 100644 zeronym-22aa9851caf68-high-medium/high/shim-submits-every-migration-to-every-configured-hub-so-an-operator-appends-their-own-and-gets-a-plaintext-copy-with-nothing-breaking.md delete mode 100644 zeronym-22aa9851caf68-high-medium/medium/a-constant-tip-offset-is-a-tunable-expiry-keyed-admission-filter-that-every-proposed-tip-rate-defence-misses.md delete mode 100644 zeronym-22aa9851caf68-high-medium/medium/attested-enclave-console-is-reopenable-from-the-parent-because-debug-mode-is-a-launch-flag-and-ssh-keys-is-not-gated-on-it.md delete mode 100644 zeronym-22aa9851caf68-high-medium/medium/attested-tls-binding-is-verified-once-by-hand-if-ever-so-operator-certificate-substitution-has-an-unbounded-undetected-window.md delete mode 100644 zeronym-22aa9851caf68-high-medium/medium/auditor-recipe-omits-the-two-checks-that-decide-where-plaintext-goes-and-names-a-defence-the-platform-does-not-rely-on.md delete mode 100644 zeronym-22aa9851caf68-high-medium/medium/classifier-unknown-consensus-branch-id-diverts-all-shielded-traffic.md delete mode 100644 zeronym-22aa9851caf68-high-medium/medium/deploy-script-tells-operators-to-expect-pcr01-failure-and-accept-pcr2-alone.md delete mode 100644 zeronym-22aa9851caf68-high-medium/medium/enclave-egress-allowlist-is-discarded-by-the-platform-and-both-enclaves-have-unrestricted-outbound.md delete mode 100644 zeronym-22aa9851caf68-high-medium/medium/hub-caution-readme-says-the-attestation-binds-the-running-binary-to-expected-sha256.md delete mode 100644 zeronym-22aa9851caf68-high-medium/medium/hub-chain-unbounded-indexer-response-body.md delete mode 100644 zeronym-22aa9851caf68-high-medium/medium/hub-chain-zaino-node-rejections-are-never-verdicts.md delete mode 100644 zeronym-22aa9851caf68-high-medium/medium/hub-flush-destroys-migration-on-single-unverifiable-verdict.md delete mode 100644 zeronym-22aa9851caf68-high-medium/medium/hub-indexer-tls-is-optional-in-code-required-in-every-document-and-unobservable-when-absent.md delete mode 100644 zeronym-22aa9851caf68-high-medium/medium/hub-liveness-probe-reads-its-own-send-backlog-as-gateway-silence-so-any-stranger-can-drive-the-fresh-identity-fleet-kill.md delete mode 100644 zeronym-22aa9851caf68-high-medium/medium/hub-lookup-fall-through-hands-every-wallets-txid-to-whichever-indexer-the-hub-is-pointed-at-which-the-shipped-config-makes-the-operator.md delete mode 100644 zeronym-22aa9851caf68-high-medium/medium/hub-nym-driver-automatic-fresh-identity-permanently-invalidates-every-shim.md delete mode 100644 zeronym-22aa9851caf68-high-medium/medium/hub-nym-identity-has-no-trust-anchor-and-the-one-the-project-already-owns-is-never-applied-to-it.md delete mode 100644 zeronym-22aa9851caf68-high-medium/medium/hub-nym-lookup-flood-starves-gettransaction-fleet-wide.md delete mode 100644 zeronym-22aa9851caf68-high-medium/medium/hub-reorg-branch-resets-last-advance-masks-stale-tip.md delete mode 100644 zeronym-22aa9851caf68-high-medium/medium/hub-tip-advance-unbounded-flush-clock.md delete mode 100644 zeronym-22aa9851caf68-high-medium/medium/hub-tip-overshoot-latches-hub-permanently-stale.md delete mode 100644 zeronym-22aa9851caf68-high-medium/medium/hub-unauthenticated-pre-publication-transaction-disclosure.md delete mode 100644 zeronym-22aa9851caf68-high-medium/medium/indexer-chooses-which-batch-members-reach-the-chain-so-the-on-chain-batch-size-is-adversary-selected-regardless-of-k.md delete mode 100644 zeronym-22aa9851caf68-high-medium/medium/nym-submit-acks-are-never-read-so-every-hub-refusal-is-invisible.md delete mode 100644 zeronym-22aa9851caf68-high-medium/medium/operator-controlled-dns-permits-a-layer-4-relay-that-every-documented-verification-step-passes.md delete mode 100644 zeronym-22aa9851caf68-high-medium/medium/publish-verdict-strings-are-zcashds-vocabulary-only-so-a-zebrad-backed-indexer-turns-not-judged-into-permanent-destruction.md delete mode 100644 zeronym-22aa9851caf68-high-medium/medium/reproduce-gate-has-never-run-on-any-commit-that-published-a-hash.md delete mode 100644 zeronym-22aa9851caf68-high-medium/medium/shim-config-no-fail-closed-mode-empty-hub-nym-silently-selects-no-privacy.md delete mode 100644 zeronym-22aa9851caf68-high-medium/medium/shim-mixnet-client-has-neither-retransmission-bound-the-hub-has-so-an-unacked-frame-retransmits-forever.md delete mode 100644 zeronym-22aa9851caf68-high-medium/medium/shim-nym-driver-every-teardown-path-silently-destroys-acknowledged-submits.md delete mode 100644 zeronym-22aa9851caf68-high-medium/medium/shim-proxy-unbounded-inbound-concurrency-enclave-memory-exhaustion.md delete mode 100644 zeronym-22aa9851caf68-high-medium/medium/shim-route-for-no-tainted-address-routing-readme-amount-overclaim.md delete mode 100644 zeronym-22aa9851caf68-high-medium/medium/widening-the-flush-window-cannot-raise-the-delivered-anonymity-set-for-todays-traffic-and-achieved-batch-size-over-counts-it.md diff --git a/zeronym-22aa9851caf68-high-medium/high/caution-verify-reproduces-from-a-repository-the-operator-nominates-so-nothing-binds-the-attested-code-to-zeronym.md b/zeronym-22aa9851caf68-high-medium/high/caution-verify-reproduces-from-a-repository-the-operator-nominates-so-nothing-binds-the-attested-code-to-zeronym.md deleted file mode 100644 index 8f68d82c..00000000 --- a/zeronym-22aa9851caf68-high-medium/high/caution-verify-reproduces-from-a-repository-the-operator-nominates-so-nothing-binds-the-attested-code-to-zeronym.md +++ /dev/null @@ -1,497 +0,0 @@ -# `caution verify` reproduces the enclave from a repository the operator nominates, so a clean `Attestation verification PASSED` proves the enclave runs *the operator's* code — nothing in the tree, the tooling or any document joins that code to zeronym - -**Severity**: High -**Validation Status**: Confirmed -**Location**: `audit-target/zeronym/shim/deploy/caution/caution.hcl.tmpl:8-18` (the "WHAT IT PROVES" claim) and `:20-27` (the `build { }` block carrying `__APP_SOURCE__`); `audit-target/zeronym/shim/deploy/caution/assemble-caution.sh:64`, `:99`, `:427-443`, `:589-604`; `audit-target/zeronym/shim/deploy/assemble.sh:60-120` (the full Rust source is what gets published); `audit-target/zeronym/deploy.sh:130-136` (`DEBUG=0` requires `APP_SOURCE`) and `:197-222` (the automated publish step); `audit-target/zeronym/README.md:71` (the auditor recipe); `audit-target/zeronym/shim/deploy/caution/README.md:124-138`; `audit-target/zeronym/shim/deploy/caution/OPERATORS.md:150-200` ("Verify"); `audit-target/zeronym/hub/deploy/caution/OPERATORS.md:97-124`; identical structure in `audit-target/zeronym/hub/deploy/caution/assemble-caution.sh` and `hub/deploy/assemble.sh` -**Found by agent:** Global, focus area G30/G32/G17 — the indexer operator's full capability sweep after the `unit.env`/certfp reversals -**In scope of audit?** Yes — `*/deploy/**` is in scope because "the reproducible-build and attestation chain **is** the trust model here", and markdown claims are in scope as security claims. - -## Description - -After the two reversals recorded in coordinator open item 6q, the attestation -chain for this product is much stronger than the audit first believed: - -- `unit.env` — every `ZIS_*`/`ZIH_*` value **and their absence** — is measured - into PCR0/PCR1, so configuration cannot be changed without changing a - measurement; -- `caution verify` binds the attested enclave to the **leaf certificate** of the - very TLS connection that carried the attestation, so an operator terminating - wallet TLS in front of the enclave fails verification. - -Both of those bind the enclave to **the tree the operator published**. Nothing -binds that tree to zeronym. - -`caution verify` reads `app_source.urls` and `app_source.commit` out of the -attested manifest, clones that repository at that commit, reads the -`caution.hcl` **from that clone**, rebuilds the whole EIF **from that clone**, -and compares the resulting PCR0/1/2 against the live attestation -(`src/cli/src/lib.rs:6432-6467`, `:6470-6473`). The URL is whatever the operator -passed to `assemble-caution.sh --app-source`. There is no allow-list, no -signature, no expected hash, and no comparison against -`github.com/ShieldedLabs/zero`. There cannot be: the published repository is a -**derived, per-deployment tree** produced by `assemble-caution.sh`, so it is -never Shielded Labs' repository by construction — `OPERATORS.md:77-78` instructs -the operator to *"Create an empty **public** git repository first"*. - -That published tree contains the complete compiled input: `zeronym/shim/**` -(including `src/classify.rs`, `src/intercept.rs`, `src/proxy.rs`), -`zebra/zebra-chain`, `zaino/packages/zaino-proto`, -`zeronym/vendor/nym-upgrade-mode-check`, and the `Containerfile` that builds -them (`shim/deploy/assemble.sh:66`, `:93-96`, `:106-108`, `:117-118`, -`:144-179`). An operator who edits any of it and publishes the edit gets an -enclave that reproduces **exactly**, because the thing being reproduced is their -edit. - -And the one artefact that could have closed the loop is absent. The attested -manifest has a `binary` field (`src/enclave-builder/src/manifest.rs:22-23`), but -the Containerfile deploy path passes `None` for it -(`src/api/src/builder.rs:925-940` — the fourth positional argument to -`EnclaveManifest::new`, whose signature is at -`src/enclave-builder/src/manifest.rs:84-91`), so **no hash of the running binary -appears anywhere in the attestation**. `EXPECTED_SHA256` is a zeronym-only file -that the Caution platform never reads — the string does not occur anywhere in -the platform source. With no binary hash in the attestation and no tie from the -app-source tree to upstream, **the chain has exactly one unbound link, and it is -the link that decides what code handles users' transactions.** - -The manifest states the opposite as the reason the deploy exists -(`shim/deploy/caution/caution.hcl.tmpl:8-15`): - -``` -# WHAT IT PROVES, which is the entire reason to deploy it. The Zeronym trust -# model asks an auditor to rebuild the shim from source, reach the published -# hash, and check that hash against the one bound into the enclave attestation. -# Reproducibility alone proves only that source and binary agree; attestation -# alone proves only that SOME binary runs in a genuine enclave. Together they -# say: the code you read is the code that is running, and the operator cannot -# see the traffic. -``` - -and `shim/deploy/caution/README.md:129-131` says verify *"rebuilds the image from -the published `app_sources` repo, and compares measurements. It is what turns -'they say this is the code' into something checkable."* It turns it into -something checkable **against the operator's own tree**. "The code you read is -the code that is running" is true only if the code an auditor reads is the code -in the operator's `app_sources` repository — and every document points them at -`github.com/ShieldedLabs/zero` instead (`OPERATORS.md:60`, `PROVENANCE`, -`README.md:71`). - -## Attack Scenario and Steps - -Attacker: the indexer operator — adversary #1, and the party the product exists -to defend against. - -1. The operator clones `github.com/ShieldedLabs/zero` at the commit everyone is - auditing and makes one edit. Any of these suffices, and none changes the - product's observable behaviour for a wallet: - - in `shim/src/intercept.rs`, on the divert path, additionally POST the - transaction bytes to an address the operator owns. The enclave has - **unrestricted outbound** — the `egress { }` allow-list is reduced by the - platform to the boolean `!egress.is_empty()` and the parent applies - `iptables -A FORWARD -i enclave0 -o $DEFAULT_IFACE -j ACCEPT` - (already filed as - `enclave-egress-allowlist-is-discarded-by-the-platform-and-both-enclaves-have-unrestricted-outbound.md`), - so there is no network control to defeat; - - or, more cheaply and with no new traffic at all, weaken - `classify::is_orchard_touching` for a chosen shape so those transactions - take `Route::PassThrough` and arrive at the operator's own indexer in the - clear; - - or simply log what `log_verdict` is careful not to log, to a destination of - their choosing. -2. They run `assemble-caution.sh --app-source https://github.com//zero-indexer-shim-deploy ...` - exactly as documented, push to Caution, and push the assembled tree to their - own public repository — which `deploy.sh:197-222` does automatically and which - `deploy.sh:134` makes **mandatory** for any non-debug deploy. -3. An auditor performs the recipe at `README.md:71` and the "Verify" section at - `OPERATORS.md:150-160`: - - `caution verify --attestation-url https:///attestation` -> - clones the operator's repository, rebuilds, and prints - `Base Nitro attestation and expected PCR0/1/2 verified`, - `TLS certificate binding verified`, `Attestation verification PASSED`. - **All of it is true.** - - `sh zeronym/shim/deploy/reproduce.sh` in a checkout of - `github.com/ShieldedLabs/zero` -> prints `REPRODUCES`, because the upstream - source really does reproduce the upstream `EXPECTED_SHA256`. **Also true, - and about a different artefact entirely.** -4. The auditor reports the endpoint as verified. Neither command ever compared - the operator's tree to Shielded Labs'. - -A second, quieter use of the same gap: because `debug.ssh_keys`, -`network.ingress` CIDRs and `resources` reach only terraform and never the EIF, -the operator can publish an `app_sources` tree whose `caution.hcl` shows -`ssh_keys = []` while the deployed manifest carried a key, and **every PCR still -reproduces**. The published manifest is therefore not evidence about any -unmeasured field. - -**Attack Requirements and Assumptions:** - -- **Access needed:** none beyond being the operator. No platform break, no - on-path position, no cryptographic weakness, no software vulnerability. -- **Cost:** one source edit and one `git push`. The publish step is already - automated by `deploy.sh`. -- **What makes this realistic:** publishing an `app_sources` URL is *required* - for an attested deploy (`deploy.sh:134`), so the operator is not doing anything - unusual; the repository they publish is expected to be theirs and new; and its - contents are a large derived tree that nobody is instructed to diff or read. -- **What limits it, stated plainly and prominently:** - - **The malicious code is public.** It sits in a public repository the operator - advertised, permanently, under a commit hash bound into an AWS-signed - attestation. For a *named* operator that is durable, attributable evidence, - and a real deterrent. - - **A cheap mechanical check does exist** — re-run `assemble-caution.sh` from - the claimed upstream commit with the operator's own parameters and `diff -r` - against the published tree. `assemble.sh` builds from `git archive HEAD`, - which stamps deterministic mtimes (`assemble.sh:13-15`), so the comparison is - expected to be exact modulo the values substituted into `caution.hcl`. - **No zeronym document describes it, and no tool performs it.** - - The `caution verify` output *does* print `App source: commit: ` - (`src/cli/src/lib.rs:7100-7104`), so an attentive auditor sees the URL — but - seeing a URL is not being told what it must contain, and no canonical value - can be published because the tree is per-deployment. - -## Impact on Users - -`README.md:71` is the entire substitute this product offers for trusting an -indexer operator: *"**Auditors** verify an endpoint without trusting its -operator."* A user is entitled to read that as "someone competent can establish -that this endpoint runs the reviewed zero-indexer." Nobody performing the -documented steps establishes that. What they establish is: *a genuine AWS Nitro -enclave, built by Caution from a tree this operator published, is answering at -this hostname, and the certificate terminating my TLS session is the one that -enclave holds.* - -The gap is not narrow. Arbitrary code inside the enclave is the strongest -position in the whole system: it sees every wallet's queries and every -Orchard-touching transaction in plaintext at the moment the wallet sends it, it -has unrestricted outbound to ship them anywhere, and it can decline to divert at -all. Every invariant in `THREATMODEL.md` §6.2 (P1-P5) and §6.3 (C1, C2) is a -property of the shipped source code and therefore falls with it. - -The failure is silent on the user's side: a wallet sees a valid certificate, -correct gRPC behaviour, and a normal txid whichever tree is running. - -It is also the *meta*-failure of this audit's whole attestation area. Every other -finding about the deployment — configuration measurement, the certfp binding, the -classifier's soundness, the fail-closed discipline — is a statement about the -code in `github.com/ShieldedLabs/zero`. If nothing binds the running enclave to -that code, none of those statements is a property of any particular endpoint. - -## Technical Details / Code Analysis - -### 1. What `caution verify` reproduces from - -Caution platform, `src/cli/src/lib.rs:6432-6467` (read from the public source at -`https://codeberg.org/caution/platform`, whose location is named by -`shim/deploy/caution/OPERATORS.md:44-46`): - -```rust - let app_source_dir = if let Some(source) = local_source { - Some(source.path.clone()) - } else if let Some(ref manifest) = external_manifest { - let app_source = manifest.app_source.as_ref().ok_or_else(|| { - anyhow::anyhow!( - "Manifest does not contain app_source - cannot reproduce without source URL" - ) - })?; - let archive_urls: Vec = app_source - .urls - .iter() - .filter_map(|url| self.git_url_to_archive_urls(url, &app_source.commit).ok()) - .flatten() - .collect(); - ... - Some( - self.download_and_extract_app_source_with_git_fallback(...).await?, - ) -``` - -and `:6470-6473`: - -```rust - let measured_config = if let Some(ref app_dir) = app_source_dir { - Some(self.read_config_from_dir(app_dir)?) -``` - -Both the *source* and the *manifest that describes how to build it* come from -`app_source.urls`. The only validation applied to that URL anywhere in the verify -path is a `git ls-remote` reachability preflight -(`preflight_app_source_ref`, `src/cli/src/lib.rs:7898-7960`). - -Note the follow-on: because `expected.domain` for the TLS certificate-fingerprint -check is derived from that same reproduced `caution.hcl` -(`tls_expectation_from_config`, `src/cli/src/lib.rs:286-312`), the strong new -binding established in open item 6q is *also* rooted in the operator's tree. It -proves the enclave serves the domain the operator's own config names. - -### 2. There is no binary hash in the attestation to fall back on - -`src/enclave-builder/src/manifest.rs:84-91` defines the constructor: - -```rust - pub fn new( - app_source: Option, - enclave_source: EnclaveSource, - framework_source: FrameworkSource, - binary: Option, - run_command: Option, - metadata: Option, - ) -> Self { -``` - -and the Containerfile deploy path calls it at `src/api/src/builder.rs:925-940`: - -```rust - let mut manifest = enclave_builder::EnclaveManifest::new( - app_source, - enclave_builder::EnclaveSource::GitArchive { ... }, - enclave_builder::FrameworkSource::GitArchive { ... }, - None, // <- `binary` - request.run_command.clone(), - None, - ); -``` - -So the `binary` field an auditor might hope to compare against `EXPECTED_SHA256` -is `None` for every zeronym deploy, and `EXPECTED_SHA256` itself is never read by -any Caution code. (For completeness, and in Caution's favour: `enclave_source` -and `framework_source` **are** pinned by the platform to -`git.distrust.co/public/enclaveos` and Caution's own `FRAMEWORK_SOURCE`, so the -platform's half of the image is not operator-nominated. Only the application half -is.) - -### 3. What the published tree contains - -`shim/deploy/assemble.sh` (invoked by `assemble-caution.sh`) copies, from -`git archive HEAD`: - -```sh -git -C "$ZERO_ROOT" archive HEAD -o "$STAGE/shim.tar" zeronym/shim # :66 -git -C "$ZERO_ROOT" archive HEAD -o "$STAGE/zebra.tar" \ # :93-96 - zebra/Cargo.toml zebra/zebra-chain zebra/zebra-test -git -C "$ZERO_ROOT" archive HEAD -o "$STAGE/zaino.tar" \ # :106-108 - zaino/Cargo.toml zaino/packages/zaino-proto -git -C "$ZERO_ROOT" archive HEAD -o "$STAGE/vendor.tar" \ # :117-118 - zeronym/vendor/nym-upgrade-mode-check -``` - -i.e. the entire compiled input, as source. `assemble-caution.sh:427-443` then -records the publication URL in the manifest's `build` block: - -```sh -if [ -n "$APP_SOURCE" ]; then - cat > "$APP_SRC_FILE" </attestation` | live PCR0/1/2 <-> the operator's published tree; attested `certfp` <-> the leaf of this TLS session; attested `domain` <-> the domain in the operator's published `caution.hcl` | the operator's tree <-> zeronym | -| `sh zeronym/shim/deploy/reproduce.sh` | upstream source <-> upstream `EXPECTED_SHA256`, on this machine | either side <-> anything the enclave contains | - -The two chains never meet. - -## Recommendations - -1. **Add the missing step to `README.md:71` and to both "Verify" sections:** - *"Re-run `assemble-caution.sh` from `github.com/ShieldedLabs/zero` at the - commit the endpoint's `PROVENANCE` names, with the parameters the published - `caution.hcl` shows, and `diff -r` the result against the `app_sources` - repository `caution verify` cloned. A `caution verify` PASS without this step - proves the enclave runs the operator's published code, not zero-indexer."* - This is the whole fix and it costs one paragraph. -2. **Ship the diff as a script.** `zeronym/shim/deploy/verify-app-source.sh - ` — clone, re-assemble from upstream, `diff -r`, exit - non-zero on any difference outside the substituted `caution.hcl` values. The - determinism `assemble.sh:13-15` already guarantees is what makes this - mechanical. -3. **Correct `caution.hcl.tmpl:8-15` and `shim/deploy/caution/README.md:129-131`.** - Replace *"the code you read is the code that is running"* and *"turns 'they say - this is the code' into something checkable"* with what the platform delivers: - *"the code **in the repository this manifest names** is the code that is - running; confirming that repository is zero-indexer is a separate step, and it - is step N of the Verify section."* The same sentence appears in the hub's - manifest and should be corrected there too. -4. **State that the published `caution.hcl` is not evidence about unmeasured - fields.** `debug.ssh_keys`, `network.ingress` CIDRs and `resources` reach only - terraform, so a published tree can differ from the deployed manifest in those - fields with every PCR still reproducing. -5. **Ask Caution for a manifest `binary` value on the Containerfile path**, or for - a documented way to bind `EXPECTED_SHA256` to a measurement. That would let an - auditor check the running binary against a value Shielded Labs publishes, and - would make recommendation 1 a fallback rather than the only control. -6. Until 1-3 exist, `README.md:71` should not say *"without trusting its - operator"*. - -Cross-references — this issue is the system-level statement the following stop -short of, and it should be reported alongside them rather than merged into any -one of them: -`auditor-recipe-omits-the-two-checks-that-decide-where-plaintext-goes-and-names-a-defence-the-platform-does-not-rely-on.md` -(the same recipe, for *configuration* rather than *code*); -`hub-caution-readme-says-the-attestation-binds-the-running-binary-to-expected-sha256.md` -and `reproduce-never-builds-the-runtime-stage-that-the-enclave-and-pcr0-are-built-from.md` -(why the hash half cannot work); -`enclave-egress-allowlist-is-discarded-by-the-platform-and-both-enclaves-have-unrestricted-outbound.md` -(why injected code has somewhere to send the plaintext); -`assemble-git-archive-honours-gitattributes-so-the-build-context-is-not-the-committed-tree.md` -(the same outcome reached through *upstream* rather than through the operator). - -## Validation Information - -**Verdict: CONFIRMED. Severity: High** (filed High with an explicit Medium -counter-argument; **the calibration is resolved in favour of High** — reasoning -below). - -**Every mechanical claim re-verified from primary sources.** The Caution platform -clone (`codeberg.org/caution/platform`) used by the filing agent was still on -disk and was re-read directly: - -- `src/cli/src/lib.rs:6432-6467` — `caution verify` builds `app_source_dir` from - `manifest.app_source.urls` @ `commit`, with only a `git ls-remote` preflight - (`preflight_app_source_ref`, `:7898-7960`). No allow-list, no signature. - Confirmed verbatim. -- `:6470-6473` — `measured_config` is read *from that clone* - (`read_config_from_dir`), so the `caution.hcl` driving the rebuild is the - operator's. Confirmed. -- `tls_expectation_from_config` (`:286-312`) consumes that same reproduced - config, so the certfp binding established by reversal 6q is itself rooted in - the operator's tree. Confirmed — this is a genuine second-order consequence and - it strengthens the finding. -- `src/enclave-builder/src/manifest.rs:84-91` — the constructor's fourth - parameter is `binary: Option`; `src/api/src/builder.rs:925-940` passes - `None` positionally. **Confirmed by reading the signature, not by counting - arguments in prose.** No binary hash enters the attestation. -- `grep -r EXPECTED_SHA256` over the whole platform source returns **nothing**: - the file exists only in zeronym (`shim/deploy/EXPECTED_SHA256`) and no Caution - code path reads it. This closes the "surely something compares the hash" - objection. -- `src/cli/src/lib.rs:7100-7104` — verify does print - `App source: commit: `. The disclosure is real; what is missing is - anything to compare it against. - -**Verified in `audit-target/zeronym/`:** - -- `deploy.sh:130-136` — `DEBUG=0` **dies** without `APP_SOURCE`, so an attested - deploy always nominates a repository; `:197-222` — `deploy.sh` pushes the - assembled tree to that repository automatically and tags it. The operator is - not doing anything unusual. -- `assemble-caution.sh:427-443` — the `app_sources = [...]` block, with the - comment stating outright that the published root *"must be THIS directory, not - the zero monorepo"*. -- `assemble-caution.sh:589-604` — `PROVENANCE`, a plain-text claim in a - repository the operator controls, pointing the reader at a different repository. -- `caution.hcl.tmpl:8-15` and `shim/deploy/caution/README.md:129-131` — the two - overclaims, quoted accurately. -- `shim/deploy/caution/OPERATORS.md:150-200` and `README.md:71` — the complete - documented Verify procedure is `caution verify` + `reproduce.sh`. A repository-wide - grep for `app_sources`/`app-source`/`APP_SOURCE` across all `*.md` returns - seven hits, **none** of which asks anyone to compare the nominated tree against - upstream. The "a diligent auditor could just diff it" objection therefore fails - on the facts: nothing tells them to, and no canonical value exists to diff - against because the tree is per-deployment by construction. - -**One strengthening added during validation** (recommendation 4 and the second -paragraph of the attack): because `debug.ssh_keys`, `network.ingress` CIDRs and -`resources` reach only terraform and never the EIF, an operator can publish a -tree whose `caution.hcl` differs from the deployed manifest in exactly those -fields and **every PCR still reproduces**. So the published manifest is not -evidence about any unmeasured field — which matters directly to the step-1 -precondition of the audit's headline finding, where `debug { enabled = false; -ssh_keys = [...] }` is one of the two routes to the parent host. - -**Why the Medium argument loses.** The filing agent asked the validator to weigh -three points; each was considered and rejected: - -1. *"The malicious source is permanently public and attributable."* True, and it - is a deterrent — but a deterrent is not a control, and it only binds an - operator with a reputation at stake. The audit's threat model explicitly - includes hostile operators and notes that anyone can run the published image - and produce a valid attestation. It also does nothing for the *user*, who has - no way to act on evidence that only becomes legible after a forensic - comparison nobody is asked to perform. -2. *"The fix is documentation, not code."* Cheapness of the fix is an argument for - fixing it, not for grading it low. Coordinator item 7a states the governing - distinction: **measurement discloses a value; it never detects a change.** - Applied here it is worse than for the configuration siblings — `ZIS_HUB_NYM` - at least has a canonical expected value that could be published and compared, - whereas the `app_sources` tree is per-deployment, so the disclosure at - `verify`'s `App source:` line terminates in a value with **no expected - counterpart at all**. That is why the sibling - `auditor-recipe-omits-...` sits at Medium and this does not. -3. *"It is really a platform property."* The platform's behaviour is reasonable - for a general-purpose service. What is a zeronym defect is that three of its - own documents assert the stronger property (`caution.hcl.tmpl:8-15`, - `shim/deploy/caution/README.md:129-131`, `README.md:71`) and its documented - procedure omits the only step that would deliver it. Markdown claims are in - scope as security claims, and under ICTM a documented property users are told - they get but do not get is itself the bug. - -**Why High rather than Critical:** it requires the endpoint's own operator to be -deliberately malicious, the malicious code is public and attributable, no funds -can be stolen, and an auditor who *does* read the nominated tree finds it -immediately. - -**Why High rather than Medium:** the capability obtained is total (arbitrary code -in the enclave nullifies every code-level invariant in `THREATMODEL.md` §6.2 and -§6.3 at once, with unrestricted outbound to exfiltrate), the cost is one commit -on a path the deploy script already automates, **every** documented check passes -and passes correctly, and — unlike the Medium siblings — there is no mechanical -check available at all until recommendation 2 is built. It also voids the single -sentence (`README.md:71`) that the product offers as its substitute for trusting -an operator. - -**False-positive checks applied.** - -- *§6 Intentional design?* The platform's reproduce-from-nominated-source model - is intentional. The finding is not that model; it is the three zeronym - documents that claim a stronger binding and the procedure that omits the - bridging step. That is a defect, not a design choice. -- *§8 Requires prior compromise?* No. The operator has this by construction. -- *§4 Test/debug only?* No — the affected path is the **attested** deploy path; - `DEBUG=0` is precisely the configuration that requires `--app-source`. -- *§1 Assumption an attacker cannot violate?* The assumption "the app_sources - repository contains zeronym" is violated by a `git commit`. - -**Double-counting guard for the report.** This must be presented as the *meta* -finding of the attestation family — the one that says why the other checks are -only as good as the tree they are run against — and **not** as another -independent count of "the operator obtains plaintext". The operator's cheaper -plaintext routes (`shim-submits-every-migration-to-every-configured-hub-...md`, -confirmed High; `log-verdict-logs-migration-value-balance-at-info.md`, confirmed -High) are separate mechanisms; the report should not sum three Highs into three -distinct losses of the same secret. - -DO NOT DEVIATE FROM THIS TEMPLATE (`docs/SECURITY-ISSUE-TEMPLATE.md`). -LEAVE THIS MESSAGE IN PLACE SO THAT ALL AGENTS KNOW NOT TO DEVIATE FROM THIS TEMPLATE (`docs/SECURITY-ISSUE-TEMPLATE.md`). diff --git a/zeronym-22aa9851caf68-high-medium/high/core-linkage-survives-in-the-attested-deployment-because-the-wallet-leg-is-unpadded-and-the-published-transaction-is-self-timestamping.md b/zeronym-22aa9851caf68-high-medium/high/core-linkage-survives-in-the-attested-deployment-because-the-wallet-leg-is-unpadded-and-the-published-transaction-is-self-timestamping.md deleted file mode 100644 index 583cc6b5..00000000 --- a/zeronym-22aa9851caf68-high-medium/high/core-linkage-survives-in-the-attested-deployment-because-the-wallet-leg-is-unpadded-and-the-published-transaction-is-self-timestamping.md +++ /dev/null @@ -1,604 +0,0 @@ -# The IP -> transaction -> amount join survives the product in the *attested* deployment: the unpadded wallet leg gives the operator `|tx|`, and the published transaction timestamps itself, so the delivered anonymity set is ~1 whatever the batch size - -**Severity**: High -**Validation Status**: Confirmed -**Location**: Whole-system. Principal loci: -`audit-target/zeronym/shim/deploy/caution/caution.hcl.tmpl:79-127` (in-enclave TLS termination, so the wallet's ciphertext is relayed across the parent host unpadded), `:186-201` (the `debug { }` block whose `ssh_keys` list zeronym asserts is inert); -`audit-target/zeronym/shim/src/intercept.rs:110-131` (divert returns before `pool.get()`, so the operator sees the *absence*), `:158-166` (the code naming `|tx|` as the secret and paying an availability cost to protect it on the next hop); -`audit-target/zeronym/shim/src/proxy.rs:744-783` (`route_for` — every sync method is `PassThrough`); -`audit-target/zeronym/hub/src/batcher.rs:40` (`FLUSH_INTERVAL_BLOCKS = 20`, a public deterministic schedule), `:55` (`MIN_WALLET_EXPIRY = 40`); -`audit-target/zeronym/hub/src/queue.rs:363-393` (`next_flush_height` / `survives_next_flush`), `:258-264` (shuffle); -`audit-target/zeronym/hub/src/chain.rs:199-210` (`broadcast_batch` — one simultaneous burst); -`audit-target/zeronym/shim/deploy/caution/assemble-caution.sh:114-123` (the "SSH is closed under attestation" note); -`audit-target/zeronym/shim/deploy/caution/OPERATORS.md:59` (the operator owns the wallet-facing DNS name), `:64-69` (managed vs BYOC); -`audit-target/zeronym/deploy.sh:156` (`caution apps create`), `:162-176` (the DNS record the operator writes); -`audit-target/zeronym/README.md:27`, `:33`, `:34`, `:54`, `:69` (the claims); `audit-target/zeronym/OPEN-QUESTIONS.md:109` (the in-tree, undisclosed-to-users statement of the size channel) -**Found by agent:** Global, focus area G29 — "link wallet IP -> on-chain Orchard-touching transaction -> balance, using only what the operator sees" -**In scope of audit?** Yes - -## Description - -This is the composed form of the attack the product exists to prevent. It is -filed separately from its components because the composition supports a -conclusion none of them supports alone: - -> **`README.md:34`'s residual — "the modal batch is zero or one... The lever is -> adoption, not code" — names a remedy that does not restore the protection.** -> Two channels re-identify an individual submitter *inside a batch of any size*, -> and both are readable by the primary adversary in a correctly attested -> deployment with `debug { enabled = false }` and a clean `caution verify PASSED`. - -The two channels: - -**(1) The wallet->shim leg is unpadded and carries `|tx|`.** Under the shipped -Caution manifest the wallet's TLS terminates on an in-enclave Caddy -(`caution.hcl.tmpl:79-127`). The enclave has no NIC: the parent host relays port -443 in over vsock with `socat`, so the wallet's ciphertext crosses the parent -byte for byte. TLS record length fields are cleartext in every TLS version, no -mainstream stack pads, and nothing in zeronym pads, chunks or normalises the -wallet-facing request. Anyone holding that socket reads the serialized length of -the transaction. The project pays a real availability cost on the *adjacent* hop -to deny the same reader the same number — `intercept.rs:161-166`: *"not fitting -the frame is the price of leaking zero bits of length... that number would -otherwise reach the parent host, which is the one reader D4 exists to keep it -from"* — and `wire.rs`'s fixed 64 KiB frame implements `hub/REVIEW.md` #12 for -exactly this reason. **The identical quantity is available for free one hop -earlier, to the same reader.** - -**(2) The published transaction timestamps itself, to one block.** For -non-ZIP-318 traffic a wallet anchors near its own tip and sets -`nExpiryHeight = build_height + 40` — ZIP 203's Blossom default, and the same 40 -the hub's `MIN_WALLET_EXPIRY` is derived from (`batcher.rs:47-55`). -`anchorOrchard` and `nExpiryHeight` are **public, immutable fields of the -published transaction**, fixed by the wallet before the shim ever sees it, so -neither the hub's queue nor its shuffle nor its batch can touch them. The -operator observes the divert at wall-clock time `T`, converts `T` to a block -height with their own node, and keeps the batch members whose expiry is -`h(T) + 40`. Batching cannot remove a timestamp the wallet baked in. - -**Why this is not fixed by growing the batch.** A `W`-block flush window admits -`k ~ lambda*W` entries **and spans ~`W` distinct expiry values**, so the members -sharing the target's expiry number `1 + (k-1)/W ~ 1 + lambda`. `W` cancels -(`audit-state/globals/G3-anonymity-set-arithmetic.md`). At the README's own -measured `lambda = 0.77` Orchard-touching transactions per block, the delivered -set on channel (2) alone is **~1.77**, and channel (1) cuts it further and -independently. Raising `k` raises the *candidate* set and leaves *selection* -untouched. - -**Both channels close only under the same unmet precondition.** A ZIP 318 -conforming migration is length-uniform, carries a shared boundary anchor and a -shared bucketed expiry, so conformance closes (1) and (2) together. -`audit-state/SPEC-NOTES.md` §5 records that ZIP 318 is `Status: Draft`, that its -reference-implementation section is a literal `TODO`, and that no shipped wallet -has been shown to implement it. `README.md:69` states the anchor half of that -precondition only as a *wallet-developer* requirement and the length half not at -all, while `README.md:27` and `:33` affirm the protection unconditionally and -volume-independently. - -## Attack Scenario and Steps - -Attacker: **the operator of the indexer the shim fronts** — adversary #1, and the -party the product exists to defend against. Target: a wallet at source IP `X`. -Goal: `IP X -> the specific on-chain transaction -> the amount it moves`. - -Every step below holds with `debug { enabled = false }`, all three PCRs -reproducing, and `caution verify` printing `Attestation verification PASSED`. - -### Step 1 — obtain the wallet-leg byte stream - -*Cost:* one manifest flag, one DNS edit, or nothing at all, depending on the -deployment shape. - -The enclave has no NIC, so **every** wallet TLS session is relayed in by a socket -on the Nitro parent host. Whoever holds that socket sees the wallet's source IP, -connection timing and every cleartext TLS record length. `deploy.sh:156` runs -`caution apps create`, which `OPERATORS.md:64` defines as *"Fully managed: in -Caution's AWS account"* — so in the shipped deployment the operator does **not** -hold that socket by default. They reach it by any of three routes, none of which -changes a PCR and none of which any documented check catches: - -- **1A — BYOC.** `OPERATORS.md:66-69` documents `caution init --byoc`, which puts - the parent host in the operator's own AWS account. Then the position is free - and inherent. `deploy.sh` does not use this path, but the runbook offers it as - a first-class option and an operator who wants their traffic in their own - account will take it. -- **1B — `debug { enabled = false; ssh_keys = [...] }`, and this is the sharpest - one.** `assemble-caution.sh --ssh-key ` without `--debug` renders the key - into the manifest and prints only *"SSH is closed when attestation is on, so - the key is recorded in the HCL but unused until a `--debug` build"* - (`assemble-caution.sh:120-123`); `caution.hcl.tmpl:196-198` repeats the claim. - **The platform disagrees.** `src/api/src/main.rs:2420-2421` reads - `debug.enabled` and `debug.ssh_keys` independently, and - `src/api/src/deployment.rs:2159-2164` emits `ingress 22 from 0.0.0.0/0` - whenever the key list is non-empty, with `user-data.sh:17-27` installing the - keys on the parent. So one flag, in a fully attested deployment, gives the - operator a root shell on the host that runs - `socat TCP-LISTEN:443,reuseaddr,fork VSOCK-CONNECT:16:443`. `tcpdump` there is - the whole of step 1. `ssh_keys` reaches only terraform, so **no PCR moves** — - and since it is unmeasured, the tree the operator publishes as `app_sources` - can carry `ssh_keys = []` while the deployment carries a key, and - `caution verify` still reproduces every PCR. -- **1C — a layer-4 relay on the DNS name the operator already owns.** - `OPERATORS.md:59` lists *"A DNS name you control for wallets"* as a - prerequisite and `deploy.sh:162-176` rewrites that record on every deploy. An - operator who points it at a TCP forwarder they run sees the wallet's IP, - timing and record lengths while the TLS session still terminates inside the - enclave. **Every documented verification step passes, and passes correctly**: - `validate_attested_tls` (`src/cli/src/lib.rs:353-378`) compares the attested - `certfp` against the leaf of the verifier's own connection, and under a pure - forwarder that leaf is unchanged. The platform's one catching check, - `dns_contains_deployment_ip` (`:261-277`), runs only on the raw-IP verify path - (`:6902`, `:6930`), which `README.md:71` and `OPERATORS.md:178` never use — and - attempting it without an out-of-band deployment address resolves the domain to - the relay and compares it to itself. Filed separately as - `operator-controlled-dns-permits-a-layer-4-relay-that-every-documented-verification-step-passes.md`. - -(A fourth route, `DEBUG=1`, is `deploy.sh:52`'s shipped default and gives the -same shell — but it also zeroes the PCRs and makes `caution verify` refuse, so it -is a *different* scenario, already filed as -`deploy-script-defaults-to-debug-mode-which-turns-attestation-off.md`. This chain -deliberately does not use it.) - -### Step 2 — recover `|tx|` - -*Cost:* arithmetic on captured TLS record headers; no decryption. - -TLS records are framed `type(1) || version(2) || length(2)` with the length in -**cleartext**, so for a connection carrying `n` records the observer computes the -application plaintext as `sum(length_i) - 17n` (16-byte AEAD tag plus the inner -content type), with `n` directly observed. The HTTP/2 plaintext of a -`SendTransaction` is `preface + SETTINGS + HEADERS + DATA`, where DATA is -`5-byte gRPC prefix + protobuf RawTransaction{ data: tx }` — i.e. `|tx|` plus a -small constant the operator calibrates once by sending a transaction of known -size through their own shim. - -Two honest notes: - -- **Direction and shape do most of the work.** A `SendTransaction` is the only - large *client->server* upload a light wallet makes; every sync method is a small - request with a large response. The broadcast burst is separable by inspection, - not by subtraction. -- **Multiplexing does not defeat it, and precision is not needed.** If the wallet - shares one h2 connection between sync and broadcast, concurrent sync *requests* - add tens to hundreds of bytes. Orchard bundle size grows in steps of - ~3.1 KB per action, so an error of a few hundred bytes never merges two - candidates. On a dedicated broadcast connection — which ZIP 318's own - sync-decoupling rule *requires* wallets to use — recovery is exact. **The - specification's privacy rule sharpens this channel.** - -### Step 3 — establish that the request was diverted - -*Cost:* zero. *Reliability:* certain; conceded by the project as unfixable. - -`intercept::send_transaction` (`intercept.rs:110-131`) returns through `divert` -**before** `pool.get()` is called, so a diverted transaction opens no TCP -connection to the backend. The operator sees a large wallet-leg upload with no -counterpart at their own indexer. `README.md:33`: *"A diverted request is the one -thing it does not see, and that asymmetry survives padding."* Size alone also -separates a diverted `SendTransaction` from the only other non-forwarded request -class, `GetTransaction`, whose body is capped at `MAX_TX_FILTER_BYTES = 1 KiB` (`intercept.rs:81`, `:249`). - -### Step 4 — bound the candidate set to one flush - -*Cost:* zero. *Reliability:* deterministic. - -`queue::survives_next_flush` (`queue.rs:380-393`) admits an entry only if it -survives the next scheduled flush, and `next_flush_height` (`:363-368`) puts -flushes on multiples of `FLUSH_INTERVAL_BLOCKS = 20`. Flush heights are therefore -publicly computable; `batcher.rs:8-17` chooses this deliberately (*"the cost (the -schedule is public) is acceptable"*). `chain::broadcast_batch` -(`chain.rs:199-210`) issues the full (transaction x endpoint) product -concurrently, so the batch enters the mempool as one simultaneous burst any node -observes. The candidate set for a divert at time `T` is exactly the batch -published at the next multiple of 20 after `T`. - -### Step 5 — resolve the candidate set to one transaction - -*Cost:* one comparison per batch member. - -- Keep the members whose `nExpiryHeight` equals `h(T) + 40` (or, for a wallet with - a different delta, whose expiry is a fixed offset from `h(T)` — the offset is a - per-wallet-implementation constant the operator learns once). Expected - survivors: `1 + lambda ~= 1.77` at the README's own measured rate. -- Independently, keep the members whose serialized length matches step 2. -- `anchorOrchard` is a third, correlated selector for a latest-anchor wallet. - -Together these are near-certain to be unique today. **They collapse to nothing -against a ZIP-318-conforming population** — that is the honest limit of the -finding, and it is also the reason it matters: the README states the protection -unconditionally while it is in fact contingent on a wallet behaviour the README -itself only *requests* at `:69` and that no shipped wallet implements. - -At batch size 0-1, the measured condition (`README.md:34`), step 5 is -unnecessary — the batch *is* the answer. - -### Step 6 — read the amount - -*Cost:* zero. *Reliability:* certain, permanent, retrospective. - -`valueBalanceOrchard` and the Ironwood value balance are public fields of the -published transaction. `README.md:54` states this as the premise of the whole -product. The join completes: **IP `X` -> this transaction -> this amount**, and it -can be run at any future date against archived captures and a permanent chain. - -**Attack Requirements and Assumptions:** - -- **Access needed:** being the operator, plus one of routes 1A/1B/1C. All three - are things the operator either already has (the DNS record, by documented - prerequisite) or obtains with one flag. Plus their own indexer logs and a Zcash - node, both definitional. -- **What makes it realistic:** no exploit, no race, no privileged bug, nothing - active on the wallet leg. The whole chain is passive and retrospective. Route - 1B in particular is a *documented* flag whose danger three zeronym texts - explicitly deny. -- **What limits it, stated plainly:** - 1. Step 5 collapses against a ZIP-318-conforming population. The finding's - force comes from that population not existing yet. - 2. The operator resolves only *their own* clients out of a batch; another - operator's wallets are just "one of these is not mine". - 3. A wallet that routes its broadcast over Tor or Nym itself — which ZIP 318 - requires wallets to *offer* — defeats step 1 outright. - 4. In the shipped fully-managed deployment the operator must take one of the - three actions in step 1; they do not hold the parent socket by default. - This is a real threat-model correction, not a refutation: all three actions - are free, undetectable by every documented check, and available to the - adversary the product is built against. - -## Impact on Users - -Every user of every zeronym endpoint, for as long as their wallet does not -implement ZIP 318 — which is every wallet today, on a migration that is -mandatory, mass and concentrated. - -A user reads `README.md:27` (*"Protected — **Source IP.** ... Volume-independent: -it holds however few others are migrating"*) and `:33` (*"the operator learns -*that* a client migrated, **though not the amount or which transaction**"*) and -concludes that broadcasting through a zeronym endpoint stops their indexer -operator linking their IP to their migration and its value. Against that -operator, in the attested deployment, it does not. - -The result is the permanent, retrospective linkage the product's own Background -section calls *"the attack"*: IP address -> the specific on-chain transaction -> -the balance it moved. It is worse than a live leak because the chain is permanent -and the captures are archivable. - -The behavioural harm is the one ICTM exists to catch. A user who believes the -claim migrates now, over a zeronym endpoint, and does not take the measure that -would actually have protected them — routing the broadcast over Tor or Nym at the -wallet. Neither the chain nor the operator's captures forget. - -The finding also removes the comfort in the disclosed residual. `README.md:34` -tells the reader the weakness is a volume problem and that *"the lever is -adoption, not code"*. That is false for both channels: `W` cancels, so the -delivered set stays at ~`1 + lambda` however wide the window and however large -the batch. The lever for these two channels is **wallet conformance and -wallet-side padding** — code, in the wallet, that the README does not ask for. - -## Technical Details / Code Analysis - -**1. Nothing on the wallet-facing leg pads, and the manifest guarantees the -ciphertext crosses the parent host.** - -`shim/deploy/caution/caution.hcl.tmpl:79-89`: - -``` - # `e2e_encryption { enabled = true }` is Caution's in-enclave TLS - # termination, shipped 2026-08-03. The platform runs a Caddy INSIDE the - # enclave: it obtains the certificate for `domain`, terminates the wallet's - # TLS in there, and forwards to our process on `port`. So the private key is - # generated and held inside the enclave and the operator never holds it, - # which is the property the whole attestation argument depends on. -``` - -and `:106-127` declares `http { domain … port = 8083; upstream_protocol = "h2c"; -e2e_encryption { mode = "tls" } }`. - -This is correct as far as it goes: the operator does not hold the key. But an -enclave has no NIC. The Caution parent relays 443 in with -`socat TCP-LISTEN:$port,reuseaddr,fork VSOCK-CONNECT:16:$port` -(`terraform/modules/aws/nitro-enclave/user-data.sh:197-215`), and TLS record -length fields are not encrypted. **The protection is of content, and `|tx|` is -not content.** - -`shim/src/intercept.rs:94-131` buffers the request whole and either diverts or -replays it verbatim; there is no padding on either arm. The synthesized -wallet-facing replies also differ in length (`grpc_send_response` at `:220-227` -vs `grpc_error` on the three divert-failure arms at `:151`, `:173` and `:209`), -handing the same observer the *verdict* as well. - -**2. The code states the exact secret being protected, and pays for it on the -other hop.** `shim/src/intercept.rs:158-166`: - -```rust - // Too large for the transport's fixed frame. RESOURCE_EXHAUSTED, not - // UNAVAILABLE: this can never succeed, and UNAVAILABLE is the status - // that tells a wallet to retry. It is never forwarded to the operator - // and never broadcast another way; not fitting the frame is the price - // of leaking zero bits of length. - // - // The log line carries the LIMIT, never the transaction's own size: - // that number would otherwise reach the parent host, which is the one - // reader D4 exists to keep it from. -``` - -Two sentences, both about `|tx|`, both naming the parent host as the reader to be -denied. The price is real: a transaction over `MAX_NYM_TX_BYTES = 65,503` is -permanently refused rather than leak its length. **The system pads and shapes the -hop whose observer is weakest, and leaves unpadded the hop whose observer is -adversary #1.** - -`OPEN-QUESTIONS.md:109` is the only place in the repository this is written down: - -> **Accepted non-defenses.** Active wallet-tagging ... and the transaction-size -> side channel (a distinctive migration size re-links via TLS ciphertext length) -> are out of scope by design. **Confirm these are acceptable**, or scope -> mitigations. - -That sentence asks the security reviewer to ratify the acceptance. This issue is -the answer: it is not acceptable while `README.md:33` tells users the opposite and -`README.md:34` tells them adoption is the fix. - -**3. The mixnet leg, for contrast, is genuinely defended — do not claim -otherwise.** Fixed 64 KiB `SubmitV1` frames (REVIEW #12) plus the nym client's own -send shaping mean a real submit *displaces* cover traffic rather than adding to -it. By the crate's own model (`shim/src/nym.rs:1089-1131`: `PACKET_BYTES = 2048`, -a shaped floor of ~8.33 packets/s, a submit = 45 packets, a lookup = 61) there is -no clean per-migration egress burst. This link was checked and refuted as a -channel and should be reported as a defence that works. - -**4. Everything a wallet needs is forwarded to the operator, and the flush clock -is public.** `route_for` (`proxy.rs:744-783`) sends `GetLatestBlock`, -`GetBlockRange`, `GetSubtreeRoots`, `GetTreeState` and `GetLatestTreeState` to -`Route::PassThrough`. `shim/ENDPOINTS.md:135` names the consequence in the -project's own words: - -> `GetLatestTreeState` ... **Anchor correlation, the strongest non-argument -> leak**: this supplies the Orchard anchor the wallet spends against, and that -> anchor root is a public field of the published tx. - -One correction to how this is usually stated, established by the platform read: -the operator's *indexer* does not get a per-wallet boundary, because the parent -relays wallet connections into the enclave and Caddy multiplexes many of them -onto few upstream connections, so pass-through traffic is **not** attributable to -a source IP. **The chain does not need it to be.** Channel (2) works off the -operator's own node clock and the transaction's own `nExpiryHeight`; the divert -timestamp comes from the wallet leg in step 1, not from the indexer. - -**5. Neither selector can be touched by the hub.** `anchorOrchard` is inside the -ZIP 244 `orchard_digest` and `nExpiryHeight` is a consensus field; both are fixed -by the wallet before the shim sees the transaction and are identical on chain -whether the hub published it or the wallet did. The queue keys on -`sha256(tx_bytes)` and `chain::broadcast_batch` publishes those bytes unmodified, -so the on-chain serialized length equals the `|tx|` measured in step 2. - -**6. The stated remedy does not act on either channel.** Raising the batch size -from 1 to `k` raises step 4's candidate set from 1 to `k` and leaves steps 2 and -5 untouched. And widening the window does not help either: a `W`-block window -admits `k ~ lambda*W` entries *and* spans ~`W` expiry values, so the delivered set -is `1 + (k-1)/W ~ 1 + lambda` and `W` cancels — a 24-hour window delivers exactly -what a 1-block window delivers (`audit-state/globals/G3-anonymity-set-arithmetic.md`, -filed as `widening-the-flush-window-cannot-raise-the-delivered-anonymity-set-...md`). -The delivered anonymity set is not `k`; it is **the number of batch members -sharing the target's `(length, anchor, expiry)` tuple**, which for today's -heterogeneous Orchard traffic is ~1. - -## Recommendations - -1. **Correct `README.md`'s Security section so the documented property matches the - delivered one.** Move the ZIP 318 conformance requirement from *Usage -> Wallet - developers* into *Not protected*, and state that until wallets conform, an - operator who observes the wallet leg can re-identify their own clients' - transactions inside a batch of any size by serialized length, anchor and - expiry. Replace *"The lever is adoption, not code"*: it is true of the - batch-size residual and false of these two channels. -2. **Disclose the transaction-size side channel to users.** It exists today only - at `OPEN-QUESTIONS.md:109`, in a list asking reviewers to confirm it is - acceptable, while `README.md:33` tells users the opposite. -3. **Specify wallet-side padding as a second hard requirement, beside aligned - anchors.** It can only be fixed there: the length is fixed by the wallet before - any zeronym code runs. HTTP/2 `DATA` frames carry a `PADDED` flag, so a wallet - can pad every `SendTransaction` to a fixed size (65,503 bytes, the hub's frame - budget, is the natural target) with no protocol change. -4. **Close the three step-1 routes, which are the only part of this chain zeronym - can fix in its own repository.** (a) Correct `assemble-caution.sh:120-123` and - `caution.hcl.tmpl:196-198`, which assert that `ssh_keys` is inert under - attestation — it is not — and make `--ssh-key` without `--debug` a hard error. - (b) Add a DNS-target check to the auditor recipe: resolve the wallet-facing - name and require it to reach the Caution deployment address obtained out of - band, or invoke `caution verify` against the raw deployment IP so - `dns_contains_deployment_ip` runs. (c) State in `THREATMODEL.md` §3 and in - `README.md` which deployment model the live endpoints use, because managed and - BYOC give the operator materially different positions. -5. **Do not attempt to fix this in the shim by padding the wallet-facing reply - alone.** That closes the verdict oracle — worth doing on its own merits — but - not the request-length channel, which is the one that matters. -6. **Measure the achieved anonymity set on the fields that discriminate.** Count - distinct `(length, anchor, expiry)` tuples in a batch, not batch cardinality; - `batcher::flush`'s `achieved <= 1` warning counts the wrong thing. -7. **Sequence the fixes.** Widening the flush window is null before wallet-side - expiry bucketing and useful only after it; do not ship it as a standalone - mitigation. - -## Validation Information - -**Verdict: CONFIRMED. Severity: High** (filed High; upheld). - -Every mechanical link was re-verified against the target and, for the platform -claims, against the Caution platform source (`codeberg.org/caution/platform`, -whose location `shim/deploy/caution/OPERATORS.md:44-46` names, cloned during this -audit) rather than against prose. - -**Verified in `audit-target/zeronym/`:** - -- `caution.hcl.tmpl:79-127` — in-enclave Caddy, `mode = "tls"`, `h2c` upstream. - Confirmed verbatim. -- `intercept.rs:110-131` — divert returns before `pool.get()`; `:158-166` — the - "zero bits of length" comment naming the parent host. Confirmed verbatim. -- `proxy.rs:744-783` — every sync method falls to `Route::PassThrough`. Confirmed. -- `batcher.rs:40,47-55` — `FLUSH_INTERVAL_BLOCKS = 20`, `MIN_WALLET_EXPIRY = 40` - derived from librustzcash's default. `queue.rs:363-368` — flushes on multiples - of 20. `chain.rs:199-210` — `join_all` over the full product, "simultaneity is - the property". All confirmed. -- `README.md:27,33,34,54,69` and `OPEN-QUESTIONS.md:109` — quoted accurately. -- `deploy.sh:156` — `caution apps create`; `:162-176` — the operator writes the - wallet-facing CNAME on every deploy. Confirmed. - -**Verified in the Caution platform source (step 1):** - -- `terraform/modules/aws/nitro-enclave/user-data.sh:197-215` — 443 is added to - `standard_ports` under `e2e_mode == "tls"` and relayed with - `socat TCP-LISTEN:443,reuseaddr,fork VSOCK-CONNECT:16:443`. **The wallet's - ciphertext demonstrably crosses the parent host.** -- `src/api/src/main.rs:2420-2421` — `debug_enabled` and `ssh_keys` are read as - two independent fields; `src/api/src/deployment.rs:2159-2164` — the terraform - template emits `ingress { from_port 22 ... cidr_blocks ["0.0.0.0/0"] }` iff the - key list is non-empty; `user-data.sh:17-27` installs them on the parent. - **Route 1B is confirmed from source, and zeronym's own text - (`assemble-caution.sh:120-123`, `caution.hcl.tmpl:196-198`) asserts the - opposite.** `ssh_keys` reaches only terraform, so no PCR moves. -- `src/cli/src/lib.rs:353-378` (`validate_attested_tls`) — the certfp check - compares the attested fingerprint against the leaf of the verifier's own - connection. Under a pure layer-4 forwarder that leaf is unchanged, so the check - **passes, and passes correctly**: it answers "did my TLS session terminate in - the attested enclave?", to which the answer under a relay is genuinely yes. - `:222-240` + `:6902` + `:6930` — `dns_contains_deployment_ip` runs only when - the verifier supplies a raw deployment IP, which the documented recipes never - do. **Route 1C is confirmed.** - -**Threat-model correction applied to the filed text.** The draft granted step 1 -to the audit's standing premise that the operator owns the parent host. That -premise is model-dependent: `deploy.sh` ships the fully-managed model, where the -parent is in Caution's AWS account. Step 1 has been rewritten as three explicit -routes and the chain now states exactly which deployment shapes it holds in: - -| deployment shape | step 1 | chain | -|---|---|---| -| managed + attested, operator takes no action | not held | **does not complete** | -| managed + attested + `debug.ssh_keys` (one flag, no PCR change) | held | **completes** | -| managed + attested + operator L4 relay on their own DNS name | held | **completes** | -| BYOC + attested | held, free | **completes** | -| `--debug` (`deploy.sh:52` default) | held | completes, but attestation is off — different scenario, separately filed | - -This is a *strengthening* of the finding's practical status, not a weakening: -before this correction, parent-host access looked like an inherent property of -Nitro with no fix. It is instead an operator action, undetectable by every -documented check, and two of the three routes are things zeronym's own repository -can close (recommendation 4). - -**Two corrections made to the filed technical argument, both against the issue:** - -1. *"The operator holds the plaintext of every non-diverted exchange on that - connection and subtracts"* was too strong as a per-wallet claim. The operator's - indexer receives multiplexed connections from the enclave with no per-wallet - boundary and no client address, so pass-through requests are not attributable - to a source IP. The step has been restated on the mechanism that actually - works: a `SendTransaction` is the only large client->server upload a light - wallet makes, so it is separable by direction and shape, and residual error - from concurrent sync requests is a few hundred bytes against a ~3.1 KB - per-action granularity. Recovery is exact on a dedicated broadcast connection, - which ZIP 318's sync-decoupling rule requires. -2. *"Match anchor and expiry against the per-IP sync tip known from pass-through - traffic"* rested on the same wrong premise. The channel is stronger without - it: `nExpiryHeight = build_height + 40` is a self-published, one-block - timestamp, matched against the divert time the operator reads off the wallet - leg using their own node's clock. No per-IP sync attribution is needed. - -**Quantification added.** The delivered anonymity set on the expiry channel alone -is `1 + (k-1)/W ~ 1 + lambda`, i.e. **~1.77** at the README's own measured -0.77 Orchard-touching transactions per block, **independent of the window `W` and -of the batch size `k`**. Per coordinator item 6s, "widen the window" must not be -offered as a mitigation: the window provably cancels. - -**False-positive checks applied.** - -- *§7 Configuration-dependent?* No. Routes 1B and 1C are not insecure - configurations a well-meaning operator stumbles into; they are deliberate acts - by the adversary the product names first, and 1C uses a DNS record - `OPERATORS.md:59` makes a prerequisite. 1B is aggravated rather than excused by - configuration, because three zeronym texts tell the operator (and any reviewer) - that the flag is inert. -- *§8 Requires prior compromise?* No. Nothing here requires reaching inside the - enclave. Every capability used is one the operator has by construction or - acquires with one flag. -- *§3 Information disclosure overstated?* No. The disclosed value is a specific - user's shielded-transaction amount joined to their IP address — the exact harm - `README.md:54` defines as "the attack". -- *§6 Intentional design?* Partly, and it is stated as such: the absence signal - (`README.md:33`) and the batch-size residual (`:34`) are disclosed. What is - **not** disclosed is that selection inside a batch defeats the disclosed - remedy, and that the size channel exists at all outside `OPEN-QUESTIONS.md`. - -**Why High and not Critical:** no funds are stolen or destroyed, the attacker must -be the endpoint's own operator, and a ZIP-318-conforming wallet population would -close both channels. **Why High and not Medium:** the harm is the product's -single reason to exist, it lands on every user of an affected endpoint, it is -passive, free and retrospective against archived captures and a permanent chain, -the three step-1 routes are all available to adversary #1 without detection, and -the project's own stated remedy ("adoption, not code") provably does not act on -it. - -**Deliberately not merged, to avoid double-counting.** This chain uses *no* -defect in the classifier, the queue, the batcher or the enclave boundary. The -operator's stronger options — appending their own hub -(`shim-submits-every-migration-to-every-configured-hub-...md`, confirmed High) and -the `DEBUG=1` log leak (`log-verdict-logs-migration-value-balance-at-info.md`, -confirmed High) — are cited as context, not folded in; the report should present -this as the chain that survives when those are fixed, and should not stack the -three severities as three separate losses of the same secret. - - ---- - -## ADDENDUM (Global auditor, focus area G4 — parent-host side channels, 2026-08-18). NOTHING ABOVE IS WITHDRAWN OR CHANGED. This is a **completeness correction to the remediation**, plus the disposition of a `[***]` brainstorm item that was explicitly deferred to the G4 pass. - -**Padding the wallet-leg REQUEST is not sufficient. The four synthesized -wallet-facing REPLIES are unpadded too, they are distinguishable by length, and -they are read by the same socket on the same connection.** - -The obvious fix for the channel this issue documents is to pad or chunk the -wallet→shim request so `|tx|` stops being recoverable from TLS record lengths. -That fix leaves a second, smaller channel in place on the return leg, and the -return leg is entirely under the shim's control — it *synthesizes* every reply on -the divert path, so unlike the request it can be made constant-size with no wallet -change at all. - -`shim/src/intercept.rs:137-213` (`divert`) produces exactly four wallet-facing -shapes, each a different number of bytes on the wire: - -| divert outcome | reply | body bytes on the wire | -|---|---|---| -| `Submit::Accepted { txid }` | `grpc_send_response(0, <64-hex txid>)` | 5-byte gRPC prefix + 66-byte `SendResponse` (`errorCode` = 0 is a proto3 default and is omitted; `errorMessage` is tag+len+64) = **71** | -| `Submit::AlreadyKnown { txid: None }` | `grpc_send_response(0, "")` | both fields default → empty message = **5** | -| `Submit::Rejected { reason }` | `grpc_send_response(-1, reason)` | `errorCode` = -1 encodes as a 10-byte varint, so ≥ **16** plus the reason | -| body unreadable / hub unreachable / too large | `grpc_error(...)`, trailers-only | **0** body, with `grpc-message` of 47, 34 and 69 characters respectively in the HEADERS frame (`intercept.rs:150-154`, `:207-212`, `:170-178`) | - -Scope, stated honestly so the report does not overclaim: this does **not** -distinguish a divert from a pass-through — a successful pass-through returns the -backend's own `SendResponse`, which is also a 71-byte body — and the `TooLarge` -arm's information (`|tx| > 65503`) is already implied by the request-length -channel this issue is about. What the reply lengths add is the **divert -outcome**: whether the migration was carried or destroyed, per wallet, per -attempt. Against today's code that is a small increment on a channel the same -reader already has, which is why it is recorded here rather than filed -separately (items 8/9 double-counting precedent). Against the *fixed* code it is -the whole channel, which is why it must be part of the fix. - -**Recommendation, to be applied together with this issue's existing ones:** pad -every synthesized wallet-facing reply on the intercepted paths to one constant -size — the same discipline `shim/src/wire.rs` already applies to the mixnet -frames for exactly this reader, and the discipline `shim/src/intercept.rs:156-166` -already cites as the reason it refuses an over-64-KiB migration rather than -leaking its length. Concretely: give `grpc_send_response` and the divert-path -`grpc_error` arms a common fixed-length `errorMessage`/`grpc-message` envelope, so -`Accepted`, `AlreadyKnown`, `Rejected` and all three failure arms are -byte-identical in length. This is a change inside one function and needs no -wallet cooperation. - -**Disposition recorded so nobody re-chases it:** this closes `BRAINSTORM.md`'s -`[***]` item "The four synthesized wallet-facing responses have distinguishable -sizes, giving the parent host a verdict oracle on the diverted path", which was -explicitly deferred to the parent-host side-channel global pass. It is disposed -of **into this issue's remediation**, not filed as a separate finding. - -DO NOT DEVIATE FROM THIS TEMPLATE (`docs/SECURITY-ISSUE-TEMPLATE.md`). -LEAVE THIS MESSAGE IN PLACE SO THAT ALL AGENTS KNOW NOT TO DEVIATE FROM THIS TEMPLATE (`docs/SECURITY-ISSUE-TEMPLATE.md`). diff --git a/zeronym-22aa9851caf68-high-medium/high/deploy-script-defaults-to-debug-mode-which-turns-attestation-off.md b/zeronym-22aa9851caf68-high-medium/high/deploy-script-defaults-to-debug-mode-which-turns-attestation-off.md deleted file mode 100644 index e41ff8fa..00000000 --- a/zeronym-22aa9851caf68-high-medium/high/deploy-script-defaults-to-debug-mode-which-turns-attestation-off.md +++ /dev/null @@ -1,361 +0,0 @@ -# `deploy.sh` defaults to `DEBUG=1`, so the repository's one-command deploy ships the configuration that zeroes the attestation PCRs, opens SSH on the operator's host, and turns on per-request wallet-method logging - -**Severity**: High -**Validation Status**: Confirmed -**Location**: `audit-target/zeronym/deploy.sh:52` (`DEBUG=${DEBUG:-1}`), `:128-134` (the branch that appends `--ssh-key … --debug`), `:206` and `:332` (the two places `DEBUG != 1` gates verifiability); `audit-target/zeronym/deploy.env.example:66` (`DEBUG=1`) and `:73` (`APP_SOURCE=` empty); the effects at `audit-target/zeronym/shim/deploy/caution/assemble-caution.sh:567-574`; the prohibitions at `audit-target/zeronym/hub/deploy/caution/OPERATORS.md:82` and `:263`, `audit-target/zeronym/shim/deploy/caution/OPERATORS.md:358-360`, and `audit-target/zeronym/shim/deploy/caution/caution.hcl.tmpl:186-199`; the claims it defeats at `audit-target/zeronym/README.md:26`, `:33` and `:71` -**Found by agent:** Local (file audit of `deploy.sh`); validated 2026-08-18 -**In scope of audit?** Yes — `deploy.sh` and `*/deploy/**` are explicitly in scope: "the reproducible-build and attestation chain **is** the trust model here, so a break in it is a security finding, not tooling noise." - -## Description - -`deploy.sh` is the repository's one-command deploy path. Its own header presents -it as *the* way to deploy (`deploy.sh:29-31`), and -`shim/deploy/caution/OPERATORS.md:129` sends operators to it ("`zeronym/deploy.sh` -automates this in the right order"). Its debug switch defaults to **on** -(`deploy.sh:52`): - -```sh -DEBUG=${DEBUG:-1} -``` - -and the operator template it tells you to copy ships the same value -(`deploy.env.example:66`): - -``` -DEBUG=1 # 1 = debug enclave (SSH console on); 0 = attested -``` - -An operator who does exactly what the script says — copy `deploy.env.example` to -`deploy.env`, fill in the names, run `./zeronym/deploy.sh` — gets a debug -enclave. Nothing else has to go wrong. - -`--debug` does three things at once -(`shim/deploy/caution/assemble-caution.sh:567-573`): - -1. flips `enabled = false` → `enabled = true` in the manifest's `debug` block, - which **disables attestation**; -2. uncomments `RUST_LOG = "zis::proxy=debug,info"`, which turns on a per-request - log line naming the gRPC method every wallet called; -3. writes the operator's SSH public key into `debug.ssh_keys`, opening port 22 on - the Nitro **parent host**, from which - `/var/log/nitro_enclaves/enclave-console.log` is readable. - -Both operator runbooks forbid this in bold, in the imperative: - -- `hub/deploy/caution/OPERATORS.md:82` and `:263`: **"Never pass `--debug`. Debug - mode disables attestation."** -- `shim/deploy/caution/OPERATORS.md:358-360`: debug "**disables attestation**, so - it is a diagnostic only, never the deployed config". -- `shim/deploy/caution/caution.hcl.tmpl:187-188`: "FALSE is the point of the - exercise: debug mode disables attestation, and an unattested shim proves - nothing that running it on a laptop would not." - -So the shipped default of the deploy script is precisely the configuration the -project's own documentation says must never be deployed. That inversion is the -defect. - -**This is default-insecure, not silent.** The operator is warned twice on their -own terminal, and an auditor who runs the documented verification procedure is -*not* fooled. Those calibrations are stated in full below and they are the reason -this is filed as High rather than Critical. What they do not do is give the -wallet user — the party who bears the loss — any signal at all. - -## Attack Scenario and Steps - -No exploitation step is required of an attacker. The insecure state is reached by -the operator following the repository's own happy path, and the system's -**primary adversary** — the indexer operator, who owns the Nitro parent host — is -handed the capabilities for free. - -1. An indexer operator follows `deploy.env.example`'s instructions: copy it to - `deploy.env`, set `COMPONENT`, `NAME`, `TLS_DOMAIN`, `DNS_DOMAIN`, `BACKEND`, - `BACKEND_TLS`, `HUB_NYM`, export `VULTR_API_KEY`, run `./zeronym/deploy.sh`. - They do not touch `DEBUG`, because it is already set to the value the file - ships with and nothing requires them to make a choice. -2. `deploy.sh:128-130` appends `--ssh-key --debug` to the - assemble invocation. -3. `assemble-caution.sh:567-573` rewrites the rendered manifest as described - above. -4. The enclave boots and serves wallets. From a wallet's point of view nothing is - different: the same hostname, the same valid Let's Encrypt certificate, the - same gRPC behaviour, the same `/healthz` `200 ok`. -5. The operator SSHes to the parent host and reads - `/var/log/nitro_enclaves/*.log` — the procedure their own runbook gives at - `shim/deploy/caution/OPERATORS.md:358-360`, and the one the project states it - used to diagnose "every previous 'boots but never serves' bug". Everything the - enclave prints is now theirs, including `intercept::log_verdict`'s per-migration - `orchard_vb` (see the companion issue). - -**Attack Requirements and Assumptions:** - -- The attacker is the **primary adversary in this system's threat model**: the - indexer operator who deploys the shim and owns the Nitro parent host. They need - no special access — they simply do nothing. -- No wallet, and no user, can observe any of this. Nothing on the wallet-facing - surface differs. -- **The friction is asymmetric and points the wrong way.** The insecure path costs - zero effort: `deploy.sh:129` needs only `~/.ssh/id_ed25519.pub`, which almost - every developer already has. The secure path is a hard stop until additional - infrastructure exists — `deploy.sh:132` refuses to proceed unless `APP_SOURCE` - is set, which means the operator must first create a **public git repository**, - arrange **push credentials** for it (`deploy.sh:206-218`), and keep it in sync. - `deploy.env.example` ships `APP_SOURCE=` empty (`:73`), so the template as - distributed cannot be deployed attested by editing one value. - -## Impact on Users - -A wallet user connecting to a debug-deployed endpoint gets **none** of the -product's security properties, and cannot tell. - -- **The attestation chain, which is the entire trust model, does not exist.** - AWS zeroes all PCR values for an enclave launched with `--debug-mode`, and - Caution documents `debug.enabled` as *"Allows reading enclave console output but - disables attestation verification"* (both fetched during this audit; see the - G10 addendum below). So there is nothing to compare, for anyone — not for a - wallet author, not for a third party, not for the operator themselves. - `README.md:71` tells auditors to *"fetch its attestation, check the PCRs against - the AWS Nitro root, reproduce the build and compare hashes"*; on a debug - deployment step 2 is vacuous and step 3 is impossible, because - `deploy.sh:128-134` also passes no `--app-source` and `caution verify` then - refuses outright (`assemble-caution.sh:439-443`: *"Cannot reproduce private code - deployment"*). -- **The parent host gets the enclave's console, which is what converts the shim's - in-enclave logging from inert to live.** Coordinator open item 7 established - from vendor documentation that enclave console output reaches the parent - **only** in debug mode. So this default is exactly what makes - `shim/src/intercept.rs::log_verdict`'s INFO line — `orchard_vb`, `ironwood_vb`, - `sapling_vb`, `expiry`, `inputs`, `outputs`, `tx_len` for **every** transaction - the shim classifies — readable by the adversary. The operator already sees the - wallet's source IP at the TCP layer. Joining that to a per-migration value - balance is the exact link — IP → transaction → amount — that - `README.md:54` calls "the attack" and that - `audit-context/AUDIT-INSTRUCTIONS.md` names as the adversary's core goal. - **This issue is the root defect; the log-discipline findings are its blast - radius.** -- **Per-request method logging.** `RUST_LOG="zis::proxy=debug,info"` records which - gRPC method each caller invoked. `caution.hcl.tmpl:178-182` describes that exact - line as *"exactly the metadata this component exists to deny an operator"*. -- **SSH on the parent.** Beyond the console, this gives the operator (and anyone - who later compromises that host, or compels it legally) a shell from which to - observe packet sizes and timing on the enclave's interfaces and to restart the - enclave at will — which for a hub destroys the RAM-only queue of migrations - wallets were already told had succeeded, and rotates its Nym identity. - -Note that the console log is a **persistent artefact on the operator's disk**, so -the exposure is not only to a deliberately hostile operator: it also reaches -anyone who breaches that host, anyone who obtains a backup, and any legal process -served on the operator. A well-meaning operator who takes the free path exposes -their own users retroactively. - -**Detectability, stated precisely.** A third-party auditor who actually runs the -`README.md:71` procedure is **not** fooled into a false positive: `caution verify` -refuses for want of `app_sources`, and the PCRs are zero. There is, however, no -positive signal on any wallet-facing endpoint — `/healthz`, `/nym-status` and the -gRPC surface are byte-identical — and a wallet user has no mechanism at all. The -residual is that an endpoint nobody happens to audit is indistinguishable, to its -users, from an attested one. - -## Technical Details / Code Analysis - -**The default and the branch it selects** (`deploy.sh:51-53`): - -```sh -DNS_TTL=${DNS_TTL:-300} -DEBUG=${DEBUG:-1} -SSH_PUBKEY_FILE=${SSH_PUBKEY_FILE:-$HOME/.ssh/id_ed25519.pub} -``` - -**The argument list it builds** (`deploy.sh:128-134`): - -```sh -if [ "$DEBUG" = 1 ]; then - [ -f "$SSH_PUBKEY_FILE" ] || die "SSH_PUBKEY_FILE not found: $SSH_PUBKEY_FILE" - set -- "$@" --ssh-key "$(cat "$SSH_PUBKEY_FILE")" --debug -else - [ -n "${APP_SOURCE:-}" ] || die "DEBUG=0 (attested) requires APP_SOURCE (public repo URL for caution verify)" - set -- "$@" --app-source "$APP_SOURCE" -fi -``` - -The asymmetry is directly visible: the debug branch needs only a file that almost -every developer already has; the attested branch is a hard stop until a public -repository exists. - -**What `--debug` does** (`shim/deploy/caution/assemble-caution.sh:559-574`): - -```sh -# --debug: flip the enclave into debug mode and turn on per-request shim logging. -# This is a DIAGNOSTIC build, not a shippable one, for two reasons stated in the -# template: debug mode disables attestation (so nothing it runs is provable), and -# RUST_LOG=zis::proxy=debug logs the gRPC method each caller invokes, which is the -# exact metadata the shim exists to deny an operator. … -if [ "$DEBUG" = "true" ]; then - sed -i.bak \ - -e 's|^ # RUST_LOG = "zis::proxy=debug,info"| RUST_LOG = "zis::proxy=debug,info"|' \ - -e 's|^ enabled = false| enabled = true|' \ - "$DEST/caution.hcl" - rm -f "$DEST/caution.hcl.bak" - echo "==> DEBUG build: attestation OFF, SSH console ON, per-request logging ON. Diagnostic only." -fi -``` - -The comment is an accurate description of a configuration the enclosing tool then -makes the default. (`assemble-caution.sh:114-118` additionally *requires* -`--ssh-key` whenever `--debug` is given, so debug and an open SSH console are -inseparable.) - -**The verifiability gates**, both of which the default fails -(`deploy.sh:206` and `:332`): - -```sh -if [ "$DEBUG" != 1 ] && [ -n "${APP_SOURCE:-}" ]; then - … - log "publishing the app-source to $APP_SOURCE_PUSH (tag $APP_SOURCE_TAG) ..." -``` - -```sh -verify : $( [ "$DEBUG" != 1 ] && [ -n "${APP_SOURCE:-}" ] && printf '%s @ %s — run: caution verify' "$APP_SOURCE" "${APP_SOURCE_TAG:-}" || printf 'n/a (debug deploy, not attested; no app-source published)' ) -``` - -**One point in the code's favour, checked and confirmed:** the comparison is -against the literal string `1`, so a mistyped value (`DEBUG=true`, `DEBUG=yes`, -`DEBUG=0 `) falls into the `else` branch and fails *toward* the attested path. The -only values that produce a debug enclave are a deliberate `1` — or the absence of -the variable entirely, which is the default under audit. - -**The template's comment names only one of the three effects** -(`deploy.env.example:66`): it states the SSH console explicitly, attestation only -by implication ("0 = attested"), and does not mention the `RUST_LOG` effect at -all. `deploy.sh:22` compounds this: *"DEBUG=1 here is now about the SSH console -alone."* Read as a statement about what `DEBUG=1` *does*, that is false, and the -same sentence's earlier clause (*"--debug turns attestation OFF"*) shows the -author knew. The charitable reading is that it is a statement about *motivation* -(the enclave console is no longer needed to read a hub's Nym address), but it is -the second place a reader is pointed at the SSH console and away from the other -two effects. - -**Vendor confirmation that debug removes attestation entirely, not merely weakens -it** (primary sources fetched during the G10/G12/G13 global audit): - -- `docs.caution.co/guides/verify-an-app/`, call-out box **"Debug mode cannot be - verified"**: *"AWS Nitro Enclaves **zero out PCR values** in debug mode. Remove - the `debug` block from `caution.hcl` and redeploy before verifying a production - app."* Prerequisites: *"You need … **A deployed Caution app running outside - debug mode**."* -- `docs.caution.co/reference/caution-hcl/`, field `debug.enabled`: *"Enable debug - mode. Allows reading enclave console output but disables attestation - verification."* -- Caution's `terraform/modules/aws/nitro-enclave/user-data.sh` installs - `capture-enclave-console.sh` as a systemd unit inside a - `%{ if debug_mode == "true" ~}` block — i.e. the parent host actively captures - the console to disk in debug mode. - -## Recommendations - -1. **Invert the default**: `DEBUG=${DEBUG:-0}`, and set `DEBUG=0` in - `deploy.env.example`. This is the single-line fix and it removes the defect. -2. **Make debug an explicit, per-invocation act rather than a config-file value.** - Require it on the command line (`./deploy.sh --debug deploy.env`) or via an - obviously-dangerous variable name such as - `I_UNDERSTAND_THIS_DISABLES_ATTESTATION=1`, so an operator cannot arrive at it - by leaving a template field alone. -3. **Remove the friction asymmetry so the secure path is not the expensive one.** - Allow `DEBUG=0` without `APP_SOURCE`, printing the "not independently - verifiable" warning `assemble-caution.sh:439-443` already emits. An - attested-but-unpublished enclave is strictly better than an unattested one, and - today the script forces the operator to choose between *doing the repo work* - and *turning attestation off*. -4. **List all three effects** in `deploy.env.example:66` and `deploy.sh:22`: - attestation off, SSH console on, per-request gRPC-method logging on. -5. **Give the deployment an externally observable signal.** Surface `debug.enabled` - (or simply the fact that the PCRs are zero) in what the endpoint publishes at - `/attestation` or `/nym-status`, so a wallet, a monitor, or a passer-by can tell - a diagnostic deployment from a real one without holding operator credentials. - -## Validation Information - -**Verdict: CONFIRMED. Severity: High.** - -Every mechanical claim was re-verified against the target: - -| Claim | Verified at | -|---|---| -| `DEBUG=${DEBUG:-1}` | `deploy.sh:52` — read directly | -| Template ships `DEBUG=1`, `APP_SOURCE=` empty | `deploy.env.example:66`, `:73` | -| `--debug` ⇒ `--ssh-key` required | `assemble-caution.sh:95`, `:114-118` | -| `--debug` flips `enabled = false` → `true` and uncomments `RUST_LOG` | `assemble-caution.sh:567-573` | -| Template's `debug { enabled = false }` and its rationale | `caution.hcl.tmpl:186-199` | -| Runbooks forbid `--debug` | `hub/…/OPERATORS.md:82`, `:263`; `shim/…/OPERATORS.md:358-360` | -| `RUST_LOG` line called "exactly the metadata this component exists to deny an operator" | `caution.hcl.tmpl:178-182` | -| `DEBUG=1` skips the app-source publish and renders `verify : n/a` | `deploy.sh:206`, `:332` | -| Console is readable on the parent in debug and *only* in debug | Coordinator open item 7 (AWS + Caution vendor docs); corroborated in-tree by `shim/deploy/README.md:289` ("neither readable on an attested enclave, which has no console") and by `shim/…/OPERATORS.md:358-360` describing the console as how every past boot bug was diagnosed | -| `deploy.sh` is a path operators are pointed at | `deploy.sh:29-31`; `shim/…/OPERATORS.md:129` | - -**The `docs/AVOIDING-FALSE-POSITIVES.md` §7 counter-argument, stated and -answered.** §7 warns against grading configuration-dependent issues highly, and -gives as its canonical example: *"Debug mode information leakage (marked High → -should be Low) — Only with `DEBUG=true` … **Production uses `DEBUG=false` by -default**."* That example does not apply here, because the polarity is inverted: -production does **not** use `DEBUG=false` by default; `deploy.sh` and the shipped -template both select `DEBUG=1`. §7's own contrasting "Real Issue" line is exactly -this case — *"insecure protocols **enabled by default**"* — as is §4's -(*"Real vulnerability would be: same issues … **enabled by default**"*). The -related §7 carve-out for *"insecure flags that communicate their own security"* -also does not apply, because that carve-out is about a flag the user -affirmatively chooses (`--dangerously-skip-permissions`); here the user chooses -nothing and inherits the dangerous state by omission. - -**Severity justification — High, and why not Critical or Medium.** - -*Impact:* total. Every security property the product advertises is void -simultaneously: attestation (PCRs zeroed, so the check is not weakened but -vacuous), reproducibility (`caution verify` refuses), console isolation (the -parent captures the enclave's stdout to disk, delivering per-migration value -balances to the primary adversary), and request-metadata privacy (`zis::proxy` -per-request method logging). The affected population is every wallet user of -every endpoint deployed this way, and the harm is retrospective and permanent -because the chain is permanent and the console log persists. - -*Likelihood:* high. It is the shipped default of the repository's one-command -deploy and of the config template operators are told to copy; the secure -alternative is gated behind creating a public git repository and provisioning -push credentials; and the audience is third-party indexer operators, onboarding -of whom began 2026-08-10. - -*Why not Critical:* no funds are stolen or destroyed; the operator is warned twice -on their terminal (`assemble-caution.sh:573` and the `verify : n/a` banner at -`deploy.sh:332`); both runbooks forbid `--debug` in bold; the canonical operator -runbook `shim/deploy/caution/OPERATORS.md` documents the manual assemble path -with `--app-source` and never with `--debug`; and any third party who runs the -documented verification is definitively **not** fooled. This is a -default-insecure, loudly-announced configuration, not a silent backdoor. - -*Why not Medium:* the two facts that would normally justify a downgrade — "the -secure setting is the default" and "the affected party can tell" — are both false -here. The insecure setting *is* the default, and the affected party (the wallet -user) has no signal whatsoever. Warnings addressed exclusively to the party who -benefits from ignoring them do not protect the party who is harmed. - -**Corrections made during validation.** The pre-validation draft's framing was -sound; three points were tightened rather than changed: - -1. The consequence of debug is stronger than "verify refuses": AWS zeroes all - PCRs, so there is nothing for anyone to compare, ever. Stated in the impact - section. -2. The "silent" adjective was removed throughout in favour of "default-insecure", - and the operator-facing warnings are now stated in the Description rather than - only in a mitigations paragraph, so the finding cannot be read as claiming a - covert failure. -3. Added the persistence angle — the debug console is written to disk on the - parent host (`capture-enclave-console.sh`), so the exposure survives the - session and reaches breach, backup and legal-process adversaries, not only a - deliberately hostile operator. - -**Cross-references.** This issue is the **parent** of every log-discipline -finding: `log-verdict-logs-migration-value-balance-at-info.md` (confirmed, High) -and `hub-per-admission-info-log-is-a-real-time-per-entry-arrival-feed.md` are -exploitable *because of* this default and are inert without it (coordinator open -item 7). The report should present them in that order. - -DO NOT DEVIATE FROM THIS TEMPLATE (`docs/SECURITY-ISSUE-TEMPLATE.md`). -LEAVE THIS MESSAGE IN PLACE SO THAT ALL AGENTS KNOW NOT TO DEVIATE FROM THIS TEMPLATE (`docs/SECURITY-ISSUE-TEMPLATE.md`). diff --git a/zeronym-22aa9851caf68-high-medium/high/gettransaction-flood-starves-migration-diversion.md b/zeronym-22aa9851caf68-high-medium/high/gettransaction-flood-starves-migration-diversion.md deleted file mode 100644 index d1648df5..00000000 --- a/zeronym-22aa9851caf68-high-medium/high/gettransaction-flood-starves-migration-diversion.md +++ /dev/null @@ -1,551 +0,0 @@ -# An unauthenticated `GetTransaction` flood converts ~100-byte gRPC requests into 61-packet mixnet emissions, saturating the shim's single egress and taking migration diversion down — loudly on a shared 32-slot channel, silently in the SDK's unbounded buffer - -**Severity**: High -**Validation Status**: Confirmed -**Location**: -`audit-target/zeronym/shim/src/intercept.rs:229-247` (routing: with a hub configured, EVERY `GetTransaction` goes to the hub), `:237-295` (`get_transaction`, the whole of its admission control), `:315-324` (fail closed), `:137-215` (`divert`), `:204-214` (the `UNAVAILABLE` arm); -`audit-target/zeronym/shim/src/nym.rs:71` (`REQUEST_TIMEOUT` = 90 s), `:80` (`SUBMIT_DISPATCH_TIMEOUT` = 5 s), `:96` / `:104` (SURB counts), `:307-309` (`is_healthy`), `:595-690` (`NymHandle::submit`, `:660` the 5 s send, `:685-689` `TransportGone`), `:695-790` (`get_transaction` / `each_target`, `:758-759` the 90 s send), `:835-905` (`correlate`), `:1085-1130` (`throughput_budget`, the crate's own emission model); -`audit-target/zeronym/shim/src/main.rs:335-336` (channel capacities 32 and 8); -`audit-target/zeronym/shim/src/nym_driver.rs:362-375` (one send in flight), `:416-436` (the crate's own statement that the SDK holds *"an unbounded transmission buffer drained at the throttled rate"* which *"may include SUBMITS ALREADY ANSWERED SUCCESS to a wallet"*), `:608-623` (`send_frame`); -`audit-target/zeronym/shim/src/hub.rs:228-249` (`Ok(()) => Submit::Accepted` at hand-off; there is no `Refused` arm); -`audit-target/zeronym/shim/src/proxy.rs:470-541` (accept loop: no connection cap), `:596-610` (h2 server: window sizes only, no `max_concurrent_streams`), `:743-748` (`route_for`); -`audit-target/zeronym/shim/src/wire.rs:455-458` and `:72` (`FRAME_BYTES` = 64 KiB, every reply padded); -`audit-target/zeronym/shim/deploy/caution/caution.hcl.tmpl:34-35` (`memory_mb = 2048`), `:141` (`ZIS_LISTEN = 0.0.0.0:8083`), and the `ingress 0.0.0.0/0` block. -The drain rate that decides the arithmetic is in the pinned SDK (`451c2aa3692fc4dc00041b74a352d4158176d9c0`): `common/client-core/src/client/base_client/mod.rs:1013` (capacity-1 input channel), `real_messages_control/mod.rs:150` (8-slot batch channel), `real_messages_control/real_traffic_stream.rs:427-490` (`poll_poisson`: one batch stored, one packet emitted, per tick), `client/transmission_buffer.rs:39-49` (no cap), `real_messages_control/message_handler.rs:500-557` (fragments prepared eagerly and pending acks armed **before** emission), `config-types/src/lib.rs:25`, `:443` (20 ms default). - -**Found by agent:** Local (file audit of `shim/src/intercept.rs`); validated 2026-08-18 -**In scope of audit?** Yes - -## Description - -With a hub configured, `intercept::get_transaction` turns **every** well-formed -`GetTransaction` request from an unauthenticated, wallet-facing, internet-reachable -listener into a full mixnet round trip to the hub (`intercept.rs:229-247`, -`:295`). The entire admission control on that path is: body ≤ 1 KiB, a decodable -`TxFilter`, and a 32-byte hash (`:249-293`). Thirty-two random bytes pass all -three. There is no authentication, no rate limit, no per-peer accounting and no -concurrency cap anywhere between the socket and the mixnet. - -A lookup costs the shim **61 sphinx packets** — 60 attached reply SURBs plus the -64-byte request (`nym.rs:104`, `:1120-1130`) — against a single serialised mixnet -client whose floor rate the crate itself fixes at **8.33 packets/s** -(`THROTTLED_PACKETS_PER_SEC`, `nym.rs:1090-1094`: `MAX_DELAY_MULTIPLIER` 6 × the -SDK's 20 ms default). So **one ~100-byte HTTP request buys ~7.3 seconds of the -shim's entire outbound mixnet capacity**, and one request every few seconds holds -that capacity permanently consumed. The shim also pays the same 61 packets for a -lookup as it pays 45 for a migration, so lookups and migrations compete for one -non-substitutable resource. - -Three separate harms follow, in increasing order of how much attacker -concurrency they need and decreasing order of how loud they are. - -**(1) Every wallet's `GetTransaction` through that shim stops working. Certain, -and essentially free.** Once the emission backlog exceeds `REQUEST_TIMEOUT` -(90 s, `nym.rs:71`) every lookup times out, sweeps its address list and fails -closed with `UNAVAILABLE` (`intercept.rs:315-324`). One request every ~7 s is -enough to keep the backlog growing. - -**(2) A genuine migration is accepted, answered `error_code 0`, and then rots.** -`shim/src/hub.rs:228-249` answers the wallet `Submit::Accepted` with a locally -computed txid the moment the frame is accepted by an in-process channel; it has -no `Refused` arm at all. Behind that hand-off is the SDK's **unbounded** -transmission buffer, which the shim's own code describes as *"an unbounded -transmission buffer drained at the throttled rate … Frames in there may include -SUBMITS ALREADY ANSWERED SUCCESS to a wallet"* (`nym_driver.rs:416-436`). Under -this flood that buffer grows without bound, so an acknowledged migration either -arrives at the hub long after its admission window (`Refusal::ExpiryTooTight`, -into an ack the shim constructs a receiver for and immediately drops, -`nym.rs:652`) or is destroyed outright by any client rebuild, rotation, redeploy -or SIGTERM. The wallet was told it succeeded and the shim keeps no record. - -**(3) With more concurrency, migrations fail closed at the wallet.** Lookups and -submits share **one** `mpsc::Sender` of capacity 32 (`main.rs:335`), but -a lookup is willing to wait 90 s to be accepted into it (`nym.rs:758-759`) while a -submit is willing to wait 5 s (`nym.rs:639`, `:660`). tokio's bounded `mpsc` is a -fair FIFO semaphore, so a submit joins the queue behind every lookup already -parked and is served in their order. When its 5 s elapses, `NymHandle::submit` -breaks with `dispatched == 0` and returns `NymError::TransportGone` -(`nym.rs:685-689`), and `divert` answers the wallet `UNAVAILABLE: hub unreachable` -(`intercept.rs:204-214`). - -Throughout all three, `/healthz` answers 200: `MixnetStatus::is_healthy` is -`configured && connected` (`nym.rs:307-309`) and the client stays connected — it -is merely backlogged. - -## Attack Scenario and Steps - -1. The attacker picks any operator's shim. The deployment listens on - `0.0.0.0:8083` behind a public TLS domain with `ingress 0.0.0.0/0` - (`caution.hcl.tmpl:141` and its network block), and the listener is - wallet-facing and unauthenticated by design. -2. The attacker opens HTTP/2 connections and issues concurrent - `POST /cash.z.wallet.sdk.rpc.CompactTxStreamer/GetTransaction` calls, each - carrying a gRPC-framed `TxFilter { hash: <32 random bytes> }` — about 45 bytes - of protobuf plus the 5-byte prefix. Random hashes pass every check the shim - makes (`intercept.rs:274-293`). -3. Each call reaches `diversion.hub.get_transaction(&filter.hash)` - (`intercept.rs:295`), becomes a `LookupV1` with 60 attached reply SURBs, and - consumes ~7.3 s of the shim's whole mixnet egress. -4. **For harms (1) and (2), a trickle is enough**: one request every few seconds - keeps the emission backlog growing without bound, since the SDK absorbs whole - messages far faster than it emits packets (see Technical Details). Cost: a few - hundred bytes per minute. -5. **For harm (3), the attacker raises concurrency** until the 32-slot channel's - waiter list is deeper than the 5 s a submit will wait. The threshold is - derived below: of order **750 concurrent live requests at the throttled rate** - (~4,500 at the SDK's unthrottled default). Each is ~100 bytes and lives ≤90 s, - so sustaining it costs roughly 5 KB/s. Nothing caps it: the accept loop has no - connection limit (`proxy.rs:470-541`) and the h2 server sets only window sizes, - never `max_concurrent_streams` (`proxy.rs:596-610`). -6. A real user's wallet then sends its migration. `intercept::send_transaction` - classifies it `Class::Migration` and calls `divert` → `HubTransport::submit` → - `NymHandle::submit`, whose `timeout_at(now + 5 s, self.requests.send(request))` - sits behind the attacker's parked lookups, expires, breaks with - `dispatched == 0`, and returns `TransportGone`. The wallet receives - `grpc-status: 14 UNAVAILABLE`. -7. The wallet, correctly reading UNAVAILABLE as "retry", retries — and fails - again, for as long as the attacker keeps going. - -**Attack Requirements and Assumptions:** -- Any host on the internet can reach the shim; no credential, no wallet, no - funds, no valid transaction, **and no Nym client of any kind**. -- Cost: a few hundred bytes per minute for harms (1) and (2); a few KB/s and - several hundred concurrent h2 streams for harm (3). -- Nothing in the shim distinguishes the flood from ordinary wallet traffic: the - requests are byte-identical to what a wallet sends. -- **The operator can run this against their own shim**, from inside their own - network, deniably — so it is available to adversary #1 in the threat model, not - only to an outsider. -- What bounds it: it must be sustained, and harm (2)'s *destruction* (as opposed - to delay) additionally needs either a wallet expiry near the librustzcash - 40-block default or a driver teardown to coincide. For ZIP 318 migrations, whose - expiry is 30-60 days, the expiry route does not apply and the teardown route - does — and there the wallet's recovery clock is 30-60 days. - -> **CORRECTION 2026-08-18 (validation of the cited file — SUPERSEDES the sentence above).** -> The wallet does **not** wait for expiry. Both official Zcash light-wallet SDKs -> automatically resubmit a sent-but-unmined transaction for as long as it remains -> unexpired — the Android SDK at the head of every ~20 s sync loop and after every -> processed block batch (`CompactBlockProcessor.kt:573,615,723`; selection -> `mined_height IS NULL AND expiry_height > ?`), the iOS SDK at most once per 300 s -> (`TxResubmitter.swift:8-15`, `TransactionDao.swift:218-228`) — and the hub's -> payload-hash dedup makes the resend free. The wallet's non-confirmation signal comes -> from compact-block scanning, which the shim does not intercept (`proxy.rs:1068-1074`). -> Expiry is therefore the **retry horizon**, not the wait: ~50 minutes for the ZIP 203 -> default traffic the shim also diverts, 30–60 days for a ZIP 318 migration. A -> *transient* loss self-heals within minutes; only a loss condition that **outlives the -> horizon** destroys the submission permanently — which is exactly what this issue's -> condition does, so this issue's severity is unaffected. Do not write "the user waits -> 30 to 60 days" in the report. Full refutation and the replacement paragraph: -> `issues/invalid/zip318-canonical-expiry-is-the-only-recovery-clock-and-a-lost-migration-freezes-the-users-notes-for-30-to-60-days.md`. - - -## Impact on Users - -`intercept.rs` fails **closed** on the lookup path and on a failed dispatch, which -is the right choice and means this is not directly a leak. What it is, is a cheap, -remote, unauthenticated switch that turns the product off, or worse, turns it into -a liar: - -- **A user trying to migrate legacy Orchard funds cannot** — either they are told - `UNAVAILABLE` (harm 3) or, worse, they are told **success with a txid** for a - transaction that will reach no mempool (harm 2). The shim keeps no per-migration - state by design (`lib.rs:30-35`), the ack carrying the hub's refusal is - discarded at construction (`nym.rs:652`), and confirmation tracking does not - exist. The user's wallet holds the notes as pending-spent until expiry. -- **Every `GetTransaction` for every wallet behind that shim fails closed** - (`intercept.rs:315-324`), so wallets cannot fetch full transaction data or - confirm anything. For a light wallet this presents as sync failure, not as one - missing screen. -- **The realistic user response is the privacy loss.** A user who switches to a - non-zeronym indexer broadcasts their Orchard-touching transaction directly, - joining their IP to it on the permanent public chain — the exact outcome the - product exists to prevent, and the attacker chooses when. -- **`/healthz` answers 200 throughout**, so the operator's monitoring does not - fire. This is precisely the "dead-client case stayed invisible" failure mode - `MixnetStatus` was introduced to eliminate, and it is not covered: the client is - connected, it is drowning. -- **The shim's memory grows while this runs.** Because the SDK stores a whole - prepared message per tick and emits one packet per tick - (`real_traffic_stream.rs:443-478`), and fragments are fully-built sphinx packets - retained together with a clone for retransmission - (`message_handler.rs:526-556`), a sustained flood adds roughly 60 packets - (~120 KiB) of retained buffer per absorbed lookup against a `memory_mb = 2048` - enclave. **[CORRECTED 2026-08-18 by the G15/G16 global auditor — see the marked - CORRECTION block at the end of the Technical Details section. The retained-memory - statement above is upheld. The "self-amplification" statement that stood here — - that `insert_pending_acks` arms retransmission timers before `forward_messages`, - so timers expire on packets still queued — is WRONG and has been struck: the SDK - starts the timer from `SentNotificationListener`, after emission, for exactly this - reason.]** The duplicate-fragment problem the project measured on the hub - (15-25 duplicate fragments per message, `hub/src/nym_driver.rs:187-201`) is real - and the shim has **no** `ack_wait_addition` equivalent, but its cause is the - enclave's ack round trip exceeding the SDK's default timer, not queue-driven - early firing. A shim OOM remains a realistic endpoint of a sustained flood - through the retained-fragment path alone, which destroys everything in the buffer - including acknowledged submits. -- Because the shim's `SendTransaction` path fails *open* (success) while its - `GetTransaction` path fails *closed*, the user's experience during an attack is - inverted from the truth: sends appear to work and lookups appear broken. - -## Technical Details / Code Analysis - -**The whole of the shim's admission control on this path** (`intercept.rs:249-295`) -is a 1 KiB body cap, a decodable `TxFilter`, and `filter.hash.len() == 32`. The -comment at `:272-273` — *"Validate the filter locally … so a bad filter never -becomes a hub round trip"* — is true and is the only throttle present; a *good* -filter always becomes a hub round trip, and 32 random bytes are always a good -filter. - -**The shared channel and the asymmetric budgets.** `main.rs:335-336`: - -```rust -335 let (req_tx, req_rx) = mpsc::channel(32); -336 let (out_tx, out_rx) = mpsc::channel(8); -``` - -`req_tx` becomes `NymHandle.requests` and is used by **both** operations. Submit -(`nym.rs:639`, `:660`, `:685-689`): - -```rust -639 let deadline = tokio::time::Instant::now() + self.dispatch_timeout; // 5 s -660 match tokio::time::timeout_at(deadline, self.requests.send(request)).await { -661 Ok(Ok(())) => dispatched += 1, -675 Ok(Err(_)) | Err(_) => break, -... -685 if dispatched > 0 { Ok(()) } else { Err(NymError::TransportGone) } -``` - -Lookup (`nym.rs:758-759`, inside `each_target`): - -```rust -758 let deadline = tokio::time::Instant::now() + self.timeout; // 90 s -759 match tokio::time::timeout_at(deadline, self.requests.send(request)).await { -``` - -tokio's bounded `mpsc` acquires a permit from an internal fair, intrusive-FIFO -semaphore, so a submit's `send` future is queued behind every lookup already -waiting and cannot overtake them. Eighteen-to-one patience on a FIFO is the whole -of harm (3). - -**What the drain rate actually is — and it is not the emission rate.** This is the -correction that fixes the filed arithmetic. `correlate` reserves an outbound -permit before accepting each request (`nym.rs:851-880`), the driver holds one send -in flight (`nym_driver.rs:362-375`), and `send_frame` awaits -`sender.send_message(...)` (`:608-623`). But that await completes when the **SDK -accepts** the message, not when it is emitted: - -- `send_message` → `ClientInput::send` → `mpsc::channel::(1)` - (`base_client/mod.rs:1013`); -- the input listener prepares fragments and awaits - `real_message_sender.send(batch)` on a channel of capacity **8** - (`real_messages_control/mod.rs:150`); -- `OutQueueControl::poll_poisson` polls that receiver **once per Poisson tick** - and stores the whole batch, emitting **one packet** per tick - (`real_traffic_stream.rs:443-478`), into a `TransmissionBuffer` with no size - limit and no byte budget (`transmission_buffer.rs:39-49`). - -So the shim absorbs of order **one whole request per packet-tick** — 8.33/s at the -crate's throttled floor, 50/s at the SDK's 20 ms default — while emitting **one -lookup per 61 ticks** (0.14-0.8/s). The 32-slot channel therefore does *not* -represent tens of seconds of backlog; it drains quickly into an unbounded internal -queue. That is what makes harm (2) automatic and harm (3) a concurrency question: - -- to be behind more than 5 s of drain, a submit needs `> 5R` waiters ahead of it: - ~42 at 8.33/s, ~250 at 50/s; -- to *sustain* a waiter list at all, arrivals must at least match `R`. Each - attacker request lives ≤ 90 s (`REQUEST_TIMEOUT` covers acceptance and reply - together, `nym.rs:758-770`), so `W` concurrent requests deliver `W/90` arrivals - per second and the condition is `W > 90R`: **~750 concurrent requests at the - throttled rate, ~4,500 at the unthrottled default.** - -Both are trivially affordable — ~100 bytes each, ≤90 s each, no connection cap -(`proxy.rs:470-541`) and no advertised `max_concurrent_streams` -(`proxy.rs:596-610`) — but they are two orders of magnitude more than the "few -tens of concurrent requests" originally filed, and the corrected figure is what a -reviewer should check the fix against. - -**The wallet is answered before anything leaves the process.** `hub.rs:228-249`: - -```rust - HubTransport::Nym(handle) => match handle.submit(tx_bytes).await { - // ... There is no Refused arm: the hub's verdict is a full round - // trip away and is deliberately not waited for ... - Ok(()) => Ok(Submit::Accepted { txid: crate::nym::local_txid(tx_bytes) }), -``` - -and what sits behind that hand-off, in the crate's own words -(`nym_driver.rs:418-432`): - -> *"be clear about what disconnect() DOES discard: everything the SDK still holds -> internally — its one-slot input, an 8-deep batch channel, and an **unbounded -> transmission buffer drained at the throttled rate**. There is no -> drain-then-disconnect in the SDK. Frames in there may include SUBMITS ALREADY -> ANSWERED SUCCESS to a wallet, and nothing upstream can protect them: a submit's -> waiter is swept the moment it is dispatched, so the supervisor's inflight count -> never sees it."* - -**The emission cost, from the crate's own constants** (`nym.rs:1090-1130`): - -```rust - const PACKET_BYTES: usize = 2 * 1024; - /// The client's own floor on sending, `MAX_DELAY_MULTIPLIER` (6) times the - /// 20 ms default `message_sending_average_delay`. - const THROTTLED_PACKETS_PER_SEC: f64 = 1000.0 / 120.0; -``` - -`packets(LOOKUP_BYTES) + LOOKUP_REPLY_SURBS = 1 + 60 = 61` packets ⇒ **7.3 s per -lookup** at the floor (1.2 s unthrottled) — for one ~100-byte request. - -**The failure arms.** `intercept.rs:204-214` (submit) and `:315-324` (lookup) both -answer `GRPC_UNAVAILABLE` and never fall back to the operator's indexer, which is -correct and is why this is a denial/integrity issue rather than a leak. And the -health endpoint that does not notice (`nym.rs:307-309`): - -```rust -307 pub fn is_healthy(&self) -> bool { -308 !self.0.configured.load(Ordering::Relaxed) || self.0.connected.load(Ordering::Relaxed) -309 } -``` - -## Recommendations - -1. **Give submits reserved capacity that lookups can never consume.** They - contend for one transport but have opposite value: a lost submit is a user who - cannot migrate and may be told they did, a lost lookup is a retry. A separate - small channel for submits, or a two-priority queue, means no lookup backlog can - take the diversion path down. This is the single fix that closes harm (3) here - **and** the same harm in - `junk-sendtransaction-flood-consumes-the-shims-whole-mixnet-egress-and-converts-acknowledged-migrations-into-silent-loss.md`, - which that issue's own remediation does not address. -2. **Cap concurrent outstanding lookups in the shim**, the way `hub/src/nym.rs` - caps them with a semaphore, and answer over-cap lookups with - `RESOURCE_EXHAUSTED` immediately rather than queueing them. A shim serving a - normal wallet population needs single-digit concurrency here. Refusing quickly - is fail-closed and costs the transport nothing. -3. **Bound the accepted backlog by what the transport can actually deliver, and - stop answering `error_code 0` for work that is only queued.** The design's own - `MAX_DELIVERY_LAG` (6 blocks, `hub/src/batcher.rs:46-48`) is asserted at - startup as if it were a constant; this issue makes it attacker-controlled. Once - the depth of `requests` + `out_frames` + the SDK's lane queue exceeds that - budget, fail closed so the wallet retries rather than accepting a migration - that will arrive too late. A visible `UNAVAILABLE` is strictly better for the - user than a silent success. -4. **Rate-limit `GetTransaction` per source address / per connection** before - `diversion.hub.get_transaction` is called. The peer address is already - available at `proxy.rs:488`, and the shim terminates TLS itself. Also set - `max_concurrent_streams` on the h2 server and cap concurrent connections in - the accept loop; both are one line each and neither refuses anything a wallet - does. -5. **Make `/healthz` (and `/nym-status`) reflect the condition that actually - matters** — whether a submit can currently be dispatched, and how deep the - emission backlog is — rather than only whether the client is connected. As - written, the one failure mode `MixnetStatus` was added for is reachable with - the client up. -6. **Make the submit budget no smaller than the lookup budget**, or drop the - asymmetry entirely, so an adversary's lookups cannot out-wait a user's - migration on a shared FIFO. This is a stopgap: item 1 is the real fix. -7. **Give the shim the `ack_wait_addition` knob the hub received** - (`hub/src/nym_driver.rs:198-217`), and cap the shim's retransmissions. The - rationale originally given here (*"because `insert_pending_acks` runs before - `forward_messages`, a deep backlog retransmits packets that were never emitted"*) - is **withdrawn as incorrect** — see the CORRECTION block below. The - recommendation itself stands on the correct mechanism, which is that the - enclave's measured ack round trip exceeds the SDK's default - `1.5 x expected_delay + 1500 ms` and that the shim's sends carry - `max_retransmissions: None`. Both are now filed separately, with full evidence, - as - `plausible/shim-mixnet-client-has-neither-retransmission-bound-the-hub-has-so-an-unacked-frame-retransmits-forever.md`. - -**CORRECTION (added 2026-08-18, G15/G16 global auditor; derived from the pinned SDK -tree at `451c2aa`, which is available locally).** Three passages in this file assert -that `insert_pending_acks` running before `forward_messages` -(`message_handler.rs:554-555`) arms retransmission timers on packets that have not -been emitted, and that this self-amplifies. The SDK does not behave that way: - -- `ActionController::handle_insert` inserts each `PendingAcknowledgement` with - `queue_key = None` and starts **no** timer - (`common/client-core/src/client/real_messages_control/acknowledgement_control/action_controller.rs:122-137`). -- A timer is created only by `Action::StartTimer` → `handle_start_timer` - (`action_controller.rs:139-162`), whose enum doc says *"Initiated by - `SentNotificationListener`"* (`:38-42`). -- `SentNotificationListener`'s own module doc states the purpose verbatim: *"Module - responsible for starting up retransmission timers. It is required because when we - send our packet to the `real traffic stream` controlled by a poisson timer, - there's no guarantee the message will be sent immediately, so we might - accidentally fire retransmission way quicker than we should have"* - (`sent_notification_listener.rs:10-13`), and it fires on the sent notification - (`:30-38`). - -So the ordering of `insert_pending_acks` and `forward_messages` is **harmless**, and -the flood does not self-amplify through that path. What survives unchanged is the -*retention* consequence: `try_split_and_send_non_reply_message` clones every fragment -before preparing it — *"we need to clone it because we need to keep it in memory in -case we had to retransmit it"* (`message_handler.rs:530-532`) — and holds the clone -in the pending-ack map until the ack arrives (`:546-547`), which is the ~120 KiB per -absorbed lookup this issue relies on. **The severity and the confirmed verdict are -not affected**; only the mechanism attributed to one contributing leg is. - -**Coordinator open item 7b is answered by this correction.** Its premise ("arms -retransmission timers before `forward_messages`, self-amplifying") does not hold. The -shim-OOM leg should **stay inside this issue** — it shares this issue's trigger, -attacker and remediation family — while the missing `ack_wait_addition` knob and the -absent retransmission cap are filed as their own issue (linked above) because they -degrade every deployed shim with no attacker present and have a distinct one-line fix. - -## Validation Information - -**Verdict: CONFIRMED. Severity: High (as filed), for reasons partly different -from those given in the filing.** The mechanism was traced end to end through the -shim and through the **pinned SDK tree at -`451c2aa3692fc4dc00041b74a352d4158176d9c0`**, which is present locally. Three -corrections were applied; one of them changes which harm carries the severity. - -### What was verified - -| Claim | Verified at | -|---|---| -| The listener is internet-reachable and unauthenticated; routing is a pure function of the path | `caution.hcl.tmpl` (`ZIS_LISTEN = 0.0.0.0:8083`, `ingress 0.0.0.0/0`), `proxy.rs:743-748` | -| No connection cap, no `max_concurrent_streams`, only window sizes | `proxy.rs:470-541`, `:596-610` | -| With a hub configured, every `GetTransaction` goes to the hub and none to the operator | `intercept.rs:229-247` | -| 32 random bytes pass every check | `intercept.rs:274-293` | -| Lookups and submits share one 32-slot channel | `main.rs:335`, used by both `nym.rs:660` and `:759` | -| 90 s vs 5 s asymmetry, and tokio `mpsc` is a fair FIFO semaphore | `nym.rs:71`, `:80`, `:639`, `:758` | -| `dispatched == 0` → `TransportGone` → `UNAVAILABLE` at the wallet | `nym.rs:685-689`, `intercept.rs:204-214` | -| The wallet is told success at hand-off, with no `Refused` arm | `hub.rs:228-249` | -| The SDK's buffer behind that hand-off is unbounded, and the crate knows it | `nym_driver.rs:416-436`; `transmission_buffer.rs:39-49` | -| 61 packets per lookup at 8.33 packets/s = 7.3 s of the shim's whole egress | `nym.rs:104`, `:1090-1130`, cross-checked against `config-types/src/lib.rs:25`, `:443` | -| `/healthz` is `configured && connected` and stays green | `nym.rs:307-309` | - -### Correction 1 — the drain-rate model was wrong, and the concurrency figure with it - -The filing said the 32-slot channel "drains at mixnet speed, not at request -speed", so "a handful of concurrent lookups is enough to keep it saturated -indefinitely" and "a few tens of concurrent requests suffice". **That is not how -the SDK behaves.** `send_message` returns when the SDK *accepts* the message: -capacity-1 `InputMessage` channel (`base_client/mod.rs:1013`) → 8-slot batch -channel (`real_messages_control/mod.rs:150`) → `poll_poisson`, which stores one -whole batch and emits one packet per tick (`real_traffic_stream.rs:443-478`) into -an uncapped `TransmissionBuffer`. The shim therefore absorbs ~8-50 requests per -second while emitting 0.14-0.8 lookups per second. - -The corrected thresholds are derived in Technical Details: `> 5R` waiters to -outlast a submit's 5 s budget (~42 at the throttled rate, ~250 unthrottled), and -`W > 90R` concurrent live requests to sustain that queue (**~750 to ~4,500**). -This is exactly the check that must be done before a claimed rate is believed — -a rate that merely matches the drain rate denies nothing. Here the corrected -figure is still trivially affordable (~100 bytes per request, ≤90 s each, no cap -on connections or streams anywhere), so the finding stands; but the filed "a few -tens" was off by two orders of magnitude and would have let a reviewer -under-specify the fix. - -### Correction 2 — the severity is carried by the silent harm, not the loud one - -The filing's headline harm is the loud one: a submit fails closed and the wallet -sees `UNAVAILABLE`. That harm is real, but it is the **most expensive to produce -and the least damaging**, because a wallet that is told UNAVAILABLE retries and -nothing is lost. - -The harm that carries the severity needs almost no concurrency at all. At any -sustained rate above roughly one request per 7.3 s the shim's emission backlog -grows without bound inside the SDK, and then: - -- every wallet's `GetTransaction` through that shim times out at 90 s and fails - closed — certain, and free; and -- a migration that *is* dispatched is answered `error_code 0` with a txid - (`hub.rs:228-249`) and then sits in an unbounded buffer the shim cannot see, - to be refused `ExpiryTooTight` on arrival (into a discarded ack) or destroyed - by the next rebuild, rotation, redeploy or SIGTERM — every one of which - discards the buffer with no drain, as `nym_driver.rs:416-436` states. - -So the attacker chooses between a loud denial and a silent destruction of a -migration the user believes is spent, and the silent one is cheaper. The body has -been restructured accordingly. - -### Correction 3 — added consequences that were not in the filing - -- **Shim memory growth toward OOM.** Fragments are fully-built sphinx packets - retained with a clone for retransmission (`message_handler.rs:526-556`), so - ~120 KiB of buffer accrues per absorbed lookup against `memory_mb = 2048`. -- ~~**Self-amplification.** `insert_pending_acks` is called *before* - `forward_messages` (`message_handler.rs:553-555`), so retransmission timers - expire on packets still queued.~~ **STRUCK 2026-08-18 — the premise is false; see - the CORRECTION block at the end of Technical Details. Timers start from - `SentNotificationListener`, after emission.** What remains true and is retained: - the project measured 15-25 duplicate fragments per message on the hub - (`hub/src/nym_driver.rs:187-201`) and mitigated it with - `ZIH_ACK_WAIT_ADDITION_MS`; the shim has no equivalent and sets no `DebugConfig` - on a production path. The cause is the ack round trip exceeding the SDK's default - timer, not early firing. -- Recommendations 3, 4 (h2/connection caps) and 7 are new and follow from these. - -The filing's amplification framing ("~60-byte request buys ~64 KiB … well over -1000×") is arithmetically right but is the least useful way to state it; the -operative figure is **7.3 seconds of a single, serialised, non-substitutable -resource per ~100-byte request**, and the body now leads with that. - -### `docs/AVOIDING-FALSE-POSITIVES.md` §5 applied - -§5's own statement of the real-vulnerability shape is *"amplification attacks -where small input causes disproportionate resource use"*, and its contrasting real -issues are *"1 KB request causing 1 GB memory allocation"* and *"single connection -consuming unbounded resources"*. This is both: ~100 bytes buys 61 sphinx packets -and ~120 KiB of retained enclave buffer, and the resource consumed is not -elastic — it is one throttled mixnet client that also carries every migration. - -*What resources would the attacker need?* A few hundred bytes per minute for the -lookup outage and the silent-loss harm; a few KB/s and several hundred concurrent -h2 streams for the loud one. No credential, no wallet, no funds, no Nym client. - -*What would stop them?* Nothing in the target and nothing in the deployment: no -authentication, no rate limit, no per-source accounting, no connection cap, no -`max_concurrent_streams`, no lookup concurrency bound, and a `/healthz` that stays -green throughout. - -*Why §5 does not cap this at Medium.* §5 would normally cap a throughput attack, -and this issue is graded above that cap for one reason that was checked rather -than assumed: the shim answers `error_code 0` at an in-process channel send -(`hub.rs:228-249`), so the denial is not visible **as** a denial. Note the -discipline this requires — that multiplier is not allowed to carry the grade on -its own. Reachability stands independently: an unauthenticated internet request -to a public DNS name becomes a metered 61-packet mixnet emission at a fixed cost -that does not depend on the request's size, and the corrected concurrency -arithmetic above shows the queue really is deniable rather than merely matched. - -### Severity: High, and why it is not a duplicate - -*Impact:* for an operator the attacker chooses, the privacy-critical divert path -and the whole `GetTransaction` path are disabled, and the `SendTransaction` -failure mode is a **false success** for a transaction that may reach no mempool, -in a pool NU6.3 has closed to new value. *Likelihood:* an unauthenticated request -to a public DNS name, at a few hundred bytes per minute, from anywhere, with no -detection surface. *Why not Critical:* no funds are stolen and no key or -plaintext is disclosed; the destruction (as opposed to delay) of an acknowledged -migration needs either a near-default wallet expiry or a teardown to coincide; -and a ZIP 318 migration's submission is destroyed rather than its funds lost (see -the CORRECTION above: the wallet keeps resubmitting for the whole expiry window). - -*Not a duplicate of -`junk-sendtransaction-flood-…-silent-loss.md` (High).* That issue reaches the same -shared resource through `SendTransaction`, and its remediation — refusing -zero-length/`Unparseable` bodies — does **nothing** here, because these requests -are perfectly well-formed lookups that a wallet also sends. Conversely this -issue's fix (reserved submit capacity) is the one that closes both. The two are -siblings on one root cause — a single 32-slot channel and one serialised emitter -shared by both operations — and both must be fixed. - -*Relationship to `hub-nym-lookup-flood-starves-gettransaction-fleet-wide.md` -(Medium).* The two were graded against each other, not in isolation, and this one -is the more severe **despite the smaller blast radius**. That issue is fleet-wide -but needs sustained free Nym clients, denies only the lookup path, fails loudly at -the wallet, destroys nothing, and heals when the backlog drains. This one needs no -mixnet capability at all, lands on the *submit* path where the shim's -success-at-hand-off converts denial into silent destruction of an acknowledged -migration, additionally consumes the hub's fleet-wide emitter through the shim -(the composition recorded as G5 §3.1, reachable by an attacker with no Nym -experience), and grows the shim's memory while it runs. This ordering upholds G5 -§2, which ranked the `GetTransaction` flood at a shim (#1) above the direct hub -lookup floods (#4, #5). - -DO NOT DEVIATE FROM THIS TEMPLATE (`docs/SECURITY-ISSUE-TEMPLATE.md`). -LEAVE THIS MESSAGE IN PLACE SO THAT ALL AGENTS KNOW NOT TO DEVIATE FROM THIS TEMPLATE (`docs/SECURITY-ISSUE-TEMPLATE.md`). diff --git a/zeronym-22aa9851caf68-high-medium/high/hub-http-lookup-path-has-no-concurrency-bound.md b/zeronym-22aa9851caf68-high-medium/high/hub-http-lookup-path-has-no-concurrency-bound.md deleted file mode 100644 index bd338766..00000000 --- a/zeronym-22aa9851caf68-high-medium/high/hub-http-lookup-path-has-no-concurrency-bound.md +++ /dev/null @@ -1,448 +0,0 @@ -# The hub's unauthenticated HTTP lookup path has no concurrency bound, so a ~100-byte request buys a fresh TCP+TLS+HTTP/2 dial to the indexer — the exact flood the sibling mixnet path caps at 64 - -**Severity**: High -**Validation Status**: Confirmed -**Location**: `audit-target/zeronym/hub/src/server.rs:442-445` (the ungated `TRANSACTION_PATH` arm of `handle`), `:487-505` (`lookup`), `:296-322` (`Hub::lookup`), `:354-408` (`serve`, unbounded `tokio::spawn` per connection); the work it triggers is `audit-target/zeronym/hub/src/chain.rs:221-266` (`get_transaction`) and `:300-337` (`unary_inner`, a fresh `TcpStream::connect` + TLS + h2 handshake **per call**, `:310`); the bound that exists on the equivalent mixnet arm is `audit-target/zeronym/hub/src/nym.rs:38-54` (`MAX_CONCURRENT_LOOKUPS = 64`) and `:171-196` (`try_acquire_owned`, drop-on-full); the collateral victims are `audit-target/zeronym/hub/src/batcher.rs:290-296` (the tip poll) and `:341-355` (`flush`); the refusal that follows is `audit-target/zeronym/hub/src/server.rs:248-254` with `audit-target/zeronym/hub/src/batcher.rs:62` (`TIP_STALE_AFTER`) and `:205-208` (`is_stale`); reachability from `audit-target/zeronym/hub/deploy/caution/caution.hcl.tmpl:51-55` and `:97-105` (`ingress 0.0.0.0/0` on 8083 and the in-enclave Caddy that maps 443 onto it), `:39-40` (`cpu = 2`, `memory_mb = 2048`) -**Found by agent:** Local (file audit of `hub/src/server.rs`); validated 2026-08-18 -**In scope of audit?** Yes - -## Description - -`hub/src/nym.rs:38-54` states the rule for the hub's lookup arm and the reason -for it, in the code's own words: - -> How many lookups may be in flight at once: dialling the operator's indexer, -> framing the reply, or waiting to hand it to the driver. ... Generous enough -> that honest polling never queues behind it, **small enough that a flood cannot -> open unbounded connections** or park an unbounded pile of 64 KiB reply frames -> behind a slow driver. -> -> ... When every slot is held the lookup is **dropped, not parked**: parking it -> in a task is the unbounded pile this bound exists to prevent. - -That bound is implemented on the **mixnet** ingress only. The **HTTP** ingress -reaches the identical core, `Hub::lookup`, with **no semaphore, no rate limit, -no connection cap and no per-peer accounting**. `handle` routes -`TRANSACTION_PATH` unconditionally — it is not behind `http_submit`, unlike -`POST /` — and `serve` spawns one unbounded `tokio::spawn` per accepted -connection. - -Each lookup that misses the in-RAM queue calls `ChainClient::get_transaction`, -which for **every** configured endpoint runs `unary_inner` -(`hub/src/chain.rs:310`): - -```rust - let stream = TcpStream::connect(addr).await?; - stream.set_nodelay(true)?; -``` - -There is no connection pool anywhere in `chain.rs`; `ChainClient` holds only -`endpoints` and `tls` (`chain.rs:123-129`). Every single call performs a fresh -TCP connect, a full TLS handshake with certificate-chain verification -(`ZIH_INDEXER_TLS` is a required production setting — -`caution.hcl.tmpl:126-132`, "WITHOUT THIS THE HOP IS PLAINTEXT"), and a fresh -HTTP/2 handshake, under a 10 s budget (`RPC_TIMEOUT`, `chain.rs:48`). - -So the cost ratio on a 2 vCPU / 2 GB enclave is: - -| Attacker pays | Hub pays, per request | -|---|---| -| ~100 bytes on an already-open connection: `POST /transaction` with a 32-byte random body | 1 outbound socket, 1 ephemeral source port held for the connection plus 60 s of `TIME_WAIT`, 1 TCP connect, **1 full TLS handshake with X.509 chain verification**, 1 HTTP/2 handshake, 1 gRPC round trip, up to 10 s of holding — *per configured indexer endpoint* | - -A random 32-byte body always misses the queue (`Queue::find_by_txid` compares -64-character hex txid strings, `queue.rs:328-348`), so the attacker chooses with -certainty that the expensive branch runs. - -**The damage is not the flood itself; it is that the hub's own outbound path -shares every ceiling the flood reaches.** `batcher::run` polls the tip through -the same `ChainClient` (`batcher.rs:290`) and `flush` publishes through it -(`batcher.rs:341-355`). Whichever limit binds first — the process's descriptor -limit, the ephemeral port range for the single indexer destination, or the two -vCPUs — the batcher's `TcpStream::connect` fails at the same moment the -attacker's lookups do. Fifteen minutes of failed tip polls make `is_stale()` -true (`batcher.rs:62`, `:205-208`), after which `Hub::admit` refuses **every** -submission on **both** ingress paths (`server.rs:248-254`), and on the deployed -mixnet transport a refusal reaches nobody: the shim answered the wallet -`error_code 0` at mixnet hand-off (`shim/src/hub.rs:231-240`). - -## Attack Scenario and Steps - -Attacker is anyone on the internet. The hub's HTTP listener is reachable at -`https:///transaction`: `caution.hcl.tmpl:97-105` declares -`http { domain, port = 8083, e2e_encryption { mode = "tls" } }`, which the -platform implements as an **in-enclave Caddy terminating TLS on 443 and -reverse-proxying to `127.0.0.1:8083`** (settled in PROGRESS.md item 5/6). Caddy -ships with no request-rate limit and no upstream concurrency cap, so it relays -whatever arrives. - -1. Open one HTTP/1.1 or HTTP/2 connection to the hub's domain, or a few hundred. - Nothing in `serve` caps the number. -2. Issue `POST /transaction` with a random 32-byte body, back to back, at - whatever concurrency is wanted. Each request is ~100 bytes on the wire and - costs the attacker no CPU. Caddy opens one upstream connection per concurrent - request, so hub-side concurrency tracks the attacker's concurrency exactly. -3. Each request misses the queue and produces one fresh TCP + TLS + HTTP/2 + - gRPC dial from the enclave to each configured indexer. -4. One of three ceilings is reached, and all three are shared with the hub's own - work: - - **Ephemeral source ports.** Every call dials the *same* destination - (the shipped deploy configures one indexer), and the hub is the side that - closes, so each completed call parks a local port in `TIME_WAIT` for ~60 s. - On stock Linux settings (`ip_local_port_range` 32768–60999 ≈ 28,232 ports, - `tcp_tw_reuse = 2`, i.e. loopback only) a sustained rate above roughly - **470 connections per second** exhausts the range for that destination and - `connect()` returns `EADDRNOTAVAIL` — for the batcher as well. 470 requests - per second of ~100-byte bodies is about 50 KB/s of attacker bandwidth. - - **File descriptors.** Each in-flight lookup holds one inbound descriptor - (from Caddy) and one outbound per endpoint. The code already anticipates - this ceiling: `serve` has an explicit `is_fd_exhaustion` arm that logs a - warning and sleeps 100 ms (`server.rs:341-352`, `:381-385`), which keeps - the listener alive while the hub is failing *its own* outbound connects. - - **CPU.** `cpu = 2`. Each request forces one client-side TLS handshake - including certificate-chain verification, on the same two workers that run - the cadence loop, the mixnet driver and the in-enclave Caddy. -5. The hub's tip poll fails. `TipTracker::observe` only refreshes `last_advance` - when a height *advances*, so after `TIP_STALE_AFTER = 15 min` — 30 - consecutive 30 s `POLL_INTERVAL` ticks — `is_stale()` becomes true. -6. `Hub::admit` then refuses every submission fleet-wide: - -```rust -// hub/src/server.rs:248-254 - pub fn admit(&self, tx_bytes: &[u8]) -> Result, Refusal> { - if self.tip.is_stale() { - return Err(Refusal::TipStale); - } -``` - -7. Because submit over Nym is dispatch-only (`shim/src/hub.rs:231-240`, and the - project's own test `shim/tests/divert_nym.rs:235-267` asserts - `error_code == 0` against a refusing hub), the wallet was already told the - migration succeeded. The refusal reaches nobody. The migration exists - nowhere. - -Steps 1–4 are certain from the code. Step 5 requires the flood to be sustained -for fifteen minutes, which is trivial. **Even without reaching step 5**, steps -1–4 alone contend with `flush`, which runs inline on the cadence task -(`batcher.rs:290-313`): a `broadcast_batch` that cannot get sockets returns -`Retryable` for the whole batch and requeues it into a later window, so an -anonymous outsider gets a lever on *when* a batch is published. - -**Attack Requirements and Assumptions:** - -- **Network access only.** No credential, no txid, no mixnet position, no Zcash - knowledge, no funds. The lookup path is not gated by `http_submit`, so it is - open in every deployment the repository ships, including the intended - mixnet-only one. -- **What makes it realistic:** the request is content-free random bytes; the - amplification is a full TLS handshake, a socket and an ephemeral port for - ~100 bytes; the target has two vCPUs; and the identical arm on the mixnet - transport is bounded at 64 with drop-on-full precisely because the authors - identified this attack there. -- **What limits it:** the flood is noisy and originates from identifiable - addresses — but there is no per-IP block list here to evade, because there is - no per-IP anything, and no operator-visible signal (see Impact). The exact - concurrency needed depends on the enclave's `RLIMIT_NOFILE` and - `ip_local_port_range`, neither of which could be measured in this environment; - the Containerfile and the manifest set neither, so both are the platform - defaults. The *direction* is not in doubt: nothing in the application bounds - concurrency, so the attacker reaches whichever ceiling is lowest, and every - one of them is shared with the batcher. - -## Impact on Users - -- **Silent loss of transactions users were told had been sent.** Under - dispatch-only submit the wallet already saw `error_code 0`. A `TipStale` - refusal, or a repeatedly-failing flush, destroys the migration with no error - surfaced anywhere the user can see. This is the same terminal state as the - confirmed High `hub-queue-unauthenticated-fill-silently-destroys-migrations.md`, - reached over clearnet at a small fraction of the cost and denying **100 %** of - submissions rather than a bandwidth-proportional fraction. -- **A privacy failure, not only an availability one.** A hub that refuses - admission makes every shim's divert fail closed; wallets retry, and a user - whose transaction will not send eventually points their wallet at a different, - unprotected light-wallet server and broadcasts the Orchard-touching - transaction in the clear. That is precisely the leak the product exists to - prevent, and the attacker chooses the moment it happens. -- **Batch timing becomes an outsider's input.** Making a flush's - `broadcast_batch` fail defers a whole batch to a later window. Choosing when a - batch is published, and keeping a hub from accepting during a chosen interval, - is an anonymity-set attack: it shrinks the set any given migration is mixed - with. -- **Nothing reports the condition.** `GET /healthz` returns 200 unconditionally - (`server.rs:450-453`); `GET /nym-status` reports only mixnet-client lifecycle; - the tip-poll failure is logged at `debug!` (`batcher.rs:293`) and - `caution.hcl.tmpl:139` leaves `RUST_LOG` at the default `info`, so it is not - even emitted. An operator monitoring the hub sees green throughout. (Filed - separately as `hub-health-surface-blind-to-the-states-that-destroy-migrations.md`.) - -## Technical Details / Code Analysis - -**The route is unconditional and the handler is unbounded.** - -```rust -// hub/src/server.rs:437-448 - match req.uri().path() { - SUBMIT_PATH if options.http_submit => match method { - Method::POST => submit(req, hub).await, - _ => Ok(text(StatusCode::METHOD_NOT_ALLOWED, "POST only")), - }, - TRANSACTION_PATH => match method { - Method::POST => lookup(req, hub).await, - _ => Ok(text(StatusCode::METHOD_NOT_ALLOWED, "POST only")), - }, -``` - -```rust -// hub/src/server.rs:487-505 -async fn lookup(req: Request, hub: Hub) -> Result>, Infallible> { - let collected = match Limited::new(req.into_body(), MAX_LOOKUP_BYTES) - .collect() - .await - { - Ok(collected) => collected, - Err(_) => return Ok(text(StatusCode::PAYLOAD_TOO_LARGE, "lookup key too large")), - }; - let wire_hash = collected.to_bytes(); - if wire_hash.is_empty() { - return Ok(text(StatusCode::BAD_REQUEST, "empty lookup key")); - } - - match hub.lookup(&wire_hash).await { -``` - -`MAX_LOOKUP_BYTES = 64` bounds the *body*, which is not the resource under -attack — it is what makes the attack cheap. There is no other bound on this -path. - -**The accept loop spawns without limit.** - -```rust -// hub/src/server.rs:390-407 - tokio::spawn(async move { - let io = TokioIo::new(stream); - if let Err(err) = http1::Builder::new() - .timer(TokioTimer::new()) - .serve_connection( - io, - service_fn(move |req| handle(req, hub.clone(), options.clone())), - ) - .await -``` - -The installed `TokioTimer` correctly re-enables hyper's 30 s header-read timeout -(the comment at `server.rs:392-396` explains why), but that only bounds *idle* -connections. A connection that keeps sending complete, valid requests is never -throttled, and there is no ceiling on how many such connections exist. - -**A queue miss is guaranteed for random input, and the miss is the expensive -branch.** - -```rust -// hub/src/server.rs:296-303 - pub async fn lookup(&self, wire_hash: &[u8]) -> LookupOutcome { - if let Some(bytes) = self.queue.find_by_txid(wire_hash) { - tracing::debug!(source = "queue", "transaction lookup answered"); - return LookupOutcome::Found { data: bytes, height: 0 }; - } - - match self.chain.get_transaction(wire_hash).await { -``` - -**Each miss is a fresh dial per endpoint.** - -```rust -// hub/src/chain.rs:221-232 - pub async fn get_transaction(&self, wire_hash: &[u8]) -> Result { - let calls = self.endpoints.iter().map(|addr| { - let filter = TxFilter { block: None, index: 0, hash: wire_hash.to_vec() }; - async move { - match self.unary::<_, RawTransaction>(*addr, GET_TRANSACTION, filter).await -``` - -```rust -// hub/src/chain.rs:310-334 - let stream = TcpStream::connect(addr).await?; - stream.set_nodelay(true)?; - ... - let response = match &self.tls { - Some(tls) => { - let stream = tls.connect(addr, stream).await?; - round_trip(stream, request).await? - } -``` - -**The bound that exists on the other transport, and what it really bounds.** - -```rust -// hub/src/nym.rs:54 -const MAX_CONCURRENT_LOOKUPS: usize = 64; -``` - -```rust -// hub/src/nym.rs:171-176 - let permit = match lookups.clone().try_acquire_owned() { - Ok(permit) => permit, - Err(_) => { - tracing::info!( - in_flight = MAX_CONCURRENT_LOOKUPS, -``` - -The permit is taken before the task is spawned and released only after the reply -has been accepted by the driver's channel (`nym.rs:169-206`), so it does bound -**concurrent indexer dials at 64** — which is the resource at issue here. (The -G21 pass established that the same permit does *not* bound mixnet *emission*, -because `outgoing` accepts within seconds; that is a different resource and a -different, already-confirmed finding.) - -`hub/src/nym.rs:17-22` claims the two ingress paths cannot drift: - -> Admission is `crate::server::Hub::admit` and lookup is -> `crate::server::Hub::lookup`, the exact calls the HTTP serving path uses, **so -> the two ingress paths cannot drift.** - -They share the *core* but not the *admission control around it*, and the -admission control is the security property. - -**The collateral victim, and why it is silent.** - -```rust -// hub/src/batcher.rs:290-296 - match chain.tip_height().await { - Ok(height) => tip.observe(height), - Err(err) => { - tracing::debug!(%err, "tip query failed on every node"); - } - } -``` - -```rust -// hub/src/batcher.rs:205-208 - pub fn is_stale(&self) -> bool { - let state = self.read(); - !state.observed || state.last_advance.elapsed() > TIP_STALE_AFTER - } -``` - -## Recommendations - -1. **Bound concurrent indexer dials across *all* ingress.** Give `ChainClient` - (or `Hub`) a single `Arc` and acquire it in `Hub::lookup`, so the - HTTP arm and the mixnet arm draw from one budget. On failure to acquire, - answer `503` immediately rather than parking the request. This is the - smallest change that closes the amplification and is the one the mixnet arm - already models. -2. **Pool or reuse indexer connections in `chain.rs`.** One long-lived HTTP/2 - connection per endpoint removes the per-call handshake, the per-call - descriptor and the per-call ephemeral port in one change, and makes the tip - poll robust under load. This is the single highest-value fix here, and it - also addresses `hub-chain-connection-per-call-fanout-and-flush-memory-amplification.md`. -3. **Gate `POST /transaction` behind a config flag the way `POST /` is gated**, - defaulting off. In the deployed topology the shim uses the mixnet `LookupV1` - path, so the clearnet lookup is transitional exactly like clearnet submit. - This also closes the confirmed - `hub-unauthenticated-pre-publication-transaction-disclosure.md`. -4. **Cap concurrent connections in `serve`** with a semaphore acquired before - `tokio::spawn`. A 2 GB, 2 vCPU enclave on an internet-facing ingress should - not accept unbounded connections. -5. **Reject lookup bodies that are not exactly 32 bytes.** It does not fix the - flood, but it removes a free variant and matches what a `TxFilter.hash` is. -6. **Reserve headroom for the hub's own egress.** Even with (1), the tip poll - and `flush` should not compete with lookups for the last descriptors: give - the batcher its own permit outside the lookup budget. -7. **Surface the condition.** Raise the tip-poll failure above `debug!`, and see - `hub-health-surface-blind-to-the-states-that-destroy-migrations.md`; without - it this attack has no detection signal at all. - -## Validation Information - -**Verdict: CONFIRMED. Severity raised from the filed Medium to High**, on the -same reasoning that put `hub-queue-unauthenticated-fill-silently-destroys-migrations.md` -at High: the attacker's marginal cost is ~100 bytes, the hub's is a socket, a -port and a TLS handshake, and the terminal state is not downtime but silent -destruction of migrations the wallet was told had succeeded. - -### Every mechanical claim re-verified against the target - -| Claim | Verified at | -|---|---| -| `POST /transaction` is routed unconditionally, with no `http_submit` guard and no auth | `hub/src/server.rs:437-448` — the `TRANSACTION_PATH` arm has no `if` | -| Body is capped at 64 bytes, so the request is tiny by construction | `hub/src/server.rs:84-87` (`MAX_LOOKUP_BYTES = 64`), `:488-494` | -| A random 32-byte body always misses the queue | `hub/src/queue.rs:328-348` — `find_by_txid` compares against 64-char hex `Entry.txid` strings | -| A miss dials the indexer once per endpoint | `hub/src/server.rs:302-306` → `hub/src/chain.rs:221-249` | -| No connection pool exists anywhere in `chain.rs` | `ChainClient` fields are `endpoints` + `tls` only (`chain.rs:123-129`); `unary_inner` calls `TcpStream::connect` on every invocation (`:310`) and `round_trip` performs a fresh h2 handshake (`:349-351`), whose sender is dropped on return | -| No concurrency bound, rate limit or per-peer accounting on the HTTP path | `hub/src/server.rs:354-408` — `tokio::spawn` per accepted connection, nothing acquired | -| The mixnet arm *does* bound concurrent dials at 64 | `hub/src/nym.rs:153` (`Semaphore::new(MAX_CONCURRENT_LOOKUPS)`), `:171-206` — permit taken before spawn, dropped only after `outgoing.send()` returns | -| The constant's own rationale names this attack | `hub/src/nym.rs:38-54` — *"small enough that a flood cannot open unbounded connections"* | -| The tip poll and the flush use the same `ChainClient` | `hub/src/main.rs:39` constructs one `Arc`, passed to `batcher::run` at `:78-85` and to `Hub` at `:111-118`; `batcher.rs:290`, `:353` | -| 15 min without a tip advance closes admission for everyone | `hub/src/batcher.rs:62`, `:205-208`, `:161-190` (`observe` refreshes `last_advance` only on advance); `hub/src/server.rs:248-254` | -| The refusal never reaches the wallet | `shim/src/hub.rs:231-240` (returns `Submit::Accepted` at hand-off, *"a hub refusal is never surfaced here"*); pinned by `shim/tests/divert_nym.rs:235-267` | -| `/healthz` is unconditionally 200 | `hub/src/server.rs:450-453` | -| The enclave is 2 vCPU / 2 GB and internet-reachable on the HTTP port | `hub/deploy/caution/caution.hcl.tmpl:39-40`, `:51-55`, `:97-105` | -| The code already treats descriptor exhaustion as a reachable condition | `hub/src/server.rs:341-352` (`is_fd_exhaustion`), `:381-385` | - -### `AVOIDING-FALSE-POSITIVES.md` §5 applied in both directions - -§5's canonical false positive is *"Unlimited concurrent connections … OS limits -(ulimit), nginx (worker_connections), firewall rules apply first; **each -connection uses minimal resources**"*. Three things take this issue out of that -shape, and they are the same three §5 names as the *real* pattern: - -1. **The connection does not use minimal resources.** One ~100-byte inbound - request causes one outbound TCP connect, one TLS handshake with chain - verification, one h2 handshake and up to 10 s of holding, *per endpoint*. - That is §5's own "Real Issue: single connection consuming unbounded - resources" / "1 KB request causing disproportionate work". -2. **The infrastructure that normally absorbs a flood is absent by - construction.** There is no proxy pooling the *outbound* leg — - `chain.rs` deliberately dials fresh every call — and the in-enclave Caddy in - front of the *inbound* leg is a plain reverse proxy with no rate limit and no - upstream connection cap in the shipped manifest. Nothing sits between the - attacker and `Hub::lookup`. -3. **"OS limits apply first" is the harm, not the mitigation.** The descriptor - limit, the ephemeral-port range and the two vCPUs are shared with the hub's - own outbound path. Reaching them stops the tip poll and the flush, which is - the whole attack. - -And the outcome is not the downtime §5 discounts. Because `shim/src/hub.rs:238-240` -acknowledges at mixnet hand-off, a hub that stops admitting is a hub that -silently destroys transactions users believe are spent — the *"violates -integrity guarantees"* row of the severity table, not the *"causes DoS"* row. - -### Corrections made to the filed text - -- **Severity Medium → High**, per the above. -- **The CPU arithmetic was replaced.** The filing's "order 100 connections … is - ~5,000 handshakes/s of demand" conflated concurrency with rate and was not - derived. The mechanisms are now stated in the order they actually bind, with - ephemeral-port exhaustion (a hard, deterministic `EADDRNOTAVAIL` at ~470 - connections/s against a *single* destination on stock Linux settings) added as - the sharpest one, and each marked as depending on platform defaults that the - Containerfile and manifest do not set and that could not be measured here. -- **The transport premise was corrected.** The filing described the ingress as - raw TCP on 8083. PROGRESS.md item 5/6 settled that `http { port = 8083 }` - makes it Caddy-proxied from 443; that does not reduce reachability (Caddy has - no rate limit or upstream cap here) but the text now says so accurately. -- **The mixnet comparison was made precise.** The G21 pass showed - `MAX_CONCURRENT_LOOKUPS` fails to bound mixnet *emission*. It does bound - concurrent *indexer dials*, which is the resource at issue here, so the - comparison stands and is now stated with that scope. -- **The "composes with a cheaper operator variant" bullet was moved out.** The - uncapped indexer response body is its own confirmed issue - (`hub-chain-unbounded-indexer-response-body.md`); the composition is noted - there rather than argued twice. -- The withdrawn amplification-loop premise recorded in PROGRESS.md item 7b-REFUTED is not relied on anywhere in this issue; the cost ratio here is a plain per-request one. - -### What this issue does *not* claim - -- It does not claim the enclave OOMs. Memory is the subject of - `hub-chain-unbounded-indexer-response-body.md`; the multiplication of the two - is real (G21 §4.3) and is noted in both, counted in neither. -- It does not re-argue the confidentiality of `POST /transaction`; that is the - confirmed `hub-unauthenticated-pre-publication-transaction-disclosure.md`. - Recommendation 3 is shared between them. -- It does not claim the flush-side `k × n` fanout, which belongs to - `hub-chain-connection-per-call-fanout-and-flush-memory-amplification.md`. - Recommendation 2 is shared. - -DO NOT DEVIATE FROM THIS TEMPLATE (`docs/SECURITY-ISSUE-TEMPLATE.md`). -LEAVE THIS MESSAGE IN PLACE SO THAT ALL AGENTS KNOW NOT TO DEVIATE FROM THIS TEMPLATE (`docs/SECURITY-ISSUE-TEMPLATE.md`). diff --git a/zeronym-22aa9851caf68-high-medium/high/hub-queue-unauthenticated-fill-silently-destroys-migrations.md b/zeronym-22aa9851caf68-high-medium/high/hub-queue-unauthenticated-fill-silently-destroys-migrations.md deleted file mode 100644 index e2dd5000..00000000 --- a/zeronym-22aa9851caf68-high-medium/high/hub-queue-unauthenticated-fill-silently-destroys-migrations.md +++ /dev/null @@ -1,450 +0,0 @@ -# An unauthenticated queue fill refuses every genuine migration fleet-wide, and the deployed submit path never tells the wallet - -**Severity**: High -**Validation Status**: Confirmed -**Location**: `audit-target/zeronym/hub/src/queue.rs:170-240` (`admit`), `:63-65` (`MAX_QUEUE_BYTES`), `:181-197` (unparseable payloads get `expiry = None`), `:29-33` (bytes-are-the-budget and never-evict rules), `:35-39` (the no-submitter-identity rule that forbids a rate limit); `audit-target/zeronym/hub/src/server.rs:248-277` (`Hub::admit`); `audit-target/zeronym/shim/src/nym.rs:563-690` (dispatch-only submit); `audit-target/zeronym/shim/tests/divert_nym.rs:235-268` (the project's own test that the refusal is not surfaced); `audit-target/zeronym/hub/deploy/caution/caution.hcl.tmpl:51-55` (`ingress 0.0.0.0/0`) -**Found by agent:** Local (`hub/src/queue.rs`) — this is AUDIT-INSTRUCTIONS "unverified lead #6", now traced end to end; validated 2026-08-18 -**In scope of audit?** Yes - -## Description - -`Queue::admit` refuses a submission with `Refusal::Full` once the queue's byte -total reaches `MAX_QUEUE_BYTES` (64 MiB). Three properties combine to make that -refusal reachable on demand by an anonymous attacker, and to make the -consequence silent: - -1. **Admission is unauthenticated and cannot be rate-limited within the stated - design.** The hub's Nym address is published at `GET /nym-address` with no - ACL, and `queue.rs:35-39` states that an entry must never carry any submitter - identifier — which is also the identifier a per-submitter quota would need. - There is no rate limit, no cost, and no ACL on any submit path. -2. **Junk is a first-class queue citizen with no expiry.** REVIEW #5's - re-parse-is-telemetry rule is implemented at `queue.rs:189-197`: a payload - that does not deserialize is admitted with `txid = None` and `expiry = None`. - `survives_next_flush` returns `true` unconditionally for `None` - (`queue.rs:386-388`), so **arbitrary bytes are admissible at any tip, forever**. - Filling the queue therefore costs nothing but bandwidth: no fee, no valid - transaction, no key material. -3. **The never-evict rule makes the fill stick.** `queue.rs:31-33` and - `queue.rs:216-225` implement "refuse at the door, never evict an admitted - entry". That is the right rule against an attacker choosing *which* entry to - remove, but it means whoever arrives first owns the budget for the rest of the - window, and the attacker can always arrive first. - -The consequence is not a fail-closed error the wallet sees. In the deployed -(mixnet) configuration the shim's submit is **dispatch-only**: it answers the -wallet `error_code 0` as soon as the frame is handed to the mixnet and never -awaits the hub's ack. The project's own test pins this: - -```rust -// shim/tests/divert_nym.rs:235-260 - async fn a_hub_refusal_is_not_surfaced_under_best_effort() { - // ... A hub that - // would refuse (queue full) therefore does not surface that refusal to the - // wallet ... - let (shim, seen) = spawn_nym_shim(backend, OnSubmit::Refuse(AckRefusal::QueueFull), ...).await; - ... - assert_eq!( - resp.error_code, 0, - "best-effort: the wallet is answered success on dispatch, not the refusal" - ); -``` - -So a `Refusal::Full` means: the wallet was told the migration was sent, the shim -holds no copy (`shim/src/nym.rs:576-590` — the ack receiver is dropped and never -awaited), and the hub discarded the bytes. **Nothing anywhere retains the -transaction.** - -## Attack Scenario and Steps - -1. Attacker fetches the hub's mixnet address from `GET /nym-address` - (`hub/src/server.rs:446-448, 469-479`), which is unauthenticated on an - `ingress 0.0.0.0/0` enclave and is published deliberately. -2. Attacker submits `SubmitV1` frames carrying arbitrary distinct bytes. - `hub/src/nym.rs:313-335` decodes them and calls `Hub::admit`, which calls - `Queue::admit`. Each payload fails `Transaction::zcash_deserialize`, so it is - admitted with `expiry = None` and never touches the expiry gate. - Distinctness is required only because dedup keys on `sha256(tx_bytes)` - (`queue.rs:206`), so varying one byte per frame suffices. -3. **Volume needed:** `MAX_QUEUE_BYTES = 67,108,864`; the mixnet submit path - carries up to `MAX_NYM_TX_BYTES = FRAME_BYTES - 33 = 65,503` bytes per frame - (`hub/src/wire.rs:76, 114, 127`). `ceil(67,108,864 / 65,503) = 1025` frames - fill the budget. Each frame is a fixed 65,536 bytes on the wire, so ~64 MiB of - frame payload. -4. **Rate needed (corrected during validation).** The queue is emptied at every - flush (`drain_shuffled` sets `bytes = 0`, `queue.rs:250-262`), every - `FLUSH_INTERVAL_BLOCKS = 20` blocks (~25 min at 75 s), so the attacker is not - filling a bucket once — they are racing the drain. With `W = 1500 s`, - `N = 1025` frames and an aggregate delivery rate `R` frames/s, the fill takes - `T = N/R` and the queue sits at the cap for `W − T`, so the fraction of each - window during which genuine migrations are refused is `f = 1 − N/(R·W)`. - Using the shim's own model of gateway throttling (`shim/src/nym.rs:1087-1115`: - `PACKET_BYTES = 2048`, `THROTTLED_PACKETS_PER_SEC = 1000/120 ≈ 8.33`), a - 64 KiB frame is 32 packets and one stock-rate client delivers ~0.26 frames/s: - - | Clients | Aggregate | `f` (window at the cap) | - |---|---|---| - | 2.6 | 0.68 frames/s ≈ 45 KB/s | **0 %** — merely matches the drain | - | 5 | 1.3 frames/s ≈ 86 KB/s (0.7 Mbit/s) | ~48 % | - | 26 | 6.8 frames/s ≈ 447 KB/s (3.6 Mbit/s) | ~90 % | - | 263 | 68 frames/s ≈ 4.5 MB/s | ~99 % | - - So this is a **bandwidth-proportional flood, not a one-shot fill**: denying - half of all migrations costs sub-megabit sustained, denying nine in ten costs - a few megabits. Nym clients are free to create and the hub cannot distinguish - or count them (`queue.rs:35-39` forbids it), so the only scarce input is - bandwidth. (An earlier draft of this issue claimed four clients saturate the - queue continuously; that is the rate that merely keeps pace with the drain - and denies nothing. Corrected at validation.) -5. While `inner.bytes` is at the cap, every genuine migration arriving from any - shim, for any operator, is refused `queue_full` (`queue.rs:223-225`). -6. The refusal is encoded as `AckRefusal::QueueFull` (`hub/src/wire.rs:280`) and - sent back to the shim, which is not listening (step above). The wallet was - told success minutes earlier. The migration is gone. -7. At each flush the hub publishes the attacker's ~1025 junk payloads to the - operator's indexer. `chain::broadcast_batch` issues the whole - (transaction x endpoint) product concurrently with no concurrency cap - (`hub/src/chain.rs:208-210`) and `unary_inner` dials a fresh TCP+TLS+h2 - connection per call (`chain.rs:310-334`), so this is ~1025 simultaneous - connections to the indexer every 25 minutes, each carrying a 64 KiB body of - garbage. That plausibly gets the hub rate-limited or banned by the only - endpoint it is allowed to reach (the enclave egress allowlist is a per-indexer - `/32`, `caution.hcl.tmpl`), which then converts every publish into - `Publish::Retryable` and feeds the unbounded-growth issue - (`hub-queue-requeue-ignores-byte-budget-unbounded-growth.md`). - -**Attack Requirements and Assumptions:** - -- **Access:** internet only. No credential, no enclave compromise, no privileged - mixnet position, no valid Zcash transaction, no fee. -- **What makes it realistic:** the hub's address is published by design; there is - no ACL, no authentication and no rate limit on any submit path (the shim→hub - channel has no authentication at all today — `OPEN-QUESTIONS.md` records STEVE - as designed-not-built); unparseable payloads are *required* to be admitted by - REVIEW #5; and the design rule at `queue.rs:35-39` structurally forbids the - per-submitter identity any quota would need. The attack is also - indistinguishable from honest load in the hub's own telemetry, which logs only - `reason = "queue_full"` counts (`server.rs:273`). -- **What limits it:** the fixed 64 KiB mixnet frame makes each byte of budget - cost ~1 byte of wire plus sphinx overhead, and gateway throttling caps a single - client, so the attacker needs several concurrent clients sustained - indefinitely rather than one burst. The clearnet `POST /` path, which would be - far cheaper, is gated behind `ZIH_HTTP_SUBMIT` and is off by default - (`hub/src/config.rs:96-97`, `hub/src/server.rs:438`), and - `hub/deploy/caution/OPERATORS.md:268,285` tells operators to leave it off. -- **Prior art to weigh:** `hub/REVIEW.md` acknowledges the *class* under - "Decisions for humans" — *"Fail-closed is a product decision ... it hands any - DoS-capable attacker a total availability kill against every participating - operator"* — and under inherent limits notes that *"Any party who can degrade - the shim-to-hub path ... chooses the moment that migration is published"*. - Neither states this vector (fill the hub's own byte budget with free junk), - its cost, or the silent-loss consequence that dispatch-only submit creates. - The review's own line 111 argues *against* a batch-size floor partly because - *"refusing ... hands any DoS-capable attacker a total availability kill"* — - yet `Refusal::Full` is exactly such a refusal. - -## Impact on Users - -- **A migration a wallet was told had been sent is destroyed with no error to - anyone.** The wallet's notes are marked spent-pending against a transaction - that will never reach a mempool. The user discovers this only by the - transaction never confirming, and only if their wallet surfaces that. -- **Fleet-wide.** One hub serves every participating operator's shims — that - shared queue *is* the anonymity set — so one attacker denies migration to all - of them at once. Orchard is closed to new value by NU6.3/ZIP 258 and the - migration is the mandatory way out, so "cannot migrate" is a substantive harm, - not a cosmetic outage. -- **It is a privacy attack, not only an availability one.** The migrations that - do slip through arrive into a batch whose only other members are the - attacker's junk — and junk is rejected by the node at publish - (`batcher.rs:368`, `Rejected` entries are dropped), so the *achieved* batch is - just the handful of genuine transactions, or one. `batcher.rs:412-419` will - duly warn `"batch provides no batching anonymity at this size"`. The attacker - chooses, per window, how many genuine migrations get in. -- **Recovery makes it worse.** A wallet that notices and retries hits the same - full queue. If the flood outlasts the transaction's expiry the wallet must - rebuild it; for a non-ZIP-318 Orchard spend carrying ZIP 203's 40-block default - that is under an hour. - -## Technical Details / Code Analysis - -The admission path, in order (`hub/src/queue.rs:170-240`): - -```rust - pub fn admit(&self, tx_bytes: &[u8], tip: u32, flush_interval: u32, mining_margin: u32) -> Admission { - if tx_bytes.len() > self.max_bytes.min(MAX_TX_BYTES) { - return Admission::Refused(Refusal::TooLarge); - } - - // Telemetry parse. A failure is never a refusal (REVIEW #5). - let (txid, expiry) = match Transaction::zcash_deserialize(&mut Cursor::new(tx_bytes)) { - Ok(tx) => (Some(tx.hash().to_string()), - tx.expiry_height().map(|h| h.0).filter(|height| *height != 0)), - Err(_) => (None, None), // <-- arbitrary bytes: no txid, NO EXPIRY - }; - - if !survives_next_flush(expiry, tip, flush_interval, mining_margin) { - return Admission::Refused(Refusal::ExpiryTooTight); - } // <-- unreachable for expiry = None - - let key: [u8; 32] = Sha256::digest(tx_bytes).into(); - let mut inner = ...; - - if inner.entries.contains_key(&key) { - return Admission::Duplicate { txid }; // <-- vary one byte to avoid - } - - if inner.bytes.saturating_add(tx_bytes.len()) > self.max_bytes { - return Admission::Refused(Refusal::Full); // <-- the target - } - inner.bytes += tx_bytes.len(); - inner.entries.insert(key, Entry { ... }); - Admission::Admitted { txid } - } -``` - -and the expiry rule that junk sails past (`queue.rs:380-393`): - -```rust -pub fn survives_next_flush(expiry: Option, tip: u32, flush_interval: u32, mining_margin: u32) -> bool { - match expiry { - None => true, - Some(expiry) => { - let deadline = next_flush_height(tip, flush_interval).saturating_add(mining_margin); - expiry >= deadline - } - } -} -``` - -Note that the arithmetic itself is correct and was verified during this audit: -`next_flush_height` is strictly-after and saturating, the `h % N == 0` and -`h % N == N-1` boundaries behave as REVIEW #2 requires, and ZIP 203's -`nExpiryHeight == 0` is correctly folded to `None` rather than to height 0. The -weakness is not in the arithmetic; it is that the *rule has no purchase on -payloads that carry no expiry at all*, and REVIEW #5 requires those to be -admitted. - -The refusal reaches the shim as a typed ack (`hub/src/nym.rs:313-326`): - -```rust - let kind = match hub.admit(&tx) { - Ok(_txid) => AckKind::Accepted, - Err(refusal) => AckKind::Refused(refusal.into()), - }; - Some(wire::encode_ack(&nonce, kind).to_vec()) -``` - -and the shim discards it, having already answered the wallet -(`shim/src/nym.rs:576-590`): - -```rust - /// A waiter is registered so the frame carries a nonce the hub CAN ack against - /// (the frame still carries reply SURBs, M6), but its receiver is dropped and - /// the reply is never awaited; the correlator sweeps the unclaimed waiter, and - /// an unmatched ack is discarded. -``` - -`Refusal::as_str()` is designed for a shim that reacts to it — -*"Typed rather than a string because the shim reacts differently to each: a tight -expiry means hold and retry, an unavailable hub means try another hub"* -(`queue.rs:67-71`) — but in the deployed mixnet configuration no shim reads it. -The typed-refusal design and the dispatch-only submit design are individually -defensible and jointly produce a silent loss. - -## Recommendations - -1. **Make the hub's byte budget robust to junk, since junk is a required - admission class.** Options that do not need a submitter identity: - - Reserve a fraction of `MAX_QUEUE_BYTES` for payloads that *parse* as - transactions, so an unparseable flood can never consume the whole budget. - This keeps REVIEW #5 (unparseable is admitted and published) while denying a - free flood the ability to displace real migrations. It costs the hub only - the ability to admit an unlimited number of unreadable payloads, which is - not a property any honest user needs. - - Refuse a *new* payload rather than an admitted one, but choose which class - to refuse by parseability rather than by arrival order. -2. **Surface `queue_full` to the wallet.** The dispatch-only trade is documented - as "the wallet learns the true outcome by confirmation", which is true for a - *queued* transaction and false for a *refused* one. Either await the ack for - the refusal-bearing cases, or have the shim retry-on-ack-refusal in the - background, or (cheapest) have the hub not refuse at all for the parseable - class per (1). A user must not be told `error_code 0` for bytes nobody holds. -3. **Bound `broadcast_batch` concurrency** (`chain.rs:208-210`) so a large batch - cannot open a thousand simultaneous connections to the one endpoint the - enclave is permitted to reach. Today the flood's second-order effect (getting - the hub banned) may be worse than its first-order effect. -4. **Raise the per-entry floor.** `hub/src/server.rs:534` rejects only *empty* - bodies and the mixnet path accepts any `declared <= MAX_NYM_TX_BYTES` - including 0; a minimum plausible transaction size would cost the attacker - nothing here (frames are fixed size) but would matter if the clearnet submit - path were ever enabled. -5. **State the residual honestly if it is accepted.** `README.md`'s "Protected" - list and `REVIEW.md`'s inherent-limits section should say that any - internet-connected party can, at low cost, prevent every participating - operator's users from migrating, and that under dispatch-only submit those - users are told the migration succeeded. - -## Validation Information - -**Verdict: CONFIRMED. Severity: High (as filed) — but the cost arithmetic in -step 4 of the attack scenario was wrong by roughly an order of magnitude and has -been corrected in place.** The mechanism, the reachability and the decisive -silent-loss leg all hold; what did not hold was "four clients saturate the queue -continuously", which described the rate that merely *keeps pace with* the drain -and therefore denies nothing. - -### Every mechanical claim re-verified against the target - -| Claim | Verified at | -|---|---| -| `MAX_QUEUE_BYTES = 64 MiB`, bytes are the budget | `hub/src/queue.rs:65`, module docs `:29-30` | -| A payload that does not deserialize is admitted with `txid = None, expiry = None` | `hub/src/queue.rs:190-197` (`Err(_) => (None, None)` at `:196`) | -| `survives_next_flush(None, ..) == true` unconditionally, so junk never meets the expiry gate at any tip | `hub/src/queue.rs:380-393` | -| `Refusal::Full` once `bytes + len > max_bytes` | `hub/src/queue.rs:222-225` | -| Never evict; first arrival owns the budget | module docs `hub/src/queue.rs:31-33`, implemented at `:216-231` | -| Dedup on `sha256(tx_bytes)`, so distinct payloads are required | `hub/src/queue.rs:206-214` | -| No submitter identity exists, and the design forbids one | `hub/src/queue.rs:35-39` — *"There is deliberately no contributor, channel or session identifier on an entry, and there must never be one."* | -| Mixnet admission is inline, unbounded, with no ACL and no rate limit | `hub/src/nym.rs:148-215`; `MAX_CONCURRENT_LOOKUPS = 64` (`:54`) bounds the *lookup* arm only, and `:41-46` says so explicitly ("Admission never waits on this") | -| The hub's Nym address is published to anyone | `hub/src/server.rs:446-449` (`NYM_ADDRESS_PATH` → `GET`, no auth), `:462-468` ("this endpoint is reachable by everyone"), on `ingress { cidr_ipv4 = "0.0.0.0/0" }` (`hub/deploy/caution/caution.hcl.tmpl`) | -| A frame carries at most `MAX_NYM_TX_BYTES = 65,536 − 33 = 65,503` bytes | `hub/src/wire.rs:114`, `:127`, enforced in `decode_submit` at `:320-322` | -| The clearnet `POST /` path really is closed | `hub/src/config.rs:82-98` (`default_value_t = false` at `:98`); grepped the whole tree — `ZIH_HTTP_SUBMIT` appears only in `OPERATORS.md:268,285` ("leave off") and in tests. The deploy never sets it, so the attack must pay mixnet bandwidth | -| The queue drains fully at each flush and the junk does not persist | `hub/src/queue.rs:246-262` (`drain_shuffled` sets `bytes = 0` at `:259`); `hub/src/batcher.rs:364-378` — a `Rejected` verdict is counted at `:368` and never requeued, so junk the node refuses does not come back | -| A flushed batch is published as one fresh TCP+TLS+h2 connection per (transaction × endpoint), uncapped | `hub/src/chain.rs:208-210`, `:300-334` | - -**The decisive leg — the refusal never reaches the wallet — is confirmed three -ways.** `shim/src/hub.rs:231-240` has no `Refused` arm and returns -`Submit::Accepted` on hand-off; `shim/src/nym.rs:652` builds the ack waiter -and immediately drops its receiver (`let (ack_tx, _drop_receiver) = oneshot::channel();`); -and the project's own test `shim/tests/divert_nym.rs:235-267` -(`a_hub_refusal_is_not_surfaced_under_best_effort`) drives the mock hub with -`OnSubmit::Refuse(AckRefusal::QueueFull)` and asserts `resp.error_code == 0` -with the comment *"best-effort: the wallet is answered success on dispatch, not -the refusal"*. A `Refusal::Full` is therefore not an error the user sees; it is -a transaction the user was told had been sent, held by nobody. - -Worth recording for the report: `hub/src/config.rs:82-88` justifies leaving the -clearnet submit path off on the grounds that *"the mixnet address IS the -credential"* — while `hub/src/server.rs:462-468` publishes that credential to -every unauthenticated caller by design. The two statements cannot both be load -bearing, and this issue is what falls out of the gap. - -### The corrected cost, and why the original figure was wrong - -The queue is emptied at every flush, so the attacker is not filling a bucket -once; they are racing the drain. Let `W = 1500 s` (20 blocks × 75 s), -`N = 1025` frames, and `R` the attacker's aggregate delivery rate in frames/s. -Filling takes `T = N/R`, so the queue sits at the cap for `W − T` of each -window, and the fraction of the window during which genuine migrations are -refused is `f = 1 − N/(R·W)`. - -At the crate's own throttled-client model (`shim/src/nym.rs:1087-1115`: a -64 KiB frame is 32 packets, 8.33 packets/s), one stock client delivers ~0.26 -frames/s: - -| Clients | Aggregate | `f` (window at the cap) | -|---|---|---| -| 2.6 | 0.68 frames/s ≈ 45 KB/s | **0 %** — this is the rate the filing called "saturating"; it only matches the drain | -| 5 | 1.3 frames/s ≈ 86 KB/s (0.7 Mbit/s) | ~48 % | -| 26 | 6.8 frames/s ≈ 447 KB/s (3.6 Mbit/s) | ~90 % | -| 263 | 68 frames/s ≈ 4.5 MB/s | ~99 % | - -So the attack is real but is a *bandwidth-proportional* flood, not a one-shot -fill: denying half of all migrations costs sub-megabit, denying nine in ten -costs a few megabits, sustained. The filing's "roughly four clients saturate the -queue continuously" has been replaced with this curve. - -### `AVOIDING-FALSE-POSITIVES.md` §5 applied rigorously - -§5 asks: *what resources would the attacker need, and what would stop them?* - -*Resources:* 0.7–3.6 Mbit/s of sustained mixnet traffic from Nym clients that -cost nothing to create and are unattributable by construction. No credential, -no valid transaction, no fee, no on-chain activity, no privileged network -position. This is a single cheap VPS. - -*What would stop them:* nothing in the target. There is no ACL, no rate limit, -no proof of work, no cost, and no per-submitter accounting — and by -`queue.rs:35-39` there must never be a submitter identifier, so the usual quota -cannot be built. The two infrastructure limiters §5 normally credits (proxy -timeouts, upstream connection caps) do not exist on this path: admission is -inline in the mixnet listener with no bound at all. The only real limiter is -mixnet throughput to the hub's single pinned gateway, which is what the curve -above prices. - -*Why this is not the §5 false-positive shape.* This is not an amplifier — the -attacker pays about one byte per byte of budget — and §5's canonical false -positives (requesting a 100 GB database, unlimited connections) are dismissed -because the *outcome* is downtime that infrastructure absorbs. Here the outcome -is not downtime. A refused submission is a transaction the wallet was told had -succeeded (`shim/src/hub.rs:231-240`), that the shim keeps no copy of -(`shim/src/lib.rs:32-34`, stateless by design), and that the hub discarded. It -therefore lands on the *"violates integrity guarantees"* row of the severity -table, not the *"causes DoS"* row. On this target the guide's usual test -inverts: the attacker's resource is bandwidth, and the damage is silent -destruction of transactions the user believes are spent. - -### `AVOIDING-FALSE-POSITIVES.md` §6 (intentional design) — stated and answered - -Three of the ingredients are deliberate and correct, and the report must not -ask for them to be reversed: - -- **Unparseable payloads must be admitted** (REVIEW #5). Refusing them would - invert the shim's fail-safe into a leak. Correct as designed. -- **Never evict an admitted entry** (`queue.rs:31-33`). Eviction would hand an - attacker the *selection* lever, which is worse. Correct as designed. -- **No submitter identifier, ever** (`queue.rs:35-39`). An operator-to-migration - map inside the enclave is precisely what the system exists to destroy. - Correct as designed — and this is why "add a rate limit" is not an available - fix and must not be the headline recommendation. - -What is *not* a design decision is the composition of those three with a single -undifferentiated byte budget: nothing in `REVIEW.md` states, or appears to have -considered, that an unparseable flood may consume 100 % of `MAX_QUEUE_BYTES` -and displace the parseable class. Recommendation 1 (reserve a fraction of the -budget for payloads that parse) closes exactly that and needs no identity, no -rate limit and no change to any of the three rules above. `hub/REVIEW.md` -acknowledges the *class* ("Fail-closed … hands any DoS-capable attacker a total -availability kill") but neither this vector, nor its cost, nor the silent-loss -consequence dispatch-only submit gives it. - -### Impact claims checked - -- *Fleet-wide:* yes. One hub serves every participating operator's shims; that - shared queue is the anonymity set. -- *"The migrations that slip through arrive into a batch whose only other - members are the attacker's junk":* correct, and in fact stronger than filed — - the node rejects the junk at publish and `flush` **drops** `Rejected` entries - (`batcher.rs:364-378`), so the junk never reaches the chain and a chain - observer sees only the genuine members. The attacker chooses, per window, how - many of those there are. -- *Recovery:* a wallet that retries re-enters the same full queue. Note the one - softening fact, recorded honestly: a wallet that resends identical bytes is - deduplicated at the hub (`queue.rs:206-214`) and will succeed if it happens to - retry during a gap in the flood, so the loss is not unconditionally - permanent — it is permanent for any migration whose expiry runs out first. - -### Severity justification — High - -*Impact:* a mandatory migration the wallet reported as sent reaches no mempool, -with no error surfaced anywhere and no component retaining the bytes; the user's -notes stay pending-spent until expiry, in a pool NU6.3/ZIP 258 has closed to new -value. It hits every operator's users at once. - -*Likelihood:* unauthenticated, internet-reachable, free, unattributable -(the mixnet is the attacker's cover too), indistinguishable from honest load in -the hub's own telemetry (`server.rs:272-274` logs only `reason = "queue_full"`), -and priced at sub-megabit to low-megabit sustained. - -*Why not Critical:* no funds are stolen and no key is compromised; the attack -must be sustained rather than executed once; a retry into a later window can -still succeed; and at today's adoption the *privacy* half of the harm is largely -redundant with the already-accepted "modal batch is 0 or 1" residual. - -*Why not Medium:* it needs no special configuration, no privileged position and -no credential — only bandwidth — and its primary consequence is the silent -destruction of a transaction the user was told had succeeded, which the severity -guidance places above a mere availability outage. - -DO NOT DEVIATE FROM THIS TEMPLATE (`docs/SECURITY-ISSUE-TEMPLATE.md`). -LEAVE THIS MESSAGE IN PLACE SO THAT ALL AGENTS KNOW NOT TO DEVIATE FROM THIS TEMPLATE (`docs/SECURITY-ISSUE-TEMPLATE.md`). diff --git a/zeronym-22aa9851caf68-high-medium/high/hub-surb-starved-lookup-replies-grow-the-sdk-pending-buffer-without-bound-and-oom-the-enclave.md b/zeronym-22aa9851caf68-high-medium/high/hub-surb-starved-lookup-replies-grow-the-sdk-pending-buffer-without-bound-and-oom-the-enclave.md deleted file mode 100644 index aa57e65d..00000000 --- a/zeronym-22aa9851caf68-high-medium/high/hub-surb-starved-lookup-replies-grow-the-sdk-pending-buffer-without-bound-and-oom-the-enclave.md +++ /dev/null @@ -1,605 +0,0 @@ -# A 64-byte anonymous lookup carrying one reply SURB makes the hub buffer 64 KiB it can never send, forever: unbounded enclave memory growth that ends in an OOM which destroys the queue and permanently strands every shim - -**Severity**: High -**Validation Status**: Confirmed -**Location**: -`audit-target/zeronym/hub/src/nym.rs:148-216` (`run_listener`, the lookup arm), `:232-234` (`is_lookup`), `:249-303` (`build_lookup_reply`, whose `hash.is_empty()` arm at `:269-272` answers a full frame with no I/O at all), `:38-54` (`MAX_CONCURRENT_LOOKUPS`), `:56-75` (`REPLY_DEADLINE`); -`audit-target/zeronym/hub/src/nym_driver.rs:452-479` (the reply arm), `:632-643` (`reply_send` → `sender.send_reply(tag, frame)`), `:262-276` (the in-RAM `Ephemeral` store that makes any process restart change the address); -`audit-target/zeronym/hub/src/wire.rs:476-500` (`encode_lookup_reply` pads **every** disposition to `FRAME_BYTES` = 64 KiB); -`audit-target/zeronym/hub/src/server.rs:62-70` and `:462-479` (`GET /nym-address` publishes the target to anyone); -`audit-target/zeronym/hub/src/main.rs:184-185` (both mixnet channels sized 64); -`audit-target/zeronym/hub/deploy/caution/caution.hcl.tmpl:39-41` (`cpu = 2`, `memory_mb = 2048`), `:50-55` (`ingress 0.0.0.0/0`). -The unbounded store is in the pinned SDK, `nym-sdk 1.21.5-rc.1` at git rev `451c2aa3692fc4dc00041b74a352d4158176d9c0` (`hub/Cargo.lock:4773-4775`): `common/client-core/src/client/replies/reply_controller/receiver_controller.rs:218-320` (`handle_send_reply`), `:685-758` (`inspect_stale_pending_data`), `:499-529` (`handle_received_surbs`), `:811-926` (`inspect_and_clear_stale_data`); `common/client-core/src/client/transmission_buffer.rs:39-49`, `:225-236`; `common/client-core/surb-storage/src/surb_storage.rs:190-192`, `:460-478`; `common/client-core/config-types/src/lib.rs:48-62`. - -**Found by agent:** Global (focus area G5 — unauthenticated ingress on both enclaves; with G21 resource exhaustion as a privacy attack and G31 the unauthenticated internet attacker) -**In scope of audit?** Yes. This closes `BRAINSTORM.md` §R10-J, which recorded the question as *"the highest-value item in R10 that I could not close"* because it needed `nym-client-core` internals. Those internals were read at the exact pinned revision during validation and the answer is: **the store is not bounded, and the SDK's own two escape hatches are both defeated by an attacker who keeps talking.** - -## Description - -The hub answers a mixnet `LookupV1` with a `LookupReplyV1` that is **always -padded to 64 KiB** (`wire.rs:476-500`), so that a reply's size cannot reveal -found-versus-not-found. The driver hands that frame to the Nym SDK with -`sender.send_reply(tag, frame)` (`nym_driver.rs:638`), which is **fire and -forget**: it returns as soon as the SDK's one-slot input channel accepts the -`InputMessage`, not when anything is emitted. - -A reply travels on **reply SURBs that the requester attached**. The hub attaches -none of its own and has no way to obtain any. So the requester decides whether -the reply can be sent at all — and the hub builds and hands over the full 64 KiB -either way. - -What the SDK does with a reply it cannot send is the whole finding. In -`ReceiverReplyController::handle_send_reply` -(`receiver_controller.rs:218-320`): - -```rust - if !self.surbs_storage.contains_surbs_for(&recipient_tag) { - // ... warn once ... - return; // dropped: SAFE - } - let mut fragments = self.message_handler.split_reply_message(data); - let available_surbs = self.surbs_storage.available_surbs(&recipient_tag); - let min_surbs_threshold = self.surbs_storage.min_surb_threshold(); // 10 - let max_to_send = if available_surbs > min_surbs_threshold { - min(fragments.len(), available_surbs - min_surbs_threshold) - } else { - 0 - }; - ... - // if there's leftover data we didn't send because we didn't have enough - // (or any) surbs - buffer it - if !fragments.is_empty() { - ... - self.insert_pending_replies(&recipient_tag, fragments, lane); // UNBOUNDED - } -``` - -`insert_pending_replies` appends into `SenderData::pending_replies`, a -`TransmissionBuffer` -(`receiver_controller.rs:29-45`, `:126-138`), which is a bare -`HashMap>` with **no size limit and no -byte budget** (`transmission_buffer.rs:39-49`). Its own -`prune_stale_connections` is never called on this buffer — the only call site in -the whole SDK is the real-traffic stream's own buffer -(`real_traffic_stream.rs:326`). - -So one attacker frame of 64 bytes causes ~41 fragments (~64 KiB of prepared -plaintext) to be retained indefinitely, **provided the sender tag has an entry in -the SURB store** so the early `contains_surbs_for` return is not taken. - -Two facts about that guard decide the whole economics of the attack, and both -were checked directly against the pinned tree: - -- **Zero attached SURBs is safe, one is enough.** A message with no attached - SURBs never reaches `send_additional_surbs` - (`received_buffer.rs:322-330`: `if !reply_surbs.is_empty()`), so no store entry - is created, `contains_surbs_for` is false, and the reply is dropped cleanly. - One attached SURB creates the entry. -- **`contains_surbs_for` is `contains_key`, not "has a SURB left"** - (`surb_storage.rs:190-192`). Once the entry exists it stays, so **every later - message from the same sender tag is buffered whether it carries a SURB or - not.** The sender tag is stable per (client, recipient) pair - (`message_handler.rs:239-250`, "using {new_tag} for all anonymous messages sent - to {recipient}"), so one client keeps one tag for the whole attack. - -The SDK has exactly two mechanisms that would eventually free the buffer, and -**both are keyed on "we have not heard from this sender", so both are defeated by -sending one SURB every few seconds** (`receiver_controller.rs:685-758`, run every -5 s from `reply_controller/mod.rs:145-147`): - -```rust - let Some(last_received_time) = - self.surbs_storage.surbs_last_received_at(pending_reply_target) else { ... }; - let diff = now - last_received_time; - ... - if vals.current_clear_rerequest_counter > max_rerequests { // 5 - to_remove.push(*pending_reply_target); continue; - } - if diff > max_rerequest_wait { // 10 s - if diff > max_drop_wait { to_remove.push(...) } // 5 min - else { vals.increment_current_clear_rerequest_counter(); ... } - } -``` - -- `surbs_last_received_at` is refreshed by **every** SURB that arrives - (`surb_storage.rs:460-478`), so `diff` never crosses the 10 s re-request wait, - let alone the 5 min drop wait. -- `handle_received_surbs` calls `reset_rerequest_counter(&from)` on **every** - arrival (`receiver_controller.rs:499-529`), so the 5-re-request give-up counter - is reset before it can reach its threshold. -- `surb_senders.remove(...)` at `receiver_controller.rs:756` is the **only** - removal site in the file; nothing else ever drops a `SenderData`. -- The other periodic sweep, `inspect_and_clear_stale_data` - (`receiver_controller.rs:811-926`), retains/purges only `surbs_storage`, and its - own eviction predicate additionally requires `possibly_abandoned` (5 min of - silence) and `pending_reception() == 0`, neither of which holds here. - -And the buffer never drains, because draining requires *more* SURBs than the -minimum threshold: `try_clear_pending_queue` returns immediately unless -`available_surbs > min_surb_threshold` (10) (`receiver_controller.rs:444-455`). -At one attached SURB per message the steady state is: the SDK spends the -attacker's SURBs on futile "send me more SURBs" requests until -`pending_reception` reaches `maximum_reply_surb_storage_threshold` (200) and it -stops asking, after which stored SURBs hover at 10-11 and each arriving SURB -clears **one** buffered fragment while that same message adds **41**. Net growth -is ~40 fragments — about 64 KiB — per attacker message, forever. - -## Attack Scenario and Steps - -Attacker: anyone with an internet connection. No credential, no Zcash knowledge, -no valid transaction, no privileged network position, no enclave access. - -1. `curl https:///nym-address`. This endpoint exists to publish the - value and answers everyone by design (`server.rs:62-70`, `:462-479`); the - enclave declares `ingress { cidr_ipv4 = "0.0.0.0/0" }` - (`hub/deploy/caution/caution.hcl.tmpl:50-55`). -2. Start a stock `nym-sdk` client. It is free, needs no registration, and - attaches to a public gateway. -3. Repeatedly `send_message(hub, frame, IncludedSurbs::new(1))` where `frame` is - 64 bytes: `b"ZNL1"` ‖ 16 random bytes ‖ `0x00` (hash length zero) ‖ 43 zero - bytes. `IncludedSurbs::Amount(1)` produces an **anonymous** message carrying - the client's stable sender tag and exactly one reply SURB - (`nym-sdk/src/mixnet/traits.rs:72-95`). -4. On the hub, `is_lookup` accepts it on size-plus-magic - (`nym.rs:232-234`), a concurrency slot is taken, and `build_lookup_reply` - reaches the empty-key arm at `nym.rs:269-272`: - - ```rust - if hash.is_empty() { - tracing::warn!(reason = "empty lookup key", "lookup refused"); - return Some(error_reply(nonce)); - } - ``` - - `error_reply` is `encode_lookup_reply(&nonce, &LookupReply::Error)` — a full - `FRAME_BYTES` = 64 KiB buffer (`wire.rs:480`). **No indexer dial, no queue - scan, no I/O of any kind**, so the hub answers these as fast as it can read - them and neither `MAX_CONCURRENT_LOOKUPS` nor `REPLY_DEADLINE` slows it: the - slot is released the moment the reply enters the 64-deep `outgoing` channel, - and the reply is minutes fresher than the deadline. -5. The driver takes it and calls `send_reply` (`nym_driver.rs:474`, `:638`), - which returns as soon as the SDK's `mpsc::channel::(1)` accepts - it (`base_client/mod.rs:1013`); the input listener does nothing with a - `Reply` but `reply_controller_sender.send_reply(...)` - (`input_message_listener.rs:60-70`), which is an **unbounded** channel - (`reply_controller/requests.rs:12-15`, `:63-77`). There is no backpressure - anywhere between the hub's driver and the buffer that grows. -6. The reply controller buffers all ~41 fragments of the 64 KiB reply, as above, - and never sends or frees them. -7. Repeat. Every packet the attacker sends adds ~64 KiB of permanently held - enclave memory. Sending at least one SURB-bearing packet every ten seconds - keeps the SDK's two cleanup paths disarmed indefinitely; everything above that - rate is pure growth, and messages sent *without* SURBs are buffered just the - same (see `contains_surbs_for` above), so the marginal cost is **one sphinx - packet per 64 KiB permanently held**. - -**Arithmetic.** The enclave has `memory_mb = 2048`, most of which the project's -own manifest attributes to EnclaveOS plus the 64 MiB queue budget -(`hub/deploy/caution/caution.hcl.tmpl:33-41`). Taking ~1 GB as headroom, the -attacker needs ~16,000 messages. A stock client's Poisson stream defaults to -`message_sending_average_delay = 20 ms` (`client-core/config-types/src/lib.rs:25`, -`:443`), i.e. ~50 packets/s, and the project's own nymnet measurement puts one -attached reply SURB at about one sphinx packet (`shim/src/nym.rs:98-104`, -`hub/src/nym.rs:60-64`: 60 SURBs ≈ 60 packets). So **one free client reaches an -OOM in roughly five to ten minutes**; even at the ~8 packets/s floor the project -measures under real gateway backpressure it is about half an hour to an hour, and -clients cost nothing to run in parallel. The attacker's own bandwidth cost is a -few kilobytes per second. - -**The attacker has a free progress indicator.** `GET /nym-status` is -unauthenticated and reports `mixnet_connected`, `client_deaths` and -`consecutive_rebuild_failures` (`server.rs:454-457`), and `GET /nym-address` -returns the current address. Polling either tells the attacker the moment the hub -died and came back with a new identity. - -**Attack Requirements and Assumptions:** -- Network access, and the hub's Nym address, which the hub publishes on purpose. -- **At least one reply SURB, once.** Zero SURBs on *every* message is not the - attack: with no store entry for the tag, `handle_send_reply` returns at - `receiver_controller.rs:225-241` and the reply is dropped cleanly. This - corrects the guess recorded in `BRAINSTORM.md` §R10-J that zero SURBs would be - the cheapest variant. One SURB creates the entry, and a SURB every ten seconds - keeps it and the buffer alive. -- No rate limit, ACL, authentication or submitter accounting exists anywhere on - this path, by design (`hub/src/nym.rs:12-15`; `hub/src/queue.rs:35-39`). -- The pinned SDK revision is exactly the tree analysed - (`hub/Cargo.lock:4775`, `shim/Cargo.lock:5065`, both - `#451c2aa3692fc4dc00041b74a352d4158176d9c0`). **Neither binary sets any - `DebugConfig` reply-SURB parameter**: the hub's only `debug_config` call sets - `acknowledgements.ack_wait_addition` (`hub/src/nym_driver.rs:202-216`) and the - shim's is gated behind the `mixnet-localnet` feature, so every value in - `config-types/src/lib.rs:48-62` is the shipped default. -- What limits it: the flood is noisy on the mixnet, and a future - encrypt-to-hub-key or STEVE layer would change the picture — but both are - listed in `README.md` as *"Designed, no code yet"*. - -## Impact on Users - -The hub is a single, shared, fleet-wide component, and this is an -unauthenticated remote kill switch for it. - -- **Migrations users were told had succeeded are destroyed.** On the deployed - transport submit is dispatch-only: the wallet is answered `error_code 0` when - the frame enters an in-process channel (`shim/src/hub.rs:226-240`, - `shim/src/nym.rs:595-690`). Everything the hub has admitted but not yet - published lives only in enclave RAM, up to `MAX_QUEUE_BYTES` = 64 MiB - (`hub/src/queue.rs:65`). An OOM is a `SIGKILL`, so **not even the shutdown - flush and not even the `unpublished … they are lost` log line - (`batcher.rs:317-330`) run** — that path is reached only from the SIGTERM / - ctrl-c handler (`main.rs:220-247`). The loss is completely silent. -- **The attacker chooses when.** Flushes happen on a 20-block cadence and are - visible on the public chain, so an attacker who can drive an OOM in ten to - sixty minutes can start early enough to land the kill just before a cadence - height, when the queue holds a whole epoch of migrations. -- **The whole shim fleet is stranded, and recovery is manual.** The hub's - identity store is `Ephemeral::default()` built inside `run_driver` - (`hub/src/nym_driver.rs:269`), so it survives client rebuilds but **not a - process restart**; the module header says so at `hub/src/nym_driver.rs:34-40`. - This holds whichever way the platform behaves: if the unit is restarted - automatically the hub comes back with a **different Nym address**, and if it is - not, the operator's redeploy produces one just the same. Every shim has the old - address baked into an immutable Caution enclave config - (`ZIS_HUB_NYM`, `shim/deploy/caution/assemble-caution.sh:345-356`; - `shim/src/config.rs:69-78`, read once at startup with no discovery mechanism - anywhere in `shim/src`). Their submits go to a recipient that no longer exists, - the SDK reports success, and every migration is silently destroyed until every - operator re-assembles and redeploys. That is the outage already filed as - `hub-nym-driver-automatic-fresh-identity-permanently-invalidates-every-shim.md`, - whose stated trigger was bad luck (~11 minutes of gateway failure) — **this - issue hands an anonymous outsider an on-demand trigger for it.** -- **Every wallet loses `GetTransaction`, not just migrating ones.** With a hub - configured the shim routes *every* lookup to the hub and fails closed on error - (`shim/src/intercept.rs:229-236`, `:315-324`), so a dead hub means no wallet - behind any zeronym shim can fetch any transaction's full data — including users - who have never touched Orchard. -- **The realistic user response is the leak.** A wallet that cannot send or - confirm is pointed at a different, unprotected indexer, where the - Orchard-touching transaction is broadcast in the clear. The attacker chooses - the moment. -- **Nothing alerts.** `GET /healthz` is unconditionally `200 ok` - (`server.rs:449-452`) right up to the OOM, and after the restart it is 200 - again — on a hub no shim can reach. -- **Repeating it burns the TLS budget.** Every enclave restart is a fresh ACME - order and the duplicate-certificate limit is 5 per week - (`hub/deploy/caution/RESTARTS.md:1-21`), so an attacker who re-triggers the OOM - a few times leaves the hub unable to obtain a certificate at all — including for - the `/nym-address` endpoint operators would need to read the new address from. - See `acme-nocache-issuance-budget-is-a-restart-triggerable-tls-outage-with-no-health-signal.md`. -- **Submissions are destroyed for as long as the outage lasts, and funds are not - frozen.** The note is not spent on chain and the wallet keeps retrying (see the - correction below); what an OOM'd hub destroys is every submission made while it - is down, plus everything its RAM queue held. - -> **CORRECTION 2026-08-18 (validation of the cited file — SUPERSEDES the sentence above).** -> The wallet does **not** wait for expiry. Both official Zcash light-wallet SDKs -> automatically resubmit a sent-but-unmined transaction for as long as it remains -> unexpired — the Android SDK at the head of every ~20 s sync loop and after every -> processed block batch (`CompactBlockProcessor.kt:573,615,723`; selection -> `mined_height IS NULL AND expiry_height > ?`), the iOS SDK at most once per 300 s -> (`TxResubmitter.swift:8-15`, `TransactionDao.swift:218-228`) — and the hub's -> payload-hash dedup makes the resend free. The wallet's non-confirmation signal comes -> from compact-block scanning, which the shim does not intercept (`proxy.rs:1068-1074`). -> Expiry is therefore the **retry horizon**, not the wait: ~50 minutes for the ZIP 203 -> default traffic the shim also diverts, 30–60 days for a ZIP 318 migration. A -> *transient* loss self-heals within minutes; only a loss condition that **outlives the -> horizon** destroys the submission permanently — which is exactly what this issue's -> condition does, so this issue's severity is unaffected. Do not write "the user waits -> 30 to 60 days" in the report. Full refutation and the replacement paragraph: -> `issues/invalid/zip318-canonical-expiry-is-the-only-recovery-clock-and-a-lost-migration-freezes-the-users-notes-for-30-to-60-days.md`. - - -## Technical Details / Code Analysis - -**1. Any 64-byte frame with the right magic is a lookup, and every lookup answer -is 64 KiB.** - -```rust -// hub/src/nym.rs:232-234 -fn is_lookup(frame: &[u8]) -> bool { - frame.len() == wire::LOOKUP_BYTES && wire::peek_lookup_nonce(frame).is_some() -} -``` - -`peek_lookup_nonce` checks only length ≥ 21 and the four magic bytes -(`hub/src/wire.rs:463-470`). The doc comment above `is_lookup` records that -dispatching on magic alone *"was an amplifier: a 21-byte message got a -65 536-byte answer"* — the size check fixed the ratio from 3000× to 1000×, but -the amplifier itself is intact, and the finding here is not about the ratio: it -is about the reply being **retained** rather than merely built. - -```rust -// hub/src/nym.rs:254-272 - let error_reply = |nonce| { - wire::encode_lookup_reply(&nonce, &LookupReply::Error) - .expect("an error reply carries no transaction and always fits") - }; - let (nonce, hash) = match wire::decode_lookup(frame) { ... }; - if hash.is_empty() { - tracing::warn!(reason = "empty lookup key", "lookup refused"); - return Some(error_reply(nonce)); - } -``` - -```rust -// hub/src/wire.rs:476-484 -pub fn encode_lookup_reply(nonce: &Nonce, reply: &LookupReply) - -> Result>, WireError> { - let mut frame = Zeroizing::new(vec![0u8; FRAME_BYTES]); // 64 KiB, every time - ... - LookupReply::Error => frame[20] = 2, -``` - -`decode_lookup` accepts `hash_len = 0` without error (`wire.rs:439-456`: it only -rejects `declared > MAX_LOOKUP_HASH_BYTES`), so the empty-key arm is reached with -a well-formed frame and no error path at all. - -**2. The bounds that exist do not bound this.** - -```rust -// hub/src/nym.rs:171-197 - let permit = match lookups.clone().try_acquire_owned() { ... }; - tokio::spawn(async move { - if let Some(frame) = build_lookup_reply(&hub, &received.frame, received_at).await { - let _ = outgoing.send(Reply { ... }).await; - } - drop(permit); - }); -``` - -`MAX_CONCURRENT_LOOKUPS` (64) and the 64-deep `outgoing` channel together bound -what the **hub** holds to about 8 MiB. They say nothing about what the hub has -already pushed into the SDK, and they do not even throttle the rate: the driver -takes one reply at a time, but `send_reply` completes in microseconds (item 3), -so `outgoing` drains at memory speed. `REPLY_DEADLINE` (60 s) drops *aged* -replies (`nym.rs:285-292`, `nym_driver.rs:465-471`) but these replies are -answered in microseconds and are never aged when handed over. - -**3. The hand-off is fire and forget, into an unbounded queue.** - -```rust -// hub/src/nym_driver.rs:632-643 -fn reply_send(sender: ..., tag: AnonymousSenderTag, frame: Vec) -> InFlight { - Box::pin(async move { - if let Err(err) = sender.send_reply(tag, frame).await { - tracing::warn!(error = %err, "mixnet reply send failed"); - } - Sent::Reply - }) -} -``` - -`send_reply` builds `InputMessage::new_reply(...)` and awaits -`input_sender.send(...)` on a channel of capacity 1 -(`nym-sdk/src/mixnet/traits.rs:122-133`, `client-core/base_client/mod.rs:1013`). -The listener's `handle_reply` does nothing but -`reply_controller_sender.send_reply(...)` -(`acknowledgement_control/input_message_listener.rs:60-70`), which is an -`UnboundedSender` (`reply_controller/requests.rs:12-15`, `:63-77`). This is the -root cause named in G5 §4.7: **the memory an enclave can be made to hold on -someone else's behalf must be bounded by something the enclave owns.** - -**4. Both SDK cleanup paths are keyed on silence from the sender.** - -`inspect_stale_pending_data` runs every 5 s (`reply_controller/mod.rs:145-147`, -`:161-163`) and is quoted in full in the Description. `handle_received_surbs` -resets the give-up counter on every arrival: - -```rust -// client-core/.../receiver_controller.rs:499-529 - pub(crate) async fn handle_received_surbs(&mut self, from: ..., reply_surbs: ..., from_surb_request: bool) { - ... - self.surbs_storage.insert_fresh_surbs(&from, reply_surbs); // refreshes surbs_last_received_at - self.reset_rerequest_counter(&from); // disarms the 5-rerequest give-up - self.try_clear_pending_retransmission(from).await; - self.try_clear_pending_queue(from).await; // no-op below the threshold - if self.should_request_more_surbs(&from) { - self.request_reply_surbs_for_queue_clearing(from).await; // spends the SURB just received - } - } -``` - -```rust -// client-core/surb-storage/src/surb_storage.rs:460-478 - pub(crate) fn insert_fresh_reply_surbs(&mut self, surbs: I) { - let received_at = OffsetDateTime::now_utc(); - ... - self.surbs_last_received_at = received_at; -``` - -**5. The buffer itself has no cap.** - -```rust -// client-core/src/client/transmission_buffer.rs:39-49 -pub(crate) struct TransmissionBuffer { - buffer: HashMap>, -} -``` - -`prune_stale_connections` (`:225-236`) would evict a lane idle for ten minutes, -but its only caller is the real-traffic stream's own buffer -(`real_traffic_stream.rs:326`), never `SenderData::pending_replies`. All replies -use `TransmissionLane::General` (`nym-sdk/src/mixnet/traits.rs:122-128`), so they -all accumulate in one `VecDeque`. - -**6. The shipped configuration is the SDK default.** The values that govern all -of this — `minimum_reply_surb_storage_threshold = 10`, -`maximum_reply_surb_storage_threshold = 200`, -`maximum_reply_surb_rerequest_waiting_period = 10 s`, -`maximum_reply_surb_drop_waiting_period = 5 min`, -`maximum_reply_surbs_rerequests = 5` — are -`client-core/config-types/src/lib.rs:48-62`, and neither enclave overrides any of -them. These defaults are sized for a chat client that will eventually stop -talking to an unresponsive peer, not for an enclave holding other people's -funds-in-flight against an adversary who has every reason to keep talking. - -**Why the sibling HTTP path is not affected:** `POST /transaction` answers -synchronously into a hyper response (`server.rs:480-517`); nothing is retained -after the connection closes. **Why the shim is not affected:** the shim never -calls `send_reply` — it only sends and awaits (`shim/src/nym_driver.rs:608-624`) -— so `handle_send_reply` is never reached in its client. - -## Recommendations - -In rough order of value, and all of them compatible with the design rule at -`hub/src/queue.rs:35-39` that forbids a *submitter-to-migration* mapping (none of -these associates a sender with a queue entry): - -1. **Do not build a reply the sender cannot receive.** Answer a `LookupV1` only - when the request arrived with enough attached reply SURBs to carry a full - frame — the honest shim already attaches `LOOKUP_REPLY_SURBS` = 60 - (`shim/src/nym.rs:98-104`), chosen precisely to clear the SDK's threshold, so - a conforming client is unaffected and every under-provisioned request is - refused before 64 KiB is allocated. If `nym-sdk` does not expose the attached - count on `ReconstructedMessage`, the equivalent is to bound outstanding - replies per `SenderTag` (next item), which needs nothing from the SDK. -2. **Bound in-flight replies per sender tag.** `Received` already carries the - tag (`hub/src/nym.rs:83-93`) and the hub already keeps it for the lifetime of - the request in order to reply at all. A small token bucket keyed on the tag — - holding a counter and a timestamp, referencing no queue entry, and forgotten - after a minute — bounds this attack to `K × 64 KiB` per distinct tag and - costs a Sybil attacker one gateway registration per bucket. **This does not - violate `queue.rs:35-39`**, which forbids an identifier *on a queue entry*; - lookups never become queue entries. -3. **Stop answering undecodable and empty-key lookups at all.** The 64 KiB - padding exists to hide `Found` from `NotFound` (`hub/src/wire.rs:476-479`) — a - genuine privacy axis. A frame that failed `decode_lookup`, or that declares a - zero-length hash, is not on that axis: a conforming shim never sends one, and - the sender already knows what it sent. Dropping those silently (as the submit - arm already does for a frame with no recoverable nonce, - `hub/src/nym.rs:327-333`) removes the I/O-free variant of this attack with no - loss of indistinguishability. -4. **Set a `DebugConfig` for the reply-SURB parameters in both binaries,** but do - not mistake it for the fix. The hub already builds one for - `ZIH_ACK_WAIT_ADDITION_MS` (`hub/src/nym_driver.rs:202-216`); extend it. - Lowering `maximum_reply_surb_rerequest_waiting_period` and - `maximum_reply_surbs_rerequests` shortens the window but does **not** close - the hole on its own, because both counters are reset by every arriving SURB — - treat this as defence in depth and say so in the comment. -5. **Report the condition.** Add the reply controller's pending-queue size, or at - minimum a resident-memory figure, to `GET /nym-status`. Today the only symptom - before the OOM is invisible, and the only symptom after it is a - `/nym-address` that quietly changed. -6. **Fix the underlying fragility, not only this instance:** an enclave must - never hand a fire-and-forget buffer to a library without a matching admission - bound of its own. `outgoing` (64) bounds what the hub holds; nothing bounds - what the hub has already given away. -7. **Make recovery survivable.** Independently of this bug, an address change - should not require every operator to redeploy an immutable enclave. Any - mechanism that lets a shim learn the hub's current address (a signed pointer - record, a second stable identity, an address list refreshed at runtime) turns - this class of outage from fleet-fatal into an interruption. - -## Validation Information - -**Status: CONFIRMED. Severity confirmed at High.** - -The decisive claim was verified line by line against the **pinned SDK tree at -`451c2aa3692fc4dc00041b74a352d4158176d9c0`**, which is present locally, not -against the report. Everything checked: - -1. **The reply is built with no I/O.** `decode_lookup` (`hub/src/wire.rs:439-456`) - accepts `hash_len = 0`; `build_lookup_reply` then takes the `hash.is_empty()` - arm (`hub/src/nym.rs:269-272`) and returns a full `FRAME_BYTES` frame without - dialling an indexer or touching the queue. Confirmed. -2. **One attached SURB is necessary and sufficient.** `received_buffer.rs:322-330` - only forwards SURBs to the reply controller `if !reply_surbs.is_empty()`, so - zero SURBs never creates a store entry and the reply is dropped at - `handle_send_reply`'s `contains_surbs_for` guard. Confirmed — and this is the - opposite of `BRAINSTORM.md` §R10-J's guess. -3. **The buffer is unbounded and never pruned.** `TransmissionBuffer` has no cap - (`transmission_buffer.rs:39-49`); `prune_stale_connections` has exactly one - caller in the whole SDK and it is the real-traffic stream's buffer, not - `pending_replies`. `surb_senders.remove` appears exactly once - (`receiver_controller.rs:756`). Confirmed. -4. **Both escape hatches are disarmed by an arriving SURB.** - `insert_fresh_reply_surbs` sets `surbs_last_received_at = now` - (`surb_storage.rs:460-478`) and `handle_received_surbs` calls - `reset_rerequest_counter` unconditionally (`receiver_controller.rs:517`). - `inspect_and_clear_stale_data`'s eviction additionally requires - `pending_reception() == 0` and 5 minutes of silence, neither of which holds. - Confirmed. -5. **The hand-off is genuinely unbackpressured.** `send_reply` → capacity-1 - `InputMessage` channel → input listener → `unbounded_send` to the reply - controller. The input listener does no work at all for a `Reply` variant, so - `send_reply` returns in microseconds and `MAX_CONCURRENT_LOOKUPS` / - `outgoing` / `REPLY_DEADLINE` throttle nothing. Confirmed. - -**Corrections made during validation** (the filed text has been updated): - -- The original text said the attack needs one SURB *per message*. It does not: - `contains_surbs_for` is `contains_key` (`surb_storage.rs:190-192`), so once the - first SURB-bearing message has created the entry, **messages with zero attached - SURBs are buffered too**. The marginal cost is therefore **one sphinx packet - per 64 KiB permanently held**, with one keep-alive SURB every ten seconds. -- Fragment count corrected from "~33" to ~41, which is the project's own measured - figure for a 64 KiB frame (`shim/src/nym.rs:98-104`). -- The steady state was refined: `pending_reception` saturates at 200 and the SDK - stops asking for more SURBs, after which each arriving SURB clears one buffered - fragment while its own message adds 41. Growth continues at ~40 fragments per - message; the attack is not defeated, only slowed by 2.5%. -- The OOM is a `SIGKILL`, so the `unpublished … they are lost` line in - `batcher.rs:317-330` does **not** run — that path is reachable only from the - SIGTERM/ctrl-c shutdown handler. The loss is quieter than the original text - claimed. -- The fleet-strand composition was checked and **holds independently of whether - the platform auto-restarts the unit**: `Ephemeral::default()` is constructed - inside `run_driver` (`nym_driver.rs:269`) in a diskless enclave, so *any* - recovery — automatic restart or operator redeploy — produces a new Nym address, - and `ZIS_HUB_NYM` is static startup configuration in an immutable enclave with - no discovery path anywhere in `shim/src`. This is what makes the severity. -- Added: repeated OOMs burn the ACME duplicate-certificate budget (5/week, - `hub/deploy/caution/RESTARTS.md`), which can leave the hub without a - certificate for the very `/nym-address` endpoint operators need to recover. - -**Applying `docs/AVOIDING-FALSE-POSITIVES.md` §5 (impractical resource -exhaustion) explicitly.** The guide's test is "what resources would the attacker -need, and what would stop them?" Here the answer is one free, unregistered -`nym-sdk` client emitting a few kilobytes per second, and **nothing stops them**: -there is no ACL, no rate limit, no per-submitter accounting (forbidden by -`queue.rs:35-39` as it is currently read), no platform-level ingress restriction -(`ingress 0.0.0.0/0`), and no alerting. One 2 KB sphinx packet causes 64 KiB of -**permanently retained** enclave memory — a 32× byte amplification into a -cumulative resource rather than a transient one. That is the guide's stated -*real* vulnerability ("1 KB request causing 1 GB memory allocation"), i.e. the -exact inverse of the pattern §5 warns about, not an instance of it. - -**Not a duplicate of `hub-nym-lookup-flood-starves-gettransaction-fleet-wide.md`.** -That issue is about the *emitter* being a single serialised fleet-wide resource -and needs the attacker to attach ~51+ SURBs so the reply is actually sent; it -costs the attacker roughly as many packets as it costs the hub, and its effect -stops when the attacker stops. This issue needs *one* SURB, costs ~40× less, and -its effect is cumulative and survives the attacker. Different mechanism, -different fix; both are confirmed and cross-referenced. - -**Severity justification (High, not Critical).** It is remotely triggerable by -the weakest adversary in the threat model, at negligible cost, with no detection; -it destroys migrations the wallet was told had succeeded, and its recovery -necessarily strands every shim in the fleet until every operator redeploys. What -holds it below Critical: no funds are stolen and no key or plaintext is -disclosed to the attacker — the destroyed migrations are recoverable by the -wallet once the outage outlives its retry horizon (loss of the submission, not of -funds; see the CORRECTION above), and -the fleet-strand is repairable by a coordinated redeploy. It is nevertheless the -single worst outcome reachable from the unauthenticated ingress surface and -should be fixed before any further operator onboarding. - -> **CORRECTION 2026-08-18 (validation of the cited file — SUPERSEDES the sentence above).** -> The wallet does **not** wait for expiry. Both official Zcash light-wallet SDKs -> automatically resubmit a sent-but-unmined transaction for as long as it remains -> unexpired — the Android SDK at the head of every ~20 s sync loop and after every -> processed block batch (`CompactBlockProcessor.kt:573,615,723`; selection -> `mined_height IS NULL AND expiry_height > ?`), the iOS SDK at most once per 300 s -> (`TxResubmitter.swift:8-15`, `TransactionDao.swift:218-228`) — and the hub's -> payload-hash dedup makes the resend free. The wallet's non-confirmation signal comes -> from compact-block scanning, which the shim does not intercept (`proxy.rs:1068-1074`). -> Expiry is therefore the **retry horizon**, not the wait: ~50 minutes for the ZIP 203 -> default traffic the shim also diverts, 30–60 days for a ZIP 318 migration. A -> *transient* loss self-heals within minutes; only a loss condition that **outlives the -> horizon** destroys the submission permanently — which is exactly what this issue's -> condition does, so this issue's severity is unaffected. Do not write "the user waits -> 30 to 60 days" in the report. Full refutation and the replacement paragraph: -> `issues/invalid/zip318-canonical-expiry-is-the-only-recovery-clock-and-a-lost-migration-freezes-the-users-notes-for-30-to-60-days.md`. - - -DO NOT DEVIATE FROM THIS TEMPLATE (`docs/SECURITY-ISSUE-TEMPLATE.md`). -LEAVE THIS MESSAGE IN PLACE SO THAT ALL AGENTS KNOW NOT TO DEVIATE FROM THIS TEMPLATE (`docs/SECURITY-ISSUE-TEMPLATE.md`). diff --git a/zeronym-22aa9851caf68-high-medium/high/junk-sendtransaction-flood-consumes-the-shims-whole-mixnet-egress-and-converts-acknowledged-migrations-into-silent-loss.md b/zeronym-22aa9851caf68-high-medium/high/junk-sendtransaction-flood-consumes-the-shims-whole-mixnet-egress-and-converts-acknowledged-migrations-into-silent-loss.md deleted file mode 100644 index c7669464..00000000 --- a/zeronym-22aa9851caf68-high-medium/high/junk-sendtransaction-flood-consumes-the-shims-whole-mixnet-egress-and-converts-acknowledged-migrations-into-silent-loss.md +++ /dev/null @@ -1,643 +0,0 @@ -# A five-byte unauthenticated request buys 45 Sphinx packets of a shim's entire mixnet egress, so at about one byte per second an anonymous attacker holds any chosen operator's divert pipeline permanently full — which breaks every wallet's `GetTransaction` through that shim, spends the design's `MAX_DELIVERY_LAG` budget, can convert admitted migrations into silent `expiry_too_tight` refusals, and removes the one stated mitigating factor of the driver-teardown loss - -**Severity**: High -**Validation Status**: Confirmed -**Location**: `audit-target/zeronym/shim/src/proxy.rs:70` (`SEND_TRANSACTION`), `:743-782` (`route_for`, path-only, no method check, no authentication); `audit-target/zeronym/shim/src/intercept.rs:94-132` (`send_transaction`), `:495-568` (`inspect`, and specifically the `RawTransaction::decode` `Ok` arm at `:559-568`), `:137-215` (`divert`), `:481-490` (`Inspection::treat_as_migration`); `audit-target/zeronym/shim/src/classify.rs:99-101` (`Class::treat_as_migration`), `:332` and `:337` (the project's own unit tests that empty and garbage input are `Unparseable`); `audit-target/zeronym/shim/src/wire.rs:273-286` (`encode_submit` always produces `FRAME_BYTES`); `audit-target/zeronym/shim/src/nym.rs:595-690` (`NymHandle::submit`), `:1085-1115` (the crate's own emission arithmetic); `audit-target/zeronym/shim/src/main.rs:335-336` (channel capacities 32 and 8); `audit-target/zeronym/shim/src/nym_driver.rs:253-272` (one send in flight), `:113-125` (the throttled rate); `audit-target/zeronym/shim/src/hub.rs:231-240` (the anchor: `Submit::Accepted` at hand-off); `audit-target/zeronym/hub/src/queue.rs:199-204`, `:380-393` (`survives_next_flush` / `ExpiryTooTight`); `audit-target/zeronym/hub/src/batcher.rs:46-55` (`MAX_DELIVERY_LAG`, `MIN_WALLET_EXPIRY`), `:93-117` (`BatchParams::validate`); `audit-target/zeronym/hub/src/nym_driver.rs:187-201` (the measured enclave retransmission behaviour) -**Found by agent:** Global (focus area G7, "loss of a wallet-acknowledged migration"); validated 2026-08-18 -**In scope of audit?** Yes - -## Description - -The shim's wallet-facing listener is unauthenticated and internet-reachable -(`caution.hcl.tmpl` declares `ingress 0.0.0.0/0`), and `route_for` is a pure -function of the request path with no method check -(`proxy.rs:743-748`). Anything posted to -`/cash.z.wallet.sdk.rpc.CompactTxStreamer/SendTransaction` therefore reaches -`intercept::send_transaction`. - -Three facts, each individually intended, compose into a very cheap amplifier on -the one resource the whole divert path is bottlenecked on: - -1. **Unparseable input is diverted, by design.** `inspect` decodes the gRPC frame - and the `RawTransaction` protobuf; if that succeeds it hands `raw.data` to the - classifier and returns `(Inspection::Classified(evidence), Some(Bytes::from(raw.data)))` - (`intercept.rs:559-568`). `classify` returns `Class::Unparseable` for bytes - that are not a transaction, and `Class::Unparseable.treat_as_migration()` is - `true` on purpose (`classify.rs:99-101`). The crate's own unit tests pin - exactly the attacker's payload: `assert_eq!(classify(&[]), Class::Unparseable)` - (`classify.rs:332`) and `assert_eq!(classify(&[0xff; 64]), Class::Unparseable)` - (`classify.rs:337`). So an *empty* `RawTransaction` is diverted, and `divert`'s - fail-closed arm does not catch it, because that arm only fires when `tx_data` - is `None` (`intercept.rs:146-155`) and here it is `Some()`. - -2. **Every diverted payload becomes a full fixed-size frame, whatever its length.** - `wire::encode_submit` allocates `vec![0u8; FRAME_BYTES]` and zero-pads - (`wire.rs:280-285`). A zero-byte transaction and a 65,503-byte transaction cost - the mixnet exactly the same: 65,536 bytes. - -3. **The mixnet egress is a single, serialized, severely rate-limited resource.** - The driver holds **one send in flight at a time** (`nym_driver.rs:253-272`) at - the rate a real gateway's backpressure imposes. The crate computes this itself - in `shim/src/nym.rs:1085-1115`: `THROTTLED_PACKETS_PER_SEC = 1000.0 / 120.0` - (8.33 packets/s), `PACKET_BYTES = 2 KiB`, so one submit is - `packets(FRAME_BYTES) + SUBMIT_REPLY_SURBS = 32 + 13 = 45` packets ≈ **5.4 s of - the shim's entire outbound mixnet capacity**. With the two-address failover - list the design assumes (`nym.rs:602-641`, submit goes to *every* address) it - is ~10.8 s. - -The attacker's request body is five bytes — `00 00 00 00 00`, a gRPC frame with -a declared length of zero wrapping a default-valued `RawTransaction`. Five bytes -in; 45 Sphinx packets and 5.4 seconds of a non-substitutable shared resource out. - -### Why this is a *loss* finding and not just a throughput one - -`NymHandle::submit` returns `Ok(())` — and therefore the wallet is answered -`error_code 0` with a txid (`hub.rs:231-240`, `intercept.rs:180-202`) — the -moment the request is accepted into the shim's in-process `requests` channel -(`nym.rs:660-661`), whose capacity is 32 (`main.rs:335`). Behind it sit the -8-slot `out_frames` channel (`main.rs:336`, from which `correlate`'s reserved -permit is taken rather than added) and one in-flight send: **~41 frames of -acknowledged-but-unemitted work**. - -At the crate's own throttled rate that is **~3.7 minutes** (single hub address) -or ~7.4 minutes (two). That figure is a floor rather than a ceiling: behind the -driver's hand-off sits the SDK's own *unbounded* transmission buffer -(`nym_driver.rs:418-432`), so to the extent the SDK accepts faster than it -emits, the backlog accumulates there with no bound and no visibility. - -At the rate the project *measured from a deployed enclave* it is far worse: `hub/src/nym_driver.rs:187-201` records that every -deployed enclave's mixnet client produced **15-25 duplicate fragments** per -32-packet message, "which is how a 32-packet reply that should take ~5 s took -45-90 s from every enclave and timed out". The hub was given -`ZIH_ACK_WAIT_ADDITION_MS` to fix that; **the shim has no equivalent knob and -sets no `DebugConfig` at all in a production build** (`nym_driver.rs:104-157`), -and its submit frame is *larger* than the reply that was measured. On those -numbers a full pipeline is **31 to 62 minutes** of acknowledged, unemitted -migrations — longer than the entire admission window in (a) below. - -Three distinct harms follow. (c) is unconditional; between (a) and (b) the -attacker chooses, and can cause both: - -**(a) Silent conversion into `expiry_too_tight`** — conditional at the crate's -optimistic emission model, certain at the rate the project measured from -deployed enclaves; see the validation section for the split. The hub admits a migration -only if it provably survives the next flush: `expiry >= next_flush_height(tip, 20) + 4` -(`queue.rs:380-393`), i.e. the transaction must reach the hub while the tip is -still 5 to 24 blocks below its expiry height. A librustzcash-default wallet sets -`expiry = build_tip + 40` (`batcher.rs:50-55` names exactly this default), so the -migration has a **16-to-35-block (20-44 minute) window** from being built to -being admitted. The design budgets 6 blocks (7.5 min) of that for delivery — -`MAX_DELIVERY_LAG` (`batcher.rs:46-48`, "wallet-to-shim lag, Nym round trips, -acknowledgement retries and hub failover"), and `BatchParams::validate` asserts -`20 + 4 + 6 <= 40` at startup as though delivery lag were a constant -(`batcher.rs:93-117`). **It is not a constant. It is an attacker-controlled -quantity**: about a byte per second of attacker traffic against a public -endpoint consumes roughly half of that budget at the crate's own optimistic rate -(a ~3.7 min pipeline against a 7.5 min allowance), and all of it and the whole -20-44 minute admission window at the 45-90 s per-frame rate measured from -deployed enclaves. When the frame finally arrives -the hub answers `Refusal::ExpiryTooTight` — into an `AckV1` that the shim -constructs a receiver for and immediately drops (`nym.rs:652`), so nothing -surfaces. The wallet was told success tens of minutes earlier. - -**(b) Removing the stated mitigation of the driver-teardown loss.** The already -filed `shim-nym-driver-every-teardown-path-silently-destroys-acknowledged-submits.md` -records as its principal *limiting* factor: *"at the documented mainnet rate of -~0.77 Orchard-touching transactions per block the pipeline is usually empty, so a -randomly-timed restart usually destroys nothing."* This flood makes the pipeline -**permanently full**, at will, from anywhere. Every ordinary redeploy, every -SIGTERM, every SDK client death, every gateway blip that triggers a rebuild — -each of which already discards everything in flight with no drain and no -accounting — now destroys up to ~41 migrations that wallets were told had -succeeded, instead of zero. An anonymous third party creates the condition; an -entirely routine operator action pulls the trigger. - -**(c) A certain, unconditional harm: every wallet's transaction lookups through -that shim stop working.** With a hub configured the shim routes EVERY -`GetTransaction` to the hub and none to the operator (`intercept.rs:237-247`), -and a lookup travels through the SAME `requests` channel as a submit -(`nym.rs:695-713` — "the request channel carries one type, so it travels in the -same buffer"). One deadline covers both acceptance and reply -(`nym.rs:756-772`) and it is `REQUEST_TIMEOUT = 90 s` (`nym.rs:71`). A -3.7-minute emission backlog exceeds 90 s, so every lookup times out, sweeps its -address list, and fails closed with `UNAVAILABLE`. This depends on no assumption -about wallet expiry, enclave retransmission, or operator restarts: it follows -from the pipeline depth and the timeout alone, at the same ~1 byte/s. - -**Loud versus silent is the attacker's choice, not automatic.** When the -`requests` channel is exactly full, a genuine submit blocks and then fails closed -after `SUBMIT_DISPATCH_TIMEOUT = 5 s` (`nym.rs:80`, `:660-678`) — the wallet sees -`UNAVAILABLE`, which is the safe outcome already described by -`gettransaction-flood-starves-migration-diversion.md`. The *silent* variant needs -the attacker to pace injections so a slot is free when the victim arrives: the -victim is then accepted, told `error_code 0`, and left to rot behind the junk. -That variant is strictly better for an attacker and is the new behaviour here. - -> **Correction, 2026-08-18, from the sibling validation of -> `gettransaction-flood-starves-migration-diversion.md`.** The silent variant does -> **not** in fact require careful pacing, and is the *default* outcome rather than -> the harder one. `send_message` returns when the SDK **accepts** the message, not -> when it is emitted: capacity-1 `InputMessage` channel -> (`client-core/base_client/mod.rs:1013`) → 8-slot batch channel -> (`real_messages_control/mod.rs:150`) → `poll_poisson`, which stores one whole -> message and emits **one packet** per Poisson tick -> (`real_traffic_stream.rs:443-478`) into an uncapped `TransmissionBuffer`. So the -> 32-slot `requests` channel drains at ~8-50 messages/s — far faster than the -> 0.19 submits/s the transport emits — and is usually **not** full. A victim -> arriving during a sustained flood is therefore normally *accepted*, told -> `error_code 0`, and left to rot; the loud `UNAVAILABLE` outcome is the one that -> needs the most attacker concurrency (of order `90 × drain rate` concurrent live -> requests), not the least. This confirms the "41 is a floor, not a ceiling" note -> in Correction 1 below and resolves it: the backlog does accumulate in the SDK's -> unbounded buffer, so the acknowledged-but-unemitted depth is bounded by nothing -> the shim owns. The arithmetic is derived in full in the sibling issue. - -## Attack Scenario and Steps - -1. The attacker picks a target: any operator running a zeronym shim. The endpoint - is a public DNS name serving wallets (`ZIS_TLS_DOMAIN`; `deploy.env.example` - uses `shieldedinfra.net`), and there is no authentication of any kind on it. -2. The attacker opens one HTTP/2 connection and issues - `POST /cash.z.wallet.sdk.rpc.CompactTxStreamer/SendTransaction` with the - five-byte body `00 00 00 00 00` and no `grpc-encoding` header. -3. `route_for` returns `Route::Intercept`; `inspect` reads `declared = 0`, - `message = &frame[5..5]`, `RawTransaction::decode(&[])` succeeds with - `data = vec![]`; `classify_with_evidence(&[])` returns `Class::Unparseable`; - `treat_as_migration()` is true; `divert` is entered with `tx_data = Some()`. -4. `NymHandle::submit` mints a nonce, `wire::encode_submit` produces a 65,536-byte - frame, the request is accepted into `requests`, and the attacker is answered - `error_code 0`. -5. The attacker repeats until ~41 frames are outstanding, then sends one more every - ~5.4 s (or ~45-90 s at the measured enclave rate) to hold the pipeline full. - Sustained cost: **five bytes per five seconds**, roughly one byte per second. -6. Every genuine migration diverted through that shim from then on is answered - `error_code 0` and then queued behind the junk. It is destroyed if - (a) it reaches the hub after its admission window has closed - (`Refusal::ExpiryTooTight`, ack discarded), or - (b) the driver tears down for any reason while it is still queued - (`Step::Stop` / `Step::Rebuild` / `Step::Died`, all of which discard without a - drain). -7. The attacker can raise the odds of (b) directly by also blackholing or simply - flooding the enclave's gateway path, or — if the attacker is the operator — - by restarting the enclave, which is a fully deniable action. - -**Attack Requirements and Assumptions:** - -- **Access needed:** the ability to make TCP connections to a shim's public - wallet-facing listener. No credentials, no wallet, no Zcash funds, no Nym - client, no knowledge of the hub's address, no on-chain activity. -- **Cost:** ~1 byte/second sustained per targeted shim. This is the cheapest - migration-destruction lever found in this audit; the previously filed hub queue - fill requires the attacker to push 64 MiB per 25-minute epoch through their own - throttled Nym clients, and the `GetTransaction` flood requires sustained - concurrency. -- **What makes it realistic:** the payload is not a corner case the developers - overlooked — the shim's own unit tests assert that exactly these bytes take the - divert path, because failing safe toward diversion is the correct privacy - decision. The amplifier is the *consequence* of that correct decision meeting a - fixed-size frame and a serialized 8-packet-per-second transport. -- **What bounds it:** harm (a) requires a wallet expiry near the librustzcash - 40-block default. A ZIP 318 migration carries a 30-to-60-*day* expiry - (`audit-state/SPEC-NOTES.md` §3) and cannot be pushed out of its admission - window this way, so for the acute use case only harm (b) applies. **The clause - that followed here — *"for ZIP 318 traffic harm (b) is worse, because the - wallet's recovery clock is then 30-60 days"* — was REFUTED 2026-08-18 and is - struck; it had the polarity backwards.** 30-60 days is the wallet's automatic- - **retry horizon**, so ZIP 318 traffic recovers from harm (b) *better* than the - ZIP 203-default traffic this shim also diverts, whose horizon is ~50 minutes. - What makes harm (b) severe here is that a *sustained* flood defeats retries of - either class — which is exactly this issue's mechanism (~1 byte/s to hold the - pipeline full). See - `issues/invalid/zip318-canonical-expiry-is-the-only-recovery-clock-and-a-lost-migration-freezes-the-users-notes-for-30-to-60-days.md`. -- The attack does not need the operator's cooperation, but the operator is - strictly better placed to run it and to trigger (b) at a moment of their - choosing. - -> **CORRECTION 2026-08-18 (validation of the cited file — SUPERSEDES the sentence above).** -> The wallet does **not** wait for expiry. Both official Zcash light-wallet SDKs -> automatically resubmit a sent-but-unmined transaction for as long as it remains -> unexpired — the Android SDK at the head of every ~20 s sync loop and after every -> processed block batch (`CompactBlockProcessor.kt:573,615,723`; selection -> `mined_height IS NULL AND expiry_height > ?`), the iOS SDK at most once per 300 s -> (`TxResubmitter.swift:8-15`, `TransactionDao.swift:218-228`) — and the hub's -> payload-hash dedup makes the resend free. The wallet's non-confirmation signal comes -> from compact-block scanning, which the shim does not intercept (`proxy.rs:1068-1074`). -> Expiry is therefore the **retry horizon**, not the wait: ~50 minutes for the ZIP 203 -> default traffic the shim also diverts, 30–60 days for a ZIP 318 migration. A -> *transient* loss self-heals within minutes; only a loss condition that **outlives the -> horizon** destroys the submission permanently — which is exactly what this issue's -> condition does, so this issue's severity is unaffected. Do not write "the user waits -> 30 to 60 days" in the report. Full refutation and the replacement paragraph: -> `issues/invalid/zip318-canonical-expiry-is-the-only-recovery-clock-and-a-lost-migration-freezes-the-users-notes-for-30-to-60-days.md`. - - -## Impact on Users - -A user's wallet is told the mandatory Orchard migration was broadcast and is -handed a txid it displays and records. The transaction reaches no mempool. -Nothing anywhere retains it: the shim keeps no per-migration state by design -(`lib.rs:30-35`), the hub either never received it or refused it and holds no -copy, confirmation tracking is "designed, no code yet", and the ack that carries -the refusal is discarded at construction. - -Because the shim's `SendTransaction` path fails *open* (success) while its -`GetTransaction` path fails *closed* (`UNAVAILABLE`), the user's experience -during an attack is precisely inverted from the truth: sends appear to work and -lookups appear broken. - -The funds are in a pool NU6.3 has closed to new value, the migration is the -user's only route out, and the wallet believes it is done — so the wallet will -not re-broadcast, and will hold the notes as pending-spent until the transaction -expires. - -## Technical Details / Code Analysis - -**The routing predicate admits anyone.** `shim/src/proxy.rs:743-748`: - -```rust -pub fn route_for(path: &str) -> Route { - if path == SEND_TRANSACTION { - return Route::Intercept; - } -``` - -There is no method check (deliberately — rule 3), no authentication, and no rate -limit anywhere between the socket and `intercept::send_transaction`. - -**An empty `RawTransaction` is diverted with `Some` bytes.** -`shim/src/intercept.rs:559-568`: - -```rust - match RawTransaction::decode(message) { - // `data` is the serialized Zcash transaction: the only value the - // classifier ever sees, and the exact bytes the hub broadcasts. - Ok(raw) => { - let evidence = classify_with_evidence(&raw.data); - ( - Inspection::Classified(evidence), - Some(Bytes::from(raw.data)), - ) - } -``` - -and `shim/src/classify.rs:99-101`: - -```rust - pub fn treat_as_migration(self) -> bool { - matches!(self, Class::Migration | Class::Unparseable) - } -``` - -with the crate's own test at `shim/src/classify.rs:332`: - -```rust - assert_eq!(classify(&[]), Class::Unparseable); -``` - -`divert`'s fail-closed guard is `let Some(tx_data) = tx_data else { ... }` -(`intercept.rs:146-155`) — `Some()` passes it. - -**The frame is fixed-size regardless.** `shim/src/wire.rs:273-286`: - -```rust -pub fn encode_submit(nonce: &Nonce, tx: &[u8]) -> Result>, WireError> { - if tx.len() > MAX_NYM_TX_BYTES { ... } - let mut frame = Zeroizing::new(vec![0u8; FRAME_BYTES]); - frame[0..4].copy_from_slice(&SUBMIT_MAGIC); - frame[4..20].copy_from_slice(nonce); - frame[20..24].copy_from_slice(&(tx.len() as u32).to_be_bytes()); - frame[SUBMIT_HEADER_BYTES..SUBMIT_HEADER_BYTES + tx.len()].copy_from_slice(tx); - Ok(frame) -} -``` - -**The crate itself computes the cost of that frame.** -`shim/src/nym.rs:1087-1115`: - -```rust - const PACKET_BYTES: usize = 2 * 1024; - const THROTTLED_PACKETS_PER_SEC: f64 = 1000.0 / 120.0; - ... - fn a_submit_fits_its_dispatch_budget_at_the_throttled_rate() { - let on_wire = packets(wire::FRAME_BYTES) + SUBMIT_REPLY_SURBS as usize; - let secs = seconds_to_emit(on_wire); - assert!(secs < 30.0, ...); - } -``` - -`packets(65536) = 32`, `+ 13 = 45`, `45 / 8.33 = 5.4 s`. The test asserts only -that this is "not absurd"; nothing anywhere asserts that the resource cannot be -consumed by a party other than a wallet. - -**The wallet is answered before anything leaves the process.** -`shim/src/nym.rs:660-689`: - -```rust - match tokio::time::timeout_at(deadline, self.requests.send(request)).await { - Ok(Ok(())) => dispatched += 1, - Ok(Err(_)) | Err(_) => break, - } - } - if dispatched > 0 { Ok(()) } else { Err(NymError::TransportGone) } -``` - -and `shim/src/hub.rs:231-240`: - -```rust - HubTransport::Nym(handle) => match handle.submit(tx_bytes).await { - // ... There is no Refused arm: the hub's verdict is a full round trip - // away and is deliberately not waited for ... so a refusal is never - // surfaced here. - Ok(()) => Ok(Submit::Accepted { txid: crate::nym::local_txid(tx_bytes) }), -``` - -**The pipeline that fills.** `shim/src/main.rs:335-336`: - -```rust - let (req_tx, req_rx) = mpsc::channel(32); - let (out_tx, out_rx) = mpsc::channel(8); -``` - -**The hub's admission rule the delay defeats.** `hub/src/queue.rs:380-393`: - -```rust -pub fn survives_next_flush(expiry: Option, tip: u32, flush_interval: u32, mining_margin: u32) -> bool { - match expiry { - None => true, - Some(expiry) => { - let deadline = next_flush_height(tip, flush_interval).saturating_add(mining_margin); - expiry >= deadline - } - } -} -``` - -**The budget the design asserts as a constant.** `hub/src/batcher.rs:46-48` and -`:93-117`: - -```rust -/// Blocks reserved for wallet-to-shim lag, Nym round trips (measured 9 to 10 s -/// unary), acknowledgement retries and hub failover. -pub const MAX_DELIVERY_LAG: u32 = 6; -``` - -```rust - pub fn validate(&self) -> Result<(), BoxError> { - let spent = self.flush_interval - .saturating_add(self.mining_margin) - .saturating_add(self.delivery_lag); - if spent > self.min_wallet_expiry { ... } -``` - -Six blocks is 7.5 minutes. A full request pipeline is ~3.7 minutes at the -crate's own optimistic rate — half the allowance — and 31-62 minutes at the rate -the project measured from a deployed enclave, which exceeds the whole 20-44 -minute admission window. The attacker sets the depth; the network sets which of -the two rates applies. - -**The measured enclave behaviour the shim never received a fix for.** -`hub/src/nym_driver.rs:187-201`: - -> *"Measured 2026-08-17: a local hub's replies reached a shim with ~1 duplicate -> fragment per lookup; every DEPLOYED hub's reached the same shim with 15-25 -- -> and each duplicate is a full send slot at the throttled rate, which is how a -> 32-packet reply that should take ~5 s took 45-90 s from every enclave and timed -> out."* - -The shim's `build_client` (`shim/src/nym_driver.rs:104-157`) applies no -`DebugConfig` in a production build, so its 45-packet submits carry the -unmitigated version of exactly this behaviour. - -## Recommendations - -1. **Do not spend a mixnet frame on a payload the shim knows is not a - transaction.** `Class::Unparseable` must keep failing safe (never forwarded to - the operator), but "fail safe" does not have to mean "spend 45 Sphinx packets". - Refuse `Unparseable` bodies below a plausible transaction size — or all of them - — with `UNAVAILABLE`/`INVALID_ARGUMENT`, which is fail-closed and costs the - transport nothing. At minimum, refuse a zero-length `RawTransaction.data` - outright; a wallet never sends one. -2. **Bound and account for the divert pipeline.** Publish the current depth of - `requests` + `out_frames` on `/nym-status`, and refuse (fail closed, so the - wallet retries) rather than accept once the depth exceeds what - `MAX_DELIVERY_LAG` can absorb. Accepting work the transport cannot deliver - inside the design's own budget is the defect; a visible `UNAVAILABLE` is - strictly better for the user than a silent success. -3. **Rate-limit the wallet-facing `SendTransaction` path per source.** The shim - already terminates TLS itself and sees the peer address; a token bucket at a - few requests per minute per source refuses nothing a wallet does. -4. **Give the shim the `ack_wait_addition` fix the hub received.** The - measurement at `hub/src/nym_driver.rs:187-201` applies verbatim to the shim's - larger frames, and without it every real-world estimate in this issue should - use the 45-90 s figure rather than the 5.4 s one. -5. **Surface the ack.** `intercept.rs:188` is the only line in the shim that can - report a hub refusal to a wallet and it is unreachable on the deployed - transport. Reading the `AckV1` — even asynchronously, into a counter on - `/nym-status` — turns every silent loss in this issue into an observable one. - The nonce, the waiter and the refusal codes all already exist. - -## Validation Information - -**Verdict: CONFIRMED. Severity: High (as filed).** This is the cheapest attack -found anywhere in the audit and every step of the mechanism was reproduced by -reading the code. Four corrections were applied to the body — the pipeline depth, -the certainty of harm (a), the loud-versus-silent choice, and the promotion of a -harm that is *certain* and was under-stated in the filing. - -### The five-byte path, step by step - -| Step | Verified at | -|---|---| -| The listener is internet-reachable and unauthenticated | `shim/deploy/caution/caution.hcl.tmpl` — `ingress { cidr_ipv4 = "0.0.0.0/0", port = 8083 }` behind a public TLS domain | -| Routing is a pure function of the path, no method check, no auth, no rate limit, no connection cap | `shim/src/proxy.rs:744-747`; `serve_connection` (`:575-609`) and `handle` (`:614-692`) contain no limiter of any kind, and hyper is configured only with window sizes | -| Body buffered under a 4 MiB cap | `shim/src/intercept.rs:101-104` | -| `frame.len() = 5 ≥ GRPC_PREFIX_LEN`, `frame[0] = 0`, `declared = 0` | `shim/src/intercept.rs:516-528` | -| `message = frame.get(5..5) = Some(&[])` and `message.len() == frame.len() − 5` | `shim/src/intercept.rs:529-551` | -| `RawTransaction::decode(&[])` succeeds — proto3 has no required fields, so prost returns the default message | `shim/src/intercept.rs:559-568` | -| `classify(&[]) == Class::Unparseable`, pinned by the crate's own test | `shim/src/classify.rs:330-332` | -| `Unparseable.treat_as_migration() == true`, on purpose | `shim/src/classify.rs:99-101` | -| `divert`'s fail-closed guard passes, because `tx_data` is `Some()` not `None` | `shim/src/intercept.rs:146-155` | -| `encode_submit` allocates and pads to `FRAME_BYTES = 65,536` regardless of `tx.len()` | `shim/src/wire.rs:273-286` and `:72` (`FRAME_BYTES`) | -| The wallet (here, the attacker) is answered `error_code 0` at mixnet hand-off | `shim/src/hub.rs:231-240`, rendered at `shim/src/intercept.rs:180-205` | - -**Five bytes really is the minimum that buys a frame**, which is worth stating -because it shows the boundary is exact rather than approximate: a *zero*-byte -body takes `frame.len() < GRPC_PREFIX_LEN` at `intercept.rs:516-521`, yields -`tx_data = None`, and hits the fail-closed arm at `:146-155`, spending no frame -at all. The 5-byte `00 00 00 00 00` is the smallest input that reaches -`NymHandle::submit`. - -### The emission arithmetic, verified against `shim/src/nym.rs:1085-1115` - -The crate's own `throughput_budget` module supplies every constant: -`PACKET_BYTES = 2 * 1024` (`:1092`), `THROTTLED_PACKETS_PER_SEC = 1000.0 / 120.0` -(`:1094`), `packets(bytes) = bytes.div_ceil(PACKET_BYTES)` (`:1096-1098`), -`SUBMIT_REPLY_SURBS = 13` (`:96`). So - - packets(65_536) = 32; 32 + 13 = 45; 45 / 8.333… = 5.4 s - -per submit — **5.4 seconds of the shim's entire outbound mixnet capacity for -five attacker bytes**, ~13,000× amplification measured in bytes. The egress is -genuinely serialized and non-substitutable: the driver holds exactly one send in -flight (`shim/src/nym_driver.rs:258-272` and the guard at `:362-369`), and -`correlate` reserves capacity ahead of accepting each request -(`shim/src/nym.rs:853-880`). - -Holding the pipeline full then costs one 5-byte request per 5.4 s ≈ **1 byte per -second** of payload (a few tens of bytes/s including HPACK-compressed HTTP/2 -headers on a single reused connection). - -### Correction 1 — the pipeline is ~41 frames, not 42 - -`requests` is `mpsc::channel(32)` and `out_frames` is `mpsc::channel(8)` -(`shim/src/main.rs:335-336`). The permit `correlate` holds -(`shim/src/nym.rs:853-863`) is taken *from* the 8, not in addition to it, so the -acknowledged-but-unemitted depth is 32 + 8 + 1 in flight = **~41 frames**, i.e. -~3.7 min at 5.4 s each. Immaterial to the conclusion; corrected for accuracy. - -Note also, in the other direction, that 41 is a *floor*, not a ceiling. The -driver's send future completes when the **SDK accepts** the message, and the -project's own comment at `shim/src/nym_driver.rs:418-432` describes what sits -behind that acceptance: *"its one-slot input, an 8-deep batch channel, and an -**unbounded transmission buffer** drained at the throttled rate. There is no -drain-then-disconnect in the SDK. Frames in there may include SUBMITS ALREADY -ANSWERED SUCCESS to a wallet."* To the extent the SDK accepts faster than it -emits, the attacker's backlog accumulates there instead, with no bound and no -visibility — which makes both harms below worse, never better. - -### Correction 2 — harm (a) is conditional at the optimistic rate, certain at the measured one - -The admission window is as filed: `survives_next_flush` requires -`expiry >= next_flush_height(tip, 20) + 4` (`hub/src/queue.rs:380-393`), so a -librustzcash-default `expiry = build + 40` must reach the hub while the tip is -16 to 35 blocks past the build height — a 20-to-44-minute window. - -At the crate's own throttled-rate model a full pipeline is **~3.7 minutes**. -That does not by itself exhaust a 20-44 minute window; what it does do is -consume about **half of `MAX_DELIVERY_LAG`** — the 6 blocks (7.5 min) -`batcher.rs:46-48` reserves for exactly this and that `BatchParams::validate` -(`:93-117`) asserts at startup as if it were a constant. So at the optimistic -rate harm (a) destroys migrations that were already within ~4 minutes of their -deadline, rather than all of them. The filing implied certainty; it is -conditional. - -At the emission rate the project **measured from deployed enclaves** it is -certain. `hub/src/nym_driver.rs:187-201` records 15-25 duplicate fragments per -32-packet message from every deployed enclave, *"which is how a 32-packet reply -that should take ~5 s took 45-90 s from every enclave and timed out."* At -45-90 s per frame a full pipeline is **31 to 62 minutes**, which exceeds the -entire admission window. Two facts make this the right rate to plan against, and -one caveat keeps it honest: - -- The mitigation the hub was given for it, `ZIH_ACK_WAIT_ADDITION_MS` - (`hub/src/nym_driver.rs:198-217`), **has no shim equivalent**: verified by - grep, the only `debug_config` call in the shim is behind - `#[cfg(feature = "mixnet-localnet")]` (`shim/src/nym_driver.rs:126-140`), so a - production shim applies no `DebugConfig` at all. -- The mechanism — a fixed retransmission timer against a slow or lossy enclave - ack path — is a property of the enclave's network path, and the shim's submits - are the same 32 packets from the same platform. -- *Caveat:* the measurement is of the hub's **replies**, not of the shim's - sends. Nobody has measured the shim's direction. The issue should not be read - as claiming they have. - -### Correction 3 — the strongest harm is certain and was under-stated - -Independent of wallet expiry and of enclave retransmission behaviour, holding -the pipeline full **breaks every wallet's transaction lookups through that -shim**: - -- With a hub configured, the shim routes **every** `GetTransaction` to the hub - and none to the operator (`shim/src/intercept.rs:237-247`). -- A lookup travels through the **same** `requests` channel as a submit - (`shim/src/nym.rs:695-713`, "the request channel carries one type, so it - travels in the same buffer"). -- One deadline covers both acceptance and reply (`shim/src/nym.rs:756-772`), and - it is `REQUEST_TIMEOUT = 90 s` (`:71`). - -A 3.7-minute emission backlog exceeds 90 s, so the lookup times out whether or -not it is accepted, `each_target` exhausts its addresses, and the shim fails -closed with `UNAVAILABLE`. This costs the attacker the same ~1 byte/s, requires -no assumption at all, and degrades ordinary wallet function for every user of -the targeted shim for as long as the attacker cares to pay. - -### Correction 4 — loud versus silent is a choice, not automatic - -When the `requests` channel is *exactly* full, a genuine submit blocks and then -fails closed after `SUBMIT_DISPATCH_TIMEOUT = 5 s` (`shim/src/nym.rs:80`, -`:660-678`), which the wallet sees as `UNAVAILABLE` — loud, and safe. The silent -variant needs the attacker to pace injections so a slot is free when the victim -arrives, so the victim is *accepted* (told `error_code 0`) and then queued behind -~41 junk frames. Both are available at the same cost and the attacker picks; the -filing read as though the silent variant were automatic. - -### `AVOIDING-FALSE-POSITIVES.md` §5 applied rigorously - -§5's own statement of the real-vulnerability shape is *"amplification attacks -where small input causes disproportionate resource use"*, and its contrasting -"Real Issue" lines are *"1 KB request causing 1 GB memory allocation"* and -*"single connection consuming unbounded resources"*. This is that shape, at an -unusually extreme ratio: **5 bytes in, 65,536 bytes and 5.4 s of a serialized -~8-packet-per-second resource out**, sustained at ~1 byte/s. - -*What resources would the attacker need?* A single TCP connection and about a -byte per second. No credential, no wallet, no Zcash funds, no Nym client, no -knowledge of the hub's address, no on-chain activity. - -*What would stop them?* Nothing in the target and nothing in the deployment. -There is no authentication, no rate limit, no per-source accounting and no -connection cap on the shim's wallet-facing listener; the platform terminates TLS -and forwards, and no limiter is configured anywhere in `caution.hcl.tmpl`. The -one structural bound is the pipeline depth itself, which converts the attack -from unbounded delay into ~3.7 min (optimistic) or ~31-62 min (measured-enclave) -of delay — and both of those are enough for the harms above. - -*And the inversion this target creates.* The guide would normally cap a -throughput attack at Medium. It cannot here, because the shim answers -`error_code 0` at an in-process channel send (`shim/src/hub.rs:231-240`) — so the -"denial" is not visible as a denial. It is a migration the user was told had -succeeded, which the shim does not retain (stateless by design, -`shim/src/lib.rs:32-34`) and which may never reach a mempool. The resource cost -is bytes per second; the damage is silent destruction of transactions the user -believes are spent. - -### The amplifier is the fail-safe working correctly — do not "fix" the classifier - -`Class::Unparseable => treat_as_migration() == true` is right and must stay -right: diverting a body the shim could not read is what prevents leaking one it -could not read. The defect is not the fail-safe; it is that the fail-safe spends -a full fixed-size mixnet frame on a payload the shim already knows is not a -transaction. Recommendation 1 in this issue (refuse a zero-length -`RawTransaction.data`, or `Unparseable` bodies below a plausible transaction -size, with a gRPC error) is fail-*closed* and costs the transport nothing, so it -does not weaken the privacy property at all. - -### Severity justification — High - -*Impact:* for a targeted operator, the privacy-critical divert path and the -whole `GetTransaction` path are disabled, and the `SendTransaction` failure mode -is a **false success** — the user's wallet records a txid for a transaction -that may reach no mempool, in a pool NU6.3/ZIP 258 has closed to new value. -Because the pipeline is held permanently full, every ordinary operator action -that tears the driver down (redeploy, SIGTERM, SDK death, gateway churn) now -discards up to ~41 acknowledged migrations instead of ~0 — a condition an -anonymous third party creates and a routine operator action triggers. - -*Likelihood:* an unauthenticated request to a public DNS name, at ~1 byte/s, -against any operator the attacker chooses, with no detection surface (the shim -publishes no queue depth; `/nym-status` is client-lifecycle only, and -`/nym-diag`'s `sends_dispatched` counts SDK acceptance, so it over-counts in -exactly the loss cases). - -*Why not Critical:* no funds are stolen and no key is compromised; the loss -requires either a wallet expiry near the librustzcash 40-block default or a -teardown to coincide; and ZIP 318 migrations — the acute use case — carry a -30-to-60-day expiry and **cannot** be pushed out of their admission window this -way, so for them only the teardown harm and the lookup outage apply. - -*Why not Medium:* the cost is ~1 byte/s from anywhere on the internet with no -credential, the target is chosen by the attacker, and the primary failure mode -is a false `error_code 0` rather than a visible outage. - -DO NOT DEVIATE FROM THIS TEMPLATE (`docs/SECURITY-ISSUE-TEMPLATE.md`). -LEAVE THIS MESSAGE IN PLACE SO THAT ALL AGENTS KNOW NOT TO DEVIATE FROM THIS TEMPLATE (`docs/SECURITY-ISSUE-TEMPLATE.md`). diff --git a/zeronym-22aa9851caf68-high-medium/high/log-verdict-logs-migration-value-balance-at-info.md b/zeronym-22aa9851caf68-high-medium/high/log-verdict-logs-migration-value-balance-at-info.md deleted file mode 100644 index c8e01a76..00000000 --- a/zeronym-22aa9851caf68-high-medium/high/log-verdict-logs-migration-value-balance-at-info.md +++ /dev/null @@ -1,458 +0,0 @@ -# `log_verdict` logs every transaction's exact value balance, size and expiry at INFO, so the shipped debug deployment hands the operator the amount the README says they never learn - -**Severity**: High -**Validation Status**: Confirmed -**Location**: `audit-target/zeronym/shim/src/intercept.rs:111` (the call site), `:577-637` (`log_verdict`), `:83-87` (`PREFIX_LOG_BYTES`); the default filter at `audit-target/zeronym/shim/src/main.rs:21-30`; the fields' provenance at `audit-target/zeronym/shim/src/classify.rs:275-302`; read against `audit-target/zeronym/shim/src/proxy.rs:812-827`, `audit-target/zeronym/shim/deploy/caution/caution.hcl.tmpl:178-182`, `audit-target/zeronym/hub/REVIEW.md` (rules "Never log a txid…" and #10), `audit-target/zeronym/README.md:33` and `:54` -**Found by agent:** Local (file audit of `shim/src/intercept.rs`); validated 2026-08-18 -**In scope of audit?** Yes — priority area 6, "Log and telemetry discipline" - -## Description - -For **every** `SendTransaction` the shim sees, `intercept::log_verdict` emits a -`tracing::info!` line on target `zis::classify`, before any routing decision. -For the `Class::Migration` arm — that is, for every transaction the shim is about -to divert to the hub — that line carries: - -* `orchard_vb` — the Orchard bundle's **value balance in zatoshis**, i.e. the exact - amount of value leaving the Orchard pool, -* `ironwood_vb`, `sapling_vb` — the same for the other two pools, -* `orchard_actions` — the Orchard action count, -* `expiry` — `nExpiryHeight`, -* `inputs`, `outputs` — transparent input/output counts, -* `tx_len` — the serialized transaction length in bytes, -* `version` — the transaction version. - -The default log filter is `info` (`shim/src/main.rs:26-30`), and the deployed -manifest leaves `RUST_LOG` unset (`caution.hcl.tmpl:178-182`), so this line is -emitted in a deployed enclave with no configuration change. Under -`--debug` the manifest sets `RUST_LOG = "zis::proxy=debug,info"`, whose trailing -`info` directive keeps `zis::classify` at INFO as well — so the line is emitted in -**both** configurations. - -**Where it goes is the whole question, and it is now settled.** Enclave console -output reaches the Nitro parent host — the operator — **only** when the enclave is -launched in debug mode (AWS: `nitro-cli console` "can be used only on an enclave -that was launched with the `--debug-mode` option"; Caution: `debug.enabled` -"Allows reading enclave console output but disables attestation verification"). -The project confirms this in-tree: `shim/deploy/README.md:289` says `/nym-diag` -exists because its data is "neither readable on an attested enclave, which has no -console." - -So the correct framing is **not** "the enclave leaks its logs". It is: - -> the shim writes the one quantity the product promises to withhold into a log -> line at default level, and the repository's shipped deploy default is the -> configuration that delivers that log line to the adversary's disk. - -`deploy.sh:52` is `DEBUG=${DEBUG:-1}` and `deploy.env.example:66` ships `DEBUG=1` -(filed as the parent issue, -`deploy-script-defaults-to-debug-mode-which-turns-attestation-off.md`). Under -that default, `assemble-caution.sh:567-573` opens SSH on the parent host and -`caution.hcl.tmpl:191-193` states this "opens port 22 on the parent so the console -can be read at `/var/log/nitro_enclaves/enclave-console.log`". The operator's -runbook additionally *instructs* operators to flip debug on to diagnose a shim -that "boots but will not serve" (`shim/deploy/caution/OPERATORS.md:358-360`) — -while that enclave is serving real wallets on its real public hostname. - -The result directly contradicts `README.md:33`, which lists under **Not -protected**: "The operator learns *that* a client migrated, **though not the -amount or which transaction**." Under the shipped default they learn both, in -real time, with no chain join required at all. - -The project's own rule forbids exactly this. `hub/REVIEW.md`: - -> Never log a txid, a transaction body, or a per-entry identifier at any level. In a -> Nitro enclave the tracing output reaches the parent via the console […] Log counts, -> reasons and aggregate timings only. - -And the shim's own `proxy.rs:812-819` applies the same premise to a **far less** -sensitive line: - -```rust - // DEBUG, not INFO, and that is a privacy decision rather than a noise one. - // A per-request line naming the method a wallet called is exactly the - // metadata this component exists to deny the operator, and it would be - // sitting in a log file on the operator's box. `RUST_LOG=zis::proxy=debug` - // turns it on for a demo or a debugging session; nothing turns it on by - // default. The classifier's own `zis::classify` lines stay at INFO, because - // in this proof of concept they are the only visible output. -``` - -That last sentence is the defect in one line: the INFO placement of -`zis::classify` is justified by the shim being *a proof of concept whose only -visible output is the verdict*. The shim is now the deployed product -(`README.md:88`: "Deployed: classify and divert … attested Nitro enclaves"), so -the justification has expired while the log level has not. A method **name** is -held at DEBUG on privacy grounds; zatoshi **amounts** sit at INFO. - -## Attack Scenario and Steps - -**The live case — the shipped deploy default.** - -1. The operator deploys with `deploy.sh`'s default (`DEBUG=1`), which passes - `--debug --ssh-key ` and therefore has SSH on the parent host and a - captured console log at `/var/log/nitro_enclaves/enclave-console.log`. -2. A wallet at IP `X` opens a TLS connection to the shim's public endpoint. The - parent host records the source IP and connection timing at the TCP/vsock - layer — a stated residual (`README.md:33`). -3. The wallet sends its Orchard-touching migration. `intercept.rs:111` runs - `log_verdict` *before* the divert at `:117-120`, writing e.g. - `version=V6 orchard_actions=2 orchard_vb=+250000 ironwood_vb=-240000 sapling_vb=+0 expiry=Some(3433150) inputs=0 outputs=0 tx_len=11994 diverted_in_production=true` - to the console. -4. The operator reads the log. **They now hold the amount directly** — no chain - join, no batch analysis, no timing correlation. Joining the log timestamp to - the connection open at the TCP layer yields `IP X → 0.0025 ZEC moved out of - Orchard`, permanently. -5. If they also want the on-chain transaction, every logged field is a **public - field of the transaction once it is on chain**, so the tuple is a near-unique - join key: `orchard_vb` is `valueBalanceOrchard`, `expiry` is `nExpiryHeight`, - `tx_len` is `raw.len()` — the serialized length (`classify.rs:299`) — and - `inputs`/`outputs` are the transparent counts. Twenty blocks later the hub - publishes a shuffled batch; the operator scans it (or the whole chain, at - leisure, forever) for the match. `README.md:34` states the measured modal - batch is zero or one, so the match is unambiguous even on `tx_len` alone. - The shuffle, the 20-block delay, the mixnet hop and the enclave are all - bypassed, because the join key was written to the adversary's own disk before - any of them ran. - -**The attested case, stated honestly and not overstated.** With -`debug { enabled = false }` the parent has **no console channel at all**, so this -line is not readable and this finding is **not live**. The defect that remains in -that configuration is latent: the moment anyone follows -`shim/deploy/caution/OPERATORS.md:358-360` and flips debug on to diagnose a -problem, every wallet using that endpoint has its migration amounts written to -the operator's disk for the duration. - -**Attack Requirements and Assumptions:** -- The attacker is the indexer operator — adversary #1 in the threat model, and the - party the whole component exists to defend against. No special position is - required beyond the one they hold by definition (they own the parent host and - deploy the enclave). -- No exploitation step is needed: the data is written to a file the operator - already has shell access to, by the default deployment path. -- The chain-side half of the join needs no privilege: the chain is public and - permanent, so it can be performed retrospectively, at any time. -- The only thing that would defeat the chain join is many simultaneous migrations - with *identical* value balance, action count, expiry and byte length. - `README.md:34` states the measured modal batch is zero or one. And under the - shipped default the join is not needed anyway, because the amount is logged - directly. - -## Impact on Users - -A user who migrates through a zero-indexer shim is told by the project's own -README that the operator does not learn the amount or which transaction. On the -repository's default deployment the operator learns both, plus the linkage to the -wallet's IP address. That is the single outcome the product is built to prevent -(`README.md:54`: *"Joining them links **IP address to on-chain transaction to -balance**"*), and it is unrecoverable because the chain is permanent and the -console log is a file on someone else's disk. - -This is worse than not deploying the shim in one respect: a user who broadcasts -directly at least knows their indexer sees them. A user behind a zeronym shim -believes they are protected, and `README.md:68` tells them they need "install -nothing and change no setting" to get that protection. - -The persistence matters as much as the operator's own intent: the console log -survives the session on the parent host, so the exposure also reaches anyone who -breaches that host, anyone who obtains a backup, and any legal process served on -the operator. - -## Technical Details / Code Analysis - -**The call site runs on every `SendTransaction`, before any routing decision** -(`shim/src/intercept.rs:99-124`): - -```rust - 99 let (parts, body) = req.into_parts(); -100 -101 // The only buffering in the entire shim, and it is bounded. -102 let collected = match Limited::new(body, MAX_SEND_TX_BYTES).collect().await { -103 Ok(collected) => collected, -104 Err(err) => return Ok(body_read_failed(err)), -105 }; -106 -107 let trailers = collected.trailers().cloned(); -108 let frame = collected.to_bytes(); -109 -110 let (inspection, tx_data) = inspect(&parts.headers, &frame); -111 log_verdict(&inspection, &frame); -112 -113 // A migration bound for the hub is diverted here, and ONLY here does the -114 // operator's indexer stay undialled: … -117 if inspection.treat_as_migration() { -118 if let Some(diversion) = diversion { -119 return divert(&diversion, tx_data).await; -120 } -``` - -**The Migration arm** (`shim/src/intercept.rs:581-601`): - -```rust -581 Inspection::Classified(evidence) => match evidence.class { -582 Class::Migration => tracing::info!( -583 target: "zis::classify", -584 version = %evidence.version, -585 // The deciding fact, first on the line. -586 orchard_actions = evidence.orchard_actions, -587 orchard_vb = %format!("{:+}", evidence.orchard_vb), -588 ironwood_vb = %format!("{:+}", evidence.ironwood_vb), -589 sapling_vb = %format!("{:+}", evidence.sapling_vb), -590 expiry = ?evidence.expiry_height, -591 inputs = evidence.inputs, -592 outputs = evidence.outputs, -593 tx_len = evidence.len, -594 diverted_in_production, -... -601 ), -``` - -The `Class::PassThrough` arm (`:602-615`) logs the identical field set for every -non-Orchard broadcast. - -**Each field is filled directly from the parsed transaction** -(`shim/src/classify.rs:275-302`): - -```rust -275 let orchard_actions = orchard_action_count(&tx); -276 let orchard_vb = tx.orchard_value_balance().orchard_amount().zatoshis(); -277 let ironwood_vb = tx.ironwood_value_balance().ironwood_amount().zatoshis(); -278 let sapling_vb = tx.sapling_value_balance().sapling_amount().zatoshis(); -... -296 expiry_height: tx.expiry_height().map(|height| height.0), -297 inputs: tx.inputs().len(), -298 outputs: tx.outputs().len(), -299 len: raw.len(), -``` - -`orchard_vb` is `valueBalanceOrchard`, a cleartext transaction-level field; -`expiry` is `nExpiryHeight`; `len` is the length of the raw serialized -transaction, i.e. exactly its on-chain serialized length. All are recoverable from -the chain, which is what makes the log line a join key rather than mere telemetry. -The project's own review says so explicitly, in rule #10: *"`orchard_vb` is public -on every diverted transaction … An observer partitions any batch by -`orchard_vb`."* - -**The default filter is `info`** (`shim/src/main.rs:21-30`): - -```rust -21 // `info` deliberately does NOT include the per-request `zis::proxy` line: -22 // that line names the method each wallet called, which is a metadata source -23 // this component exists to deny the operator, and it would live in a log -24 // file on the operator's box. `RUST_LOG=zis::proxy=debug,info` turns it on -25 // when someone is debugging or demoing. `zis::classify` stays at info. -26 tracing_subscriber::fmt() -27 .with_env_filter( -28 tracing_subscriber::EnvFilter::try_from_default_env().unwrap_or_else(|_| "info".into()), -29 ) -30 .init(); -``` - -and the deployed enclave leaves `RUST_LOG` unset -(`shim/deploy/caution/caution.hcl.tmpl:178-182`): - -``` - # Default is `info`, which deliberately omits the per-request zis::proxy - # line naming the method each wallet called. That line is exactly the - # metadata this component exists to deny an operator, so it stays off in - # a deployed enclave. Turn it on only for a local demo, never here. - # RUST_LOG = "zis::proxy=debug,info" -``` - -That comment is the project stating, in the deployment template, that a *method -name* is too sensitive for a deployed enclave's log. `zis::classify` at `info` is -not excluded by that filter and carries orders of magnitude more. Note also that -under `--debug` `assemble-caution.sh:568` uncomments exactly that line, and its -`,info` directive keeps `zis::classify` emitting — so debug mode turns the smaller -leak on *without* turning the larger one off. - -**The smaller instance on the fail-safe arms.** `log_verdict`'s `Unparseable` and -`Failsafe` arms (`shim/src/intercept.rs:616-635`) log `frame_len` and - -```rust -621 body_prefix = %hex_prefix(frame, GRPC_PREFIX_LEN + PREFIX_LOG_BYTES), -``` - -i.e. 13 raw bytes of the request frame plus its exact length, at `warn` (so also -emitted under the default filter). Both arms are reachable **on demand by any -wallet** — send a truncated frame, or a `RawTransaction` whose `data` does not -parse — so this is attacker-triggerable transaction-derived material going to the -same reader. The doc comment at `:83-87` claims those eight bytes "carry the -version and version group id"; on the `Failsafe` arm they are the leading bytes of -the *protobuf* message (tag plus length varint) rather than of the transaction, so -the comment is also inaccurate. - -**Stale message text.** The `Migration` and `PassThrough` arms still say "this PoC -still forwards it; production diverts it to the hub" (`:598-600`). The shim *is* -production. (Filed separately as -`forward-only-log-claims-migration-was-diverted.md`, which concerns the -`diverted_in_production` field's meaning rather than the leaked values.) - -## Recommendations - -1. **Reduce the `Class::Migration` and `Class::PassThrough` arms to what - `hub/REVIEW.md` already requires of the hub: counts, reasons and aggregate - timings.** Concretely, drop `orchard_vb`, `ironwood_vb`, `sapling_vb`, - `expiry`, `inputs`, `outputs` and `tx_len` from both arms, keeping at most - `version`, `orchard_actions` and the verdict. Better still, emit only a - periodic aggregate counter (`migrations_diverted_total`). -2. **Drop `body_prefix` and `frame_len` from the fail-safe arms**, or move both - arms behind the same gate as `zis::proxy`, so a wallet cannot cause raw request - bytes to be logged on demand. -3. **If any per-transaction evidence is genuinely needed for operations, put it - behind the `zis::classify=debug` gate**, so `RUST_LOG` is the single switch and - the deployment template's existing warning covers it. Then remove the "in this - proof of concept they are the only visible output" rationale at - `proxy.rs:818-819`, which no longer holds. -4. **Fix `deploy.sh:52` so `DEBUG` defaults to `0`** (tracked as the parent issue), - so the configuration in which this line is certainly readable is not the - shipped default. -5. **Document the console premise in `README.md`.** Every logging decision in both - binaries rests on "does enclave stdout reach the parent?", the answer is - "only in debug mode", and stating it would let operators reason about the - diagnostic procedure in `OPERATORS.md:358-360` correctly. - -## Validation Information - -**Verdict: CONFIRMED. Severity: High.** - -Every mechanical claim was re-verified against the target: - -| Claim | Verified at | -|---|---| -| `log_verdict` called on every `SendTransaction`, before routing | `intercept.rs:110-111`, divert at `:117-120` | -| Migration arm logs all nine fields at INFO | `intercept.rs:582-601` | -| PassThrough arm logs the same set at INFO | `intercept.rs:602-615` | -| Fail-safe arms log 13 raw frame bytes + length at WARN, wallet-triggerable | `intercept.rs:616-635`, `:66`, `:87` | -| `orchard_vb` = `valueBalanceOrchard`; `len` = raw serialized length | `classify.rs:276`, `:299` | -| Default filter is `info`; deployed manifest leaves `RUST_LOG` unset | `main.rs:26-30`; `caution.hcl.tmpl:182` | -| `--debug`'s `RUST_LOG="zis::proxy=debug,info"` keeps `zis::classify` at INFO | `assemble-caution.sh:568`; `EnvFilter` semantics (trailing `info` is the global default directive) | -| Method name held at DEBUG for privacy, `zis::classify` left at INFO because "in this proof of concept" | `proxy.rs:812-819` | -| Project's counts-only rule | `hub/REVIEW.md`, "Never log a txid, a transaction body, or a per-entry identifier at any level" | -| `orchard_vb` publicly partitions a batch — the project's own words | `hub/REVIEW.md` #10 | -| README claims the operator does not learn the amount | `README.md:33` | -| Console readable on the parent **only** in debug | Coordinator open item 7 (AWS + Caution vendor docs); corroborated in-tree at `shim/deploy/README.md:289` and `caution.hcl.tmpl:191-193` | -| Runbook instructs operators to enable debug for diagnosis | `shim/deploy/caution/OPERATORS.md:358-360` | - -**Corrections made during validation.** The draft's "Case B" left the attested-mode -console premise open. It is no longer open: coordinator open item 7 resolved it -from two vendor sources, and `shim/deploy/README.md:289` states the same thing -in-tree. The issue has been rewritten so that: - -- the attested configuration is described as **not** leaking this line, plainly - and without hedging — this finding must not be cited as evidence that an - attested enclave leaks logs; -- the live exposure is attributed to its actual cause, the `DEBUG=1` deploy - default, with the parent issue cited; -- a second live route is added that the draft missed and that survives a fix to - the default: `shim/deploy/caution/OPERATORS.md:358-360` **instructs** operators - to enable debug on a shim that boots but will not serve, and that shim is - serving real wallets on its real hostname while the console is open. - -Two claims were also strengthened, both verified: under `--debug` the manifest's -`RUST_LOG="zis::proxy=debug,info"` does **not** suppress `zis::classify`, so debug -mode adds the method-name leak on top of this one; and the fail-safe arms' -`body_prefix` is reachable **on demand by any wallet**, which makes that sub-case -attacker-triggerable rather than incidental. - -**Severity justification — High, and how it moves if the parent issue is fixed.** - -*Impact:* the operator obtains the exact zatoshi amount, size and expiry of every -transaction a wallet broadcasts through the shim, in real time, next to the TCP -connection that carries the wallet's source IP. That is the precise linkage -`README.md:54` calls "the attack" and the precise quantity `README.md:33` promises -is not learned. It affects every user of the endpoint, it is retrospective and -permanent, and it needs no chain analysis at all. - -*Likelihood:* high **as the system ships today**, because the exposure is gated on -debug mode and debug mode is the default of `deploy.sh` and -`deploy.env.example`. `docs/AVOIDING-FALSE-POSITIVES.md` §4 and §7 both warn -against grading debug-only leaks highly — and both name the identical exception -that applies here: the leak is real when the insecure mode is **enabled by -default**, which it is. - -*If `deploy.sh` is fixed to `DEBUG=0` but this log line is left alone*, the -correct grade becomes **Medium**: the sanctioned diagnostic procedure in -`OPERATORS.md:358-360` still opens the console on a shim serving live wallet -traffic, so the exposure remains reachable by a documented operator action rather -than by a default. That conditional is stated here so the report can present the -two fixes as independent and both necessary. - -*Why not Critical:* no funds move; the exposure requires the debug configuration -rather than being present in the attested deployment; and it is bounded to users -of the affected endpoint. - -**Relationship to the parent issue.** This is the blast radius of -`deploy-script-defaults-to-debug-mode-which-turns-attestation-off.md`, but it is a -**separate defect with a separate fix**: inverting the `DEBUG` default does not -remove these fields from the log line, and removing the fields does not restore -attestation, close SSH, or suppress the `zis::proxy` method log. Both fixes are -needed and neither substitutes for the other. Per coordinator open item 7 the -report should present them as parent and blast radius, in that order. - - ---- - -## [ADDENDUM — Global auditor, focus area G18 (log and telemetry discipline as one policy), 2026-08-18. The CONFIRMED verdict and severity are untouched. Two corrections to the RECOMMENDATIONS, and one to the scenario bound.] - -**(a) Recommendation 1 is incomplete: the minimal fix leaves a per-migration -arrival feed intact.** Stripping the fields from the `Class::Migration` arm -removes the *amount*, not the *event*. Two timestamped INFO lines are emitted per -diverted migration, not one: - -- `shim/src/intercept.rs:582-601` — `log_verdict`'s `Class::Migration` arm, the - subject of this issue; and -- `shim/src/intercept.rs:197-201`, inside `divert`, after the hub answers: - -```rust - tracing::info!( - target: "zis::classify", - accepted = error_code == 0, - "migration diverted to the hub" - ); -``` - -The second line survives every fix proposed above and is, on its own, exactly -what `hub/REVIEW.md` #157 forbids on the other binary: a per-entry event with a -timestamp. It is the shim-side twin of -`hub-per-admission-info-log-is-a-real-time-per-entry-arrival-feed.md`, on the side -where the parent host is the primary adversary rather than a third party, and it -additionally carries the hub's verdict (`accepted`) — so a reader of the console -learns not only that a migration was diverted at time T but whether the hub took -it. Recommendation 1 should be extended to: **reduce both call sites to a single -periodic aggregate** (`migrations_diverted_total`, `diverts_refused_total`), which -is the form REVIEW #157 permits and which `hub/src/batcher.rs:396-403` already -uses correctly. - -**(b) A scoping refinement that makes the fix cheaper, and which should be stated -so it is not lost: the `Class::PassThrough` arm is not a leak.** A pass-through -transaction is forwarded to the operator's own indexer in full -(`shim/src/intercept.rs:127-131`), so its value balances, expiry, counts and -length are already in the operator's hands from the transaction itself. Logging -them adds nothing. The arm that must lose its fields is `Class::Migration`, plus -the two fail-safe arms (which are also diverted). Removing the fields from the -`PassThrough` arm as well is still worth doing for uniformity and to keep the two -arms from drifting, but it is not the security-bearing half, and a maintainer -weighing the diagnostic value of the line should know which half is which. - -**(c) The Case B bound is narrower than "not live in an attested deployment".** -This issue correctly records that with `debug { enabled = false }` the parent has -no console. What was not established at the time is that the console is closed by -a **parent-side launch flag**, not by attestation: `--debug-mode` is appended to -`nitro-cli run-enclave` by a systemd unit on the parent's disk -(Caution platform, `terraform/modules/aws/nitro-enclave/user-data.sh:166`), the -EIF is at `/opt/nitro/enclave.eif` (`:52`), `aws-nitro-enclaves-cli` is installed -unconditionally (`:15`), and `debug.ssh_keys` — which is **not** gated on -`debug.enabled` (`src/api/src/deployment.rs:2158-2164`) — puts the operator's key -on the parent's default administrative account. So an operator who set an SSH key -in an otherwise fully attested manifest can re-launch the same image in debug -mode and read this exact line, with `RUST_LOG` still at the measured `info`. That -is filed separately as -`attested-enclave-console-is-reopenable-from-the-parent-because-debug-mode-is-a-launch-flag-and-ssh-keys-is-not-gated-on-it.md` -(Medium) rather than being folded in here, so the harm is not counted twice — but -the two must be read together, because it is the reason Recommendation 1 should -not be deferred on the grounds that attestation contains this line. - - -DO NOT DEVIATE FROM THIS TEMPLATE (`docs/SECURITY-ISSUE-TEMPLATE.md`). -LEAVE THIS MESSAGE IN PLACE SO THAT ALL AGENTS KNOW NOT TO DEVIATE FROM THIS TEMPLATE (`docs/SECURITY-ISSUE-TEMPLATE.md`). diff --git a/zeronym-22aa9851caf68-high-medium/high/shim-submits-every-migration-to-every-configured-hub-so-an-operator-appends-their-own-and-gets-a-plaintext-copy-with-nothing-breaking.md b/zeronym-22aa9851caf68-high-medium/high/shim-submits-every-migration-to-every-configured-hub-so-an-operator-appends-their-own-and-gets-a-plaintext-copy-with-nothing-breaking.md deleted file mode 100644 index 0a06268c..00000000 --- a/zeronym-22aa9851caf68-high-medium/high/shim-submits-every-migration-to-every-configured-hub-so-an-operator-appends-their-own-and-gets-a-plaintext-copy-with-nothing-breaking.md +++ /dev/null @@ -1,535 +0,0 @@ -# `NymHandle::submit` sends every diverted transaction to **every** address in `ZIS_HUB_NYM`, so an operator appends a hub they run and receives a real-time plaintext copy of every migration while the canonical hub keeps publishing normally — nothing breaks, and the one check the audit recommends elsewhere does not catch it - -**Severity**: High -**Validation Status**: Confirmed -**Location**: `audit-target/zeronym/shim/src/nym.rs:595-689` (`NymHandle::submit`, the fan-out), `:602`, `:618-634` (the code's own statement of what makes it safe); `audit-target/zeronym/shim/src/nym_driver.rs:180` (`targets` = number of configured addresses); `audit-target/zeronym/shim/src/config.rs:65-77` (`ZIS_HUB_NYM` is an unbounded list) and `:262-289` (the only validation); `audit-target/zeronym/shim/src/hub.rs:228-240` (the divert call site); `audit-target/zeronym/shim/deploy/caution/assemble-caution.sh:228-235` (per-entry shape check, no cap) and `:345-356` (the rendering, and a comment that misdescribes the fan-out); `audit-target/zeronym/README.md:90` (the public claim that submit *rotates*); `audit-target/zeronym/README.md:70` (operators may run a hub) -**Found by agent:** Global, focus area G30/G32/G17 — the indexer operator's full capability sweep after the `unit.env`/certfp reversals -**In scope of audit?** Yes — `shim/src/nym.rs` is priority area 4, `README.md` claims are in scope as security claims, and `AUDIT-INSTRUCTIONS.md` unverified lead #10 names send-to-all-targets directly. - -## Description - -`ZIS_HUB_NYM` is a comma-separated **list** of hub Nym addresses, and -`NymHandle::submit` sends the transaction to **all** of them -(`shim/src/nym.rs:602`, `:642`): - -```rust - // Send to EVERY configured hub address, not one (REVIEW #6). - … - for target in 0..targets { -``` - -The list exists for failover: a diskless hub mints a new Nym address on every -restart, so a shim carries "the current address and the one it just rotated away -from" (`shim/src/config.rs:70-75`, `shim/src/nym.rs:622-626`). The code is -explicit that this is only safe under that assumption -(`shim/src/nym.rs:622-634`): - -> Sending to every address is therefore safe **only while the other addresses -> are DEAD** … If two hubs were ever live at once, both would broadcast the same -> transaction. On-chain that is harmless (the second is a known txid), but it -> would publish the migration in two different batches at two different moments -> … and **it doubles the number of enclaves holding the plaintext**. Running a -> hot standby therefore needs an explicit passive mode in the hub, which does not -> exist today. - -Nothing enforces the assumption. `Config::hub_selection` rejects only malformed -entries and exact duplicates (`shim/src/config.rs:276-282`); there is no cap on -the count, no notion of a primary, and no way for the shim to know whether an -address is live. `assemble-caution.sh:228-235` applies the same shallow shape -check per entry and never counts them. - -The operator writes that list. So an operator who wants the plaintext does not -need to *repoint* the shim at a hostile hub — a repoint is a functional change, -and a hub that fails to batch or publish correctly produces visible breakage -that a user, a wallet developer or a monitoring third party could notice. They -**append**: - -``` -ZIS_HUB_NYM = "," -``` - -Every diverted transaction is then framed twice, with a fresh nonce each time -(`shim/src/nym.rs:645`), and dispatched to both. The canonical hub queues, -batches, shuffles and publishes exactly as before, so **the wallet's transaction -confirms on the normal cadence and nothing anywhere behaves differently**. The -operator's hub receives the identical `SubmitV1` frame, decodes it, and holds the -raw transaction in plaintext — which is what a hub is *for*; it needs no -modification, no vulnerability and no enclave. It is `zero-indexer-hub`, running -on a VPS, from the public repository. - -Two independent harms follow, and the second does not need the operator to know -any wallet's IP address: - -1. **Real-time per-migration content disclosure to adversary #1.** The operator - holds the full transaction bytes of every Orchard-touching transaction their - users send, milliseconds after the wallet sends it — the exact thing - `THREATMODEL.md` P1 says they do not receive and `README.md:83` says "goes to - the hub instead of the operator". -2. **Anonymity-set subtraction against every *other* operator's users.** The - batching design's whole product is that a published batch is the union of many - operators' migrations and nobody can attribute a member. An operator with a - listening hub knows precisely which members of every published batch came from - *their own* shim, and can subtract them. For the other operators' users the - effective anonymity set shrinks by exactly that number, permanently and - retrospectively, with no signal to anyone. This is the SYBIL set-subtraction - attack in `THREATMODEL.md` §3, executed for free by a party the design already - admits to the system. - -And when the operator also holds the wallet leg — which they do, either by owning -the Nitro parent host (BYOC) or by the layer-4 relay on the DNS record they -control (`operator-controlled-dns-permits-a-layer-4-relay-that-every-documented-verification-step-passes.md`) -— the two halves compose into the product's headline threat with no analysis -required: the wallet's source IP and connection timestamp on one side, the exact -transaction bytes arriving at their own hub milliseconds later on the other, -matched by time and by length. No batch reasoning, no chain join, no statistics. - -## Attack Scenario and Steps - -Attacker: the indexer operator. `README.md:70` already contemplates that they -*"run the shim in front of their indexer, and **optionally a hub**"*. - -1. The operator runs `zero-indexer-hub` — the published binary, unmodified, on - an ordinary VM — and reads its address from `GET /nym-address`. The hub has - no submitter allow-list and no authentication by design, so it will accept - frames from any shim, including their own. -2. In `deploy.env` they set - `HUB_NYM=","`. `deploy.sh:113` passes it through as - `--hub-nym`; `assemble-caution.sh:228-235` accepts both entries; - `:352` renders `ZIS_HUB_NYM = ","` into `unit.env`. -3. They deploy attested, publish the app-source, and run `caution verify`. All - three PCRs reproduce, the TLS certificate binding verifies, and - `✓ Attestation verification PASSED` prints — correctly. The shim is the - genuine shim; it is simply configured with two hubs, which is a configuration - the code, the CLI and the runbook all accept. -4. Every wallet that migrates through this endpoint has its transaction - delivered to the canonical hub **and** to the operator's hub. -5. The canonical hub publishes on the normal 20-block cadence. The migration - confirms. The operator's hub also publishes it; the network answers "already - known" and `classify_publish_error` records that as achieved, so even the - duplicate broadcast produces no error anywhere. Because both hubs use the same - `FLUSH_INTERVAL_BLOCKS = 20` cadence, the two publications are essentially - simultaneous, so a chain observer sees nothing anomalous either. -6. The operator reads the plaintext off their own hub — from `POST /transaction` - (`hub/src/queue.rs:328-348`, `Queue::find_by_txid`, which returns the raw - transaction bytes of an entry that has not yet been published), or simply from - a one-line patch, since it is their process on their own machine. (Not from the - hub's logs: `hub/src/server.rs:269` logs `parseable = ` and nothing else, - which is the counts-only discipline working as intended.) - -**Attack Requirements and Assumptions:** - -- **Access needed:** the operator's own configuration, plus one VM. No enclave, - no Caution account for the second hub, no code change to either binary, no - mixnet position, and no vulnerability. -- **Detectability — the important part, stated precisely.** The value **is** - measured into PCR0/PCR1 (open item 6q) and **is** served verbatim in - `.manifest.run_command` of every `/attestation` response, so this is - detectable *in principle*. But: - - No zeronym document tells anyone to read `ZIS_HUB_NYM`. This is already - filed as - `auditor-recipe-omits-the-two-checks-that-decide-where-plaintext-goes-and-names-a-defence-the-platform-does-not-rely-on.md`. - - **The remediation that issue recommends does not catch this.** Its - recommendation 2 is to publish the canonical hub address so that checking - becomes "a string comparison rather than a judgement call". A checker asking - *"does `ZIS_HUB_NYM` name the canonical hub?"* answers **yes** for an - appended list. Only exact equality of the whole list catches it, and only if - the checker knows that a list of two is not the ordinary failover state the - code documents as normal. - - **The public documentation actively points the wrong way.** - `README.md:90` states: *"multi-hub failover. The shim **rotates which hub - address each submit targets**"*. That is false for `submit` — rotation - (`each_target`, `shim/src/nym.rs:733-760`) applies to *lookups*; submits go - to all addresses unconditionally. A reader who checks the README before - judging a two-address list concludes each migration reaches one hub. - `assemble-caution.sh:351` repeats the error inside the manifest the auditor - reads: *"The driver tries each address until one acks."* Submit awaits no - ack at all. - - A two-address list is exactly the shape the design says to expect - (`shim/src/config.rs:70-75`), so its presence is not itself suspicious. -- **What makes it realistic:** it is the *cheapest* way for an operator to - obtain migration plaintext, and uniquely it has **no functional side effect** — - unlike repointing, it cannot break a user's transaction, cannot be caught by - any smoke test, and leaves the system's end-to-end behaviour identical. -- **What limits it:** the operator's own hub is a second enclave-less process, so - the transaction plaintext also sits outside any TEE; if that is discovered it is - unambiguous. And an auditor who reads the manifest *and* knows to require exact - equality does catch it. - -## Impact on Users - -For every wallet using an endpoint configured this way: - -- The operator obtains the complete bytes of every Orchard-touching transaction - the wallet sends, at the moment it sends it. `THREATMODEL.md` P1 ("OPERATOR - does not receive the contents of your Orchard-touching transaction") and - `README.md:83` ("goes to the hub instead of the operator") do not hold. -- `THREATMODEL.md` C2 — *"The party that sees your transaction in the clear and - the party that sees your IP address are two different parties"* — is the - invariant this destroys most directly, and it is the one the whole two-component - architecture exists to create. -- With the wallet leg (parent host or DNS relay), the IP → transaction → amount - linkage is direct and certain rather than statistical. - -For **every other operator's users**, whose wallets never touched this endpoint: -the anonymity set of every published batch shrinks by the number of members this -operator contributed, because those members are known to them and can be -subtracted. At the project's own measured rate (0.77 Orchard-touching -transactions per block, modal batch 0–1) that is frequently the whole batch. - -## Technical Details / Code Analysis - -### The fan-out - -`shim/src/nym.rs:595-601`: - -```rust - pub async fn submit(&self, tx_bytes: &[u8]) -> Result<(), NymError> { - let targets = self.targets.load(Ordering::Relaxed); - if targets == 0 { - // No hub address to send to: nothing was dispatched. Fail closed. - return Err(NymError::TransportGone); - } -``` - -`shim/src/nym.rs:639-676`, the loop, in full: - -```rust - let deadline = tokio::time::Instant::now() + self.dispatch_timeout; - let mut dispatched = 0usize; - - for target in 0..targets { - // A FRESH nonce per address: two hubs answering the same nonce would be - // indistinguishable to the correlator, and the ack is unread anyway. - let nonce = fresh_nonce(); - … - let frame = wire::encode_submit(&nonce, tx_bytes).map_err(NymError::Encode)?; - let (ack_tx, _drop_receiver) = oneshot::channel(); - let request = Request { - nonce, - frame, - reply_surbs: SUBMIT_REPLY_SURBS, - waiter: Waiter::Ack(ack_tx), - target, - }; - match tokio::time::timeout_at(deadline, self.requests.send(request)).await { - Ok(Ok(())) => dispatched += 1, - Ok(Err(_)) | Err(_) => break, - } - } -``` - -`targets` is set once from the configured list -(`shim/src/nym_driver.rs:180`): - -```rust - targets.store(hub_addresses.len(), Ordering::Relaxed); -``` - -so `targets` is exactly the number of comma-separated entries in `ZIS_HUB_NYM`. -`tx_bytes` is the same buffer for every iteration; each address receives the -whole transaction. - -### The only validation - -`shim/src/config.rs:272-286`: - -```rust - let mut seen: Vec<&str> = Vec::new(); - for addr in &addresses { - if !is_nym_address(addr) { - return Err(ConfigError::MalformedNymAddress((*addr).to_owned())); - } - if seen.contains(addr) { - return Err(ConfigError::DuplicateNymAddress((*addr).to_owned())); - } - seen.push(addr); - } - Ok(HubSelection::Nym( - addresses.iter().map(|addr| (*addr).to_owned()).collect(), - )) -``` - -`is_nym_address` is a shape check for `identity.encryption@gateway` -(`shim/src/config.rs:297-315`). Distinct well-formed addresses are accepted in -any number. - -### The deploy tooling agrees, and misdescribes the result - -`shim/deploy/caution/assemble-caution.sh:228-235`: - -```sh - OLDIFS=$IFS; IFS=',' - for addr in $HUB_NYM; do - case "$addr" in - ?*.?*@?*) : ;; - *) echo "error: --hub-nym entry '$addr' is not identity.encryption@gateway" >&2; exit 2 ;; - esac - done - IFS=$OLDIFS -``` - -and `:345-352`, which writes both the value and the incorrect gloss into the -manifest an auditor reads: - -```sh - # ZIS_HUB_NYM is the address list; the driver picks a live one and fails over - # (D10). … - { - printf '\n # Divert Orchard-touching transactions over the Nym mixnet to these hub\n' - printf ' # addresses. The mixnet is the confidentiality boundary; there is no TLS\n' - printf ' # name to verify on this hop. The driver tries each address until one acks.\n' - printf ' ZIS_HUB_NYM = "%s"\n' "$HUB_NYM" -``` - -*"picks a live one and fails over"* and *"tries each address until one acks"* are -both descriptions of `each_target` (the **lookup** path, -`shim/src/nym.rs:733-760`), not of `submit`. `submit` neither picks nor waits. -`PROVENANCE` likewise records only `hub(s): $HUB_NYM` -(`assemble-caution.sh:595`) with no comment on multiplicity. - -### Why the second hub is invisible downstream - -- **Dedup is per-hub.** `shim/src/nym.rs:618-621` states it: each hub - deduplicates its *own* queue on the payload hash; there is no cross-hub - dedup, and the hub "has no notion of being active or standby, so any hub that - RECEIVES a migration will queue and broadcast it." -- **The duplicate broadcast is benign and silent.** Both hubs flush on heights - ≡ 0 mod `FLUSH_INTERVAL_BLOCKS` (20), so the two publications land at the same - boundary; whichever is second gets an "already known" response, which - `chain::classify_publish_error` maps to `AlreadyKnown` and `batcher::flush` - counts as achieved. -- **The shim's own telemetry cannot show it.** `/healthz` and `/nym-status` are - mixnet-client lifecycle only; the shim's per-migration log line reports - `accepted = error_code == 0`, which on the mixnet path is always true. -- **The wallet cannot show it.** The txid it is shown is computed by the shim - from its own bytes (`THREATMODEL.md` N2), and the transaction confirms. - -## Recommendations - -1. **Make multi-hub an explicit, loud decision rather than a silent one.** At - startup, if `hub_selection()` yields more than one Nym address, emit a - `warn!` naming every address and stating that **each migration is sent to all - of them in plaintext**. Better: require an explicit - `--allow-multiple-hubs` / `ZIS_ALLOW_MULTIPLE_HUBS=true` before accepting a - list of length > 1, so that the failover shape the design intends is chosen - deliberately and appears in the measured `unit.env` where an auditor can see - the intent as well as the addresses. -2. **Correct `README.md:90` and `assemble-caution.sh:345`/`:351`.** Submits go to - every address; only lookups rotate. Both texts currently tell a reader the - opposite, and both are the texts someone would consult when judging a - two-address list. -3. **Publish the canonical hub address *and* state that the whole list must - equal it.** This is the strengthening of recommendation 2 of - `auditor-recipe-omits-…`: the check must be exact-list equality, not - membership. Add it to `README.md:71` and to both `OPERATORS.md` "Verify" - sections, as: - `curl -sX POST https:///attestation -d '{"nonce":"…"}' | jq -r '.manifest.run_command' | grep ZIS_HUB_NYM` - with the expected line published verbatim. -4. **Implement the passive/standby mode the code says is missing** - (`shim/src/nym.rs:633-634`), so the failover list can be satisfied by one - *active* address and the fan-out stops being the mechanism. This removes the - capability rather than documenting it. -5. Consider having the hub refuse to broadcast a transaction it can see is - already published, so a second live hub is at least detectable on chain — a - weaker measure than 4, listed because it needs no shim change. - -Cross-references: -`auditor-recipe-omits-the-two-checks-that-decide-where-plaintext-goes-and-names-a-defence-the-platform-does-not-rely-on.md` -(the missing check; this issue shows its recommended form is insufficient); -`shim-config-hub-identity-is-unattested-unobservable-operator-configuration.md` -(the repoint variant, with its post-reversal correction); -`nym-submit-fanout-always-starts-at-address-zero-and-reports-success-on-a-partial-sweep.md` -(the same function, for *partial* fan-out causing loss — disjoint from this); -`operator-controlled-dns-permits-a-layer-4-relay-that-every-documented-verification-step-passes.md` -(the wallet-leg half of the composition). - -## Validation Information - -**Validated 2026-08-18. CONFIRMED at High.** Every load-bearing claim was -re-derived from the audit target and, for the platform half, from the Caution -platform's own public source (cloned during this audit from -`codeberg.org/caution/platform`, `git.distrust.co/public/bootproof.git`, and -`aws-nitro-enclaves-image-format 0.4.0`). This issue was written *after* the -`unit.env`/certfp reversals and was re-tested against them specifically, because -those reversals invalidated the sibling repoint finding. - -### 1. The fan-out is real, and it is unconditional - -Verified line by line in `shim/src/nym.rs:595-689`. `submit` reads -`self.targets` (set once at `shim/src/nym_driver.rs:180` to -`hub_addresses.len()`) and loops `for target in 0..targets`, framing the same -`tx_bytes` with a fresh nonce for each and handing every frame to the driver. -There is no primary, no ack wait, and no early exit on success — the only `break` -arms are transport failure. `shim/src/nym_driver.rs:613` resolves `out.target` -into `hub_addresses[target]` and sends anonymously. - -Verified that `each_target` (`shim/src/nym.rs:733-784`), the *rotating* sweep, is -called from exactly one place — `get_transaction` at `:696`. **Submits never -rotate.** This is the fact that makes `README.md:90` and -`assemble-caution.sh:345`/`:351` false, and both were read in place. - -### 2. The list is uncapped at every layer - -- `shim/src/config.rs:65-77`: `ZIS_HUB_NYM` is `Vec` with - `value_delimiter = ','`. -- `shim/src/config.rs:250-289` (`hub_selection`): shape check via - `is_nym_address` plus a reject-exact-duplicates pass. No cap, no primary, no - identity check. Confirmed by reading the function in full. -- `shim/deploy/caution/assemble-caution.sh:228-235`: per-entry - `case "$addr" in ?*.?*@?*)` shape check inside a comma-split loop. No count, - no cap. -- `deploy.sh:113`: `HUB_NYM` is passed straight through as one `--hub-nym` - argument. - -### 3. Nothing breaks downstream — re-verified - -- `hub/src/chain.rs:513-533` (`classify_publish_error`) folds hyphens and maps - `already known` / `already in mempool` / `already in block chain` / `duplicate` - to `Publish::AlreadyKnown`, and `hub/src/chain.rs:94` states that - `AlreadyKnown` is a success. The hub's own test at `:559-574` is named - `duplicate_submissions_are_success_not_failure` and its comment reads - *"Every shim submits to every hub, so duplicates are normal operation."* -- The hub has no submitter ACL on the Nym ingress (`hub/src/nym.rs:305-343` - decodes a `SubmitV1` and calls `hub.admit` with no notion of who sent it), - which is a stated design property, so the operator's second hub accepts their - own shim's frames without modification. - -### 4. THE KEY TEST — does the attack survive the `unit.env` reversal? **Yes, completely.** - -The reversal is real and was re-derived here from source rather than taken on -trust: - -- `src/caution-config/src/lib.rs:233-274` — `UnitConfig::run_command_string()` - emits `export KEY=` for every `Expression::String` entry in - `unit.env`. -- `src/api/src/main.rs:2413-2417` → `src/enclave-builder/src/build.rs:355-361`, - `:468` — that string becomes `{{USER_CMD}}` in the generated `run.sh`. -- `src/enclave-builder/templates/Containerfile.eif` — `COPY run.sh /build/run.sh`, - `RUN cp /build/run.sh /build/initramfs/run.sh`, `cpio … | gzip > rootfs.cpio.gz`, - then `eif_build … --ramdisk /build/rootfs.cpio.gz` with **exactly one** - `--ramdisk`. -- `aws-nitro-enclaves-image-format-0.4.0/src/utils/mod.rs:660-691` — that ramdisk - is written into `image_hasher` (PCR0) and, being index 0, into - `bootstrap_hasher` (PCR1). - -So appending an address **does** change PCR0/PCR1. **That changes nothing about -this attack**, for one decisive reason: `deploy.sh:206-219` publishes the -*deployed* tree — the one containing `ZIS_HUB_NYM = ","` — to -`APP_SOURCE`, and `caution verify` reproduces PCRs from *that* tree. The -measurement therefore reproduces exactly and `✓ Attestation verification PASSED` -prints (`src/cli/src/lib.rs:7255-7305`, read in full). **The measurement -discloses the appended address; it never detects it.** Detection requires a human -to read a value and compare it against a reference — and see §5. - -The value is disclosed in three places, all of which were confirmed: the -published `caution.hcl` in the app-source tree; `.manifest.run_command` in every -`/attestation` response (`bootproofd`'s `NoncedAttestationResponse` carries the -build manifest alongside the signed document); and `PROVENANCE` -(`assemble-caution.sh:595` writes `hub(s): $HUB_NYM`, plural, without comment). -None of these is read by any check the project runs or documents. - -### 5. The distinguishing result: this defeats the fix proposed for the sibling issue - -`auditor-recipe-omits-the-two-checks-that-decide-where-plaintext-goes-…` -recommends publishing the canonical hub address so the check becomes a string -comparison. Confirmed that this does **not** catch an appended list, and for -three independent reasons rather than one: - -1. A *membership* test ("does `ZIS_HUB_NYM` name the canonical hub?") answers - **yes**. Only exact whole-list equality catches it. -2. A two-address list is the documented normal state, twice over. - `shim/src/config.rs:70-75` gives two separate innocent reasons for a list — - *"a diskless hub mints a new address on every restart, so shims carry the - current and the just-rotated one at once"* **and** *"one hub is hosted at - several gateways for uptime"*. A reviewer seeing two entries has two - documented benign explanations before they reach a hostile one. -3. **The documentation tells the reviewer the wrong thing about what a list - means.** `README.md:90` — *"multi-hub failover. The shim rotates which hub - address each submit targets"* — is false for `submit`, verified above. - `assemble-caution.sh:351` writes *"The driver tries each address until one - acks"* into the manifest the auditor reads, which is also false for `submit` - (no ack is awaited at all). A reader who checks either text before judging a - two-address list concludes each migration reaches **one** hub. - -There is currently no published canonical value to compare against in any case, -which is why the sibling issue asks for one; this issue constrains the form that -publication must take. - -### 6. The attack is cheaper and quieter than the issue text says - -Two strengthenings found during validation, both folded into the text above only -as notes here rather than rewriting the scenario: - -- **The second endpoint need not be a hub.** Submit is dispatch-only and the ack - is never read (`shim/src/nym.rs:578-594`, `:639-676`), so a bare Nym client - that receives the frame and runs `wire::decode` is sufficient. That variant - publishes nothing, so there is **no duplicate broadcast at all** and step 5's - reasoning about `AlreadyKnown` becomes unnecessary rather than merely - satisfied. -- **No telemetry anywhere counts targets.** The mixnet startup arm - (`shim/src/main.rs:228-239`) logs neither the addresses nor their count — the - clearnet arm at `:126-130` logs `hub = %hub_addr`, the mixnet arm logs a bare - sentence. `/nym-status` reports a `diversion_configured` boolean - (`shim/src/nym.rs:207-215`). Nothing in the running system is a function of the - list length. - -### 7. Deflations applied, so the report does not double-count - -- **Against `core-linkage-survives-…` (G29).** For the operator's *own* wallets, - G29 already yields near-certain linkage from the unpadded wallet leg. This - issue's marginal gain there is exactness and the raw bytes rather than a new - class of harm — but it is **not** redundant, because the two are - anti-correlated in time: G29's channel closes under wallet-side padding and - ZIP 318 conformance, which the project is actively asking wallet developers - for, and **this one does not close under anything a wallet can do.** State both, - do not stack them. -- **The set-subtraction half is real but second-order.** The operator learns the - exact bytes, hence the exact txid, of every migration *their own* shim - contributed, and can subtract those members from every published batch. That - harms *other* operators' users. It is an upgrade from G29's statistical - subtraction to a certain one, not a new capability, and should be reported as - the upgrade. -- **Not deniable.** Unlike the on-path and inference attacks elsewhere in this - audit, this one leaves a durable public artefact: `deploy.sh:206-219` pushes the - tree and a `deploy-` tag to a public repository, so the appended address - is recorded permanently. This is the single strongest thing arguing the - severity down, and it is why the recommendation to publish the expected list is - worth so much: it converts a permanent record nobody reads into a permanent - record that fails a check. - -### Why High rather than Medium - -Impact is the maximum the product can suffer: the primary adversary named in the -threat model obtains the complete plaintext of every Orchard-touching transaction -their users send, in real time, with certainty, and `THREATMODEL.md` C2 ("the -party that sees your transaction in the clear and the party that sees your IP -address are two different parties") — the invariant the whole two-component -architecture exists to create — is destroyed for that endpoint. Exploitation -needs one comma in a config file and one VM: no vulnerability, no code change to -either binary, no on-path position, no mixnet position. Nothing observable -changes anywhere — the canonical hub keeps batching and publishing, wallets -confirm on the normal cadence, and no smoke test, health endpoint or telemetry is -a function of the list length. Every check the project documents passes on its -own terms, including `caution verify` with all three PCRs reproducing and the TLS -certificate binding verified, and the one remediation the audit proposes -elsewhere for this trust boundary does not catch it. That combination — -maximal impact, trivial cost, zero functional signature, and a defence that -exists but is aimed one inch to the side — is High. - -It is held below Critical because it requires a deliberately hostile operator -(not an accident or a stranger), it affects one endpoint's users at a time rather -than everyone at once, and it leaves the durable public record described above. - -### Also confirmed, and worth a sentence in the report - -The code comment at `shim/src/nym.rs:618-634` is an accurate and well-reasoned -statement of exactly this hazard — *"safe only while the other addresses are -DEAD"*, *"it doubles the number of enclaves holding the plaintext"*, *"needs an -explicit passive mode in the hub, which does not exist today"*. The design's own -author identified the property and the missing control. What is missing is not -the analysis; it is (a) any enforcement or warning at the configuration layer, -and (b) two public documents that describe the mechanism correctly. That makes -recommendations 1 and 2 the cheap, high-value fixes here. - -DO NOT DEVIATE FROM THIS TEMPLATE (`docs/SECURITY-ISSUE-TEMPLATE.md`). -LEAVE THIS MESSAGE IN PLACE SO THAT ALL AGENTS KNOW NOT TO DEVIATE FROM THIS TEMPLATE (`docs/SECURITY-ISSUE-TEMPLATE.md`). diff --git a/zeronym-22aa9851caf68-high-medium/medium/a-constant-tip-offset-is-a-tunable-expiry-keyed-admission-filter-that-every-proposed-tip-rate-defence-misses.md b/zeronym-22aa9851caf68-high-medium/medium/a-constant-tip-offset-is-a-tunable-expiry-keyed-admission-filter-that-every-proposed-tip-rate-defence-misses.md deleted file mode 100644 index 36968142..00000000 --- a/zeronym-22aa9851caf68-high-medium/medium/a-constant-tip-offset-is-a-tunable-expiry-keyed-admission-filter-that-every-proposed-tip-rate-defence-misses.md +++ /dev/null @@ -1,413 +0,0 @@ -# A constant tip offset turns admission control into an attacker-tunable filter on transaction expiry, silently destroying a chosen fraction of migrations — and no rate-based tip defence can see it - -**Severity**: Medium -**Validation Status**: Confirmed -**Location**: `audit-target/zeronym/hub/src/queue.rs:199-204` (the expiry gate in `Queue::admit`) and `:357-393` (`next_flush_height`, `survives_next_flush`); `audit-target/zeronym/hub/src/batcher.rs:160-197` (`TipTracker::observe`, especially the unconditional first-observation branch at `:163-169`), `:200-232` (`is_stale`, `observed_height`, `cadence_height`), `:40-56` (`FLUSH_INTERVAL_BLOCKS`, `MINING_MARGIN`, `MIN_WALLET_EXPIRY`); `audit-target/zeronym/hub/src/main.rs:60-68` (the startup seed); `audit-target/zeronym/hub/src/server.rs:248-277` (`Hub::admit` and its refusal log); `audit-target/zeronym/hub/src/chain.rs:155-173` (`tip_height`, `max()` over endpoints, `as u32`); `audit-target/zeronym/shim/src/nym.rs:595-690` (dispatch-only submit, so the refusal is never surfaced); `audit-target/zeronym/deploy.env.example:16-17`, `:22-23`. Claims contradicted: `audit-target/zeronym/hub/REVIEW.md:40` (#2), `audit-target/zeronym/hub/src/queue.rs:11-19`. -**Found by agent:** Global (focus area G2 "the flush clock as an attack surface" / G14 "admission→flush→publish as one system"); see `audit-state/globals/G2-G14-G26-flush-clock-pipeline-and-isolation.md` and coordinator item 6u(d) -**In scope of audit?** Yes — priority area 2, "the anonymity mechanism" - -## Description - -`REVIEW.md` #2 replaced the early-expiry flush trigger with admission control, and -`queue.rs:11-19` states the property this is supposed to buy: - -> **Admission control instead of an early-expiry flush (#2).** […] an entry is -> admitted only if it provably survives the next scheduled flush, which makes -> urgency unreachable rather than rate-limited. - -The rule is `expiry >= next_flush_height(tip, N) + mining_margin` -(`queue.rs:379-392`). `tip` is `TipTracker::observed_height()`, which is whatever -the indexer said (`server.rs:256-261`, `batcher.rs:210-215`, `chain.rs:155-173`). - -So the admission rule is a **threshold on the transaction's expiry height, and the -threshold's position is set by a number the adversary writes.** Raising the -reported tip by a constant Δ raises the threshold by Δ, refusing every transaction -whose expiry falls below it. That is not a side effect of an attack on the flush -clock; it is a second, independent lever with a different effect, a different -victim set, and — critically — a different detection signature. - -**Why it is invisible to every defence proposed for the flush-clock lever.** The -already-confirmed `hub-tip-advance-unbounded-flush-clock.md` is about *racing* the -tip, and its remedies are all rate-based: bound the advance by -`last_advance.elapsed() / NOMINAL_BLOCK_SECS`, enforce a minimum wall-clock -interval between flushes, add an absolute plausibility floor. **A constant offset -violates none of them.** - -- The offset can be established at the hub's very first tip query, because the - `!state.observed` branch of `observe` adopts any height unconditionally - (`batcher.rs:163-169`, seeded from `main.rs:60-68`). No jump ever occurs, so - there is nothing for a rate bound to catch. -- Thereafter the reported height advances by exactly one per real block, so - `last_advance` refreshes every block, `is_stale()` never fires, `cadence_height()` - never free-runs, and the flush *rate* is exactly the designed one. -- The flush *phase* shifts by `Δ mod 20` real blocks, which is one of twenty - equally ordinary phases and is not observable as anomalous. -- A "height must be plausible" floor (`> 3_000_000`, the constant - `hub/tests/live_chain.rs:33-37` already uses) passes trivially. - -The only feedback the system produces is a per-refusal -`tracing::info!(reason = "expiry_too_tight", …)` at `server.rs:273`, which reads as -*wallets are submitting transactions with too-tight expiries* — a wallet-side -problem — and which, on the deployed mixnet transport, never reaches the shim or -the wallet at all because submit is dispatch-only and the ack receiver is dropped -at construction (`shim/src/nym.rs:652`; filed as -`nym-submit-acks-are-never-read-so-every-hub-refusal-is-invisible.md`). - -## Attack Scenario and Steps - -Attacker: whoever answers the hub's `GetLightdInfo`. With -`INDEXERS=66.241.124.200:443` (`deploy.env.example:22`) that is one party, and -`tip_height`'s `max()` over a one-element list is the identity function. - -1. From the hub's first tip query onward, the indexer answers `block_height = - real_height + Δ` for a fixed Δ of its choosing. Everything else about its - behaviour is honest. -2. `Hub::admit` computes `survives_next_flush(expiry, real + Δ, 20, 4)`, i.e. - admits iff `expiry >= next_flush_height(real + Δ, 20) + 4`. -3. A transaction built by an ordinary librustzcash-based wallet carries - `expiry = build_height + 40` — the default `batcher.rs:49-55` names as - `MIN_WALLET_EXPIRY` and sizes the whole design against. Writing - `r = (real + Δ) mod 20`, such a transaction built at the current height is - admitted iff `r >= Δ - 16`: - - | Δ | fraction of ZIP 203-default traffic admitted | - |---|---| - | 0 | 20/20 = 100 % | - | 20 | 16/20 = 80 % | - | 26 | 10/20 = 50 % | - | 30 | 6/20 = 30 % | - | **36 or more** | **0/20 = 0 %** | - - (A wallet whose own chain view lags by `d` blocks builds `expiry = real - d + 40` - and is refused at a correspondingly smaller Δ, so these are upper bounds on the - Δ the attacker needs.) -4. A conforming ZIP 318 migration is untouched: its expiry is a bucketed absolute - height 34,561–69,120 blocks ahead of broadcast (`audit-state/SPEC-NOTES.md` §3), - so refusing it needs Δ ≈ 34,500 — at which point *everything* is refused. An - unparseable payload is untouched at any Δ, because `expiry = None` always - survives (`queue.rs:189-197`, `:385-387`). -5. Every refused submission is a migration the wallet was already told had been - sent (`shim/src/hub.rs:231-240`), and the hub keeps no copy of it: the shim is - stateless and the ack is not read, so nothing retries and nothing reports. - -The attacker therefore holds a **continuous dial from "everything through" to -"nothing through"**, moved at zero cost, from a party the threat model already -designates as untrusted, with no jump, no rate anomaly, no staleness, no change in -flush cadence, and a log line that blames the wallets. - -**Attack Requirements and Assumptions:** -- **Control of, or a bug in, a configured `ZIH_INDEXERS` endpoint.** Nothing else: - no submissions, no fees, no mixnet position, no shim, no chain observation. This - is **not** reachable by an anonymous party on the internet — see the severity - bound in the validation section. -- Also reachable **accidentally**: an indexer serving a different network, a - misconfigured height offset, an indexer that is permanently behind the chain at - the moment the hub boots, or the `info.block_height as u32` truncation at - `chain.rs:158` all produce a wrong tip that is adopted unconditionally at - startup and never questioned afterwards. -- **What makes it less severe than it first looks:** the primary harm is - destruction and denial, not deanonymisation. Refused transactions never enter a - batch, so they are not published in a batch of one — they are not published at - all. It is an anonymity attack only in the weaker sense that the surviving - population is filtered along an axis (expiry, hence construction height) the - attacker chooses. - -## Impact on Users - -Per `audit-state/SPEC-NOTES.md` §5, **no shipped wallet has been shown to implement -ZIP 318**, and the shim's interception predicate is `is_orchard_touching`, which -diverts every Orchard-touching transaction regardless of shape. So the diverted -population today is essentially all ZIP 203-default traffic, and Δ = 36 refuses -**all of it**, fleet-wide, silently, for as long as the attacker chooses — while -every wallet is told `error_code 0`, every hub log line attributes the refusals to -the wallets, and the hub's own health and status endpoints stay green. - -A user in that state has an Orchard note they believe they have spent. They cannot -spend it again until the transaction they think is in flight expires — about 50 -minutes at the 40-block default — and their retry then meets the same filter. The -funds are not lost (the note was never spent on-chain, and a syncing wallet -eventually sees non-confirmation), but the migration cannot complete for as long as -the attacker holds the dial, and nothing tells the user, the shim operator or the -hub operator why. Orchard is closed to new value by NU6.3/ZIP 258 and the migration -is the mandatory way out, so a sustained "cannot migrate" is a substantive harm. - -At intermediate Δ the harm is worse in one respect and better in another: some -migrations succeed, so the failure looks like flaky infrastructure rather than an -outage, and nobody investigates. - -## Technical Details / Code Analysis - -The whole of admission control (`hub/src/queue.rs:379-392`): - -```rust -pub fn survives_next_flush( - expiry: Option, - tip: u32, - flush_interval: u32, - mining_margin: u32, -) -> bool { - match expiry { - None => true, - Some(expiry) => { - let deadline = next_flush_height(tip, flush_interval).saturating_add(mining_margin); - expiry >= deadline - } - } -} -``` - -with - -```rust -// hub/src/queue.rs:362-368 -pub fn next_flush_height(h: u32, n: u32) -> u32 { - if n == 0 { - return h; - } - ((h / n).saturating_add(1)).saturating_mul(n) -} -``` - -`tip` reaches it from the tracker, unfiltered (`hub/src/server.rs:252-261`): - -```rust - if self.tip.is_stale() { - return Err(Refusal::TipStale); - } - - match self.queue.admit( - tx_bytes, - self.tip.observed_height(), - self.params.flush_interval, - self.params.mining_margin, - ) { -``` - -and `observed_height` is the raw last-observed value (`hub/src/batcher.rs:210-215`), -whose own doc comment states the property this attack breaks: *"This is what -admission checks expiry against, because admission must never be more optimistic -than the chain actually is."* - -The branch that lets the offset exist from the first query, with no comparison -against anything (`hub/src/batcher.rs:161-175`): - -```rust - pub fn observe(&self, height: u32) { - let mut state = self.write(); - - if !state.observed { - state.height = height; - state.last_advance = Instant::now(); - state.observed = true; - return; - } - - if height > state.height { - state.height = height; - state.last_advance = Instant::now(); - return; - } -``` - -Note that the *backwards* direction is bounded by `REORG_ALLOWANCE` and logged at -`warn!` (`:177-196`), while the forward direction and the initial adoption are -neither bounded nor logged. `observed_height` is never emitted in any log line -anywhere in the crate, so an operator cannot compare the hub's belief about the -chain with the chain. - -The refusal's only trace (`hub/src/server.rs:272-275`): - -```rust - Admission::Refused(refusal) => { - tracing::info!(reason = refusal.as_str(), "submission refused at admission"); - Err(refusal) - } -``` - -`Refusal::ExpiryTooTight` renders as the fixed string `"expiry_too_tight"` -(`queue.rs:90-92`), which carries no indication that the tip it was measured -against might be wrong. - -Finally, the design statement this defeats (`hub/REVIEW.md:40`, #2): - -> Admission control makes the trigger unreachable rather than rate-limited: if -> every admitted entry provably survives the next scheduled flush, no entry can -> ever be urgent. - -The conclusion holds — no admitted entry becomes urgent. What the argument does not -establish, and what this issue is about, is that the *admission predicate itself* -is now an adversary-positioned gate, so the attacker's lever moved from "make an -entry urgent" to "decide which entries exist". - -### The offset is signed, and the negative direction inverts the failure (found during validation) - -`observe`'s unconditional first-observation branch adopts a tip that is too **low** -just as readily as one that is too high, and a permanently-lagging indexer — or an -attacker choosing Δ < 0 — then holds that offset for the process's life, because -every subsequent report still advances at the chain rate. - -With reported tip `T = real - L`, the deadline is computed `L` blocks too early -while the flush itself still happens within 20 real blocks of admission (the -cadence runs on the same shifted clock, so only its phase moves). The check -therefore **passes entries it should have refused**: - -- honest tip: a transaction with `expiry = real + 5` is measured against - `next_flush_height(real, 20) + 4` and is **refused** — correctly, and the wallet - can immediately rebuild with a fresh expiry; -- lagging tip with `L = 1000`: the same transaction is measured against - `next_flush_height(real - 1000, 20) + 4`, a number ~1000 blocks in its past, is - **admitted**, is held until the next cadence boundary, is published after it has - expired, and is refused by the node — which `flush` classes as a verdict and - **drops permanently** (`batcher.rs:366-368`, `chain.rs:459-474`, - `chain.rs:513-533`). - -So in this direction admission control fails **open** rather than closed, and -REVIEW #2's "every admitted entry provably survives the next scheduled flush" is -false for the entries it matters most for. The affected population is narrower than -the positive-Δ case — it is submissions whose expiry is already tight (a wallet -whose chain view is >16 blocks stale, a resend near expiry, or a wallet with a -shorter expiry default such as the 20 blocks `batcher.rs:51-53` explicitly puts out -of scope) — which is exactly the population admission control exists to protect -from silent destruction. It is recorded here rather than filed separately because -the root cause, the code site and the fix are identical. - -## Recommendations - -- **Bound the tip against something the reporting party does not control.** A - wall-clock rate bound (the fix proposed for `hub-tip-advance-unbounded-flush-clock.md`) - is necessary but not sufficient, because it cannot see a constant offset. The - offset is only detectable by comparison: against a second, independent endpoint - (see `tip-and-verdict-aggregation-scale-in-opposite-directions-so-adding-indexers-fixes-one-lever-and-aggravates-three.md` - for why the current `max()` fold makes that comparison useless), or against the - hub's own wall clock anchored at a startup height an operator supplies out of - band. A sanity floor and a sanity *ceiling* on the very first observation — - `main.rs:60-68` is the one place a human-supplied expectation could be checked — - would close the accidental half at near-zero cost. -- **Log the observed height and the admission threshold.** Both are aggregates - carrying no per-entry information, so the counts-only rule (#157) permits them, - and either one would have made this visible. Today neither is emitted anywhere in - the crate. This is the cheapest instrumentation fix in the hub and it serves both - tip findings at once. -- **Alarm on the refusal *rate*, not just the refusal.** A sustained non-zero - `expiry_too_tight` rate has exactly two explanations — wallets with genuinely - tight expiries, or a tip that is ahead of the chain — and the hub is currently - unable to distinguish them or to report either. -- **Surface the refusal to the wallet.** On the deployed transport the wallet is - told success and the refusal is discarded, which is what turns a refusal into a - destruction. This is the same root cause as - `nym-submit-acks-are-never-read-so-every-hub-refusal-is-invisible.md` and fixing - it there fixes the silent half of this issue. -- **Correct `queue.rs:11-19` and `REVIEW.md` #2** to state that admission control - makes urgency unreachable *given a trustworthy tip*, and that the tip is supplied - by a party the same document treats as adversarial elsewhere. - -## Validation Information - -**Verdict: CONFIRMED. Severity: Medium (as filed), for consistency with the bound -already applied to the sibling tip issue.** - -### The arithmetic was re-derived independently and it is correct - -With `tip = real + Δ`, `flush_interval = 20`, `mining_margin = 4`, and a wallet -transaction built at the true tip with `expiry = real + 40`: - -`next_flush_height(tip, 20) = 20·(⌊tip/20⌋+1)`. Writing `tip = 20q + r` with -`0 ≤ r < 20`, admission holds iff `real + 40 ≥ 20q + 24`, and `20q = tip − r = -real + Δ − r`, giving **`r ≥ Δ − 16`**. The fraction of block phases satisfying -that is `(20 − max(0, Δ−16))/20`, which reproduces every row of the table: -Δ=0→100 %, Δ=20→80 %, Δ=26→50 %, Δ=30→30 %, Δ≥36→0 %. Confirmed. - -### Every mechanical claim re-verified against the target at HEAD - -| Claim | Verified at | -|---|---| -| `!state.observed` adopts any height with no bound, no floor, no log | `hub/src/batcher.rs:163-169` | -| Forward advance is likewise unbounded; only the backward direction is guarded and logged | `hub/src/batcher.rs:171-196` (`REORG_ALLOWANCE = 10` at `:59`) | -| The startup seed feeds that branch directly from one indexer call | `hub/src/main.rs:60-68` | -| `observed_height()` is the raw value and is what admission uses | `hub/src/batcher.rs:210-215`, `hub/src/server.rs:252-261` | -| `survives_next_flush` is the whole of admission control, and `expiry = None` always passes | `hub/src/queue.rs:379-392`, applied at `:199-204` | -| `next_flush_height` is the next strict multiple of the interval | `hub/src/queue.rs:362-368` | -| Constants: interval 20, margin 4, `MIN_WALLET_EXPIRY = 40` (librustzcash's default) | `hub/src/batcher.rs:40-56` | -| Cadence runs on `cadence_height()/interval`, so a constant offset shifts phase only | `hub/src/batcher.rs:300-311`, `:219-232` | -| `is_stale()` never fires while the reported tip advances | `hub/src/batcher.rs:205-211` (`TIP_STALE_AFTER = 15 min` at `:63`) | -| `tip_height` is `max()` over endpoints, and the shipped list has one entry | `hub/src/chain.rs:155-173`; `deploy.env.example:22` | -| `u64 → u32` truncation is unchecked | `hub/src/chain.rs:158` (`info.block_height as u32`) | -| Refusal renders as a fixed string that blames the wallet | `hub/src/queue.rs:90-92`, logged at `hub/src/server.rs:273` | -| The ack is never read, so no refusal reaches the wallet | `shim/src/nym.rs:652` — `let (ack_tx, _drop_receiver) = oneshot::channel();` | -| `HubTransport::submit` has no `Refused` arm on the Nym path | `shim/src/hub.rs:230-249` | -| Nothing logs the observed height anywhere in the crate | grep: `observed_height` appears only at its definition and two call sites, never in a `tracing!` | -| `BACKEND` and `INDEXERS` are the same host/port in the shipped example | `deploy.env.example:16-17` vs `:22-23` — so the composition needs no collusion | - -### The severity bound (coordinator item 6p), applied consistently - -The G7 pass's bound is explicit and governs here: **all tip manipulation requires -control of a configured `ZIH_INDEXERS` endpoint; it is a hub-trust / robustness -defect, not an internet-reachable weapon.** The prior validator applied exactly -that bound to the sibling `hub-tip-advance-unbounded-flush-clock.md`, correcting it -High → Medium. The same bound applies unchanged here — nothing in this attack is -reachable by an anonymous party — so **Medium**, not High. - -What holds it *at* Medium rather than lower: - -1. **Against the party who can reach it, the attack is certain, free and - invisible.** One integer per 30 s poll. At the shipped `n = 1` there is no - aggregation to defeat, no rate check, no plausibility check, and no log line to - alarm on. -2. **That party is adversary #1 by name.** `AUDIT-INSTRUCTIONS.md`'s trust - boundaries state *hub → indexer … can lie about the tip and about publish - verdicts*, and `deploy.env.example` points `BACKEND` and `INDEXERS` at the same - host, so in the shipped example the hub's admission gate is held by a party who - is simultaneously a shim's backing indexer. -3. **It is reachable with no attacker at all**, by a lagging, misconfigured or - wrong-network indexer — including the negative-offset direction recorded above. -4. **The harm is fleet-wide and silent**: at Δ ≥ 36 every migration in the system - is destroyed while every wallet is told `error_code 0` and every health surface - stays green. - -### Anti-double-counting - -- **Against `hub-tip-advance-unbounded-flush-clock.md` (confirmed Medium).** That - issue owns the *race*: an advancing tip that fires flushes early and collapses - the batch to size 0 or 1 — an anonymity harm. Its validation section explicitly - delegates the constant-offset form to this file (*"The constant-offset form of - this … is filed separately as …; this issue's race form reaches the same state as - a side effect"*), so the allocation is already agreed and is preserved here. The - two are not the same finding: the remedies proposed for the race (rate bound, - minimum flush interval, plausibility floor) **all pass** a constant offset, which - is this issue's entire point, and the harm here is destruction rather than batch - collapse. -- **Against `indexer-chooses-which-batch-members-reach-the-chain-…` (confirmed - High).** That is the *publish* gate; this is the *admission* gate. Different - code, different verdict path, different fix. -- **The isolation/batch-of-one framing is deliberately NOT claimed here.** A - temporal admission window (open the gate for one epoch, close it for the rest) - would isolate a target, but it requires either a jump or a stall, both of which - belong to the sibling race issue and to `is_stale`'s existing fail-closed - behaviour. Per coordinator item 6u(b) the isolation harm must not be stacked; - this issue claims destruction and denial only. -- **The silent-loss half** is owned by - `nym-submit-acks-are-never-read-so-every-hub-refusal-is-invisible.md`; it is - cited here as the reason a refusal becomes a destruction, not re-counted. - -### Corrections applied against the filed text - -- *"The transaction is gone"* and *"no copy of it exists anywhere"* were softened. - The user's note is not spent on-chain, and per coordinator item 7n a syncing - wallet does eventually observe non-confirmation and can rebuild after expiry - (~50 minutes at the 40-block default). What is genuinely destroyed is the - submission, silently, with a false success already delivered — and the retry - meets the same filter. The Impact section now says this precisely. -- The line references to `TipTracker::observe`, `survives_next_flush`, - `Refusal::as_str` and the startup seed were re-derived and corrected - (`batcher.rs:163-169`, `queue.rs:379-392`, `queue.rs:90-92`, `main.rs:60-68`). -- A **second direction** was added: the offset is signed, and a negative offset - makes the same predicate fail *open*, admitting entries that are then published - after expiry and dropped as a verdict. This was verified by working the cadence - arithmetic through in both units; an earlier form of this claim (that *any* - lagging tip delays publication past expiry) is **false** and is not stated — - publication always occurs within 20 real blocks of admission because the cadence - runs on the same shifted clock. The correct statement is the narrower one now in - the text. - -DO NOT DEVIATE FROM THIS TEMPLATE (`docs/SECURITY-ISSUE-TEMPLATE.md`). -LEAVE THIS MESSAGE IN PLACE SO THAT ALL AGENTS KNOW NOT TO DEVIATE FROM THIS TEMPLATE (`docs/SECURITY-ISSUE-TEMPLATE.md`). diff --git a/zeronym-22aa9851caf68-high-medium/medium/attested-enclave-console-is-reopenable-from-the-parent-because-debug-mode-is-a-launch-flag-and-ssh-keys-is-not-gated-on-it.md b/zeronym-22aa9851caf68-high-medium/medium/attested-enclave-console-is-reopenable-from-the-parent-because-debug-mode-is-a-launch-flag-and-ssh-keys-is-not-gated-on-it.md deleted file mode 100644 index 7135855e..00000000 --- a/zeronym-22aa9851caf68-high-medium/medium/attested-enclave-console-is-reopenable-from-the-parent-because-debug-mode-is-a-launch-flag-and-ssh-keys-is-not-gated-on-it.md +++ /dev/null @@ -1,471 +0,0 @@ -# Attestation does not close the enclave console — a parent-side launch flag does, and `debug.ssh_keys` (ungated on `debug.enabled`) hands that flag to the operator, so an "attested" shim's per-migration INFO log can be turned back on at will - -**Severity**: Medium -**Validation Status**: Confirmed -**Location**: `audit-target/zeronym/shim/deploy/caution/caution.hcl.tmpl:186-201` and `audit-target/zeronym/hub/deploy/caution/caution.hcl.tmpl:143-157` (the `debug` block and the claim "SSH is closed under attestation", shim `:198` / hub `:154`); `audit-target/zeronym/shim/deploy/caution/OPERATORS.md:354-357` ("Reading state from an attested enclave: **there is no SSH**"); `audit-target/zeronym/hub/deploy/caution/OPERATORS.md:138` ("an attested enclave has **no SSH**"); `audit-target/zeronym/deploy.sh:227-229` (the same coupling used as the design rationale for the unauthenticated `/nym-address` endpoint); the renderer that produces the combination — `audit-target/zeronym/shim/deploy/caution/assemble-caution.sh:120-123` (the NOTE) and `:444-462` (the `ssh_keys` block, branched on `$SSH_KEYS` and never on `$DEBUG`), identical in `audit-target/zeronym/hub/deploy/caution/assemble-caution.sh:141-144` and `:359-376`; the log stream this reopens is `audit-target/zeronym/shim/src/intercept.rs:577-637` (`log_verdict`, `Class::Migration` arm at `:582-601`) and `audit-target/zeronym/hub/src/server.rs:265-275` (the per-admission line at `:269`), both at INFO under `audit-target/zeronym/shim/src/main.rs:21-30` / `audit-target/zeronym/hub/src/main.rs:22-26`. Platform mechanism (outside `audit-target/`, re-read during validation from the public Caution clone `codeberg.org/caution/platform` @ `1f8d8cb`): `terraform/modules/aws/nitro-enclave/user-data.sh:15` (nitro-cli installed on every parent), `:17-27` (keys written to `ec2-user`, guarded by `length(ssh_keys) > 0` only), `:53` (`aws s3 cp … /opt/nitro/enclave.eif`), `:127-154` + `:174-190` (console-capture service, `debug_mode`-gated), `:157-173` (`nitro-enclave.service`, `--debug-mode` appended iff `debug_mode`, at `:166`), `:194-212` (the `socat` 443 vsock relay, `ExecStart` at `:210`), `src/api/src/main.rs:2420-2421` (`debug_enabled` and `ssh_keys` read independently), `src/api/src/deployment.rs:2158-2164` (`debug_mode` and `ssh_ingress` are independent template fields; `ssh_ingress` emits `ingress 22 … 0.0.0.0/0`), `:2018-2032` (that ingress lands on `aws_security_group.enclave`, attached to the instance at `:2075`, which carries a public EIP at `:2118-2126`), `:1871-1879` (AL2023 AMI), `src/enclave-builder/src/manifest.rs:14-45` (`EnclaveManifest` — no `debug` field, so the key list is in no attestation) -**Found by agent:** Global, focus area G18 (log and telemetry discipline as one policy) — re-deriving, as coordinator open item 6x requires, who can actually read these logs in each deployment shape -**In scope of audit?** Yes — priority area 6 ("Log and telemetry discipline… Every log line… is an egress channel to the adversary") and priority area 7 (the attestation chain); the two `caution.hcl.tmpl` files, both `OPERATORS.md` and `deploy.sh` are in scope as security claims - -## Description - -The audit's bounding of every log-discipline finding rests on one premise, settled -as coordinator open item 7: with `debug { enabled = false }` the parent host has -**no console channel**, so `log_verdict`'s per-migration value balances are "not -live in a correctly attested deployment". **That premise is true as stated, and -this issue does not contradict it.** What it corrects is the inference drawn from -it — that *attestation* is what closes the channel. - -It is not. Attestation and the console are two independent switches on the -**parent host**, and only one of them is cryptographic: - -1. `debug.enabled` decides whether `nitro-cli run-enclave` is invoked with - `--debug-mode`. That is a **string in a systemd unit file on the parent's - local disk**, rendered once at instance boot - (`user-data.sh:166`, `%{if debug_mode == "true"}--debug-mode%{endif}`). -2. `debug.ssh_keys` decides whether port 22 is open on the parent and the - operator's key is installed on `ec2-user`. It is read **independently**: - `src/api/src/main.rs:2420-2421` reads `enabled` and `ssh_keys` as two separate - fields, `deployment.rs:2158-2164` renders `debug_mode` and `ssh_ingress` as two - separate fields of the same terraform template, and `user-data.sh:17-27` is - guarded by `%{ if length(ssh_keys) > 0 ~}`, never by `debug_mode`. - -So `debug { enabled = false; ssh_keys = [ "ssh-ed25519 …" ] }` is a **fully -attested** enclave — `caution verify` passes, PCRs reproduce — whose parent host -the operator can log into. Coordinator item 6x established that much and graded it -on packet capture, enclave restarts and vsock access; it also recorded, following -the filed `hub-manifest-debug-block-…` issue, that **"Note what is not at risk: -enclave console output."** That is the part that is wrong, and it is wrong in the -direction that matters: the console is not protected by attestation, it is -protected by a flag on a machine the operator is now logged into. - -**The bound this issue must preserve, because several confirmed issues depend on -it.** `nitro-cli console` **refuses** on an enclave that was launched without -`--debug-mode` (AWS: *"can be used only on an enclave that was launched with the -`--debug-mode` option"*), and the console-capture service that would write -`/var/log/nitro_enclaves/enclave-console.log` is template-gated on `debug_mode` -and pinned by the platform's own unit test -`test_enclave_console_capture_is_debug_only` (`src/api/src/deployment.rs:554-571`). -So the *deployed* attested-with-`ssh_keys` state does **not** leak logs, and **no -log finding may be escalated by this route directly.** The escalation costs one -extra, deliberate act — terminate the enclave and relaunch the same signed EIF -with `--debug-mode` — during which the enclave's attestation document carries -zeroed PCRs. - -Everything that act needs is already on the parent, installed unconditionally by -the platform's own bootstrap: - -- `aws-nitro-enclaves-cli` is installed on every parent, debug or not - (`user-data.sh:15`). -- The signed enclave image is at `/opt/nitro/enclave.eif` (`user-data.sh:53`). -- The launch command is a plain systemd unit at - `/etc/systemd/system/nitro-enclave.service` (`user-data.sh:157-173`). -- The AMI is Amazon Linux 2023 (`deployment.rs:1871-1879`) and the keys land on - `ec2-user`, that AMI's default administrative account with passwordless `sudo`. - The 535-line `user-data.sh` contains no `sudo`, `sshd`, `usermod` or - `PermitRootLogin` statement of any kind, so the AMI defaults stand unmodified. - -**The binary and its environment are unchanged**, so what appears on the reopened -console is exactly the shipped INFO stream — on the shim `log_verdict`'s -`orchard_vb` / `ironwood_vb` / `sapling_vb` / `expiry` / `inputs` / `outputs` / -`tx_len`, one line per diverted migration (`intercept.rs:582-601`), and on the hub -the per-admission line (`server.rs:269`). - -**Five zeronym texts assert the opposite**, and they are the texts an operator or -a reviewer would consult: - -- `shim/deploy/caution/caution.hcl.tmpl:198` and the identical - `hub/deploy/caution/caution.hcl.tmpl:154`: *"without `--debug` this list renders - empty and is moot (**SSH is closed under attestation**)."* -- `shim/deploy/caution/OPERATORS.md:354`: *"Reading state from an attested - enclave: **there is no SSH.**"* -- `hub/deploy/caution/OPERATORS.md:138`: *"That endpoint exists because an - attested enclave has **no SSH**…"* -- `shim/deploy/caution/caution.hcl.tmpl:189-193`, which presents debug mode as the - thing that "opens port 22 on the parent so the console can be read", conflating - the two independent switches into one. -- `deploy.sh:227-229`, where the coupling is used as a **design rationale**: *"the - enclave console is only open with --debug, and --debug disables attestation, so - an address that could only be read from the console could never belong to a hub - that had also been proven."* - -## Attack Scenario and Steps - -The actor is adversary #1, the indexer operator the shim fronts. They are not -required to breach anything; they configure their own deployment. - -1. The operator assembles an **attested** shim, passing `--ssh-key` without - `--debug`. `shim/deploy/caution/assemble-caution.sh:449-462` renders the key - into the `debug` block next to `enabled = false` — the branch is on - `$SSH_KEYS` and never on `$DEBUG` — and `:120-123` prints a note reassuring - them that *"SSH is closed when attestation is on, so the key is recorded in the - HCL but unused until a --debug build"*. Nothing in the tooling refuses this - combination; the filed `hub-manifest-debug-block-…` issue owns this half. -2. They deploy. `caution verify` passes. `/attestation` serves a COSE-signed - document with reproducing PCRs and a `.manifest.run_command` showing the - expected environment. The `debug` block is **not** part of the served manifest - (`src/enclave-builder/src/manifest.rs:14-45` has no such field), so no - attestation consumer sees the key list at all. Wallets connect; real migrations - flow. -3. At a moment of their choosing — for one batch, one hour, or one target's - session — the operator SSHes to the parent's public address and runs: - - ``` - sudo nitro-cli describe-enclaves # note EnclaveID, memory, cpus - sudo nitro-cli terminate-enclave --all - sudo nitro-cli run-enclave --eif-path /opt/nitro/enclave.eif \ - --memory --cpu-count --enclave-cid 16 --debug-mode - sudo nitro-cli console --enclave-id - ``` - - (Editing `--debug-mode` into `/etc/systemd/system/nitro-enclave.service` and - restarting the unit is equivalent.) The parent's `socat` relays - (`user-data.sh:194-212`, `ExecStart=/usr/bin/socat TCP-LISTEN:$port,reuseaddr,fork VSOCK-CONNECT:16:$port` - at `:210`) fork and re-dial CID 16 per connection, so the shim keeps serving on - the same hostname. The in-enclave Caddy re-runs ACME and republishes its - `certfp`. -4. Every subsequent `SendTransaction` prints one INFO line naming that - transaction's Orchard, Ironwood and Sapling value balances, its expiry, its - transparent input/output counts and its exact length — beside the TCP - connection the operator's own parent is relaying. This is the confirmed High - `log-verdict-logs-migration-value-balance-at-info.md`, delivered in a - deployment whose manifest says `enabled = false`. -5. They terminate and relaunch without `--debug-mode` to restore reproducing - PCRs. - -**Attack Requirements and Assumptions:** - -- **Access needed:** the operator must have put an SSH key in their own manifest. - That is a supported, documented option of the project's own assembler, and the - manifest comment tells them it is inert. `deploy.sh` does not produce the - combination on its own (`deploy.sh:128-135` passes `--ssh-key` only together - with `--debug`), so the reachable route is running `assemble-caution.sh` - directly — which both `deploy/README.md` files document as the primary - interface — or hand-editing the assembled `caution.hcl`, which nothing - re-validates (`shim-assemble-never-verifies-the-manifest-it-rendered.md`). -- **Verified during validation, not assumed:** the security group rule the key - list produces is `from_port = 22, to_port = 22, cidr_blocks = ["0.0.0.0/0"]` - (`deployment.rs:2161-2164`), attached to the enclave instance's own security - group (`:2018-2032`), and the instance carries an Elastic IP - (`:2118-2126`). So the parent's sshd is exposed to the **whole internet**, not - to the operator's address — the attacker set is "whoever holds that private - key", which includes anyone who steals or compels it, and the exposure is a - standing one for the life of the deployment. -- **Stated assumption:** that `ec2-user` on the AL2023 parent reaches root. This - is the AWS default for that AMI and the platform's `user-data.sh` does not - disable it (checked: the file contains no sudo/sshd hardening). The same - assumption is already load-bearing for the other filed parent-host capabilities - (tcpdump on the 443 relay, enclave restart). -- **What it costs the operator:** during the debug window `/attestation` returns - zeroed PCRs, so anyone who fetches and checks it *in that window* sees the - enclave is unattested; and each relaunch spends one of the domain's five weekly - Let's Encrypt issuances and shows a new certificate in CT - (`acme-nocache-issuance-budget-…`). Neither is observed by anything zeronym - ships: the confirmed `attested-tls-binding-is-verified-once-by-hand-if-ever-…` - establishes that no check is scheduled anywhere, and the runbook's own advice is - to watch CT for *certificates you cannot account for* — an operator restarting - their own enclave accounts for it. -- **Why this is not "requires the system to already be compromised":** nothing is - compromised. The operator exercises a configuration option on infrastructure - they are entitled to configure, and the security claim being broken is the one - made to *users and auditors* — that an attested enclave's logs cannot reach the - operator. In the deployment `deploy.sh` actually performs (fully managed, in - **Caution's** AWS account — `shim/deploy/caution/OPERATORS.md:64`, coordinator - item 6x) the operator does **not** hold the parent host by default, so - `debug.ssh_keys` is how they *obtain* the position, not a restatement of one - they already have. - -## Impact on Users - -- It removes the mitigation that bounds the audit's log-discipline findings. The - correct statement is no longer "the amount leak is live only in the `DEBUG` - deployment"; it is "the amount leak is live in the `DEBUG` deployment, and - **re-armable on demand** in an attested deployment that carries an SSH key". -- For a user whose migration is logged during such a window, the harm is the one - the confirmed High already describes: the operator holds their TCP source - address at the parent's relay socket and, on the same host, a line giving the - exact zatoshi value balances, expiry and length of the transaction that wallet - just diverted — which `README.md:33` says the operator does not learn. -- The window is chosen by the adversary and is invisible to the user, because a - wallet has no way to see PCRs and the shim's TLS identity is unchanged from its - point of view. -- Second-order, and worth stating because it is a standing risk rather than an - operator choice: the manifest option opens **22/tcp to `0.0.0.0/0`** on a - machine in Caution's account for the lifetime of the deployment. - -## Technical Details / Code Analysis - -The two switches, in the platform's own template. From -`terraform/modules/aws/nitro-enclave/user-data.sh` (Caution SEZC, AGPL-3.0; -quoted for analysis): - -```bash -# :15 — unconditional, on every parent -dnf install -y aws-nitro-enclaves-cli aws-nitro-enclaves-cli-devel docker socat dnsmasq iptables iproute - -# :17-27 — guarded by ssh_keys ONLY -%{ if length(ssh_keys) > 0 ~} -mkdir -p /home/ec2-user/.ssh -%{ for key in ssh_keys ~} -echo "${key}" >> /home/ec2-user/.ssh/authorized_keys -%{ endfor ~} -%{ endif ~} - -# :53 — unconditional -aws s3 cp "${eif_s3_path}" /opt/nitro/enclave.eif - -# :166 — the entire console gate, in a unit file on the parent's disk -ExecStart=/bin/bash -c 'nitro-cli run-enclave --eif-path /opt/nitro/enclave.eif \ - --memory ${memory_mb} --cpu-count ${cpu_count} --enclave-cid 16 \ - %{if debug_mode == "true"}--debug-mode%{endif} && tail -f /dev/null' -``` - -The console-capture service that writes -`/var/log/nitro_enclaves/enclave-console.log` is `debug_mode`-gated (`:127-154`, -`:174-190`) and the platform even has a unit test pinning that — -`src/api/src/deployment.rs:554-571`, `test_enclave_console_capture_is_debug_only`, -which asserts every reference to `capture-enclave-console.sh`, -`nitro-enclave-console.service`, `/var/log/nitro_enclaves/enclave-console.log` and -`nitro-cli console` sits inside a `%{ if debug_mode == "true" ~}` block. **That -test is exactly the evidence for this finding**: the platform guarantees the -*file* is not written without debug mode; it guarantees nothing about a shell on -the parent invoking `nitro-cli console` itself, because the gate it enforces is a -template gate, not a privilege boundary. - -The independence of the two fields, from `src/api/src/main.rs:2420-2421`: - -```rust - let debug_enabled = ec_debug.and_then(|d| d.enabled).unwrap_or(false); - let ssh_keys = ec_debug.map(|d| d.ssh_keys.clone()).unwrap_or_default(); -``` - -and from `src/api/src/deployment.rs:2158-2164`: - -```rust - debug_mode = if request.debug_mode { "true" } else { "false" }, - ssh_keys_json = - serde_json::to_string(&request.ssh_keys).unwrap_or_else(|_| "[]".to_string()), - ssh_ingress = if request.ssh_keys.is_empty() { - "# SSH ingress disabled (no ssh_keys in Procfile)".to_string() - } else { - "…ingress {\n from_port = 22\n to_port = 22\n protocol = \"tcp\"\n cidr_blocks = [\"0.0.0.0/0\"]…" - }, -``` - -Nothing anywhere in the platform validates the pair: a repository-wide search for -`ssh_keys` finds parsing (`src/caution-config/src/lib.rs:504-576`, where -`has_debug = debug.is_some() || !ssh_keys.is_empty()`), rendering, and tests — -and no rule relating it to `enabled`. - -What the reopened console carries. `shim/src/main.rs:21-30` installs a -process-wide subscriber defaulting to `info`, with a comment that is precise about -the policy: - -```rust - // `info` deliberately does NOT include the per-request `zis::proxy` line: - // that line names the method each wallet called, which is a metadata source - // this component exists to deny the operator, and it would live in a log - // file on the operator's box. -``` - -and `shim/src/intercept.rs:582-601` puts, at that same INFO level, the fields the -same reasoning would forbid: - -```rust - Class::Migration => tracing::info!( - target: "zis::classify", - version = %evidence.version, - orchard_actions = evidence.orchard_actions, - orchard_vb = %format!("{:+}", evidence.orchard_vb), - ironwood_vb = %format!("{:+}", evidence.ironwood_vb), - sapling_vb = %format!("{:+}", evidence.sapling_vb), - expiry = ?evidence.expiry_height, - inputs = evidence.inputs, - outputs = evidence.outputs, - tx_len = evidence.len, - diverted_in_production, - "MIGRATION detected: …" - ), -``` - -**What attestation still buys, stated precisely.** `RUST_LOG` is not free for the -operator to raise at this point: it is an `export` line inside the measured -`run.sh` (`src/caution-config/src/lib.rs:253-276` emits `export KEY=value` for -every literal `env` entry; item 6q), so the reopened console shows the INFO stream -and not the `zis::proxy` debug stream. That constraint is real, but it is enforced -by **PCR0/PCR1 and by `.manifest.run_command`** — and `deploy.sh:220` instructs -verifiers to *expect PCR0/1 to fail and accept PCR2 alone* (confirmed -`deploy-script-tells-operators-to-expect-pcr01-failure-and-accept-pcr2-alone.md`), -while no zeronym document tells anyone to read `.manifest.run_command` at all -(confirmed `auditor-recipe-omits-…`). So the honest form of the slogan is: -**attestation constrains what is logged; it never constrains who reads it; and the -first half is enforced only by checks the project's own instructions skip.** - -The zeronym texts that are falsified, verbatim -(`shim/deploy/caution/caution.hcl.tmpl:186-201`): - -```hcl - debug { - # … - # If the enclave boots but never serves, flip this to true and redeploy; - # that opens port 22 on the parent so the console can be read at - # /var/log/nitro_enclaves/enclave-console.log. … - # With --debug the flip is one boolean and the key is already listed; without - # --debug this list renders empty and is moot (SSH is closed under attestation). - enabled = false - __DEBUG_SSH_KEYS__ - } -``` - -`shim/deploy/caution/OPERATORS.md:354-357`: - -> **Reading state from an attested enclave**: there is no SSH. Use -> `https:///attestation` and, on the hub, `/healthz` and -> `/nym-address`. - -and `deploy.sh:227-229`, which uses the same belief as a design rationale: - -> the enclave console is only open with --debug, and --debug disables -> attestation, so an address that could only be read from the console could never -> belong to a hub that had also been proven. - -## Recommendations - -1. **Refuse the combination in the assembler.** Both `assemble-caution.sh` scripts - should `exit 2` on `--ssh-key` without `--debug`, exactly as they already do - for other unsafe input combinations (shim `:185-192`, `:221-225`). This is a - three-line change and it closes the whole issue for the deploy path. -2. **Correct the five texts.** Delete "SSH is closed under attestation" and "there - is no SSH" from both `caution.hcl.tmpl`s, both `OPERATORS.md`s - (shim `:354`, hub `:138`) and the rationale at `deploy.sh:227-229`, and replace - them with the true statement: a non-empty `ssh_keys` list opens port 22 on the - parent to `0.0.0.0/0` regardless of `debug.enabled`, and a shell on the parent - can relaunch the enclave in debug mode and read the console. -3. **Give auditors a check that can see it.** The `debug` block is not in the - attested manifest, so the only external evidence is the open port. Publish the - deployment's public IP alongside the attestation URL and state that `22/tcp` - open is a finding; the platform's own `dns_contains_deployment_ip` check on the - raw-IP verify path already establishes the pattern. -4. **Reduce what the console is worth.** Remove `orchard_vb` / `ironwood_vb` / - `sapling_vb` / `expiry` / `inputs` / `outputs` / `tx_len` from the - `Class::Migration` and fail-safe arms of `log_verdict` (keeping the - counts-and-verdict form the hub already uses), so that a reopened console - yields a count rather than an amount. This is the recommendation of the - confirmed `log-verdict-logs-migration-value-balance-at-info.md`; this issue is - the reason it should not be deferred on the grounds that attestation contains - it. -5. **Restrict the SSH ingress upstream.** Independently of zeronym, the platform's - `0.0.0.0/0` SSH rule should be narrowed to an operator-supplied CIDR; worth - raising with Caution and recording in `OPEN-QUESTIONS.md`. - -## Validation Information - -**Verdict: CONFIRMED at Medium** (severity as filed). Every mechanical claim was -re-derived during validation from the Caution platform clone -(`codeberg.org/caution/platform` @ `1f8d8cb`) and from the target, not inherited -from earlier passes. - -**Re-verified from platform source:** - -- `src/api/src/main.rs:2420-2421` — `debug_enabled` and `ssh_keys` are two - independent reads of the same `debug` block; `debug_enabled` is not consulted - when building `ssh_keys`. -- `src/api/src/deployment.rs:2158-2164` — `ssh_ingress` is predicated on - `request.ssh_keys.is_empty()` alone and emits `from_port = 22 … cidr_blocks = - ["0.0.0.0/0"]`; `debug_mode` is a separate template variable. This is inside - `generate_nitro_deployment_main_tf`, i.e. the AWS Nitro path the fully-managed - deploy uses (the on-prem generator at `:2458-2464` is identical). -- The ingress lands on `aws_security_group.enclave` (`:2018-2032`) which is - attached to `aws_instance.enclave` (`:2075`), and the instance has an - `aws_eip.enclave` (`:2118-2126`) — so the port is reachable from the internet. -- `user-data.sh:15` (nitro-cli unconditional), `:17-27` (keys under - `length(ssh_keys) > 0` only), `:53` (EIF path), `:157-173` (the unit), - `:166` (`--debug-mode` iff `debug_mode`), `:127-154` + `:174-190` (capture - service, debug-gated), `:210` (the 443 `socat` relay, `fork`, per-connection). - **Three citations in the filing were off and are corrected here:** the EIF copy - is `:53` not `:52`, the unit is `:157-173` not `:155-172`, and the - `assemble-caution.sh` line numbers in the original Location were the **hub's** - copy while the text named the shim's — both files are now cited explicitly. -- `src/enclave-builder/src/manifest.rs:14-45` — `EnclaveManifest` has no `debug` - field, so the key list appears in no attestation response. Confirms the - "unmeasured" half of item 6x. -- No sshd/sudo/user hardening anywhere in the 535-line `user-data.sh`, so the - AL2023 defaults (ec2-user, passwordless sudo, pubkey sshd) stand. This is the - one assumption the finding rests on that could not be executed here; it is the - documented AWS default for that AMI family and is already load-bearing for the - audit's other parent-host findings. - -**Re-verified in the target:** the five falsified texts at the exact lines quoted; -the assembler branch (`shim/deploy/caution/assemble-caution.sh:449-462`) tests -`$SSH_KEYS` and never `$DEBUG`, while `:120-123` prints the reassuring NOTE and -`:567-574` flips only `enabled` under `--debug`; `deploy.sh:52` (`DEBUG=${DEBUG:-1}`) -and `:128-135` (key only with `--debug`); `log_verdict` at `intercept.rs:577` -with the `Class::Migration` arm at `:582`; the hub's per-admission line at -`server.rs:269`; both subscribers defaulting to `info`. - -**THE BOUND, which the report must carry verbatim and which no other finding may -quietly drop:** `debug.ssh_keys` on its own does **not** hand over the console. -The image was launched without `--debug-mode`, so `nitro-cli console` refuses and -the capture service was never installed. **No log finding may be escalated via -this route directly.** The confirmed -`hub-per-admission-info-log-is-a-real-time-per-entry-arrival-feed.md` (Low) and -the Case B bound of the confirmed -`log-verdict-logs-migration-value-balance-at-info.md` are correct as they stand and -must not be re-inflated on account of this issue. What this issue establishes is -the *conditional*: the closure is a parent-side launch flag, and an operator with -`ssh_keys` can flip it, at the price of zeroed PCRs for the duration. - -**Anti-double-counting, checked against four neighbours:** - -- `hub-manifest-debug-block-claims-ssh-keys-render-empty-and-ssh-is-closed-under-attestation.md` - (plausible; item 6x recommends Low → High) owns **obtaining the parent-host - shell** — the doc/renderer contradiction and the packet-capture, DNS-log, - iptables and restart capabilities that follow. This issue owns **only the - console leg**: that the shell also reaches the enclave's `tracing` stream, which - that file explicitly and wrongly excludes. Its "Note what is not at risk: - enclave console output" paragraph should be struck when it is graded; its - severity argument should not be increased on account of this file, nor this - file's on account of it. -- `core-linkage-survives-in-the-attested-deployment-…` (confirmed High) uses - parent-host access as **step 1**. That step is the *shell*, owned by the sibling - above — not the console. This issue adds nothing to the linkage chain and must - not be counted into it. -- `deploy-script-defaults-to-debug-mode-which-turns-attestation-off.md` (confirmed) - owns the shipped `DEBUG=1` default, which delivers the same console with no - ssh key and no relaunch. This issue is the *attested* deployment's version of - the same exposure and is strictly narrower. -- `log-verdict-logs-migration-value-balance-at-info.md` (confirmed High) owns the - content of the leak. This issue owns only the reachability of the channel. - -**One consequence explored during validation and deliberately NOT claimed here.** -The worst thing a root shell on the parent can do is not read the console: it is -replace `/opt/nitro/enclave.eif` and run a different image entirely. That -capability belongs to the parent-shell issue above, and its detectability is -already owned by two confirmed issues — `caution verify` would flag PCR0/PCR1, and -`deploy.sh:220` tells verifiers to expect exactly those two to fail and accept -PCR2 alone, which item 6q established is a universal constant. It is recorded here -so the report can state the residual once, in the right place, rather than three -times. - -**Severity justification — Medium.** -*Why not Low:* the capability is the primary adversary position obtained inside a -deployment the product presents as operator-blind; it is invisible in the -attestation (no `debug` field in the manifest), invisible to wallets, and denied -five times in the project's own text, including once as a load-bearing design -rationale. The harm delivered — per-migration value balance beside the wallet's -TCP source address — is the exact property `README.md:33` promises. -*Why not High:* it needs a deliberate operator act that no shipped script -performs; the deployed attested state does not leak (the bound above); the -relaunch zeroes PCRs for its duration, so a verifier checking *at that moment* -sees it; and every downstream harm is separately graded, including two cheaper -routes to a worse outcome for the same adversary (the `DEBUG=1` default, and -`shim-submits-every-migration-to-every-configured-hub-…` at High). - -**Nothing in the filing was found to be false.** The changes are: three corrected -platform citations, two additional falsified texts (`hub/deploy/caution/OPERATORS.md:138` -and `deploy.sh:227-229`), the verified `0.0.0.0/0` reach of the SSH rule and the -public EIP, the `EnclaveManifest` confirmation, the precise re-statement of what -attestation buys (PCR0/PCR1 and `.manifest.run_command`, both skipped by the -project's own recipe), and the explicit bound and anti-double-count map above. - -DO NOT DEVIATE FROM THIS TEMPLATE (`docs/SECURITY-ISSUE-TEMPLATE.md`). -LEAVE THIS MESSAGE IN PLACE SO THAT ALL AGENTS KNOW NOT TO DEVIATE FROM THIS TEMPLATE (`docs/SECURITY-ISSUE-TEMPLATE.md`). diff --git a/zeronym-22aa9851caf68-high-medium/medium/attested-tls-binding-is-verified-once-by-hand-if-ever-so-operator-certificate-substitution-has-an-unbounded-undetected-window.md b/zeronym-22aa9851caf68-high-medium/medium/attested-tls-binding-is-verified-once-by-hand-if-ever-so-operator-certificate-substitution-has-an-unbounded-undetected-window.md deleted file mode 100644 index 02944bfb..00000000 --- a/zeronym-22aa9851caf68-high-medium/medium/attested-tls-binding-is-verified-once-by-hand-if-ever-so-operator-certificate-substitution-has-an-unbounded-undetected-window.md +++ /dev/null @@ -1,456 +0,0 @@ -# The enclave's TLS certificate binding is real but time-of-check-only: zeronym schedules no re-verification, wallets never check it, and the continuous signal `README.md:71` offers instead — Certificate Transparency — cannot carry the signal on a diskless enclave - -**Severity**: Medium -**Validation Status**: Confirmed -**Location**: `audit-target/zeronym/README.md:26` (the unconditional "Protected → Broadcast contents" claim) and `:30-36` (the "Not protected" list, which omits certificate substitution); `audit-target/zeronym/README.md:71` (the auditor recipe, which names Certificate Transparency as the anti-substitution defence); `audit-target/zeronym/shim/deploy/caution/OPERATORS.md:151-217` (the "Verify" section — one-shot), `:132-137` (the documented NXDOMAIN window), `:341-343` (the diskless re-issuance rule) and `:345-347` (the CT watch, assigned to the operator); `audit-target/zeronym/hub/deploy/caution/OPERATORS.md:97-124` and `:244`; `audit-target/zeronym/deploy.sh:220` (the only verification instruction emitted at the moment of use, which runs nothing); `audit-target/zeronym/shim/deploy/caution/RESTARTS.md` (the issuance ledger the CT check would have to be compared against); `audit-target/zeronym/OPEN-QUESTIONS.md:108` (a real endpoint for which the CT signal is zero) -**Found by agent:** Local (file audit of `shim/src/tls.rs`), re-scoped by the validator after the G10/G12/G13 primary-source reversal -**In scope of audit?** Yes — `README.md` and both `OPERATORS.md` files are in scope "as security claims"; under ICTM a property users are told they get but do not get is itself a bug. - -## Description - -The single headline protection this product offers is stated unconditionally at -`README.md:26`, under **Protected**: - -> - **Broadcast contents.** An Orchard-touching transaction is hidden from the -> operator: the wallet's TLS terminates inside an attested enclave, not at the -> operator's indexer. - -The residual exposure is well known: the operator holds the DNS zone and the -public IP for the wallet-facing name, so they can obtain a second, entirely -legitimate CA certificate for that name and terminate the wallet's TLS themselves, -in front of the enclave, on a host they own. The wallet cannot tell — both -certificates chain to a public root in `webpki-roots`, both carry the expected -name. - -**A cryptographic defence against exactly this exists and ships today.** The -platform's in-enclave Caddy publishes the SHA-256 of its own served leaf -certificate; `bootproofd` places that fingerprint *inside* the COSE-signed Nitro -attestation as `user_data.tls.certfp`; and `caution verify` compares it against -the leaf of the very TLS connection that carried the `/attestation` response. An -operator terminating 443 with their own certificate therefore **fails** -verification, with the message *"attested TLS certfp does not match the live leaf -certificate"*. (Derivation in Technical Details. This corrects the original -version of this issue, which asserted that no such binding existed — see -Validation Information.) - -The finding is what surrounds that control, and it is four things that compound: - -1. **Wallets get nothing from it.** The platform vendor states this as a design - property of the mode: *"Attested TLS deliberately preserves ordinary browser - HTTPS expectations, so the client does not validate Nitro evidence."* A wallet - sees a valid WebPKI certificate for the right name in both the honest and the - substituted case. So the protection at `README.md:26` is not a property of a - wallet's connection; it is a property of somebody else having recently checked. - -2. **Nobody has been asked to check on a schedule, and nothing does.** The same - vendor documentation states the precondition of the mode plainly: *"To rely on - Attested TLS, carefully verify fresh Nitro evidence against reviewed source and - expected PCR0, PCR1, and PCR2 **on a regular schedule** and after relevant - deployment, DNS, or certificate changes"*, and ships `Caution Canary - --e2e-mode tls` for continuous enforcement. zeronym runs no canary. `deploy.sh` - verifies nothing — its entire contribution is one printed log line - (`deploy.sh:220`). Both `OPERATORS.md` "Verify" sections are one-shot manual - steps performed once at deploy time, and the hub's runbook explicitly - deprioritises verification after an incident (`:244`: *"`caution verify` is a - further ~7 min but does NOT belong on the critical path — restore service - first, verify after"*). A case-insensitive sweep of the whole target - for `periodic|re-verif|reverif|on a schedule|scheduled verif|cron|canary` - returns exactly two hits, and **neither is about verifying a live endpoint**: - `hub/src/queue.rs:245` (OS entropy reseeding) and - `shim/deploy/README.md:834` ("Independent re-verification, 2026-07-31"), which - is a one-off *build*-reproducibility exercise against a since-superseded binary - hash. No dated `caution verify` transcript exists for any deployed endpoint. **The window during which a substitution goes - undetected is therefore unbounded in the shipped operating model.** - -3. **The only continuously-available signal the README offers cannot carry the - signal.** `README.md:71` names Certificate Transparency as the anti-substitution - check. Two independent properties of this deployment make CT unreadable here: - - **The enclave is the loudest source of noise in the log the auditor is asked - to read.** The enclave is diskless: the platform's Caddy writes its ACME - state to `/var/lib/caddy` inside the initramfs, which is tmpfs, so every - restart is a fresh Let's Encrypt order, and Caddy renews on its own schedule - besides. `shim/deploy/caution/OPERATORS.md:341-343` states the consequence in - its own words — *"the enclave is diskless, so every restart is a fresh Let's - Encrypt order, and every push spends one of the hostname's 5 weekly - production issuances"*. A shadow certificate is one more row among rows - identical in issuer, SANs and validity, appearing at times no external - observer can predict. - - **The signal can be zero, not merely noisy.** The shim is a drop-in behind - an operator's existing public URL (`README.md:60`). An operator who already - holds an unexpired certificate for that name — the normal case, since they - ran the endpoint before the shim existed — issues nothing and generates **no - CT record at all**. `OPEN-QUESTIONS.md:108` records exactly this for a real - endpoint: *"The `zec.rocks` certificate. Its existing TLS cert is valid - through October, so the scheme is ineffective for that domain until then."* - - The CT watch is additionally assigned, in the one place it becomes an - operating instruction, to the operator — i.e. to the adversary the check exists - to catch (`shim/deploy/caution/OPERATORS.md:345-347`) — and the artefact an - independent party would have to compare CT against does not cover the fleet. - `shim/deploy/caution/RESTARTS.md:114-118` states its own security property: - *"a certificate for these names that does not appear in this file is either an - unrecorded deploy or someone else's certificate for our domain. The Auditor - Role in the Zeronym design exists partly to watch for exactly that."* But the ledger holds tables for exactly two hostnames, - `zis-zaino.shieldedinfra.net` (`:76-83`) and `zis-lwd.shieldedinfra.net` - (`:86-91`), both Shielded Labs' own; it has **no rows at all** for - `test-shim-nym-1.shieldedinfra.net`, the hostname `deploy.env.example:13` - ships; and it is already incomplete within its own document, since `:29` - records a deploy to `zis-lwd-test-1.shieldedinfra.net` that appears in neither - table. A third-party operator publishes no such ledger at all. - -4. **`README.md` never states the residual.** Certificate substitution appears - nowhere in the "Not protected" list at `README.md:30-36`, and `README.md:26` - is unconditional. A reader is not told that the protection they are promised - depends on a verification practice, let alone that the practice is unscheduled. - -## Attack Scenario and Steps - -The adversary is the indexer operator — adversary #1 in -`audit-context/AUDIT-INSTRUCTIONS.md` and the reason the product exists. - -1. The operator deploys the shim honestly. `caution apps create` allocates an - Elastic IP; the operator points `` at it in their own DNS zone - (`deploy.sh` writes the record into their Vultr zone). An auditor runs - `caution verify --attestation-url https:///attestation` at time - `T0`, gets `✓ Base Nitro attestation and expected PCR0/1/2 verified`, - `✓ TLS certificate binding verified` and `✓ Attestation verification PASSED`, - and publishes that the endpoint is good. Wallets are pointed at it. -2. At any later time `T1` the operator obtains a second certificate for - `` from any public CA. They control the DNS zone (DNS-01) and the - public IP (HTTP-01), so this is one `certbot` invocation and no privilege they - do not already hold. If they already have an unexpired certificate for the - name, they skip this step and generate no CT record. -3. On the parent host they stop forwarding 443 to the enclave and terminate it - themselves, re-originating to the enclave behind. They own the parent host; - this is a firewall rule and a Caddy/nginx config, not an exploit. -4. From `T1` onward every wallet request is readable in the clear at the - operator, including every `SendTransaction` body — every Orchard-touching - migration the product exists to hide — together with the source IP that sent - it. That is the exact IP → transaction → balance join `README.md:54` - identifies as "the attack". -5. Detection, in order of what is actually available: - - **`caution verify` would catch it**, at any moment anyone chose to run it. - Nothing runs it. There is no canary, no cron, no deploy-time gate, no - published transcript, and no document asking anyone to repeat it. The - auditor's `T0` verdict carries no expiry and is the only one on record. - - **The wallet cannot catch it**, by the mode's design. - - **Certificate Transparency cannot catch it**, for the two reasons in - Description point 3. -6. If and when someone does re-verify, the substitution is exposed — permanently - and unambiguously. The operator's exposure is therefore a function of how - often anyone re-verifies, which today is: never, on the evidence in the - repository. - -**Attack Requirements and Assumptions:** - -- The attacker must be the operator of the endpoint, or anyone who can obtain a - CA-issued certificate for the name and get on path — which, for the holder of - the name's DNS zone and IP, is the same person. -- No software vulnerability is required. Every step uses authority the operator - already holds by construction. -- No wallet change, no user action, and no wallet-visible error is involved. -- **What limits the attack, and why this is Medium rather than High:** unlike the - configuration-level attacks elsewhere in this audit, a working detector exists, - it is cheap (~7 minutes), it needs no Caution account and no checkout of the - operator's, and anyone in the world may run it at any time. A rational operator - must weigh permanent exposure against the gain. The defect is that nothing - converts that latent detector into an actual one. - -## Impact on Users - -Every user of an affected endpoint loses the product's headline protection -silently and completely, for as long as nobody re-verifies. The operator recovers -exactly what the shim was deployed to deny them: source IP, timing, and the full -plaintext of every Orchard-touching broadcast — joinable against the public chain -retrospectively and permanently. - -Under ICTM the finding stands independently of whether any operator ever does it: -`README.md:26` tells users this is **Protected**, unconditionally, and the "Not -protected" list immediately below does not mention certificate substitution. The -true statement is *"protected as long as somebody re-runs `caution verify`"*, and -the shipped operating model contains no such somebody. A user who reads the -README holds a belief about their protection that the deployment does not support. - -## Technical Details / Code Analysis - -### The binding that exists (and which the original version of this issue denied) - -Established from the Caution platform's own public source -(`codeberg.org/caution/platform`) and `bootproofd` -(`git.distrust.co/public/bootproof.git`), both cloned during this audit. - -1. **In the enclave.** `src/enclave-builder/templates/caddy-certfp.sh` runs on a - 60-second loop, connects to the enclave's own Caddy with - `openssl s_client -connect 127.0.0.1:443 -servername - -verify_return_error -verify_hostname -purpose sslserver -CAfile …`, - takes the SHA-256 of the served leaf, and writes - `{"tls":{"mode":"tls","domain":"","certfp":""}}` to - `/metadata.json`. -2. **`bootproofd`** passes `/metadata.json` as `user_data` into - `Nitro.generate(user_data, nonce)`, i.e. **inside the COSE-signed document** - (`crates/bootproofd/src/routes/nonced_attestation.rs:88-118`). -3. **`caution verify`** enforces all four fields - (`src/cli/src/lib.rs:354-384`): - - ```rust - anyhow::ensure!(user_data.tls.mode == "tls", "attested TLS mode is not tls"); - anyhow::ensure!(user_data.tls.domain == expected.domain, …); - anyhow::ensure!(user_data.tls.certfp.len() == 64 && …lowercase hex…, - "attested TLS certfp is not lowercase SHA-256 hex"); - anyhow::ensure!(user_data.tls.certfp == observed_certfp, - "attested TLS certfp does not match the live leaf certificate"); - ``` - - `observed_certfp` is `sha256(leaf DER)` of the **same** WebPKI-validated, - redirect-disabled HTTPS response that carried `/attestation` - (`src/cli/src/lib.rs:6892-6912`), and `expected.domain` is read from the - *reproduced* `caution.hcl` and is therefore itself PCR-bound - (`tls_expectation_from_config`, `:291-312`). - -`shim/deploy/caution/OPERATORS.md:189-194` records the project observing this -work on the 2026-08-14 attested pair, "with the TLS certificate binding -verified". - -**Why `shim/src/tls.rs` is not the relevant code.** `ServerTls` never runs on any -deployment this repository ships: `shim/deploy/caution/caution.hcl.tmpl:163-176` -leaves `ZIS_TLS_DOMAIN` deliberately unset (*"the in-enclave Caddy declared in the -http block above owns the certificate"*). Wallet-facing TLS is terminated by the -platform's Caddy, which does the binding above. See -`servertls-is-unreachable-in-every-attested-deployment-the-repo-ships.md`. - -### The two documented paths that print PASSED without the binding - -`src/cli/src/lib.rs:7261-7291`: - -```rust -let tls = if pcr_only { - TlsVerification::PcrOnly -} else if let Some(ref expected) = expected_tls { - self.verify_tls_binding(expected, &payload, &attestation_url, attestation_leaf.as_deref()) - .await.context("TLS certificate binding failed")? -} else { - TlsVerification::NotApplicable -}; -``` - -- `--pcrs` yields `TlsVerification::PcrOnly`, printed as *"TLS certificate - binding: not performed (--pcrs)"*, and `✓ Attestation verification PASSED` - still prints below it. -- `TlsVerification::SkippedNoDns` (`:6913-6932`) is reached only on the **raw-IP** - attestation-URL flow, when the configured domain has no DNS answer or does not - resolve to the deployment IP. `shim/deploy/caution/OPERATORS.md:132-137` - documents an NXDOMAIN window of about a minute after every deploy, which is when - an operator is most likely to reach for the raw IP. - -Caution's own documentation warns about both (*"Do not treat that result as -Attested TLS verification"*; *"Do not use `--pcrs` for this check"*). **No zeronym -document mentions either**, nor tells a verifier to require the -`✓ TLS certificate binding verified` line. (The README half of that omission is -filed separately as -`auditor-recipe-omits-the-two-checks-that-decide-where-plaintext-goes-and-names-a-defence-the-platform-does-not-rely-on.md`.) - -### Nothing in the tree performs or schedules a verification - -Every occurrence of `caution verify` in the target is a one-shot manual -instruction or a comment about one: `deploy.sh:132`, `:198`, `:220`, `:332`; -`shim/deploy/caution/OPERATORS.md:32`, `:92`, `:158`, `:178`; -`shim/deploy/caution/README.md:30`, `:126`, `:132`; -`hub/deploy/caution/OPERATORS.md:100`, `:244`; `hub/deploy/caution/README.md:57`; -plus the assembler's `--app-source` warnings. `deploy.sh:220` — the only -verification guidance emitted at the moment of use — prints advice and runs -nothing: - -```sh -log "verify with: caution verify (expect PCR0/1 FAILED on Caution's floating framework; PCR2 is the check that matters)" -``` - -(That line's *content* is separately wrong and separately filed as -`deploy-script-tells-operators-to-expect-pcr01-failure-and-accept-pcr2-alone.md`; -what matters here is that it is a `log` call, not a check.) - -### Why the CT channel is unreadable, mechanically - -`src/enclave-builder/templates/run.sh.template:104-126` starts Caddy with -`HOME=/var/lib/caddy XDG_DATA_HOME=/var/lib`, directories created by `mkdir -p` -inside the initramfs. The enclave has no persistent storage, so certificate and -ACME account state do not survive a restart: every restart is a fresh order. -`caddy-certfp.sh`'s 60-second re-publication loop exists precisely because the -served certificate changes underneath the attestation. The project's own runbook -budgets for this (5 production issuances per name per week) and `RESTARTS.md` -exists to manage it. So the enclave emits a stream of CT rows for the same name, -at times an external observer cannot predict, and a shadow certificate is one more -such row. - -## Recommendations - -In order of value: - -1. **Make the check happen on a schedule and publish the result.** Adopt - `Caution Canary --e2e-mode tls`, or an equivalent cron'd - `caution verify --attestation-url https:///attestation` run by a party - other than the operator, and publish a dated, per-endpoint transcript that - includes the `✓ TLS certificate binding verified` line. Attested TLS is only as - strong as the frequency of that check, and today the frequency is zero. This is - the platform vendor's own stated precondition for the mode zeronym has chosen, - and relaying it is the whole fix. -2. **Correct `README.md:26` and the "Not protected" list.** Move certificate - substitution into "Not protected", and state the residual as it is: *the - operator can terminate the wallet's TLS with their own certificate for the same - name; the wallet cannot detect it; `caution verify` can and does, but only when - somebody runs it, and nothing schedules a run.* -3. **Have `deploy.sh` run the check rather than describe it,** and fail the - attested deploy unless all three success lines appear — including - `✓ TLS certificate binding verified`. A verification step that is only ever - printed is one that frequently does not happen. -4. **Stop assigning the detection to the adversary, and demote CT.** - `shim/deploy/caution/OPERATORS.md:345-347` should name who *other than the - operator* watches, and against what published record. If CT is kept at all it - should be labelled supplementary, with an honest note that the enclave's own - re-issuance makes the channel noisy and that a drop-in deployment on a name the - operator already certifies produces no signal whatsoever. If the answer is the - "Auditor Role", the issuance ledger must be a published, machine-readable, - per-endpoint artefact that third-party operators maintain — not a hand-edited - table in Shielded Labs' own repo that is already missing rows. -5. **Warn about the two skip paths.** Any documented verification step must state - that `TLS certificate binding: not performed (--pcrs)` and - `TLS certificate binding: skipped because the configured domain has no DNS - answer` are **failures for this purpose**, notwithstanding the - `✓ Attestation verification PASSED` printed beneath them. - -Cross-references: -`auditor-recipe-omits-the-two-checks-that-decide-where-plaintext-goes-and-names-a-defence-the-platform-does-not-rely-on.md` -(the `README.md:71` half — CT named instead of the binding); -`operator-controlled-dns-permits-a-layer-4-relay-that-every-documented-verification-step-passes.md` -(the *weaker but wholly undetectable* interposition: no certificate, no CT trace, -and the certfp binding still verifies, because TLS is never terminated); -`shim-operators-runbook-tls-name-instruction-names-the-wrong-adversary-and-cannot-be-satisfied.md` -(the only pre-deployment defence, and why it cannot be satisfied); -`open-questions-operator-error-alarm-asks-operators-to-accept-a-signal-that-fires-on-every-restart-by-design-and-to-apply-mitigations-they-cannot.md` -(`OPEN-QUESTIONS.md:110`'s "auto-renewal disabled" mitigation, which a diskless -enclave cannot implement); -`servertls-is-unreachable-in-every-attested-deployment-the-repo-ships.md` -(why `shim/src/tls.rs` is not the code that terminates wallet TLS). - -## Validation Information - -**Validated 2026-08-18. CONFIRMED at Medium, after a full re-scope.** The issue -as originally filed was titled *"Nothing binds the wallet-facing TLS identity to -the enclave attestation, and the Certificate-Transparency check the README gives -auditors as the defence cannot detect operator substitution"*, was graded High, -and rested on an analysis of `shim/src/tls.rs`. **Its first clause is false and -its attack scenario does not succeed as it was written.** Those parts have been -removed rather than softened, because leaving them would have put a false positive -in `confirmed/`. What is above is the subset that survives primary-source -verification, plus the residual that the reversal exposed. - -### What was refuted, and how - -Re-derived by the validator from the platform's public source, not accepted from -the earlier correction note: - -- **A binding exists, and it is precisely the original issue's own - recommendation 1.** `validate_attested_tls` (`src/cli/src/lib.rs:354-384`) was - read in full, together with its caller `verify_tls_binding` (`:6892-6961`) and - the `TlsConnection::AttestationResponse` branch that supplies `observed_certfp` - from the leaf of the same connection that carried `/attestation`. The original - Description point 3 (*"The attestation is fetched over the channel under attack, - with no channel binding"*) is **withdrawn**: that binding is exactly what - exists. -- **Attack-scenario steps 3-5 therefore fail as written.** An operator - terminating 443 with a second legitimate certificate cannot produce the - enclave's private key and cannot alter the signed `user_data`, so - `caution verify` returns *"attested TLS certfp does not match the live leaf - certificate"* and verification fails. -- **The analysis targeted dead code.** `ZIS_TLS_DOMAIN` is deliberately unset in - the shipped manifest (`shim/deploy/caution/caution.hcl.tmpl:163-176`, read in - place), so `ServerTls` never runs and `NoCache`'s effects are the platform - Caddy's, not `tls.rs`'s. The original issue's `NoCache` argument survives only - because the platform Caddy is diskless for the same reason — that translation - was verified against - `src/enclave-builder/templates/run.sh.template:104-126` rather than assumed. -- **The `ZIS_CAUTION_ATTESTATION` addendum is retired.** The in-enclave Caddyfile - generated at `run.sh.template:107-121` gives `/attestation` its own `handle` - block reverse-proxying `127.0.0.1:49502` (bootproofd), with the bare `handle` - as fallback. Caddy sorts `handle` blocks by path specificity, so `/attestation` - is answered by Caddy and never reaches the shim; `shim/src/proxy.rs`'s - `Route::CautionAttestation` arm is unreachable in any `mode = "tls"` Caution - deployment, which is every deployment this repository ships. - -### What survives, and why it is a finding rather than an inherent limitation - -Four claims were checked and all four hold: - -1. **Wallets never validate the binding.** Verified against the platform vendor's - own documentation: *"Attested TLS deliberately preserves ordinary browser HTTPS - expectations, so the client does not validate Nitro evidence."* This is a - property of the mode, not a defect — but `README.md:26` states the resulting - protection unconditionally and `README.md:30-36` omits the residual, and that - is the defect. -2. **The mode's own precondition is unimplemented.** The same vendor - documentation: *"To rely on Attested TLS, carefully verify fresh Nitro - evidence … on a regular schedule and after relevant deployment, DNS, or - certificate changes"*, and *"For continuous enforcement, Caution Canary - supports an Attested TLS profile configured with `--e2e-mode tls`."* Every - `caution verify` reference in the target was enumerated (fifteen of them, listed - in Technical Details); all are one-shot, and none mentions repetition. This is - what makes it a defect rather than an inherent limitation: zeronym selected a - mode with a stated operational precondition and relays none of it. -3. **CT cannot carry the signal here.** Both legs verified in the target: the - diskless-enclave churn (`run.sh.template:104-126` writes Caddy state into the - initramfs; `shim/deploy/caution/OPERATORS.md:341-343` states the 5-issuances-per-week - consequence in the project's own words) and the pre-existing-certificate case - (`README.md:60` drop-in; `OPEN-QUESTIONS.md:108` for a named real endpoint). - This half of the original issue is unchanged and is the part three other - plausible issues cite as their reason for treating CT as unavailable — which is - the second reason this issue is confirmed rather than invalidated. -4. **The two skip paths are real but thin.** `--pcrs` is a self-describing flag - whose help text reads *"Compare against PCRs from file without TLS certificate - binding"*, so it belongs at the bottom of the recommendation list rather than in - the finding. `SkippedNoDns` is reachable only on the raw-IP flow - (`tls_connection`, `src/cli/src/lib.rs:221-241` — an HTTPS URL whose host equals - the configured domain always takes the binding branch), which narrows it - considerably from how the correction note framed it. Both are recorded at their - true weight above. - -### Why Medium - -Impact is maximal for the affected endpoint's users — complete, silent loss of the -product's headline protection, with the source IP joined to the transaction -plaintext. Two things hold it below High. First, a working cryptographic detector -exists, is cheap, and is runnable by anyone in the world without an account or a -checkout, so the operator's exposure is permanent and unbounded in time rather -than nil; that is a materially different risk calculus from -`shim-submits-every-migration-to-every-configured-hub-…` (High), where no check -anyone is asked to run catches the attack at all. Second, the surrounding -documentation defects — the recipe naming CT instead of the binding, the -unsatisfiable name-choice instruction, the layer-4 relay variant — are each filed -and graded separately, and this issue is deliberately scoped to what none of them -owns: that the binding is time-of-check-only, that nothing schedules the check, -and that `README.md:26` states the resulting protection unconditionally. - -It is held above Low because the undetected window is not merely long but -*unbounded*, no verification transcript exists for any endpoint on the record, and -the ICTM gap is in the product's single headline claim rather than in a -supporting document. - -### Considered and rejected during validation - -- **Marking the whole issue Invalid and letting the recipe issue absorb the - remainder.** Rejected on two grounds. (a) After the reversal, no other issue - owns the certificate-substitution attack itself; the recipe issue owns the - omission in `README.md:71`, the DNS issue owns the *non*-terminating relay, and - the name-choice issue owns the pre-deployment mitigation — none of them states - the attack, its detector, and the absence of a schedule. (b) Three other - plausible issues cite this file by name as the establishment of "CT cannot work - here"; invalidating it would strand that reference while the underlying claim is - true. -- **Keeping the High.** Rejected: the control the original issue asked for exists - and works, so the grade must reflect a scheduling and disclosure gap rather than - a missing control. -- **Splitting "no scheduled verification" from "CT cannot work".** Rejected as - over-fragmentation: they are one argument — the continuous signal offered is - unreadable, and the readable signal is not continuous — and separating them - would leave two issues each of which is unpersuasive alone. - -DO NOT DEVIATE FROM THIS TEMPLATE (`docs/SECURITY-ISSUE-TEMPLATE.md`). -LEAVE THIS MESSAGE IN PLACE SO THAT ALL AGENTS KNOW NOT TO DEVIATE FROM THIS TEMPLATE (`docs/SECURITY-ISSUE-TEMPLATE.md`). diff --git a/zeronym-22aa9851caf68-high-medium/medium/auditor-recipe-omits-the-two-checks-that-decide-where-plaintext-goes-and-names-a-defence-the-platform-does-not-rely-on.md b/zeronym-22aa9851caf68-high-medium/medium/auditor-recipe-omits-the-two-checks-that-decide-where-plaintext-goes-and-names-a-defence-the-platform-does-not-rely-on.md deleted file mode 100644 index 3831781b..00000000 --- a/zeronym-22aa9851caf68-high-medium/medium/auditor-recipe-omits-the-two-checks-that-decide-where-plaintext-goes-and-names-a-defence-the-platform-does-not-rely-on.md +++ /dev/null @@ -1,509 +0,0 @@ -# `README.md:71`'s auditor recipe omits both checks that decide where a user's plaintext actually goes, and substitutes Certificate Transparency for the certificate-binding check the platform actually provides - -**Severity**: Medium -**Validation Status**: Confirmed -**Location**: `audit-target/zeronym/README.md:71` (the four-step recipe, and the only verification instruction any user-facing document gives); `audit-target/zeronym/shim/deploy/caution/OPERATORS.md:151-217` (the "Verify" section that implements it) and `:345-347` (the CT watch); `audit-target/zeronym/hub/deploy/caution/OPERATORS.md:97-124`; `audit-target/zeronym/deploy.sh:220` (the same instruction at the moment of use); `audit-target/zeronym/shim/deploy/caution/caution.hcl.tmpl:129-184` and `audit-target/zeronym/hub/deploy/caution/caution.hcl.tmpl:110-141` (the `unit "default" { env = { … } }` blocks the recipe never reads) -**Found by agent:** Global (focus areas G10 / G12 / G13 — the attestation and reproducibility chain, end to end) -**In scope of audit?** Yes — markdown claims are in scope as security claims; under ICTM a property users are told they get but do not get is itself a bug, and `README.md:71` is the substitute the product offers for trusting the operator. - -## Description - -`README.md:71` is the whole of what a wallet user is offered in place of trusting -their indexer operator: - -> - **Auditors** verify an endpoint without trusting its operator: fetch its -> attestation, check the PCRs against the AWS Nitro root, reproduce the build -> and compare hashes, and check Certificate Transparency for a shadow -> certificate. - -Read against what the Caution platform actually does — established during this -audit by reading the platform's own public source at -`https://codeberg.org/caution/platform` and `https://git.distrust.co/public/bootproof.git` -rather than from in-repo comments — the recipe is wrong in four separate places, -and each error points the auditor away from a check that exists and toward one -that does not do the job. - -**1. It never reads the configuration, which is the thing that decides where the -plaintext goes.** `ZIS_HUB_NYM` names the hub every diverted migration is sent to. -`ZIH_INDEXER_TLS` is the only thing standing between the hub's outbound batch and -a plaintext read. `ZIS_CAUTION_ATTESTATION` decides who answers `/attestation`. -All three live in `unit "default" { env = { … } }` in the deployed `caution.hcl`, -all three **are** measured (see Technical Details), and **all three are served, -verbatim, by the endpoint itself** in the `manifest.run_command` field of every -`/attestation` response. The recipe contains no step that looks. Every step in it -passes on a shim pointed at a hub the operator runs. - -**2. "check the PCRs" does not say which, and one of the three is a constant.** -On Caution's EIF layout PCR2 is the measurement of an *absent* application -ramdisk: `sha384(0^48 ‖ sha384(""))` = `21b9efbc18480766…`, identical for every -Caution enclave that has ever booted, whatever code it runs — and PCR0 and PCR1 -are computed over the identical byte stream, so they are necessarily equal to each -other. `deploy.sh:220` then tells the operator, at the moment of use, to expect -PCR0/1 to fail and to accept PCR2. (That instruction is separately filed as -`deploy-script-tells-operators-to-expect-pcr01-failure-and-accept-pcr2-alone.md`; -this issue is about the README sentence that leaves the door open for it by not -naming a criterion at all.) - -**3. "reproduce the build and compare hashes" compares a hash nothing else uses.** -`caution verify` never reads `EXPECTED_SHA256`; `reproduce.sh` never computes a -PCR. The platform builds the app with `docker build -f .` and no -`--target`, i.e. the Containerfile's **last** stage (`runtime`), and measures that -filesystem into PCR0/PCR1; `reproduce.sh` builds the **`export`** stage and hashes -its tar. The two artefacts are different and nothing in the tree or the tooling -joins them. - -**4. It names Certificate Transparency as the anti-substitution defence, when the -platform ships a cryptographic one and the recipe omits it.** Caution's Attested -TLS puts the SHA-256 of the DER-encoded **leaf certificate** into the -Nitro-signed `user_data.tls.certfp`, and `caution verify` compares it against the -leaf of the very WebPKI-validated, redirect-disabled TLS connection that carried -the `/attestation` response. An on-path operator terminating 443 with their own -certificate for the same name therefore **fails** verification. That is the check -an auditor should be told to require — by name, as the output line `✓ TLS -certificate binding verified`. Instead they are told to read crt.sh, which -`attested-tls-binding-is-verified-once-by-hand-if-ever-so-operator-certificate-substitution-has-an-unbounded-undetected-window.md` -shows cannot work here. (Precisely stated after validation: both runbooks do -*mention* the binding, but only as narration of expected output — -`shim/deploy/caution/OPERATORS.md:192-193` and -`hub/deploy/caution/OPERATORS.md:108`. Neither they nor `README.md:71` tell a -verifier to **require** that line, and the string `certfp` appears nowhere in -zeronym at all.) The binding also has two paths on which -`caution verify` prints `Attestation verification PASSED` **without** performing -it — `--pcrs` mode, and the **raw-IP** flow (`--attestation-url https:///…`) -when the configured domain has no DNS answer or does not resolve to the pinned -deployment IP (`TlsVerification::SkippedNoDns`). Corrected during validation: -`SkippedNoDns` is reachable **only** from the `TlsConnection::PinnedIp` branch -(`src/cli/src/lib.rs:221-241`, `:6902-6932`), **not** from the -`--attestation-url https:///attestation` flow both runbooks prescribe — -on that flow the certfp comparison is unconditional. So the NXDOMAIN window -`OPERATORS.md` documents after each deploy does **not** produce a silent skip on -the documented flow; it produces a failed HTTPS fetch. The residual is narrower -than originally written but real: a verifier who follows the raw-IP shortcut, or -whose domain has been repointed at a relay, gets `PASSED` with a warning line -instead of a failure, and no zeronym document tells them that line is -disqualifying. - -## Attack Scenario and Steps - -The adversary is the indexer operator: adversary #1 in the threat model and the -reason the product exists. - -1. The operator stands up their own `zero-indexer-hub`. It is public code in the - same repository, needs no consortium membership, and publishes its own Nym - address at `GET /nym-address`. -2. They deploy the genuine, reproducible shim image with - `HUB_NYM=` in `deploy.env`. `assemble-caution.sh:345-355` writes - `ZIS_HUB_NYM = ""` into the manifest's `env` block; the enclave is - built, attested and healthy. They publish the assembled tree as `--app-source`, - because they must: without it verification is impossible and the endpoint looks - suspicious. -3. An auditor performs, in order, exactly the four steps of `README.md:71`: - - *fetch its attestation* — succeeds, on a genuine Nitro document; - - *check the PCRs against the AWS Nitro root* — **succeeds**, because the PCRs - are reproduced from the published tree, which is the tree that was deployed; - - *reproduce the build and compare hashes* — succeeds, because - `EXPECTED_SHA256` is a hash of the binary, and the binary is byte-identical - whichever hub it names (`shim/src/config.rs:110-114` states this as a design - goal); - - *check Certificate Transparency* — nothing anomalous, because nothing - anomalous happened to the certificate. -4. The auditor reports the endpoint as verified. Every Orchard-touching - transaction from every wallet using it is delivered in plaintext into an - enclave the operator provisioned. -5. **The evidence was one field away the whole time.** `caution verify` staged and - rebuilt a tree containing `ZIS_HUB_NYM = ""`, and the endpoint - itself returns that string in the `manifest.run_command` of every - `/attestation` response. Nothing in the recipe, in either `OPERATORS.md` - "Verify" section, or in `deploy.sh`'s printed advice tells anyone to look at it, - or that it is there. - -**Attack Requirements and Assumptions:** - -- The operator must be willing to run a second enclave and publish the repointed - tree. Both are the documented deploy path; `deploy.env.example` ships the hub's - configuration alongside the shim's. -- The attack is **detectable** — that is what distinguishes this issue from an - attestation-coverage gap. It is detectable only by a step the product never - asks for, and the omitted step costs one HTTP request. -- No software vulnerability, platform break, or on-path position is required. -- An auditor who happens to read the reproduced `caution.hcl` in the build - artefacts directory that `caution verify` prints — which Caution's own - documentation tells them to do ("Inspect the staged source, **configuration**, - generated build recipe, and manifest there before deciding whether to trust what - the verified workload does") — catches it. zeronym's documents never relay that - instruction. - -## Impact on Users - -`README.md:71` is the only thing standing between "trust your indexer operator" -and "you do not have to". A user reading it believes that someone performing those -four steps has established that this endpoint diverts their migration away from -the operator. Nobody performing those four steps has established that. What they -have established is narrower and should be stated as such: *some* attested Caution -enclave, built from a published tree, is answering — with the contents of that -tree, and therefore the destination of the plaintext, unexamined. - -The same omission covers the hub's `ZIH_INDEXER_TLS`, whose absence would let the -hub's parent host read every batch before publication, and -`ZIS_CAUTION_ATTESTATION`, whose absence changes who answers the attestation -endpoint on a non-Caution deployment. - -The failure is silent from the user's side: the wallet sees a valid certificate -and correct gRPC behaviour either way. - -## Technical Details / Code Analysis - -### The configuration is measured, and it is served - -**Measured.** The chain, read in Caution's platform source: - -`src/caution-config/src/lib.rs:253-274` — every literal string in `unit.env` -becomes a shell `export` line: - -```rust - pub fn run_command_string(&self) -> Result { - let mut out = String::new(); - if let Some(env) = &self.env { - for (key, expr) in env { - let value = match expr { Expression::String(s) => s, _ => continue }; - if !is_valid_env_key(key) { return Err(FromStrError::InvalidEnvKey(key.clone())); } - let quoted = shlex::try_quote(value).map_err(|_| FromStrError::UnquotableCommand)?; - out.push_str("export "); out.push_str(key); out.push('='); - out.push_str("ed); out.push('\n'); - } - } - ... -``` - -`src/api/src/main.rs:2413-2417` passes it as `run_command`; -`src/enclave-builder/src/build.rs:355-361` and `:468` substitute it into -`run.sh.template` at `{{USER_CMD}}`; `src/enclave-builder/templates/Containerfile.eif` -copies `run.sh` into `/build/initramfs/run.sh`, cpio-archives the initramfs into -`rootfs.cpio.gz`, and passes it as the EIF's single `--ramdisk`. Per -`aws-nitro-enclaves-image-format-0.4.0/src/utils/mod.rs:664-692`, that ramdisk is -measured into PCR0 and PCR1. It is measured a second time via `manifest.json`, -which carries `run_command` as a field (`src/enclave-builder/src/manifest.rs:14-45`) -and is also copied into the initramfs. - -So `shim/src/config.rs:110-114`'s design goal — - -```rust - /// ... Tuning it here changes the enclave config, NOT the binary, - /// so `EXPECTED_SHA256` and the reproducibility trail stay put. ... -``` - -— is true of `EXPECTED_SHA256` and **false of the PCRs**. The configuration is -inside the attestation; it is simply not inside the artefact the README's third -step compares. - -**Served.** `bootproofd` returns the manifest alongside the signed document -(`crates/bootproofd/src/routes/nonced_attestation.rs:20-29, 78-120`): - -```rust -pub struct NoncedAttestationResponse { - /// A base64 encoded attestation document. - pub document: String, - /// The manifest used to build the enclave. - pub manifest: serde_json::Value, -} -``` - -so the missing check is: - -```sh -curl -s -X POST "https:///attestation" \ - -H 'content-type: application/json' -d '{"nonce":"'"$(head -c32 /dev/urandom | base64)"'"}' \ -| jq -r '.manifest.run_command' -``` - -which prints, for a shim, the literal `export ZIS_HUB_NYM='…'` line. The response -manifest is deliberately outside the COSE payload and is therefore unsigned on its -own; a completed `caution verify` is what certifies it, because the same string is -inside the ramdisk the PCRs cover. - -### The certificate binding the recipe does not mention - -In the enclave, `src/enclave-builder/templates/caddy-certfp.sh` polls the served -certificate every 60 s and publishes its fingerprint: - -```sh - if /usr/bin/openssl s_client -connect "${tls_address}" -servername "${caddy_domain}" \ - -showcerts -verify_return_error -verify_hostname "${caddy_domain}" \ - -purpose sslserver -CAfile "${ca_file}" ... ; then - certfp="$( /usr/bin/openssl x509 -in "${served_leaf}" -noout -fingerprint -sha256 \ - | cut -d= -f2 | tr -d ': ' | tr 'A-F' 'a-f' )" - ... - if printf '{"tls":{"mode":"tls","domain":"%s","certfp":"%s"}}\n' \ - "${caddy_domain}" "${certfp}" >"${metadata_tmp}" ... -``` - -`bootproofd` passes `/metadata.json` as the attestation's `user_data` -(`nonced_attestation.rs:88-118` → `crates/bootproof/src/format/nitro.rs:22-62`), -and `caution verify` enforces it (`src/cli/src/lib.rs:353-384`): - -```rust - anyhow::ensure!(user_data.tls.mode == "tls", "attested TLS mode is not tls"); - anyhow::ensure!(user_data.tls.domain == expected.domain, ...); - anyhow::ensure!(user_data.tls.certfp.len() == 64 && ... , "attested TLS certfp is not lowercase SHA-256 hex"); - anyhow::ensure!(user_data.tls.certfp == observed_certfp, - "attested TLS certfp does not match the live leaf certificate"); -``` - -where `observed_certfp` is `sha256` of the leaf DER of the same -redirect-disabled, WebPKI-validated response -(`src/cli/src/lib.rs:6892-6963`). `expected.domain` comes from the *reproduced* -`caution.hcl` (`tls_expectation_from_config`, `:286-312`), so it too is -PCR-bound. The project has observed this working: -`shim/deploy/caution/OPERATORS.md:189-194` records the 2026-08-14 pair verifying -"with the TLS certificate binding verified". - -The two paths that skip it and still print PASSED are -`TlsVerification::PcrOnly` (`--pcrs`) and `TlsVerification::SkippedNoDns` -(`src/cli/src/lib.rs:6919-6932`), reported as -`"TLS certificate binding: not performed (--pcrs)"` and -`"TLS certificate binding: skipped because the configured domain has no DNS answer"`. - -### PCR2, and why "check the PCRs" needs a criterion - -Caution's `Containerfile.eif` calls `eif_build` with exactly one `--ramdisk`. -`aws-nitro-enclaves-image-format-0.4.0/src/utils/mod.rs:674-691` sends ramdisk 0 -to the bootstrap hasher and only ramdisks 1..n to the customer-app hasher, so the -customer-app hasher receives nothing; `defs/eif_hasher.rs`'s -`tpm_extend_finalize_reset()` then returns `sha384(0^48 ‖ sha384(""))`, which -computes to `21b9efbc18480766…` — the exact value both `OPERATORS.md` files record -as "PCR2, identical across shim and hub". By the same token PCR0 and PCR1 are fed -the identical byte stream (kernel ‖ cmdline ‖ ramdisk0) and are therefore equal, -which is also what those tables record. "Check the PCRs" without a criterion is -satisfied by checking a constant. - -### The hash that is compared to nothing - -`src/enclave-builder/src/docker.rs:69` — `format!("docker build -f {} .", containerfile)`, -no `--target`, so the platform builds the final stage (`runtime`). -`*/deploy/reproduce.sh` builds `--target export`. `caution verify` never reads -`EXPECTED_SHA256`; `reproduce.sh` never produces a PCR. - -## Recommendations - -1. **Rewrite `README.md:71` as five steps, with criteria.** Suggested: - *"Auditors verify an endpoint without trusting its operator: run `caution - verify --attestation-url https:///attestation` from a Linux/x86_64 - checkout and require all of `✓ Base Nitro attestation and expected PCR0/1/2 - verified`, `✓ TLS certificate binding verified` and `✓ Attestation verification - PASSED` — a run that says `TLS certificate binding: skipped` or `not performed` - has not verified the endpoint; confirm the commit it staged is the commit you - reviewed; read `unit.env` in the reproduced `caution.hcl` (or - `.manifest.run_command` from the `/attestation` response) and confirm - `ZIS_HUB_NYM` is a hub you trust and `ZIS_CAUTION_ATTESTATION` is absent or - true; and separately run `sh zeronym/shim/deploy/reproduce.sh` against - `EXPECTED_SHA256`, noting that nothing joins that hash to the attestation."* -2. **Publish the expected `ZIS_HUB_NYM` value.** The hub's Nym address is public by - design (`GET /nym-address`). Shielded Labs should publish the canonical value - next to the README so "is this the right hub?" is a string comparison rather - than a judgement call, and note that a shim naming any other hub is not a - zero-indexer deployment. -3. **Add the same steps to both `OPERATORS.md` "Verify" sections** and replace - `deploy.sh:220`'s printed advice with them. -4. **Drop Certificate Transparency from the recipe, or demote it.** It is the - weakest of the available defences here and its presence displaces the strong - one. Keep it as a supplementary monitoring recommendation with an honest note - that the enclave's own re-issuance makes the channel noisy. -5. **Consider having `deploy.sh` run `caution verify` on the attested path** and - fail the deploy unless all three success lines appear. A verification step that - is only ever described is one that frequently does not happen. - -Cross-references: `operators-runbook-attributes-the-hub-destination-to-the-binary-hash-and-egress-rules-neither-of-which-binds-it.md` -(the same trust boundary — confirmed Low, renamed from -`shim-config-hub-identity-is-unattested-unobservable-operator-configuration.md`); -`deploy-script-tells-operators-to-expect-pcr01-failure-and-accept-pcr2-alone.md` -(the PCR criterion, with the PCR2 derivation appended); -`attested-tls-binding-is-verified-once-by-hand-if-ever-so-operator-certificate-substitution-has-an-unbounded-undetected-window.md` -(the CT half — confirmed Medium, renamed from -`wallet-facing-tls-identity-is-not-bound-to-the-attestation-and-the-ct-defence-cannot-work.md`); -`reproduce-never-builds-the-runtime-stage-that-the-enclave-and-pcr0-are-built-from.md` -and `hub-caution-readme-says-the-attestation-binds-the-running-binary-to-expected-sha256.md` -(the hash half). - ---- - -## ADDENDUM (Global audit G30/G32/G17, 2026-08-18) — two amendments to the recommendations, and a fifth omission the recipe has - -Nothing above is retracted. Three things found while sweeping adversary #1's -full capability set bear directly on this issue's recommendations. - -**1. Recommendation 2 must require exact list equality, not membership.** -`NymHandle::submit` sends every diverted transaction to **every** address in -`ZIS_HUB_NYM` (`shim/src/nym.rs:602`, `:642`), and the list is uncapped -(`shim/src/config.rs:262-289`). An operator who *appends* their own hub rather -than repointing gets a plaintext copy of every migration while the canonical hub -keeps publishing normally, so nothing observable changes anywhere — and a -checker asking "does `ZIS_HUB_NYM` name the canonical hub?" answers yes. -`README.md:90` compounds it by stating that submit *rotates* which address it -targets, which is true only of lookups. Filed as -`shim-submits-every-migration-to-every-configured-hub-so-an-operator-appends-their-own-and-gets-a-plaintext-copy-with-nothing-breaking.md`. - -**2. There is a fifth omission, and it is larger than the four above: the recipe -never establishes that the tree `caution verify` reproduced from is -zero-indexer.** `caution verify` clones `app_source.urls` at `app_source.commit` -from the attested manifest, reads the `caution.hcl` from that clone, and rebuilds -the EIF from that clone (`src/cli/src/lib.rs:6432-6473`). Those URLs are the -operator's `--app-source`, and the repository is a per-deployment assembled tree -containing the full Rust source, so it is never `ShieldedLabs/zero` by -construction. No allow-list, no signature, no expected hash, and — since the -Containerfile deploy path passes `binary: None` -(`src/api/src/builder.rs:925-940`) — no binary hash in the attestation to fall -back on. Filed as -`caution-verify-reproduces-from-a-repository-the-operator-nominates-so-nothing-binds-the-attested-code-to-zeronym.md` -(High). Recommendation 1's rewritten `README.md:71` should gain a step for it. - -**3. Two `caution.hcl` blocks are outside every measurement and should be named -in recommendation 1 as things to read rather than to verify by tool:** -`network.ingress` (the CIDRs reach only the security group — narrowing them can -force wallets through an operator-run relay) and `debug.ssh_keys` (which the -platform honours **irrespective of `debug.enabled`**, opening port 22 to -`0.0.0.0/0` on the enclave's parent host — see the addendum on -`hub-manifest-debug-block-claims-ssh-keys-render-empty-and-ssh-is-closed-under-attestation.md`). -Neither changes a PCR, so `caution verify` cannot speak to either; both are -visible in the published `caution.hcl` to a reader who is told to look. - -## Validation Information - -**Validated 2026-08-18. CONFIRMED at Medium.** All four omissions were checked in -the target at HEAD and against a **fresh** clone of the Caution platform -(`https://codeberg.org/caution/platform`, HEAD `6051734a`, 2026-08-18), the -`bootproof` repository, and `aws-nitro-enclaves-image-format 0.4.0` — not taken -on trust from the filing global audit. - -### The recipe, and what the tree actually contains - -`README.md:71` is verbatim as quoted, and it is the only verification instruction -in any user-facing document. - -The two decisive greps over the whole `audit-target/zeronym` tree: - -- **`run_command` and `.manifest`: zero occurrences.** No document, script, test - or comment anywhere in zeronym mentions the field that carries the enclave's - entire environment, even though `bootproofd` returns it in the body of every - `/attestation` response (`crates/bootproofd/src/routes/nonced_attestation.rs:20-29`, - `:82-85`, `:120` — `NoncedAttestationResponse { document, manifest }`). **Nobody - is told the value is there, let alone told to read it.** -- **`certfp`: zero occurrences.** The string `TLS certificate binding` appears - twice, both as *expected output* in a runbook narrative - (`shim/deploy/caution/OPERATORS.md:192-193`, `hub/deploy/caution/OPERATORS.md:108`), - never as a criterion to require and never with a warning that a `skipped` or - `not performed` line invalidates the run. `Certificate Transparency` / `crt.sh` - appears **seven** times, including in the `README.md:71` recipe itself. - -### The measurement chain, re-derived - -`ZIS_HUB_NYM` really is inside the attestation: `unit "default" { env = { … } }` -(`shim/deploy/caution/caution.hcl.tmpl:129-186`, filled by -`shim/deploy/caution/assemble-caution.sh:345-355`) → -`UnitConfig::run_command_string()` emits `export ZIS_HUB_NYM='…'` -(`src/caution-config/src/lib.rs:253-274`) → `{{USER_CMD}}` in the generated -`run.sh` (`src/enclave-builder/src/build.rs:468`, -`templates/run.sh.template:180`) → `/build/initramfs/run.sh` → the EIF's **single** -`--ramdisk` (`templates/Containerfile.eif:297-302`) → PCR0/PCR1 -(`aws-nitro-enclaves-image-format-0.4.0/src/utils/mod.rs:660-691`, `:176-178`). -So the value is measured **and** disclosed, and the recipe still contains no step -that reads it. That is precisely open item 7a's sentence — *measurement discloses -a value; it never detects a change* — and this is the issue that owns the -"nobody is told to look" half of it. - -The certfp binding also exists exactly as described: -`templates/caddy-certfp.sh` publishes `{"tls":{"mode","domain","certfp"}}` to -`/metadata.json`; `bootproofd` passes it as the COSE-signed `user_data`; and -`validate_attested_tls()` (`src/cli/src/lib.rs:354-385`) requires -`mode == "tls"`, `domain ==` the domain from the *reproduced* `caution.hcl`, a -64-char lowercase-hex `certfp`, and `certfp == sha256(leaf DER)` of the same -redirect-disabled, WebPKI-validated response. The recipe names crt.sh instead. - -### The attack scenario holds under the 6q reversal - -Checked specifically, because 6q reversed two neighbouring findings: -`deploy.sh:206-219` publishes the **deployed** tree as `--app-source`, and -`caution verify` reproduces PCRs from that tree -(`src/cli/src/lib.rs:7235-7305`), so an operator who repoints or appends -`ZIS_HUB_NYM` produces an enclave that verifies against itself and prints -`✓ Attestation verification PASSED`. Every one of the four documented steps -passes on a shim pointed at a hub the operator runs. The evidence is one field -away in the response body of a request the auditor already made. - -### Corrections made during validation - -1. **The `SkippedNoDns` claim was narrowed** (see the Description). It is - reachable only on the raw-IP flow, not on the `https://` flow both - runbooks prescribe, so the NXDOMAIN-window remark does not apply to the - documented path. The residual — `PASSED` printed alongside a skip warning that - no zeronym document calls disqualifying — survives. -2. **Omission 4 was softened from "the recipe omits it" to "the recipe omits it - and nothing requires it".** Both runbooks *do* mention the binding, as - expected output. What is missing is the requirement and the failure criterion. - -### Severity: Medium — and this is the explicit ownership decision the coordinator asked for - -One harm — *an operator can point users' plaintext somewhere else and every -documented check still passes* — is spread across four files. The allocation set -by the earlier validation of -`operators-runbook-attributes-the-hub-destination-to-the-binary-hash-and-egress-rules-neither-of-which-binds-it.md` -is retained deliberately, and **this issue absorbs the "no document tells anyone -to perform the check" harm at Medium while the siblings stay where they are**: - -| Issue | Owns | Severity | -|---|---|---| -| `shim-submits-every-migration-to-every-configured-hub-…-appends-their-own-…` | the **capability** (fan-out to every address, ack unread) | High (confirmed) | -| `caution-verify-reproduces-from-a-repository-the-operator-nominates-…` | the **capability** (nothing binds the reproduced tree to zeronym) | High (confirmed) | -| **this issue** | the **procedure gap**: `README.md:71` never reads the configuration, and names CT instead of the binding | **Medium** | -| `operators-runbook-attributes-the-hub-destination-…` | the **false assurance** ("an operator cannot silently repoint it") | Low (confirmed) | -| `attested-tls-binding-is-verified-once-by-hand-if-ever-…` | the **time-of-check-only** binding | Medium (confirmed) | - -The reverse allocation (this issue Low, the runbook issue Medium) was considered -and rejected: `README.md:71` is the *only* verification instruction any -user-facing document gives, it is what the product offers in place of trusting -the operator, and the omission it contains is the enabling condition for **both** -confirmed Highs — whereas the runbook sentence is a false assurance in a document -only operators read. Medium here and Low there is the right way round. - -**Not High.** It is not itself an attack; the exploitable capabilities are filed -and graded above, and grading this High would count the same harm twice. **Not -Low.** The omitted check costs one HTTP request, the omitted criterion is one -output line, and per the G10/G12/G13 sweep this is the cheapest high-value fix in -the entire attestation area. - -### What is *not* counted in this severity - -- **Omission 2** (no PCR criterion) — the wrong-criterion half is owned by - `deploy-script-tells-operators-to-expect-pcr01-failure-and-accept-pcr2-alone.md` - (confirmed Medium). This issue keeps only the narrower point that the recipe - names no criterion at all, which is what leaves the door open for it. -- **Omission 3** (the hash joined to nothing) — owned by - `reproduce-never-builds-the-runtime-stage-that-the-enclave-and-pcr0-are-built-from.md` - and `hub-caution-readme-says-the-attestation-binds-the-running-binary-to-expected-sha256.md`. -- **The addendum's "fifth omission"** (`caution verify` reproduces from an - operator-nominated repository) — since filed and **confirmed High** as - `caution-verify-reproduces-from-a-repository-the-operator-nominates-so-nothing-binds-the-attested-code-to-zeronym.md`. - It remains here only as a step the rewritten recipe must gain. -- **The addendum's `debug.ssh_keys` and `network.ingress` points** — owned by - `hub-manifest-debug-block-claims-ssh-keys-render-empty-and-ssh-is-closed-under-attestation.md` - and `attested-enclave-console-is-reopenable-from-the-parent-because-debug-mode-is-a-launch-flag-and-ssh-keys-is-not-gated-on-it.md`. - -So this issue's Medium rests on **omissions 1 and 4 only** — the two that are -unique to `README.md:71` and owned nowhere else. - -### One constraint on Recommendation 1 that the report must carry - -The check this issue recommends — read `ZIS_HUB_NYM` out of -`.manifest.run_command` — terminates in an **unanchored value**. A Nym address is -not a name that gets authenticated; it *is* the recipient's key material, so an -auditor who performs the check obtains a base58 string with nothing to compare it -against (open item 6z, and -`hub-nym-identity-has-no-trust-anchor-and-the-one-the-project-already-owns-is-never-applied-to-it.md`). -Recommendation 2 (publish the canonical value) is therefore **not optional and -must ship first or together** — and the comparison must be **exact whole-list -equality**, not membership, because `NymHandle::submit` sends to every address in -the list; and a **missing** `ZIS_HUB_NYM` line in `.manifest.run_command` must be -treated as a failure, not a pass, since `run_command_string()` silently skips -`env::vault(...)` entries (`src/caution-config/src/lib.rs:253-263`). - -DO NOT DEVIATE FROM THIS TEMPLATE (`docs/SECURITY-ISSUE-TEMPLATE.md`). -LEAVE THIS MESSAGE IN PLACE SO THAT ALL AGENTS KNOW NOT TO DEVIATE FROM THIS TEMPLATE (`docs/SECURITY-ISSUE-TEMPLATE.md`). diff --git a/zeronym-22aa9851caf68-high-medium/medium/classifier-unknown-consensus-branch-id-diverts-all-shielded-traffic.md b/zeronym-22aa9851caf68-high-medium/medium/classifier-unknown-consensus-branch-id-diverts-all-shielded-traffic.md deleted file mode 100644 index 12ea46e2..00000000 --- a/zeronym-22aa9851caf68-high-medium/medium/classifier-unknown-consensus-branch-id-diverts-all-shielded-traffic.md +++ /dev/null @@ -1,419 +0,0 @@ -# The next Zcash network upgrade silently turns every un-redeployed shim into a divert-everything box: an unrecognised `nConsensusBranchId` makes the classifier fail-safe on 100% of v5/v6 traffic, and no health signal, alert field or document says so - -**Severity**: Medium -**Validation Status**: Confirmed -**Location**: `audit-target/zeronym/shim/src/classify.rs:235-241` (the parse-failure arm), `:99-101` (`Class::treat_as_migration`), `:27-31` (the Ironwood pass-through boundary the failure erases), `:44` (the "a false positive is merely a wasted diversion" cost claim); `audit-target/zeronym/shim/src/intercept.rs:117-119`, `:137-215` (`divert`), `:616-624` (the only signal emitted); `audit-target/zeronym/shim/src/nym.rs:1363-1378` (`local_txid`), `audit-target/zeronym/shim/src/hub.rs:238-239`; `audit-target/zeronym/hub/src/queue.rs:189-197`, `:328-348` (`find_by_txid`); the branch-ID table in `audit-context/zero/zebra/zebra-chain/src/parameters/network_upgrade.rs:230-249` and its readers at `audit-context/zero/zebra/zebra-chain/src/transaction/serialize.rs:1086-1087` (v5) and `:1150-1151` (v6) -**Found by agent:** Local (file audit of `shim/src/classify.rs`); validated 2026-08-18 -**In scope of audit?** Yes - -## Description - -The shim decides where a `SendTransaction` goes by parsing the transaction with the -vendored `zebra-chain`. Bytes that do not parse are `Class::Unparseable`, and -`Class::treat_as_migration()` folds that into "divert to the hub" -(`classify.rs:99-101`). That is the right direction for privacy, and the module -argues it is cheap: *"A false positive is merely a wasted diversion."* -(`classify.rs:44`). - -`zebra-chain` rejects a v5 or v6 transaction whose `nConsensusBranchId` is not in a -**compile-time table**. `NetworkUpgrade::try_from` is a linear search of -`CONSENSUS_BRANCH_IDS` returning `InvalidConsensusBranchId` for anything absent -(`network_upgrade.rs:75-85`), which `SerializationError::from` maps to -`Parse("invalid consensus branch id")` (`serialization/error.rs:98`). Both the v5 and -the v6 deserializer call it as the **second** field they read, before any bundle: - -```rust -// zebra-chain/src/transaction/serialize.rs:1150-1151 (v6; the v5 arm at :1086-1087 is identical) - let network_upgrade = - NetworkUpgrade::try_from(limited_reader.read_u32::()?)?; -``` - -The table the production build compiles ends at NU6.3. NU7 exists in the enum and in -the table only as a **test placeholder behind a cfg gate**, with the real branch ID -still a TODO: - -```rust -// zebra-chain/src/parameters/network_upgrade.rs:241-244 - (Nu6_3, ConsensusBranchId(0x37a5165b)), - // TODO: set below to (Nu7, ConsensusBranchId(0x77190ad8)), once the same value is set in librustzcash - #[cfg(any(test, feature = "zebra-test"))] - (Nu7, ConsensusBranchId(0xfffffffe)), -``` - -`zebra-chain` is a plain (featureless) dependency of the shim -(`shim/Cargo.toml`, `zebra-chain = { path = "../../zebra/zebra-chain" }`), so the -enclave binary carries no NU7 row at all. - -**Consequence.** A Zcash transaction is only valid in a block whose network upgrade -matches its `nConsensusBranchId`, so from the moment the next upgrade activates every -wallet builds transactions carrying the new ID. On any shim whose image predates that -upgrade: - -* **every v5 and every v6 transaction fails to parse** and is diverted — not only - Orchard-touching ones. Since NU5 the v5 format is what ordinary wallets emit for - *everything*, so `Class::PassThrough` becomes an empty class in practice: Ironwood - payments, Sapling payments and transparent payments are all diverted; -* the diverted transaction is **still published**, by the hub, at the next flush - (`hub/src/queue.rs:189-197` folds an unparseable payload to `expiry = None`, which - always survives admission, and `batcher::flush` broadcasts the raw bytes), so this - is a correctness-and-availability failure rather than a privacy leak; -* but the wallet is answered with an **empty txid**, every `GetTransaction` for a - still-queued transaction becomes unanswerable, everything over 65,503 bytes is - **permanently refused**, and every existing silent-loss path in the divert - pipeline widens from "migrations" to "all of this shim's send traffic". - -**Nothing detects it.** `/healthz` reports only mixnet liveness, `/nym-status`'s four -alert fields (`shim/deploy/caution/OPERATORS.md:413-418`) are all unchanged, the only -signal is a per-transaction `warn!` that `shim/src/proxy.rs:111` says nobody can read -(*"an attested enclave has no console"*), and a grep for `"network upgrade"` across -the whole of `audit-target/zeronym/` returns **zero** hits — no README, runbook or -`RESTARTS.md` mentions that the image embeds a consensus-branch table with an -expiry date. - -## Attack Scenario and Steps - -**No attacker is required.** The trigger is a scheduled protocol event, and the -distribution's own history shows this class of staleness has already shipped once. - -1. An operator deploys the attested shim image. Attestation pins it, a redeploy moves - PCR2 and spends a certificate issuance (`shim/deploy/caution/RESTARTS.md`; - confirmed `restarts-ledger-budget-model-omits-hub-forced-redeploys-…`), so images - are long-lived by design. -2. Zcash activates its next network upgrade. The mainnet cadence in the vendored - table is *accelerating* — NU6 → NU6.1 = 420,000 blocks (~1 year), NU6.1 → NU6.2 = - 218,200 (~189 days), NU6.2 → NU6.3 = 63,543 (~55 days) - (`zebra-chain/src/parameters/constants.rs:90-99`) — and NU7 is already carried in - the enum with a reserved branch ID awaiting only librustzcash. The next activation - is a matter of weeks-to-months, not years. -3. A wallet sends an ordinary Ironwood payment. `intercept::inspect` decodes the gRPC - frame and the `RawTransaction` fine and calls `classify_with_evidence(&raw.data)` - (`intercept.rs:563`). -4. `Transaction::zcash_deserialize` hits the unknown branch ID and returns - `Parse("invalid consensus branch id")`; `classify.rs:238-241` returns - `Evidence::unparseable`. -5. `Inspection::treat_as_migration()` is `true` (`intercept.rs:481-489`), so - `divert` runs (`intercept.rs:117-119`). The operator's indexer is never dialled. - Correct for a migration; wrong for this transaction. -6. On the mixnet transport `NymHandle::submit` is dispatch-only (`nym.rs:595-690`; - `let (ack_tx, _drop_receiver) = oneshot::channel();` at `:652`), so the wallet is - answered `SendResponse { error_code: 0, error_message: local_txid(bytes) }` — and - `local_txid` returns `String::new()` for bytes it cannot parse - (`nym.rs:1370-1377`, via `hub.rs:238-239`). **Every send returns success with an - empty txid.** -7. The transaction sits in the hub's queue for up to one flush interval (20 blocks, - ~25 minutes) and is then published. During that window the hub cannot answer a - lookup for it: `admit` stored `txid: None`, and `find_by_txid` only matches - entries whose txid is `Some` (`queue.rs:189-197`, `:328-348`). Every - `GetTransaction` — and with a hub configured, *all* of them go to the hub - (`intercept.rs:229-236`) — returns `NOT_FOUND` for the user's own just-sent - transaction. - -**Attack Requirements and Assumptions:** -- Requires only that an operator is running an image whose vendored `zebra-chain` - predates the active network upgrade. That is a maintenance lapse that nothing in - the product warns about, measures, or alerts on. -- Not reachable today: mainnet is on NU6.3 (activated at height 3,428,143, tip - ~3,451,298 on 2026-08-17), which the vendored table does contain — the `37 a5 16 5b` - little-endian bytes at offset 8 of every committed V6 fixture. -- **No test can catch it, and a test would give the wrong answer.** `proptest-impl` - enables `zebra-test` (`zebra-chain/Cargo.toml:47`), and `proptest-impl` is a - dev-dependency of the shim, so under feature unification a `cargo test` build has - the `(Nu7, 0xfffffffe)` row the release build does not. The test build and the - shipped build therefore have *different* consensus-branch tables. (Filed separately - as `shim-tests-parse-with-a-zebra-chain-whose-consensus-branch-table-differs-from-the-shipped-one.md`.) -- **The distribution has already shipped this exact class of staleness.** Its own - changelog records it: *"The z3 smoke probes assert that the deployed zebrad reports - the NU6.3 branch ID and pins activation at 3,428,143. A stale image passed every - other probe while silently lacking both the consensus rules and the grace window - above; the 2026-07-17 cached-layer incident shipped that exact class of mismatch."* - (`audit-context/zero/CHANGELOG.md:138-142`). A probe was added for **zebrad**. - No equivalent probe exists for the shim, whose branch table is just as pinned. - -## Impact on Users - -Every wallet behind an affected shim, for every transaction it sends — not only -migrations: - -- **Up to ~25 minutes of added latency on all traffic**, including the Ironwood - payments the design deliberately exempted *because* they are time-sensitive - commerce (`classify.rs:27-31`: *"A transaction with only Ironwood actions must - still pass through"*). That boundary becomes unreachable, because an Ironwood-only - transaction never gets as far as `is_orchard_touching`. -- **An empty txid in every `SendResponse`**, where lightwalletd and the shim's own - non-upgraded behaviour return the real one. -- **`GetTransaction` blind for the whole flush window.** The hub's queue lookup keys - on a txid it could not compute, so the one feature that lets a wallet see its - diverted transaction before publication stops working for everything. -- **Hard, non-retryable refusal above 65,503 bytes.** `wire::encode_submit` refuses - (`shim/src/wire.rs:124`, `:274-279`) and `divert` answers `RESOURCE_EXHAUSTED` - (`intercept.rs:167-179`), never forwarding and never broadcasting another way. - Pre-upgrade such a transaction would have passed through untouched. -- **Hub unavailability becomes a total send outage.** Today an unreachable hub fails - closed only for Orchard-touching sends (`intercept.rs:208-214`); afterwards it - fails closed for every send the shim sees. -- **Every silent-loss path in the divert pipeline widens to all traffic.** The - wallet is told `error_code: 0` at hand-off, so the confirmed - `shim-nym-driver-every-teardown-path-silently-destroys-acknowledged-submits.md` - (Medium), `nym-rotation-deferral-cannot-protect-submits-so-rotating-destroys-acknowledged-migrations.md` - (Low) and `junk-sendtransaction-flood-…` (High) stop being scoped to a small, - self-selected population that knows migrations are slow, and start applying to - ordinary payments. The shim's whole mixnet egress is ~45 Sphinx packets ≈ 5.4 s per - submit at the crate's own throttled rate (`shim/src/nym.rs:1085-1115`), which the - design sized against ~0.77 Orchard-touching transactions per block, not against a - shim's entire send volume. - -Privacy is **not** harmed: the failure direction is toward diversion, and nothing -reaches the operator's indexer that would not have before. This is an availability -and integrity finding. - -## Technical Details / Code Analysis - -**The parse-failure arm** (`shim/src/classify.rs:235-241`) — the only place a -protocol-version failure and random junk are distinguished, and they are not: - -```rust -pub fn classify_with_evidence(raw: &[u8]) -> Evidence { - let mut cursor = Cursor::new(raw); - - let tx = match Transaction::zcash_deserialize(&mut cursor) { - Ok(tx) => tx, - Err(err) => return Evidence::unparseable(raw.len(), err.to_string()), - }; -``` - -**The routing policy** (`shim/src/classify.rs:99-101`): - -```rust - pub fn treat_as_migration(self) -> bool { - matches!(self, Class::Migration | Class::Unparseable) - } -``` - -**The boundary this erases** (`shim/src/classify.rs:27-31`): - -```rust -//! ORCHARD ONLY, NOT IRONWOOD, and that is deliberate. Ironwood is the NEW pool, -//! where ordinary time-sensitive commerce will live, and the time-insensitivity -//! half of the rationale does not hold for it. A transaction with only Ironwood -//! actions must still pass through, so there is no Ironwood arm in this -//! predicate and none should be added. -``` - -**The only signal emitted** (`shim/src/intercept.rs:616-624`). The reason string is -there — `error = "parse error: invalid consensus branch id"` — but it is the same -`warn!` line random junk produces, there is no counter, and in an attested enclave -there is no console to read it from: - -```rust - Class::Unparseable => tracing::warn!( - target: "zis::classify", - error = evidence.error.as_deref().unwrap_or("(none)"), - tx_len = evidence.len, - frame_len = frame.len(), - body_prefix = %hex_prefix(frame, GRPC_PREFIX_LEN + PREFIX_LOG_BYTES), - diverted_in_production, - "MIGRATION-FAILSAFE: unparseable SendTransaction body, treating as migration" - ), -``` - -**The empty txid** (`shim/src/nym.rs:1370-1377` and `shim/src/hub.rs:238-239`): - -```rust -pub fn local_txid(tx_bytes: &[u8]) -> String { - use zebra_chain::serialization::ZcashDeserialize; - match zebra_chain::transaction::Transaction::zcash_deserialize(&mut std::io::Cursor::new( - tx_bytes, - )) { - Ok(tx) => tx.hash().to_string(), - Err(_) => String::new(), - } -} -``` - -```rust - Ok(()) => Ok(Submit::Accepted { - txid: crate::nym::local_txid(tx_bytes), - }), -``` - -**The hub side, which is why this is not a destruction finding.** `admit`'s parse is -telemetry, and a failure folds expiry to `None` (`hub/src/queue.rs:189-197`), which -`survives_next_flush` always admits (`:380-392`). So the `ExpiryTooTight` refusal -**cannot** fire for this class, and the transaction is queued and published: - -```rust - let (txid, expiry) = match Transaction::zcash_deserialize(&mut Cursor::new(tx_bytes)) { - Ok(tx) => ( - Some(tx.hash().to_string()), - tx.expiry_height() - .map(|h| h.0) - .filter(|height| *height != 0), - ), - Err(_) => (None, None), - }; -``` - -The same `txid: None` is what blinds `find_by_txid` (`hub/src/queue.rs:337-345`), -because the match arm requires a txid to exist: - -```rust - .find(|entry| { - entry - .txid - .as_deref() - .is_some_and(|txid| txid == forward || txid == reversed) - }) -``` - -**Why the rest of the version handling is sound** (recorded so this is not mistaken -for a wider claim): `orchard_shielded_data()` is version-agnostic and returns `None` -for V1–V4 (`zebra-chain/src/transaction.rs:1065-1083`); `orchard_actions()` derives -from it (`:1085-1091`); an unknown transaction *version* falls to -`(_, _) => Err(SerializationError::Parse("bad tx header"))` (`serialize.rs:1220`) and -therefore also diverts; a future format appending fields to v6 would leave -`cursor.position() != raw.len()` and be caught at `classify.rs:248-257`. Every one of -those is fail-safe. Note also that a **v4** transaction is unaffected: v4 carries no -`nConsensusBranchId` on the wire, so Sapling-only v4 traffic keeps parsing. The -finding is that the file's stated cost model for the fail-safe direction is wrong, -and that one certain, scheduled, attacker-free event routes essentially the entire -traffic mix into it. - -## Recommendations - -1. **Give the shim a branch-ID health signal.** Distinguish "unrecognised protocol - version" from "unparseable garbage" in `classify.rs` (the reason string is already - in `Evidence.error`), count it, and add it to `/nym-status` and the - `OPERATORS.md:413-418` alert table. If the unrecognised-branch rate is non-trivial - over a window, fail `/healthz`. The zebrad half of this distribution already has - exactly this probe (`CHANGELOG.md:138-142`); the shim needs the equivalent. -2. **Pin the newest known branch ID as build metadata** and publish it beside - `deploy/EXPECTED_SHA256`, so an operator can compare it against the chain's active - branch without reading source. -3. **Document the coupling.** `shim/deploy/caution/OPERATORS.md` and - `RESTARTS.md` must state that the enclave image embeds a consensus-branch table - and must be redeployed before the next network upgrade activates. Today the string - "network upgrade" appears nowhere in `audit-target/zeronym/`. -4. **Correct `classify.rs:44`.** A false positive is not "merely a wasted diversion": - it costs up to a flush interval of latency, returns an empty txid, blinds - `GetTransaction` for the window, is a hard `RESOURCE_EXHAUSTED` above 65,503 bytes, - and is silently destroyed by every driver teardown because submit is - dispatch-only. Stating the real cost is what lets a future maintainer weigh - widening the fail-safe correctly. -5. **Test the state, not the placeholder.** A test that builds a transaction with an - unknown branch ID must exercise the *release* table; today `proptest-impl` pulls in - `zebra-test` and gives the test build an NU7 row production does not have. - -## Validation Information - -**Verdict: CONFIRMED. Severity raised from Low to Medium.** - -Every mechanical claim was checked against primary sources. - -**Confirmed as filed:** - -1. The v5 and v6 deserializers call `NetworkUpgrade::try_from` on the branch ID - before reading any bundle (`serialize.rs:1086-1087`, `:1150-1151`), and - `try_from` is a lookup in the const `CONSENSUS_BRANCH_IDS` table returning - `InvalidConsensusBranchId` on a miss (`network_upgrade.rs:75-85`). The error - surfaces as `SerializationError::Parse("invalid consensus branch id")` - (`serialization/error.rs:98`). -2. The production build has no NU7 row. The only `Nu7` entry is - `#[cfg(any(test, feature = "zebra-test"))] (Nu7, ConsensusBranchId(0xfffffffe))` - — a placeholder, not the real ID, which is still a TODO - (`network_upgrade.rs:241-244`). Cross-checked against the distribution's - librustzcash, where `BranchId::Nu7 => 0xffff_ffff` is likewise a placeholder - (`librustzcash/components/zcash_protocol/src/consensus.rs:751`, `:772`). Grepping - the whole distribution for the real value `0x77190ad8` returns exactly one hit: - the TODO comment. -3. `Unparseable` routes to divert (`classify.rs:99-101`), `divert` never falls back - to the operator (`intercept.rs:137-215`), and the >65,503-byte case is a - non-retryable `RESOURCE_EXHAUSTED` (`wire.rs:124` = `65536 - 33`; - `intercept.rs:167-179`). -4. The documentation gap is real. `grep -rni "network upgrade"` over - `audit-target/zeronym/` returns **no hits**; `grep -i upgrade` over - `shim/deploy/caution/OPERATORS.md`, `RESTARTS.md` and `shim/README.md` returns no - hits. Every "NU6.3" mention in the tree is about pool semantics, never about the - table going stale. - -**Corrected — the filed issue's step 9 was wrong, and its arithmetic was inverted:** - -The filed version claimed the state causes **silent destruction** via -`Refusal::ExpiryTooTight` for wallets with tight expiries. It does not, and cannot. -In this state the **hub cannot parse the transaction either** (same vendored -`zebra-chain`, same missing row), so `admit` folds expiry to `None` -(`queue.rs:189-197`) and `survives_next_flush(None, …)` returns `true` -unconditionally (`queue.rs:380-392`). `ExpiryTooTight` is unreachable for exactly -this class. The transaction is queued and published. - -The filed text also inverted the expiry arithmetic it used to support that claim: -with a 20-block expiry delta the admission test `expiry >= next_flush_height(tip,20)+4` -fails when `tip mod 20 < 4` — i.e. when the tip is **fewer** than four blocks past a -flush boundary (20% of heights), not "more than four blocks past" as filed. That -branch is out of scope here in any case; it belongs to the general expiry story owned -by `a-constant-tip-offset-is-a-tunable-expiry-keyed-admission-filter-…` (Medium). - -**Corrected — scope was both too narrow and too broad in places:** - -- "All *shielded* traffic" understates it. The branch ID sits in the v5/v6 header, so - **transparent-only v5 transactions are diverted too**. Post-NU5 wallets emit v5 for - everything, so `Class::PassThrough` becomes an empty class in practice. -- "Every Sapling-only transaction" needs a qualifier: a **v4** Sapling transaction - still parses, because v4 carries no `nConsensusBranchId` on the wire. Only v5/v6 - Sapling traffic is affected. - -**Added — four consequences the filed issue missed, all verified:** - -- **Empty txid to every wallet.** `local_txid` returns `String::new()` on a parse - failure (`nym.rs:1370-1377`) and `HubTransport::submit` puts it straight into - `Submit::Accepted` (`hub.rs:238-239`), so every `SendResponse` carries - `error_code: 0` with an empty `error_message`. -- **The hub's queue lookup goes blind.** `admit` stores `txid: None`, and - `find_by_txid` requires `Some` (`queue.rs:328-348`), so a wallet polling for its own - just-sent transaction gets `NOT_FOUND` for the entire flush window — the one thing - the queue-first lookup exists to prevent. -- **Hub downtime becomes a total send outage** rather than a migration-only one. -- **Every existing silent-loss path widens from migrations to all send traffic**, - because the wallet is acknowledged at mixnet hand-off. - -**Added — the loud-or-silent question, answered:** silent. `/healthz` tracks only -`MixnetStatus::is_healthy`; `/nym-status`'s four documented alert fields -(`OPERATORS.md:413-418`: `diversion_configured`, `mixnet_connected`, `client_deaths`, -`consecutive_rebuild_failures`) are all unchanged and all green; the per-transaction -`warn!` is the only trace, and `shim/src/proxy.rs:111` states outright that *"an -attested enclave has no console"*. In the shipped `DEBUG=1` default the console does -land in a file on the parent host (established in -`attested-enclave-console-is-reopenable-from-the-parent-…`), but nothing alerts on it -and no runbook tells anyone to look. - -**Added — reachability, which is the reason for the re-grade.** This needs no -attacker at all; it is produced by an ordinary scheduled network upgrade. The -mainnet cadence in the vendored constants is accelerating (420,000 → 218,200 → -63,543 blocks between the last three activations, -`zebra-chain/src/parameters/constants.rs:90-99`, i.e. ~1 year → ~189 days → ~55 -days), NU7 is already carried in the enum awaiting only its branch ID, and the -distribution's own changelog records that a stale image lacking the *current* branch -rules already shipped once and was caught only after a probe was added for zebrad -(`CHANGELOG.md:138-142`). No such probe exists for the shim. - -**Severity: Medium, raised from Low.** Likelihood is effectively certain over the -lifetime of a long-lived attested image, and the failure is invisible to every signal -the product exposes. Impact stops short of High because the transactions are still -published — the direction of the failure is safe, and nothing leaks — so what users -actually suffer is universal added latency, an empty txid, a blind lookup window, a -hard refusal above 64 KiB, and a large widening of the population exposed to the -already-confirmed silent-loss paths. That places it alongside -`publish-verdict-strings-are-zcashds-vocabulary-only-…` (Medium): an environmental -change, not an attacker, converts a working component into a broken one, with no -signal. - -**What this issue uniquely owns:** the coupling between the enclave image's frozen -consensus-branch table and a scheduled network upgrade, and the resulting -divert-everything state with its four downstream consequences. It does **not** own: -the release-profile question (`shim-manifest-pins-no-release-profile-…`), the -test/production table divergence -(`shim-tests-parse-with-a-zebra-chain-whose-consensus-branch-table-differs-from-the-shipped-one.md`), -dispatch-only acknowledgement (`nym-submit-acks-are-never-read-…`, Medium), or any of -the teardown/flood loss paths it composes with. - -DO NOT DEVIATE FROM THIS TEMPLATE (`docs/SECURITY-ISSUE-TEMPLATE.md`). -LEAVE THIS MESSAGE IN PLACE SO THAT ALL AGENTS KNOW NOT TO DEVIATE FROM THIS TEMPLATE (`docs/SECURITY-ISSUE-TEMPLATE.md`). diff --git a/zeronym-22aa9851caf68-high-medium/medium/deploy-script-tells-operators-to-expect-pcr01-failure-and-accept-pcr2-alone.md b/zeronym-22aa9851caf68-high-medium/medium/deploy-script-tells-operators-to-expect-pcr01-failure-and-accept-pcr2-alone.md deleted file mode 100644 index 02ac9ba8..00000000 --- a/zeronym-22aa9851caf68-high-medium/medium/deploy-script-tells-operators-to-expect-pcr01-failure-and-accept-pcr2-alone.md +++ /dev/null @@ -1,492 +0,0 @@ -# `deploy.sh` tells the operator to expect a PCR0/PCR1 verification failure and to accept a PCR2 match alone — advice both `OPERATORS.md` files now say is wrong, because PCR2 is byte-identical across two entirely different binaries - -**Severity**: Medium -**Validation Status**: Confirmed -**Location**: `audit-target/zeronym/deploy.sh:220` (the only verification guidance the deploy script emits); the same claim repeated at `audit-target/zeronym/shim/deploy/caution/README.md:131-136`; contradicted by `audit-target/zeronym/shim/deploy/caution/OPERATORS.md:189-209` and `audit-target/zeronym/hub/deploy/caution/OPERATORS.md:107-124`; the user-facing promise at `audit-target/zeronym/README.md:71` -**Found by agent:** Local (file audit of `deploy.sh`) -**In scope of audit?** Yes — `deploy.sh` and `*/deploy/**` are in scope, and markdown/operator claims are in scope "as security claims": under ICTM a documented property users are told they get but do not get is itself a bug. - -## Description - -`deploy.sh` never verifies anything about the artifact it deploys — it does not -run `build.sh`, does not run `reproduce.sh`, does not compare `EXPECTED_SHA256` -to anything, and does not run `caution verify`. Its **entire** contribution to -verification is one log line, printed at the end of every attested deploy -(`deploy.sh:220`): - -```sh -log "verify with: caution verify (expect PCR0/1 FAILED on Caution's floating framework; PCR2 is the check that matters)" -``` - -That single line makes two assertions, and the project's own operator runbooks — -which were updated after measurements taken on 2026-08-14 — say both are wrong: - -1. **"expect PCR0/1 FAILED"** — `shim/deploy/caution/OPERATORS.md:189-193`: - - > Older copies of this guide warned that verify reports PCR0/1 FAILED on a - > healthy enclave, because Caution's builder fetched its framework from a - > floating `main.tar.gz`. **That is fixed**: on the attested pair deployed - > 2026-08-14 the manifest pinned both the enclave and framework sources to - > commits, and **all three PCRs reproduced** on both … Expect a clean - > `✓ Attestation verification PASSED`. - -2. **"PCR2 is the check that matters"** — `shim/deploy/caution/OPERATORS.md:195-209`: - - > **Do not fall back to "PCR2 is the one that matters".** That advice - > circulated while PCR0/1 were failing, and it is wrong on this platform. - > Measured 2026-08-14 across the attested shim and hub — two entirely - > different binaries: - > - > | | shim | hub | - > |---|---|---| - > | PCR0 | `accb679a…` | `218d1f64…` | - > | PCR1 | `accb679a…` | `218d1f64…` | - > | PCR2 | `21b9efbc…` | `21b9efbc…` **(identical)** | - > - > **PCR2 does not distinguish the application.** PCR0/PCR1 are what change - > with it. So an attestation accepted on a PCR2 match alone would prove only - > that *some* Caution enclave is running, not that it is running your - > reviewed code — which is the entire claim. Require **all three** to - > reproduce, and treat a PCR0/1 mismatch as a real finding about the - > application until proven otherwise. - -`hub/deploy/caution/OPERATORS.md:112-124` says the same thing in the same words. - -So the deploy script trains the operator, at the exact moment they are about to -verify, to (a) **pre-accept the failure of the only two measurements that -distinguish one application from another**, and (b) **treat as sufficient the one -measurement the project has empirically shown to be identical between two -entirely different binaries**. This is the classic "train the user to ignore the -alarm" failure, and here the alarm is the whole product's trust root. - -## Attack Scenario and Steps - -1. An operator, or a third party acting as the "Auditor" role `README.md:71` - defines, deploys or inspects an endpoint and runs `caution verify - --attestation-url https:///attestation`. -2. The application differs from the reviewed one, so PCR0/PCR1 differ. `caution - verify` requires **all three** PCRs together — `Nitro::new(attestation_bytes, - expected_nitro_pcrs)` is built from PCR0, PCR1 and PCR2 and `nitro.verify(...)` - is a single all-or-nothing check (Caution platform `src/cli/src/lib.rs:7235-7256`). - It therefore prints `✗ Attestation verification FAILED` and exits non-zero — - and then prints a per-index breakdown (`src/cli/src/lib.rs:7313-7337`): - - ``` - ✗ Attestation verification FAILED - PCR comparison: - PCR0: MISMATCH - PCR1: MISMATCH - PCR2: match - ``` - - That breakdown is exactly the shape the bad advice was written for. -3. The verifier has been told in advance, by the project's own tooling - (`deploy.sh:220`), to **expect** `PCR0/1 FAILED` and that `PCR2 is the check - that matters`. They read `PCR2: match`, override the `FAILED` banner, and - record the endpoint as verified. -4. Nothing about the application was actually checked. Per the project's own - measurement, PCR2 `21b9efbc…` is produced by both the shim and the hub, so a - PCR2 match is consistent with *any* Caution enclave — including a shim built - with diversion disabled, with `RUST_LOG` raised, with the classifier - predicate weakened, or an entirely unrelated binary. - -**Attack Requirements and Assumptions:** - -- No special attacker capability is needed to produce the *misleading guidance*; - it is printed unconditionally on every attested deploy. -- To *exploit* it, a party who can influence what image runs — the operator, or - the platform — deploys code other than the reviewed code. The operator is the - primary adversary in this system's threat model and controls the deploy, so - this is not a hypothetical actor. -- The verifier must be someone who follows the tooling's own printed advice - rather than cross-reading `OPERATORS.md`. That is the realistic case: the - deploy script's line is what appears on the terminal at the moment of use, and - `shim/deploy/caution/README.md:133-135` independently repeats the same wrong - advice, so two of the four in-tree sources agree with the script. -- **Honest counterweight:** an auditor who reads either `OPERATORS.md` gets the - correct instruction, in bold, with the measurement table. The defect is a - contradiction inside the tree, not a uniform overclaim — but the contradiction - is resolved the wrong way by the artifact that speaks last and loudest. - -## Impact on Users - -`README.md:71` sells the following to users as the reason they do not have to -trust their indexer operator: - -> **Auditors** verify an endpoint without trusting its operator: fetch its -> attestation, check the PCRs against the AWS Nitro root, reproduce the build and -> compare hashes, and check Certificate Transparency for a shadow certificate. - -That promise is the substitute for trusting the operator, and it is the only one -a wallet user has. If the PCR check as *instructed by the deploy tooling* cannot -distinguish the reviewed shim from an arbitrary other enclave, then a "verified" -endpoint carries no more assurance than an unverified one, and the user's -protection reduces to trusting the operator after all — which is precisely what -the product exists to remove. - -The failure is silent from the user's side: a wallet sees a valid certificate and -correct gRPC behaviour whether or not the enclave runs the reviewed code. - -This composes with `wallet-facing-tls-identity-is-not-bound-to-the-attestation-and-the-ct-defence-cannot-work.md`: -that issue shows an auditor may be talking to a proxy rather than the enclave; -this one shows that even when they *are* talking to the enclave, the check they -are told to run does not bind the application. - -## Technical Details / Code Analysis - -**The line, in context** (`deploy.sh:206-221`, the attested-deploy publish -block): - -```sh -if [ "$DEBUG" != 1 ] && [ -n "${APP_SOURCE:-}" ]; then - APP_SOURCE_PUSH=${APP_SOURCE_PUSH:-$APP_SOURCE} - APP_SOURCE_TAG=${APP_SOURCE_TAG:-deploy-$APP_ID} - log "publishing the app-source to $APP_SOURCE_PUSH (tag $APP_SOURCE_TAG) ..." - ( cd "$DEST" && - { git remote remove app-source 2>/dev/null || true; } && - git remote add app-source "$APP_SOURCE_PUSH" && - git push app-source HEAD:main && - git tag -f "$APP_SOURCE_TAG" && - git push -f app-source "refs/tags/$APP_SOURCE_TAG" - ) >&2 || die "app-source publish FAILED. …" - log "app-source published: $APP_SOURCE @ $APP_SOURCE_TAG" - log "verify with: caution verify (expect PCR0/1 FAILED on Caution's floating framework; PCR2 is the check that matters)" -fi -``` - -This is the last thing an operator reads before the DONE banner, and it is the -only verification instruction the script gives. - -**The corrected instruction** (`hub/deploy/caution/OPERATORS.md:107-124`): - -``` -Expect `✓ Attestation verification PASSED` with all three PCRs reproducing and -the TLS certificate binding verified. Measured on the first attested hub -(2026-08-14): all of PCR0/1/2 matched … - -**Require all three PCRs. Do not accept a PCR2 match alone.** … measured -2026-08-14, the attested hub and the attested shim — two entirely different -binaries — produced **byte-identical PCR2** (`21b9efbc…`), while PCR0/PCR1 -differed per application (`218d1f64…` hub, `accb679a…` shim). PCR2 does not -distinguish the application, so an attestation accepted on PCR2 alone would prove -only that *some* Caution enclave is running, not that it is running the reviewed -hub — which for the component holding plaintext migrations is the whole point. -``` - -**The full set of in-tree PCR claims, which do not agree with one another** -(established by `SURVEY.md` §9 observation 11 and re-checked here): - -| Source | Claim | -|---|---| -| `deploy.sh:220` | expect PCR0/1 FAILED; **PCR2 is the check that matters** | -| `shim/deploy/caution/README.md:133-135` | **PCR2 (the application layer) is the check that matters** | -| `shim/deploy/caution/OPERATORS.md:189-209` | all three reproduce; **require all three**; PCR2 alone proves nothing | -| `hub/deploy/caution/OPERATORS.md:107-124` | all three reproduce; **require all three**; PCR2 alone proves nothing | -| `OPEN-QUESTIONS.md:86` | PCR0/PCR1 not reproducible; "PCR2 is the measurement that carries weight today" | -| `shim/deploy/README.md:1055-1056` | no PCR0, PCR1 or PCR2 has ever been computed from this image | - -Two of these tell a verifier to accept a result the other two say proves nothing. -`deploy.sh:220` is the one that is executed. - -**Note on what `deploy.sh` does *not* do.** It never binds the reproducible hash -to the attestation either: `EXPECTED_SHA256` is not read, compared, or mentioned -by this script, and per BRAINSTORM §R12-F nothing in the tree binds -`EXPECTED_SHA256` to any PCR — `caution verify` clones `app_sources`, rebuilds, -and compares PCRs without ever consulting it. So the two halves of the chain that -`README.md:71` asks an auditor to join ("check the PCRs … reproduce the build and -compare hashes") are joined nowhere, and the only PCR guidance the deploy path -gives is the incorrect one above. - -## Recommendations - -1. **Replace `deploy.sh:220` with the runbook's own text**, e.g.: - `log "verify with: caution verify --attestation-url https://$TLS_DOMAIN/attestation — expect PASSED with ALL THREE PCRs reproducing. A PCR0/1 mismatch is a real finding; PCR2 alone does not distinguish the application (see OPERATORS.md)."` -2. **Fix the same claim at `shim/deploy/caution/README.md:131-136`**, and repair - its cross-reference: the sentence sends the reader to `OPERATORS.md` "for the - current PCR0/1 caveat", and `OPERATORS.md:189-194` is the document that - *retracts* that caveat. (`OPEN-QUESTIONS.md:86` carries the same superseded - rule and needs the same fix, but it is owned by a separate issue — - `open-questions-pcr-entry-inverts-which-measurement-distinguishes-the-application.md` - — and is deliberately **not** counted in this issue's severity.) -3. Consider having `deploy.sh` *run* `caution verify` on the attested path and - fail the deploy if it does not report `PASSED`, rather than printing advice - about it. A verification step that is only ever described is one that - frequently does not happen. -4. Add a single authoritative statement of the PCR policy in one place and make - every other document reference it, so this class of drift is structurally - prevented — the tree currently carries six mutually inconsistent statements. - -## Addendum (added by the `shim/deploy/caution/README.md` local audit, 2026-08-18) — three facts about the second location; nothing here changes the finding or its severity - -This issue's Location already names `shim/deploy/caution/README.md:133-135`, so -the coordinator's question ("does the filed issue name this file too?") is -answered **yes** and no second issue was filed. Three corrections/additions: - -1. **The exact range is `:131-136`, not `:133-135`.** The sentence begins on line - 131 and the caveat runs to 136. Verbatim: - - ``` - Third parties can run - `caution verify --attestation-url https:///attestation` with no Caution - account and no checkout. See OPERATORS.md for the current PCR0/1 caveat: a - Caution-side unpinned-framework bug makes verify report FAILED on healthy - enclaves, and PCR2 (the application layer) is the check that matters until - their fix lands. - ``` - -2. **The cross-reference is worse than the claim.** *"See OPERATORS.md for the - current PCR0/1 caveat"* sends the reader to - `shim/deploy/caution/OPERATORS.md:189-194`, which does not contain that caveat — - it **retracts** it, in those words: *"Older copies of this guide warned that - verify reports PCR0/1 FAILED on a healthy enclave, because Caution's builder - fetched its framework from a floating `main.tar.gz`. **That is fixed** … - **all three PCRs reproduced** on both … Expect a clean `✓ Attestation - verification PASSED`."* So the README does not merely repeat stale advice; it - attributes that advice to the document that repudiates it. A reader who trusts - the summary is misinformed; a reader who follows the pointer gets the right - answer and a contradiction. That asymmetry is worth one sentence in the report, - because it is the same shape as `OPERATORS.md:33` → this file (recorded in - BRAINSTORM §R18-A): the two documents each point at the other's wrong half. - -3. **This file supplies the *operative* criterion, because the other reference - document supplies none.** The `shim/deploy/README.md` audit established - (`deploy-readme-says-no-enclave-no-eif-and-no-pcr-exist-which-its-own-current-row-refutes.md`, - BRAINSTORM §R31-D) that the reproducible-build reference text asserts no PCR of - any index has ever been computed, and that the strings `caution verify`, - `PCR0`, `PCR1`, `PCR2` appear there **only** inside sentences denying any of it - happened. `shim/deploy/caution/README.md` is therefore the only one of the two - deploy reference documents that gives an auditor a PCR rule at all — and the - rule it gives is the retracted one. Net effect on a document-following auditor: - the reproducible-build half offers no criterion, the attestation half offers - the wrong criterion, and only the two `OPERATORS.md` runbooks (which the - attestation half misquotes) carry the correct one. `deploy.sh:220` then prints - the wrong criterion at the moment of use. - -Everything else in this issue was re-verified against `shim/deploy/caution/README.md` -at HEAD and is accurate as written. - - -## ADDENDUM (added 2026-08-18 by the G10/G12/G13 global audit) — PCR2 on Caution is a universal constant, and here is the derivation. This does not change the finding; it makes it far worse and answers the caveat `OPERATORS.md` itself flags as unresolved. - -`shim/deploy/caution/OPERATORS.md:210-211` ends its measurement table with an -honest caveat: *"(The observation is empirical; we have not confirmed with Caution -which layer each index measures on their EIF layout.)"* That question is now -answered, from Caution's own public source -(`https://codeberg.org/caution/platform`, cloned during this audit) and -`aws-nitro-enclaves-image-format 0.4.0` from crates.io. - -**Caution builds its EIF with exactly one ramdisk.** -`src/enclave-builder/templates/Containerfile.eif`, final stage: - -``` -RUN eif_build \ - --kernel /build/kernel/bzImage \ - --kernel_config /build/kernel/linux.config \ - --ramdisk /build/rootfs.cpio.gz \ - --output /build/enclave.eif \ - --pcrs_output /build/enclave.pcrs \ - --cmdline "reboot=k panic=1 pci=off nomodules console=ttyS0 ... nit.target=/run.sh" -``` - -**`EifBuilder::measure`** (`aws-nitro-enclaves-image-format-0.4.0/src/utils/mod.rs:664-692`): - -```rust -self.image_hasher.write_all(&buffer[..]).unwrap(); // kernel -> PCR0 -self.bootstrap_hasher.write_all(&buffer[..]).unwrap(); // kernel -> PCR1 -self.image_hasher.write_all(&self.cmdline[..]).unwrap(); -self.bootstrap_hasher.write_all(&self.cmdline[..]).unwrap(); -for (index, mut ramdisk) in self.ramdisks.iter().enumerate() { - ... - self.image_hasher.write_all(&buffer[..]).unwrap(); // PCR0: all ramdisks - if index == 0 { self.bootstrap_hasher.write_all(&buffer[..]).unwrap(); } // PCR1: the first - else { self.customer_app_hasher.write_all(&buffer[..]).unwrap(); } // PCR2: the rest -} -``` - -With one ramdisk, three things follow **mathematically**: - -1. **PCR0 and PCR1 are computed over the identical byte stream** — kernel ‖ - cmdline ‖ ramdisk0 — so **PCR0 == PCR1 on every Caution enclave**. That is - exactly what this issue's own quoted table shows (`accb679a…` in both shim - rows, `218d1f64…` in both hub rows). The table is not a transcription error. -2. **`customer_app_hasher` receives nothing at all.** - `defs/eif_hasher.rs` with `new_without_cache` (`block_size == 0`) makes - `finalize_reset()` return `sha384()` and - `tpm_extend_finalize_reset()` return `sha384(0^48 ‖ that)`. Therefore - - **PCR2 = sha384( 48 zero bytes ‖ sha384("") )** - - which computes to **`21b9efbc18480766…`** — the exact value both `OPERATORS.md` - files record as PCR2 for *both* components. (Reproduce in one line: - `python3 -c "import hashlib;print(hashlib.sha384(b'\0'*48+hashlib.sha384(b'').digest()).hexdigest())"`.) -3. So PCR2 is not "a measurement that happens not to distinguish the shim from the - hub". **It is the measurement of an absent application ramdisk: a fixed - constant, identical for every Caution enclave that has ever booted, whatever - code it runs.** - -**What this does to the severity of the advice at `deploy.sh:220`.** The line does -not tell an auditor to rely on a weak check. It tells them to pre-authorise the -failure of the only two measurements that exist, and to accept in their place a -value that is *the same for every Caution deployment on earth*. A verifier who -follows it runs a check that **cannot fail**, for any enclave, ever. The same -applies verbatim to `shim/deploy/caution/README.md:131-136` and to -`OPEN-QUESTIONS.md:86`, which asks security reviewers to ratify it. - -**One correction to the runbooks' otherwise-correct advice, worth a line in the -report.** "Require **all three** to reproduce" is the right operational rule but -implies threefold redundancy that does not exist: there is **one** independent -measurement (PCR0, duplicated as PCR1) plus one constant. The right statement is: -*Caution's attestation carries a single measurement, over the kernel, the kernel -command line, and one ramdisk containing the whole enclave — EnclaveOS `init`, -`bootproofd`, `caddy`, `caddy-certfp.sh`, busybox, socat, the generated `run.sh` -(which carries the entire `unit.env`), `manifest.json`, and the application -filesystem. It appears twice, as PCR0 and PCR1.* - -That last point also settles what the measurement **covers**, which is broader -than any zeronym document says: `run.sh` embeds every `ZIS_*`/`ZIH_*` value from -`caution.hcl`, so the configuration is inside the attestation even though it is -outside `EXPECTED_SHA256`. See `BRAINSTORM.md` §G10-A and §G12-A. - -*(End of addendum.)* - -## Validation Information - -**Validated 2026-08-18. CONFIRMED at Medium.** Every load-bearing fact was -re-derived from primary sources rather than accepted from the file: the target -tree at HEAD, a fresh clone of the Caution platform -(`https://codeberg.org/caution/platform`, HEAD `6051734a`, 2026-08-18, plus the -earlier `1f8d8cb3` clone the global audit used), and -`aws-nitro-enclaves-image-format 0.4.0` from crates.io. - -### The four claims, each checked - -1. **`deploy.sh:220` is verbatim as quoted, and it is the only verification - guidance the script emits.** Confirmed by reading `deploy.sh` in full: the - strings `EXPECTED_SHA256`, `build.sh` and `reproduce.sh` appear nowhere as - invocations, and `caution verify` appears only at `:198` (a comment), `:220` - (this line) and `:332` (the DONE banner, which says only `run: caution - verify` with no criterion). So the script prints advice about verification - and performs none. - -2. **The runbooks say the opposite, in bold, and they are the correct source.** - `shim/deploy/caution/OPERATORS.md:189-193` retracts the PCR0/1-will-fail - caveat ("**That is fixed** … **all three PCRs reproduced** on both … Expect a - clean `✓ Attestation verification PASSED`"); `:195-209` retracts the - PCR2-only rule and carries the measurement table; - `hub/deploy/caution/OPERATORS.md:107-124` states the same in prose. Both - verified in the target at HEAD. - -3. **The "floating framework" premise the line relies on is false in the - platform today.** Caution pins the framework to a commit: - `require_platform_framework_commit()` (`src/api/src/builder.rs:331`) and - `pin_archive_url_to_commit(..., &request.framework_commit, ...)` (`:1030`), - with `framework_commit` recorded in the manifest - (`src/api/src/builder.rs:174`, `main.rs:1959`, `:2702`). The only surviving - `main.tar.gz` reference is a unit-test fixture (`build.rs:1090`). So a - PCR0/PCR1 mismatch today is not the benign condition the line describes. - -4. **PCR2 is a universal constant — the addendum's derivation is correct, and I - reproduced it independently.** - - `src/enclave-builder/templates/Containerfile.eif:297-302` calls `eif_build` - with exactly **one** `--ramdisk /build/rootfs.cpio.gz` (identical in both - platform clones). - - `aws-nitro-enclaves-image-format-0.4.0/src/utils/mod.rs:660-691` - (`EifBuilder::measure`) feeds kernel ‖ cmdline to both `image_hasher` and - `bootstrap_hasher`, then feeds ramdisk index 0 to `bootstrap_hasher` and - only indices `1..n` to `customer_app_hasher`. With one ramdisk the - customer-app hasher is written **zero bytes**, and PCR0 and PCR1 are - computed over the identical byte stream. - - `src/utils/mod.rs:233-239` constructs all three with - `EifHasher::new_without_cache(...)`, i.e. `block_size == 0`, so - `finalize_reset()` returns `sha384()` - (`src/defs/eif_hasher.rs:85-90`) and `tpm_extend_finalize_reset()` returns - `sha384(0^48 ‖ that)` (`:97-104`). - - `src/utils/mod.rs:176-178` maps `image_hasher → PCR0`, - `bootstrap_hasher → PCR1`, `app_hash → PCR2`. - - Computed here: - `sha384(0^48 ‖ sha384("")) =` - `21b9efbc184807662e966d34f390821309eeac6802309798826296bf3e8bec7c10edb30948c90ba67310f7b964fc500a`. - Both `OPERATORS.md` files record PCR2 as `21b9efbc…` for **both** - components; the recorded prefix matches the derived value. The tables also - record PCR0 == PCR1 per component (`accb679a…` shim, `218d1f64…` hub), - which is exactly what the one-ramdisk layout forces. The tables are not - transcription errors, and `OPERATORS.md:210-211`'s own caveat ("we have not - confirmed with Caution which layer each index measures") is now answered. - - **Consequence, and it is the finding's real weight:** `deploy.sh:220` does - not recommend a weak check. It recommends a check that **cannot fail for any - Caution enclave**, because the value it tells the verifier to accept is - identical for every enclave the platform has ever built, whatever code is - inside it. - -### The attack path is real, with one mechanical correction now folded into the body - -`caution verify` is **all-or-nothing**: PCR0, PCR1 and PCR2 go into -`Nitro::new(...)` together and a mismatch on any of them produces -`✗ Attestation verification FAILED` plus a non-zero exit -(`src/cli/src/lib.rs:7235-7256`, `:7313-7338`). So following `deploy.sh:220` -requires a human to **override an explicit tool-level failure** — it is not a -silent pass. The original Attack Scenario implied verify would report a -per-PCR verdict without an overall failure; that has been corrected in place. - -This does not weaken the finding, because the tool prints the per-index -breakdown (`PCR0: MISMATCH / PCR1: MISMATCH / PCR2: match`) directly beneath the -failure banner, which is precisely the reading `deploy.sh:220` pre-authorises. -Training a verifier to expect and discount an alarm is the whole mechanism here, -and the alarm in question is the only application-distinguishing measurement the -platform produces. - -### Why Medium, and not higher or lower - -- **Not High.** It grants an attacker nothing on its own; a second actor (an - operator or platform deploying code other than the reviewed code) is required, - and the two `OPERATORS.md` runbooks give the correct rule in bold with the - measurement table. `caution verify` also fails loudly. Exploitation needs the - verifier to follow the wrong one of two conflicting in-tree instructions. -- **Not Low.** The instruction is emitted by the tooling **at the moment of - use**, on every attested deploy, and the second carrier - (`shim/deploy/caution/README.md:131-136`) is the only one of the two deploy - *reference* documents that gives an auditor any PCR rule at all — the other - (`shim/deploy/README.md:1055-1056`) asserts no PCR has ever been computed. The - check being disabled is the one that distinguishes the reviewed shim from any - other enclave, and `README.md:71` sells that check to users as the substitute - for trusting their operator. A defect that switches off the trust root's only - discriminating measurement, in the project's own tooling, is squarely Medium. - -### Scope, and how double-counting was avoided - -The superseded PCR2-only rule survives in **three** carriers. This issue owns -**two**: - -- `deploy.sh:220` (the tooling, at the moment of use), and -- `shim/deploy/caution/README.md:131-136` (the deploy reference document, whose - cross-reference additionally attributes the retracted rule to the file that - retracts it). - -The third, `OPEN-QUESTIONS.md:86`, is owned by -`open-questions-pcr-entry-inverts-which-measurement-distinguishes-the-application.md` -and is **excluded** from this issue's severity — it is referenced here only so -the census is complete. Recommendation 2 was rewritten to say so. - -Two neighbouring facts are also owned elsewhere and are **not** counted here: -that nothing joins `EXPECTED_SHA256` to any measured value -(`reproduce-never-builds-the-runtime-stage-that-the-enclave-and-pcr0-are-built-from.md`, -`hub-caution-readme-says-the-attestation-binds-the-running-binary-to-expected-sha256.md`), -and that `README.md:71` names no PCR criterion at all -(`auditor-recipe-omits-the-two-checks-that-decide-where-plaintext-goes-and-names-a-defence-the-platform-does-not-rely-on.md`). -This issue is specifically about a document naming the **wrong** criterion, not -about documents naming **none**. - -### One correction to the runbooks that the report should carry - -"Require **all three** to reproduce" is the right operational rule but implies a -threefold redundancy that does not exist. Caution's attestation carries **one** -independent measurement — over the kernel, the kernel command line, and the single -ramdisk containing the entire enclave (EnclaveOS `init`, `bootproofd`, `caddy`, -`caddy-certfp.sh`, busybox, socat, the generated `run.sh` carrying the whole -`unit.env`, `manifest.json`, and the application's `runtime`-stage filesystem) — -reported twice as PCR0 and PCR1, plus one constant. That TCB is considerably -larger than the "single 4.4 MB static binary" the manifests describe. - -DO NOT DEVIATE FROM THIS TEMPLATE (`docs/SECURITY-ISSUE-TEMPLATE.md`). -LEAVE THIS MESSAGE IN PLACE SO THAT ALL AGENTS KNOW NOT TO DEVIATE FROM THIS TEMPLATE (`docs/SECURITY-ISSUE-TEMPLATE.md`). diff --git a/zeronym-22aa9851caf68-high-medium/medium/enclave-egress-allowlist-is-discarded-by-the-platform-and-both-enclaves-have-unrestricted-outbound.md b/zeronym-22aa9851caf68-high-medium/medium/enclave-egress-allowlist-is-discarded-by-the-platform-and-both-enclaves-have-unrestricted-outbound.md deleted file mode 100644 index 52da6b64..00000000 --- a/zeronym-22aa9851caf68-high-medium/medium/enclave-egress-allowlist-is-discarded-by-the-platform-and-both-enclaves-have-unrestricted-outbound.md +++ /dev/null @@ -1,405 +0,0 @@ -# The Caution platform discards every `egress` rule in `caution.hcl` and grants the enclave unrestricted outbound internet, so the "exfiltration is a network-level impossibility" property both manifests state as a security property does not exist in any configuration - -**Severity**: Medium -**Validation Status**: Confirmed -**Location**: `audit-target/zeronym/shim/deploy/caution/caution.hcl.tmpl:57-77` (the claim, and the `__HUB_EGRESS__` marker at `:78`); `audit-target/zeronym/hub/deploy/caution/caution.hcl.tmpl:57-73` (the same claim, and the `__EGRESS_BLOCKS__` marker at `:82`); `audit-target/zeronym/shim/deploy/caution/assemble-caution.sh:265-320` and `audit-target/zeronym/hub/deploy/caution/assemble-caution.sh:252-273` (the rule renderers); `audit-target/zeronym/deploy.env.example:63`; `audit-target/zeronym/shim/deploy/caution/OPERATORS.md:240-242` and `audit-target/zeronym/hub/deploy/caution/OPERATORS.md:77-78` ("the enclave's entire allowlist"); `audit-target/zeronym/hub/REVIEW.md` (the containment reasoning the allowlist supports). Platform side, read directly during this audit: `codeberg.org/caution/platform` `src/caution-config/src/lib.rs:322-327`, `src/api/src/main.rs:2451`, `src/api/src/deployment.rs:2055-2061`, `terraform/modules/aws/nitro-enclave/user-data.sh:69-95`. -**Found by agent:** Global (focus areas G10 / G12 / G13 — the attestation and reproducibility chain, end to end) -**In scope of audit?** Yes — `*/deploy/**` is explicitly in scope; `AUDIT-INSTRUCTIONS.md` priority area 7 asks in terms "whether the enclave's egress allowlists actually constrain exfiltration"; and manifest claims are in scope as security claims under ICTM. - -## Description - -Both enclave manifests state, explicitly as a *security property* rather than as -tidiness, that the enclave is structurally incapable of sending the plaintext it -holds anywhere except a named `/32`. The shim's, verbatim -(`shim/deploy/caution/caution.hcl.tmpl:57-72`): - -```hcl - # Egress to exactly one host and port: the backing indexer, nothing else. - # - # Deliberately narrower than the platform's example, which allows all - # egress. This narrowness is a security property rather than tidiness. The - # shim sees every wallet's queries in the clear, which is precisely the - # exposure Zeronym exists to contain, so the enclave should be structurally - # incapable of shipping that anywhere except the one indexer it fronts. A - # /32 and a single port make exfiltration to a third party a network-level - # impossibility instead of a promise about the code. - # - # Note what is absent: no port 53, even though the backend is authenticated - # by NAME. That combination is the point. ZIS_BACKEND stays a literal - # address, so the enclave dials an IP and never resolves DNS; - # ZIS_BACKEND_TLS names what the certificate must say. A poisoned DNS answer - # has nothing to poison, and a hijacked address cannot present a valid - # certificate for the name. Update the CIDR whenever the backend IP moves. -``` - -The hub's manifest says the same thing about the enclave that holds every -migration in plaintext (`hub/deploy/caution/caution.hcl.tmpl:57-73`). - -**The Caution platform never applies these rules.** The entire `egress { }` list -— every `cidr_ipv4`, every `port`, every `ip_protocol` — is parsed, validated -against the schema, and then reduced to a single boolean: *is the list empty or -not?* If it is non-empty the enclave is given a NATted TAP bridge to the parent -host with an unconditional `iptables ... -j ACCEPT`, an AWS security group whose -egress rule is `0.0.0.0/0` on all protocols and all ports, and a DHCP-supplied DNS -resolver on the parent. - -So in the deployed configuration: - -- Both enclaves can open a TCP or UDP connection to **any host on the internet, on - any port**, not to a `/32` on one port. -- DNS is available and configured (`/etc/resolv.conf` → `nameserver 10.0.100.1`, - the parent's `dnsmasq`), regardless of whether a port-53 rule was written. The - "no port 53 … a poisoned DNS answer has nothing to poison" argument does not - hold, and the `1.1.1.1/32:53:udp` rule the project does write buys nothing - either. -- The parent host's `dnsmasq` runs with `--log-queries --log-dhcp`, so every - hostname either enclave resolves is written to a log on the parent. **Who reads - that log is deployment-model dependent** (open item 6x): `deploy.sh:156` runs - `caution apps create`, which `shim/deploy/caution/OPERATORS.md:64` defines as - "fully managed: in Caution's AWS account", so in the shipped shape the log is - Caution's; on the BYOC model the same file documents at `:66-71`, it is the - operator's — the threat model's adversary #1. Note this is a *convenience*, not - a new channel: the parent is the enclave's NAT router, so it sees every - destination IP either way, and a UDP/53 query to an allowlisted resolver would - have crossed the same NAT in plaintext. The point is that the manifests reason - as if neither were true. - -This is not the same finding as -`hub-egress-allowlist-is-not-the-exfiltration-barrier-the-manifest-claims.md`, -which argues that the *shipped rule set* is too broad (`0.0.0.0/0:9000:tcp` plus -DNS). That issue's conclusion is right and its mechanism understates the problem: -the breadth of the rules is irrelevant, because **a perfectly narrow rule set -would be discarded in exactly the same way.** An operator who wrote a single -`/32:443:tcp` and nothing else would still get unrestricted outbound. - -## Attack Scenario and Steps - -The egress allowlist is positioned by both manifests as the control that holds -*"instead of a promise about the code"* — i.e. the one that survives the code not -being what the reader believes. The scenario is therefore the one it was written -for. - -1. An attacker obtains code execution inside either enclave, or arranges for the - attested build to contain code the reviewers did not see. Routes that require - no platform break: - - A memory-safety or logic defect on an untrusted-input path. Both enclaves - parse attacker-supplied Sphinx packets and attacker-supplied transaction - bytes from the whole internet, over `nym-sdk 1.21.5-rc.1` and its legacy - transitive stack (`rustls 0.21.12`, `hyper 0.14.32`, `h2 0.3.27`, - `sha2 0.9.9`). - - A supply-chain compromise of any dependency. Attestation binds PCRs to a - build; a malicious dependency reproduces and attests exactly as cleanly as a - benign one. - - `assemble-git-archive-honours-gitattributes-so-the-build-context-is-not-the-committed-tree.md`, - which is the in-tree route to exactly this: source that reproduces and - attests cleanly and is not what a reviewer read. -2. The code opens a TCP connection to any host the attacker chooses and writes out - what the enclave holds: for the shim, every wallet's query stream and source - address plus the plaintext of every diverted transaction; for the hub, the - whole unpublished queue. -3. Nothing blocks it. There is no per-destination policy anywhere in the path — - not in the security group, not in the parent's `iptables`, not in the enclave. - -**Attack Requirements and Assumptions:** - -- **Requires a second failure**, exactly as a defence-in-depth control implies: - code execution, a subverted dependency, or a build the reviewers did not read. - It is not directly exploitable by a network attacker on its own. -- **That is the whole point of the control.** Both manifests say the narrowness - makes exfiltration "a network-level impossibility instead of a promise about the - code". A defence-in-depth control that is entirely absent is a finding even - though a second defect is needed to reach it. -- No AWS, Nitro, or Caution compromise is required, and no operator - misconfiguration: this is how the platform behaves for a correctly written - manifest. -- Under ICTM the documentation half stands on its own, with no second failure at - all: a stated structural property that does not exist is itself the bug, and - this one is stated in the artefact `caution verify` reproduces and an auditor is - pointed at. - -## Impact on Users - -- The containment argument that both components' "we hold your plaintext but we - cannot ship it anywhere" claim rests on does not exist. For the hub — the single - point that sees every migration in the clear — the barrier is back to being "a - promise about the code", which is precisely what the manifest says it is not. -- An auditor or operator who reads the deployed `caution.hcl` (the artefact - published to `--app-source` and rebuilt by `caution verify`) is told the enclave - can only reach one `/32` and cannot resolve DNS. Both statements are false. - `shim/deploy/caution/OPERATORS.md:240-242` repeats it operationally: *"The - `--nym-egress` rules are the enclave's **entire allowlist** for reaching the - mixnet: the nym-api, a gateway, and a DNS resolver."* They are not an allowlist at all. -- Because a successful exfiltration from the hub is total and retrospective — - every migration that ever passed through, joinable against the permanent public - chain — the value of the missing control is high even though its exercise - requires a second defect. -- Secondary, and live without any second defect: the enclave resolves DNS through - the parent host's `dnsmasq --log-queries`, which the project believes it - constrained by allowlisting exactly one resolver (`1.1.1.1/32:53:udp`). On BYOC - that log is the operator's; on the shipped fully-managed path it is Caution's. - This is a smaller point than it first appears — the parent is the enclave's NAT - router and sees every destination IP either way — but it is a second place where - the tree reasons from an allowlist that does not exist. - -## Technical Details / Code Analysis - -**1. Where the rules go to die.** `codeberg.org/caution/platform`, -`src/caution-config/src/lib.rs:322-327` — the *only* place `NetworkConfig::egress` -is consumed: - -```rust -impl NetworkConfig { - /// Outbound internet access is enabled iff at least one egress rule is declared. - pub fn egress_enabled(&self) -> bool { - !self.egress.is_empty() - } -} -``` - -`src/api/src/main.rs:2451`: - -```rust -let egress = ec_network.map(|n| n.egress_enabled()).unwrap_or(false); -``` - -That `bool` is the whole of the egress configuration from this point on. It is -threaded into the Terraform variables (`src/api/src/deployment.rs:2105`, -`:2170` — `egress = if request.egress { "true" } else { "false" }`) and into the -enclave's `run.sh` as a template-block toggle -(`src/enclave-builder/src/build.rs:411-413` — `if egress { enabled_blocks.push("EGRESS"); }`). -No CIDR, port or protocol is carried anywhere. - -**2. What the security group actually says** (`src/api/src/deployment.rs:2055-2061`, -rendered into the enclave instance's Terraform): - -```hcl - egress { - from_port = 0 - to_port = 0 - protocol = "-1" - cidr_blocks = ["0.0.0.0/0"] - description = "Allow all outbound" - } -``` - -**3. What the parent host does** -(`terraform/modules/aws/nitro-enclave/user-data.sh:69-95`, inside -`%{ if egress == "true" ~}`): - -```sh -ip tuntap add mode tap name enclave0 -ip addr add 10.0.100.1/24 dev enclave0 -ip link set enclave0 up -socat TUN,tun-type=tap,iff-no-pi,tun-name=enclave0 VSOCK-LISTEN:3,fork,reuseaddr & -iptables -t nat -A POSTROUTING -s 10.0.100.0/24 -o "$DEFAULT_IFACE" -j MASQUERADE -iptables -A FORWARD -i enclave0 -o "$DEFAULT_IFACE" -j ACCEPT -iptables -A FORWARD -i "$DEFAULT_IFACE" -o enclave0 -m state --state RELATED,ESTABLISHED -j ACCEPT -dnsmasq \ - --interface=enclave0 --bind-interfaces \ - --dhcp-range=10.0.100.10,10.0.100.50,12h \ - --dhcp-option=3,10.0.100.1 --dhcp-option=6,10.0.100.1 \ - --no-daemon --log-queries --log-dhcp & -``` - -`-j ACCEPT` with no destination match is the entire forwarding policy. The last -five lines are the DNS resolver. - -**4. What the enclave does with it** (`src/enclave-builder/templates/run.sh.template`, -`# {EGRESS` block): - -```sh -/bin/socat TUN,tun-type=tap,iff-no-pi,iff-up,tun-name=eth0 VSOCK-CONNECT:3:3 & -... -/bin/busybox udhcpc -i eth0 -n -q -s /bin/udhcpc-script ... -echo "nameserver 10.0.100.1" > /etc/resolv.conf -``` - -So a resolver is configured unconditionally whenever egress is on, which makes -both manifests' "no port 53, and that is the point" paragraphs inoperative. - -**5. What zeronym believes it is doing.** `hub/deploy/caution/assemble-caution.sh:270-272` -writes each rule into the manifest: - -```sh - printf '\n # Nym mixnet egress (gateway / nym-api / DNS / Nyx), operator-allowlisted.\n' >> "$EGRESS" - printf ' egress {\n cidr_ipv4 = "%s"\n port = %s\n ip_protocol = "%s"\n }\n' \ - "$cidr" "$port" "$proto" >> "$EGRESS" -``` - -and `shim/deploy/caution/assemble-caution.sh` does the same. Both scripts spend -substantial effort on the rule set — the shim's warns when the nym-api or DNS rule -count is 0 or 1, and `deploy.env.example:56-59` refuses to allowlist the -Fastly-fronted nym-api on the explicit ground that shared anycast edges "would let -the enclave reach every origin fronted by that CDN". That refusal is reasoning -about a control that is not in force. - -**6. The same fate for ingress CIDRs** (`src/api/src/main.rs:2452-2471` keeps only -port numbers; `src/api/src/deployment.rs:2044-2052` opens each to `0.0.0.0/0` -unconditionally). No security delta for zeronym, which declares `0.0.0.0/0` -anyway, but it means an operator who narrows an ingress CIDR gets no narrowing. - -**Method note.** These are lines of the Caution platform's own public source -(`https://codeberg.org/caution/platform`, cloned during this audit), not vendor -prose. The vendor documentation is silent on the question, which is why -`audit-context/EXTERNAL-CONTEXT.md` §7 and `BRAINSTORM.md` §R2-O carried it as an -unanswerable premise. It was answerable. - -## Recommendations - -1. **Ask Caution whether per-destination egress filtering is planned, and until it - exists, delete the claim.** The three paragraphs at - `shim/deploy/caution/caution.hcl.tmpl:57-72` and - `hub/deploy/caution/caution.hcl.tmpl:57-73` should say plainly that the - `egress` block enables outbound access and does not restrict it, that the - enclave can reach any host on the internet, and that a DNS resolver on the - parent host is configured whenever egress is enabled. This is the cheapest fix - and removes the ICTM overclaim immediately. -2. **Correct the two operator runbooks.** `shim/deploy/caution/OPERATORS.md:240-242` - and `hub/deploy/caution/OPERATORS.md:77-78` both call `--nym-egress` "the - enclave's entire allowlist"; it is not an allowlist. -3. **Remove the dependency on it from the threat model and from every document - that leans on it**, including `hub/REVIEW.md`'s containment reasoning and - `deploy.env.example:56-59`'s reasoning about which nym-api endpoints are safe - to allowlist. -4. **If containment matters — and for the hub it does — implement it where it can - be implemented.** Either (a) obtain per-destination filtering from Caution and - make it part of what `caution verify` covers, or (b) accept that the only - in-scope control is the code, and treat every "structurally impossible" claim - in the tree as a claim about the code instead. -5. **Re-examine the DNS posture.** Both enclaves currently resolve through the - parent host. The nym-sdk's `no_hostname` switch (already identified in - `hub/deploy/caution/assemble-caution.sh:328-331` as "driver work") would remove - zeronym's own reliance on it; nothing removes the resolver from the enclave. - -## Validation Information - -**Validated 2026-08-18. CONFIRMED at Medium.** The platform half of this issue is -an external, moving target, so it was re-verified from a **fresh clone** of -`https://codeberg.org/caution/platform` taken during validation (HEAD -`6051734af680bcb1a96e6034a0d6409af57891f1`, 2026-08-18) as well as against the -earlier clone the filing global audit used (`1f8d8cb39b29a09530b0c3087b5da9198eb3d295`, -2026-08-13). **The line numbers cited in the body are correct for the 2026-08-13 -clone**; at 2026-08-18 HEAD the same code has moved slightly (the sole -`egress_enabled()` consumer is `src/api/src/main.rs:2471`, the security group -`egress {}` block is `src/api/src/deployment.rs:2077-2083` and `:2345-2351`). -The behaviour is identical in both. - -### Every mechanism checked, in both clones - -1. **The rules are reduced to a boolean.** `NetworkConfig::egress_enabled()` is - literally `!self.egress.is_empty()` (`src/caution-config/src/lib.rs:322-327`). - I enumerated **every** reader of `NetworkConfig::egress` across the whole - platform tree: `src/api/src/main.rs` (the boolean), `src/cli/src/lib.rs:1157-1160` - / `:6237` / `:6751` / `:6759` (the same boolean, `config_egress_enabled`), and - `src/cli/src/apps/migrate_procfile.rs:217,248` (Procfile→HCL *conversion*, not - enforcement). **No CIDR, port or protocol from an `egress` block is read - anywhere else in the platform.** - -2. **The security group allows everything outbound.** - `egress { from_port = 0, to_port = 0, protocol = "-1", cidr_blocks = ["0.0.0.0/0"], description = "Allow all outbound" }`, - rendered unconditionally into the enclave instance's Terraform in both - deployment templates. - -3. **The parent host forwards everything.** - `terraform/modules/aws/nitro-enclave/user-data.sh:69-95`, inside - `%{ if egress == "true" ~}`: a TAP device `enclave0` (10.0.100.1/24) bridged - to the enclave over vsock by `socat`, then - `iptables -t nat -A POSTROUTING -s 10.0.100.0/24 -o "$DEFAULT_IFACE" -j MASQUERADE` - and `iptables -A FORWARD -i enclave0 -o "$DEFAULT_IFACE" -j ACCEPT` — **no - destination match of any kind** — plus `dnsmasq … --log-queries --log-dhcp`. - -4. **Nothing filters inside the enclave either.** `iptables`/`nft` appear nowhere - in `src/enclave-builder/`; the only occurrences in the platform are the three - parent-side rules above plus the `dnf install` line. The enclave's - `run.sh.template:53-78` (`# {EGRESS`) brings up `eth0` over vsock, runs - `udhcpc`, and writes `nameserver 10.0.100.1` — so a resolver is configured - **unconditionally whenever egress is enabled**, which is what makes both - manifests' "no port 53, and that is the point" paragraphs inoperative. - -5. **The zeronym half is as quoted.** `shim/deploy/caution/caution.hcl.tmpl:57-72` - and `hub/deploy/caution/caution.hcl.tmpl:57-73` both state the "network-level - impossibility instead of a promise about the code" property verbatim; - `shim/deploy/caution/OPERATORS.md:240-242` and - `hub/deploy/caution/OPERATORS.md:77-78` both call the `--nym-egress` rules "the - enclave's entire allowlist"; `deploy.env.example:56-63` refuses the - Fastly-fronted nym-api on the explicit ground that shared anycast edges "would - let the enclave reach every origin fronted by that CDN"; and - `shim/deploy/caution/assemble-caution.sh:317-325` reasons at length that "DNS - cannot be dropped from this deployment by editing egress; it needs driver - work". **All four are engineering decisions taken about a control that is not - in force.** - -### Why this is a finding and not a defence-in-depth quibble - -It is filed and graded as two things, and only the first needs a second failure: - -- **Live, no second failure required (the ICTM half).** Both manifests state a - *structural* security property, in the artefact that `caution verify` stages - and rebuilds and that auditors are pointed at, and it does not exist in any - configuration. Two runbooks restate it operationally. This is exactly the class - of defect this engagement treats as first-class: a property a reader is told - they have and does not have. `AUDIT-INSTRUCTIONS.md` priority area 7 asks the - question directly ("whether the enclave's egress allowlists actually constrain - exfiltration"); the answer is **no, for every rule set, on both components.** -- **Latent (the containment half).** The missing control is defence-in-depth by - construction, so realising it needs code execution, a subverted dependency, or - a build reviewers did not read. That is the scenario the control was written - for, and its complete absence is a real reduction in the hub's containment: a - hub with a working `/32` allowlist could only exfiltrate through the indexer - hop, which is a materially harder covert channel than an outbound socket to - any host on the internet. - -### Why Medium, not High and not Low - -- **Not High.** No attacker gains anything directly; there is no reachable attack - path that this issue alone opens, and the *confidentiality* boundaries the - system actually depends on (TLS with name pinning to the backend/indexer, the - mixnet) are untouched by it. -- **Not Low.** It affects **both** enclaves, including the one that holds every - migration in the world in plaintext; it is unfixable inside zeronym (the honest - remedies are to delete the claim or to obtain the feature from the vendor); and - it has already misdirected at least four documented design decisions in the - tree. A stated structural guarantee that is absent in every configuration, in - the enclave that holds the plaintext, is Medium. - -### How double-counting was avoided - -This issue is the **root cause and the owner** of "the enclave egress allowlist -provides no containment". Three neighbours touch it and none is counted here: - -- `hub-egress-allowlist-is-not-the-exfiltration-barrier-the-manifest-claims.md` - (still in `plausible/`, filed Medium) argues the narrower point that the - *shipped hub rule set* is too broad (`0.0.0.0/0:9000:tcp` plus DNS). Its - conclusion is correct but is strictly **subsumed** by this one, which shows a - perfectly narrow rule set would be discarded identically, and which covers the - shim as well. **Recommendation to whoever validates it: merge it into this - issue, or confirm it at Info as the rule-breadth documentation half — do not - confirm it at Medium, because that would count one harm twice.** I have not - changed its status; it was not assigned to me. -- `assemble-git-archive-honours-gitattributes-so-the-build-context-is-not-the-committed-tree.md` - already claims the *composition* ("injected code has unrestricted egress") as - an argument for its **own** severity. This issue therefore lists the injection - route as one attack path but does **not** take credit for the compound - outcome; the compound harm belongs there. -- `shim-assemble-nym-egress-cidr-unvalidated-and-no-breadth-advisory.md` and - `hub-assemble-indexers-validation-is-line-oriented-so-a-newline-injects-egress-blocks.md` - both reason about the consequences of *wrong egress rules*. Whoever validates - them should note that an injected or over-broad `egress { }` block changes - nothing about the enclave's actual reachability — their surviving content is - about manifest integrity and about the HCL injection primitive, not about - network exposure. - -### One claim in the body deliberately softened during validation - -The original "the operator's own `dnsmasq` logs every hostname the enclave -resolves" was **qualified rather than kept as written**. Per open item 6x, the -shipped deploy path (`deploy.sh:156`, `caution apps create`) is fully managed — -`shim/deploy/caution/OPERATORS.md:64` puts that parent host in *Caution's* AWS -account, not the operator's — so the log lands with the operator only on BYOC. It -is also not a new channel: the parent is the NAT router and sees every -destination IP regardless, and a UDP/53 query to an allowlisted external resolver -would have crossed the same NAT in the clear. The corrected text is in the -Description and Impact sections. - -DO NOT DEVIATE FROM THIS TEMPLATE (`docs/SECURITY-ISSUE-TEMPLATE.md`). -LEAVE THIS MESSAGE IN PLACE SO THAT ALL AGENTS KNOW NOT TO DEVIATE FROM THIS TEMPLATE (`docs/SECURITY-ISSUE-TEMPLATE.md`). diff --git a/zeronym-22aa9851caf68-high-medium/medium/hub-caution-readme-says-the-attestation-binds-the-running-binary-to-expected-sha256.md b/zeronym-22aa9851caf68-high-medium/medium/hub-caution-readme-says-the-attestation-binds-the-running-binary-to-expected-sha256.md deleted file mode 100644 index 772b3e69..00000000 --- a/zeronym-22aa9851caf68-high-medium/medium/hub-caution-readme-says-the-attestation-binds-the-running-binary-to-expected-sha256.md +++ /dev/null @@ -1,410 +0,0 @@ -# The verification criterion the documentation names does not exist: eleven passages tell an auditor to check the reproduced binary hash "against the one bound into the enclave attestation", and no attestation contains such a value - -**Severity**: Medium -**Validation Status**: Confirmed -**Location**: Worst instance and exemplar: `audit-target/zeronym/hub/deploy/caution/README.md:3-6` (the claim) together with `:55-63` (the "Verify the attestation" procedure built on it). Ten sibling passages carrying the same criterion: `hub/deploy/README.md:18-23`; `hub/deploy/caution/caution.hcl.tmpl:10-17`; `hub/deploy/build.sh:7-9`; `hub/deploy/Containerfile:21-29`; `shim/README.md:179-184`; `shim/deploy/README.md:19-21`; `shim/deploy/README.md:1172-1173`; `shim/deploy/Containerfile:33-37`; `shim/deploy/build.sh:8-10`; `shim/deploy/caution/caution.hcl.tmpl:8-18`. Supporting evidence: `hub/deploy/reproduce.sh:25-35`, `:47-53`, `:79-91`; `hub/deploy/Containerfile:148-149`, `:154-183`; `hub/deploy/caution/assemble-caution.sh:504-519`; Caution platform `src/enclave-builder/src/docker.rs:69`, `src/api/src/builder.rs:1034-1049`, `src/enclave-builder/src/manifest.rs:84-91`, `src/cli/src/lib.rs:6432-6467`. -**Found by agent:** Local (audit of `hub/deploy/caution/README.md`), extended by the `hub/deploy/README.md` and `EXPECTED_SHA256` local audits and by the G10/G12/G13 global audit. **Merged during validation (2026-08-18) with `hub-caution-readme-verify-step-presents-a-bare-elf-hash-as-confirming-the-enclave-measurement.md`, which described the same defect one paragraph further down the same file.** -**In scope of audit?** Yes — `*/deploy/**` and documented security claims are both in scope; AUDIT-INSTRUCTIONS states that "the reproducible-build and attestation chain **is** the trust model here". - -> **Note on the filename.** It is narrower than the issue: eight other files (three confirmed issues, two plausible ones, `globals/G10-G12-G13-…`, `PROGRESS.md`, `BRAINSTORM.md`) reference this file by name, so it was deliberately not renamed. Use the title, not the filename. - -## Description - -Zeronym's trust model is stated the same way in eleven places across ten files: -rebuild the binary from source, get the published hash, and **check that hash -against the hash bound into the enclave attestation**. The hub's rationale -document states it in the strongest form of all: - -> `hub/deploy/caution/README.md:3-6` -> ``` -> The hub receives diverted migrations in plaintext and broadcasts them to the -> Zcash network. That is exactly why it runs as an attested enclave: the -> attestation binds the running binary to `../EXPECTED_SHA256`, so an auditor who -> reproduces the build knows the code holding the plaintext is the code they read. -> ``` - -**No attestation this platform produces contains a binary hash.** Caution builds -the app image with `docker build -f .` and no `--target` -(`src/enclave-builder/src/docker.rs:69`), stages the resulting filesystem into the -EIF's single ramdisk, and measures that ramdisk into PCR0/PCR1. The attested -manifest has a `binary` field (`src/enclave-builder/src/manifest.rs:88`) and the -Containerfile deploy path passes **`None`** for it (`src/api/src/builder.rs:1047`). -`grep -r EXPECTED_SHA256` over the entire Caution platform source returns -**nothing**. So there is no value in any attestation that `EXPECTED_SHA256` could -be compared against, and no tool anywhere performs such a comparison. - -The mechanism that does connect code to measurement is different: -`caution verify` clones the manifest's `app_sources` repository, rebuilds the EIF -from it, and compares PCR0/1/2 against the live attestation -(`src/cli/src/lib.rs:6432-6467`). **Nine of the ten files carrying the claim never -mention it.** In the hub's `deploy/` half outside `caution/` — `README.md`, -`build.sh`, `reproduce.sh`, `Containerfile`, `assemble.sh` — the strings `PCR`, -`PCR0`, `PCR1`, `PCR2` and `caution verify` appear **zero** times (grep-verified), -while `hub/deploy/README.md:18` tells the reader the trust model "hands the -auditor **one job**", and that job is the non-existent comparison. - -The hub's rationale document then converts the claim into a procedure: - -> `hub/deploy/caution/README.md:55-63` -> ``` -> ## Verify the attestation -> -> `caution verify` (from the assembled directory; or `POST /attestation`) returns -> the measurement bound to the running EIF. Confirm it against a local reproduce: -> -> ```sh -> git checkout -> sh zeronym/hub/deploy/reproduce.sh # must print the hash in ../EXPECTED_SHA256 -> ``` -> ``` - -The step it presents as the confirmation is a **self-contained determinism check -on the auditor's own checkout**: `reproduce.sh` reads `EXPECTED_SHA256` out of the -very tree it just built (`:25-35`), and reads no attestation, no PCR and no byte -from any deployment. It prints `zero-indexer-hub: REPRODUCES` for a commit that -has never been deployed anywhere. And the commit it is run against is named by the -party being audited: `` comes from an unsigned text file the -operator's own assemble run wrote (`assemble-caution.sh:504-519`). - -Under ICTM this is the defect the audit instructions call out directly: a property -the reader is told they get, which the mechanism named cannot deliver. - -## Attack Scenario and Steps - -The relevant adversary is **whoever causes the deployed enclave to run code other -than the published tree** — a careless operator who deployed a different commit or -a hand-patched image, a third party who substitutes the image, or Caution's own -build path substituting a different `caution.hcl`. The victim is the auditor or -shim operator whose job is to notice. - -1. The deployed hub runs an image that does not correspond to the published - source — say a build from an uncommitted local edit, or a redeploy of an older - commit than the one published. -2. An auditor is pointed at the documentation. `hub/deploy/README.md:34` sends the - reader to `caution/README.md` for the attested deploy; `shim/deploy/caution/OPERATORS.md:33` - sends the shim's reader to its own `deploy/caution/README.md`. -3. The auditor performs the check the documents name: `git checkout `, `sh reproduce.sh`, confirm it prints the hash in `EXPECTED_SHA256`. -4. It passes — because it is a check on the auditor's own checkout. Two cold builds - of a clean commit agree with each other and with the hash that commit publishes, - regardless of what is deployed. -5. The auditor concludes, in the document's own words, that "the code holding the - plaintext is the code they read". The divergence is not detected, and would have - been detected by `caution verify`, which rebuilds the EIF and compares PCRs. - -**Attack Requirements and Assumptions:** - -- No network attacker and no privileged position is involved. This is a - false-assurance defect: the documented check cannot observe the deployment. -- **Bounded honestly: against a *deliberate* operator the correct check is also - insufficient**, because `caution verify` reproduces from the repository the - operator nominated (`caution-verify-reproduces-from-a-repository-the-operator-nominates-….md`, - confirmed **High**). What this issue uniquely costs is detection of *accidental* - divergence and of substitution by anyone other than the operator — exactly the - cases `caution verify` does catch. -- **Mitigations that must travel with this finding.** The correct procedure exists - in three places and is stated well: `hub/deploy/caution/OPERATORS.md:97-122` and - `shim/deploy/caution/OPERATORS.md:151-211` both prescribe - `caution verify --attestation-url https:///attestation` and require all - three PCRs; `shim/deploy/caution/README.md:17-26` states the decomposition - correctly and **explicitly denies this issue's claim** (see Technical Details - §5). `zeronym/README.md:71`'s auditor bullet also names PCRs. An operator or - auditor who reaches a runbook is not misled by the criterion — they are misled - only about which artefact is load-bearing. -- The exemplar file compounds this: its own worked assemble block omits - `--app-source`, so for the enclave *it* builds, `caution verify` refuses outright - and the reproduce is all the reader has left. Filed separately as - `hub-caution-readme-worked-deploy-builds-an-unverifiable-mixnetless-hub-and-offers-debug.md` - (Low). - -## Impact on Users - -A wallet user's entire protection against the hub — the one component that by -design sees their migration in plaintext alongside every other user's — rests on -the claim that third parties can check what code it runs. Eleven passages tell -those third parties to check something that says nothing about a deployment, and -seven of the eleven present it as the auditor's whole task — the two "one job" -sentences (`hub/deploy/README.md:18`, `shim/deploy/README.md:19`), the two -"gives the auditor the job of" Containerfile comments, the two "the claim that -matters" `build.sh` headers, and `shim/deploy/README.md:1172-1173`'s "the -load-bearing claim". A hub or shim running code nobody reviewed would pass -every check those documents describe. - -The harm is not a live exploit. It is that the audit step the product's threat -model depends on does not close, and that a reader who does the documented work -believes it has. - -## Technical Details / Code Analysis - -**1. What `EXPECTED_SHA256` is, and what the enclave is built from.** - -`hub/deploy/reproduce.sh:47-53` builds `--target export` twice and hashes the -result. That stage is, in its entirety (`hub/deploy/Containerfile:148-149`): - -```dockerfile -FROM scratch AS export -COPY --from=builder /usr/local/bin/zero-indexer-hub /zero-indexer-hub -``` - -The platform builds the Containerfile's **last** stage, `runtime` -(`Containerfile:154-183`), because it passes no `--target`: - -```rust -format!("docker build -f {} .", containerfile) -``` -(Caution platform `src/enclave-builder/src/docker.rs:69`) - -That image becomes the EIF's single ramdisk, and PCR0/PCR1 are computed over -kernel ‖ cmdline ‖ that one ramdisk — which also contains the generated `run.sh` -(carrying the whole `unit.env`), `manifest.json`, `bootproofd`, `caddy`, -`caddy-certfp.sh`, EnclaveOS `init`, busybox and socat. - -**Fairness point, and it is the reason the claim reads plausibly:** the runtime -stage copies the binary *through* the export stage -(`Containerfile:175-177`, "so the bytes an auditor hashes and the bytes that ship -are provably the same file"), so the deployed binary really is bit-identical to the -hashed one. The claim is wrong about the **mechanism**, not about the bytes: the -attestation commits to a filesystem containing that binary, but publishes no digest -of it, so the commitment is only checkable by rebuilding the whole EIF. - -**2. Nothing joins the two values — verified on both sides.** - -- Zeronym side: the only functional readers of `EXPECTED_SHA256` are - `hub/deploy/reproduce.sh:34` (compares against a *locally built* ELF) and - `hub/deploy/caution/assemble-caution.sh:504`, which interpolates the string into a - plain-text `PROVENANCE` file (`:505-519`). `PROVENANCE` is written to the - assembled repo root, and `Containerfile:107-109` copies only `zebra/`, `zaino/` - and `zeronym/` into the builder — so it enters no image layer, no EIF and no PCR. - It is an unsigned, operator-written text file whose last three lines are the same - loop the README prescribes. -- Platform side: `grep -r EXPECTED_SHA256` over `codeberg.org/caution/platform` - (HEAD `6051734`) returns **no matches**, and the manifest's binary hash is never - populated: - -```rust -let mut manifest = enclave_builder::EnclaveManifest::new( - app_source, - enclave_builder::EnclaveSource::GitArchive { … }, - enclave_builder::FrameworkSource::GitArchive { … }, - None, // <-- `binary: Option` - request.run_command.clone(), - None, -); -``` -(`src/api/src/builder.rs:1034-1049`; signature at `src/enclave-builder/src/manifest.rs:84-91`) - -**3. The reproduce step is self-referential and its commit is operator-chosen.** - -`hub/deploy/reproduce.sh:25-35`: - -```sh -ZERO_ROOT="$(git rev-parse --show-toplevel)" -HERE="$ZERO_ROOT/zeronym/hub/deploy" -… -if [ "${EXPECTED+set}" != "set" ]; then - EXPECTED=$(cat "$HERE/EXPECTED_SHA256" 2>/dev/null || echo "") -fi -``` - -After `git checkout `, both the sources built and the hash compared come -from that commit. The script's own header is honest about this (`:2-9`: "build … -twice from cold, check that the two binaries are byte-identical, AND check that -they equal the hash this repo publishes"); it never claims to say anything about a -deployment. Note also `:79-81`: if `EXPECTED_SHA256` is empty or missing the -comparison is skipped with a `NOTE:` and the script still exits 0, so the README's -`# must print the hash in ../EXPECTED_SHA256` comment describes an outcome the -script does not enforce in that state (owned separately by -`reproduce-reports-reproduces-and-exits-zero-when-the-published-hash-comparison-is-skipped.md`). - -`$SHA` in `PROVENANCE` is `git rev-parse HEAD` of the operator's tree, and -`$EXPECTED` is read from the same tree (`assemble-caution.sh:504-519`). The auditor -is asked to check out a commit the audited party named and confirm it is internally -consistent. - -**4. The census — eleven passages, ten files, all verified verbatim at HEAD.** - -| # | File and lines | Wording | -|---|---|---| -| 1 | `hub/deploy/caution/README.md:3-6` | "the attestation **binds the running binary** to `../EXPECTED_SHA256`" | -| 2 | `hub/deploy/README.md:18-23` | "hands the auditor **one job** … check that hash against the one bound into the enclave attestation" | -| 3 | `hub/deploy/caution/caution.hcl.tmpl:10-17` | "check that hash against the one bound into the enclave attestation" **and** "The binary under audit is the one recorded in `deploy/EXPECTED_SHA256`" | -| 4 | `hub/deploy/build.sh:7-9` | "The binary hash is the claim that matters, because that is what gets bound into the enclave attestation" | -| 5 | `hub/deploy/Containerfile:21-29` | "matching it against the hash bound into the enclave attestation" | -| 6 | `shim/README.md:179-184` | "so an auditor can match it against the hash bound into an enclave attestation" | -| 7 | `shim/deploy/README.md:19-21` | "hands the auditor **one job** … check that hash against the one bound into the enclave attestation" | -| 8 | `shim/deploy/README.md:1172-1173` | "The binary hash is the load-bearing claim, because that is what an enclave attestation binds" | -| 9 | `shim/deploy/Containerfile:33-37` | "matching it against the hash bound into the enclave attestation" | -| 10 | `shim/deploy/build.sh:8-10` | "that is what gets bound into the enclave attestation" | -| 11 | `shim/deploy/caution/caution.hcl.tmpl:8-18` | "check that hash against the one bound into the enclave attestation" **and** "The binary under audit is the one recorded in `deploy/EXPECTED_SHA256`" | - -Two of these (#3, #11) are the **manifests that are rendered and pushed to the -platform**, i.e. the worst possible home for a claim about a file the platform -never reads. Two more (#5, #9) are the files the project itself calls the entire -definition of the build. - -The grep that finds them all — the phrase breaks across a newline in shell and -Containerfile comments, so single-line greps miss half the set: - -``` -grep -rniE "bound into|binds? (the )?(running )?binary|attestation binds|hash (is )?bound" audit-target/zeronym -``` -plus a whitespace-flattening pass for the comment-wrapped instances. - -**5. The one artefact in the tree that states it correctly, quoted so the fix has a -model.** `shim/deploy/caution/README.md:17-26`: - -``` -A Nitro attestation binds a measurement of the loaded image into a signed -document … - -| | proves | does not prove | -|---|---|---| -| reproducible build | source and published hash agree | that hash is what runs | -| attestation alone | *some* image runs in a real enclave | which source produced it | -| both | the code you read is the code serving you | | -``` - -followed by "`caution verify` rebuilds from source and compares". The row -"reproducible build … **does not prove** that hash is what runs" is the exact -negation of the eleven passages in the table above. (Its Verify section at `:29-30` then names `caution verify` as the mechanism, which is correct.) - -**6. Boundaries — what this issue does NOT own.** Stated explicitly so the report -does not count one harm twice. - -- *That no binary hash enters the attestation, and that `caution verify` rebuilds - from a repository the operator nominates* — owned by the confirmed **High** - `caution-verify-reproduces-from-a-repository-the-operator-nominates-so-nothing-binds-the-attested-code-to-zeronym.md`. - This issue owns the **documentation census** and the substituted criterion. -- *The retracted "PCR2 is the check that matters" advice* (`deploy.sh:220`, - `shim/deploy/caution/README.md:131-136`) — owned by the confirmed **Medium** - `deploy-script-tells-operators-to-expect-pcr01-failure-and-accept-pcr2-alone.md`, - which also establishes that PCR2 is `sha384(0^48 ‖ sha384(""))`, a constant - across every Caution enclave. Disjoint passages from the eleven above. -- *`README.md:71`'s auditor recipe and its omissions 1 and 4* — owned by the - confirmed **Medium** `auditor-recipe-omits-the-two-checks-that-decide-where-plaintext-goes-and-names-a-defence-the-platform-does-not-rely-on.md`, - whose severity explicitly excludes the `EXPECTED_SHA256` omission and reserves it - for this issue. -- *`shim/deploy/caution/caution.hcl.tmpl:12-14`'s separate "the operator cannot see - the traffic" clause* — owned by `shim-manifest-header-tells-an-auditor-that-attestation-proves-the-operator-cannot-see-the-traffic.md`. - Only the `:9-10` and `:17-18` clauses belong here. -- *That `reproduce.sh` builds `export` rather than `runtime`* — owned by - `reproduce-never-builds-the-runtime-stage-that-the-enclave-and-pcr0-are-built-from.md`. -- *The exemplar file's worked deploy* (`:34-53`) — owned by - `hub-caution-readme-worked-deploy-builds-an-unverifiable-mixnetless-hub-and-offers-debug.md`. - -## Recommendations - -1. Replace `hub/deploy/caution/README.md:3-6` with the mechanism that exists: - *"…that is exactly why it runs as an attested enclave: `caution verify` clones - the manifest's `app_sources`, rebuilds the EIF, and compares all three PCRs - against the attestation the running enclave produces. `EXPECTED_SHA256` is a - separate, weaker claim — that this source tree builds deterministically to a - known binary — and is not part of any attestation measurement."* -2. Replace `:55-63` with the runbook's procedure verbatim - (`hub/deploy/caution/OPERATORS.md:99-110`): - `caution verify --attestation-url https:///attestation`, expect - `✓ Attestation verification PASSED` **and** `✓ TLS certificate binding - verified`. Do not instruct the auditor to take the commit under test from - `PROVENANCE`; the binding commit is the one the manifest pins in `app_sources` - (branch **and** commit), which `caution verify` reads from the attested manifest. - Keep the reproduce, relabelled honestly: *"Separately, `reproduce.sh` checks that - the published source builds deterministically to `EXPECTED_SHA256`. That is a - property of the repository, not of any deployment."* -3. Apply the same correction to all ten sibling passages in the table above — this - is a sweep, not an edit. In particular, delete the `EXPECTED_SHA256` sentence - from **both** `caution.hcl.tmpl` files, and repair the two "one job" sentences - (`hub/deploy/README.md:18`, `shim/deploy/README.md:19`), which foreclose the real - criterion rather than merely omitting it. -4. If a real hash-to-measurement binding is wanted, two options exist and neither is - implemented: (a) publish the reproduced PCR triple next to `EXPECTED_SHA256` at - the same commit, so a single recorded value is comparable to something the - enclave emits; or (b) have `reproduce.sh` build the `runtime` stage and record a - hash of the same filesystem the platform stages. - -## Validation Information - -**Verdict: CONFIRMED, Medium.** Validated 2026-08-18 against the target at HEAD -(`62baea8`, confirmed byte-identical to `audit-context/zero` at that commit), the -monorepo git history, and a local clone of the Caution platform -(`codeberg.org/caution/platform`, HEAD `6051734`). Every line reference in the -census table was re-read verbatim; every platform claim was re-derived from source -rather than taken from the earlier audit notes. - -**This issue absorbed a second filed issue.** -`hub-caution-readme-verify-step-presents-a-bare-elf-hash-as-confirming-the-enclave-measurement.md` -described `:55-63` — the procedure — while this file described `:3-6` — the claim -the procedure implements. They are one harm in one 70-line document, with one fix, -and reporting them separately would double-count. That file has been moved to -`invalid/` carrying a MERGED banner; **it is not a refuted finding**, and its -surviving content is §3 and Recommendation 2 above. - -**What was verified (all independently re-derived):** - -- `docker build -f .` with no `--target` (`src/enclave-builder/src/docker.rs:69`) ⇒ the platform builds the `runtime` stage; `reproduce.sh:51` builds `--target export`. -- `EnclaveManifest::new(…, None, …)` on the Containerfile path (`src/api/src/builder.rs:1047`) ⇒ **no binary hash in any attestation**. -- `grep -r EXPECTED_SHA256` over the whole platform tree: **zero matches**. -- `caution verify` without `app_sources` bails with *"Manifest does not contain app_source - cannot reproduce without source URL"* (`src/cli/src/lib.rs:6432-6437`). -- `reproduce.sh` reads `EXPECTED_SHA256` from the working checkout (`:33-34`) and exits 0 with only a `NOTE:` when it is empty (`:79-81`, `fail` unset on that branch). -- `PROVENANCE` is written to the assembled repo root (`assemble-caution.sh:505`) while the build context copies only `zebra/ zaino/ zeronym/` (`Containerfile:107-109`) ⇒ unmeasured. -- Zero occurrences of `PCR`/`PCR0`/`PCR1`/`PCR2`/`caution verify` in `hub/deploy/README.md`, `build.sh`, `reproduce.sh`, `Containerfile`, `assemble.sh` and in `shim/README.md`; the only three hits in `shim/deploy/README.md` (`:664`, `:1055-1056`) are stale "no PCR has ever been computed" statements, not verification instructions. - -**Five corrections applied against the filed texts. Do not restore them.** - -1. **"Eleven passages across nine files" → ten files.** Recounted: hub carries five (in five files), shim six (in five files, because `shim/deploy/README.md` carries two). The passage count of eleven is right; the file count was not. -2. **The hub's manifest carries *both* clauses too.** The filed text said only the shim's `caution.hcl.tmpl` carried the "binary under audit is the one recorded in `deploy/EXPECTED_SHA256`" sentence alongside the binding sentence. `hub/deploy/caution/caution.hcl.tmpl:16-17` carries it verbatim as well. -3. **The absorbed issue's claim that the prescribed verification "consumes nothing from the running enclave" was too strong and has been dropped.** `:57` does tell the reader to run `caution verify`, which *is* the correct check. The accurate defect is that the document **subordinates** it — "Confirm it against a local reproduce" makes the local determinism check the arbiter — and that the same file's assemble block removes verify's ability to run at all. -4. **The absorbed issue's point (d) — "the invocation form given is the one the runbook says fails" — is not supportable as stated and was struck.** `caution apps create` **does** write `.caution/deployment.json` (`src/cli/src/lib.rs:5297`), so a bare `caution verify` in the assembled directory can resolve the app, contradicting `hub/deploy/caution/OPERATORS.md:101-105`. The real, smaller defect in that invocation is different: `get_attestation_url()` returns `http:///attestation` (`:6094-6108`), which takes the raw-IP branch of `tls_connection` (`:221-241`); on that branch the certificate binding can be **skipped** with only a warning while `✓ Attestation verification PASSED` still prints (`:6902-6932`, `:7284-7306`). The runbook's `--attestation-url https:///attestation` form takes the strong branch. That residual is **already owned** by the confirmed Medium `attested-tls-binding-is-verified-once-by-hand-if-ever-….md` and is recorded here only as a cross-reference; it is not counted in this issue's severity. -5. **The claim is wrong about the mechanism, not about the bytes.** The filed addendum's "the attestation and `EXPECTED_SHA256` are computed over different artefacts" is true but incomplete: `Containerfile:175-177` copies the binary *through* the export stage, so the deployed binary is bit-identical to the hashed one. Stated in Technical Details §1 so a developer reading the fix is not told something they can immediately falsify. - -**Exploitability, assessed honestly.** No attacker capability is created. Against a -deliberate operator the documented check and the correct check are both -insufficient, because verification reproduces from the operator's own published -tree (confirmed High, item 7a: *measurement discloses a value; it never detects a -change*). The residual this issue owns is real but narrower than the filed text -implied: `caution verify` detects **accidental** divergence (wrong commit, -hand-patched image, a substituted `caution.hcl` from the platform's own build path) -and substitution by any party other than the operator; the documented criterion -detects neither, and in seven of the eleven passages it is presented as the -auditor's complete job. - -**Severity: Medium, and the reasoning for not grading it higher or lower.** - -- *Not High.* No user data is exposed and no attacker is enabled; the fix is prose. - The exploitable mechanism is owned at High elsewhere. -- *Not Low, unlike the closest precedent.* - `operators-runbook-attributes-the-hub-destination-to-the-binary-hash-and-egress-rules-neither-of-which-binds-it.md` - was deliberately held to Low because it was one property in one runbook whose - exploitable halves were owned at Medium and High. This is different in two ways - that matter: the footprint is eleven passages in ten files including both pushed - manifests and both Containerfiles, and in nine of those ten files **no correct - criterion is stated anywhere**, so the reader is not merely given a weak reason — - they are given a complete and wrong instruction set. -- *Consistent with the sibling Mediums.* - `deploy-script-tells-operators-to-expect-pcr01-failure-and-accept-pcr2-alone.md` - is Medium for two carriers of a check that cannot fail; this is the same shape at - five times the footprint for a check that does not exist. -- *No double count.* The boundary list in Technical Details §6 was written before - the severity was set, and each neighbouring issue's owned scope was re-read to - confirm the reserved allocation (in particular `auditor-recipe-omits-…`'s - explicit exclusion of omission 3, recorded in `PROGRESS.md` item 7k). - -**Positives that must survive into the report, so this does not read as a -condemnation of the hub's documentation.** - -- `hub/deploy/README.md` is, on the record of this audit, the **strongest deploy - document in the tree** (`PROGRESS.md` row for that file): it diagnoses its own - siblings' build-context errors, names the vendored `nym-upgrade-mode-check` that - the shim's reference omits, transcribes no literal hash and is therefore immune to - the shim's three-way hash drift, and gives the most honest statement of the mixnet - network relaxation in the tree. It carries passage #2 and nothing else here. -- `hub/deploy/caution/OPERATORS.md` — the runbook sitting in the same directory as - the exemplar — gets this entirely right, including the "require all three PCRs, - do not accept a PCR2 match alone" argument (`:112-122`). -- `shim/deploy/caution/README.md:17-26` is the one artefact in the tree that states - the decomposition correctly, and it should be the template for the sweep. -- The exemplar file is also right about things the audit should credit: `:18-21` - gets `--indexer-tls` right and emphatically, and `:22-25` gets the HTTP/1.1 - vs h2c distinction right for a non-obvious platform reason. - -DO NOT DEVIATE FROM THIS TEMPLATE (`docs/SECURITY-ISSUE-TEMPLATE.md`). -LEAVE THIS MESSAGE IN PLACE SO THAT ALL AGENTS KNOW NOT TO DEVIATE FROM THIS TEMPLATE (`docs/SECURITY-ISSUE-TEMPLATE.md`). diff --git a/zeronym-22aa9851caf68-high-medium/medium/hub-chain-unbounded-indexer-response-body.md b/zeronym-22aa9851caf68-high-medium/medium/hub-chain-unbounded-indexer-response-body.md deleted file mode 100644 index cc3856e7..00000000 --- a/zeronym-22aa9851caf68-high-medium/medium/hub-chain-unbounded-indexer-response-body.md +++ /dev/null @@ -1,353 +0,0 @@ -# The hub buffers an indexer response with no size ceiling, so a ~100-byte unauthenticated lookup buys a multi-megabyte allocation it then throws away - -**Severity**: Medium -**Validation Status**: Confirmed -**Location**: `audit-target/zeronym/hub/src/chain.rs:373-377` (`round_trip`, the unbounded `collect()`) and `:415-430` (`unframe`, whose length check runs *after* the bytes are resident), reached from `:221-266` (`get_transaction`), `:176-198` (`broadcast`) and `:155-173` (`tip_height`). Consumers: `audit-target/zeronym/hub/src/server.rs:296-322` (`Hub::lookup`), `:487-505` (the HTTP handler), `:508-517` (`found`, which copies the body again into the response), `audit-target/zeronym/hub/src/nym.rs:249-303` (`build_lookup_reply`, which **discards** anything over ~64 KiB after buffering it). Enclave size: `audit-target/zeronym/hub/deploy/caution/caution.hcl.tmpl:39-40`. Contrast: every *inbound* body in the same crate is wrapped in `http_body_util::Limited` (`server.rs:488`, `:523`). -**Found by agent:** Local (file audit of `hub/src/chain.rs`); validated 2026-08-18 -**In scope of audit?** Yes - -## Description - -`chain.rs` is the hub's only outbound network client, and every reply it reads -comes from a party the engagement's threat model designates untrusted: - -> `hub → indexer`: sees the whole batch seconds before it is public; **can lie -> about the tip and about publish verdicts** -> — `audit-context/AUDIT-INSTRUCTIONS.md`, "Trust boundaries" - -`round_trip` reads that reply with an unbounded `collect()` -(`chain.rs:373-377`): - -```rust - let collected = tokio::time::timeout(RPC_TIMEOUT, response.into_body().collect()) - .await - .map_err(|_| -> BoxError { "reading the gRPC response timed out".into() })??; - let trailers = collected.trailers().cloned(); - let body = collected.to_bytes(); -``` - -There is no ceiling of any kind on how many bytes are accumulated: no -`Limited`, no `content-length` check, and no cap derived from the gRPC frame -header. `unframe` does check the declared length (`chain.rs:424-428`), but only -*after* the entire body is already resident, so that check cannot bound the -allocation. HTTP/2 flow control does not bound the total either: `collect()` -consumes eagerly, which releases receive-window capacity as fast as it arrives, -so the window governs bytes *in flight*, not bytes *accumulated*, and the peer -may keep sending for the whole 10 s `RPC_TIMEOUT`. - -**Two facts make this a defect rather than a design choice.** - -First, it is a gap specific to this file, not a house style. The same crate -applies `http_body_util::Limited` to *every* body it accepts from an untrusted -peer — `server.rs:488` (`MAX_LOOKUP_BYTES`, 64 bytes) and `server.rs:523` -(`MAX_TX_BYTES`, 64 KiB) — and `queue.rs:61` caps a stored transaction at -64 KiB. The indexer's response is the one untrusted input read without a limit. - -Second, **the hub can never use more than ~64 KiB of it.** `nym.rs:295-302` -says so in as many words: - -> The reply budget is nine bytes under the submit cap, so **an indexer can return -> a transaction that fits nowhere in a reply frame.** Fail closed rather than -> truncate. - -So on the mixnet lookup arm the hub buffers a multi-megabyte body, decodes it, -discovers it will not fit a `LookupReplyV1` frame, and answers `error`. The -memory is spent to produce a refusal. The possibility of an oversized answer is -understood one file away; what is missing is the bound that would stop it being -buffered. - -## Attack Scenario and Steps - -Two paths reach it, and they have very different reachability. The distinction -matters for severity and is stated up front. - -**Path A — an unauthenticated internet client and an entirely honest indexer. -This is the one that makes the finding.** - -1. `POST /transaction` is served unconditionally — it is *not* behind - `--http-submit` (`server.rs:442-445`) — on an enclave whose domain the - in-enclave Caddy maps onto port 8083 from `0.0.0.0/0` - (`caution.hcl.tmpl:51-55`, `:97-105`), with no ACL, no authentication and no - rate limit. The same core is also reachable over the mixnet, whose address - `GET /nym-address` publishes to anyone (`server.rs:446-449`). -2. The attacker picks the txid of a large Zcash mainnet transaction. This is - public data. A transaction must fit in a block, and - `MAX_BLOCK_BYTES = 2_000_000` (`zebra-chain/src/block/serialize.rs:24`), so a - single transaction can be ~2 MB — **31× the hub's own `MAX_TX_BYTES` and - 31× the largest reply a Nym lookup can carry.** If no suitable transaction - already exists on chain, the attacker can mint one once, for a few dollars of - ZIP 317 fees, and reuse its txid forever. -3. The attacker sends `POST /transaction` carrying that hash — ~100 bytes. -4. Every request misses the queue (`server.rs:297`; it is not a diverted - migration) and falls through to `chain.get_transaction()`, which fetches the - full ~2 MB transaction from the honest indexer, once per configured endpoint. -5. Each in-flight lookup materialises the payload several times over: the - collected body, the contiguous `to_bytes()` copy (`chain.rs:377`), the - `RawTransaction.data` `Vec` prost decodes into (`chain.rs:429`, carried to - `TxLookup::Found` at `:233-236`), and finally - `Bytes::copy_from_slice(tx_bytes)` when the HTTP response is built - (`server.rs:509`). Call it 4–6 MB transiently per concurrent lookup, of which - ~2 MB is *retained* in the response until it has been written to the client. -6. `serve` spawns one task per accepted connection with **no concurrency - semaphore** (`server.rs:354-408`), so the multiplier is limited only by the - attacker's connection count. That missing bound is the separate confirmed - issue `hub-http-lookup-path-has-no-concurrency-bound.md`; the defect owned - *here* is the per-response byte ceiling, which is missing independently of it - and which would bound the damage under any concurrency limit eventually - chosen. -7. There is **no response-write timeout** anywhere on this path: the - `TokioTimer` installed at `server.rs:397-398` enables hyper's *header-read* - timeout only. A client that stops reading therefore parks its ~2 MB response - buffer in the hub indefinitely, and the in-enclave Caddy streams rather than - fully buffering, so the back-pressure lands on the hub. - -Even the **bounded** mixnet arm is uncomfortable: 64 concurrent lookups -(`nym.rs:54`) × ~2 MB fetched and then discarded is ~128 MB of churn on top of a -64 MiB queue and whatever the HTTP arm is doing, on a 2 GB enclave. - -**Path B — a hostile or compromised indexer.** The indexer answers any of the -three calls (`GetLightdInfo`, `SendTransaction`, `GetTransaction`) with an -arbitrarily large body inside the 10 s window. `GetLightdInfo` needs no attacker -action at all: `batcher.rs:290` polls it every 30 seconds. This is the only path -on which the response is *truly* unbounded. It requires control of a configured -`ZIH_INDEXERS` endpoint, which the audit treats as a hub-trust/robustness defect -rather than an internet-reachable weapon, and such a party already holds cheaper -levers (they receive every batch in plaintext). It is recorded because the -indexer is frequently a *third party* the hub operator does not control, and -because it is what makes "no ceiling" more than an aesthetic complaint. - -**Attack Requirements and Assumptions:** -- Path A needs only network reach to the hub's published address and one public - txid. No shim, no mixnet position, no credential, no funds beyond an optional - one-off fee. -- Path B needs the configured indexer to be hostile or compromised — a party the - threat model already treats as untrusted and able to lie. -- What makes Path A realistic: `POST /transaction` is reachable by design (it is - how the shim answers `GetTransaction`), the ingress is `0.0.0.0/0`, there is - no rate limit anywhere on it, and the in-enclave Caddy imposes none either. -- What limits it: with an honest indexer the per-response size is bounded by - consensus at ~2 MB, so reaching a 2 GB enclave's ceiling needs several hundred - concurrent lookups *and* an indexer willing and able to serve several hundred - megabytes concurrently. Neither the enclave's `RLIMIT_NOFILE` nor the - indexer's throughput could be measured in this environment. - -## Impact on Users - -Killing the hub process is not a restart, it is data loss. The queue is -deliberately RAM-only and the enclave is diskless -(`caution.hcl.tmpl:34-38`: "The queue is deliberately NOT persisted"), and -`server.rs:371-374` records exactly what that costs: - -> the RAM-only queue -- every migration already acked to a wallet and waiting for -> the next flush -- went with it. - -Every wallet whose migration is in the queue was already told `error_code 0` at -mixnet hand-off (`shim/src/hub.rs:231-240`, dispatch-only submit), and the shim -keeps no copy. An OOM therefore silently destroys migrations their owners -believe are on their way to the network, with no retry anywhere in the system. -An OOM is a `SIGKILL`, so even the hub's own `unpublished … they are lost` -accounting (`batcher.rs:325-331`) does not run. - -Short of an OOM, the same requests are a straightforward memory- and -bandwidth-amplifier against an enclave that has 2 GB for everything: ~100 bytes -in, up to ~2 MB allocated and up to ~2 MB pulled across the indexer link that -the *operator* pays for. - -Beyond the loss, an on-demand hub kill is a *privacy* primitive: with the hub -down, every shim in the fleet walks its retry/failover ladder simultaneously and -the whole fleet is unprotected in a window the attacker chose. A restart also -mints a fresh Nym identity (`nym_driver.rs`, `Ephemeral::default()` in a -diskless enclave), which is the terminal state of the confirmed -`hub-nym-driver-automatic-fresh-identity-permanently-invalidates-every-shim.md`. - -## Technical Details / Code Analysis - -The read path, in full (`hub/src/chain.rs:344-402`, abridged to the relevant -lines): - -```rust -async fn round_trip(stream: IO, request: hyper::Request>) -> Result -{ - let (mut sender, conn) = http2::Builder::new(TokioExecutor::new()) - .handshake(TokioIo::new(stream)) - .await?; - ... - let collected = tokio::time::timeout(RPC_TIMEOUT, response.into_body().collect()) - .await - .map_err(|_| -> BoxError { "reading the gRPC response timed out".into() })??; - let trailers = collected.trailers().cloned(); - let body = collected.to_bytes(); -``` - -`http2::Builder::new(...)` is used with defaults; hyper's HTTP/2 client exposes -no maximum response body size, and no window setting is applied here — but no -window setting would help, because a window bounds bytes in flight and -`collect()` releases it continuously. - -The bound that does exist is applied too late (`hub/src/chain.rs:415-430`): - -```rust -fn unframe(body: &[u8]) -> Result { - if body.len() < GRPC_PREFIX_LEN { - return Err("gRPC response shorter than its frame header".into()); - } - ... - let declared = u32::from_be_bytes([body[1], body[2], body[3], body[4]]) as usize; - let message = GRPC_PREFIX_LEN - .checked_add(declared) - .and_then(|end| body.get(GRPC_PREFIX_LEN..end)) - .ok_or_else(|| -> BoxError { "gRPC frame length overruns the body".into() })?; - Ok(M::decode(message)?) -} -``` - -`unframe` itself is memory-safe and panic-free (`checked_add`, `body.get`), and -its unit tests at `chain.rs:674-682` pin that. The defect is not in `unframe`; it -is that `body` is fully materialised before `unframe` ever sees it. - -The lookup path an unauthenticated client drives -(`hub/src/server.rs:296-322`): - -```rust - pub async fn lookup(&self, wire_hash: &[u8]) -> LookupOutcome { - if let Some(bytes) = self.queue.find_by_txid(wire_hash) { ... } - - match self.chain.get_transaction(wire_hash).await { - Ok(TxLookup::Found { data, height }) => { - LookupOutcome::Found { data: Zeroizing::new(data), height } - } - ... -``` - -the handler that reaches it, with an inbound limit but no outbound one -(`hub/src/server.rs:487-494`): - -```rust - let collected = match Limited::new(req.into_body(), MAX_LOOKUP_BYTES) - .collect() - .await - { ... }; -``` - -and the point at which the buffered bytes are proved useless on the mixnet arm -(`hub/src/nym.rs:295-302`): - -```rust - match wire::encode_lookup_reply(&nonce, &reply) { - Ok(frame) => Some(frame), - // The reply budget is nine bytes under the submit cap, so an indexer - // can return a transaction that fits nowhere in a reply frame. Fail - // closed rather than truncate. - Err(err) => { ...; Some(error_reply(nonce)) } - } -``` - -Note that `chain.rs` fans out one such read per endpoint per call -(`chain.rs:222-249` for lookups, `:181-195` for publishes), so the per-call -figure multiplies by `endpoints.len()`. - -## Recommendations - -- **Wrap the response body in `http_body_util::Limited` in `round_trip`**, with a - ceiling passed in by the caller: a few kilobytes for `GetLightdInfo`, and - `queue::MAX_TX_BYTES + GRPC_PREFIX_LEN` for `GetTransaction` and - `SendTransaction`. Anything larger is already destined to be discarded - (`nym.rs:295-302`), so the cap costs the hub nothing it can use. This is a - handful of lines and closes both paths. -- **Reject on `content-length` before reading**, where the peer supplies one, so - the common case never allocates at all. -- **Bound concurrency on the HTTP serving path** the way the mixnet path already - is — see the confirmed `hub-http-lookup-path-has-no-concurrency-bound.md`. - That issue and this one multiply; either fix alone leaves the product of the - other two factors. -- **Add a response-write deadline** so a client that stops reading cannot park a - materialised response body in the hub indefinitely. - -## Validation Information - -**Verdict: CONFIRMED. Severity confirmed at Medium**, and the two paths have -been separated because they do not have the same reachability and the filed text -ran them together. - -### Every mechanical claim re-verified against the target - -| Claim | Verified at | -|---|---| -| The response body is collected with no ceiling | `hub/src/chain.rs:373-377` — `response.into_body().collect()` with only a `tokio::time::timeout` around it | -| The only bound is time, not bytes | `hub/src/chain.rs:48` (`RPC_TIMEOUT = 10 s`), applied at `:291` and `:373` | -| `unframe`'s length check runs after full materialisation | `hub/src/chain.rs:415-430`, called at `:336` on the already-collected `response` | -| `to_bytes()` makes a second contiguous copy when the body arrived in multiple frames | `hub/src/chain.rs:377` | -| prost decodes `RawTransaction.data` into a fresh `Vec`, carried out as `TxLookup::Found` | `hub/src/chain.rs:232-236`, `:429` | -| The HTTP response copies it again | `hub/src/server.rs:508-509` — `Bytes::copy_from_slice(tx_bytes)` | -| Every *inbound* body in the same crate is `Limited`; this one is not | `hub/src/server.rs:488`, `:523`; `hub/src/queue.rs:61` | -| The hub cannot use more than ~64 KiB of the answer and discards the rest | `hub/src/nym.rs:295-302`; `hub/src/wire.rs` reply budget | -| `POST /transaction` is unauthenticated, ungated and internet-reachable | `hub/src/server.rs:442-445`; `hub/deploy/caution/caution.hcl.tmpl:51-55`, `:97-105` | -| No concurrency bound on the HTTP arm; 64 on the mixnet arm | `hub/src/server.rs:354-408`; `hub/src/nym.rs:54`, `:153` | -| No response-write timeout exists; the installed timer only re-enables the header-read timeout | `hub/src/server.rs:392-398` and its own comment | -| A Zcash transaction may be ~2 MB | `zebra-chain/src/block/serialize.rs:24` — `MAX_BLOCK_BYTES = 2_000_000`; a transaction must fit in a block | -| `GetLightdInfo` is polled every 30 s with no attacker involvement | `hub/src/batcher.rs:71` (`POLL_INTERVAL`), `:290` | -| The enclave has 2 GB and a RAM-only queue | `hub/deploy/caution/caution.hcl.tmpl:34-40` | -| A hub death destroys acked migrations | `hub/src/server.rs:370-374`; `shim/src/hub.rs:231-240` | - -### `AVOIDING-FALSE-POSITIVES.md` §5 applied in both directions - -§5's test is *what resources would the attacker need, and what would stop them?* - -*What the attacker needs:* a TCP connection and a ~100-byte request naming a -public txid. That is the **inverse** of §5's canonical false positive -("must send 100 GB"): the ratio is roughly 1 byte in to 20,000 bytes allocated, -which is §5's own stated *real* pattern — *"1KB request → 1GB memory -allocation"*, *"single connection consuming unbounded resources"*. - -*What would stop them:* nothing in the target and nothing in the platform. There -is no ACL, no rate limit, no per-response cap, no `content-length` check, no -concurrency bound on the HTTP arm, and no response-write deadline. The -in-enclave Caddy is a plain reverse proxy with none of these configured. - -*Honest deflation, and it is why this is Medium and not High.* With an **honest** -indexer the per-response size is bounded by Zcash consensus at ~2 MB, so this -alone is a large-but-finite allocation; converting it into an OOM requires the -*separate* missing concurrency bound plus an indexer able to serve hundreds of -megabytes concurrently, and the enclave's descriptor limit may bind first. -Neither could be measured here. The **unbounded** case is genuinely unbounded but -needs a hostile configured indexer, which PROGRESS.md item 6p classifies as a -hub-trust defect rather than an internet-reachable weapon, and such a party -already sees every batch in plaintext. Medium reflects the union: a certain, -cheap, unauthenticated amplifier with an uncertain path to the worst outcome. - -### Corrections made to the filed text - -- **The two paths were re-ordered and re-scoped.** The filing led with the - hostile indexer ("Path A") and called the whole thing unbounded. With an - honest indexer it is bounded at ~2 MB by consensus; the text now says so, cites - `MAX_BLOCK_BYTES`, and applies item 6p's bound to the hostile-indexer leg - explicitly. -- **The HTTP/2 flow-control paragraph was corrected.** The filing asserted a - specific 64 KiB default window. hyper 1.11.0 does not use h2's bare defaults - and the exact figure is not load-bearing; the correct statement is that a - window bounds bytes *in flight* while `collect()` releases capacity as it - consumes, so no window setting bounds the total. Restated that way. -- **A mechanism the filing missed was added, and it is the strongest one:** - there is no response-write deadline, so a client that stops reading parks the - materialised response in the hub indefinitely rather than merely transiently. -- **The sharpest argument was added:** on the mixnet arm the buffered bytes are - provably useless — `build_lookup_reply` discards anything that will not fit a - reply frame — so the allocation is spent to produce a refusal. -- **The "a few hundred concurrent lookups is enough to pass 2 GB" claim was - qualified** rather than asserted, because it depends on the enclave's - descriptor limit and the indexer's throughput, neither of which is knowable - from the tree. -- Severity left at Medium; the filed value was right. - -### Relationship to other issues (stated so nothing is counted twice) - -- The missing concurrency bound is `hub-http-lookup-path-has-no-concurrency-bound.md` - (Confirmed, High). The two multiply — G21 §4.3 records the composition — and - each file states it once. Neither claims the other's defect as its own. -- The flush-side `k × n` fanout and the per-flush copies of the batch belong to - `hub-chain-connection-per-call-fanout-and-flush-memory-amplification.md`. -- The confidentiality of `POST /transaction` is the confirmed - `hub-unauthenticated-pre-publication-transaction-disclosure.md`. - -DO NOT DEVIATE FROM THIS TEMPLATE (`docs/SECURITY-ISSUE-TEMPLATE.md`). -LEAVE THIS MESSAGE IN PLACE SO THAT ALL AGENTS KNOW NOT TO DEVIATE FROM THIS TEMPLATE (`docs/SECURITY-ISSUE-TEMPLATE.md`). diff --git a/zeronym-22aa9851caf68-high-medium/medium/hub-chain-zaino-node-rejections-are-never-verdicts.md b/zeronym-22aa9851caf68-high-medium/medium/hub-chain-zaino-node-rejections-are-never-verdicts.md deleted file mode 100644 index d703d8c8..00000000 --- a/zeronym-22aa9851caf68-high-medium/medium/hub-chain-zaino-node-rejections-are-never-verdicts.md +++ /dev/null @@ -1,396 +0,0 @@ -# Against a Zaino indexer the hub can never obtain a `Rejected` verdict, so a refused transaction is re-published at every flush forever — the queue has no other eviction path, and one unauthenticated junk fill becomes permanent - -**Severity**: Medium -**Validation Status**: Confirmed -**Location**: `audit-target/zeronym/hub/src/chain.rs:491-501` (`classify_publish_failure`), the claim at `:432-437` (`classify_send_response`'s doc: "lightwalletd's convention, **which zaino follows**"), the rationale at `:447-473` (`best_of`) and `:503-512`, and the in-file test at `:589-607` that pins `"13"` as retryable. Consumed at `audit-target/zeronym/hub/src/batcher.rs:365-390`; the requeue with no eviction rule is `audit-target/zeronym/hub/src/queue.rs:279-295`. Backend behaviour (read directly from the vendored source, not inferred): `audit-context/zero/zaino/packages/zaino-state/src/indexer/node_backed_indexer.rs:1561-1570`, `.../chain_index.rs:2912-2922`, `.../error.rs:118-168` and `:525-531`, `audit-context/zero/zaino/packages/zaino-serve/src/rpc/grpc/service.rs:25-40,159-177`. -**Found by agent:** Local (file audit of `hub/src/chain.rs`) -**In scope of audit?** Yes - -## Description - -`chain.rs` splits publish failures into two classes and the batcher acts on the -split: `Rejected` is dropped permanently, `Retryable` is requeued for the next -flush. The split rests on an assumption stated as fact in the code -(`chain.rs:434-437`): - -> lightwalletd's convention, **which zaino follows**: `error_code == 0` means -> success and `error_message` carries the txid. A non-zero code carries the -> node's rejection text … - -The first half is true of Zaino. **The second half is not.** In the whole of -Zaino's non-generated source there is exactly **one** construction of a -`SendResponse`, and it hardcodes success -(`zaino/packages/zaino-state/src/indexer/node_backed_indexer.rs:1561-1570`): - -```rust - /// Submit the given transaction to the Zcash network - async fn send_transaction(&self, request: RawTransaction) -> Result { - let hex_tx = hex::encode(request.data); - let tx_output = self.send_raw_transaction(hex_tx).await?; - - Ok(SendResponse { - error_code: 0, - error_message: tx_output.hash().to_string(), - }) - } -``` - -The `?` sends every node refusal down the error path instead, and that path -flattens to gRPC status **13 (INTERNAL)** with a fixed message. Traced end to -end in the vendored tree: - -`ChainIndex::send_raw_transaction` maps *any* backing-node error through -`ChainIndexError::backing_validator` (`chain_index.rs:2912-2922`), which is -`kind: InternalServerError, message: "InternalServerError: error receiving data -from backing node", source: Some()` (`error.rs:525-531`). -The `From for tonic::Status` impl then renders -that as `tonic::Status::internal(err.message)` (`error.rs:132-137`) — **using the -fixed message and discarding the `source` chain**. The gRPC handler propagates it -verbatim (`zaino-serve/src/rpc/grpc/service.rs:29-40`, `:159-177`). - -So against Zaino a node rejection arrives at the hub as a `GrpcStatusError` with -`code == "13"` and `message == "InternalServerError: error receiving data from -backing node"`, and `classify_publish_failure` treats only codes 3 and 9 as -verdicts (`chain.rs:491-501`): - -```rust -fn classify_publish_failure(err: &BoxError) -> Publish { - let reason = err.to_string(); - match err.downcast_ref::() { - Some(status) - if status.code == GRPC_INVALID_ARGUMENT || status.code == GRPC_FAILED_PRECONDITION => - { - Publish::Rejected { reason } - } - _ => Publish::Retryable { reason }, - } -} -``` - -`13` falls to the `_` arm and becomes `Retryable`. The file's own unit test -asserts exactly this (`chain.rs:598`: `for code in ["14", "4", "8", "2", "13"]` -… *"must be retryable"*). - -**Consequence: with a Zaino backend, `Publish::Rejected` is unreachable for a -node rejection**, by both routes — `classify_send_response`'s non-zero branch is -dead because Zaino never builds a non-zero `SendResponse`, and -`classify_publish_failure` cannot produce a verdict because Zaino's -`send_transaction` path emits neither code 3 nor code 9. Every transaction the -network refuses is requeued and offered again at every subsequent flush. - -**Nothing else evicts it.** `Queue::requeue` (`queue.rs:279-295`) re-inserts -unconditionally and re-charges the bytes; it applies no expiry check, no retry -counter and no deadline. `Entry::received_height` exists and its doc comment says -*"Drives the confirmation deadline"* — a repository-wide grep shows it is read -**nowhere** (`queue.rs:135` and `:235` are its only two occurrences). `admit`'s -`survives_next_flush` check runs at admission only. So the *only* thing that ever -removes an entry from the queue is a terminal verdict from the indexer, and -against Zaino no terminal verdict except success exists. - -The code says so itself, and the sentence is false against this backend -(`batcher.rs:383-389`): - -> The entry keeps its original expiry; when the indexer answers again a stale one -> gets the node's verdict and leaves. - -Against Zaino a stale entry gets `INTERNAL` and stays. - -## Attack Scenario and Steps - -1. The hub is configured against a Zaino `CompactTxStreamer`. This is a - first-class supported configuration: the hub links `zaino-proto` as its RPC - types (`hub/Cargo.toml:52`), speaks - `/cash.z.wallet.sdk.rpc.CompactTxStreamer/SendTransaction` - (`chain.rs:53`), the product is described as sitting in front of the - operator's existing indexer, and `chain.rs:434` names Zaino explicitly. -2. An attacker submits junk over the unauthenticated Nym submit path. A payload - that does not deserialize is admitted with `txid = None, expiry = None` - (`queue.rs:189-204`, REVIEW #5 — deliberate, because the shim diverts what it - cannot parse). Distinct payloads defeat the payload-hash dedup at - `queue.rs:206-218`. -3. At every flush, `flush` drains the junk, `broadcast_batch` offers it, Zaino's - node refuses it, Zaino answers `INTERNAL`, `chain.rs` says `Retryable`, and - `batcher.rs:390` puts it straight back. -4. The entry is immortal. Not because `expiry == None` — that only matters at - admission — but because **no code path anywhere consults expiry after - admission**. A *parseable but permanently invalid* migration is equally - immortal. -5. Once `inner.bytes` reaches `MAX_QUEUE_BYTES`, `admit` returns - `Refusal::Full` for every genuine migration (`queue.rs:223-225`) **until the - process restarts** — and per the confirmed - `hub-nym-driver-automatic-fresh-identity-…` / failover runbook, restarting the - hub changes its Nym address and strands every shim for "well over an hour". - Each flush's outbound fan-out also grows with the junk (one connection per - transaction per endpoint, `chain.rs:176-210`). - -**This is what the issue owns that its siblings do not:** on a lightwalletd/zebra -backend the same junk is answered with a node error string, classified -`Rejected`, and **dropped at the first flush** — which is why coordinator item -6u(f) could record "junk never reaches the chain and there is nothing to -subtract", and why the confirmed High `hub-queue-unauthenticated-fill-…` requires -the attacker to *sustain* the flood across every epoch. Against Zaino the same -attack is **one-shot and permanent**. - -The non-attacker form needs no attacker at all: any genuinely invalid or expired -migration — a stale anchor, a bad signature, a double-spend — is re-offered to -the operator's indexer at every flush, forever, at one connection and one publish -per endpoint per flush. - -**Attack Requirements and Assumptions:** -- Requires the hub to be pointed at Zaino rather than lightwalletd. Both are - supported and the code names both; the shipped example endpoint - (`deploy.env.example:22-23`, `INDEXER_TLS=na.zec.rocks`) was established by the - sibling issue's validation to be **lightwalletd+zebra** today, with Zaino on - separate `zaino.*.zec.rocks` hosts. So this is a supported-configuration - finding, not a shipped-default one — but the operator has no way to learn it, - because the code asserts the opposite. -- The junk form requires nothing else: hub submission is unauthenticated by - design. -- The mapping was read directly from the vendored Zaino source at - `audit-context/zero/zaino` (monorepo commit `62baea8`), not inferred from - zeronym's comments about it. - -## Impact on Users - -**Genuine migrations are refused, permanently.** Once the byte budget is consumed -by immortal entries, `admit` answers `Full` and — because submit is dispatch-only -on the deployed transport — the shim has already told the wallet `errorCode 0` -with a txid (`shim/src/hub.rs:231-241`, `shim/src/intercept.rs:186`). Per the -threat model that is not merely availability: a wallet that cannot migrate -through the hub is a wallet whose user retries, changes indexer, or broadcasts in -the clear — the exact outcome the product exists to prevent — and with this bug -the attacker chooses when it starts and it does not end without a hub restart -that is itself a fleet-wide outage. - -**The hub re-publishes the same refused bytes to the operator's indexer every 20 -blocks, indefinitely.** For a *genuinely invalid but real* migration — a user's -transaction that failed for a stale anchor, say — that is a repeating, -per-transaction signal delivered to the indexer operator once per cadence for as -long as the process lives, which is precisely the "fresh timing signal tied to -one transaction" that `chain.rs:513-520` says this component exists to avoid -emitting. - -## Technical Details / Code Analysis - -The seam, with its own description of the property it is meant to hold -(`hub/src/chain.rs:475-490`): - -```rust -/// Map a failed `SendTransaction` call (no `SendResponse` came back) onto -/// [`Publish`]. -/// -/// This is the seam between "the indexer judged the transaction" and "the -/// indexer was never really asked", and the batcher's requeue depends on it -/// being drawn honestly. Only INVALID_ARGUMENT and FAILED_PRECONDITION are -/// verdicts here: they are what a gRPC service returns when it read the request -/// and refuses its content. ... -``` - -"they are what a gRPC service returns when it read the request and refuses its -content" is the assumption that fails. Zaino returns `INTERNAL` for a node -refusal because, from Zaino's point of view, the failure came from its upstream -JSON-RPC call, not from the request's arguments. - -Zaino *does* preserve the node's legacy error code — but only in the `source()` -chain, and only its **JSON-RPC** front end walks it -(`zaino-serve/src/rpc/jsonrpc/service.rs:525`, -`sendrawtransaction_error_object_from_indexer_error`, pinned by -`chain_index/tests/mockchain_tests.rs:1293-1327`). The gRPC front end the hub -talks to converts through `Into` and drops it. This matters for -the fix: text-matching the gRPC message is **not** a workaround here, because the -message the hub receives is the constant `"InternalServerError: error receiving -data from backing node"`. - -The batcher end of the seam (`hub/src/batcher.rs:365-390`): - -```rust - for (i, entry) in batch.into_iter().enumerate() { - match outcomes.get(i) { - Some(Publish::Accepted { .. }) | Some(Publish::AlreadyKnown) => achieved += 1, - Some(Publish::Rejected { .. }) => rejected += 1, - Some(Publish::Retryable { reason }) => { - sample_failure.get_or_insert_with(|| reason.clone()); - unplaced.push(entry); - } - None => unplaced.push(entry), - } - } - ... - let requeued = queue.requeue(unplaced); -``` - -and the requeue that has no eviction rule (`hub/src/queue.rs:279-295`): - -```rust - pub fn requeue(&self, entries: Vec) -> usize { - ... - for entry in entries { - if inner.entries.contains_key(&entry.key) { continue; } - inner.bytes = inner.bytes.saturating_add(entry.tx_bytes.len()); - inner.entries.insert(entry.key, entry); - reinserted += 1; - } -``` - -`best_of`'s doc (`chain.rs:447-458`) makes the *absence* of this bug its own -premise: - -> if an unreachable endpoint could outvote a live one's verdict, a single dead -> endpoint in the list would keep every doomed transaction resident until it -> expired, and **an unparseable payload never expires** (queue.rs, REVIEW #5): -> junk plus one dead endpoint would fill the byte budget for everyone. - -Against Zaino that safeguard never engages, because the verdict it protects never -exists — and the failure mode it describes is exactly what happens, without -needing a dead endpoint. - -**A second, opposite defect in the same two lines, recorded as latent.** -`GRPC_FAILED_PRECONDITION` ("9") is classified as a permanent verdict, for which -the migration is *destroyed*. That reads the gRPC contract backwards: -FAILED_PRECONDITION denotes "the system is not in a state required for the -operation", i.e. a condition that may clear, whereas INVALID_ARGUMENT is about -the arguments themselves. Zaino uses code 9 for exactly such transient service -state — `tonic::Status::failed_precondition("zaino not yet synced")` -(`error.rs:163-165`) and `ChainIndexErrorKind::InvalidSnapshot` -(`error.rs:134-136`). Every producer of `UnavailableNotSyncedEnough` was checked -(`node_backed_indexer.rs:1051, 1089, 1175, 1544`) and **none is on the -`send_transaction` path**, so this is not live against Zaino today. It is -recorded because it is a real inversion that any front end or future backend -answering `9` for a transient condition would turn into permanent destruction of -a whole batch of migrations whose wallets were already told `error_code 0`. - -Neither behaviour has any test coverage against a real indexer: -`hub/tests/live_chain.rs` is entirely `#[ignore]`d, gated on `ZIH_TEST_INDEXER`, -and exercises only `tip_height`; the in-module tests at `chain.rs:535-622` assert -the mapping the code implements rather than the mapping the backends produce; and -per coordinator item 6n no CI runs `cargo test` at all. - -## Recommendations - -1. **Give the queue a GC path, independently of the classification.** This is the - fix that closes the issue whatever the backend does: no entry may be immortal. - `Entry::received_height` is already stored and already documented as *"drives - the confirmation deadline"* — read it. A bounded requeue count, or a - `received_height + kN` deadline, satisfies `REVIEW.md`'s existing requirement - that `expiry == None` entries need a deadline as their only GC path. -2. **Do not infer "the node judged this transaction" from a gRPC status code - alone, and do not rely on the message either.** Against Zaino the message is a - constant. If a verdict is wanted from Zaino, the honest options are to use its - JSON-RPC front end (which does preserve the node's code), or to ask upstream - for a status/message that distinguishes "the node refused these bytes" from - "zaino could not reach its node" — a change worth requesting regardless, since - any gRPC client of Zaino has this problem. -3. **Correct the doc comment at `chain.rs:434-437`:** Zaino follows lightwalletd's - *success* convention only; it reports every failure as a gRPC status, so - `classify_send_response`'s non-zero branch and `classify_publish_error` are - dead code behind a Zaino backend. -4. **Reclassify `FAILED_PRECONDITION` as retryable**, matching the gRPC contract - and Zaino's actual usage of it. -5. **Add an integration test that runs a flush against each supported backend and - asserts the disposition of an invalid transaction.** The three sibling issues - in this function exist because no such test does. - -## Validation Information - -**Verdict: CONFIRMED. Severity: Medium (as filed).** - -### The backend behaviour was re-derived from the vendored source, end to end - -Every hop was read in `audit-context/zero/zaino` (monorepo commit `62baea8`), -matching the standard set by the zebra sibling's validation: - -1. **`SendResponse` has exactly one construction site in Zaino outside the - generated proto crate.** `grep -rn "SendResponse" packages/ --exclude - zaino-proto` returns seven hits: two `use` lines, the trait declaration - (`indexer.rs:696-699`), the macro entry (`zaino-serve/.../grpc/service.rs:177`), - and the single implementation at `node_backed_indexer.rs:1562-1570`, which - hardcodes `error_code: 0`. **`classify_send_response`'s non-zero branch is - therefore unreachable behind Zaino** — which is the same partition the zebra - sibling's validation recorded from the other side. -2. **The error path.** `node_backed_indexer.rs:538-546` → `?` → - `NodeBackedChainIndexSubscriber::send_raw_transaction` - (`chain_index.rs:2912-2922`, `type Error = ChainIndexError` via - `ChainIndexRpcExt: ChainIndex`, `chain_index.rs:547`, `:1660-1662`) → - `.map_err(ChainIndexError::backing_validator)` → `error.rs:525-531`: - `kind = InternalServerError`, `message = "InternalServerError: error receiving - data from backing node"`, real cause in `source`. -3. **The conversion.** `error.rs:118-137`: - `NodeBackedIndexerServiceError::ChainIndexError(err) => match err.kind { - InternalServerError => tonic::Status::internal(err.message), … }`. **Status 13, - constant message, `source` discarded.** Every other variant in that impl is - also `internal(...)` except two `failed_precondition` cases. -4. **The handler.** `zaino-serve/src/rpc/grpc/service.rs:25-40` (`client_method_ - helper!`) does `.map_err(Into::into)?` and `:159-177` binds - `send_transaction`, so the status reaches the wire unmodified. -5. **The hub's receiving end.** `chain.rs:356-401` reads `grpc-status` from a - trailers-only response *and* from trailers after a body, so a tonic error - status is correctly captured as `GrpcStatusError { code: "13", … }`; - `classify_publish_failure` (`:491-501`) sends it to the `_` arm; the in-file - test at `:589-607` pins `"13"` as retryable. -6. **`FAILED_PRECONDITION` reachability was checked, not assumed.** All four - producers of `UnavailableNotSyncedEnough` are on `get_latest_block`, - `get_block`-family and `get_transaction` paths, never `send_transaction`; and - `InvalidSnapshot` is `#[allow(dead_code)]` and never constructed - (`error.rs:483-484`). The filed "latent inversion" framing is correct and has - been kept as such — it carries **no** severity here. - -### The eviction claim was checked exhaustively, and is stronger than filed - -The filing said the entry is immortal "because `expiry == None`". That reasoning -is wrong and has been corrected: `survives_next_flush` is consulted **only** at -`queue.rs:202`, inside `admit`. `requeue` (`:279-295`) has no expiry test, `flush` -(`batcher.rs:341-390`) has no expiry test, and `Entry::received_height` — whose -own doc comment says *"Drives the confirmation deadline"* — is read **nowhere** -in the crate (`grep -rn received_height hub/src/` returns exactly `queue.rs:135` -and `:235`). So against Zaino a **parseable, expired, permanently invalid** -migration is just as immortal as junk. The queue's only eviction is a terminal -verdict, and the backend that never produces one has no GC at all. That is why -recommendation 1 is now first: it is the backend-independent fix. - -### Ownership — the three issues in this function partition cleanly - -The zebra sibling's validation already recorded the partition from its side and -this pass confirms it from Zaino's: - -| issue | backend | direction of the error | -|---|---|---| -| `publish-verdict-strings-are-zcashds-vocabulary-only-…` (**confirmed, Medium**) | lightwalletd + zebrad | zebra's strings match none of the four substrings, so *"never judged"* answers become `Rejected` → **destroyed** | -| `hub-chain-duplicate-nullifier-rejection-counted-as-published.md` (plausible, Low) | zcashd | `"duplicate"` over-matches a double-spend, so a refusal is counted `AlreadyKnown` → **reported as achieved** | -| **this issue** | **Zaino** | no answer is ever a verdict, so a refusal becomes `Retryable` → **immortal** | - -**This issue owns the Zaino backend, and nothing else.** It must not be graded as -if it also covered the shipped lightwalletd+zebra endpoint, and the report should -present all three together as *"the function is wrong for every backend, in a -different direction each time"*. - -### Severity: why Medium - -*Impact, given Zaino:* one unauthenticated 64 MiB fill becomes **permanent** -rather than sustained, so the confirmed High queue-fill attack gets strictly -cheaper and its recovery becomes a hub restart, which is itself the fleet-kill -outage the runbook budgets at "well over an hour". Plus a no-attacker path (any -invalid migration re-broadcast forever) that leaks a per-transaction cadence -signal to the indexer operator. - -*Likelihood:* bounded by the backend. Zaino is a supported, code-named, -same-monorepo indexer and the hub links its proto crate, but the shipped example -endpoint is lightwalletd+zebra (established by the sibling's validation from -zec.rocks' own published stack and operator statements). An operator choosing -Zaino has no way to discover the problem: `chain.rs:434` asserts the convention -holds, and no test in the repository exercises any real backend. - -*Why not High:* it needs a configuration choice the shipped example does not -make, and the terminal harm — fleet-wide silent destruction after `Refusal::Full` -— is already owned at High by `hub-queue-unauthenticated-fill-silently-destroys- -migrations.md`. Item 6u(b)'s "do not stack the severities" applies: what is -graded here is the *escalation* (sustained → permanent) and the no-attacker GC -failure, not the terminal state. - -*Why not Low:* the defect is in the one function that decides whether a -wallet-acknowledged migration is kept or destroyed; it disables the queue's only -garbage-collection mechanism entirely; it is invisible to every test and every -telemetry surface the project has; and it is asserted to be safe by a comment in -the code that is factually wrong about the backend it names. - -DO NOT DEVIATE FROM THIS TEMPLATE (`docs/SECURITY-ISSUE-TEMPLATE.md`). -LEAVE THIS MESSAGE IN PLACE SO THAT ALL AGENTS KNOW NOT TO DEVIATE FROM THIS TEMPLATE (`docs/SECURITY-ISSUE-TEMPLATE.md`). diff --git a/zeronym-22aa9851caf68-high-medium/medium/hub-flush-destroys-migration-on-single-unverifiable-verdict.md b/zeronym-22aa9851caf68-high-medium/medium/hub-flush-destroys-migration-on-single-unverifiable-verdict.md deleted file mode 100644 index fb447a94..00000000 --- a/zeronym-22aa9851caf68-high-medium/medium/hub-flush-destroys-migration-on-single-unverifiable-verdict.md +++ /dev/null @@ -1,390 +0,0 @@ -# `flush` treats one endpoint's unverifiable `Rejected` as final and deletes the last copy of the migration, so a hostile or buggy indexer permanently destroys a transaction the wallet was already told had been sent - -**Severity**: Medium -**Validation Status**: Confirmed -**Location**: `audit-target/zeronym/hub/src/batcher.rs:361-390` (the requeue-vs-drop split in `flush`, the `Rejected` arm at `:368`, the rationale at `:379-389`) and `:333-340` (the claim about what the returned count measures); verdicts produced at `audit-target/zeronym/hub/src/chain.rs:438-445` (`classify_send_response`), `:446-474` (`best_of`), `:476-497` (`classify_publish_failure`), `:503-533` (`classify_publish_error`); the queue side at `audit-target/zeronym/hub/src/queue.rs:252-264` (`drain_shuffled`) and `:279-295` (`requeue`); the acknowledgement boundary at `audit-target/zeronym/shim/src/hub.rs:225-249`. Deployed endpoint count: `audit-target/zeronym/deploy.env.example:22`. -**Found by agent:** Local (file audit of `hub/src/batcher.rs`) -**In scope of audit?** Yes - -## Description - -`flush` splits the batch three ways after `broadcast_batch` returns -(`batcher.rs:361-390`): - -* `Accepted` / `AlreadyKnown` — counted as achieved and **the entry is dropped**; -* `Rejected` — counted as rejected and **the entry is dropped**; -* `Retryable` / no verdict — the entry is requeued for the next cadence flush. - -Only the third arm keeps a migration alive. Every one of those verdicts is -derived from a string or a gRPC status the indexer chose, and the hub verifies -none of it. With the shipped configuration there is exactly one indexer -(`deploy.env.example:22`, `INDEXERS=66.241.124.200:443`), so `best_of` -(`chain.rs:459-473`) folds a single endpoint's answer and there is no second -opinion. The threat model for this engagement names that party explicitly: -*"hub → indexer: sees the whole batch seconds before it is public; **can lie -about the tip and about publish verdicts**"*. - -**One `Rejected` answer destroys a migration, unrecoverably, at the first -flush.** The entry is removed from the queue by `drain_shuffled` -(`queue.rs:252-264`, which drains the whole map) and is simply not put back, so -its `Zeroizing` buffer is dropped. No copy exists anywhere else in the system: -the shim answered the wallet `error_code 0` at mixnet hand-off and keeps no -per-migration state (`shim/src/hub.rs:230-240`), and the enclave is diskless. -`flush`'s own comment (`batcher.rs:379-389`) explains why a *transport* failure -must be requeued — "once the entry left this queue there is no other copy -anywhere that anyone will retry" — and then applies the opposite rule to a -verdict it cannot verify. `chain.rs:487-490` states the same principle even more -sharply, for the transport case only: - -> re-offering an entry costs one call per flush and stops at its expiry, while -> **dropping a valid migration on a misread error is unrecoverable, because the -> shim has already told the wallet it was sent**. - -Two properties make this worse than "the node said no": - -1. **It is not fixed by adding endpoints, and gets a new failure mode from - them.** `Rejected` (rank 1) outranks `Retryable` (rank 0) in `best_of`, and - `Retryable` is the only outcome that keeps an entry alive. So during any - outage of the honest endpoints — precisely when the requeue is the mechanism - protecting the entry — one endpoint answering `Rejected` outranks their - `Retryable` and the entry is destroyed. The publish path is n-of-n for - *delivery* and **1-of-n for the retry**. The confirmed - `tip-and-verdict-aggregation-scale-in-opposite-directions-…` names this - "recovery veto" and delegates it here by name. -2. **It reaches the population the product exists for.** The other route to a - `Rejected` drop — holding a transaction back until it expires, owned by the - confirmed `indexer-chooses-which-batch-members-reach-the-chain-…` — works only - against short-expiry traffic; that issue states plainly that ZIP 318 - migrations, with 34,561–69,120 blocks of slack, "survive indefinitely" under - it. A direct `Rejected` answer destroys them on the first flush. - -A second, smaller defect sits in the same function. `batcher.rs:335-336` calls -the returned count *"the honest measure of the privacy the flush actually -delivered"*, and `REVIEW.md` #9 makes the distribution of that number the -**launch gate** for the whole product. It is not a measure of what reached the -network; it is a tally of what the indexer *said*. Nothing ever checks the chain -afterwards — the confirmation tracking `REVIEW.md` #7 specifies is a documented -not-built item — and `flush` does not even compare the txid the endpoint returns -in `Publish::Accepted { txid }` against `Entry::txid`, which the hub computed -from the same bytes at admission (`queue.rs:125`, `Publish::Accepted`'s payload -discarded at `batcher.rs:367`). - -The design's stated position (`REVIEW.md` #5) is that "`sendrawtransaction` at the -node is the only authority on validity". The implementation extends that authority -from a *node* to a third-party *indexer relay* in front of it — `chain.rs:17-21` -acknowledges the difference ("an indexer is a single funnel in front of a single -node") — and that relay is a member of the adversary class the product exists to -defend against. - -## Attack Scenario and Steps - -Attacker: the operator of the hub's indexer, or anyone who compromises or compels -it. It sees every batch member in plaintext seconds before publication (a stated -residual), so it can select by content — a value balance, an action count, a -length, an `anchorOrchard`, or a txid supplied out of band. - -Targeted destruction: - -1. The hub flushes a batch; the indexer receives all `k` members concurrently - over `SendTransaction`, each on its own connection with a 10 s budget - (`chain.rs:198-210`, `:269-337`), so it holds the whole batch before it has to - answer any of it. -2. For the chosen member, the indexer answers gRPC `INVALID_ARGUMENT` (status 3), - or `OK` carrying `SendResponse { error_code: -26, error_message: "16: - bad-txns-orchard-binding-signature-invalid" }`. The first maps to - `Publish::Rejected` at `chain.rs:491-496`; the second at `chain.rs:513-533`, - because it matches none of the four `AlreadyKnown` substrings. -3. `best_of` has one outcome to fold, so the transaction's verdict is `Rejected`. -4. `flush` increments `rejected` and does **not** put the entry in `unplaced`, so - `queue.requeue` never sees it (`batcher.rs:367-390`). The entry's `Zeroizing` - buffer is dropped when the loop iteration ends. -5. It broadcasts every other member normally, so the hub's log line reads - `flush_size = k, achieved_batch_size = k-1, rejected = 1` — a single count with - no identifier, indistinguishable from one genuinely invalid transaction, and - written to a console that does not exist in an attested deployment. - -The wallet is never told. A resend by the user *would* be admitted (dedup is on -`sha256(tx_bytes)` and the entry is gone, so it is a fresh admission, and -`queue.rs`'s dedup makes a resend safe) — but nothing produces a -non-confirmation signal to prompt one, and a resend meets the same endpoint and -the same answer. - -**Attack Requirements and Assumptions:** - -- Requires control of, or a bug in, an indexer the hub publishes through — the - shipped configuration has exactly one. **No mixnet access, no shim, and no - chain observation are needed, and there is no internet-reachable path to this - behaviour** (coordinator item 6p: the `Rejected` drop is a hub-trust / - robustness defect, not a remote weapon). -- The same outcome arises without an adversary. An indexer that returns - `INVALID_ARGUMENT` for a transaction version it does not recognise — a live - concern across an NU6.3/v6 rollout — causes the hub to destroy every such - migration rather than hold it. The parallel defect on the *other* verdict path, - where the shipped example backend (lightwalletd + zebrad) produces text that no - `AlreadyKnown` substring matches and "not judged" becomes destruction, is - separately confirmed as - `publish-verdict-strings-are-zcashds-vocabulary-only-so-a-zebrad-backed-indexer-turns-not-judged-into-permanent-destruction.md` - and is not re-claimed here. -- Users cannot detect it: the wallet was told success, and the hub reports counts - only, to a console that does not exist under `debug { enabled = false }`. - -## Impact on Users - -- A migration the wallet reported as sent never reaches the network and is never - retried by anything. The user's value stays in a pool NU6.3 closed to new - value, and the only automatic clock that can ever surface the failure is the - transaction's own expiry — roughly 50 minutes for an ordinary Orchard spend and - **30–60 days for the ZIP 318 migration the product exists for** - (`zip318-canonical-expiry-is-the-only-recovery-clock-…`). Non-confirmation - monitoring is asked of wallet vendors in `hub/deploy/caution/OPERATORS.md:190-195` - and is guaranteed by nothing in this system. -- This is loss of a *submission*, not of funds: the note is not spent on chain and - the user can build and send a replacement once they notice. Against a hostile - endpoint the replacement is destroyed the same way, so the practical outcome is - targeted, indefinite, silent censorship of one user's migration through the - private path. -- Separately, the number `REVIEW.md` #9 designates as the launch gate for the - product's privacy claim is a tally of the measured party's own assertions. - -> **CORRECTION 2026-08-18 (validation of the cited file — SUPERSEDES the sentence above).** -> The wallet does **not** wait for expiry. Both official Zcash light-wallet SDKs -> automatically resubmit a sent-but-unmined transaction for as long as it remains -> unexpired — the Android SDK at the head of every ~20 s sync loop and after every -> processed block batch (`CompactBlockProcessor.kt:573,615,723`; selection -> `mined_height IS NULL AND expiry_height > ?`), the iOS SDK at most once per 300 s -> (`TxResubmitter.swift:8-15`, `TransactionDao.swift:218-228`) — and the hub's -> payload-hash dedup makes the resend free. The wallet's non-confirmation signal comes -> from compact-block scanning, which the shim does not intercept (`proxy.rs:1068-1074`). -> Expiry is therefore the **retry horizon**, not the wait: ~50 minutes for the ZIP 203 -> default traffic the shim also diverts, 30–60 days for a ZIP 318 migration. A -> *transient* loss self-heals within minutes; only a loss condition that **outlives the -> horizon** destroys the submission permanently — which is exactly what this issue's -> condition does, so this issue's severity is unaffected. Do not write "the user waits -> 30 to 60 days" in the report. Full refutation and the replacement paragraph: -> `issues/invalid/zip318-canonical-expiry-is-the-only-recovery-clock-and-a-lost-migration-freezes-the-users-notes-for-30-to-60-days.md`. - - -## Technical Details / Code Analysis - -`hub/src/batcher.rs:361-390` — the split, in full: - -```rust - let mut achieved = 0usize; - let mut rejected = 0usize; - let mut sample_failure: Option = None; - let mut unplaced = Vec::new(); - for (i, entry) in batch.into_iter().enumerate() { - match outcomes.get(i) { - Some(Publish::Accepted { .. }) | Some(Publish::AlreadyKnown) => achieved += 1, - Some(Publish::Rejected { .. }) => rejected += 1, // entry falls out of scope here - Some(Publish::Retryable { reason }) => { - sample_failure.get_or_insert_with(|| reason.clone()); - unplaced.push(entry); - } - None => unplaced.push(entry), - } - } - ... - let requeued = queue.requeue(unplaced); -``` - -and the comment that decides it (`batcher.rs:386-389`): - -```rust - // A Rejected verdict is not put back: the node said no, and re-offering the - // same bytes buys the same answer every flush until expiry. -``` - -That reasoning is sound for an honest node and unsound for the party actually -answering, which is an indexer relay the threat model classes as able to lie. - -`hub/src/chain.rs:459-473` — with one endpoint, `best_of` is the identity function -on that endpoint's verdict, and with several it is a 1-of-n veto on the retry: - -```rust -fn best_of(outcomes: Vec) -> Publish { - fn rank(outcome: &Publish) -> u8 { - match outcome { - Publish::Accepted { .. } => 3, - Publish::AlreadyKnown => 2, - Publish::Rejected { .. } => 1, - Publish::Retryable { .. } => 0, - } - } - outcomes.into_iter().max_by_key(rank).unwrap_or(Publish::Rejected { … }) -} -``` - -Its doc comment (`chain.rs:446-458`) argues the ordering carefully — a dead -endpoint must not keep a doomed transaction resident until expiry, and an -unparseable payload never expires — and closes with *"A verdict from any endpoint -that answered is final, exactly as it is today with one endpoint."* The argument -is correct for the case it considers and never considers a hostile answer. - -`hub/src/chain.rs:491-496` and `:513-533` — the two roads to `Rejected`, one a -status code the endpoint picks, one a free-text string it writes: - -```rust -fn classify_publish_failure(err: &BoxError) -> Publish { - let reason = err.to_string(); - match err.downcast_ref::() { - Some(status) - if status.code == GRPC_INVALID_ARGUMENT || status.code == GRPC_FAILED_PRECONDITION => - { - Publish::Rejected { reason } - } - _ => Publish::Retryable { reason }, - } -} - -fn classify_publish_error(message: &str) -> Publish { - let m = message.to_ascii_lowercase().replace('-', " "); - if m.contains("already in block chain") || m.contains("already known") - || m.contains("already in mempool") || m.contains("duplicate") { - Publish::AlreadyKnown - } else { - Publish::Rejected { reason: message.to_string() } - } -} -``` - -`hub/src/queue.rs:135` declares `received_height` as "The tip when this was -admitted. **Drives the confirmation deadline.**" A repository-wide search shows -the field is written at `queue.rs:235` and **read nowhere** in `hub/src` or -`hub/tests`: there is no confirmation deadline and no awaiting-confirmation set of -the kind `REVIEW.md`'s implementation rules describe (`REVIEW.md:145`). The -absence of confirmation tracking is a documented limitation and is not the -finding; the dead field asserting otherwise, and the decision to delete the last -copy of an entry in its absence, are. (The same dead field is noted in the -validation of `hub-chain-zaino-node-rejections-are-never-verdicts.md`; it belongs -to the family, not to this file alone.) - -## Recommendations - -- **Do not treat a single unverifiable verdict as final.** Either requeue - `Rejected` entries until their expiry, or require corroboration from at least - two independent endpoints before destroying an entry. The bounded cost is one - call per flush per held entry, which is what `chain.rs:487-490` already accepts - for the transport case. -- **Sequence this with a GC path, or it regresses into the opposite defect.** - Retaining rejected entries is only safe once the queue can evict on its own: - `survives_next_flush` is consulted **only** in `admit` (`queue.rs:202`), so - `requeue` and `flush` apply no expiry test, and an unparseable payload has - `expiry = None` and would become immortal. `REVIEW.md:145` already specifies the - mechanism (`received_height + 2N`, its own byte budget), and - `hub-chain-zaino-node-rejections-are-never-verdicts.md` asks for the same GC - path from the other direction. **Implement the GC first, then the retention.** -- **Note for the report's remediation ordering:** the standing advice "deploy more - than one independent indexer" — correct for the hold-back and blackhole attacks - — does **not** fix this direction and introduces the 1-of-n recovery veto. Ship - the two together, and see - `hub-one-tls-name-for-the-whole-indexer-list-…` for why "independent" is not - expressible in today's configuration surface. -- **Cross-check the txid the endpoint returns** against `Entry::txid`, which the - hub already computed from the same bytes. It is free and it catches a lazy or - broken endpoint; it does not defeat a competent liar, which is why it is a - sanity check and not the fix. -- Log rejections with enough aggregate detail for an operator to notice a pattern - (a per-flush rejection *rate*, an `observed vs expected` counter), while keeping - to the counts-only rule — and give it an egress path, since a `tracing` line - reaches no console in an attested enclave. -- Either implement the confirmation check `REVIEW.md` #7 specifies, or change - `batcher.rs:335-336` to say that `achieved_batch_size` measures what the indexer - reported rather than "the privacy the flush actually delivered". -- Remove or implement `Entry::received_height`'s "drives the confirmation - deadline" claim. - -## Validation Information - -**Verdict: CONFIRMED at Medium** (severity as filed). - -**Every mechanical claim was re-derived in the target during validation:** - -- `batcher.rs:361-390` — `unplaced` is fed by the `Retryable` and `None` arms - only; the `Rejected` arm at `:368` increments a counter and lets `entry` fall - out of scope. `queue.requeue` at `:390` therefore never sees it. -- `queue.rs:252-264` — `drain_shuffled` drains the entire map and zeroes `bytes`, - so an entry not returned by `requeue` is gone from the process. -- `shim/src/hub.rs:230-240` — on the deployed mixnet transport a successful - hand-off is reported to the wallet as `Submit::Accepted { txid: local_txid(...) }`, - with an explicit comment that the hub's verdict is "deliberately not waited - for". The shim retains nothing. -- `chain.rs:459-473` — `rank` is `Accepted 3 / AlreadyKnown 2 / Rejected 1 / - Retryable 0`, so `Rejected` beats `Retryable` at any `n`; `max_by_key` on a - one-element vector is the identity. -- `chain.rs:491-496` — status 3 and status 9 are the only codes mapped to - `Rejected`; `chain.rs:513-533` — the four substrings are the only escapes from - `Rejected` on the `OK`-with-non-zero-code path. -- `deploy.env.example:22-23` — one endpoint, `INDEXER_TLS=na.zec.rocks`. -- `received_height`: `grep -rn received_height hub/ shim/` returns exactly - `queue.rs:135` (declaration, with the doc comment) and `queue.rs:235` (write), - plus two `REVIEW.md` lines. It is read nowhere. -- `Publish::Accepted { txid }` is discarded at `batcher.rs:367`; `Entry::txid` - exists at `queue.rs:125`. No comparison is made anywhere. - -**Corrections applied against the filing:** - -1. **"the migration ceases to exist" was softened.** What is destroyed is the - hub's only copy and any prospect of the system retrying it. The wallet still - holds the transaction and a user-initiated resend is admissible and safe (item - 7r established the same point for the runbook issue). The real defect is that - nothing in *this system* produces a signal — but see the CORRECTION above: - the wallet's own sync does, and both official SDKs resend automatically for the - whole expiry window, so what this defect actually destroys is every submission - made while the condition holds, not a single one. -2. **Loss-of-funds framing removed**, per the precedent set in item 7o: the note - is not spent on chain, so this is silent destruction of a *submission*, not loss - of funds. **The "30–60 day recovery clock" half of this sentence was REFUTED - 2026-08-18** (PROGRESS item 8a): 30–60 days is the wallet's automatic-**retry - horizon**, during which it keeps resubmitting, not a period of waiting. -3. **The "silent blackhole" leg was demoted to evidence and its harm delegated.** - An indexer answering `error_code: 0` for everything while broadcasting nothing - is the *blackhole* attack, which the confirmed - `indexer-chooses-which-batch-members-reach-the-chain-…` owns and which - additional endpoints genuinely fix (an honest endpoint's real relay makes the - lie irrelevant). What survives here is narrower and unowned elsewhere: the - count `REVIEW.md` #9 makes the launch gate is a tally of the indexer's - assertions, the code comment says the opposite, and a free txid cross-check is - not performed. -4. **The "ordinary bug" leg was scoped.** `classify_publish_error`'s - "anything unrecognised is a rejection" default against a real backend's real - vocabulary is the confirmed zebrad Medium; this file keeps only the - `classify_publish_failure` (status 3) instance and cross-references it. -5. Recommendation 1 now carries the **ordering constraint** (GC before retention) - that the Zaino sibling makes necessary, and the "deploy more endpoints" advice - is explicitly qualified rather than repeated unqualified as in the filing. - -**What this issue owns after reconciliation with its four siblings — checked file -by file, because the risk of double-counting here is high:** - -| Sibling | Direction | What it owns | -|---|---|---| -| `publish-verdict-strings-are-zcashds-vocabulary-only-…` (confirmed Medium) | zebrad: "not judged" → `Rejected` | the **vocabulary** of `classify_publish_error` against one real backend | -| `hub-chain-zaino-node-rejections-are-never-verdicts.md` (confirmed Medium) | Zaino: every refusal → `Retryable` | the **immortal-entry** direction and the GC ask | -| `hub-chain-duplicate-nullifier-rejection-counted-as-published.md` (plausible) | zcashd: double-spend → `AlreadyKnown` | the **false-success** direction | -| `indexer-chooses-which-batch-members-reach-the-chain-…` (confirmed) | `Retryable` → hold-back | **anonymity**: which members reach the chain, and destruction *via expiry* for short-expiry traffic only | -| **this file** | `Rejected` → immediate, terminal drop | the **decision that a single unverifiable verdict is final**, its 1-of-n recovery-veto form at `n > 1`, immediate destruction of the **ZIP 318 population** (which the hold-back route explicitly cannot reach), and `achieved_batch_size`'s misdescription | - -**It is not subsumed, and the audit's own accounting already depends on it.** The -confirmed `tip-and-verdict-aggregation-scale-in-opposite-directions-…` was -deflated to Low with an ownership map that assigns "One unverifiable `Rejected` -destroys an entry" **to this file by name**, and notes that this file's -recommendation is also the fix for the `n > 1` retry-veto form. Merging it away -would orphan that harm and leave the recovery-veto unowned. - -**Severity justification — Medium.** -*Why not High:* item 6p's bound applies exactly as it does to the four confirmed -tip findings — the attacker must control a configured `ZIH_INDEXERS` endpoint, so -this is a hub-trust and robustness defect, not an internet-reachable weapon; the -harm is loss of a submission with a recovery path the user can take once they -notice; and the accidental instances against real backends are separately graded. -*Why not Low:* it destroys wallet-acknowledged migrations belonging to the exact -traffic class the product exists for, on the first flush, chosen per victim from -plaintext the endpoint is handed by design; it is undetectable by the wallet, the -user and (in an attested enclave) the operator; the shipped configuration makes -one party's word final; and the code's own stated principle — never drop a valid -migration on an unverifiable error — is applied to the transport path and not to -the verdict path four lines away. - -DO NOT DEVIATE FROM THIS TEMPLATE (`docs/SECURITY-ISSUE-TEMPLATE.md`). -LEAVE THIS MESSAGE IN PLACE SO THAT ALL AGENTS KNOW NOT TO DEVIATE FROM THIS TEMPLATE (`docs/SECURITY-ISSUE-TEMPLATE.md`). diff --git a/zeronym-22aa9851caf68-high-medium/medium/hub-indexer-tls-is-optional-in-code-required-in-every-document-and-unobservable-when-absent.md b/zeronym-22aa9851caf68-high-medium/medium/hub-indexer-tls-is-optional-in-code-required-in-every-document-and-unobservable-when-absent.md deleted file mode 100644 index 2bb21bd2..00000000 --- a/zeronym-22aa9851caf68-high-medium/medium/hub-indexer-tls-is-optional-in-code-required-in-every-document-and-unobservable-when-absent.md +++ /dev/null @@ -1,591 +0,0 @@ -# The hub's indexer hop is authenticated only if an environment variable happens to be set: unset `ZIH_INDEXER_TLS` fails OPEN into plaintext h2c under a comment that says the hub refuses to run, and when it *is* set nothing in the attested binary constrains what it is set to - -**Severity**: Medium -**Validation Status**: Confirmed -**Location**: `audit-target/zeronym/hub/src/config.rs:47-53` (the field), `:100-108` (`Config::indexer_tls`), `:142-144` (the unit test that pins the optionality); `audit-target/zeronym/hub/src/main.rs:30-39` (the warn-and-continue, and the comment that says otherwise) versus `:41-45` (the invariant that *does* abort startup); `audit-target/zeronym/hub/src/chain.rs:124-128` (`tls: Option>`), `:316-334` (the live plaintext branch), `:438-445` (`classify_send_response`), `:459-467` (`best_of`), `:212-266` (relayed lookups); `audit-target/zeronym/hub/src/batcher.rs:342` (`drain_shuffled`), `:365-390` (an `Accepted` verdict consumes the entry); `audit-target/zeronym/hub/src/tls.rs:1-26` (the "why this is not optional" module doc) and `:59-84` (`IndexerTls::new` — no pin, no allowlist). Enforcement that exists only *outside* the attested binary: `audit-target/zeronym/deploy.sh:115`, `audit-target/zeronym/hub/deploy/caution/assemble-caution.sh:126-131`, `audit-target/zeronym/deploy.env.example:23`. Claims bearing on it: `audit-target/zeronym/README.md:77`, `audit-target/zeronym/hub/src/lib.rs:29`, `audit-target/zeronym/hub/deploy/caution/caution.hcl.tmpl:128-132`, `audit-target/zeronym/hub/deploy/caution/OPERATORS.md:73-76` and `:282`, and — stating the opposite and false — `audit-target/zeronym/hub/deploy/Containerfile:166-167` and `audit-target/zeronym/hub/deploy/README.md:48` -**Found by agent:** Local (two concurrent file audits: `hub/src/config.rs` and `hub/src/tls.rs`), merged into one issue by the Issue Validator on 2026-08-18 -**In scope of audit?** Yes — operator-supplied configuration and environment (`ZIH_*`) is a declared trust boundary in `audit-context/AUDIT-INSTRUCTIONS.md`, priority area 5 is fail-closed discipline, and markdown/comment claims are in scope as security claims under ICTM - -> **Note on the filename.** This file keeps its original name so the twenty -> cross-references elsewhere in `audit-state/` stay valid. Two phrases in that -> filename are **legacy artefacts and are wrong**: the hop is *not* "required in -> every document" (two build documents deny the TLS stack exists at all — see -> Description §5), and the absent state is *not* "unobservable" (it is measured -> into PCR0/PCR1 and served at `.manifest.run_command` — see Description §4). The -> title above is the corrected statement of the finding. This file also **merges** -> `hub-indexer-tls-is-optional-operator-chosen-and-unobservable.md`, filed -> concurrently by the `hub/src/tls.rs` audit, which is retained under `invalid/` -> for bookkeeping only. - -## Description - -The hub's outbound hop to the indexer is the channel by which every migration the -hub holds leaves the enclave on its way to the network. Whether that channel is -authenticated at all is decided by the *presence* of one environment variable, -and the absent case fails **open**. - -**1. Unset means no authentication, and the process keeps running.** -`indexer_tls` is an `Option` with no default (`config.rs:47-53`). -`Config::indexer_tls()` maps `None` to `Ok(None)` (`:100-108`), `main` emits a -`tracing::warn!` and continues (`main.rs:34-38`), and `ChainClient` carries a -complete, live plaintext branch for that case (`chain.rs:316-334`). "Plaintext" -understates it: with `tls: None` no certificate is requested, presented or -checked, so the hop has **no peer authentication of any kind**. Whoever is on the -path *is* the indexer, as far as the hub can tell. - -The comment directly above says the opposite of what the code does -(`main.rs:30-32`): *"Refuse to run blind."* Nothing refuses. There is no -`Config::validate()` in the crate. Five lines further down, `main` demonstrates -the discipline it did not apply here, and applies it to a set of four -compile-time constants that no operator can change (`main.rs:41-45`, -`params.validate()?`). Verified for this audit: `BatchParams` is constructed only -via `Default::default()`, and `grep -rn "std::env\|env::var" hub/src` returns -exactly one hit (`nym_driver.rs:202`), so that startup invariant genuinely is not -operator-settable — the hub hard-fails at boot for constants and warn-continues -on the one operator-settable value that decides whether its entire egress is -authenticated. - -The same struct applies the opposite discipline to its two booleans eight lines -away: `--nym` and `--http-submit` take an explicit `true`/`false` -(`config.rs:59-65`, `:93-97`) precisely *because* "an environment variable's mere -presence is a bad way to express a security-relevant choice". That care is spent -on the two flags whose worst-case misreading is a mixnet client that does not -start and an endpoint that 404s. It is not spent on the one variable whose -**absence** is the dangerous state. A unit test pins the optionality as intended -behaviour (`config.rs:142-144`, `fn tls_is_optional_but_a_bad_name_is_refused`, -first assertion `assert!(cfg.indexer_tls().expect("no tls configured").is_none())`). - -**2. The harm is integrity, not confidentiality.** The confidentiality loss is -genuinely bounded — the batch is public in a mempool seconds later — and this -issue does not lead with it. What is unbounded is what an unauthenticated peer -can *do*: - -- **Destroy every migration in the batch, silently.** A forged - `SendResponse { error_code: 0 }` is enough. `classify_send_response` - (`chain.rs:438-445`) returns `Accepted { txid: resp.error_message }` with **no - check that the string is a txid, or is a hash, or is anything at all**; - `best_of` ranks `Accepted` highest (`chain.rs:459-467`); `flush` counts it - achieved and does **not** requeue it (`batcher.rs:365-390`). The entry was - already removed from the queue by `drain_shuffled` (`batcher.rs:342`), the - queue is RAM-only, and the shim answered the wallet `error_code 0` at mixnet - hand-off (`shim/src/hub.rs:232-240`), keeping no copy. There is no copy - anywhere and nothing retries. -- **Or, worse, capture and re-publish at a chosen moment.** The forged rejection - is not the strongest move. An on-path party holds valid, signed transaction - bytes: it can answer `error_code 0`, drop the broadcast, and publish that one - transaction itself at any later instant. The migration *does* confirm, so - nothing anywhere looks wrong, and it appears on chain **alone, at a time the - attacker chose** — which is precisely the linkage the batch exists to prevent. -- **Drive the flush clock.** `tip_height` runs over the same hop and takes the - max over endpoints (`chain.rs:155-175`), so a forged `GetLightdInfo` height - moves the cadence. That is the lever `hub-tip-advance-unbounded-flush-clock.md` - (confirmed, Medium) describes, here available without being the indexer. -- **Answer the wallets' lookups.** The hub relays every shim's `GetTransaction` - to the same endpoints and returns whatever comes back; `chain.rs:212-266` - never checks that the returned `RawTransaction` matches the requested - `TxFilter.hash`. - -**3. Even when it *is* set, nothing in the attested binary constrains the value.** -`IndexerTls::new(sni_name)` accepts any parseable name and verifies against the -compiled-in WebPKI roots (`tls.rs:59-84`). There is no compiled-in expected name, -no SPKI pin and no allowlist, and `ZIH_INDEXERS` is operator-supplied too. So -`ZIH_INDEXER_TLS=indexer.` plus `ZIH_INDEXERS=` is a fully TLS-"protected" hop that the operator terminates and -reads — and it looks *more* correct in a manifest review than deleting the line. -The TLS verification itself is correct and should be reported positively -(compiled-in `webpki-roots`, no `dangerous()`, no custom `ServerCertVerifier`, no -filesystem or environment root source anywhere in the crate); "correctly verified -against a name the adversary chose" is simply not a confidentiality property. - -**4. Bounding facts — all of them verified, and they are why this is Medium and -not High.** - -- The repository's own deploy path refuses the unset state **three times**: - `deploy.env.example:23` ships `INDEXER_TLS=na.zec.rocks`, `deploy.sh:115` uses - `: "${INDEXER_TLS:?set INDEXER_TLS for a hub}"` (verified by execution under - `dash`: `:?` aborts on unset **and** on empty), and - `assemble-caution.sh:126-131` refuses again with `exit 2`. Every worked example - in every runbook passes `--indexer-tls`. **This finding must therefore be - argued from the code, never from the deploy path.** -- An **empty** value fails closed. This one is derived from the two libraries' - sources rather than executed (there is no Rust toolchain in the audit - environment): clap 4.6.5 stores `env::var_os(name)` at `Arg::env()` - (`clap_builder/src/builder/arg.rs`, `let value = env::var_os(&name);`), so a - present-but-empty variable is applied as a value; `""` then reaches - `ServerName::try_from("")` (`tls.rs:76`), and rustls-pki-types 1.15.1's - DNS-name validator rejects the empty name (it is not a valid `DnsName` and does - not parse as an `IpAddr`), so `main.rs:33`'s `?` exits the process. Nothing in - this finding depends on that case; it is recorded so the *unset* case is not - confused with it. -- The value **is** covered by the attestation. `unit.env` string values become - `export KEY=` lines in the generated `run.sh` - (`caution-config/src/lib.rs:253-274`), `run.sh` is copied into the single EIF - ramdisk, and that ramdisk is measured into PCR0/PCR1. It is additionally served - to anyone at `.manifest.run_command` of every `/attestation` response. So the - filed claim "no party outside the operator can tell whether it is on" is - **struck**. What replaces it is narrower and still true: measurement - **discloses** a value, it never **detects** a change (`deploy.sh:206-219` - publishes the *deployed* tree as `--app-source`, so an edited manifest - reproduces its own PCRs and `caution verify` prints PASSED), and **no document - in zeronym tells anyone to look** — that omission is owned by the confirmed - Medium `auditor-recipe-omits-the-two-checks-that-decide-where-plaintext-goes-and-names-a-defence-the-platform-does-not-rely-on.md`. -- Against the **shipped** endpoint an accidental omission would not silently - succeed. `deploy.env.example:22-23` pairs `66.241.124.200:443` with the - certificate name `na.zec.rocks`, i.e. a TLS-terminating 443, and a TLS listener - does not answer plain h2c — so every tip query and every flush would fail and - the hub would be visibly broken (to the extent anything about an attested hub is - visible). This is therefore not a quiet accident on the shipped configuration: - it is a deliberate, deniable downgrade, or a genuine accident on a hub pointed - at a plaintext endpoint by other means. Note the corollary, which cuts the other - way: an on-path attacker who *does* answer h2c turns that visibly broken hub - into a working one under its own control. - -**5. Correction to the premise this file was originally titled on.** "Required in -practice in *every* document" is wrong. Two build documents state something -stronger and **false** — that the mechanism does not exist in the binary at all: -`hub/deploy/Containerfile:166-167` (*"The hub speaks plaintext HTTP/1.1 and has no -TLS stack in its dependency graph today"*) and `hub/deploy/README.md:48` (*"no TLS -stack"*, listed among the attested build's ingredients). Both are refuted by -`hub/Cargo.toml:95-102`, where `rustls`, `tokio-rustls`, `rustls-pki-types` and -`webpki-roots` are unconditional dependencies under a comment stating the exact -opposite. The reader of those two documents — the operator who configures the -image from them — is not merely under-warned about a variable; they are told the -mechanism behind it does not exist. (The `README.md` instance is filed separately -as `hub-deploy-readme-lists-no-tls-stack-among-the-attested-builds-ingredients.md`; -the `Containerfile` instance is Recommendation 3 of -`shim-containerfile-plants-a-system-ca-trust-store-in-the-enclave-that-tls-rs-says-cannot-exist.md`.) - -## Attack Scenario and Steps - -Two routes reach the unauthenticated state. Neither runs through `deploy.sh`. - -**Route A — a hub that is not deployed by the repository's scripts.** -`README.md:70` invites it (*"Operators run the shim in front of their indexer, -and optionally a hub"*), and the code and docs explicitly support hubs outside an -enclave (`config.rs:50-51` "correct only for a test or a trusted local path"; -`deploy.env.example:40` "Leave unset for a local hub"). `caution.hcl.tmpl` is a -*template*: an operator writing their own manifest, a container start, a systemd -unit or a `docker run` has no `assemble-caution.sh` in the path. Omitting -`ZIH_INDEXER_TLS` yields a running, healthy-looking hub — `/healthz` answers a -constant `ok` (`server.rs:450-452`) — with an unauthenticated egress. Note that -the neighbouring mistake fails loudly: omitting `ZIH_INDEXERS` is a clap -`required = true` error and the process never starts. Only the security-relevant -omission boots. - -**Route B — a one-line edit between assembly and deployment.** -`assemble-caution.sh` renders `$DEST/caution.hcl`; `deploy.sh` then pushes that -same directory to `APP_SOURCE` (`deploy.sh:206-219`). Deleting the -`ZIH_INDEXER_TLS` line from the deployed copy — or, in the leg-3 form, changing -its value to a name the operator holds a certificate for — is re-published as the -app source, so the enclave reproduces its **own** PCRs and `caution verify` -prints PASSED. Nothing re-validates the manifest, and no zeronym document tells -any verifier to read `.manifest.run_command`. - -Once in that state, the hub dials every `ZIH_INDEXERS` address over plain h2c -(`chain.rs:316-334`). The Nitro **parent host** is unavoidably on that path: the -enclave has no NIC, and every packet it sends is tunnelled to the parent over -vsock and forwarded by it (Caution's `run.sh.template`: -`socat TUN,tun-type=tap,…,tun-name=eth0 VSOCK-CONNECT:3:3`). So does any network -element between that host and the indexer. Any of them can now: - -1. Read every `SendTransaction` body of every flush — the whole batch, seconds - before it is public — plus every `GetTransaction` the hub relays for a shim. -2. Answer `SendResponse { error_code: 0 }` for chosen members and drop the - broadcast, destroying them permanently (§2 above) while the wallets were told - `error_code 0` at dispatch; or hold and re-publish a chosen member alone at a - chosen instant, which confirms normally and defeats the batch. -3. Answer `GetLightdInfo` with a height of its choosing and move the flush - cadence. - -**Attack Requirements and Assumptions:** - -- **No internet-side or wallet-side attacker can reach this.** The vulnerable - configuration is chosen by whoever deploys the hub. The exploiting party is - whoever is on the hub's egress path — for a Nitro deployment that is, first and - unavoidably, the parent host, i.e. `AUDIT-INSTRUCTIONS.md` attacker 8 and - exactly the party the enclave exists to exclude. -- **The repository's own path refuses it** (three checks, all verified). The - honest framing is a fail-open *default of the attested binary*, whose only - enforcement lives in two shell scripts that are neither measured, nor run at - boot, nor part of what an auditor verifies. -- Leg 3 (no pin on the value) needs no deviation from the deploy path at all — - only an operator who names a host they control. -- The state is disclosed in the attestation manifest, so it is detectable *by - someone who thinks to look*; nothing tells anyone to look, and nothing in the - running system refuses. - -## Impact on Users - -`hub/deploy/caution/caution.hcl.tmpl:10-17` states the deal the whole product -offers: attestation plus reproducibility together say *"the code you read is the -code that is holding your migration in plaintext, and it does nothing with it but -broadcast it."* Both legs of this finding falsify the second half — not by -changing the code, but by choosing who the code talks to. - -- **Integrity (the headline).** Migrations that a wallet reported as sent are - destroyed with no error, no retry and no diagnostic anywhere — the hub counts - them achieved. Past `nExpiryHeight` the loss is permanent, and there is no - confirmation tracking (designed, not built), so a user learns only by noticing - much later that funds never moved. The blast radius is fleet-wide: the hub - holds migrations from every shim. -- **Anonymity.** A party who both reads the batch and controls the tip can - isolate a chosen migration into a batch of one, or hold one back and publish it - alone later. That is the exact outcome the cadence, the shuffle, the mixnet and - the enclave were all built to prevent. -- **Confidentiality (bounded, and stated as such).** A few seconds' lead time on - data that becomes public anyway, plus the relayed `GetTransaction` lookups, - which never appear on chain at all. -- **Assurance.** `README.md:77` tells users the attested enclaves are "the only - things that ever see a migration in cleartext" and `hub/src/lib.rs:29` tells a - code reader that `tls` "verifies that connection, which an enclave deployment - requires". Neither is enforced by the artefact the attestation covers. - -## Technical Details / Code Analysis - -The field, in full (`hub/src/config.rs:47-53`): - -```rust - /// The DNS name the indexer's certificate must carry. - /// - /// Unset means PLAINTEXT h2c, which is correct only for a test or a trusted - /// local path. A deployed enclave must set this: without it the enclave's - /// parent host reads every batch in the clear moments before it is public. - #[arg(long = "indexer-tls", env = "ZIH_INDEXER_TLS")] - pub indexer_tls: Option, -``` - -The accessor (`hub/src/config.rs:100-108`): - -```rust -impl Config { - /// The TLS verifier for indexer connections, if one is configured. - pub fn indexer_tls(&self) -> Result, BoxError> { - match &self.indexer_tls { - Some(name) => Ok(Some(IndexerTls::new(name)?)), - None => Ok(None), - } - } -} -``` - -The whole of the startup enforcement (`hub/src/main.rs:30-39`): - -```rust - // Refuse to run blind. A hub without indexer TLS lets the enclave's parent - // host read every batch in the clear moments before it is public, so this is - // announced loudly rather than left to a config review. - let tls = config.indexer_tls()?; - if tls.is_none() { - tracing::warn!( - "no --indexer-tls: the hop to the indexer is PLAINTEXT and the host can read every batch" - ); - } - let chain = Arc::new(ChainClient::new(config.indexers.clone(), tls)?); -``` - -versus the invariant five lines below that *does* abort (`hub/src/main.rs:41-45`): - -```rust - // The expiry budget is asserted at startup, not trusted. A parameter change - // that overspends it must fail here rather than be discovered in production - // as a percentage of real traffic quietly expiring. - let params = BatchParams::default(); - params.validate()?; -``` - -"Announced loudly" is also false in the deployment this component targets. The -warning is a `tracing` line on the enclave's console, and an attested enclave has -no console: Caution's own Terraform starts the enclave with `--debug-mode` and -installs the console-capture unit **only** under `debug_mode == "true"` -(`terraform/modules/aws/nitro-enclave/user-data.sh:166`, and the -`%{ if debug_mode == "true" }` block around `capture-enclave-console.sh`), while -`debug { enabled = true }` is documented by the platform as disabling attestation -verification. In the attested case the warning is discarded; in the debug case it -is delivered to the parent host, i.e. to the one party the plaintext benefits. - -The live plaintext branch (`hub/src/chain.rs:316-334`): - -```rust - let authority = match &self.tls { - Some(tls) => tls.authority().to_owned(), - None => addr.to_string(), - }; - - let request = hyper::Request::builder() - .method("POST") - .uri(format!("http://{authority}{path}")) - … - - let response = match &self.tls { - Some(tls) => { - let stream = tls.connect(addr, stream).await?; - round_trip(stream, request).await? - } - None => round_trip(stream, request).await?, - }; -``` - -with the field's own comment stating an assumption about operator behaviour as -though it were an invariant (`chain.rs:126-128`): *"`None` means plaintext h2c … -A deployed enclave always sets this."* - -The verdict a forged reply reaches (`hub/src/chain.rs:438-445`): - -```rust -fn classify_send_response(resp: &SendResponse) -> Publish { - if resp.error_code == 0 { - return Publish::Accepted { - txid: resp.error_message.clone(), - }; - } - classify_publish_error(&resp.error_message) -} -``` - -and what `flush` then does with it (`hub/src/batcher.rs:342`, `:365-390`): - -```rust - let batch = queue.drain_shuffled(); - … - for (i, entry) in batch.into_iter().enumerate() { - match outcomes.get(i) { - Some(Publish::Accepted { .. }) | Some(Publish::AlreadyKnown) => achieved += 1, - Some(Publish::Rejected { .. }) => rejected += 1, - Some(Publish::Retryable { reason }) => { … unplaced.push(entry); } - None => unplaced.push(entry), - } - } - … - let requeued = queue.requeue(unplaced); -``` - -Only `Retryable` survives. An `Accepted` — which any on-path party can produce by -answering `error_code = 0` — consumes the entry, and `flush`'s own comment -records why that is unrecoverable: *"the shim answered the wallet error_code 0 -the moment the frame reached the mixnet and keeps no record, so once the entry -left this queue there is no other copy anywhere that anyone will retry."* - -On leg 3, the entirety of what the attested binary knows about who it will talk -to (`hub/src/tls.rs:59-84`): - -```rust - pub fn new(sni_name: &str) -> Result { - install_crypto_provider(); - - let roots = RootCertStore { - roots: webpki_roots::TLS_SERVER_ROOTS.to_vec(), - }; - let mut config = ClientConfig::builder() - .with_root_certificates(roots) - .with_no_client_auth(); - config.alpn_protocols = vec![ALPN_H2.to_vec()]; - - let server_name = ServerName::try_from(sni_name.to_owned()) - .map_err(|_| -> BoxError { format!("invalid indexer TLS name {sni_name:?}").into() })?; - … -``` - -`sni_name` is whatever the environment said, and the addresses come from -`ZIH_INDEXERS`. Both are operator-supplied; neither is compared against anything. - -Finally, the enforcement everyone relies on, in the two places it actually lives -— both outside the measured artefact, neither run at boot: - -```sh -# deploy.sh:115 - : "${INDEXERS:?set INDEXERS for a hub}"; : "${INDEXER_TLS:?set INDEXER_TLS for a hub}" - -# hub/deploy/caution/assemble-caution.sh:126-131 -[ -n "$INDEXER_TLS" ] || { - echo "error: --indexer-tls is required (the DNS name the indexer's cert carries)." >&2 - echo " Without it the hop is plaintext and the parent host reads every batch." >&2 - exit 2 -} -``` - -Degenerate-value behaviour, checked rather than assumed: an **IP literal** is -accepted (`ServerName::IpAddress`) and then fails at the handshake — fail-closed -with a confusing symptom, the hub-side instance of the trap filed for the shim in -`backend-tls-ip-literal-defeats-the-startup-validation-and-debug-formats-into-the-authority.md` -(the hub's `authority()` stores the original string, so that issue's -`Debug`-into-protocol half does not apply here) and, for the hub's own copy of -that gap, `hub-indexer-tls-drops-two-guards-its-sibling-backend-tls-carries.md` -(plausible, Low — a separate finding about the *set* value, not the unset one); -and `ZIH_INDEXERS` empty or -whitespace-only fails closed twice over (clap `required = true` plus -`SocketAddr` parsing, and `ChainClient::new`'s own `endpoints.is_empty()` check -at `chain.rs:132-136`). - -## Recommendations - -1. **Make the safe state the default and the unsafe one explicit, in the - binary.** Either mark `--indexer-tls` `required = true`, or keep the `Option` - and require an explicit opt-in — `--allow-plaintext-indexer` / - `ZIH_INSECURE_PLAINTEXT_INDEXER=true` — before `indexer_tls()` may return - `Ok(None)`, aborting startup otherwise exactly as `params.validate()?` already - does one screen below. This is the same discipline `--nym` and - `--http-submit` already use, applied to the variable that needs it most. A - test-only convenience should cost a flag that says "insecure", not a missing - variable. Update `config.rs:142-144` so the test pins the new behaviour. -2. Failing that, refuse `None` unless every `--indexer` address is loopback or - RFC1918. That preserves the stated legitimate use ("a test or a trusted local - path") and removes the deployed one. -3. **Correct the two comments and the two build documents that are false.** - `main.rs:30-32` says "Refuse to run blind" and "announced loudly"; neither is - true, and the second cannot be true in an enclave with no console. - `chain.rs:126-128`'s "A deployed enclave always sets this" is an assumption, - not an invariant. `hub/deploy/Containerfile:166-167` and - `hub/deploy/README.md:48` state that the hub has **no TLS stack**, which - `hub/Cargo.toml:95-102` refutes. - - *Added 2026-08-18 when `hub-deploy-readme-lists-no-tls-stack-among-the-attested-builds-ingredients.md` - was merged into this issue (see the merge record below).* Two details the - merged filing established, which the fixing commit should carry: - (a) `hub/deploy/README.md:46-48` asserts "no TLS stack" of the **shim** as - well — the sentence ends "is identical to the shim" — and that half is false - too (`shim/Cargo.toml:71-79` takes `rustls`, `tokio-rustls`, - `rustls-pki-types` and `rustls-acme` unconditionally, and the shim's own - `deploy/README.md:299` records the build where they were linked in as - "TLS on both hops … the binary grows 4.4 MB to 7.6 MB, which is the TLS - stack"); (b) the sentence sits in this document's list of **determinism - ingredients**, so the correct replacement is not a deletion but a positive - statement: the hub links `rustls` with the `ring` provider and a compiled-in - `webpki-roots` trust store, so a `webpki-roots` bump both moves the published - hash and changes which CAs the enclave trusts. Provenance for the fix - message: the phrase was written by `d5f4687` (2026-08-09), when the hub really - had no TLS and spoke plaintext JSON-RPC; `a746496` (2026-08-09, the same day) - added `hub/src/tls.rs` and the four dependencies, rewrote the bullet - immediately above it in the same file, and left it standing. -4. **Give leg 3 something to check against.** Publish the expected - `ZIH_INDEXER_TLS`/`ZIH_INDEXERS` alongside the expected measurements so a - verifier reading `.manifest.run_command` has a reference; better, compile the - expected name (or an SPKI pin) into the measured binary, so changing the peer - changes what an auditor reproduces rather than only what they could have read. -5. Optionally, report the effective indexer-hop mode (authenticated under which - name, or plaintext) on `GET /nym-status`. It discloses nothing an adversary - does not already have — the same string is in the public manifest and the - endpoint addresses are in the egress rules — and it converts "clone the - operator's `app_sources` repo, or POST for an attestation, and know to read - `run_command`" into one `curl` any shim operator or wallet author can run. - Note this is a convenience, not the fix: recommendation 1 is the fix. -6. Add the manifest check to the auditor recipe at `README.md:71` — tracked in - `auditor-recipe-omits-the-two-checks-that-decide-where-plaintext-goes-and-names-a-defence-the-platform-does-not-rely-on.md` - (confirmed, Medium), which owns that half and should not be double-counted - here. - -Cross-references: `hub-tip-advance-unbounded-flush-clock.md` (confirmed, Medium — -what a forged tip on this hop buys); `hub-chain-zaino-node-rejections-are-never-verdicts.md` -and `publish-verdict-strings-are-zcashds-vocabulary-only-…md` (confirmed, Medium — -the other ways a verdict destroys an entry); -`hub-lookup-fall-through-hands-every-wallets-txid-to-whichever-indexer-the-hub-is-pointed-at-…md` -(confirmed, Medium — the lookups this hop also carries); -`hub-chain-indexer-tls-hop-has-no-test-coverage-and-declares-scheme-http.md` -(plausible, Info — the same hop from the `chain.rs` side); and -`operators-runbook-attributes-the-hub-destination-to-the-binary-hash-and-egress-rules-neither-of-which-binds-it.md` -(confirmed, Low — the same "configuration, not the binary, decides where plaintext -goes" shape on the shim side). - -## Validation Information - -**Verdict: CONFIRMED. Severity: Medium** (both filings proposed Medium; that -grade survives, for different reasons than either gave). Validated 2026-08-18 by -the Issue Validator, which also **merged** the two concurrently filed copies of -this finding into this file. - -### What was verified, directly against the target - -| Claim | Verified at | -|---|---| -| `indexer_tls` is `Option`, unset → `Ok(None)` | `config.rs:47-53`, `:100-108` | -| Unset produces a `warn!` and the process continues | `main.rs:30-39`; no `Config::validate` anywhere in `hub/src` | -| A live, complete plaintext branch consumes it | `chain.rs:316-334`; field comment `:126-128` | -| A unit test pins the optionality as intended | `config.rs:142-144` | -| The two booleans nearby take an explicit `true`/`false` for this exact reason | `config.rs:59-65`, `:93-97` | -| `params.validate()?` aborts startup for constants an operator cannot change | `main.rs:41-45`; `BatchParams` built only by `Default::default()`; one `env::var` in `hub/src` (`nym_driver.rs:202`) | -| `error_code == 0` ⇒ `Accepted { txid: error_message }`, no validation of the string | `chain.rs:438-445` | -| `Accepted` outranks everything and consumes the entry; only `Retryable` is requeued | `chain.rs:459-467`; `batcher.rs:342`, `:365-390` | -| The wallet was already told success at mixnet hand-off | `shim/src/hub.rs:232-240` | -| Relayed lookups are not checked against the requested hash | `chain.rs:212-266` | -| No pin/allowlist of any kind on the configured name | `tls.rs:59-84`; no `dangerous()`, no custom verifier, no filesystem/env root source in the crate | -| `deploy.sh:115`'s `${VAR:?}` aborts on unset **and** on empty | executed under `dash`: both cases exit 2 | -| `assemble-caution.sh:126-131` refuses an empty/absent value with `exit 2` | read directly | -| `deploy.env.example:23` ships a value; every worked runbook example passes `--indexer-tls` | `deploy.env.example:23`; `hub/deploy/caution/README.md:38`; `OPERATORS.md:61` | -| `deploy.sh` publishes the **deployed** tree as `--app-source` | `deploy.sh:206-219` | -| An attested enclave has no console; console capture exists only under `--debug-mode` | Caution `terraform/modules/aws/nitro-enclave/user-data.sh:166` and its `debug_mode` guard | -| Every enclave packet transits the parent host | Caution `src/enclave-builder/templates/run.sh.template` (`socat TUN … VSOCK-CONNECT:3:3`) | -| `unit.env` is measured and served at `.manifest.run_command` | `caution-config/src/lib.rs:253-274` → `run.sh` → single EIF ramdisk; PROGRESS item 6q, re-derived here from the same source | -| The two build documents deny the TLS stack exists | `hub/deploy/Containerfile:166-167`, `hub/deploy/README.md:48`, refuted by `hub/Cargo.toml:95-102` | - -### What was struck or corrected from the two filings - -1. **"Unobservable when absent" / "no party outside the operator can tell whether - it is on" — STRUCK.** `unit.env` is measured into PCR0/PCR1 and the whole - environment is served at `.manifest.run_command` of every `/attestation` - response. The replacement statement is "measurement discloses, it does not - detect, and no document tells anyone to look" — and the *documentation* half of - that is owned by the confirmed Medium recipe issue, not by this one. -2. **"Required in practice in every document" — CORRECTED.** Two build documents - say the opposite and are false. The filename is retained for cross-reference - stability with this correction recorded at the top of the file. -3. **"The attestation chain is untouched / verification still passes with the - variable removed" — REPLACED with the accurate mechanism.** Verification does - not pass against a *previously published* tree; it passes because `deploy.sh` - publishes whatever was deployed, so the modified enclave reproduces its own - PCRs. Same practical outcome, different and checkable reason. -4. **The confidentiality framing was demoted and the integrity framing promoted** - (this was filing 1's distinct contribution, and it is the part that carries the - severity). One consequence neither filing had is added: an on-path party can - *hold and later re-publish* a single transaction, which confirms normally and - is a cleaner anonymity break than destroying it. -5. **"Global focus area 48 is unresolved" — REMOVED**; it was closed by the G10 - platform pass and the answer is folded in above. -6. **The `main.rs` warning's fate was sharpened rather than assumed**: verified - from Caution's Terraform that console capture exists only in `--debug-mode`, - i.e. the warning is discarded exactly when attestation is on. - -### Why Medium — not High, not Low - -- **Not High.** The unauthenticated state is not reachable through any documented - path: three independent refusals stand between an operator and it, and against - the shipped TLS endpoint an accidental omission breaks the hub loudly rather - than working insecurely. Every downstream harm (tip manipulation, verdict-driven - destruction, lookup exposure) is already separately graded, and the "nobody is - told to check the manifest" half is owned at Medium elsewhere; stacking those - severities here would count one harm three times, against the precedent set in - PROGRESS item 7a. -- **Not Low.** This is the definition case for Medium in - `docs/AUDIT-PROCESS.md`: *"a serious vulnerability that only exists if the user - has configured the application in a specific, uncommon way."* The defect is in - the attested binary — the artefact the entire reproducibility and attestation - apparatus points at — and its only enforcement lives in two shell scripts that - are not measured, do not run at boot, and are not what an auditor verifies. The - dangerous state is selected by *omission* rather than by an explicitly named - insecure flag, which is precisely the pattern - `docs/AVOIDING-FALSE-POSITIVES.md` §7 identifies as a real issue rather than a - configuration excuse. Leg 3 needs no deviation from the deploy path at all. - And the code-side consequence is not a leak of already-public data: it is - permanent, silent destruction of migrations that wallets have already reported - as sent, plus a lever that isolates a chosen migration on chain. - -### Merge record - -`hub-indexer-tls-is-optional-operator-chosen-and-unobservable.md` (filed by the -`hub/src/tls.rs` local audit) is the same finding with a different emphasis. Its -distinct contribution — leg 3, that nothing in the attested binary constrains the -value when it *is* set — is carried here in Description §3, Route B, Impact and -Recommendation 4. That file has been moved to `invalid/` **for bookkeeping only**, -with a header stating that its substance was validated and found real and that it -must not be reported as a refuted claim. - -**Second merge, 2026-08-18.** -`hub-deploy-readme-lists-no-tls-stack-among-the-attested-builds-ingredients.md` -(filed by the `hub/deploy/README.md` local audit) is the document-side half of -the same defect: it audits `hub/deploy/README.md:46-48`, the second of the two -build documents already named in this issue's Location line, verification table -and Recommendation 3. Every claim in it was re-verified and holds; its two -details that were not already here are folded into Recommendation 3 above. It is -under `invalid/` **for bookkeeping only**, with the same MERGED-NOT-REFUTED -header. The report must present this defect once, here. - -DO NOT DEVIATE FROM THIS TEMPLATE (`docs/SECURITY-ISSUE-TEMPLATE.md`). -LEAVE THIS MESSAGE IN PLACE SO THAT ALL AGENTS KNOW NOT TO DEVIATE FROM THIS TEMPLATE (`docs/SECURITY-ISSUE-TEMPLATE.md`). diff --git a/zeronym-22aa9851caf68-high-medium/medium/hub-liveness-probe-reads-its-own-send-backlog-as-gateway-silence-so-any-stranger-can-drive-the-fresh-identity-fleet-kill.md b/zeronym-22aa9851caf68-high-medium/medium/hub-liveness-probe-reads-its-own-send-backlog-as-gateway-silence-so-any-stranger-can-drive-the-fresh-identity-fleet-kill.md deleted file mode 100644 index c2aa4f84..00000000 --- a/zeronym-22aa9851caf68-high-medium/medium/hub-liveness-probe-reads-its-own-send-backlog-as-gateway-silence-so-any-stranger-can-drive-the-fresh-identity-fleet-kill.md +++ /dev/null @@ -1,681 +0,0 @@ -# The hub's inbound-liveness verdict is measured through the SDK send queue an anonymous stranger controls, so a pulsed lookup burst makes a healthy hub declare itself undelivered-to — and `short_lives` never decays, so five pulses reach the fresh-identity fallback - -**Severity**: Medium -**Validation Status**: Confirmed -**Location**: -`audit-target/zeronym/hub/src/nym_driver.rs:417-451` (the probe arm), specifically `:421` (the silence predicate, with **no** send-backlog conjunct), `:432` (`Step::Silent`), `:484-493` (`Sent::Probe => inbound_at_probe = Some(inbound_total)`), `:494-503` (`inbound_total += 1`), `:530-570` (the short-life accounting, especially `:558` `if silent || lived < STABLE_LIFE` and `:568` the reset), `:240-245` (`Failures::exhausted`), `:279-306` (the fresh-identity fallback), `:112` (`STABLE_LIFE`), `:124` (`SHORT_LIVES_BEFORE_NEW_IDENTITY`), `:133` (`PROBE_INTERVAL`), `:138` (`SILENT_ROUNDS_BEFORE_REBUILD`), `:632-643` (`reply_send`), `:647-671` (`probe_send` / `send_probe`); -the guard the hub is missing: `audit-target/zeronym/shim/src/nym_driver.rs:284-311` (the rationale) and `:312` (`Some(mark) if seen == mark && out_frames.len() == 0`); -the cheap reply that manufactures the backlog: `audit-target/zeronym/hub/src/nym.rs:167-196` (the lookup arm), `:232-234` (`is_lookup`), `:265-272` (the `hash.is_empty()` arm), `audit-target/zeronym/hub/src/wire.rs:44-49` (the `LookupV1` layout) and `:86` (`LOOKUP_BYTES = 64`); -the emission arithmetic: `audit-target/zeronym/shim/src/nym.rs:96-104` (`LOOKUP_REPLY_SURBS`; a full frame is 41 reply packets) and `:1066-1152` (`throughput_budget`, and the project's own measurement that a deployed enclave pegs at the throttled rate); -the ingress: `audit-target/zeronym/hub/src/server.rs:62-70` (`GET /nym-address` publishes the target), `:76-83` (`GET /nym-status` publishes `client_deaths`), `audit-target/zeronym/hub/deploy/caution/caution.hcl.tmpl:51-55` (`ingress 0.0.0.0/0`); -the pinned SDK (`nym-sdk` git rev `451c2aa3692fc4dc00041b74a352d4158176d9c0`): `sdk/rust/nym-sdk/src/mixnet/traits.rs:80` and `:126` (both `send_message` and `send_reply` use `TransmissionLane::General`), `common/client-core/src/client/real_messages_control/real_traffic_stream.rs:427-500` (one whole message stored per Poisson tick, one packet emitted per tick), `common/client-core/src/client/transmission_buffer.rs:39-49` (unbounded) and `:170-178` (`pop_front_from_lane`, strict FIFO within a lane), `common/client-core/src/client/received_buffer.rs:66-73` (loop cover traffic is filtered out and therefore never advances `inbound_total`), `common/client-libs/gateway-client/src/packet_router.rs:39-73` (acks are routed on a separate channel and never surface as messages), `common/client-core/src/client/replies/reply_controller/receiver_controller.rs:179-216` (`should_request_more_surbs`) and `common/client-core/config-types/src/lib.rs:48-50` (the SURB thresholds that decide whether the hub asks the attacker for more SURBs). -**Found by agent:** Global (focus area G21, resource exhaustion as a privacy attack — dedicated re-run) -**In scope of audit?** Yes — priority area #4 ("the mixnet transport ... identity rotation, the liveness probe"), and the `*/src/nym_driver.rs` row of "Code Areas That Should Get Extra Attention" - -## Description - -The hub's mixnet driver has exactly one self-repair mechanism: every 60 seconds it -sends an empty message to its **own** Nym address and counts the probe rounds in -which **no inbound message of any kind** arrived. Two such rounds tear the client -down (`Step::Silent`); five teardowns mint a fresh Nym identity, which permanently -invalidates the `ZIS_HUB_NYM` value baked into every shim's immutable enclave -configuration. - -That terminal state is already confirmed as -`hub-nym-driver-automatic-fresh-identity-permanently-invalidates-every-shim.md` -(Medium) and is reached by two other confirmed issues. **This issue is not about -the terminal state. It is about a defect in the predicate that decides to go -there: the silence test does not measure whether the gateway is delivering to the -hub. It measures whether the hub's own outbound emission queue has drained far -enough for the probe to have left.** Those are the same thing only while the queue -is short, and any anonymous stranger can make it long for a few thousand sphinx -packets. - -The predicate is (`hub/src/nym_driver.rs:419-421`): - -```rust - match inbound_at_probe { - // A probe was outstanding and nothing at all has arrived - // since. - Some(mark) if inbound_total == mark => { - silent_rounds += 1; -``` - -The sibling binary has the identical mechanism **with a backlog conjunct** -(`shim/src/nym_driver.rs:312`): - -```rust - Some(mark) if seen == mark && out_frames.len() == 0 => { -``` - -and the shim's author wrote out exactly why, at `shim/src/nym_driver.rs:289-311`: - -> But "since" only means anything if the probe actually LEFT. The mark is -> stamped when the SDK accepts the probe into its one-slot input channel, not -> when it is emitted, and behind that slot sit an 8-deep batch channel and an -> unbounded FIFO drained at the throttled rate. Under a send backlog the probe -> is still queued behind every frame ahead of it, nothing has been asked of the -> gateway yet, and **"silent" is a statement about OUR queue, not about -> delivery.** Rebuilding on it would disconnect a healthy client and discard -> that whole queue. - -The hub has no such conjunct, the hazard is cheaper to induce on the hub than on -the shim, and it costs far more when it fires. - -### Why the hub's SDK queue can be made to hold minutes of emission - -Three facts, all verified against the pinned SDK tree during validation: - -1. **Handing a reply to the SDK is ~41x faster than emitting it.** - `OutQueueControl::poll_poisson` polls its 8-slot `real_receiver` **once per - Poisson tick** and, on `Ready`, stores the *entire* fragment vector of one - message into the transmission buffer while emitting exactly **one packet** - (`real_traffic_stream.rs:461-479`). A 64 KiB `LookupReplyV1` is 41 packets by - the project's own measurement (`shim/src/nym.rs:99-104`), and the throttled - emission rate is 8.33 packets/s — `MAX_DELAY_MULTIPLIER = 6` times the 20 ms - default `message_sending_average_delay`, which the project states its deployed - enclaves actually sit at (`shim/src/nym.rs:1071-1077`: *"the same code against - a shared public gateway pegged at multiplier 6"*). So the General lane grows by - **40 packets per tick** for as long as the hub has replies to hand over. -2. **The buffer is unbounded, un-aged and shared with the probe.** - `TransmissionBuffer` is a bare `HashMap>` with - no cap (`transmission_buffer.rs:39-49`), `pop_front_from_lane` is strict FIFO - within a lane (`:170-178`), and **both** `send_reply` and `send_message` use - `TransmissionLane::General` (`nym-sdk/src/mixnet/traits.rs:80`, `:126`). The - liveness probe is therefore queued strictly behind every reply fragment already - stored. -3. **Nothing else advances `inbound_total`.** It is incremented once per - reconstructed inbound message (`hub/src/nym_driver.rs:501`). Loop cover traffic - is discarded before reconstruction (`received_buffer.rs:66-73`) and - acknowledgements are routed on a separate channel that never reaches the - message buffer (`gateway-client/src/packet_router.rs:39-73`), so neither - surfaces through `wait_for_messages`. When the attacker stops sending, the - hub's own probe echo is the **only** thing that can move the counter. - -Points 1 and 2 are not new: the confirmed -`hub-nym-lookup-flood-starves-gettransaction-fleet-wide.md` established them and -showed that `REPLY_DEADLINE` and the `in_flight` guard are inoperative for the -same reason. What is new is point 3's consequence: **the mechanism that is -supposed to detect a dead gateway is measured through the very queue the attacker -controls.** - -### Why five teardowns need not be five consecutive minutes - -`failures` is reset in exactly one place (`hub/src/nym_driver.rs:567-569`): the -`else` of `if silent || lived < STABLE_LIFE`, i.e. only when a client **dies -non-silently** after living at least three minutes. A `Step::Silent` teardown -always increments `short_lives` regardless of how long the client lived (`:558`, -and the comment at `:553-557` says so deliberately), and a client that simply -keeps running never resets anything. So `short_lives` has **no decay**: a hub that -runs healthily for a week between attacker pulses still carries whatever -`short_lives` it had. The counter's own doc comment calls these "CONSECUTIVE -short-lived clients" (`:114-115`); in practice the only thing that clears them is -a rare event the attacker is not required to avoid for long. - -## Attack Scenario and Steps - -Attacker: anyone on the internet. No credential, no funds, no enclave compromise, -no gateway position, no ability to degrade any Nym node. - -1. `curl https:///nym-address`. The endpoint exists to publish the - value and answers everybody (`hub/src/server.rs:62-70`). -2. Start a few stock `nym-sdk` clients. They are free: neither enclave sets - `enabled_credentials_mode` and the SDK default is `false`. -3. **Burst.** Send anonymous 64-byte `LookupV1` frames - (`b"ZNL1"` ‖ 16 random bytes ‖ `0x00` ‖ 43 zero bytes — `hub/src/wire.rs:44-49`), - each carrying **>= 51 reply SURBs**, which is the provisioning a conforming shim - uses (`LOOKUP_REPLY_SURBS = 60`, `shim/src/nym.rs:99-104`). Each takes the - `hash.is_empty()` arm (`hub/src/nym.rs:269-272`), does **no** indexer dial and - **no** queue scan, and yields a full 64 KiB `error` reply that the hub hands to - the SDK within milliseconds. Two properties of the SURB count matter and are - both required: - - >= 51 (41 for the reply plus the SDK's `minimum_reply_surb_storage_threshold` - of 10) means the whole 41-packet reply is prepared and queued at once, which - is what builds the backlog; - - it also leaves the hub's SURB pool above threshold afterwards, so - `should_request_more_surbs` stays false - (`receiver_controller.rs:179-216`) and the hub never asks the attacker for - more SURBs. That matters because a SURB top-up would arrive as an inbound - message and reset `silent_rounds` — see step 4. -4. **Go quiet for ~135 seconds.** Send nothing further, from any client. -5. Once the hub's own `outgoing` channel drains (it drains almost instantly — see - Technical Details), the deferred probe tick fires. The first quiet tick sees - `inbound_total != mark` (traffic arrived during the burst), resets - `silent_rounds` and queues probe P1, stamping the mark at the final count. P1 - is appended to the General lane behind the attacker's backlog. Sixty seconds - later nothing has arrived — P1 has not been emitted, the attacker is silent, - cover traffic does not count and acks do not count — so `silent_rounds` reaches - 1; sixty seconds after that it reaches `SILENT_ROUNDS_BEFORE_REBUILD` and the - driver takes `Step::Silent` (`:423-432`). -6. The hub calls `client.disconnect().await` (`:537`), which discards everything - the SDK still holds — **including any honest shim's reply queued behind the - attacker's** — records `status.set_died()`, increments `failures.short_lives` - (`:559`), backs off 5 s and rebuilds on the same storage at the same address. -7. Repeat steps 3-6 four more times, at any spacing. On the fifth, - `failures.exhausted()` is true at the top of the outer loop (`:286`), and the - hub executes `storage = Ephemeral::default(); gateway = None;` — **a fresh Nym - identity and therefore a new address.** - -**Cost, recomputed during validation** (the original filing understated the burst -and overstated the role of `MAX_CONCURRENT_LOOKUPS`; see Technical Details §3): - -- The probe must stay unemitted for two probe intervals plus the round trip, so - the General lane must hold roughly `8.33 x 130 s` ~= **1,100 packets ~= 27 - full-frame replies** at the moment P1 is queued. -- Backlog grows at `41 x lambda - 8.33` packets/s, where `lambda` is the rate at - which lookups reach the hub. The attacker therefore needs - **`lambda > 0.21 lookups/s`**, i.e. more than ~11 sphinx packets/s of their own - emission at ~52 packets per well-SURBed lookup. -- One stock client throttled to 8.33 packets/s cannot do it; **two can (~4 - minutes of bursting), three do it in ~100 s, five in ~45 s.** A client on an - unthrottled gateway (multiplier 1, ~50 packets/s) does it alone, and an attacker - writing raw sphinx traffic has no client-side shaping at all. -- Per pulse that is **~45-75 lookups, ~2,400-3,900 sphinx packets (~5-8 MB)**; - five pulses is **~12,000-20,000 packets (~25-40 MB)**, spread over any interval - the attacker likes. - -**Attack Requirements and Assumptions:** - -- **What makes it realistic.** The hub's Nym address is published by design and - the enclave declares `ingress 0.0.0.0/0`. There is no ACL, no rate limit and no - submitter identity to key one on. Every step uses stock SDK clients sending - traffic that is byte-for-byte the shape a conforming shim sends — the SURB - provisioning in step 3 is *exactly* what an honest lookup carries, so nothing - distinguishes it. -- **The honest bound: the attacker needs a ~135-second window in which no shim - traffic reaches the hub, five times in total.** Any inbound message — a real - submit, a real lookup, or a SURB-replenishment artefact — advances - `inbound_total` and resets `silent_rounds`. This is why the attack is a *pulsed* - burst and not a sustained flood; a sustained flood keeps the hub looking alive. - Two things make the window obtainable today: traffic is sparse (`README.md` - records that no migration has yet been observed crossing Nym; the first - third-party operator dates from 2026-08-10), and `short_lives` never decays, so - the attacker needs five such windows *in total* rather than five in a row. It - gets harder as the fleet grows, and every wallet `GetTransaction` behind any - shim is a reset. -- **The other bound: an intervening non-silent client death after `STABLE_LIFE` - resets the count.** The attacker cannot prevent that, but can race it by pulsing - every few minutes, and failed attempts cost only packets. -- The attacker gets a free progress counter: `GET /nym-status` publishes - `client_deaths` unauthenticated (`hub/src/server.rs:182-190`), which increments - on every `Step::Silent`, and `GET /nym-address` confirms the final rotation. - -## Impact on Users - -**Scope note, so this is not double-counted.** The terminal state and everything -that follows from it are already owned by -`hub-nym-driver-automatic-fresh-identity-permanently-invalidates-every-shim.md` -(Confirmed, Medium) and are reached independently by -`hub-surb-starved-lookup-replies-grow-the-sdk-pending-buffer-without-bound-and-oom-the-enclave.md` -(Confirmed, High). **What this issue contributes is the trigger**: a deliberate, -cheap, stranger-reachable path to that state that needs no memory exhaustion, no -process restart and no position on the mixnet — and, critically, one that the -confirmed High's recommended fix does not close. - -The consequences that belong to *this* file are: - -1. **A stranger can tear down the hub's mixnet client on demand, repeatedly.** - Each cycle costs the whole fleet a `disconnect()` plus a 5 s backoff plus a - reconnect, during which `mixnet_connected` reads false and lookups fail closed - at every shim. `client.disconnect()` at `:537` discards the SDK's whole - transmission buffer, so every honest shim's reply queued behind the attacker's - burst is destroyed too and those wallets see `UNAVAILABLE`. This is a - *chosen-moment* outage: an adversary who wants a particular shim's lookups to - fail, or wants to shape when the hub is carrying traffic, can produce it. -2. **The hub charges itself an irreversible strike for it.** Because - `short_lives` never decays, each successful pulse is permanent progress toward - an action — a fresh Nym identity — whose recovery the project's own runbook - budgets at *"well over an hour"* across every operator, with *"no discovery - mechanism; the handoff is a human message"* - (`hub/deploy/caution/OPERATORS.md:230-243`). -3. **The log line asserts a conclusion the predicate cannot support.** `:426-431` - says *"the hub is registered but not being delivered to"*. What the predicate - actually established is "our own probe has not come back yet", and the probe - may never have been emitted. An operator reading that line is told the wrong - thing about their own system. -4. **The monitoring is misleading rather than absent, and the misleading part is - what this trigger exploits.** `mixnet_connected` does flap false during each - silent round and `client_deaths` does climb, so the attack is not invisible — - but `OPERATORS.md:167-174` characterises climbing `client_deaths` as benign - *"gateway churn"*, and the counter the same table names as the fresh-identity - predictor, `consecutive_rebuild_failures`, reads **0** for the entire walk, - because every cycle connects successfully and `NymAddress::set` zeroes it - (`hub/src/server.rs:127-137`). The runbook's own early-warning field cannot - move on this path. (This instrumentation gap is shared with, and primarily - owned by, the confirmed fresh-identity issue.) - -## Technical Details / Code Analysis - -**1. The hub's probe arm, in full (`hub/src/nym_driver.rs:417-451`):** - -```rust - _ = probe.tick(), if in_flight.is_none() => { - match inbound_at_probe { - // A probe was outstanding and nothing at all has arrived - // since. - Some(mark) if inbound_total == mark => { - silent_rounds += 1; - if silent_rounds >= SILENT_ROUNDS_BEFORE_REBUILD { - tracing::error!( - silent_rounds, - gateway = %own.gateway(), - "no inbound mixnet traffic across consecutive probes; the hub \ - is registered but not being delivered to. Rebuilding on the \ - same registration; if that stays silent the fresh-identity \ - fallback follows." - ); - Step::Silent - } else { - tracing::warn!( - silent_rounds, - "no inbound mixnet traffic since the last probe; watching" - ); - in_flight = Some(probe_send(sender.clone(), own)); - Step::Ferried - } - } - // Either the first round, or traffic HAS arrived since the - // last probe ... - _ => { - silent_rounds = 0; - in_flight = Some(probe_send(sender.clone(), own)); - Step::Ferried - } - } - }, -``` - -**2. The mark is stamped at SDK acceptance, not at emission -(`hub/src/nym_driver.rs:484-493`):** - -```rust - sent = drive(&mut in_flight), if in_flight.is_some() => { - in_flight = None; - match sent { - Sent::Reply => {} - // The mark means "inbound seen as of the probe going out", - // so it is read here and not when the probe was queued. - Sent::Probe => inbound_at_probe = Some(inbound_total), - } - Step::Ferried - }, -``` - -The comment says *"as of the probe going out"*. `Sent::Probe` resolves when -`sender.send_message(...).await` returns, which is when the SDK has taken the -message off its one-slot input channel and pushed its single fragment into the -8-slot `real_sender` — not when the packet is emitted. - -**3. Why `MAX_CONCURRENT_LOOKUPS` and `outgoing` are not the limiting factor -(correction to the original filing).** The hub's reply hand-off is *not* -rate-limited. `sender.send_reply(tag, frame)` puts an `InputMessage::Reply` on the -capacity-1 input channel; `InputMessageListener::on_input_message` handles that -variant with a single `reply_controller_sender.send_reply(...)` on an **unbounded** -channel and returns (`acknowledgement_control/input_message_listener.rs:60-71`, -`:150-158`). So the driver empties its 64-deep `outgoing` channel in microseconds, -each lookup task's `MAX_CONCURRENT_LOOKUPS` permit is released almost immediately -(`hub/src/nym.rs:183-196`), and the pipeline recycles far faster than the attacker -can feed it. **The limiting factor is the attacker's own delivery rate**, and the -backlog is built inside the SDK's `ReplyController`/transmission buffer, not -inside zeronym's channels. The corrected arithmetic is in the Attack Scenario; the -original "one burst of 64 lookups puts 2,624 packets into the SDK" was right about -the packet count of 64 replies but wrong to treat 64 as the cap. - -**4. The 41x fill/drain mismatch, from the pinned SDK -(`real_traffic_stream.rs:461-479`):** - -```rust - match Pin::new(&mut self.real_receiver).poll_recv(cx) { - Poll::Ready(Some((real_messages, conn_id))) => { - self.transmission_buffer.store(&conn_id, real_messages); - let real_next = self.pop_next_message().expect("Just stored one"); - Poll::Ready(Some(StreamMessage::Real(Box::new(real_next)))) - } - Poll::Pending => { - if let Some(real_next) = self.pop_next_message() { ... } - } - } -``` - -One tick stores a whole message and emits one packet. `pop_next_message` -> -`pop_next_message_at_random` -> `pop_front_from_lane` -(`transmission_buffer.rs:170-178`), strict FIFO within a lane, and there is one -lane: `TransmissionLane::General` for both `send_reply` -(`nym-sdk/src/mixnet/traits.rs:126`) and `send_message` (`:80`). -`TransmissionBuffer` has no size limit (`transmission_buffer.rs:39-49`); -`prune_stale_connections` only evicts a lane idle for ten minutes, which the -General lane is not during the two-minute window that matters. - -**5. Nothing but a real inbound message can clear the state -(`hub/src/nym_driver.rs:494-503`):** - -```rust - messages = client.wait_for_messages() => match messages { - Some(messages) => { - for message in messages { - // Counted BEFORE `deliver` filters ... - inbound_total += 1; - deliver(&incoming, message).await; - } -``` - -Loop cover traffic never reaches this point -(`received_buffer.rs:66-73`, `if nym_sphinx::cover::is_cover(fragment_data) { ... -return None }`), and acknowledgements are routed to a separate channel by the -gateway client (`packet_router.rs:39-73`), so they never become reconstructed -messages either. - -**6. Why the attacker must over-provision SURBs, and why that is free.** If the -hub's SURB pool for a tag falls below `min_surb_threshold + buffer` it queues an -`AdditionalReplySurbs` request (`receiver_controller.rs:179-216`, -`:322-345`) on its own transmission lane, and -`pick_random_small_lane` (`transmission_buffer.rs:149-157`, "small" = fewer than -100 items) makes that short lane preempt the General backlog — so it would be -emitted promptly, and a stock client answering it would produce an inbound message -at the hub, resetting `silent_rounds`. With the defaults -(`min = 10`, `max = 200`, `buffer = 0`, `config-types/src/lib.rs:48-50`), a lookup -carrying 51-60 SURBs leaves 10-19 in the pool after its 41-packet reply, which is -at or above the threshold, so no request is made. This is precisely the -provisioning an honest shim uses, so the attack traffic is indistinguishable from -conforming traffic. - -**7. A `Step::Silent` teardown always counts, and the counter is monotone in -practice (`hub/src/nym_driver.rs:530-570`):** - -```rust - Step::Died | Step::Silent => { - let silent = matches!(step, Step::Silent); - if silent { - client.disconnect().await; - } else { ... } - status.set_died(); - let lived = connected_at.elapsed(); - if silent || lived < STABLE_LIFE { - failures.short_lives += 1; - ... - } else { - failures = Failures::default(); - } -``` - -**8. And five of them is the fleet kill (`hub/src/nym_driver.rs:240-245`, -`:286-306`):** - -```rust - fn exhausted(&self) -> bool { - self.rebuilds >= REBUILDS_BEFORE_NEW_IDENTITY - || self.short_lives >= SHORT_LIVES_BEFORE_NEW_IDENTITY // 5 - } -``` - -```rust - if failures.exhausted() { - storage = Ephemeral::default(); - gateway = None; -``` - -**9. The reply that manufactures the backlog costs the hub 41 packets and the -attacker ~52 (`hub/src/nym.rs:232-234`, `:265-272`):** - -```rust -fn is_lookup(frame: &[u8]) -> bool { - frame.len() == wire::LOOKUP_BYTES && wire::peek_lookup_nonce(frame).is_some() -} -``` - -```rust - if hash.is_empty() { - tracing::warn!(reason = "empty lookup key", "lookup refused"); - return Some(error_reply(nonce)); - } -``` - -`error_reply` is `encode_lookup_reply(&nonce, &LookupReply::Error)`, which pads to -`FRAME_BYTES` = 65,536 (`hub/src/wire.rs:476-501`) with **no indexer dial and no -queue scan**, so the concurrency semaphore is never the constraint. - -### Relationship to the already-filed issues, and why this is not a duplicate - -- `hub-nym-lookup-flood-starves-gettransaction-fleet-wide.md` (confirmed, Medium) - owns points 1-2 above and the **transient starvation** they cause. Its Impact - section correctly states, for the attack shape it describes, that *"migrations - themselves keep working during this attack, which bounds the severity"*. That - bound holds for a **sustained** flood, which keeps `inbound_total` moving and so - keeps the hub alive. It does not hold for the **pulsed** shape described here. -- `hub-nym-driver-automatic-fresh-identity-permanently-invalidates-every-shim.md` - (confirmed, Medium) owns the terminal state and the instrumentation gap. Its - triggers are environmental (a gateway outage, a nym-api outage, an adversary who - can degrade one publicly-addressed Nym node). This issue supplies a strictly - weaker adversary — no mixnet position at all — and a code defect (the predicate - measured on the wrong queue) that its triggers do not involve. -- `hub-surb-starved-lookup-replies-...-oom-the-enclave.md` (confirmed, High) - reaches the same terminal state via OOM and process restart. **Its headline - remediation does not fix this one, and makes this attack's input the only - answerable one — see the Recommendations.** -- `shim-nym-driver-liveness-selfheal-resets-instead-of-deferring.md` (plausible) - is the same mechanism on the shim, failing in the **opposite** direction (the - guard is present but resets instead of deferring, so a flood *prevents* the - self-heal). - -### Checked and does *not* apply to the shim - -The shim's equivalent burst — fill its submit pipeline with junk -`SendTransaction` bodies, then go quiet — puts hundreds of packets of emission -into the shim's SDK, past the shim's own stated bound (*"two silent rounds is -120 s of drain at the throttled rate -- far more than any residual it could be -holding"*, `shim/src/nym_driver.rs:305-311`; a full pipeline is closer to 270 s, -so that premise is quantitatively wrong by ~2.3x and the comment should be -corrected). But it does **not** produce a false Silent on the shim, because each -junk submit carries 13 reply SURBs and the hub answers each with a 1-packet -`AckV1`; that backflow arrives throughout the drain window and advances the shim's -`inbound_total` (`shim/src/nym_driver.rs:393`). The hub has no analogous backflow: -a SURB reply generates nothing in return. Separately, a shim that did go falsely -Silent would only rebuild through its supervisor and reroll an identity it rotates -by design, so there is no fleet-invalidating consequence on that side. - -## Recommendations - -**These must ship together with the SURB requirement recommended by the confirmed -SURB-starvation OOM issue. The two remediations are anti-correlated:** that -issue's headline fix — answer a lookup only if it carried enough reply SURBs to -carry its own padded reply — closes the SURB-starved OOM **by making -well-provisioned lookups the only answerable ones**, and a well-provisioned lookup -is exactly the input this attack sends. Fixing either alone moves the attack -rather than removing it. Present the SURB requirement, a per-tag in-flight reply -bound, and a backlog-aware liveness verdict as **one** change, not three options. - -1. **Make the silence verdict about the right queue.** `outgoing.len() == 0` (the - conjunct the shim has) is necessary but not sufficient here, because the hub's - own channel drains in microseconds while the SDK holds minutes. The SDK exposes - `MixnetClient::shared_lane_queue_lengths()` - (`sdk/rust/nym-sdk/src/mixnet/native_client.rs:259-261`); gate the silence - verdict on the General lane being empty or below a small threshold, so "silent" - once again means "the gateway is not delivering to us". -2. **Defer, do not count, and do not reset.** Under a backlog the round should be - skipped entirely, leaving `silent_rounds` unchanged — the fix - `shim-nym-driver-liveness-selfheal-resets-instead-of-deferring.md` recommends, - applied to both binaries. -3. **Stop handing the SDK more than it can emit.** The `in_flight` boolean is not - backpressure: it clears in microseconds while the emission it authorised takes - ~5 s. Gate `outgoing.recv()` on the SDK's General lane length instead. This is - the single change that also makes `REPLY_DEADLINE` and `MAX_CONCURRENT_LOOKUPS` - operative again and bounds the transmission-buffer growth that the confirmed - lookup-flood and OOM issues both end in. -4. **Do not spend a full 64 KiB frame on a request no conforming shim sends.** The - `decode_lookup` failure arm and the `hash.is_empty()` arm - (`hub/src/nym.rs:258-272`) are the cheapest way to manufacture backlog. The - padding exists to hide `Found` from `NotFound`; these two are not on that axis. - Drop them silently, exactly as the submit arm already drops a frame with no - recoverable nonce (`hub/src/nym.rs:327-333`). -5. **Decay `short_lives`, or bound the fallback in wall-clock time.** A counter - that only resets on a rare event is a counter that only goes up. Reset or age - it out after any period longer than `STABLE_LIFE` in which the client stayed - connected, and require the five strikes to fall inside a bounded window. -6. **Make the identity change require a human, or at least make it visible.** A - fresh identity is a fleet-wide, irreversible action; gate it behind an explicit - operator opt-in (`ZIH_ALLOW_FRESH_IDENTITY`), and in either case put an - `identity_generation` counter on `/nym-status` beside `client_deaths`. -7. **Correct the log line at `:426-431`.** It asserts "registered but not being - delivered to" from a predicate that cannot distinguish that from "our own queue - is long". -8. **Test it.** Nothing in `hub/tests/` exercises `run_driver`'s probe arm or the - short-life accounting (`grep -rn "run_driver\|silent_rounds\|short_lives" - hub/tests/` returns nothing); `hub/tests/nym_identity.rs` only pins that - `Ephemeral::default()` yields new keys. A test that feeds `outgoing` and asserts - `silent_rounds` across ticks would catch both this and the shim's mirror defect. - -## Validation Information - -**Verdict: CONFIRMED as a real defect. Severity: High -> Medium (top of Medium), -for the double-counting reason set out below. The mechanism is confirmed in full; -the arithmetic and two impact claims were corrected.** - -### The defect, verified line by line - -- `hub/src/nym_driver.rs:421` is `Some(mark) if inbound_total == mark =>` with no - second conjunct. `shim/src/nym_driver.rs:312` is - `Some(mark) if seen == mark && out_frames.len() == 0 =>` with the twelve-line - rationale quoted above. **The asymmetry is real and verbatim.** -- `:558` `if silent || lived < STABLE_LIFE { failures.short_lives += 1 }` and - `:568` `else { failures = Failures::default() }`. `failures` is assigned in - exactly three places in the file: `Failures::default()` at initialisation - (`:269`), `+= 1` on a failed connect (`:314`) and on a short/silent life - (`:559`), and the two resets — after the fallback fires (`:305`) and in that - `else`. **`short_lives` has no decay path.** A hub whose client stays connected - for a week, or whose clients only ever die silently, never clears it. Confirmed - as claimed, and it is what turns a transient annoyance into permanent progress. -- `:240-245` `exhausted()` is an OR, so five short lives is independent of the - sixty-connect-failures path. `:286-306` replaces `storage` with - `Ephemeral::default()` and drops the gateway pin. Confirmed. - -### The SDK mechanism, verified against the pinned tree at `451c2aa` - -The tree was read directly, not inferred from zeronym's comments about it. - -1. `traits.rs:80` and `:126`: `send_message` and `send_reply` both set - `let lane = TransmissionLane::General;`. **Probe and replies share one lane.** -2. `transmission_buffer.rs:170-178`: `pop_front_from_lane` is a `VecDeque` - `pop_front`. **Strict FIFO within the lane**, so a probe appended after 1,100 - reply packets is emitted after them. `:39-49`: no capacity anywhere. - `:149-157`: `pick_random_small_lane` prefers lanes with fewer than 100 items, - which is what makes a SURB-request lane preempt the backlog (§6 above) and is - why the attacker must over-provision. -3. `real_traffic_stream.rs:461-479`: one `real_receiver` poll per Poisson tick, - `transmission_buffer.store(&conn_id, real_messages)` for the whole fragment - vector, one packet emitted. `config-types/src/lib.rs:25` gives the 20 ms base - delay and `sending_delay_controller.rs:23` `MAX_DELAY_MULTIPLIER = 6`, so the - throttled rate is 8.33 packets/s and the unthrottled one 50/s. The project's - own comment (`shim/src/nym.rs:1071-1077`) records that deployed enclaves sit at - multiplier 6, which is the case favourable to the attacker. -4. `received_buffer.rs:66-73`: cover traffic returns `None` before reconstruction. - `gateway-client/src/packet_router.rs:39-73`: acks go to `ack_sender`, messages - to `mixnet_message_sender`; they never mix. **Nothing but a real inbound - message advances `inbound_total`.** -5. `receiver_controller.rs:179-216` plus `config-types/src/lib.rs:48-50` - (`min = 10`, `max = 200`, `buffer = 0`): with >= 51 SURBs attached, the pool - after a 41-packet reply is >= 10, `is_below_required_surbs` is false and no - `AdditionalReplySurbs` request is emitted. `check_surb_refresh` - (`:760-808`) only fires on a key-rotation change, which is an epoch-scale - event. **So a well-SURBed attacker genuinely generates no inbound traffic at - the hub**, which is the precondition the whole attack rests on. This was the - most likely way for the attack to fail and it does not. -6. `nym-node/src/node/mixnet/handler.rs:281-320`: irrelevant to this issue but - checked while validating the sibling — acks are forwarded by the destination - gateway regardless of client state. - -### Corrections made to the filing - -- **The cost arithmetic was wrong and has been replaced.** The original said - "~1,400 sphinx packets" per pulse from "two or three clients ... under two - minutes", and treated `MAX_CONCURRENT_LOOKUPS` + `outgoing` as a 128-reply - reservoir. In fact the hub's reply hand-off is unbackpressured - (`input_message_listener.rs:60-71` forwards `InputMessage::Reply` on an - unbounded channel and returns), so neither bound limits anything; the limiting - factor is the attacker's own emission. Corrected model: - `backlog_growth = 41*lambda - 8.33 packets/s`, needing `lambda > 0.21 - lookups/s`; **one stock throttled client is not enough**, two to five are, and - the per-pulse cost is ~2,400-3,900 packets rather than 1,400. The total (five - pulses, ~12,000-20,000 packets) is close to the original estimate. -- **"Nothing warns anybody" was too strong and has been replaced.** During each - pulse `mixnet_connected` flaps to false and `client_deaths` increments, and - `OPERATORS.md:186` tells operators to alert on `/nym-address` changing. The - accurate and still-damaging statement, now in the Impact section, is that the - runbook classifies climbing `client_deaths` as benign gateway churn and names - `consecutive_rebuild_failures` as the fresh-identity predictor — and that - counter provably reads 0 for the whole walk because every cycle connects and - `NymAddress::set` zeroes it (`hub/src/server.rs:127-137`). -- **The claim that the missing conjunct is the fix was softened.** Adding - `outgoing.is_empty()` to the hub would not stop this attack, because that - channel drains in microseconds. The filing already said so in its - recommendations; the Description now says it up front so no reader takes the - one-line diff as sufficient. -- The "~130 s of quiet" figure was checked against the tick sequence and is right: - the first quiet tick resets (traffic arrived during the burst) and re-marks, and - two further ticks are needed, so the window is ~120-180 s depending on tick - phase. - -### Exploitability assessment - -A real-world attacker can meet every precondition. The target address is published -by design at an endpoint whose purpose is to publish it; the enclave takes -`ingress 0.0.0.0/0`; there is no ACL, no rate limit and no credential requirement; -the traffic is byte-identical in shape and SURB provisioning to what a conforming -shim sends; and `/nym-status` gives the attacker a free progress counter. The -resource cost is tens of megabytes across a handful of free clients. - -The two things the attacker does *not* control are the honest bounds, and they are -real: five ~135-second windows with no shim traffic reaching the hub, and no -intervening non-silent client death after `STABLE_LIFE`. On today's deployment -(one hub, a pilot fleet, no observed migrations over Nym) those are easy; on a -busy fleet where every wallet `GetTransaction` behind every shim is a reset, they -are materially harder and force the attacker to work at quiet hours and retry. -That is what separates this from the confirmed SURB-starvation OOM, which needs -one client, no timing, and works regardless of traffic. - -### Severity: why Medium, and why not High or Low - -Per the coordinator's instruction, the **trigger** is graded here and the -**consequence** is not re-counted. - -- The terminal state — fresh identity, fleet-wide strand, silently destroyed - acknowledged migrations, >1 h multi-party recovery — is already owned by - `hub-surb-starved-...-oom` (High) and - `hub-nym-driver-automatic-fresh-identity-...` (Medium). Grading a third route at - High would count one outage three times, which is exactly the reasoning the - validator of the fresh-identity issue applied when deflating it. -- Graded on its own, the trigger is: an unauthenticated stranger can, for tens of - megabytes, tear down the hub's mixnet client at a moment of their choosing, - destroy whatever replies were queued behind their burst, and bank irreversible - progress toward a fleet-invalidating action. That is a repeatable, - chosen-moment denial of service with collateral loss — Medium under this audit's - scale ("Causes DoS with significant effort"). -- Not Low, because the effort is genuinely small, the progress is permanent, the - attacker gets a progress oracle, and the code defect is unambiguous. - -**Remediation priority is higher than the severity label implies, and the report -must say so.** The confirmed High's recommended fix (require a lookup to carry -enough reply SURBs for its own padded reply) makes well-provisioned lookups the -only answerable ones, which is precisely this attack's input. Shipping that fix -alone removes one route to the fleet kill and leaves this one fully open. The SURB -requirement, a per-tag in-flight reply bound, and a backlog-aware liveness verdict -are one change. - -### False-positive checks applied - -- *§1 Assumption an attacker cannot violate?* No — the assumption "two probe - rounds with no inbound means the gateway is not delivering to us" is violated by - an ordinary send backlog, and the sibling binary's source says so explicitly. -- *§5 Impractical resource exhaustion?* No. The cost is tens of megabytes from - free clients against a target with no rate limit, and the amplification is - ~1:41 in packets (52 attacker packets buy 41 packets of a scarce, serialised, - shaped emitter). -- *§6 Intentional design?* The fresh-identity escape hatch is by design; the - predicate that walks to it is not, and the shim's copy of the same predicate - documents the defect as a defect. -- *§9 Obviously broken functionality?* No — the false Silent requires a backlog - that ordinary sparse traffic never produces, which is why the deployment works. -- *Double counting?* Explicitly avoided; see the severity section and the scope - note in Impact. - -DO NOT DEVIATE FROM THIS TEMPLATE (`docs/SECURITY-ISSUE-TEMPLATE.md`). -LEAVE THIS MESSAGE IN PLACE SO THAT ALL AGENTS KNOW NOT TO DEVIATE FROM THIS TEMPLATE (`docs/SECURITY-ISSUE-TEMPLATE.md`). diff --git a/zeronym-22aa9851caf68-high-medium/medium/hub-lookup-fall-through-hands-every-wallets-txid-to-whichever-indexer-the-hub-is-pointed-at-which-the-shipped-config-makes-the-operator.md b/zeronym-22aa9851caf68-high-medium/medium/hub-lookup-fall-through-hands-every-wallets-txid-to-whichever-indexer-the-hub-is-pointed-at-which-the-shipped-config-makes-the-operator.md deleted file mode 100644 index 46324aad..00000000 --- a/zeronym-22aa9851caf68-high-medium/medium/hub-lookup-fall-through-hands-every-wallets-txid-to-whichever-indexer-the-hub-is-pointed-at-which-the-shipped-config-makes-the-operator.md +++ /dev/null @@ -1,300 +0,0 @@ -# The hub's `GetTransaction` fall-through delivers every wallet's txid to whichever indexer the hub is pointed at, and the shipped example points it at the same host the shim fronts - -**Severity**: Medium -**Validation Status**: Confirmed -**Location**: `audit-target/zeronym/hub/src/server.rs:296-322` (`Hub::lookup`, queue-then-indexer, no cache); `audit-target/zeronym/hub/src/chain.rs:212-249` (`get_transaction` sends the wallet's `TxFilter.hash` to every configured endpoint); `audit-target/zeronym/shim/src/intercept.rs:229-236, 295-323` (every lookup is routed to the hub, no cache); `audit-target/zeronym/deploy.env.example:16-17` (`BACKEND`) versus `:22-23` (`INDEXERS`); `audit-target/zeronym/deploy.sh:110-116`; `audit-target/zeronym/smoke-local.sh:39-40`; `audit-target/zeronym/hub/deploy/caution/OPERATORS.md:281`. Claims contradicted: `audit-target/zeronym/README.md:28`, `audit-target/zeronym/shim/src/intercept.rs:231-232`, `audit-target/zeronym/hub/REVIEW.md:97-99` (REVIEW #7). -**Found by agent:** Global (focus area G26 "isolate a target's migration into a batch of one", reached from G14's end-to-end pipeline trace); validated and re-scoped by the Issue Validator -**In scope of audit?** Yes - -## Description - -The shim intercepts **every** `GetTransaction` and routes it to the hub. Its own -doc comment states the property this is supposed to buy -(`shim/src/intercept.rs:231-235`): - -> With a hub configured, EVERY `GetTransaction` is answered via the hub and NONE -> reaches the operator's indexer. - -and `README.md:28` sells it to users: - -> **`GetTransaction`.** The shim routes every lookup to the hub, so the operator -> cannot watch a wallet fetch back the transaction it just diverted. This moves -> the lookup to the hub rather than making it private outright. - -The hub does not hold the lookup. `Hub::lookup` checks its in-RAM queue first, and -**on a miss forwards the wallet's `TxFilter.hash` verbatim to every endpoint in -`ZIH_INDEXERS`** (`server.rs:305`, `chain.rs:221-249`). The queue answers only in -the window between admission and the next flush; after publication — the entire -period a wallet spends confirming, and the only period in which librustzcash's -`TransactionDataRequest::Enhancement(txid)` fires at all (`hub/REVIEW.md:99`) — -every lookup misses and is forwarded. Neither side caches: the shim is stateless -by design (`intercept.rs:193`) and `Hub::lookup` dials the indexer on every miss, -so `n` wallet polls produce exactly `n` upstream queries naming the same txid. - -`chain.rs:212-220` states the consequence itself, accurately, while being silent -about who the endpoint is: - -> The bytes are passed through unmodified … so behaviour is identical to a direct -> query by construction. - -That is the problem. If the endpoint is an indexer the operator runs, the query -the shim refused to send them arrives at them anyway — byte-identical, one Nym -round trip later, with only the source IP removed. - -**The shipped example makes the two the same host.** `deploy.env.example` — the -single file `deploy.sh` reads for *both* components (`deploy.sh:110-116`) — sets -`BACKEND=66.241.124.200:443` / `BACKEND_TLS=na.zec.rocks` (`:16-17`) and -`INDEXERS=66.241.124.200:443` / `INDEXER_TLS=na.zec.rocks` (`:22-23`). Same IPv4, -same port, same certificate name. `smoke-local.sh:39-40` defaults identically. -Nothing in `deploy.sh`, `hub/src/config.rs`, or any document requires them to be -different parties; `hub/deploy/caution/OPERATORS.md:281` describes `ZIH_INDEXERS` -only as *"your indexers; every batch member goes to all of them"* — a broadcast -description for a value that also receives every wallet's lookups. - -This uses **no defect**. Every component behaves exactly as written. - -## Attack Scenario and Steps - -Attacker: the party that operates the indexer named in `ZIH_INDEXERS`, in the case -where that party also fronts a participating shim — i.e. owns the shim's Nitro -parent host and therefore sees wallet source IPs at the TCP layer, which the -threat model records as the primary adversary's standing capability. Wholly -passive; nothing is injected and no fee is paid. - -1. Wallet `W` submits an Orchard-touching transaction. The shim diverts it, so the - operator's indexer sees **no** `SendTransaction`. That absence, against a - connection whose source IP the parent host reads, is the residual `README.md:33` - and `hub/REVIEW.md:181` both concede: *the operator learns that client IP C - submitted an Orchard-touching transaction at time T.* -2. The hub publishes the batch at the next cadence height — through - `ZIH_INDEXERS`, i.e. through this same operator, who therefore already holds - every batch member's plaintext (`chain.rs:176-210`; a conceded residual). - What they do **not** have is which member is `W`'s. -3. `W`'s wallet begins fetching the transaction. Each fetch is a `GetTransaction` - to the shim, diverted to the hub over Nym; post-flush every one misses the - queue and becomes `GetTransaction(TxFilter { hash: T })` from the hub's fixed - IP to the operator's indexer, over TLS the operator terminates. -4. The operator now holds two time series they own both ends of: on their own - parent host, `W`'s request/response bursts inside the wallet's TLS; on their own - indexer, per-txid lookup arrivals offset by one Nym round trip (measured 9–10 s - unary, `batcher.rs:45-47`). Matching the two by **phase and count over repeated - polls** names `T` as `W`'s transaction. -5. `T` is on chain; the operator reads `valueBalanceOrchard` off it. IP → transaction - → balance, complete. - -**Attack Requirements and Assumptions:** -- The operator must be both the shim's `BACKEND`/parent host **and** (one of) the - hub's `ZIH_INDEXERS`. The shipped example sets `BACKEND` and `INDEXERS` to the - same host, port and TLS name, and the product is aimed at exactly the handful of - public indexer operators, so the conjunction is realistic — but it is a - **deployment property under the hub operator's control**, not something the code - forces. -- Conditional on wallets issuing `GetTransaction` at all. `server.rs:288-296` is - the project's own model (*"Wallets poll on multi-second intervals and tolerate a - transient NOT_FOUND"*), and `hub/REVIEW.md:99` cites librustzcash's - `Enhancement(txid)` firing once the scan sees the transaction mined. Either way - the query happens; the number of samples decides how sharp step 4 is. -- Retrospective: the operator's own logs plus the public chain suffice, so the join - can be made at any later date. - -**Honest bound on step 4, corrected during validation.** The filed issue claimed -the two envelopes share *"the same start, the same stop, the same period and the -same count"*. Start and stop are **not** discriminating: every member of a batch -becomes queryable at the same flush and confirms in the same block or nearby, so -all `k` lookup streams begin and end together. The discriminators are **poll phase -and period**, matched through Nym's jitter. With a handful of polls and low -concurrency this is decisive; with a large batch of same-app wallets it requires -averaging over many samples. It is a correlation channel with an error rate, not a -naming primitive. - -## Impact on Users - -- **The product's core linkage, reached without any defect.** Where the - co-location holds, the operator obtains IP → on-chain transaction → balance for - their own wallets, passively and retrospectively. `hub/REVIEW.md:97-99` (REVIEW - #7) states the harm in the project's own words: *"the txid completes the link and - bypasses Nym, the TEE and the batch in one query"*, and mandates that *"the shim - must never issue a txid-specific query to its backing indexer for a diverted - migration."* The shim obeys; the hub then issues that query to the same indexer - on the shim's behalf. -- **Independent of the mechanisms the project is relying on to improve.** - `README.md:34` states the remedy for the batching residual as *"the lever is - adoption, not code"*. This channel names one transaction out of the batch by - query correlation rather than by set-membership arithmetic, so raising `k` - weakens it only through concurrency, not through anonymity-set size. It also - **survives ZIP 318 conformance**: the already-filed core-linkage chain selects - within a batch by `(length, anchor, expiry)`, and those become uniform exactly - when wallets do what `README.md:69` asks; this channel uses no transaction shape - at all. -- **`README.md:28` is an overclaim for the shipped configuration.** "The operator - cannot watch a wallet fetch back the transaction it just diverted" is false when - `ZIH_INDEXERS` is the operator: they see the fetch, just not the wallet's IP and - not on the wallet's connection. The hedge that follows ("rather than making it - private outright") does not tell a reader that the query may be delivered back to - the party it was routed away from. -- **Undisclosed third-party export.** A shim run by operator `B` sends its wallets' - lookups to the hub, which forwards them to whatever `ZIH_INDEXERS` names — a - value `B` cannot observe, that is not in `B`'s attestation, and that no document - discloses. `shim/ENDPOINTS.md:167-175` reasons about this trust shift for a - *planned, Zeronym-operated* indexer and argues it is acceptable because it is *"a - gain versus the operator (who additionally sees the on-chain publication and - could correlate)"*. The shipped `ZIH_INDEXERS` is not that indexer, and can be - precisely the operator that analysis excluded. - -## Technical Details / Code Analysis - -The fall-through (`hub/src/server.rs:296-312`): - -```rust - pub async fn lookup(&self, wire_hash: &[u8]) -> LookupOutcome { - if let Some(bytes) = self.queue.find_by_txid(wire_hash) { - tracing::debug!(source = "queue", "transaction lookup answered"); - return LookupOutcome::Found { data: bytes, height: 0 }; - } - - match self.chain.get_transaction(wire_hash).await { - Ok(TxLookup::Found { data, height }) => { ... } -``` - -and what `chain.get_transaction` does with it (`hub/src/chain.rs:221-232`): - -```rust - pub async fn get_transaction(&self, wire_hash: &[u8]) -> Result { - let calls = self.endpoints.iter().map(|addr| { - let filter = TxFilter { - block: None, - index: 0, - hash: wire_hash.to_vec(), - }; -``` - -`self.endpoints.iter()` — the hash goes to **every** configured endpoint, so -adding endpoints widens the disclosure linearly. This is the same `endpoints` list -`broadcast_batch` publishes through (`chain.rs:176-210`), so the lookup recipient -is by construction the party that already saw the whole batch in plaintext; the -new information is the *per-txid, per-poll timing*, not the content. - -The window in which the queue answers, and why it is the minority of a -transaction's life (`hub/src/server.rs:288-296`): - -> Note the flush-in-flight gap: `flush()` drains the queue before `broadcast_batch` -> has reached the indexer, so a lookup in that window gets a queue miss then an -> indexer NOT_FOUND … Wallets poll on multi-second intervals and tolerate a -> transient NOT_FOUND - -No cache exists on either side. The shim's `get_transaction` -(`shim/src/intercept.rs:295`) calls `diversion.hub.get_transaction(&filter.hash)` -unconditionally for every request, and `intercept.rs:193` states *"Nothing is -recorded: the shim keeps no map of what it diverted."* `Hub::lookup` has no -memoisation of any kind. - -The configuration, from the single file both components are deployed from -(`audit-target/zeronym/deploy.sh:110-116`): - -```sh - : "${BACKEND:?set BACKEND for a shim}"; : "${BACKEND_TLS:?set BACKEND_TLS for a shim}" - set -- "$@" --backend "$BACKEND" --backend-tls "$BACKEND_TLS" - ... - : "${INDEXERS:?set INDEXERS for a hub}"; : "${INDEXER_TLS:?set INDEXER_TLS for a hub}" - set -- "$@" --indexers "$INDEXERS" --indexer-tls "$INDEXER_TLS" -``` - -with `deploy.env.example:16-17` and `:22-23` both naming `66.241.124.200:443` / -`na.zec.rocks`. There is no check anywhere that `INDEXERS` and `BACKEND` are -disjoint, and no warning that they should be. - -## Recommendations - -- **Separate the publish endpoint from the lookup endpoint in configuration.** The - hub needs a broadcast path (inherently operator-visible network) and a - chain-query path (which is not). Two configuration values let an operator publish - through a public indexer while answering lookups somewhere the fleet's wallets - are not exposed to. This is the structural fix. -- **Say, where an operator will read it, that `ZIH_INDEXERS` must not be an indexer - that fronts a participating shim.** At minimum make `deploy.env.example`'s - `BACKEND` and `INDEXERS` visibly different values with a comment explaining why, - and add the requirement to the `ZIH_INDEXERS` row of - `hub/deploy/caution/OPERATORS.md:281` and to `shim/deploy/caution/OPERATORS.md`. -- **Coalesce the fall-through.** Answer repeated lookups for the same txid from a - short-lived in-hub cache so `n` polls produce one upstream query; optionally add - jitter or cross-wallet batching. All of these attack the phase-matching step, - which is what makes step 4 work. -- **Correct `README.md:28`.** The accurate statement is that the operator never - sees the query *on the wallet's connection*, and whether they see it at all - depends on a hub-side configuration value the shim operator cannot observe. -- **Disclose `ZIH_INDEXERS` to participating shim operators**, since their wallets' - lookup traffic terminates there and their own attestation says nothing about it. - -## Validation Information - -**Verdict: CONFIRMED at Medium** (downgraded from the filed High). - -**Every factual claim was checked against the target:** - -- **The fall-through exists and is verbatim.** `server.rs:296-303` returns from the - queue only on a hit; `server.rs:305` calls `self.chain.get_transaction(wire_hash)` - on every miss; `chain.rs:221-232` builds `TxFilter { hash: wire_hash.to_vec() }` - and issues it to `self.endpoints.iter()` — every configured endpoint. -- **No cache on either side.** Shim: `intercept.rs:237-323` has no lookup memo and - calls the hub on every request; `intercept.rs:193` states the shim records - nothing. Hub: `Hub::lookup` (`server.rs:296-322`) has no memoisation. Confirmed - as filed. -- **The shipped configuration claim is exactly right.** `deploy.env.example:16-17` - = `BACKEND=66.241.124.200:443` / `BACKEND_TLS=na.zec.rocks`; `:22-23` = - `INDEXERS=66.241.124.200:443` / `INDEXER_TLS=na.zec.rocks` — same host, port and - TLS name. `smoke-local.sh:39-40` defaults to the same pair. `deploy.sh:110-116` - reads both from that one file. No disjointness check exists anywhere. - `OPERATORS.md:281` documents `ZIH_INDEXERS` in broadcast terms only. -- **It is independent of batch size and survives ZIP 318** — upheld. The channel is - a per-transaction query correlation, not a set-membership argument, and uses no - transaction shape, so the wallet-side `(length, anchor, expiry)` uniformity that - closes the filed core-linkage chain leaves it open. -- **It uses no defect** — upheld. Every component does what its comments say. - -**Corrections made during validation (the filed issue overstated the correlation):** - -1. **"Same start, same stop, same period, same count" is wrong on two of four.** - All batch members become queryable at the same flush and confirm at nearly the - same height, so start and stop are shared across the whole batch and discriminate - nothing. The real discriminators are poll **phase** and **period**, recovered - through Nym's 9–10 s jittery offset over repeated samples. The issue now says so. -2. **The lookup recipient is always also the publish recipient**, because - `ZIH_INDEXERS` is the same list used by `broadcast_batch`. So the fall-through - never discloses migration *content* the endpoint did not already have; its - marginal contribution is the per-wallet timing correlation. The "cross-operator - export" harm is correspondingly narrower than filed and has been restated. -3. **`shim/ENDPOINTS.md:62-64` is not a shipped claim.** That sentence sits inside - the design section for a *planned, Zeronym-operated, non-enclaved* indexer that - does not exist in the code. It is still relevant — the residual analysis at - `:167-175` explicitly assumes the query recipient is **not** the operator, which - the shipped `ZIH_INDEXERS` can violate — but it is cited as an assumption the - deployment breaks, not as a shipped promise that is false. - -**Why Medium and not High:** - -- The severe leg (complete IP → transaction → balance) needs the conjunction - *"the hub's `ZIH_INDEXERS` is operated by the party that also fronts the victim's - shim"*. That conjunction is invited by the shipped example and is realistic given - how few public Zcash indexers exist, but it is a deployment choice the hub - operator controls and can fix in one line — it is not forced by the code. -- The unconditional leg (a third-party indexer learns "someone asked about txid T", - IP-blinded) is a genuine, undisclosed trust shift, but it is close to the residual - `shim/ENDPOINTS.md:167-175` already writes down, and `README.md:28` does hedge - ("rather than making it private outright"). -- Step 4 is a traffic-correlation step with an error rate that grows with - concurrency, not a direct read of the linkage. -- At today's adoption the modal batch is 0 or 1 (`hub/REVIEW.md:175`), so the - incremental harm now is near zero; the finding matters for the adoption regime - the project is designing toward. - -**Why not Low, and why not invalid:** - -- The mechanism is fully verified in code and in the shipped configuration file, - requires no attacker capability beyond passively reading traffic the design - delivers to them, and defeats a property the README lists under **Protected** and - that the project's own REVIEW #7 identifies as the query that "completes the - link". A privacy product handing a wallet's txid lookups back to the operator it - routed them away from is a real architectural finding, not a theoretical one. - -DO NOT DEVIATE FROM THIS TEMPLATE (`docs/SECURITY-ISSUE-TEMPLATE.md`). -LEAVE THIS MESSAGE IN PLACE SO THAT ALL AGENTS KNOW NOT TO DEVIATE FROM THIS TEMPLATE (`docs/SECURITY-ISSUE-TEMPLATE.md`). diff --git a/zeronym-22aa9851caf68-high-medium/medium/hub-nym-driver-automatic-fresh-identity-permanently-invalidates-every-shim.md b/zeronym-22aa9851caf68-high-medium/medium/hub-nym-driver-automatic-fresh-identity-permanently-invalidates-every-shim.md deleted file mode 100644 index 86207f1b..00000000 --- a/zeronym-22aa9851caf68-high-medium/medium/hub-nym-driver-automatic-fresh-identity-permanently-invalidates-every-shim.md +++ /dev/null @@ -1,458 +0,0 @@ -# The hub mints a fresh Nym identity after ~10.5 minutes of inbound mixnet silence, permanently invalidating every shim's baked-in configuration — and the counter the operator runbook names as the warning sign never moves on that path - -**Severity**: Medium -**Validation Status**: Confirmed -**Location**: `audit-target/zeronym/hub/src/nym_driver.rs:69` (`REBUILD_BACKOFF`), `:70-99` (`REBUILDS_BEFORE_NEW_IDENTITY`), `:101-112` (`STABLE_LIFE`), `:114-124` (`SHORT_LIVES_BEFORE_NEW_IDENTITY`), `:127-138` (`PROBE_INTERVAL`, `SILENT_ROUNDS_BEFORE_REBUILD`), `:225-245` (`Failures` / `exhausted`), `:279-320` (the fallback block and the connect-failure accounting), `:329-338` (address announcement), `:417-449` (the probe arm), `:525-578` (the short-life accounting), `:647-670` (`probe_send` / `send_probe`); with `audit-target/zeronym/hub/src/main.rs:205-214` (`NymAddress::set` on every published address) and `audit-target/zeronym/hub/src/server.rs:118-190` (`set`, `set_died`, `set_rebuild_failed`, `status_json`); the documentation mismatch at `audit-target/zeronym/hub/deploy/caution/OPERATORS.md:168-186` and `:221-245`; reachability and cost from `audit-target/zeronym/deploy.env.example:43-63`; user impact from `audit-target/zeronym/shim/src/hub.rs:231-241` and `audit-target/zeronym/shim/src/config.rs:252-289` (`--hub-nym` is static startup configuration) -**Found by agent:** Local (file audit of `hub/src/nym_driver.rs`) -**In scope of audit?** Yes — priority area #4 ("the mixnet transport ... identity rotation"), and the `*/src/nym_driver.rs` row of "Code Areas That Should Get Extra Attention" - -## Description - -The hub's whole reason for pinning one `Ephemeral` store across client rebuilds is -that its Nym address is *configuration for other people's enclaves*. The module -says so itself (`hub/src/nym_driver.rs:24-32`): - -> **The address survives a client death.** ... That matters because the address is -> baked into every shim's enclave config and a Caution managed app is immutable, -> so an address change costs every operator a re-assemble and redeploy. - -The file then provides an escape hatch that throws that away automatically, on a -timer, with no operator in the loop (`hub/src/nym_driver.rs:279-306`): - -```rust - loop { - // The stored registration is not coming back, whether the gateway - // refuses us outright or registers us and then drops or starves every - // client. Give up on it and mint a new identity, ... - if failures.exhausted() { - storage = Ephemeral::default(); - gateway = None; - tracing::warn!( - after_failed_connects = failures.rebuilds, - after_short_lived_clients = failures.short_lives, - "the hub's gateway registration is unrecoverable; taking a FRESH \ - identity and dropping the gateway pin. The hub's Nym address WILL \ - change: read the new one from /nym-address and re-point every shim, \ - or migrations keep failing closed" - ); - failures = Failures::default(); - } -``` - -with (`hub/src/nym_driver.rs:240-245`): - -```rust -impl Failures { - fn exhausted(&self) -> bool { - self.rebuilds >= REBUILDS_BEFORE_NEW_IDENTITY - || self.short_lives >= SHORT_LIVES_BEFORE_NEW_IDENTITY - } -} -``` - -`SHORT_LIVES_BEFORE_NEW_IDENTITY` is **5** (`:124`). A "short life" is a client -torn down for inbound silence *however long it lived* (`:558`: -`if silent || lived < STABLE_LIFE`), and inbound silence is declared after -`SILENT_ROUNDS_BEFORE_REBUILD = 2` rounds of `PROBE_INTERVAL = 60 s`. Each cycle -is ~120 s of probing plus a 5 s backoff, so **about ten and a half minutes of "no -inbound mixnet message of any kind reached this client" is sufficient to -irreversibly change the hub's Nym address and disconnect every shim in the -fleet.** - -Three properties make that expensive: - -- **automatic and unattended** — no flag, no confirmation, no operator action; -- **irreversible** — `Ephemeral::default()` mints new identity *and* encryption - keys (`hub/tests/nym_identity.rs:63-86` pins exactly this), and the old ones are - in RAM only, so the previous address can never be recovered; -- **fleet-wide** — every shim's `--hub-nym` is validated once at startup - (`shim/src/config.rs:252-289`) and never re-read; nothing in `shim/src/` fetches - `/nym-address` (verified by grep: the path string appears nowhere under - `shim/src/`), and the value is baked into an immutable Caution enclave. - -**And the signal the operator runbook tells them to watch does not move on this -path.** `hub/deploy/caution/OPERATORS.md:172` describes -`consecutive_rebuild_failures` as *"growing — it is down and not recovering; **at -60 it takes a NEW identity** and every shim needs re-pointing"*, and `:182-184` -repeats *"After 60 consecutive failed rebuilds the hub also deliberately takes a -fresh identity"*. Neither the field table nor the failure-mode list mentions the -five-short-lives path at all. But on that path every cycle **connects -successfully**, so `main.rs:213` calls `NymAddress::set`, which zeroes -`consecutive_failures` (`server.rs:127-137`). The counter the runbook names as the -predictor therefore reads **0 for the entire ten-and-a-half-minute walk to the -identity change.** The one field that does move, `client_deaths`, is characterised -in the same table as benign: *"climbing: gateway churn"*. - -## Attack Scenario and Steps - -The dominant path is **not** an attacker; it is an ordinary transient. - -**Path A — a transient inbound outage of ~11 minutes.** - -1. The hub is connected. Its inbound path stops delivering for eleven minutes. - Any of these does it: the pinned entry gateway restarts or is upgraded; the - gateway registers the client and stops delivering to it (the exact fault this - project *measured in production* — `hub/src/nym_driver.rs:345-353` and - `shim/src/nym_driver.rs:221-231` record that on 2026-08-14, **two of four** - deployed clients on identical config never received a single inbound message, - one broken three minutes after boot and still broken hours later); a mix node - on the probe's route churns; the free-tier bandwidth allowance runs out, which - `hub/deploy/caution/OPERATORS.md:197-199` documents as *"reception stops"*. -2. The self-probe (`:663-670`) is a message to the hub's own address, so it - traverses gateway -> mix hops -> gateway -> client. Nothing on that path - returning means `inbound_total` does not advance. -3. `silent_rounds` reaches 2 at ~t+120 s -> `Step::Silent` -> `client.disconnect()` - -> `failures.short_lives += 1` -> 5 s backoff -> rebuild on the same storage. -4. The condition persists, so the cycle repeats. Each cycle is ~125 s. -5. At the fifth, `exhausted()` is true at the top of the outer loop. Fresh - identity, gateway pin dropped, new address. -6. **The outage clears — or, more likely, the fresh identity's new gateway draw - simply works, because dropping the pin is what the fallback does.** The hub is - perfectly healthy, at an address nobody has. - -Step 6 is the sting: the fallback is *most* likely to succeed exactly when the -fault was localised to one gateway — i.e. exactly when the old identity would -also have recovered on its own. - -**Path B — a transient connect failure.** `REBUILDS_BEFORE_NEW_IDENTITY = 60` -with `REBUILD_BACKOFF = 5 s` per cycle; the constant's own doc (`:83-90`) measures -~370 s against a locally refused port and warns the real figure is longer against -a remote gateway that times out. `deploy.env.example:43-63` makes the timing-out -case the deployed one: the nym-api endpoints are allowlisted **by IP snapshot**, -blocked packets are *dropped* rather than refused so each blocked attempt "waits -out the SDK's full 30 s timeout", and the same comment says those IPs *"are DNS -snapshots, not configuration — re-derive them ... on every redeploy **or they rot -silently**"*. A nym-api IP rotation or outage therefore produces a run of 30-90 s -connect failures; 60 of them is ~35-90 minutes, and then the hub mints a new -identity **which cannot help**, because the blockage is the topology fetch, not -the registration — and it will mint another one every 35-90 minutes until the -blockage clears. The marginal harm is on the transient case: without the -fallback, a nym-api outage that clears leaves the hub reconnecting at the *same* -address and the fleet recovering with zero human action; with it, the hub returns -at a new one. - -**Path C — an adversary.** Attacker class #5 in `AUDIT-INSTRUCTIONS.md` ("mixnet -parties: gateways, mix nodes, and anyone able to degrade the shim->hub path"). The -hub deliberately pins **one** entry gateway for the life of its identity -(`:170-177`), and that gateway is a public Nym node with a published IP the -enclave's egress rule names explicitly. Eleven minutes of interference against -that one IP — or *being* that gateway, accepting the registration and dropping -inbound, which is free and undetectable — forces the fleet-kill, repeatably. The -attacker needs no credentials, no ACL bypass and no access to any enclave. -(Separately, `hub-surb-starved-lookup-replies-...-oom.md`, already confirmed, -hands an anonymous outsider an on-demand trigger for the same terminal state via -a process restart. That issue owns that impact; it is cited here only so the -report does not present the two as unrelated.) - -**Attack Requirements and Assumptions:** - -- Paths A and B require **no attacker** — only a Nym gateway or nym-api endpoint - unavailable for a window the project's own field notes show is routine. -- Path C requires only the ability to degrade one publicly-addressed Nym node's - delivery path for ~11 minutes, or to be that node. -- Recovery requires **every third-party operator** to read the hub's new - `/nym-address`, re-run `assemble-caution.sh`, and redeploy an immutable - enclave. `hub/deploy/caution/OPERATORS.md:226-245` budgets this at ~25 minutes - per component and *"well over an hour"* end to end, and step 4 of its own - runbook is *"Send it to every shim operator. There is no discovery mechanism; - the handoff is a human message."* - -## Impact on Users - -While the fleet holds a stale address: - -1. **Every diverted migration is silently destroyed, and the wallet is told it - succeeded.** `shim/src/hub.rs:231-241` returns `Submit::Accepted { txid }` the - moment the frame is handed to the in-process mixnet transport, with its own - comment recording that a hub refusal *"is never surfaced here"*. The frame goes - to a `Recipient` whose identity key no longer exists and is dropped. The wallet - holds `error_code 0` and a txid for a transaction that exists nowhere, on a - migration path that is mandatory and whose funds sit in a pool closed to new - value. -2. **Every wallet's `GetTransaction` fails.** With a hub configured the shim routes - *every* lookup to the hub (`shim/src/intercept.rs:229-236`) and fails closed on - timeout, so **no** wallet behind **any** zeronym shim can fetch a transaction's - full data — including users who have never touched Orchard. -3. **Users are pushed off the protected path.** A wallet that cannot broadcast or - confirm gets pointed at a different, unprotected indexer, where the - Orchard-touching transaction is broadcast in the clear — converting an - availability failure into the permanent, on-chain privacy leak the product - exists to prevent. In this system availability *is* a privacy property. -4. **The documented alarm does not fire in time.** The runbook's named predictor - (`consecutive_rebuild_failures` approaching 60) stays at 0; the runbook's other - instruction, *"poll `/nym-address` and alert on change"* (`:186`, `:221-222`), - is detection after the fact and is a manual instruction nothing in the - repository implements. `mixnet_connected` does flap false during each silent - round, so an operator polling it frequently would see *something* — but the - runbook's guidance for that state is "the hub is receiving nothing", not "you - have ten minutes before the fleet's configuration is void". - -## Technical Details / Code Analysis - -**The accounting, in full.** The counters are cleared only by a client that both -lived out `STABLE_LIFE` *and* was not torn down for silence -(`hub/src/nym_driver.rs:556-569`): - -```rust - let lived = connected_at.elapsed(); - if silent || lived < STABLE_LIFE { - failures.short_lives += 1; - tracing::warn!( - lived_secs = lived.as_secs(), - silent, - consecutive_short_lives = failures.short_lives, - "hub mixnet client did not prove the stored registration; \ - counting it against it" - ); - } else { - failures = Failures::default(); - } -``` - -and by the fallback itself (`:305`). A failed connect only ever increments -(`:313-320`). There is no decay and no half-life — only consecutiveness, which -`:94-98` deliberately widened to include clients that connected successfully. - -**The silence timer, step by step** (`:417-449`, with -`probe.set_missed_tick_behavior(Delay)` at `:376` and the first tick immediate): - -| t | arm | state | -|---|---|---| -| ~0 s | first tick, `inbound_at_probe == None` -> catch-all arm | `silent_rounds = 0`, probe sent; on completion (`:490`) `inbound_at_probe = Some(inbound_total)` | -| ~60 s | tick, `inbound_total == mark` | `silent_rounds = 1`, warn, probe re-sent | -| ~120 s | tick, `inbound_total == mark` | `silent_rounds = 2 >= SILENT_ROUNDS_BEFORE_REBUILD` -> `Step::Silent` | - -`Step::Silent` then runs `client.disconnect().await` (`:537`), `status.set_died()` -(`:550`), `failures.short_lives += 1` (`:559`), and `REBUILD_BACKOFF` (`:574`). -Five iterations is ~625 s. - -**Why the runbook's counter cannot warn.** On this path `build_client` *succeeds* -every cycle, so `address_out.send(address)` (`:338`) reaches `main.rs:213`'s -`nym_address.set(...)`, which sets `connected = true` **and stores 0 into -`consecutive_failures`** (`server.rs:127-137`). `set_rebuild_failed` -(`server.rs:160-167`) — the only thing that increments that counter — is reached -only from the connect-failure arm. `status_json` (`server.rs:182-190`) carries -`mixnet_connected`, `address_published`, `client_deaths` and -`consecutive_rebuild_failures`, and **no identity-change counter of any kind**: - -```rust - serde_json::json!({ - "mixnet_connected": self.0.connected.load(Ordering::Relaxed), - "address_published": self.get().is_some(), - "client_deaths": self.0.deaths.load(Ordering::Relaxed), - "consecutive_rebuild_failures": self.0.consecutive_failures.load(Ordering::Relaxed), - }) -``` - -So `/nym-status` cannot distinguish "reconnected, same address" from "the fleet's -configuration is now void" — the single most important distinction this endpoint -could make, on an enclave whose console does not exist. - -**The evidence the decision is taken on is about the whole mixnet, not the -registration.** `inbound_total` is incremented before any filtering (`:494-503`), -so SURB-replenishment artifacts and the self-probe all count — the right choice -for the liveness question, but it makes the silence verdict a statement about the -entire round trip (outbound leg, mix hops, return leg, gateway delivery) while the -action taken on it is scoped to *this client's registration* (`:363-369`). A -mixnet-wide or route-level fault is therefore attributed to the hub's own -registration and paid for with the fleet's configuration. - -Worse, an **outbound-only** fault counts as inbound silence: `send_probe` -(`:660-670`) logs and swallows a send error, and `probe_send` (`:647-652`) returns -`Sent::Probe` regardless, so `inbound_at_probe` is stamped at `:490` even for a -probe that never left. - -**Contrast with the shim, which has the guard the hub lacks.** The shim's -identical arm refuses to declare silence while its own send queue is non-empty -(`shim/src/nym_driver.rs:312`: `Some(mark) if seen == mark && out_frames.len() == 0`), -because rebuilding on it *"would disconnect a healthy client and discard that -whole queue"*. The hub's arm (`hub/src/nym_driver.rs:421`) has no such condition. -For the shim a false rebuild costs one identity that was meant to rotate anyway; -for the hub it is a strike toward invalidating every operator's enclave. The -asymmetry runs the wrong way. - -**Why "a hub restart changes the address anyway" does not excuse this.** The -constant's rationale says (`:79-82`) *"a hub restart would change the address -anyway, so this only automates what an operator would otherwise do by hand"*. A -restart is an operator action taken with knowledge, at a chosen time, with the -operator ready to notify the fleet. This is an unattended action taken by a timer -on evidence that is (a) indirect, (b) unable to distinguish a transient from a -permanent fault, and (c) unable to distinguish a fault in the hub's registration -from a fault anywhere else on the mixnet. In Path B it is additionally futile and -repeating. - -**No executable evidence exists for any of this.** `hub/src/nym_driver.rs` has no -`#[cfg(test)] mod tests` (the shim's does, at `shim/src/nym_driver.rs:657`), and -nothing in `hub/tests/` references `run_driver`, `Failures`, `short_lives` or -`probe_send`. The only thing that drives the fallback is `nymnet/probe`, which -needs a human to kill and restart a localnet gateway by hand — and -`nymnet/localnet.sh` exposes only -`up|down|status|smoke|lookup|wire|e2e|e2e-driver|clean|env` (`:248`), with **no -gateway-kill subcommand**, contradicting the citation at -`hub/src/nym_driver.rs:90-91`. Per the coordinator's verified facts no CI runs -`cargo test` at all. Both thresholds, the short-life/connect-failure split, and -the ~370 s measurement are held by prose. - -## Recommendations - -In rough order of value: - -1. **Make `/nym-status` able to say the one thing that matters.** Add - `identity_changes` (a monotone counter) and `address_generation` to - `NymAddress::status_json` (`hub/src/server.rs:182-190`) and bump them from the - fallback block (`nym_driver.rs:286-306`). This is client-lifecycle data of - exactly the kind the endpoint already carries, the address itself is public, so - it costs nothing in anonymity terms — and it is the cheapest fix in the file. -2. **Require persistence, not just consecutiveness.** Gate `exhausted()` on - wall-clock as well as counts — e.g. "the stored registration has produced no - stable client for at least N hours". Ten minutes is far shorter than the time - it takes a human to notice, and two orders of magnitude shorter than the - recovery it triggers. -3. **Make the fresh-identity fallback opt-in** (`ZIH_NYM_ALLOW_NEW_IDENTITY`, - default off) or operator-triggered. "The hub is down and will come back at the - same address" is strictly recoverable; "the hub is up at a new address" is not. -4. **Correct `hub/deploy/caution/OPERATORS.md:168-186`.** It documents only the - 60-connect-failure path and names `consecutive_rebuild_failures` as the - predictor, which stays at 0 on the faster path. Document the five-short-lives - path, and change the `client_deaths` row from "climbing: gateway churn" to - state that five consecutive silence teardowns mint a new identity. -5. **Do not attribute a whole-mixnet fault to the local registration.** At minimum - do not count a round in which `send_probe` returned an error, and add the - shim's `out_frames.len() == 0` backlog guard. -6. **Give the fleet a way to follow an address change.** `--hub-nym` already - accepts a list (`shim/src/config.rs:252-289`); a pre-published successor - address, or an address record shims could pin by hub public key rather than by - gateway-bound `Recipient`, would turn a fleet-kill into a reconnection. Until - something like that exists the address must be treated as immutable state, not - as something a timer may replace. - -## Validation Information - -**Verdict: CONFIRMED as a real defect. Severity: DEFLATED High -> Medium.** - -**Every mechanical claim re-verified against `audit-target/zeronym/`:** - -- `hub/src/nym_driver.rs:69` `REBUILD_BACKOFF = 5s`; `:99` - `REBUILDS_BEFORE_NEW_IDENTITY = 60`; `:112` `STABLE_LIFE = 3 min`; `:124` - `SHORT_LIVES_BEFORE_NEW_IDENTITY = 5`; `:133` `PROBE_INTERVAL = 60s`; `:138` - `SILENT_ROUNDS_BEFORE_REBUILD = 2`. All confirmed verbatim. The code's own - comment at `:121-123` says *"Five short lives is ~10 min"* — the project agrees - with the timeline. -- `:240-245` `exhausted()` is an **OR** over the two counters, so the - five-short-lives path is independent of the sixty-connect-failures path. - Confirmed. -- `:279-306` the fallback replaces `storage` with `Ephemeral::default()` and drops - the gateway pin; `hub/tests/nym_identity.rs:67-86` asserts two independent - stores have different identity keys. **Irreversibility confirmed.** -- `:417-449` and `:525-578` — the probe/silence timeline and the short-life - accounting are exactly as described; ~125 s per cycle, five cycles. -- `:647-670` — `send_probe` swallows the send error and `probe_send` returns - `Sent::Probe` unconditionally, so `:490` stamps the mark even for a probe that - never left. **An outbound-only fault does count as inbound silence.** Confirmed. -- `shim/src/nym_driver.rs:312` has the `out_frames.len() == 0` guard; the hub's - `:421` does not. Confirmed. -- `shim/src/config.rs:252-289` validates `--hub-nym` once at startup; - `grep -rn "nym-address" shim/src/` returns **no hits**, so no shim ever refetches - it. Confirmed. -- `shim/src/hub.rs:231-241` returns `Submit::Accepted` on hand-off to the - transport, with the comment *"a refusal is never surfaced here"*. Confirmed. -- `hub/src/nym_driver.rs:345-353` and `shim/src/nym_driver.rs:221-231` — the - 2026-08-14 field note ("two of four ... never answered") is in the source, so - the base rate of the triggering condition is the project's own measurement, not - the auditor's speculation. Confirmed. -- `nymnet/localnet.sh:248` — the subcommand list is - `up|down|status|smoke|lookup|wire|e2e|e2e-driver|clean|env`, with no - gateway-kill arm; `hub/src/nym_driver.rs` has no `#[cfg(test)]`. Confirmed. - -**One claim in the filed draft was too strong and has been replaced with a -sharper, verified one.** The draft said the identity change is *"invisible to the -health surface"*. That is not accurate: `GET /nym-address` serves the **new** -address, `client_deaths` increments once per cycle, `mixnet_connected` flaps -false during each silent round, and `hub/deploy/caution/OPERATORS.md:186` and -`:221-222` explicitly tell operators to *"poll `/nym-address` and alert on -change"*. The correct and more damaging statement, verified during validation, is -the **documentation/instrumentation mismatch**: - -- `OPERATORS.md:172` names `consecutive_rebuild_failures` and says *"at 60 it - takes a NEW identity"*; `:182-184` repeats the 60-failure story. **Neither - mentions the five-short-lives path**, which is roughly five times faster. -- On that path every cycle connects, so `main.rs:213` -> `NymAddress::set` -> - `server.rs:132-137` **stores 0 into `consecutive_failures`**. The runbook's - named predictor reads 0 for the entire walk. -- `status_json` (`server.rs:182-190`) carries no identity or generation counter, - so even after the fact `/nym-status` cannot distinguish a benign reconnect from - a fleet-invalidating identity change. - -The issue body and recommendations have been rewritten around this; recommendation -4 (correcting the runbook) is new and follows from it. - -**Why the severity is Medium and not High.** Four reasons, applied in the order -the audit's own precedents require: - -1. **The terminal harm is already owned at High by two confirmed issues, and - stacking it a third time would triple-count one outage.** - `hub-surb-starved-lookup-replies-...-oom.md` (Confirmed, High) reaches exactly - this end state — a process restart, a new address, every shim stranded, - `GetTransaction` dead fleet-wide, acknowledged migrations silently destroyed — - and its own Impact section names this file as *"the outage already filed as ... - whose stated trigger was bad luck"*. `hub-queue-unauthenticated-fill-silently-destroys-migrations.md` - (Confirmed, High) owns the silent-destruction-after-ack harm. This issue's - **marginal** contribution is a *further* trigger for the same terminal state, - plus the instrumentation gap. -2. **The marginal harm is confined to the transient case.** If the inbound silence - is genuinely persistent, the fallback is defensible and the code's rationale - holds: keeping the old identity leaves the hub off the mixnet and the fleet - equally dead, and the eventual hub redeploy changes the address anyway. The - defect is real but bounded — it is the case where the fault would have cleared - between ~10.5 minutes and whenever a human would have acted, which the fallback - converts from a self-healing outage into a multi-hour, multi-party - reconfiguration. -3. **The mechanism's existence is a documented design decision** - (`nym_driver.rs:70-99`, `OPERATORS.md:182-184`), so `AVOIDING-FALSE-POSITIVES` - §6 covers the *escape hatch*. What is **not** covered, and is the actual - finding, is the threshold (5 short lives / ~10.5 min, with no wall-clock - persistence and no decay), the attribution error (a whole-mixnet round-trip - measurement driving a local-registration decision, including probes that never - left), and the runbook naming a counter that cannot move. -4. **No confidentiality is breached directly.** The privacy harm (users pushed to - an unprotected indexer) is second-order and depends on user behaviour. - -**Why Medium and not Low.** The trigger threshold is ~10.5 minutes against a -condition the project's own production notes show hitting **two of four** deployed -clients; the blast radius is every user of every zeronym endpoint at once, -including non-Orchard users who lose `GetTransaction`; the change is irreversible -and recovery is documented by the project itself at *"well over an hour"* with a -human message as the only distribution channel; and the operator's documented -early-warning signal provably never fires. This sits at the top of Medium, and the -two highest-value fixes (a counter in `status_json`, a wall-clock gate on -`exhausted()`) are each a handful of lines. - -**One composition worth recording without re-counting it.** Coordinator item 6z -established that the hub's Nym address is an **unanchored** value: it is the -recipient's key material, published over plain WebPKI on operator-controlled DNS -at `GET /nym-address`, with the alternative distribution channel being verbatim -"a human message", and `hub/deploy/caution/OPERATORS.md:244-245` telling operators -that `caution verify` *"does NOT belong on the critical path — restore service -first, verify after"*. So a party who can force this fleet-kill also manufactures -the exact circumstance in which an unauthenticated address handoff happens under -time pressure across many organisations. That composition belongs in the report's -narrative; the substitution harm itself is already owned by -`shim-config-hub-identity-is-unattested-unobservable-operator-configuration.md` -and must not be counted again here. - -**False-positive checks applied.** - -- *§6 Intentional design?* Partly — see reason 3 above. The escape hatch is by - design; the threshold, the attribution and the missing signal are defects. -- *§5 Impractical resource exhaustion?* Not applicable to paths A and B, which - need no attacker at all. Path C is cheap rather than expensive. -- *§9 Obviously broken functionality?* No — the fallback only fires on a - condition that is uncommon per deployment even if common across a fleet, and the - project has observed the *related* address-instability problem in production and - written a failover runbook for it, which is consistent with the code as written. -- *§1 Assumption an attacker cannot violate?* The assumption "eleven minutes of - inbound silence means this registration is unrecoverable" is violated by an - ordinary gateway restart. - -DO NOT DEVIATE FROM THIS TEMPLATE (`docs/SECURITY-ISSUE-TEMPLATE.md`). -LEAVE THIS MESSAGE IN PLACE SO THAT ALL AGENTS KNOW NOT TO DEVIATE FROM THIS TEMPLATE (`docs/SECURITY-ISSUE-TEMPLATE.md`). diff --git a/zeronym-22aa9851caf68-high-medium/medium/hub-nym-identity-has-no-trust-anchor-and-the-one-the-project-already-owns-is-never-applied-to-it.md b/zeronym-22aa9851caf68-high-medium/medium/hub-nym-identity-has-no-trust-anchor-and-the-one-the-project-already-owns-is-never-applied-to-it.md deleted file mode 100644 index cbd37783..00000000 --- a/zeronym-22aa9851caf68-high-medium/medium/hub-nym-identity-has-no-trust-anchor-and-the-one-the-project-already-owns-is-never-applied-to-it.md +++ /dev/null @@ -1,384 +0,0 @@ -# The hub's Nym address is the recipient's key, and nothing anchors it: the canonical copy is fetched over plain WebPKI on operator-controlled DNS and is never compared against the attestation the project already publishes - -**Severity**: Medium -**Validation Status**: Confirmed -**Location**: `audit-target/zeronym/shim/src/config.rs:65-77` (`ZIS_HUB_NYM` is a list), `:255-289` (`hub_selection`), `:292-309` (`is_nym_address`); `audit-target/zeronym/shim/src/nym.rs:595-690` (`NymHandle::submit`), `:733-784` (`each_target`); `audit-target/zeronym/shim/src/nym_driver.rs:608-624` (`send_frame`); `audit-target/zeronym/shim/src/wire.rs:28-34` (`SubmitV1` layout); `audit-target/zeronym/hub/src/server.rs:62-70`, `:469-479` (`GET /nym-address`); `audit-target/zeronym/deploy.sh:255-259` (the bare `curl`), `:283-318`; `audit-target/zeronym/hub/deploy/caution/OPERATORS.md:131-145` (the handoff), `:186-189` (poll and alert), `:236-245` (the rotation runbook and "verify after"); `audit-target/zeronym/shim/deploy/caution/OPERATORS.md:237-239` ("the authoritative copy"), `:312`, `:370-378`; `audit-target/zeronym/README.md:71`, `:77`, `:92` -**Found by agent:** Global with focus area of G8 (hub-as-adversary / shim-as-adversary), taken with G27; recorded as coordinator item 6z -**In scope of audit?** Yes - -## Description - -Every diverted migration in the system is delivered to one value: the Nym address in -`ZIS_HUB_NYM`. That address is not a name that gets resolved and then authenticated — it -**is** the recipient's public key material (`identity.encryption@gateway`). The Sphinx -payload is encrypted end to end to the encryption key inside it, which is a genuine -strength of the transport: no mix node and no gateway can read a submission. But it -means the transport's confidentiality is confidentiality *to whoever generated that -address*. Substituting the string substitutes the key, and there is no certificate, no -signature and no attestation over it. - -The project knows this hop needs peer authentication: `OPEN-QUESTIONS.md` §3 records -that under STEVE "**the shim verifies the hub**, extracts its key, and derives a session -key. That is enough for privacy." STEVE is designed, not built (`README.md:92`), and its -absence is a stated residual that this issue does not re-report. - -What this issue reports is the **interim substitute the deployed system actually runs -on**: the address is handed over out of band, and the only mechanism the documentation -offers for checking it is unauthenticated — even though the project already owns an -authenticated one and applies it two lines away. - -1. **The documented authoritative source is plain WebPKI over operator-controlled - DNS.** `shim/deploy/caution/OPERATORS.md:237-239` calls - `https:///nym-address` "the authoritative copy"; `deploy.sh:256` fetches - it with a bare `curl -s ... "https://$TLS_DOMAIN$1"`. The presented leaf is **never** - compared against `user_data.tls.certfp` in the hub's COSE-signed attestation — the - binding coordinator item 6q established **does** exist and **does** work on this - platform, and which `caution verify` performs when it is run. -2. **No shim runbook tells a shim operator to verify the hub at all.** Every - `caution verify` instruction in `shim/deploy/caution/` targets the *shim's own* - endpoint; the hub's `caution verify` line - (`hub/deploy/caution/OPERATORS.md:100`) is a self-check for the hub operator. -3. **The alternative source is verbatim "a human message."** - `hub/deploy/caution/OPERATORS.md:143-145`: *"Hand that string to each shim operator … - There is no discovery mechanism — the handoff is a human message."* -4. **The value is required to change, on an operational schedule, with no key - continuity.** A diskless hub mints a fresh identity on restart, so - `shim/deploy/caution/OPERATORS.md:370-378` and `hub/deploy/caution/OPERATORS.md:186-189` - instruct operators to poll `/nym-address` and alert on change, then re-bake. Nothing - signs the new identity with the old one, so a shim operator has no way to distinguish - a legitimate rotation from a substitution. The rotation runbook - (`hub/deploy/caution/OPERATORS.md:236-245`) ends with "**`caution verify` … does NOT - belong on the critical path — restore service first, verify after**", i.e. the entire - fleet re-points at a fresh, unverified identity by design and the one available check - is explicitly deferred. -5. **The attestation cannot rescue it.** The shim's manifest faithfully attests whatever - `ZIS_HUB_NYM` was baked in; if the value was poisoned at source, the attestation is - *valid* and proves only that the shim will send to that address. An auditor who - performs the check that - `auditor-recipe-omits-the-two-checks-that-decide-where-plaintext-goes-and-names-a-defence-the-platform-does-not-rely-on.md` - recommends obtains a base58 string and has nothing to compare it against except the - same unauthenticated endpoint. - -Consequence: a party who controls the hub's DNS name (or obtains a certificate for it), -and the hub operator themselves, can publish an address whose key they hold and thereby -receive **every migration on the network, in plaintext, at divert time, with per-shim -sender tags**, indefinitely, with an entirely honest shim operator and nothing -observable changing anywhere. - -## Attack Scenario and Steps - -**Variant A — substituted address (the primary case).** - -1. The attacker obtains control of the hub's DNS zone or registrar, or a certificate for - the name (which follows from DNS control via ACME) — or *is* the hub operator, who - needs neither, since they author both the endpoint and the handoff message. -2. The attacker runs a stock `nym-sdk` client and serves its address `R_att` at - `https:///nym-address` (and/or sends it as the handoff message). -3. Shim operators re-assemble with `--hub-nym R_att` — which is routine, because the - address is *supposed* to change on every hub restart, and a forced restart is itself - cheap (see `hub-liveness-probe-reads-its-own-send-backlog-as-gateway-silence-so-any-stranger-can-drive-the-fresh-identity-fleet-kill.md`, - confirmed). `Config::hub_selection` (`shim/src/config.rs:255-289`) shape-checks the - string and boots. -4. Every divert now encrypts the `SubmitV1` frame to `R_att`'s key. `R_att` decodes it - (`shim/src/wire.rs:28-58`, byte-identical in `hub/src/wire.rs`, and the wire format is - published in-tree), obtaining the raw Zcash transaction, its arrival time and the - shim's stable `AnonymousSenderTag`. -5. `R_att` re-frames each transaction and forwards it to the real hub from its own - client, and relays `LookupV1`/`LookupReplyV1` in both directions. Everything downstream - is unchanged: the real hub queues, batches and publishes; wallets get their txids; - `/healthz` is 200; `/nym-status` reports `mixnet_connected: true`. The only cost is a - small added latency, well inside the shim's 90 s lookup budget. - A rogue that does not want to relay can broadcast each transaction itself and answer - lookups from its own copy; the on-chain result is a smaller batch, which is - indistinguishable from today's expected batch size of 0-1 - (`hub/src/batcher.rs:412-419` already warns about that on every honest flush). - -**Variant B — appended address.** `ZIS_HUB_NYM` is a comma-separated list and a submit -goes to **every** entry (`shim/src/nym.rs:642`, `for target in 0..targets`), so handing -out `R_real,R_att` gives the attacker a silent copy with no relaying at all. This variant -is *mechanically* the same capability that -`shim-submits-every-migration-to-every-configured-hub-…` (confirmed High) owns for a -hostile shim operator; it is noted here only because an unanchored reference value lets -someone who is **not** the shim operator plant it. One caveat that keeps Variant A the -primary case: with two addresses, `each_target` (`shim/src/nym.rs:733-784`) sends lookups -to each in turn and only a **timeout** sweeps on (`:770-781`), so a silent second address -costs ~90 s on half of all lookups — degraded latency an operator might notice, whereas -Variant A has no such tell. - -**Attack Requirements and Assumptions:** - -- The attacker needs **one** of: control of the hub's DNS zone or registrar; a - certificate for the hub's name; or to be the hub operator. It does **not** need the - shim operator to be malicious, the enclave to be compromised, the mixnet to be - compromised, or any code defect. -- A **spoofed handoff message alone is weaker than it looks** and is not the load-bearing - vector: the shim runbook designates the `/nym-address` endpoint "the authoritative - copy", so an operator who cross-checks there defeats a message-only spoof. The - substitution has to reach the endpoint (or the operator has to skip the cross-check, - which the rotation runbook's time pressure makes likelier). -- The opportunity is **recurring by design**: it exists at every hub restart, and - `hub/deploy/caution/OPERATORS.md:244-245` explicitly moves verification off the - critical path during exactly that window. -- What makes it hard to notice: there is no key continuity between an old and a new hub - identity, so a legitimate rotation and a substitution are the same event to a shim - operator; and both variants are indistinguishable from correct operation at every - surface the system exposes. - -## Impact on Users - -If it happens, every user of every shim in the fleet loses the whole protection at once: -the attacker holds the plaintext of each Orchard-touching transaction, its exact length, -its arrival instant, and a stable per-shim label. Joined against the public chain that is -`IP → transaction → balance` for any user whose shim the attacker also observes, and -`operator → transaction → balance` for the rest. Nothing a wallet, a user, an auditor -following `README.md:71`, or a shim operator following their own runbook can see -distinguishes it from correct operation. - -It also makes the reachability precondition of -`hub-unauthenticated-pre-publication-transaction-disclosure.md` real: the attacker holds -candidate txids for unpublished migrations, which per coordinator item 6z is otherwise -not reachable by any adversary other than the wallet itself. - -`README.md:77` ("Green boxes are attested enclaves, the only things that ever see a -migration in cleartext") is stated unconditionally, and `README.md:71`'s auditor recipe -lists four checks, none of which is about the destination. The "Not protected" list at -`README.md:30-36` does not mention that the hub's identity is an unauthenticated -configured value; the closest disclosure is `README.md:92`'s "Designed, no code yet: the -STEVE handshake, the encrypt-to-hub-key layer", from which a reader must infer the -consequence themselves. - -## Technical Details / Code Analysis - -**The address is key material, and only its shape is checked.** `shim/src/config.rs:292-309`: - -```rust -fn is_nym_address(addr: &str) -> bool { - let Some((keys, gateway)) = addr.split_once('@') else { return false; }; - let Some((identity, encryption)) = keys.split_once('.') else { return false; }; - !gateway.is_empty() && !identity.is_empty() && !encryption.is_empty() - && !gateway.contains('@') && !encryption.contains('.') -} -``` - -Its own doc comment says it is "deliberately shallow ... leaves the real parse (base58, -key lengths) to the SDK". Nothing above it establishes *whose* keys these are. -`Config::hub_selection` (`:255-289`) accepts any number of such entries, rejecting only -empty and duplicate ones. - -**Confidentiality is to the key in the address, not to the hub.** The Sphinx payload is -encrypted to the recipient's encryption key (verified in the pinned `nym-sdk` tree by the -G15 pass: `common/nymsphinx/src/preparer/mod.rs:165-180`), so the mixnet leg is genuinely -private — against everyone except the holder of the key named in `ZIS_HUB_NYM`. Inside -that envelope the frame is cleartext; there is no application-layer AEAD, because that -is the unbuilt STEVE (`shim/src/wire.rs:28-34`): - -```text -SubmitV1, exactly FRAME_BYTES: - 0 magic 4 b"ZNS1" - 4 nonce 16 request nonce, from OsRng - 20 tx_len 4 u32 big-endian - 24 tx tx_len bytes - .. padding zeros to FRAME_BYTES -``` - -**The provenance chain, end to end.** `deploy.sh:255-259`: - -```sh - hub_get() { - _resp=$(curl -s --max-time 20 -w ' %{http_code}' "https://$TLS_DOMAIN$1") || true - _code=${_resp##* } - _body=${_resp% *} - } -``` - -Default WebPKI verification of `$TLS_DOMAIN`, nothing else. The loop at `:283-308` checks -only the HTTP status, the *shape* of the body (`is_nym_address`, `:265-280`), and that -`/nym-status` reports `mixnet_connected == true`; on success `:316` prints the string to -stdout for a caller to bake into a shim. No attestation is fetched and `caution verify` -is never invoked on this path. The hub side of the handoff -(`hub/deploy/caution/OPERATORS.md:131-145`) is a bare `curl https:///nym-address` -followed by the human-message sentence; the shim side -(`shim/deploy/caution/OPERATORS.md:237-239`, `:312`) designates that URL authoritative and -the config-table default. - -**Why the attestation does not cover this, even after coordinator item 6q.** 6q -established that `unit.env` is measured into PCR0/PCR1 and that the whole environment is -served at `.manifest.run_command` of `/attestation`, and that `caution verify` genuinely -binds the enclave to its TLS leaf via `user_data.tls.certfp`. Both are true and both are -beside the point here: - -- the shim's attestation proves the shim sends to `R_att` — which is precisely the harm, - not a defence against it; -- `caution verify` against the **hub** *would* detect Variant A, because it compares the - leaf of the connection that served `/attestation` against the COSE-signed `certfp`. - It is never run by a shim operator, appears in no shim runbook, and is explicitly - deferred in the one procedure during which every shim in the fleet adopts a new - identity. - -**What a fix looks like with today's parts.** `caution verify --attestation-url -https:///attestation` first, pin the leaf, then read `/nym-address` over a -connection presenting that same leaf. Because the attested hub binary serves only the -address its own in-enclave Nym client minted (`hub/src/server.rs:62-70`, `:469-479`, -filled by the driver), that sequence *does* anchor the address to the attested enclave. -The machinery exists; nothing in the repository composes the two steps. - -## Recommendations - -In rough order of cost: - -1. **Make the fetch attested.** Change `deploy.sh`'s `hub_get` and both OPERATORS - runbooks so `/nym-address` is only ever read over a connection whose leaf has been - compared against the hub's `/attestation` `user_data.tls.certfp`. Remove "verify - after" from `hub/deploy/caution/OPERATORS.md:244-245`: the whole fleet is adopting a - new key in that window, which is the one moment verification is load-bearing. -2. **Bind the Nym address to the attestation itself.** Have the platform place the hub's - current Nym address in the COSE-signed `user_data` (alongside `tls.certfp`), the same - way `caddy-certfp.sh` gets the leaf fingerprint there, so the address is signed by the - Nitro key rather than merely served over TLS. A shim operator then verifies the - address, not the endpoint that served it. -3. **Give the rotation key continuity.** Have the hub sign its new Nym address with the - previous identity (or with a long-lived offline consortium key) and publish the - signature at `/nym-address`, so a legitimate rotation is distinguishable from a - substitution without re-running the whole attestation flow. -4. **Add a shim-side check on the list.** At minimum, log the configured addresses and - their count at startup (today `shim/src/main.rs`'s mixnet arm logs neither), and - refuse to start with more than one `--hub-nym` entry unless an explicit - `--hub-nym-failover` flag is passed. -5. **Correct the claims.** `README.md:77` and `README.md:71` should state that the hub's - identity is an unauthenticated configured value today, and what an auditor can and - cannot conclude from an attestation that contains it. - -## Validation Information - -**Verdict: CONFIRMED. Severity: Medium — the filed grade is upheld and the case for High -argued inside the issue is decided against, for the reasons below.** - -### Every mechanical and documentary claim re-verified against the target at HEAD - -| Claim | Verified at | -|---|---| -| Only the shape of a `ZIS_HUB_NYM` entry is checked; any number of entries accepted | `shim/src/config.rs:292-309`, `:255-289` | -| A submit fans out to every entry, with no ack and no rotation | `shim/src/nym.rs:642` (`for target in 0..targets`), `:652` (ack receiver dropped at construction) | -| Lookups rotate a cursor and only a **timeout** sweeps to the next address | `shim/src/nym.rs:743-781` | -| A `LookupReply::Error` fails closed immediately with no sweep | `shim/src/hub.rs:264-266` | -| The frame carries the transaction with no application-layer AEAD | `shim/src/wire.rs:28-34` | -| Sphinx encrypts to the key **in the address**, so substituting the string substitutes the recipient | `nym-sdk` `common/nymsphinx/src/preparer/mod.rs:165-180` (established by the G15 pass) | -| `deploy.sh` fetches `/nym-address` with a bare `curl` over WebPKI and never touches the attestation | `deploy.sh:255-259`, `:283-318` | -| "the authoritative copy" | `shim/deploy/caution/OPERATORS.md:239` | -| "There is no discovery mechanism — the handoff is a human message" | `hub/deploy/caution/OPERATORS.md:144` and again at `:238` | -| "`caution verify` … does NOT belong on the critical path — restore service first, verify after" | `hub/deploy/caution/OPERATORS.md:244-245` | -| Poll-and-alert is the documented response to rotation; no key continuity anywhere | `shim/deploy/caution/OPERATORS.md:376`, `hub/deploy/caution/OPERATORS.md:186-189` | -| `certfp` appears **nowhere** in the target | grep: zero occurrences (also recorded under coordinator item 7k) | -| README states the cleartext claim unconditionally and lists four unrelated auditor checks | `README.md:77`, `:71`; the "Not protected" list at `:30-36` does not mention the destination | -| A stranger can force the rotation that triggers the fleet-wide re-handoff | `hub-liveness-probe-…-fresh-identity-fleet-kill.md` (confirmed Medium) | - -### Why this is a real, separate finding and not a re-report of a stated residual - -The audit instructions record "there is no authentication on the shim→hub channel today" -as a self-declared limitation, and `AUDIT-INSTRUCTIONS.md` says a stated residual must not -be re-reported. That rule was applied deliberately, and this issue survives it: - -- The **residual** is that the channel has no cryptographic peer authentication (STEVE, - encrypt-to-hub-key). This issue does not claim otherwise and does not ask for STEVE. -- The **finding** is about the *interim substitute the project chose instead*: a value - distributed out of band whose only documented reference is unauthenticated, while an - authenticated reference (`certfp`, verified by 6q to exist and work) is used two lines - away for a different purpose and never composed with this one. That is a concrete - deploy-script and runbook defect with a concrete fix, not a restatement of "STEVE is - not built". Recommendation 1 costs a few lines of `deploy.sh` and two paragraphs of - runbook. -- The audit instructions' own carve-out applies too: `README.md:77` claims more than the - residual allows, and `README.md:30-36` omits it from "Not protected". - -### Why Medium and not High — the case for High, decided - -The filed text argued for High on impact (total defeat of the core guarantee, fleet-wide, -undetectable) plus a cheap vector (spoofing a base58 string in a chat message). The impact -half is correct. The likelihood half does not carry: - -1. **The cheapest vector is defeated by a check the documentation already prescribes.** - The shim runbook designates `https:///nym-address` "the authoritative - copy", so a spoofed handoff *message* fails against any operator who does the - cross-check the same document tells them to do. What is left as the load-bearing - precondition is control of the hub's DNS/PKI, or being the hub operator. -2. **DNS/registrar/certificate control over the hub's domain is a genuine position**, not - a capability a stranger has. This audit has consistently held findings whose - precondition is a configured-endpoint or infrastructure position at Medium rather than - High — coordinator item 6p's bound, applied to both tip issues — and the same standard - governs here. -3. **The hub-operator route is real but is the cheapening of an already-stated residual.** - The audit's own threat model records W9 ("the hub sees every diverted transaction in - cleartext, and this is total ... compromise of the hub, or a legal order served on - whoever runs it, exposes every migration"). Substitution makes that outcome reachable - without touching the enclave, which is a genuine and unowned observation — it is why - this issue is confirmed rather than merged away — but it does not create a new class - of victim for that party. -4. **Nothing here is reachable by an anonymous party.** Compare the confirmed High it is - most often confused with, whose attacker is adversary #1 (the shim's own operator, a - party whose hostility is the product's founding premise) and needs only one comma in a - config file. - -Medium is therefore the honest grade: maximal impact, real and recurring opportunity, -undetectable by any documented procedure — held below High by a precondition that is a -real infrastructure position, and by the fact that the underlying absence of peer -authentication is disclosed by the project. - -### Anti-double-counting — checked against all four neighbours - -- **`shim-submits-every-migration-to-every-configured-hub-…` (confirmed High)** — a - *different mechanism* (append at the shim, by the shim's own operator) and a different - victim set (that endpoint's users). This issue is *substitute at the source*, by - someone who is not the shim operator, affecting every shim at once. Neither fix - addresses the other: exact whole-list equality against a published value does nothing - if the published value is the attacker's, and anchoring the published value does - nothing about an operator appending an extra entry. Variant B is described here only to - show the reference-value defect also enables the *other* issue's capability from - outside; the append capability itself stays credited there. -- **`auditor-recipe-omits-…` (confirmed Medium)** — owns "no document tells anyone to - look". This issue owns the orthogonal half: *even when someone looks, the comparison - terminates in an unanchored value*. That constraint is already recorded in the recipe - issue's validation as a constraint on its Recommendation 1; this file is where the - constraint's own fix (Recommendations 1-3 above) lives. The recipe issue's severity is - not moved by this one. -- **`operators-runbook-attributes-the-hub-destination-to-…` (confirmed Low)** — owns two - operator-facing sentences that falsely assure the reader the destination is bound by - the binary hash and egress rules. Documentation-only, no independent attack path. This - issue does not restate those lines. -- **`attested-tls-binding-is-verified-once-by-hand-if-ever-…` (confirmed Medium)** — owns - the certfp binding being *time-of-check-only*. This issue is the case where the binding - is **never applied at all** to a different value. Different artefact, different fix. -- **`hub-nym-driver-automatic-fresh-identity-permanently-invalidates-every-shim.md` - (confirmed)** — owns the *availability* consequence of the same mandatory rotation; - this is its *authenticity* consequence. The forced-rotation primitive - (`hub-liveness-probe-…-fleet-kill.md`, confirmed Medium) is cited here as the thing that - manufactures the handoff window, not re-counted. - -### Corrections applied against the filed text - -- **The "cheapest vector" argument was struck.** The filed text leaned on spoofing the - human message; the shim runbook's own "authoritative copy" instruction defeats that in - isolation, and saying otherwise would overstate the finding. The load-bearing - preconditions are now stated as DNS/PKI control or the hub operator. -- **"`SubmitV1` is plaintext" was made precise.** As written it could be read as - contradicting a positive the report must state — that the mixnet leg is end-to-end - encrypted and no mix node or gateway can read a submission (G15). The frame is - cleartext *inside* the Sphinx envelope, i.e. readable only by the holder of the key - named in `ZIS_HUB_NYM`, which is exactly why substituting that string is the whole - attack. -- **Variant ordering was inverted.** The filed text led with the appended-address variant, - which is mechanically the confirmed High's capability and carries a ~90 s lookup-latency - tell on half of all lookups. Substitution is the variant this issue actually owns and is - now primary. -- **The claim that a silent rogue "need not send a single packet to stay invisible" was - softened** to name the latency cost, verified at `shim/src/nym.rs:743-781` and - `shim/src/nym.rs:48-71` (`REQUEST_TIMEOUT = 90 s`, overridable via - `ZIS_LOOKUP_TIMEOUT_SECS` at `shim/src/config.rs:115-116`, and documented at - `shim/deploy/caution/OPERATORS.md:316` as *multiplying* by the number of addresses). -- **The fix was re-grounded in parts that exist.** The filed Recommendation 1 asked the - platform to add the address to `user_data`; validation established that a - certfp-pinned fetch of `/nym-address` from the attested hub achieves the same binding - with today's artefacts, so that is now Recommendation 1 and the platform change is - Recommendation 2. - -DO NOT DEVIATE FROM THIS TEMPLATE (`docs/SECURITY-ISSUE-TEMPLATE.md`). -LEAVE THIS MESSAGE IN PLACE SO THAT ALL AGENTS KNOW NOT TO DEVIATE FROM THIS TEMPLATE (`docs/SECURITY-ISSUE-TEMPLATE.md`). diff --git a/zeronym-22aa9851caf68-high-medium/medium/hub-nym-lookup-flood-starves-gettransaction-fleet-wide.md b/zeronym-22aa9851caf68-high-medium/medium/hub-nym-lookup-flood-starves-gettransaction-fleet-wide.md deleted file mode 100644 index a3c5428c..00000000 --- a/zeronym-22aa9851caf68-high-medium/medium/hub-nym-lookup-flood-starves-gettransaction-fleet-wide.md +++ /dev/null @@ -1,555 +0,0 @@ -# An anonymous lookup flood consumes the hub's single fleet-wide mixnet emitter, so `GetTransaction` stops working for every wallet on every shim - -**Severity**: Medium -**Validation Status**: Confirmed -**Location**: -`audit-target/zeronym/hub/src/nym.rs:38-54` (`MAX_CONCURRENT_LOOKUPS`), `:56-75` (`REPLY_DEADLINE`), `:148-216` (`run_listener`, the lookup arm), `:218-234` (`is_lookup`), `:249-303` (`build_lookup_reply`); -`audit-target/zeronym/hub/src/nym_driver.rs:384-403` (the one-reply-in-flight guard), `:452-479` (the reply arm), `:632-643` (`reply_send` → `send_reply`); -`audit-target/zeronym/hub/src/wire.rs:476-501` (`encode_lookup_reply` pads every disposition to 64 KiB), `:439-470` (`decode_lookup`, `peek_lookup_nonce`); -`audit-target/zeronym/hub/src/main.rs:184-185` (both mixnet channels sized 64); -`audit-target/zeronym/hub/src/server.rs:62-70`, `:446-479` (`GET /nym-address` publishes the target to anyone); -`audit-target/zeronym/hub/deploy/caution/caution.hcl.tmpl:50-55` (`ingress 0.0.0.0/0`); -`audit-target/zeronym/shim/src/intercept.rs:229-247` (every `GetTransaction` goes to the hub), `:315-324` (fail closed with `UNAVAILABLE`); -`audit-target/zeronym/shim/src/nym.rs:45-71` (`REQUEST_TIMEOUT` = 90 s, and that it multiplies over the hub-address list), `:98-104` (`LOOKUP_REPLY_SURBS` = 60 and the measured 41-packet reply), `:1085-1130` (`throughput_budget`, the crate's own emission model). -The emission FIFO that actually starves honest lookups is in the pinned SDK (`nym-sdk` at git rev `451c2aa3692fc4dc00041b74a352d4158176d9c0`, `hub/Cargo.lock:4773-4775`): `common/client-core/src/client/real_messages_control/real_traffic_stream.rs:427-490` (`poll_poisson`), `:156-157` and `real_messages_control/mod.rs:150` (the 8-slot batch channel), `client/transmission_buffer.rs:39-49`, `:190-250` (`pop_next_message_at_random`, FIFO within a lane); `client/replies/reply_controller/requests.rs:12-15` (the unbounded reply-controller channel); `client/replies/reply_controller/receiver_controller.rs:218-270` (`handle_send_reply`'s SURB arithmetic); `common/client-core/config-types/src/lib.rs:25`, `:48-62`, `:443` (all shipped defaults). - -**Found by agent:** Local (`hub/src/nym.rs`); validated 2026-08-18 -**In scope of audit?** Yes — priority area 4 ("the mixnet transport") and `AUDIT-INSTRUCTIONS.md` unverified lead #7. - -## Description - -The hub answers every shim's `GetTransaction` over the Nym mixnet, and every -answer is a `LookupReplyV1` padded to a fixed 64 KiB (`wire.rs:476-501`) so its -size cannot reveal found-versus-not-found. The project's own nymnet measurement -puts that frame at **41 sphinx packets** (`shim/src/nym.rs:98-104`), and the -project's own emission model puts one mixnet client's floor rate at -**8.33 packets/s** — `MAX_DELAY_MULTIPLIER` (6) against the SDK's 20 ms default -`message_sending_average_delay` (`shim/src/nym.rs:1090-1094`, confirmed against -`client-core/config-types/src/lib.rs:25`, `:443`). - -The hub has exactly **one** mixnet client, pinned to one gateway and one -identity by design (D10, address stability). So the hub's whole reply capacity — -for every shim, every operator, every wallet — is: - -| regime | packets/s | seconds per 64 KiB reply | replies per minute, fleet-wide | -|---|---|---|---| -| the project's throttled floor (`THROTTLED_PACKETS_PER_SEC`) | 8.33 | ~4.9 | **~12** | -| the SDK's unthrottled default (20 ms) | 50 | ~0.8 | ~73 | - -A stranger can consume all of it. The hub's Nym address is published on purpose -and answers everyone (`GET /nym-address`, `server.rs:62-70`; the enclave declares -`ingress 0.0.0.0/0`), a lookup is a 64-byte frame that only has to start with -`b"ZNL1"` (`is_lookup`, `nym.rs:232-234`), and **every** failure arm of -`build_lookup_reply` — an undecodable frame, an empty hash, an unavailable -indexer — still returns a full 64 KiB frame (`nym.rs:254-302`). There is no -authentication, no ACL, no rate limit, and no submitter identity to key one on: -`nym.rs:12-15` and `queue.rs:35-39` forbid holding one. - -Because the shim routes **every** `GetTransaction` to the hub — it is stateless -and cannot recognise its own migrations (`intercept.rs:229-247`) — and fails -**closed** with `UNAVAILABLE` when the hub does not answer in 90 s -(`intercept.rs:315-324`, `nym.rs:45-71`), starving this one emitter breaks -transaction lookup for every wallet behind every shim pointing at that hub, not -only for migrating ones. - -### The mechanism is not the one the hub's own comments describe - -This was the central correction of validation, and it matters because it changes -which mitigation is load-bearing and which is inert. - -The hub's design intends `outgoing` (64 deep) plus the driver's one-reply-in-flight -guard to be the queue, and `REPLY_DEADLINE` (60 s) to drop replies that have -outlived the shim's budget *before* they cost an emission -(`nym.rs:56-75`, `nym_driver.rs:384-403`, `:452-471`). **On the deployed SDK none -of that binds**, because handing a reply over is non-blocking: - -`nym_driver.rs:638` calls `sender.send_reply(tag, frame)`, which is -`ClientInput::send` into an `mpsc::channel::(1)` -(`base_client/mod.rs:1013`); `InputMessageListener::on_input_message` does -nothing at all for a `Reply` variant except -`reply_controller_sender.send_reply(...)` -(`acknowledgement_control/input_message_listener.rs:60-70`), and that channel is -`mpsc::unbounded` (`reply_controller/requests.rs:12-15`). So `send_reply` -returns in microseconds, `in_flight` is never held, `outgoing` drains at memory -speed, and a reply is essentially never in `outgoing` long enough for -`is_dead()` to fire. - -The FIFO that really forms is one layer down, inside the SDK, and it is -strictly worse than the one the hub designed: - -- it is `OutQueueControl::transmission_buffer`, a bare - `HashMap>` with **no size limit and no byte - budget** (`transmission_buffer.rs:39-49`); -- it is drained one packet per Poisson tick and *filled* one whole message per - tick (`real_traffic_stream.rs:443-478` polls `real_receiver` once per tick and - stores the entire batch), so under overload it absorbs ~41× faster than it - emits; -- **all** replies travel on `TransmissionLane::General` - (`nym-sdk/src/mixnet/traits.rs:122-133`), and within a lane - `pop_next_message_at_random` is `pop_front` on a `VecDeque` - (`transmission_buffer.rs:190-208`) — strict FIFO; -- nothing in it ages anything out. `prune_stale_connections` only evicts a lane - that has been idle for ten minutes, and under a flood the General lane is never - idle. - -So an honest shim's reply is enqueued behind the attacker's entire backlog and -emitted whenever that backlog clears — which under a sustained flood is never -inside the shim's 90 s budget. Two consequences follow that the filed text did -not have: the hub's only anti-starvation mitigation (`REPLY_DEADLINE`) is -**inoperative**, and the outage **outlives the attacker** by however long the -accumulated backlog takes to drain. - -`MAX_CONCURRENT_LOOKUPS = 64` is also not the constraint here. It bounds memory -and concurrent indexer dials, which it does correctly, but the cheapest attack -frame (`hash_len = 0`) does no I/O at all (`nym.rs:269-272`), so its slot is held -for microseconds and the semaphore is never near full. - -## Attack Scenario and Steps - -Attacker: anyone with an internet connection. No credential, no enclave -compromise, no privileged network position, no Zcash knowledge, no funds. - -1. `curl https:///nym-address`. The endpoint exists to publish the value and - answers everyone (`server.rs:62-70`, `:462-479`). -2. Start two or three stock `nym-sdk` clients. They are free and need no - registration: neither enclave sets `enabled_credentials_mode`, and the SDK - default is `false` (`nym-sdk/src/mixnet/client.rs:281-288`), so an attacker's - client runs on the same terms the target's own does. -3. From each, send a stream of 64-byte messages — `b"ZNL1"` ‖ 16 random bytes ‖ - `0x00` ‖ 43 zero bytes — as **anonymous** messages carrying **≥51 reply - SURBs**. Fifty-one is the threshold: `handle_send_reply` sends - `min(fragments, available_surbs − min_surb_threshold)` fragments with - `min_surb_threshold = 10` (`receiver_controller.rs:248-256`, - `config-types/src/lib.rs:48`), so 41 fragments need 51 SURBs to go out in one - pass. The honest shim already attaches 60 (`shim/src/nym.rs:98-104`). -4. Each frame reaches `is_lookup` on size-plus-magic, takes the empty-key arm - (`nym.rs:269-272`) with **no indexer dial and no queue scan**, and produces a - full 64 KiB `error` reply that the hub hands to the SDK immediately. -5. The SDK emits those replies FIFO on lane General at 8-50 packets/s. Every - honest shim reply enqueued afterwards waits behind them. -6. Each starved wallet lookup costs the wallet the shim's `REQUEST_TIMEOUT` - (90 s), multiplied by the number of configured hub addresses - (`each_target`, `shim/src/nym.rs:695-790`), and then returns `UNAVAILABLE`. - -**Attack Requirements and Assumptions:** - -- **The cost is symmetric per packet, and that is the point.** To have a reply - emitted at all the attacker must attach ~51 reply SURBs, which the project - measures at roughly one sphinx packet each (60 SURBs ≈ 60 packets, - `shim/src/nym.rs:98-104`, `:1120-1130`). So ~51 attacker packets buy ~41 - packets of hub egress. **There is no bandwidth amplification.** The asymmetry - is structural: the hub is one throttled client serving the entire fleet, and - the attacker may run as many clients as they like. At equal rates ~1.3 - attacker clients match the hub's whole capacity; three swamp it; ten leave an - honest lookup ~8 % of the emitter. -- Nothing stops them. There is no ACL, no rate limit and no per-submitter - accounting anywhere on this path, and the anonymity requirement - (`nym.rs:12-15`, `queue.rs:35-39`) is why — see the Recommendations for the - identity-free controls that are nevertheless available. -- The attacker is unattributable: there is no source address, and the sender tag - is never interpreted by design. -- Nothing alerts. `/healthz` is unconditionally 200 (`server.rs:449-452`), - `/nym-status` reports `mixnet_connected: true, client_deaths: 0`, and the only - in-process signal is a `warn!` that under `debug { enabled = false }` reaches - no console. The SDK's own `log_status` does warn once the transmission buffer - passes 1,000 packets (`real_traffic_stream.rs:570-578`) — into the same absent - console. -- **What bounds it:** the flood is noisy on the mixnet, it must be sustained, - and it costs the attacker real (if free) mixnet egress. It destroys nothing and - discloses nothing. - -## Impact on Users - -- **Every wallet behind every shim pointing at this hub loses `GetTransaction`, - fleet-wide, for as long as the attack runs plus the drain time afterwards.** - Not only migrations: the shim routes *all* lookups to the hub. Each attempt - hangs 90 s per configured hub address and then fails `UNAVAILABLE`. -- **This is more than cosmetic for a light wallet.** `GetTransaction` is the - call a wallet makes to fetch a full transaction after finding it in a compact - block ("enhancement"), so a fleet-wide outage presents to users as sync - failures, not as one missing detail screen. -- **The realistic user response is the privacy loss.** A wallet that cannot - fetch its transactions gets pointed at a different, unprotected indexer — which - is the deanonymisation this product exists to prevent, and the attacker chooses - the moment. -- **Migrations themselves keep working during this attack**, which bounds the - severity and is worth stating plainly: admission is inline and never waits on - the lookup bound (`nym.rs:120-144`), the shim's submit is dispatch-only and - never awaits the ack, and the batcher publishes over HTTP, not the mixnet. - Nothing is lost or corrupted by this issue on its own. -- **Sustained, it degenerates into the confirmed OOM.** Because the SDK absorbs - ~41× faster than it emits, a sustained flood also grows the hub's transmission - buffer without bound — the same terminal outcome as - `hub-surb-starved-lookup-replies-grow-the-sdk-pending-buffer-without-bound-and-oom-the-enclave.md`, - reached ~40× more expensively. That outcome is graded there, not here. - -## Technical Details / Code Analysis - -**1. Any 64-byte frame with the lookup magic buys a full 64 KiB frame, with no -I/O.** - -`hub/src/nym.rs:232-234`: - -```rust -fn is_lookup(frame: &[u8]) -> bool { - frame.len() == wire::LOOKUP_BYTES && wire::peek_lookup_nonce(frame).is_some() -} -``` - -`peek_lookup_nonce` (`hub/src/wire.rs:459-470`) checks only length ≥ 21 and the -4-byte magic, so the added condition is `len == 64`. `decode_lookup` -(`wire.rs:438-457`) accepts `hash_len = 0` without error, so a well-formed frame -reaches: - -```rust -// hub/src/nym.rs:269-272 - if hash.is_empty() { - tracing::warn!(reason = "empty lookup key", "lookup refused"); - return Some(error_reply(nonce)); - } -``` - -and `error_reply` is `encode_lookup_reply(&nonce, &LookupReply::Error)`, which -allocates and pads `FRAME_BYTES` = 65,536 every time (`wire.rs:480-499`). - -The doc comment above `is_lookup` (`nym.rs:218-231`) presents the size check as -having closed an amplifier — *"a 21-byte message got a 65 536-byte answer, 41 -sphinx packets of the hub's own metered egress"*. On the mixnet it closes -nothing: the smallest message a sender can put on the mixnet is one sphinx -packet, so a 21-byte payload and a 64-byte payload cost the sender exactly the -same, and the answer is the same 65,536 bytes. The real cost driver — the -attached reply SURBs — is untouched by a length check. `hub/tests/nym.rs`'s -`a_runt_lookup_shaped_message_buys_no_reply_at_all` pins a property that does not -hold on the transport that ships. - -**2. The hand-off is non-blocking, so the hub's own queue and deadline never -engage.** - -```rust -// hub/src/nym_driver.rs:632-643 -fn reply_send(sender: ..., tag: AnonymousSenderTag, frame: Vec) -> InFlight { - Box::pin(async move { - if let Err(err) = sender.send_reply(tag, frame).await { ... } - Sent::Reply - }) -} -``` - -`send_reply` → `ClientInput::send` → `mpsc::channel::(1)` -(`base_client/mod.rs:1013`) → `InputMessageListener::handle_reply`, whose entire -body is: - -```rust -// client-core/.../input_message_listener.rs:60-70 - async fn handle_reply(&mut self, recipient_tag, data, lane, max_retransmissions) { - // offload reply handling to the dedicated task - let _ = self.reply_controller_sender - .send_reply(recipient_tag, data, lane, max_retransmissions); - } -``` - -and `ReplyControllerSender` wraps `futures::channel::mpsc::unbounded` -(`reply_controller/requests.rs:12-15`, `:38-45`). Nothing between -`nym_driver.rs:474` and an unbounded queue applies backpressure, so the guard at -`nym_driver.rs:455` (`if in_flight.is_none()`) and the drop at `:465-471` -(`reply.is_dead()`) are both inert in practice. - -**3. The FIFO that does form is unbounded, un-aged and shared.** - -```rust -// client-core/.../real_traffic_stream.rs:443-478 (poll_poisson, one tick) - match Pin::new(&mut self.real_receiver).poll_recv(cx) { - Poll::Ready(Some((real_messages, conn_id))) => { - self.transmission_buffer.store(&conn_id, real_messages); - let real_next = self.pop_next_message().expect("Just stored one"); - Poll::Ready(Some(StreamMessage::Real(Box::new(real_next)))) - } - Poll::Pending => { /* pop one, else Cover */ } - } -``` - -One tick stores a whole message (41 fragments) and emits **one packet**. -`pop_next_message_at_random` picks a lane and then `pop_front`s -(`transmission_buffer.rs:190-208`); all replies are -`TransmissionLane::General` (`nym-sdk/src/mixnet/traits.rs:122-133`), so honest -and attacker replies are one FIFO. `TransmissionBuffer` has no cap -(`transmission_buffer.rs:39-49`). - -**4. The rate, from the project's own constants.** - -```rust -// shim/src/nym.rs:1090-1094 - const PACKET_BYTES: usize = 2 * 1024; - /// The client's own floor on sending, `MAX_DELAY_MULTIPLIER` (6) times the - /// 20 ms default `message_sending_average_delay`. - const THROTTLED_PACKETS_PER_SEC: f64 = 1000.0 / 120.0; -``` - -41 packets ÷ 8.33 packets/s ≈ **4.9 s per reply ⇒ ~12 replies/minute fleet-wide**; -at the unthrottled 20 ms default (`config-types/src/lib.rs:25`, `:443`) it is -~0.8 s ⇒ ~73/minute. `hub/src/nym.rs:62-70` and `nym_driver.rs:456-464` state the -same arithmetic in prose. Both figures are ceilings for the whole fleet, not per -shim. - -**5. What the attacker must pay, exactly.** - -```rust -// client-core/.../receiver_controller.rs:248-256 - let available_surbs = self.surbs_storage.available_surbs(&recipient_tag); - let min_surbs_threshold = self.surbs_storage.min_surb_threshold(); // 10 - let max_to_send = if available_surbs > min_surbs_threshold { - min(fragments.len(), available_surbs - min_surbs_threshold) - } else { 0 }; -``` - -so ≥51 attached SURBs to get all 41 fragments emitted. Fewer than 11 is a -different attack entirely (the reply is buffered forever — that is the confirmed -OOM issue). Zero is safe: `contains_surbs_for` fails and the reply is dropped -(`receiver_controller.rs:225-241`). - -**6. The shim's fail-closed arm turns this into a wallet-visible outage.** - -```rust -// shim/src/intercept.rs:315-324 - Err(err) => { - tracing::warn!(target: "zis::classify", %err, "hub lookup failed; failing closed"); - Ok(grpc_error(GRPC_UNAVAILABLE, "zero-indexer-shim: hub unreachable")) - } -``` - -with `REQUEST_TIMEOUT = 90 s` (`shim/src/nym.rs:71`), and `each_target` -multiplying it by the number of configured hub addresses -(`shim/src/nym.rs:695-790`). - -## Recommendations - -In order of value. Items 1 and 2 are the same two controls G5 §4.1/§4.2 -identified, and both are compatible with the strictest reading of -`hub/src/queue.rs:35-39`, which forbids an identifier *on a queue entry*; a -lookup never becomes a queue entry. - -1. **Answer a lookup only when the request carried enough reply SURBs to carry a - full frame.** The conforming shim already attaches 60 precisely so the hub - never has to re-request (`shim/src/nym.rs:98-104`), so no honest client is - affected, and every under-provisioned request is refused before 64 KiB is - allocated. This raises this attack's cost floor to what an honest client - already pays and simultaneously removes the confirmed OOM. -2. **Bound in-flight replies per sender tag.** `Received` already carries the - tag (`nym.rs:83-93`) and the hub already holds it for the life of the request - in order to reply at all. A token bucket keyed on the tag — a counter and a - timestamp, referencing no queue entry, forgotten after a minute — caps any one - tag at `K` replies and costs a Sybil attacker one gateway registration per - bucket. -3. **Do not spend a full frame on requests a conforming shim never sends.** The - `decode_lookup` failure arm and the `hash.is_empty()` arm - (`nym.rs:258-272`) each cost 41 packets today. The 64 KiB padding exists to - hide `Found` from `NotFound`, an axis these two are not on. Drop them - silently, exactly as the submit arm already drops a frame with no recoverable - nonce (`nym.rs:327-333`). -4. **Fair-share the emitter across sender tags.** The tag is already on `Reply` - (`nym.rs:98-105`). Round-robin across distinct tags in `outgoing` gives an - honest shim `1/(T+1)` of the emitter against `T` attacker tags instead of a - vanishing share, turning a total outage into graceful degradation. This is - again a counter, not an identity. -5. **Make `REPLY_DEADLINE` actually bind, or delete it and say why.** As shipped - it can almost never fire, because `send_reply` returns before anything is - emitted. If the hub is to keep this mitigation it must stop handing the SDK - more than the SDK can emit — e.g. gate `outgoing.recv()` on the SDK's own - lane-queue length (`ClientState::lane_queue_lengths` is exposed by the SDK) - rather than on a boolean `in_flight`. Leaving an inert mitigation in place with - a 20-line comment explaining the starvation it prevents is worse than having - none, because it stops the next reader looking further. -6. **Give the reply path more than one emitter.** One pinned client is deliberate - for address stability (D10), but nothing stops the hub from holding one stable - *inbound* identity and a small pool of additional clients used only for - outbound replies, multiplying reply capacity by the pool size. -7. **Surface the condition.** Export the SDK's pending lane length, or at minimum - a delayed aggregate count of lookups answered and lookups dropped, on - `/nym-status`, so an operator can tell "the hub is being flooded" from "the - mixnet is slow". Today both look identical and both look healthy. -8. **Correct the `is_lookup` doc comment and - `a_runt_lookup_shaped_message_buys_no_reply_at_all`** so neither is read as - having closed an amplification gap; on the mixnet the size check does not - change the sender's cost. -9. Revisit `REQUEST_TIMEOUT` × `each_target` on the shim side, so a saturated hub - does not cost a wallet 90 s per configured address per lookup. - -Related, do not duplicate: - -- `hub-surb-starved-lookup-replies-grow-the-sdk-pending-buffer-without-bound-and-oom-the-enclave.md` - (High) is the **under**-provisioned-SURB case: one SURB, the reply is buffered - forever, the effect is cumulative and permanent. This issue is the - **over**-provisioned case: ≥51 SURBs, the reply is emitted, the effect is - transient starvation of everyone else. Different cost, different consequence, - different fix — though recommendation 1 above closes both. -- `hub-unauthenticated-pre-publication-transaction-disclosure.md` covers the - *disclosure* properties of the same lookup path. -- `hub-http-lookup-path-has-no-concurrency-bound.md` is the complementary finding - on the clearnet leg (the bound is missing there). This issue is the opposite - observation on the mixnet leg: the bound is present and correct for what it was - written for, and still does not prevent starvation, because the scarce resource - is the single emitter. -- `gettransaction-flood-starves-migration-diversion.md` (High) is the shim-side - analogue and the cheaper route to a subset of this effect (G5 §3.1): a plain - HTTP request at any shim becomes a mixnet lookup at the shim's expense which - then lands on this hub, so an attacker with no mixnet capability at all reaches - both enclaves with one request. - -## Validation Information - -**Verdict: CONFIRMED. Severity corrected from High to Medium.** - -Everything decisive was checked against the code and against the **pinned SDK -tree at `451c2aa3692fc4dc00041b74a352d4158176d9c0`**, which is available locally, -rather than against the filing. - -### What was verified - -1. **Reachability.** `GET /nym-address` is unauthenticated and answers everyone - (`server.rs:446-479`); the enclave declares `ingress 0.0.0.0/0` - (`caution.hcl.tmpl:50-55`). `is_lookup` accepts on size-plus-magic - (`nym.rs:232-234`), `decode_lookup` accepts `hash_len = 0` (`wire.rs:438-457`), - and the empty-key arm returns a full `FRAME_BYTES` frame with **no indexer - dial and no queue scan** (`nym.rs:269-272`, `wire.rs:480-499`). Confirmed. -2. **Attacker clients are free.** Neither binary enables - `enabled_credentials_mode`, and the SDK default is `false` - (`nym-sdk/src/mixnet/client.rs:281-288`), so an attacker's stock client runs on - exactly the terms the target's own does. Confirmed. -3. **The packet arithmetic, recomputed independently.** 41 packets per 64 KiB - reply is the project's measured figure (`shim/src/nym.rs:98-104`); 8.33 - packets/s is the project's own `THROTTLED_PACKETS_PER_SEC` - (`shim/src/nym.rs:1090-1094`), correctly derived as - `MAX_DELAY_MULTIPLIER (6) × DEFAULT_MESSAGE_STREAM_AVERAGE_DELAY (20 ms)`, - which matches `config-types/src/lib.rs:25`, `:443`. So **~12 replies/min at the - floor and ~73/min unthrottled, fleet-wide**, from a single client the hub pins - by design. The filed "≈12 answered lookups per minute" is right at the floor; - the range is now stated so the finding does not rest on the worst case alone. - Confirmed. -4. **The attacker must pay ≥51 SURBs**, i.e. roughly 51 packets of their own - egress for 41 of the hub's (`receiver_controller.rs:248-256` with - `min_surb_threshold = 10`). The filing already said this and it is correct: - **there is no bandwidth amplification here.** The finding rests entirely on - the structural asymmetry of one fleet-wide emitter against N attacker clients, - and that asymmetry is real. Confirmed. -5. **The starvation is total, not proportional.** All replies use - `TransmissionLane::General` (`nym-sdk/src/mixnet/traits.rs:122-133`) and a lane - is a FIFO `VecDeque` (`transmission_buffer.rs:190-208`), so an honest reply is - emitted only after the attacker's entire backlog. Confirmed. - -### Corrections made to the filing - -- **The stated mechanism was wrong and has been replaced.** The filing modelled a - 128-deep pipeline (`outgoing` 64 + 64 permits) that "accepts ~2.1 lookups/s and - emits ~0.2/s, so ~90 % of everything accepted is discarded as dead". That - equilibrium does not exist: `send_reply` returns as soon as an **unbounded** - channel accepts the reply (`nym_driver.rs:638` → capacity-1 `InputMessage` - channel → `input_message_listener.rs:60-70` → `reply_controller/requests.rs:12-15`), - so `outgoing` drains at memory speed, the one-in-flight guard is never held, and - `REPLY_DEADLINE` essentially never fires. The real queue is the SDK's unbounded, - un-aged, FIFO `transmission_buffer`. **This is a strengthening correction**: the - hub's only designed defence against exactly this failure is inoperative, and the - backlog survives the attacker instead of being discarded at 60 s. -- **`MAX_CONCURRENT_LOOKUPS` is not the contested resource and the "first-come - drop hits honest lookups" claim has been removed.** The cheapest attack frame - holds a slot for microseconds, so the semaphore is nowhere near full; it is the - emitter that is exhausted. (The filing already said this in its headline; the - supporting paragraph contradicted it.) -- **Recommendation 5 is new** and follows directly from the mechanism correction: - an inert `REPLY_DEADLINE` with a twenty-line comment describing the starvation - it prevents is actively misleading to the next reader. -- Recommendation 1 was promoted to first place (it also closes the confirmed OOM), - and the per-tag bucket and fair-share round-robin were added from G5 §4.2/§4.6, - since the filing's original recommendation set led with measures that either - trade against the padding property or need new infrastructure. - -### `docs/AVOIDING-FALSE-POSITIVES.md` §5 applied - -*What resources would the attacker need?* Two to three free, unregistered -`nym-sdk` clients emitting a few tens of KB/s. *What would stop them?* Nothing in -the target and nothing in the deployment: no ACL, no rate limit, no per-submitter -accounting (currently read as forbidden by `queue.rs:35-39`), `ingress 0.0.0.0/0`, -and no alerting that distinguishes a flood from a slow mixnet. - -§5's caution applies only in part, and the issue is graded accordingly. There is -**no amplification** — the attacker pays ~51 packets for ~41 of the hub's — so -this is not the guide's "1 KB request causing 1 GB allocation" shape. What makes -it a real finding rather than "the attacker must out-resource the target" is that -the target's capacity is *structurally* one throttled client for the entire fleet, -so "out-resourcing it" costs about 1.3 free clients. The correct grade for that is -neither dismissal nor High. - -### Severity: Medium, downgraded from the filed High - -*Why not High.* This issue destroys nothing, discloses nothing, and corrupts -nothing. Migrations continue to be diverted, admitted, batched and published -throughout: admission is inline and unbounded (`nym.rs:120-144`), the shim's -submit is dispatch-only and never awaits the ack, and the batcher publishes over -HTTP. The failure is loud at the wallet (`UNAVAILABLE` after 90 s) rather than a -false success. It requires the attacker to keep paying continuously, and it heals -once the backlog drains. That places it below the three confirmed High findings on -this same surface — `hub-queue-unauthenticated-fill-silently-destroys-migrations`, -`hub-surb-starved-…-oom-the-enclave` and `junk-sendtransaction-flood-…` — each of -which destroys migrations a wallet was told had succeeded. It is graded level with -`hub-unauthenticated-pre-publication-transaction-disclosure` (Medium), which is -the calibration point for "serious, cheap, unauthenticated, but not -loss-of-funds-in-flight". - -*Why not Low.* The blast radius is every wallet on every shim in the fleet, the -attacker is anonymous and free, `GetTransaction` failure presents to a light -wallet as sync failure rather than a missing detail view, and the realistic user -response — switching to an unprotected indexer — is precisely the deanonymisation -the product exists to prevent. Nothing in the current design can rate-limit it, -and nothing reports it. - -*Relationship to `gettransaction-flood-starves-migration-diversion.md` (High).* -The two were graded against each other, not in isolation. That issue is **worse -despite the smaller blast radius**: it needs no mixnet capability at all -(~100-byte HTTP requests), it lands on the *submit* path where -`shim/src/hub.rs:231-240` answers the wallet `error_code 0` at hand-off — so -denial there becomes silent destruction of an acknowledged migration — and it -additionally reaches this hub through the shim (G5 §3.1). This issue is broader -but confined to the lookup path, fails loudly, loses nothing, and heals. G5's cost -ranking (S3 above H9/H10) is therefore upheld, and the severities follow it. - ---- - -### ADDENDUM (Global Auditor, focus area G21 dedicated re-run, 2026-08-18) — a scope qualification, no leg withdrawn, no verdict or severity changed - -This issue's mechanism was re-derived independently from the same pinned SDK tree -and is confirmed in every particular. Three notes. - -1. **The stated severity bound holds only for the SUSTAINED attack shape.** The - Impact section says *"migrations themselves keep working during this attack, - which bounds the severity and is worth stating plainly."* That is correct for a - sustained flood — and is correct for the reason given, that admission is inline - and never waits on the lookup bound. It does **not** hold for a **pulsed** - variant (burst, then be silent for ~130 s), which is a different attack with a - permanent outcome: the hub's liveness probe cannot leave while the backlog this - issue describes is draining, so the hub concludes its gateway has stopped - delivering to it, tears its client down and charges a `short_life`; five such - pulses mint a fresh Nym identity and permanently strand the fleet. Filed - separately as - `hub-liveness-probe-reads-its-own-send-backlog-as-gateway-silence-so-any-stranger-can-drive-the-fresh-identity-fleet-kill.md`, - because the code defect is a different one (a missing conjunct in - `hub/src/nym_driver.rs:421` that `shim/src/nym_driver.rs:312` has) and the fix - is different. Please keep the sentence, and qualify it with "during a sustained - flood". - -2. **A third mitigation belongs on the inoperative list.** The issue already shows - `REPLY_DEADLINE` and the `in_flight` guard do not bind. `MAX_CONCURRENT_LOOKUPS` - does not bind *emission* either, and for the same unit mismatch: the permit is - released once `outgoing` accepts the reply (`hub/src/nym.rs:184-197`), which - happens within seconds, so the semaphore recycles at ~8.33/s and never limits - the total emission an attacker has committed the hub to. It bounds concurrent - indexer dials and resident reply frames correctly, which is what it was written - for; it is worth saying explicitly that it is the third bound denominated in the - wrong unit. - -3. **Recommendation 1 must not ship alone.** Requiring a lookup to arrive with - enough reply SURBs to carry its own padded reply closes the SURB-starved OOM by - making well-provisioned lookups the only answerable ones — and a - well-provisioned lookup is exactly the input the pulsed variant in note 1 needs. - Recommendations 1, 2 and 5 are one change, not three options. - -DO NOT DEVIATE FROM THIS TEMPLATE (`docs/SECURITY-ISSUE-TEMPLATE.md`). -LEAVE THIS MESSAGE IN PLACE SO THAT ALL AGENTS KNOW NOT TO DEVIATE FROM THIS TEMPLATE (`docs/SECURITY-ISSUE-TEMPLATE.md`). diff --git a/zeronym-22aa9851caf68-high-medium/medium/hub-reorg-branch-resets-last-advance-masks-stale-tip.md b/zeronym-22aa9851caf68-high-medium/medium/hub-reorg-branch-resets-last-advance-masks-stale-tip.md deleted file mode 100644 index 7d1720d7..00000000 --- a/zeronym-22aa9851caf68-high-medium/medium/hub-reorg-branch-resets-last-advance-masks-stale-tip.md +++ /dev/null @@ -1,362 +0,0 @@ -# Following a reorg refreshes `last_advance`, so an oscillating tip masks staleness and freezes the flush cadence forever - -**Severity**: Medium -**Validation Status**: Confirmed -**Location**: `audit-target/zeronym/hub/src/batcher.rs:177-196` (the reorg arm of `TipTracker::observe`), interacting with `:204-208` (`is_stale`), `:217-230` (`cadence_height`) and `:299-312` (the epoch flush trigger); admission gate at `audit-target/zeronym/hub/src/server.rs:248-255`. -**Found by agent:** Local (file audit of `hub/src/batcher.rs`); validated 2026-08-18 -**In scope of audit?** Yes - -## Description - -`TipTracker` decides the tip is stale from one timestamp, `last_advance`, which its -own field comment and `REVIEW.md` #8 define as *the last time a node advanced*: - -> Declare the tip stale only when **no node has advanced** for 15 minutes. -> (`hub/REVIEW.md` #8) - -The implementation refreshes `last_advance` not only on an advance but also when it -*follows a regression* inside `REORG_ALLOWANCE` (`batcher.rs:184-186`). A backwards -move is not an advance. Consequently a tip source whose reported height oscillates -inside the reorg allowance — `h, h-1, h, h-1, …` — refreshes `last_advance` on -**every single observation**, in both directions, while the effective height never -moves. - -The result is that both halves of the stale-tip response specified in `REVIEW.md` #8 -are disabled at once: - -* `is_stale()` is never true, so `Hub::admit` keeps admitting migrations - (`server.rs:252`) instead of failing closed with `Refusal::TipStale`; and -* `cadence_height()` returns the frozen `state.height` rather than free-running - (`batcher.rs:224-227`), so `epoch` stops increasing and **`flush` stops being - called** (`batcher.rs:300-309`) — permanently, because `last_flush_epoch` is a - high-water mark that the non-matching arm leaves untouched. - -The two effects share one root: `is_stale()` and `cadence_height()` are both pure -functions of `last_advance` (`batcher.rs:205-208`, `:222-230`), and `last_advance` -is the field the reorg arm wrongly stamps. Note that `is_stale()` has exactly one -consumer in the whole crate — `Hub::admit` (`server.rs:252`) — so nothing else in -the hub can notice. - -The hub therefore accepts migrations it will never publish, for as long as the -oscillation continues, and the wall-clock fallback that exists precisely to keep -publishing through a stalled tip never engages because the tip does not look -stalled. - -## Attack Scenario and Steps - -1. The hub's tip source returns heights that alternate between `h` and `h-1` on - successive 30-second polls. (Both values are inside `REORG_ALLOWANCE = 10`, so - both are followed.) -2. `observe(h-1)` takes the reorg arm: `state.height = h-1`, `last_advance = now`. - `observe(h)` takes the advance arm: `state.height = h`, `last_advance = now`. - `last_advance` is thus refreshed roughly every 30 seconds, forever. -3. `is_stale()` never fires. `Hub::admit` continues to admit: `observed_height()` is - `h` or `h-1`, and `survives_next_flush(expiry, h, 20, 4)` passes easily for any - real wallet transaction, so migrations flow into the queue normally and shims are - acknowledged. -4. `cadence_height()` returns `h` (or `h-1`). `epoch` therefore takes at most two - adjacent values, and `last_flush_epoch` is only ever *raised* — the non-matching - arm is `_ => {}` (`batcher.rs:308`), which leaves it alone. So after **at most - one** further flush (the one that fires if the loop happens to see the higher of - the two epochs first), `epoch > previous` never matches again. **Publication - stops.** -5. Migrations accumulate until `MAX_QUEUE_BYTES` (64 MiB) is reached, after which - further submissions are refused `Full`. Everything held expires unpublished. On - the deployed mixnet path none of this reaches the wallet: `shim/src/hub.rs:231-240` - answers `Submit::Accepted` at mixnet hand-off and its comment states that a hub - refusal "is never surfaced here", so the wallet was told `error_code 0` for every - one of them. -6. The only signal to the hub operator is the *absence* of `"flush published"` lines - and the presence of `"chain tip moved backwards within the reorg allowance"` - warnings — a message whose own comment says it is expected to be rare, and which - is not wired to any alarm. - -**Attack Requirements and Assumptions:** -- Deliberate: the hub's single indexer (`deploy.env.example:22`) simply alternates - two heights. Costs one integer per poll, needs no transactions and no mixnet - position, and — unlike inflating the tip — produces no implausible values that a - future sanity check would catch, because a 1-block regression is exactly what a - real reorg looks like. -- Accidental, and this is the path that makes the issue worth fixing regardless of - any adversary. What is required is that the **observed** height oscillate within - the allowance while making no net progress. Two realistic shapes produce that: - - **One endpoint, two backends.** An indexer address behind a load balancer over - two nodes a block apart answers from whichever backend it picks, so the observed - height alternates. While both backends advance this is harmless jitter; once - they stop advancing (their node halts, or their sync stalls) the alternation - continues at a fixed pair of heights and the hub silently stops publishing. - - **Several endpoints, one of them flaky — and note this gets *worse* with the - multi-endpoint configuration the project recommends.** `tip_height` returns the - `max()` over the endpoints that answered (`chain.rs:161-173`), so if the - highest endpoint intermittently times out the max itself alternates between - `h` and `h-1`. Combine that with a stalled set and the same freeze follows. See - `tip-and-verdict-aggregation-scale-in-opposite-directions-…`. -- **Which stall is destructive matters, and the filed text did not separate them.** - If the *chain* has halted, nothing expires (expiry is a height), the queue simply - waits, and publication resumes when the chain does — the harm is bounded to - delay and to holding plaintext longer than designed. If instead the *indexers' - view* has stalled while the chain keeps advancing — a routine indexer failure, and - precisely the case `REVIEW.md` #8's free-running clock exists for — then admission - keeps accepting transactions that are aging toward their expiry inside a queue - that will not be flushed, and they are destroyed. This second case is the graded - one. -- No shim, no submitter and no chain observer capability is required. - -## Impact on Users - -Every migration admitted during a freeze that outlives its expiry is destroyed: the -wallet was told it was sent, the transaction is never broadcast, and it ages out -inside the hub's RAM. When the tip source eventually recovers, the first flush -offers the expired entries to the node, which refuses them; `flush` classes those -`Rejected` and **drops** them (`batcher.rs:368`) rather than requeueing. The user -has no error to react to and no copy to retry. - -**Stated precisely, so it is not overstated.** The user's *funds* are not lost — the -transaction was never broadcast, so the note is not spent — but the submission is -gone and the user believes it succeeded. The delay before they can discover that is -~50 minutes for ordinary traffic and **30 to 60 days** for the ZIP 318 migration the -product exists for; that delay is owned by -`zip318-canonical-expiry-is-the-only-recovery-clock-…` and is cited, not re-counted. - -> **CORRECTION 2026-08-18 (validation of the cited file — SUPERSEDES the sentence above).** -> The wallet does **not** wait for expiry. Both official Zcash light-wallet SDKs -> automatically resubmit a sent-but-unmined transaction for as long as it remains -> unexpired — the Android SDK at the head of every ~20 s sync loop and after every -> processed block batch (`CompactBlockProcessor.kt:573,615,723`; selection -> `mined_height IS NULL AND expiry_height > ?`), the iOS SDK at most once per 300 s -> (`TxResubmitter.swift:8-15`, `TransactionDao.swift:218-228`) — and the hub's -> payload-hash dedup makes the resend free. The wallet's non-confirmation signal comes -> from compact-block scanning, which the shim does not intercept (`proxy.rs:1068-1074`). -> Expiry is therefore the **retry horizon**, not the wait: ~50 minutes for the ZIP 203 -> default traffic the shim also diverts, 30–60 days for a ZIP 318 migration. A -> *transient* loss self-heals within minutes; only a loss condition that **outlives the -> horizon** destroys the submission permanently — which is exactly what this issue's -> condition does, so this issue's severity is unaffected. Do not write "the user waits -> 30 to 60 days" in the report. Full refutation and the replacement paragraph: -> `issues/invalid/zip318-canonical-expiry-is-the-only-recovery-clock-and-a-lost-migration-freezes-the-users-notes-for-30-to-60-days.md`. - - -Secondarily, the enclave holds a growing set of plaintext migrations (up to 64 MiB) -for an unbounded period rather than the few minutes the design assumes, which -enlarges the blast radius of the stated residual "the hub sees every migration in -plaintext" — compromise or compulsion against the hub during the freeze now yields -hours or days of accumulated traffic rather than one flush window. - -## Technical Details / Code Analysis - -`hub/src/batcher.rs:161-197`, with the two arms that both refresh the timestamp: - -```rust - if height > state.height { - state.height = height; - state.last_advance = Instant::now(); // an advance: correct - return; - } - - if height < state.height { - let drop = state.height - height; - if drop <= REORG_ALLOWANCE { - tracing::warn!(drop, "chain tip moved backwards within the reorg allowance; following it"); - state.height = height; - state.last_advance = Instant::now(); // NOT an advance - } else { /* ignored */ } - } -``` - -Note also that `height == state.height` falls through both arms and does *not* -refresh `last_advance` — so a tip source that reports a **constant** height is -correctly detected as stale after 15 minutes. It is only the oscillating case that -defeats detection, which is why this is a live gap rather than a redundant one: the -straightforward failure is handled, the near-miss is not. - -`hub/src/batcher.rs:217-230` — the free-running fallback is gated on the same -timestamp, so masking staleness also disables the fallback: - -```rust - fn cadence_height(&self) -> u32 { - let state = self.read(); - if !state.observed { return 0; } - let elapsed = state.last_advance.elapsed(); - if elapsed <= TIP_STALE_AFTER { - return state.height; // <-- always taken while oscillating - } - let estimated_blocks = (elapsed.as_secs() / NOMINAL_BLOCK_SECS) as u32; - state.height.saturating_add(estimated_blocks) - } -``` - -`hub/src/batcher.rs:299-312` — with a constant `cadence_height`, `epoch` is constant -and the flush arm is unreachable. - -The module documentation at `batcher.rs:23-25` states the intended invariant that -this breaks: *"Staleness is a wall-clock fact (no node has advanced for -`TIP_STALE_AFTER`) …"* — the code's condition is "no node has been observed at all", -not "no node has advanced". - -## Recommendations - -- Do not refresh `last_advance` in the reorg arm. Follow the regression (adopting the - lower height is the right call), but leave the advance timestamp alone: the tip has - not advanced, and 15 minutes without an advance is exactly the condition the design - wants to detect. If liveness evidence from a reorg is considered valuable, track it - in a *separate* timestamp that does not gate `is_stale` or `cadence_height`. -- Add a test that drives `observe` with an oscillating sequence and asserts that the - tracker becomes stale, and a test that asserts a frozen `cadence_height` eventually - free-runs. Neither behaviour is covered today (see the companion test-coverage - issue). -- Consider alarming when a flush has not happened for more than ~2 cadence intervals - of wall-clock time, which would catch this class of freeze regardless of cause. - -## Validation Information - -**Verdict: CONFIRMED. Severity: Medium (as filed).** The defect is real, the "both -halves at once" claim is the distinguishing one and it holds, and the reachability -bound is the same one this audit applies to every tip finding. - -### Both halves were verified independently, because that claim is what this issue turns on - -**Half 1 — the fail-closed admission gate never fires.** `is_stale()` is -`!observed || last_advance.elapsed() > TIP_STALE_AFTER` (`batcher.rs:204-208`) and -nothing else. Under `h, h-1, h, h-1, …` at the 30 s poll interval -(`batcher.rs:71`), every observation lands in one of the two arms that stamp -`last_advance` — the advance arm at `:171-175` and the reorg arm at `:179-187` — so -`elapsed()` is bounded by 30 s forever and `TIP_STALE_AFTER` (15 min, `:62`) is -never reached. `Hub::admit` (`server.rs:248-254`) therefore never returns -`Refusal::TipStale`, and admission proceeds into `queue.admit(..., -observed_height(), ...)` where `survives_next_flush` (`queue.rs:380-393`) compares -the transaction's expiry against `next_flush_height(h, 20) + 4 <= h + 24`. Because -`h` is frozen *below* the advancing real chain, a freshly built wallet transaction -(`expiry = build + 40`) clears that bar by an ever-growing margin — so admission does -not merely stay open, it gets **easier** the longer the freeze runs. - -**Half 2 — the free-running cadence never engages, and publication stops for good.** -`cadence_height()` (`batcher.rs:217-231`) returns `state.height` unchanged whenever -`last_advance.elapsed() <= TIP_STALE_AFTER`, which half 1 guarantees. `epoch` is -therefore `h/20` or `(h-1)/20`. In `run_with_poll_interval` the flush arm is -`Some(previous) if epoch > previous` and the fall-through is `_ => {}` -(`batcher.rs:300-309`), which does **not** lower `last_flush_epoch`. So once the -higher of the two epochs has been seen, no later observation can match. Verified by -enumerating both phases: -- `h mod 20 != 0`: both values map to the same epoch; the arm never matches; **zero** - further flushes. -- `h mod 20 == 0`: the values map to `e` and `e-1`; at most **one** further flush - (whichever ordering the loop sees first), then never again. - -Either way the hub admits indefinitely and publishes nothing. Both halves of -`REVIEW.md` #8's stale-tip response — *"stop admitting"* and *"keep the cadence -running off a free-running wall-clock clock"* — are disabled by the same one-line -mistake. That simultaneity is exactly what distinguishes this issue from its -siblings, and it is confirmed. - -### The `height == state.height` case is the proof that this is a bug, not a design choice - -An equal reading falls through both `if` blocks and does **not** stamp -`last_advance`, so a tip source reporting a *constant* height is correctly detected -as stale after 15 minutes. The author's model is therefore "advance", not "answered" -— and the reorg arm silently departs from it. `REVIEW.md` #8 (`hub/REVIEW.md:103`) -says *"Declare the tip stale only when no node has advanced for 15 minutes"*, and -`batcher.rs:23-24` restates it. The straightforward failure is handled; the -near-miss is not. - -Removing the stamp was checked for regressions and is safe: after a legitimate -reorg the chain rebuilds within `REORG_ALLOWANCE = 10` blocks, i.e. ~12 minutes at -the nominal rate, and the first rebuilt block is an advance that re-stamps -`last_advance` normally. A deeper or slower rebuild *should* register as stale. - -### Reachability, and the same bound as every other tip finding - -- **Deliberate:** requires control of a configured `ZIH_INDEXERS` endpoint. That is - item 6p's bound, applied here exactly as it was to - `hub-tip-advance-unbounded-flush-clock.md` (corrected High → Medium) and to - `a-constant-tip-offset-…` (confirmed Medium). This is a hub-trust / robustness - defect, **not** an internet-reachable weapon. Worth noting that it is the *least* - detectable of the three: a 1-block regression is indistinguishable from a real - reorg, so no plausibility check on the *value* can catch it. -- **Accidental:** genuinely reachable with no adversary, but the preconditions are - narrower than the filed text implied and were tightened during validation. The - observed height must oscillate within the allowance *while making no net - progress*, which needs a stalled and heterogeneous tip source; and the destructive - variant additionally needs the real chain to keep advancing while the indexers' - view does not. Both are ordinary indexer-operations failures, and the second is - precisely the scenario `REVIEW.md` #8's free-running clock was written for. -- **Not reachable by an anonymous party**, by a shim, or by a submitter. - -### Corrections applied against the filed text - -1. *"`flush` is never called"* → at most one further flush then never again, with the - `h mod 20` case analysis. The filed absolute was very nearly right but not - exactly. -2. *"Everything held expires unpublished"* → separated into the chain-halted variant - (benign: nothing expires, because expiry is a height) and the indexer-stalled - variant (destructive). The filed text conflated them, which would have overstated - the harm in one case and understated the mechanism in the other. -3. *"Every migration admitted during the freeze is lost"* → the submission is - destroyed; the funds are not. Same softening the validator of - `a-constant-tip-offset-…` applied, for consistency. -4. Added the multi-endpoint accidental path — a flaky *highest* endpoint makes - `max()` itself oscillate — which also makes this the fourth item on - `tip-and-verdict-aggregation-…`'s "adding endpoints makes it worse" list. - -### Checked and NOT claimed - -- **"A stale tip causes an early flush" is REFUTED and is not asserted here.** This - issue's whole point is the opposite: the free-run never engages at all. Nothing in - the text claims an early publication, and nothing should be added that does. -- **The detection gap is cited, not re-counted.** `/healthz` is unconditionally 200 - and no health surface reads `is_stale()`; that belongs to - `hub-health-surface-blind-to-the-states-that-destroy-migrations.md`. -- **The missing tests are cited, not re-counted** — - `hub-batcher-staleness-and-free-run-paths-are-untested.md` owns the fact that no - test drives `observe` with an oscillating sequence. Verified: the five - `TipTracker` tests (`batcher.rs:473-516`) cover fresh-tracker, advance, small - regression, large regression and cadence-while-fresh, and **none of them exercises - the passage of time** — `is_stale()` and `cadence_height()` are only ever asserted - immediately after an observation, so no test can distinguish a refreshed - `last_advance` from a stale one. There is no clock seam to write such a test - through. -- **The plaintext-accumulation residual is a secondary note, not the grade.** The - 64 MiB ceiling is real (`queue.rs:65`) and there is no expiry-based eviction, but - the queue-growth harm is owned by `hub-queue-requeue-ignores-byte-budget-…`. - -### Boundary against the three siblings in `TipTracker::observe` - -This is the **mirror image** of `hub-tip-overshoot-latches-hub-permanently-stale.md` -(confirmed Medium): that issue pins `is_stale()` permanently **true** and the -free-running cadence permanently **on** (admission dead, empty flushes forever); this -one pins them permanently **false** and **off** (admission wide open, no flush ever). -Both come from the same root confusion in `observe` — treating *receipt of an -observation* as *evidence the chain advanced* — but they are triggered by opposite -inputs and need different fixes: - -| Issue | Trigger | `is_stale()` | Cadence | Fix | -|---|---|---|---|---| -| `hub-tip-advance-unbounded-flush-clock.md` (Medium) | repeated advances | false | runs fast | bound the forward advance | -| `a-constant-tip-offset-…` (Medium) | constant offset from first observation | false | normal | plausibility / cross-endpoint check | -| `hub-tip-overshoot-latches-…` (Medium) | one advance beyond the allowance | **stuck true** | free-runs from a bogus height | forward bound + self-healing regression arm | -| **this issue** (Medium) | oscillation within the allowance | **stuck false** | **frozen** | do not stamp `last_advance` on a regression | - -The forward-bound fix recommended by the first issue does **not** fix this one, and -this one's fix does not fix any of the other three. That is why they are four files. - -### Severity justification — Medium - -*Impact:* the hub silently stops publishing while continuing to accept, so every -migration admitted during an indexer-side stall long enough to outlive its expiry is -destroyed after the wallet was told it succeeded — and the mechanism specifically -defeats the safety net the design built for that exact scenario. - -*Likelihood:* the deliberate form needs a configured indexer endpoint (item 6p's -bound). The accidental form needs a stalled, oscillating tip source, which is an -ordinary operations failure but not an everyday one. - -*Why not High:* same bound as its three siblings — not reachable by an anonymous -attacker; grading it above `hub-tip-advance-unbounded-flush-clock.md` would be -inconsistent. - -*Why not Low:* it disables **both** halves of a specified safety response at once, -in the failure mode that response was written for; it needs no adversary; nothing in -the hub's telemetry or health surface can see it; and the loss it produces is the -wallet-acknowledged silent destruction the whole design works to avoid. - -DO NOT DEVIATE FROM THIS TEMPLATE (`docs/SECURITY-ISSUE-TEMPLATE.md`). -LEAVE THIS MESSAGE IN PLACE SO THAT ALL AGENTS KNOW NOT TO DEVIATE FROM THIS TEMPLATE (`docs/SECURITY-ISSUE-TEMPLATE.md`). diff --git a/zeronym-22aa9851caf68-high-medium/medium/hub-tip-advance-unbounded-flush-clock.md b/zeronym-22aa9851caf68-high-medium/medium/hub-tip-advance-unbounded-flush-clock.md deleted file mode 100644 index 03d332bd..00000000 --- a/zeronym-22aa9851caf68-high-medium/medium/hub-tip-advance-unbounded-flush-clock.md +++ /dev/null @@ -1,348 +0,0 @@ -# A single indexer can drive the hub's flush clock arbitrarily fast, collapsing every batch to size 0 or 1 - -**Severity**: Medium -**Validation Status**: Confirmed -**Location**: `audit-target/zeronym/hub/src/batcher.rs:160-197` (`TipTracker::observe`), `:299-312` (the epoch-crossing flush trigger in `run_with_poll_interval`), consuming `audit-target/zeronym/hub/src/chain.rs:148-173` (`ChainClient::tip_height`). Deployed configuration: `audit-target/zeronym/deploy.env.example:22-23` (`INDEXERS=66.241.124.200:443`, a single endpoint). -**Found by agent:** Local (file audit of `hub/src/batcher.rs`); validated 2026-08-18 -**In scope of audit?** Yes - -## Description - -The hub publishes a batch when, and only when, the chain height crosses a multiple -of `FLUSH_INTERVAL_BLOCKS`. `batcher.rs` is explicit that this clock must not be -influenceable by anyone (`batcher.rs:8-22`): - -> Every conditional trigger is a lever someone else can pull. […] The height is -> taken as the MAX over all nodes that answer […] because a single lagging or -> hostile node would otherwise be a second independent lever on the flush clock: -> a stalled tip freezes flushes, **an advanced tip drains the queue**. - -`TipTracker::observe` implements a careful guard in the *backwards* direction — -`REORG_ALLOWANCE = 10`, beyond which a regression is refused — but there is **no -guard whatsoever in the forwards direction**. Any height greater than the current -one is accepted verbatim, with no bound on how far, and no reference to how much -wall-clock time has passed since the last advance, even though `last_advance` is -right there in the same struct. - -The max-over-endpoints defence in `chain.rs` only removes the lever when there is -more than one *independent* endpoint. The shipped configuration -(`deploy.env.example:22`) has exactly one, and `chain.rs:17-21` acknowledges this -("an indexer is a single funnel in front of a single node"). With one endpoint, -`max()` over one value is that value: the endpoint is the clock. - -The endpoint in the shipped configuration is `na.zec.rocks` — a light-wallet -indexer operator, i.e. a member of the exact adversary class this product exists -to defend users against (`AUDIT-INSTRUCTIONS.md`, attacker #1), and one the audit's -own threat model already names as able to "lie about the tip" -(`AUDIT-INSTRUCTIONS.md`, "Trust boundaries": *hub → indexer … can lie about the -tip and about publish verdicts*). - -## Attack Scenario and Steps - -The attacker is whoever controls the `GetLightdInfo` answers the hub receives: the -indexer operator, or anyone who compromises that indexer. (A third case — anyone -on the network path, if `--indexer-tls` is unset, which `main.rs:33-38` only -*warns* about — is noted for completeness but is **not** part of the graded -finding: the deployed configuration sets `ZIH_INDEXER_TLS`, so per -`docs/AVOIDING-FALSE-POSITIVES.md` §7 it does not carry severity. See the -validation section.) - -1. The hub polls `GetLightdInfo` every `POLL_INTERVAL = 30 s` (`batcher.rs:70-71, - 290-297`). -2. On each poll the attacker returns `block_height = previous + 20` instead of the - true height. -3. `observe()` accepts it unconditionally (`batcher.rs:171-175`), because - `height > state.height`. -4. Immediately below, in the same loop iteration, - `epoch = tip.cadence_height() / 20` has advanced by one, so - `Some(previous) if epoch > previous` matches and `flush(&queue, &chain)` runs - (`batcher.rs:300-309`). -5. A flush therefore happens **every 30 seconds** instead of every ~25 minutes, - publishing whatever arrived in the last 30 s. At the project's own stated - arrival rate (README: "the modal batch is 0 or 1" even at 25 minutes), every - published batch is 0 or 1 transactions, i.e. every migration is published alone. -6. Nothing logs the observed height, so the hub's telemetry shows only - `flush_size` / `achieved_batch_size` falling — which at current adoption is - indistinguishable from normal operation, and the `achieved <= 1` warning at - `batcher.rs:412-420` is *expected* to fire today. - -**Admission control does not throttle this attack for the traffic that matters.** -The natural objection is that racing the tip ahead makes -`queue.rs::survives_next_flush` refuse submissions (`expiry >= next_flush_height(tip) -+ 4`), which would be self-limiting. That is true only for transactions with a -tight expiry. Per `audit-state/SPEC-NOTES.md` §3 (verified against the ZIP 318 -source), a conforming ZIP 318 migration carries a *bucketed absolute* expiry -34,561–69,120 blocks ahead of its broadcast height. At today's tip that is roughly -**39,000 blocks of runway** for the attacker before ZIP 318 traffic starts being -refused. At +20 blocks per 30 s poll that is ~16 hours of continuous maximum-rate -attack; paced at +20 blocks every five minutes it is over a week, and still cuts -the effective batching window from 25 minutes to 5. For the acute use case the -product exists to serve, admission control provides essentially no back-pressure -against an inflated tip. - -**Attack Requirements and Assumptions:** -- Requires control of, or a bug in, the single indexer the hub queries — a party - the threat model already designates as untrusted and able to lie about the tip. - No mixnet position, no shim, and no submission of the attacker's own - transactions are needed. -- Costs nothing: it is one integer in one protobuf field per 30 seconds. -- Also reachable *accidentally*: `chain.rs:158` casts the protobuf `u64` - `block_height` to `u32` with `as`, so any nonsense or sentinel value from a buggy - or misconfigured indexer becomes an arbitrary accepted height, and a - misconfiguration pointing at a different network's indexer produces a wrong - height too. -- The attack is invisible: the observed height is never logged, and there is no - plausibility check to alarm on. - -## Impact on Users - -The batch **is** the anonymity set — it is the entire privacy mechanism the hub -provides. Reducing every batch to a single transaction removes it completely, for -every user of every shim pointed at that hub, while the system continues to report -itself healthy. The transactions are still published, so nothing fails visibly; -the users simply do not get the property they were promised, and the loss is -permanent and retrospective because it is recorded on a public chain. - -Concretely, the indexer operator running this attack learns, for each published -transaction, the exact 30-second window in which it arrived at the hub. Combined -with a colluding or identical shim-side operator — who already knows "client IP C -submitted an Orchard-touching transaction at time T", the residual `REVIEW.md` -accepts because the batch was supposed to break the rest of the link — this -completes IP → txid → balance, which is the exact linkage the product exists to -destroy. - -## Technical Details / Code Analysis - -`hub/src/batcher.rs:160-197` — the whole of `observe`. Note the asymmetry: the -`height < state.height` arm has a bound and a loud log, the `height > state.height` -arm has neither. - -```rust - /// Record a height observed from the network (already the max over nodes). - pub fn observe(&self, height: u32) { - let mut state = self.write(); - - if !state.observed { - state.height = height; - state.last_advance = Instant::now(); - state.observed = true; - return; - } - - if height > state.height { - state.height = height; // <-- no bound of any kind - state.last_advance = Instant::now(); - return; - } - - if height < state.height { - let drop = state.height - height; - if drop <= REORG_ALLOWANCE { // <-- bounded, logged - tracing::warn!(drop, "chain tip moved backwards within the reorg allowance; following it"); - state.height = height; - state.last_advance = Instant::now(); - } else { - tracing::warn!(drop, "ignoring a tip regression larger than the reorg allowance"); - } - } - } -``` - -`hub/src/batcher.rs:290-312` — observation and flush decision in one loop -iteration, so an accepted advance fires a flush on the same poll: - -```rust - match chain.tip_height().await { - Ok(height) => tip.observe(height), - Err(err) => { tracing::debug!(%err, "tip query failed on every node"); } - } - - if tip.is_ready() { - let epoch = tip.cadence_height() / params.flush_interval.max(1); - match last_flush_epoch { - None => last_flush_epoch = Some(epoch), - Some(previous) if epoch > previous => { - flush(&queue, &chain).await; - last_flush_epoch = Some(epoch); - } - _ => {} - } - } -``` - -`hub/src/chain.rs:148-173` — `tip_height` takes `.max()` over the endpoints that -answer. This is exactly what `REVIEW.md` #8 asked for, and it is the right -mitigation against a *lagging* endpoint; it is a no-op against an *advancing* one, -and with `endpoints.len() == 1` it is a no-op against both. - -`hub/src/config.rs:36-46` requires at least one `--indexer` and allows a list, but -`deploy.env.example:22` ships one: `INDEXERS=66.241.124.200:443`. - -`hub/src/queue.rs:380-392` (`survives_next_flush`) is the only feedback path from -an inflated tip, and per `audit-state/SPEC-NOTES.md` §4(a) it is ~1,150× too loose -to constrain the attacker for ZIP 318 traffic. - -Why this is not "the tip is trusted infrastructure, out of scope": `REVIEW.md` #8 -identifies precisely this attacker and this lever, and specifies a mitigation for -it. The mitigation as built (max-over-nodes plus a backwards-only monotonicity -bound) is incomplete: it removes the lagging-node lever and leaves the advancing-node -lever untouched, and the deployed single-endpoint configuration removes even the -lagging-node half. - -## Recommendations - -Bound the forward advance the same way the backward one is bounded, using the -`last_advance` timestamp the struct already keeps: - -- In `observe`, reject (or clamp, and log loudly) any advance larger than - wall-clock permits — e.g. `last_advance.elapsed().as_secs() / NOMINAL_BLOCK_SECS` - plus a generous burst allowance for catch-up after a partition. A chain running - at 75 s/block cannot legitimately deliver 20 blocks in a 30 s poll. -- Independently, enforce a minimum wall-clock interval between flushes in - `run_with_poll_interval`, so that no sequence of tip observations can produce - flushes faster than the cadence is designed to run. -- Log the observed height (an aggregate, not per-entry information) so that tip - manipulation is visible in the hub operator's telemetry at all. -- Fix the `info.block_height as u32` truncation at `chain.rs:158` to a checked - conversion, and treat an out-of-range height as a failed tip query. -- Deploy more than one *independent* indexer, or the max-over-nodes rule from - `REVIEW.md` #8 has nothing to work with. - ---- - -**MARKED ADDENDUM (LocalAuditor, `hub/tests/live_chain.rs` audit, 2026-08-18) — an -in-tree precedent for the missing plausibility check, and a ready-made constant.** -This issue notes that "there is no plausibility check to alarm on". The project has -in fact written one down, in the one place it cannot act: -`hub/tests/live_chain.rs:34-37` asserts `height > 3_000_000` and explains why — -*"mainnet is far past this and it will not regress, so a plausible height proves we -parsed a real answer rather than a default."* That is exactly the floor -`TipTracker::observe`'s first-observation branch lacks, and the reasoning behind it -is the same hazard: `unframe` + prost decode an empty body to `LightdInfo::default()`, -so `tip_height` returns `Ok(0)`, which the tracker adopts unconditionally at startup. -Two consequences for this issue's recommendations, neither changing its substance: -(a) recommendation "reject any advance larger than wall-clock permits" can be paired -with a cheap absolute floor whose value the codebase has already chosen and justified; -and (b) the floor lives in a `#[ignore]`d, environment-gated test that no CI runs, so -it is a comment about production behaviour rather than a check on it. Note also that -`> 3_000_000` does **not** discriminate mainnet from testnet, so it is a defence -against a *default* answer, not against this issue's wrong-network aside on the -`info.block_height as u32` cast at `chain.rs:158`. - -## Validation Information - -**Verdict: CONFIRMED. Severity corrected: High → Medium.** The mechanism is -real, deterministic, and free for the attacker who can reach it; the filed -severity was too high because that attacker is not a stranger on the internet. - -### Every mechanical claim re-verified against the target - -| Claim | Verified at | -|---|---| -| `observe` accepts any forward advance with no bound and no clock reference | `hub/src/batcher.rs:161-193` — the `height > state.height` arm is three lines: assign, stamp `last_advance`, return | -| The backward direction *is* bounded and logged | `hub/src/batcher.rs:178-192` (`REORG_ALLOWANCE = 10` at `:59`) — the asymmetry is exactly as filed | -| One accepted advance fires a flush in the same loop iteration | `hub/src/batcher.rs:287-311`: `tip.observe(height)` then `epoch = tip.cadence_height() / flush_interval`, then `Some(previous) if epoch > previous => flush(...)` | -| Maximum forced flush rate is one per poll | `hub/src/batcher.rs:71` (`POLL_INTERVAL = 30 s`) and the single `flush` call per iteration — a 400-block jump still yields one flush, so 30 s is the floor | -| `tip_height` is `max()` over endpoints | `hub/src/chain.rs:155-173` | -| Shipped configuration has exactly one endpoint | `deploy.env.example:22` — `INDEXERS=66.241.124.200:443`. `max()` over one value is the identity function, so that endpoint **is** the flush clock | -| The `u64 → u32` truncation | `hub/src/chain.rs:158` — `info.block_height as u32`, unchecked | -| A degenerate answer decodes to height 0 and is adopted unconditionally | `hub/src/chain.rs:415-430` (`unframe`: a 5-byte all-zero body gives `declared = 0`, and `LightdInfo::decode(&[])` is `Ok(default)`) feeding `batcher.rs:163-168`, the `!state.observed` branch | -| The `> 3_000_000` plausibility floor exists only in an ignored test | `hub/tests/live_chain.rs:34-37` — the addendum is accurate | -| Nothing logs the observed height | Confirmed by grep: `observe` has no `tracing` call on the accept path, and no other site logs a height | -| The threat model designates this party untrusted | `AUDIT-INSTRUCTIONS.md`, "Trust boundaries": *hub → indexer … can lie about the tip and about publish verdicts*. So this is untrusted input reaching a security-critical decision, not a trusted-infrastructure assumption | -| `REVIEW.md` #8 named this exact lever | `hub/src/batcher.rs:19-22`: *"a stalled tip freezes flushes, **an advanced tip drains the queue**"* — the code documents the attack it does not defend against | -| ZIP 318 expiry gives ~34,561–69,120 blocks of runway, so admission control does not throttle the attack for the traffic that matters | `audit-state/SPEC-NOTES.md:46-47`, `:235` (confirmed against the ZIP source) | - -### The bound that fixes the severity (PROGRESS open item 6p) - -The global audit's severity bound applies verbatim: **all tip manipulation -requires control of a configured `ZIH_INDEXERS` endpoint. This is a -hub-trust / robustness defect, not an internet-reachable weapon.** Nothing in -this issue is reachable by an anonymous party. Concretely: - -- The hub dials literal IPv4 addresses supplied by its own operator - (`hub/src/config.rs:36-46`), with no DNS egress at all - (`hub/deploy/caution/caution.hcl.tmpl`, "no port 53, even though the indexer - is authenticated by NAME"), so there is no name-resolution path to hijack. -- The network-path variant the issue raises ("if `--indexer-tls` is unset — - `main.rs:33-38` only *warns*") is **not** part of the graded finding. The - deployed configuration sets `ZIH_INDEXER_TLS` (`caution.hcl.tmpl`, - `deploy.env.example:23`), and `AVOIDING-FALSE-POSITIVES.md` §7 correctly - discounts a vulnerability that exists only under an explicitly insecure - configuration the shipped one does not use. It remains worth the one-line - hardening, but it does not raise the severity. - -What survives the bound, and why it is still a real finding: - -1. **Given a hostile or buggy indexer, the attack is certain, free and - invisible.** One integer per 30 s. At `n = 1` there is no aggregation to - defeat, no rate check, no plausibility check, and no log line to alarm on. -2. **The party in question is in the adversary set by name.** The audit - instructions rank "the indexer operator the shim fronts" as attacker #1, and - `deploy.env.example` points `BACKEND` (`:16-17`) and `INDEXERS` (`:22-23`) at - the *same host, port and TLS name* — so in the shipped example the hub's - flush clock is held by a party who is simultaneously a shim's backing - indexer. (The full composition is filed separately as - `single-party-composition-the-operator-who-sees-the-wallet-leg-also-owns-the-flush-clock-and-the-publish-gate-in-the-shipped-configuration.md`.) -3. **It is reachable with no attacker at all.** A buggy indexer, one pointed at - the wrong network, or one whose `block_height` overflows 32 bits - (`chain.rs:158`) produces the same effect. That half is a plain robustness - defect. - -### Impact, stated precisely - -Two distinct consequences, and the attacker gets both: - -- **The batching anonymity set collapses.** A 30 s cadence publishes whatever - arrived in 30 s, which at any realistic arrival rate is 0 or 1 transactions. - The batch *is* the anonymity set; there is no other mechanism behind it. The - loss is silent (`achieved_batch_size` falling is indistinguishable from low - adoption, and `batcher.rs:412-420`'s warning is expected to fire today - anyway) and permanent, because it is recorded on a public chain. -- **Beyond ~36 blocks of accumulated offset, ZIP 203-default traffic is - silently destroyed.** `survives_next_flush` (`queue.rs:380-393`) is evaluated - against the inflated `observed_height`, so a wallet's `expiry = build + 40` - stops clearing the bar; the hub answers `Refusal::ExpiryTooTight`, and by - `shim/src/hub.rs:231-240` that refusal never reaches the wallet, which was - told `error_code 0` minutes earlier. The *constant-offset* form of this - (which no rate-based defence proposed here would catch, because no jump ever - occurs) is filed separately as - `a-constant-tip-offset-is-a-tunable-expiry-keyed-admission-filter-that-every-proposed-tip-rate-defence-misses.md`; - this issue's race form reaches the same state as a side effect. - -### Severity justification — Medium - -*Impact if exploited:* severe. The product's entire privacy mechanism is voided -for every user of every shim pointed at that hub, silently and irreversibly, -and past a threshold the same lever destroys migrations the wallet was told had -succeeded. - -*Likelihood:* bounded. It needs control of, or a fault in, an endpoint the hub -operator explicitly configures and could replace. It is not reachable by -"anyone on the internet" (attacker #2), which is what separates this from the -two shim/hub flood findings validated alongside it. - -*Why not High:* the audit's own global pass established that tip manipulation -is a hub-trust defect rather than an internet-reachable weapon, and the -severity guidance reserves High for issues that are *likely* to have severe -impact on many users. Additionally, at today's adoption the achieved batch is -already 0 or 1 — an accepted, documented residual — so the marginal privacy -loss *today* is small. The finding's value is that the defence `REVIEW.md` #8 -specified is not buildable as written, so the property will not appear when -adoption arrives. - -*Why not Low:* the code names this exact attack in its own module docs and -ships a mitigation that is a no-op against it; the deployed `n = 1` removes -even the half that does work; the defect is also reachable by accident; and -there is no telemetry that would let an operator notice either case. - -### Note on `AVOIDING-FALSE-POSITIVES.md` §1 - -§1 warns against flagging code for relying on properties an attacker cannot -violate. It does not apply: the tip is not an externally-enforced invariant here -but a single unauthenticated integer read from a party the project's own threat -model lists as able to lie about it, and `batcher.rs:19-22` states the -consequence of that lie in the code itself. - -DO NOT DEVIATE FROM THIS TEMPLATE (`docs/SECURITY-ISSUE-TEMPLATE.md`). -LEAVE THIS MESSAGE IN PLACE SO THAT ALL AGENTS KNOW NOT TO DEVIATE FROM THIS TEMPLATE (`docs/SECURITY-ISSUE-TEMPLATE.md`). diff --git a/zeronym-22aa9851caf68-high-medium/medium/hub-tip-overshoot-latches-hub-permanently-stale.md b/zeronym-22aa9851caf68-high-medium/medium/hub-tip-overshoot-latches-hub-permanently-stale.md deleted file mode 100644 index 6c56c9b1..00000000 --- a/zeronym-22aa9851caf68-high-medium/medium/hub-tip-overshoot-latches-hub-permanently-stale.md +++ /dev/null @@ -1,374 +0,0 @@ -# One implausibly high tip reading latches the hub into "stale forever": admission stops for the whole fleet and nothing short of a process restart clears it - -**Severity**: Medium -**Validation Status**: Confirmed -**Location**: `audit-target/zeronym/hub/src/batcher.rs:160-197` (`TipTracker::observe`), `:204-208` (`is_stale`), `:217-230` (`cadence_height`); consumed at `audit-target/zeronym/hub/src/server.rs:248-255` (`Hub::admit`). Trigger reachable via `audit-target/zeronym/hub/src/chain.rs:155-173` and the `info.block_height as u32` cast at `chain.rs:158`. -**Found by agent:** Local (file audit of `hub/src/batcher.rs`); validated 2026-08-18 -**In scope of audit?** Yes - -## Description - -`TipTracker::observe` accepts any forward height without bound, and refuses any -regression larger than `REORG_ALLOWANCE = 10`. The two rules compose into a latch: -once the tracker has recorded a height meaningfully above the real chain, **every -subsequent truthful observation is discarded as an implausible regression**, so -`last_advance` never updates again, `is_stale()` becomes permanently true 15 -minutes later, and `Hub::admit` refuses every submission with `Refusal::TipStale` -from then on. The hub keeps running, keeps flushing an empty queue on the -free-running clock, and never recovers until either the real chain climbs to within -10 blocks of the bogus height or the enclave is restarted. - -There is no operator override and no runtime control surface of any kind: -`TipTracker` has no reset, `observe` is the only writer of its state and is called -from exactly two places (`main.rs:63` at boot and `batcher.rs:291` in the poll -loop), and the indexer list is baked into an immutable attested enclave. Whether or -not the console is open (it is not, under the canonical `debug { enabled = false }` -runbook) changes nothing: the console carries output, not commands. - -**Framing.** This is primarily a *robustness* defect, and it needs no adversary. Any -single implausibly high reading — from a lying indexer, a buggy one, a proxy serving -someone else's cached answer, an endpoint pointed at a different network, or a `u64` -that does not fit the unchecked `as u32` at `chain.rs:158` — is permanent. The -adversarial version is the same mechanism aimed deliberately, and is bounded exactly -as its two confirmed siblings are: it requires control of a configured -`ZIH_INDEXERS` endpoint, so it is a hub-trust defect, not something a stranger on -the internet can reach. - -## Attack Scenario and Steps - -1. The hub's single indexer (`deploy.env.example:22`) answers one `GetLightdInfo` - with `block_height = H + K` for any `K > 10` — say `K = 5000` (~4 days of - blocks). No other lie is ever needed. -2. `observe(H+K)` takes it (`batcher.rs:171-175`): `state.height = H+K`, - `last_advance = now`. -3. The indexer reverts to telling the truth. Every later observation is - `real_height < H+K` by more than 10, so it hits the else arm at - `batcher.rs:188-195`, logs "ignoring a tip regression larger than the reorg - allowance", and **leaves `last_advance` untouched**. -4. Fifteen minutes later `is_stale()` is true (`batcher.rs:205-208`) and stays true. - `Hub::admit` (`server.rs:252-254`) now returns `Refusal::TipStale` for every - submission, from every shim, indefinitely. -5. Meanwhile `cadence_height()` free-runs upward from `H+K` - (`batcher.rs:224-229`), so `run` keeps crossing epochs and calling `flush` on an - empty queue every ~25 minutes. From the outside — and from the hub's own logs, - which print `flush_size = 0, "flush: nothing held"` (`batcher.rs:345-351`) — this - is indistinguishable from "no one is migrating today". -6. Recovery requires `real_height >= H + K - 10`, i.e. `K - 10` blocks ≈ 4 days for - `K = 5000`. With a large `K` (for example the `u32::MAX` that - `chain.rs:158`'s unchecked `as u32` cast produces from a nonsense `u64`), it never - happens. -7. And "recovery" is not the end of it. During the latch `cadence_height()` has been - free-running from `H+K` for the whole `K - 10` blocks of real time, so - `last_flush_epoch` has climbed to roughly `(H + 2K)/20` while the real epoch is - `(H + K)/20`. `last_flush_epoch` is a monotone high-water mark - (`batcher.rs:272`, `:300-309`), so when admission reopens **no flush happens for - a further ~`K` blocks**, while admission is wide open. That second-order effect is - the separately filed `hub-free-run-overshoot-suppresses-flushes-after-recovery.md` - and is not counted again here; it is noted because it means the natural recovery - path does not restore service either. - -**Why the refusal is silent rather than a visible failure.** On the deployed mixnet -path the shim's submit is dispatch-only: `shim/src/hub.rs:231-240` returns -`Submit::Accepted` to the wallet the moment the frame reaches the mixnet, and its own -comment says *"There is no Refused arm: the hub's verdict is a full round trip away -and is deliberately not waited for … so a refusal is never surfaced here."* The -`TipStale` refusal therefore never reaches the wallet. The wallet has been told -`error_code 0`, keeps no retry state for it, and the transaction is simply gone. - -**Attack Requirements and Assumptions:** -- **No adversary is required.** One wrong high reading from the hub's own indexer, - whatever its cause, is sufficient and permanent. This is the primary path. -- The deliberate version needs **control of a configured `ZIH_INDEXERS` endpoint** - — a party the audit's threat model already designates as able to lie about the - tip, and the party `deploy.env.example:22` points at. It is not reachable by an - anonymous party. This is the same bound applied to the two confirmed siblings - (`hub-tip-advance-unbounded-flush-clock.md`, `a-constant-tip-offset-…`) and it is - what caps the severity at Medium. -- The plaintext-hop variant ("reachable by a network attacker if `--indexer-tls` is - unset, since `main.rs:33-38` only warns") is **not part of the graded finding**: - the deployed configuration sets it (`deploy.env.example:23`), and - `docs/AVOIDING-FALSE-POSITIVES.md` §7 correctly discounts a vulnerability that - exists only under an explicitly insecure configuration nobody ships. -- With more than one endpoint configured the precondition gets *weaker*, not - stronger: `tip_height` returns the `max()` (`chain.rs:161-173`), so **any one of - `n` endpoints** can latch the hub with one packet. See - `tip-and-verdict-aggregation-scale-in-opposite-directions-…`. - -## Impact on Users - -Every wallet that submits a migration while the hub is latched is told its -transaction was sent, and it never is. There is no error, no retry, and no record -anywhere: the shim keeps none, the hub refuses admission so the queue keeps none, -and the frame dies at the hub. - -**Stated precisely, so it is not overstated.** The user's *funds* are not lost: the -transaction was never broadcast, so the note is not spent and the wallet's own -expiry handling eventually returns it to a spendable state. What is destroyed is the -*submission*, silently, together with the user's belief that the migration happened. -For ordinary (non-ZIP-318) traffic that discovery comes at ~50 minutes; for the ZIP -318 migration the product exists for, the canonical expiry is **30 to 60 days** -away, and nothing in the system produces a non-confirmation signal before then. That -delay is owned by -`zip318-canonical-expiry-is-the-only-recovery-clock-and-a-lost-migration-freezes-the-users-notes-for-30-to-60-days.md` -and is cited here, not re-counted. - -> **CORRECTION 2026-08-18 (validation of the cited file — SUPERSEDES the sentence above).** -> The wallet does **not** wait for expiry. Both official Zcash light-wallet SDKs -> automatically resubmit a sent-but-unmined transaction for as long as it remains -> unexpired — the Android SDK at the head of every ~20 s sync loop and after every -> processed block batch (`CompactBlockProcessor.kt:573,615,723`; selection -> `mined_height IS NULL AND expiry_height > ?`), the iOS SDK at most once per 300 s -> (`TxResubmitter.swift:8-15`, `TransactionDao.swift:218-228`) — and the hub's -> payload-hash dedup makes the resend free. The wallet's non-confirmation signal comes -> from compact-block scanning, which the shim does not intercept (`proxy.rs:1068-1074`). -> Expiry is therefore the **retry horizon**, not the wait: ~50 minutes for the ZIP 203 -> default traffic the shim also diverts, 30–60 days for a ZIP 318 migration. A -> *transient* loss self-heals within minutes; only a loss condition that **outlives the -> horizon** destroys the submission permanently — which is exactly what this issue's -> condition does, so this issue's severity is unaffected. Do not write "the user waits -> 30 to 60 days" in the report. Full refutation and the replacement paragraph: -> `issues/invalid/zip318-canonical-expiry-is-the-only-recovery-clock-and-a-lost-migration-freezes-the-users-notes-for-30-to-60-days.md`. - - -Secondarily, this is a complete availability kill on the hub, which -`REVIEW.md`'s "Decisions for humans" section already flags as the cost of -fail-closed design ("it hands any DoS-capable attacker a total availability kill"). -The new part here is that the kill is (a) permanent rather than lasting as long as -the attack, (b) triggerable by a single packet or a single bug, (c) *silent* to the -user because of the dispatch-only submit, and (d) invisible to the operator, because -`is_stale()` is read at exactly one place in the whole crate — `Hub::admit` -(`server.rs:252`) — and by nothing on the health surface (that gap is owned by -`hub-health-surface-blind-to-the-states-that-destroy-migrations.md`). - -**And the only remedy is itself costly.** Clearing the latch requires terminating and -relaunching the hub process. The hub's Nym identity lives in an `Ephemeral` store in -RAM, and the module says so itself (`nym_driver.rs:33-36`): *"What still changes the -address: a real process restart (the store is in RAM, and a diskless enclave has -nowhere to persist it)."* So the restart hands every shim in the fleet a dead -`ZIS_HUB_NYM`, and a shim is an immutable managed app — each one needs a re-assemble -and redeploy, spending one of its five weekly certificate issuances -(`restarts-ledger-budget-model-omits-hub-forced-redeploys-…`, confirmed). The -restart also destroys whatever the RAM queue still holds. A one-packet fault -therefore costs a fleet-wide reconfiguration. - -## Technical Details / Code Analysis - -The latch is the interaction of three pieces that are individually reasonable. - -`hub/src/batcher.rs:177-196` — regressions beyond the allowance are ignored *and do -not refresh `last_advance`*: - -```rust - if height < state.height { - let drop = state.height - height; - if drop <= REORG_ALLOWANCE { - /* follow it, refresh last_advance */ - } else { - tracing::warn!(drop, "ignoring a tip regression larger than the reorg allowance"); - } - } -``` - -`hub/src/batcher.rs:204-208` — staleness is a function of `last_advance` only: - -```rust - pub fn is_stale(&self) -> bool { - let state = self.read(); - !state.observed || state.last_advance.elapsed() > TIP_STALE_AFTER - } -``` - -`hub/src/server.rs:248-255` — a stale tip stops admission entirely: - -```rust - pub fn admit(&self, tx_bytes: &[u8]) -> Result, Refusal> { - if self.tip.is_stale() { - return Err(Refusal::TipStale); - } -``` - -`hub/src/chain.rs:155-159` — the cast that lets an out-of-range `u64` become an -arbitrary `u32` height: - -```rust - let queries = self.endpoints.iter().map(|addr| async move { - let info: LightdInfo = self.unary(*addr, GET_LIGHTD_INFO, Empty {}).await?; - Ok::(info.block_height as u32) - }); -``` - -**The root error, stated in one line.** `observe` conflates *"I received an -observation"* with *"the chain advanced"*. The large-regression arm treats an -observation as evidence to discard and thereby latches; the small-regression arm -(`batcher.rs:179-187`, the separately filed -`hub-reorg-branch-resets-last-advance-masks-stale-tip.md`) treats an observation as -evidence of liveness and thereby masks. The two are exact opposites of each other: -this issue pins `is_stale()` permanently **true** and the free-running cadence -permanently **on**; that one pins them permanently **false** and **off**. - -Note that this is not the same defect as the unbounded-forward-advance issue filed -separately (`hub-tip-advance-unbounded-flush-clock.md`, confirmed Medium), although -it shares the same missing guard: that one is about *repeated* advances driving the -flush clock, this one is about a *single* advance making the tracker permanently -unable to accept reality. Fixing the forward bound fixes both; fixing only the -flush-rate limit fixes neither. - -The free-running clock is not a mitigation here. `REVIEW.md` #8 specifies it so that -"the cadence keeps running off a free-running wall clock" during a stall — and it -does, but it free-runs from the *bogus* height, so it publishes nothing (admission -is closed) and drifts further from reality with every hour. - -## Recommendations - -- Bound forward advances in `observe` against `last_advance.elapsed()` (see the - companion issue); an advance that wall-clock cannot justify should be rejected and - logged, exactly as an implausible regression already is. -- Make the "implausible regression" arm self-healing: if the same lower height (or a - monotone sequence of them) persists for longer than `TIP_STALE_AFTER`, the - tracker's own recorded height is the outlier, not the network's — adopt it and log - loudly. A tracker that can never be corrected by the only source of truth it has - is a latch by construction. -- Replace `info.block_height as u32` with a checked conversion, treating an - out-of-range height as a failed tip query. -- Surfacing hub refusals to the wallet on the mixnet path would turn this from - silent destruction into a visible failure. That fix is **owned by - `nym-submit-acks-are-never-read-so-every-hub-refusal-is-invisible.md`** (confirmed - Medium) and should be taken there, not here. *Correction applied during - validation:* the filed parenthetical "have the shim retry against the other hub" - does not work — `nym.rs:642` already submits every migration to **every** - configured hub address (`shim-submits-every-migration-to-every-configured-hub-…`, - confirmed High), so there is no unused hub to fail over to. -- Add the aggregate telemetry that would make this state visible at all: the - observed height and the fact of a latched tracker are logged nowhere, and - `is_stale()` has exactly one consumer in the crate. Both are aggregates, so the - counts-only rule (#157) permits them. (Owner for the health-surface half: - `hub-health-surface-blind-to-the-states-that-destroy-migrations.md`.) - -## Validation Information - -**Verdict: CONFIRMED. Severity: Medium (as filed).** The latch is real, permanent, -and re-derivable from three short functions; it needs no attacker at all; and the -adversarial form of it carries the same precondition as its two confirmed siblings, -which is what holds it at Medium rather than High. - -### Every mechanical claim re-verified against the target at HEAD - -| Claim | Verified at | -|---|---| -| A forward advance of any size is accepted verbatim and stamps `last_advance` | `hub/src/batcher.rs:171-175` | -| A regression larger than `REORG_ALLOWANCE` is ignored **and does not stamp `last_advance`** | `hub/src/batcher.rs:177-196`; `REORG_ALLOWANCE = 10` at `:59` | -| `is_stale()` reads nothing but `observed` and `last_advance` | `hub/src/batcher.rs:204-208` | -| A stale tip refuses **every** admission, on both ingress paths | `hub/src/server.rs:248-254`; `Hub::admit` is the single funnel for HTTP (`server.rs:541`) and mixnet (`nym.rs:321`) | -| `cadence_height()` free-runs from the *stored* height, which is the bogus one | `hub/src/batcher.rs:217-231` — the estimate is never written back into `state.height` | -| The `u64 → u32` truncation that can manufacture an arbitrary high reading | `hub/src/chain.rs:158` — `info.block_height as u32`, unchecked | -| `blockHeight` is `uint64` field 7 of `LightdInfo`, so an out-of-range value is representable on the wire | `zaino/packages/zaino-proto/proto/service.proto:105` | -| The shipped configuration is a single endpoint | `deploy.env.example:22` | -| The refusal is never surfaced to the wallet on the deployed transport | `shim/src/hub.rs:231-240` — no `Refused` arm, `Submit::Accepted` on mixnet hand-off | - -### The decisive question: is recovery genuinely impossible short of a restart? YES - -This was checked exhaustively rather than assumed. - -- **`observe` is the only writer of `TipState`.** A repository-wide grep over - `hub/src` and `hub/tests` returns exactly two production call sites: - `main.rs:63` (the boot seed) and `batcher.rs:291` (the 30 s poll). `TipTracker` - exposes no reset, no setter, and no constructor path reachable after startup. -- **No configuration change helps.** `--indexers` is startup-only - (`hub/src/config.rs:36-46`) and the enclave image is immutable, so the operator - cannot even repoint the hub at a different indexer without a redeploy. -- **No console command exists.** Under the canonical runbook - (`debug { enabled = false }`) the parent has no console at all; under - `deploy.sh`'s `DEBUG=1` default it has a read-only one. Neither is a control - channel. There is no admin endpoint: `server.rs` serves `/healthz`, - `/nym-status`, `/nym-address`, lookup and (optionally) submit, and nothing else. -- **The three ways out, all verified:** - 1. the real chain climbs to within `REORG_ALLOWANCE` of the bogus height — for - `K = 5000` that is ~4.3 days, and for a `K` produced by the `as u32` - truncation it is effectively never (`u32::MAX` ≈ 4.29e9 blocks ≈ 10,000 years - at 75 s/block); - 2. the *same* source reports a still-higher value, which is an advance and does - stamp `last_advance` — i.e. only the party who broke it can fix it, and doing - so gives them an on/off switch over fleet-wide admission rather than a repair; - 3. a process restart, whose cost is the confirmed fleet-wide consequence recorded - in Impact above. - -So the filed claim survives, and it is stronger than filed: even route 1 does not -restore *publication* promptly, because `last_flush_epoch` has free-run ahead in the -meantime (step 7, added during validation). - -### Reachable by a buggy endpoint, not only a hostile one — and that is the primary framing - -The coordinator's question was whether this is an adversary-only finding. It is not. -`observe` has no plausibility floor, no ceiling, and no clock reference on the -forward path, so a single erroneous high reading from *any* cause is terminal. The -project has already written the plausibility check it needs, in the one place it -cannot act: `hub/tests/live_chain.rs:34-37` asserts `height > 3_000_000` with the -comment *"a plausible height proves we parsed a real answer rather than a default"*. -That test is `#[ignore]`d and environment-gated, and in any case a floor only catches -the *low* direction; nothing anywhere catches the high one. - -Two accidental triggers were checked and are real: the unchecked `as u32` at -`chain.rs:158` (any `u64` ≥ 2^32 becomes an arbitrary `u32`), and an endpoint -serving a height for a chain other than the one the hub publishes to. A third — a -degenerate/empty body decoding to `LightdInfo::default()` — was checked and yields -`block_height = 0`, i.e. the **low** direction, which is the sibling's concern and -not this one. Stating that explicitly so it is not mis-cited later. - -### What was checked and is NOT claimed - -- **"A stale tip causes an early flush" is REFUTED and is not asserted here.** - `cadence_height()` extrapolates from `last_advance.elapsed()` at the nominal block - rate, so the free-running estimate lands back in phase with where the chain would - have been; the observable direction is *late* (the cadence is frozen for the first - `TIP_STALE_AFTER` and then steps forward by ~12 blocks at once), never early. - `REVIEW.md` #8 also says "Never flush early because the tip is stale". The filed - step 5 says only that empty flushes continue at ~25-minute intervals, which is - correct. -- **The batch-collapse / isolation harm is not claimed here.** The one immediate - off-cadence flush that the overshoot itself triggers belongs to - `hub-tip-advance-unbounded-flush-clock.md`; this issue's harm is the state the hub - is left in *afterwards*. -- **Loss of funds is not claimed.** Corrected in Impact above, consistently with the - same softening applied to `a-constant-tip-offset-…` during its validation. -- **The detection gap is cited, not re-counted** (owner: - `hub-health-surface-blind-to-the-states-that-destroy-migrations.md`), and so is the - 30-to-60-day discovery delay (owner: the ZIP 318 expiry issue). - -### Boundary against the three siblings in `TipTracker::observe` - -`observe` is 37 lines and carries four distinct filed defects. The split is -deliberate and each has a different fix: - -| Issue | Input pattern | Effect | Fix | -|---|---|---|---| -| `hub-tip-advance-unbounded-flush-clock.md` (confirmed, Medium) | *repeated* advances | flush clock runs fast, batches collapse to 0 or 1 | bound the forward advance against wall clock | -| `a-constant-tip-offset-…` (confirmed, Medium) | a *constant* offset adopted at first observation | admission threshold moves; no jump, no staleness | absolute/cross-endpoint plausibility; rate-based fixes miss it | -| **this issue** | *one* advance beyond `REORG_ALLOWANCE` of reality | tracker can never re-converge; admission closed forever | the same forward bound, **plus** a self-healing regression arm | -| `hub-reorg-branch-resets-last-advance-…` | oscillation *within* `REORG_ALLOWANCE` | staleness masked; cadence frozen; nothing ever published | stop stamping `last_advance` on a regression | - -The forward bound recommended by the first issue also fixes this one; it does **not** -fix the third or fourth. That is why the four are filed separately rather than merged. - -### Severity justification — Medium - -*Impact:* severe and total. Every migration submitted to a latched hub is silently -discarded, for every shim pointed at it, for as long as the latch holds — which is -"forever" in the accidental `as u32` case. The repair is a fleet-wide redeploy. - -*Likelihood:* bounded by the same precondition as both confirmed siblings — control -of, or a fault in, a configured `ZIH_INDEXERS` endpoint. Not reachable by an -anonymous party (item 6p). Raised somewhat above the siblings by the fact that no -adversary is needed at all, and lowered by the fact that a single wrong high reading -is not an everyday event. - -*Why not High:* the audit applies one consistent bound to every tip-manipulation -finding — it is a hub-trust / robustness defect, not an internet-reachable weapon — -and `hub-tip-advance-unbounded-flush-clock.md` was corrected High → Medium on -exactly this basis. Grading this one higher would be inconsistent with its siblings. - -*Why not Low:* the state is permanent, has no runtime remedy, is invisible on every -health surface the hub exposes, and destroys wallet-acknowledged submissions the -whole time it holds. - -DO NOT DEVIATE FROM THIS TEMPLATE (`docs/SECURITY-ISSUE-TEMPLATE.md`). -LEAVE THIS MESSAGE IN PLACE SO THAT ALL AGENTS KNOW NOT TO DEVIATE FROM THIS TEMPLATE (`docs/SECURITY-ISSUE-TEMPLATE.md`). diff --git a/zeronym-22aa9851caf68-high-medium/medium/hub-unauthenticated-pre-publication-transaction-disclosure.md b/zeronym-22aa9851caf68-high-medium/medium/hub-unauthenticated-pre-publication-transaction-disclosure.md deleted file mode 100644 index a738e893..00000000 --- a/zeronym-22aa9851caf68-high-medium/medium/hub-unauthenticated-pre-publication-transaction-disclosure.md +++ /dev/null @@ -1,367 +0,0 @@ -# Unauthenticated transaction lookup discloses queued, not-yet-published migrations and lets a third party steal the broadcast - -**Severity**: Medium -**Validation Status**: Confirmed -**Location**: `audit-target/zeronym/hub/src/server.rs:429-458` (`handle`, routing), `:487-505` (`lookup`), `:508-517` (`found`), `:296-322` (`Hub::lookup`); `audit-target/zeronym/hub/src/queue.rs:328-348` (`Queue::find_by_txid`); `audit-target/zeronym/hub/src/nym.rs:249-300` (`build_lookup_reply`, the same core over the mixnet); `audit-target/zeronym/hub/src/chain.rs:513-533` (`classify_publish_error`); `audit-target/zeronym/hub/src/batcher.rs:361-378` (`flush`); `audit-target/zeronym/hub/deploy/caution/caution.hcl.tmpl:51-55, 96-105` (`ingress 0.0.0.0/0` + the Caddy-proxied `http` block) -**Found by agent:** Brainstorming Agent; validated and substantially revised by the Issue Validator -**In scope of audit?** Yes - -## Description - -The hub answers a transaction lookup from its **queue first**, and a queue hit -returns the **raw bytes of a diverted migration that has not yet been broadcast -anywhere**, with `x-tx-height: 0`. The lookup is unauthenticated on **both** -transports: - -* **Clearnet.** `POST /transaction` is served unconditionally. Unlike the submit - path `POST /`, which is gated behind `ServeOptions::http_submit` (off by - default) and falls through to a `404`, the lookup path has **no gate at all** - (`server.rs:437-445`). The enclave declares `ingress { cidr_ipv4 = "0.0.0.0/0" }` - and the platform maps the hub's TLS domain onto it, so the endpoint is - reachable by anyone on the internet. -* **Mixnet.** `LookupV1` frames are answered through *exactly the same* - `Hub::lookup` core (`nym.rs:249-300`), the hub's Nym address is published by - design at `GET /nym-address`, and there is no submitter ACL. **Gating the - clearnet path therefore does not close this**; it only raises the cost from - `curl` to a Nym client. - -Three distinct capabilities follow, with very different preconditions: - -1. **Pre-publication disclosure of the transaction body**, to anyone who holds a - candidate txid, up to ~25 minutes before the transaction exists anywhere else - on the network. -2. **A residency oracle.** `200` + `x-tx-height: 0` means "queued here, - unbroadcast"; `404` means "not here"; the 200→404 transition timestamps the - flush. A queue hit also returns in microseconds while a miss pays a full - indexer round trip (up to `RPC_TIMEOUT` = 10 s, `chain.rs:48`), so a latency - oracle survives even if the body and the header were removed. This leg needs - **no txid at all** if the attacker probes with their own submissions. -3. **Broadcast theft.** Whoever retrieves the bytes can broadcast them - themselves, from a node of their choosing, at an instant of their choosing. - The hub's later flush is then answered "already known", which - `classify_publish_error` maps to `Publish::AlreadyKnown`, which `flush` counts - as **achieved** and drops. Nothing anywhere records that the transaction did - not enter the network from the hub, at the hub's cadence, with the batch. - -Leg 3 is the one that defeats a claimed protection: the product's mechanism is -that the transaction is published **by the hub, on a fixed cadence, together -with others**. This lets a third party choose the publisher and the moment, -silently. - -> **CORRECTION carried forward (coordinator, 2026-08-17), and now the settled -> reading.** An earlier revision of this issue also claimed the Nitro **parent -> host** could read these bodies off the wire with a packet capture. **That claim -> is withdrawn.** `caution.hcl.tmpl:96-105` declares -> `http { domain = ...; port = 8083; e2e_encryption { mode = "tls" } }`, which is -> Caution's *in-enclave* TLS termination: the platform runs Caddy inside the -> enclave, holds the private key there, and forwards plaintext only over the -> enclave-internal loopback to `ZIH_LISTEN=0.0.0.0:8083`. The parent host sees -> ciphertext on this hop. The attacker set for this issue is therefore -> **callers** — which is still the entire internet, unauthenticated. The -> withdrawn claim does **not** affect legs 2 and 3, which never depended on it. -> It also does **not** apply to the hub's *outbound* hop to its indexer, where -> `ZIH_INDEXER_TLS` really is the only thing keeping the parent host out of every -> batch (see `caution.hcl.tmpl:128-131` and the separate issues on that hop). - -## Attack Scenario and Steps - -**Leg 3 (the damaging one): steal the broadcast of a user whose txid you hold.** - -1. The wallet is handed the txid the instant it submits: the shim answers - `SendTransaction` with a synthesized `SendResponse { error_code: 0, - error_message: }` computed locally (`shim/src/hub.rs:238-240`, - `shim/src/intercept.rs:179-186`). The user therefore has, and can share, the - txid ~25 minutes before the transaction is public. -2. The attacker obtains that txid out of band — a counterparty or merchant given - it as a payment reference, a support channel, a screenshot, wallet telemetry, - or another app on the device. -3. `POST /transaction` to the hub with the 32-byte wire-order hash (or the - equivalent `LookupV1` frame over Nym). The hub answers `200`, - `content-type: application/octet-stream`, `x-tx-height: 0`, body = the full - raw transaction. -4. The attacker submits those exact bytes to a Zcash node of their choosing, at a - moment of their choosing, before the hub's next 20-block flush. -5. At the flush, `chain::broadcast_batch` offers the same bytes to the hub's - indexer, the node answers already-known, `classify_publish_error` returns - `AlreadyKnown`, `flush` increments `achieved` and drops the entry - (`batcher.rs:367`). `achieved_batch_size` logs a normal, healthy flush. - *(If the node's wording is one `classify_publish_error` does not match, the - entry is classed `Rejected` and dropped instead — the outcome for the user is - identical; only the hub's counter differs.)* - -**Leg 2 (no txid needed): measure the flush.** Submit a transaction of your own -(the mixnet submit path is unauthenticated and the hub's address is published), -compute its txid locally from your own bytes, and poll the lookup. The 200→404 -transition gives the flush instant to the second. - -**Attack Requirements and Assumptions:** - -- **Network access only.** No credential, no enclave compromise, no mixnet - position for the clearnet leg; a Nym client and the published hub address for - the mixnet leg. -- **Legs 1 and 3 require a candidate txid**, which is not guessable - (256-bit) and cannot be enumerated. Under ZIP 244 a txid is computable *from* - the transaction bytes, so nobody derives it from the chain before publication. - The realistic holders are the user and anyone the user tells within the ≤25 - minute window. **This is a targeted attack, not a mass one.** -- **Leg 2 requires nothing** beyond the ability to submit and to look up. -- **What makes it realistic:** the endpoint is on by default, ungated, and - reachable from the whole internet; it has **no legitimate caller in the - deployed topology at all**, because `HubTransport` is clearnet XOR mixnet - (`shim/src/hub.rs:219-227`) and `deploy.env.example:18` configures the mixnet; - and `smoke.sh:315-327` asserts only that `POST /` is closed, never checking - `POST /transaction`. -- **What bounds it:** the marginal value of the *body* over the txid alone is - small — the same bytes are public on chain within ~25 minutes, and - `hub/REVIEW.md:177` already concedes that batch membership is publicly - enumerable after the fact. The genuinely new capability is leg 3. - -## Impact on Users - -- **For a targeted user whose txid reaches a hostile party inside the window:** - the core protection is nullified and nothing notices. Their transaction is - published by an attacker-chosen node at an attacker-chosen instant, so - "published by the hub, on a cadence, mixed with others" does not hold for them, - and the ~25-minute delay that decorrelates their submission from the on-chain - appearance can be collapsed to seconds. Because `AlreadyKnown` counts as - achieved, **no component reports anything wrong**: not the wallet (told - `error_code 0` at submit), not the hub's telemetry, not the operator, not the - user. -- **The batch loses that member**, so every other user in that flush gets a - smaller on-chain anonymity set than the hub believes it delivered. -- **For anyone:** a live residency and flush-timing oracle. Its practical value - is bounded — the cadence is public by design and `hub/REVIEW.md:177` states - that the batch is publicly identifiable from chain data anyway — but it - contradicts the hub's own stated invariant that read-only endpoints must not - become an anonymity-set oracle (`server.rs:22-25`, `queue.rs:297-299`). -- **Pre-publication body disclosure** is real but mostly costs earliness: value - balance, anchor, expiry and action count all become public on chain shortly - afterwards. - -## Technical Details / Code Analysis - -**1. The lookup path is ungated while the submit path is gated** — `hub/src/server.rs:437-445`: - -```rust - match req.uri().path() { - SUBMIT_PATH if options.http_submit => match method { - Method::POST => submit(req, hub).await, - _ => Ok(text(StatusCode::METHOD_NOT_ALLOWED, "POST only")), - }, - TRANSACTION_PATH => match method { - Method::POST => lookup(req, hub).await, - _ => Ok(text(StatusCode::METHOD_NOT_ALLOWED, "POST only")), - }, -``` - -`hub/src/server.rs:487-505` buffers up to `MAX_LOOKUP_BYTES` (64) and rejects -only an *empty* key — no length check, no authentication, no rate limit: - -```rust -async fn lookup(req: Request, hub: Hub) -> Result>, Infallible> { - let collected = match Limited::new(req.into_body(), MAX_LOOKUP_BYTES).collect().await { ... }; - let wire_hash = collected.to_bytes(); - if wire_hash.is_empty() { - return Ok(text(StatusCode::BAD_REQUEST, "empty lookup key")); - } - match hub.lookup(&wire_hash).await { - LookupOutcome::Found { data, height } => Ok(found(&data, height)), - LookupOutcome::NotFound => Ok(text(StatusCode::NOT_FOUND, "transaction not found")), - LookupOutcome::Unavailable => Ok(text(StatusCode::BAD_GATEWAY, "indexer unavailable")), - } -} -``` - -**2. The queue is consulted first and a hit returns unpublished plaintext** — -`hub/src/server.rs:296-303`: - -```rust - pub async fn lookup(&self, wire_hash: &[u8]) -> LookupOutcome { - if let Some(bytes) = self.queue.find_by_txid(wire_hash) { - tracing::debug!(source = "queue", "transaction lookup answered"); - return LookupOutcome::Found { data: bytes, height: 0 }; - } -``` - -`hub/src/queue.rs:328-348` linear-scans the queue, matches **both** byte orders -of the supplied hash, and returns a copy of the entry's bytes. The height -sentinel is an explicit response header (`hub/src/server.rs:508-517`): - -```rust -fn found(tx_bytes: &[u8], height: u64) -> Response> { - let mut resp = Response::new(Full::new(Bytes::copy_from_slice(tx_bytes))); - resp.headers_mut().insert(CONTENT_TYPE, HeaderValue::from_static("application/octet-stream")); - resp.headers_mut().insert(TX_HEIGHT_HEADER, HeaderValue::from(height)); - resp -} -``` - -`server.rs:90-91` documents `0` as "mempool (a queued, unflushed transaction), -matching lightwalletd's sentinel". - -**3. The mixnet path is the same core, also unauthenticated** — -`hub/src/nym.rs:275-279`: - -```rust - let reply = match hub.lookup(&hash).await { - LookupOutcome::Found { data, height } => LookupReply::Found { height, tx: data }, - LookupOutcome::NotFound => LookupReply::NotFound, - LookupOutcome::Unavailable => LookupReply::Error, - }; -``` - -This is why gating `POST /transaction` alone is not a fix: the module header at -`nym.rs:17-18` states outright that admission and lookup are "the exact calls the -HTTP serving path uses", and the hub publishes its Nym address to everyone. - -**4. Already-known is counted as success and the entry is dropped** — -`hub/src/chain.rs:513-533`: - -```rust -fn classify_publish_error(message: &str) -> Publish { - let m = message.to_ascii_lowercase().replace('-', " "); - if m.contains("already in block chain") || m.contains("already known") - || m.contains("already in mempool") || m.contains("duplicate") - { Publish::AlreadyKnown } else { Publish::Rejected { reason: message.to_string() } } -} -``` - -and `hub/src/batcher.rs:365-378`: - -```rust - for (i, entry) in batch.into_iter().enumerate() { - match outcomes.get(i) { - Some(Publish::Accepted { .. }) | Some(Publish::AlreadyKnown) => achieved += 1, - Some(Publish::Rejected { .. }) => rejected += 1, - Some(Publish::Retryable { reason }) => { ...; unplaced.push(entry); } - None => unplaced.push(entry), - } - } -``` - -Only `Retryable`/`None` are requeued, so an `AlreadyKnown` entry is neither -retried nor flagged; it increments the number the hub logs as -`achieved_batch_size`. The comment at `batcher.rs:356-358` explains why -already-known must count as success ("with every shim submitting to every hub, -the second hub's publish is already-known by construction") — sound reasoning -that cannot distinguish a sibling hub from an attacker who published first. - -**5. Exposure** — `hub/deploy/caution/caution.hcl.tmpl:51-55` and `:96-105`: - -```hcl - ingress { cidr_ipv4 = "0.0.0.0/0"; port = 8083; ip_protocol = "tcp" } - ... - http { domain = "__TLS_DOMAIN__"; port = 8083; e2e_encryption { mode = "tls" } } -``` - -The hub process itself speaks plain HTTP/1.1 (`hub/src/tls.rs` is a *client* to -the indexer only); the TLS a caller sees is terminated by the in-enclave Caddy. -The exposure is the open, unauthenticated ingress — not a cleartext wire. - -**6. The stated invariant this sits against** — `hub/src/server.rs:22-25` requires -that these read-only endpoints never become a live anonymity-set oracle, and -`hub/src/queue.rs:297-299` refuses to return queue depth for the same reason. - -## Recommendations - -In order of value: - -1. **Do not serve an unpublished transaction's bytes to an unauthenticated - lookup.** A queue hit can answer `x-tx-height: 0` with an empty body (enough - for a wallet to render "pending"), or `NotFound`. This closes legs 1 and 3 on - **both** transports at once and is the only recommendation that does. -2. **Break the "already known ⇒ achieved" equivalence for an entry this hub has - not previously offered.** Record whether a prior flush of the same entry - returned `Accepted`/`Retryable`, and surface a *first-flush* `AlreadyKnown` as - a counter-level anomaly. Today that event is indistinguishable from success. -3. **Gate `POST /transaction` behind an explicit flag, defaulted off**, the way - `POST /` already is. Note this is a *reduction in exposure, not a fix* — the - mixnet lookup path remains open — but the clearnet endpoint has no legitimate - caller in the deployed mixnet topology, so it is free to remove. -4. **Authenticate the hub's ingress** (the STEVE / mutual-attestation work - `OPEN-QUESTIONS.md` §3 records as designed-not-built). Until it exists, every - hub endpoint is hostile-reachable and should disclose nothing. -5. **Extend `smoke.sh`** to assert the disposition of `POST /transaction` - alongside its existing `POST /` assertion (`smoke.sh:315-327`). - -## Validation Information - -**Verdict: CONFIRMED at Medium** (downgraded from the filed High). - -**What was verified directly in the target:** - -- `handle` routes `TRANSACTION_PATH` with no `options.*` guard while `SUBMIT_PATH` - carries `if options.http_submit` — `server.rs:437-445`. `ServeOptions::http_submit` - defaults false (`server.rs:198`), and `hub/tests/endpoints.rs:170-181` - (`the_lookup_path_is_not_gated_by_the_submit_flag`) pins the ungated behaviour - as intended. -- `Hub::lookup` consults `Queue::find_by_txid` before the indexer and returns - `height: 0` on a hit — `server.rs:296-303`; `queue.rs:328-348` returns a copy of - `entry.tx_bytes`, matching either byte order. -- `found` writes the raw bytes as the body plus `x-tx-height` — `server.rs:508-517`. -- The enclave manifest opens `0.0.0.0/0` on 8083 and the platform maps the TLS - domain onto that same port — `caution.hcl.tmpl:51-55`, `:96-105`. `smoke.sh` - reaches `/nym-status`, `/nym-address`, `/healthz` and `POST /` over the public - URL, so the listener is demonstrably internet-reachable. -- `classify_publish_error` → `AlreadyKnown` and `flush` counting it as `achieved` - without requeue — `chain.rs:513-533`, `batcher.rs:365-378`. -- The shim hands the wallet a locally computed txid at submit time — - `shim/src/hub.rs:231-240` (`Submit::Accepted { txid: crate::nym::local_txid(...) }`), - rendered into the `SendResponse` at `shim/src/intercept.rs:179-190`. This is the - mechanism by which a txid exists, and can leak, before publication. - -**Corrections made to the issue during validation:** - -- The withdrawn parent-host packet-capture claim has been **removed from the body - entirely** rather than left as struck-through text, and the surviving legs - restated on their own merits. The correction note is retained once, as the - record of why. -- **The filed recommendation "gate `POST /transaction` … closes the whole issue - for the deployed topology in one line" was wrong and has been corrected.** - `hub/src/nym.rs:249-300` answers `LookupV1` through the identical `Hub::lookup` - core, the hub's Nym address is published at `GET /nym-address` by design, and - there is no submitter ACL — so the disclosure and the oracle both survive the - clearnet gate. Only refusing to serve unpublished bytes closes them. -- Added that a mis-matched already-known message classes as `Rejected` and is - *also* dropped, so leg 3 does not depend on `classify_publish_error`'s string - list matching the node's wording. - -**Why Medium and not High:** - -- Legs 1 and 3 are gated on holding a 256-bit txid within a ≤25 minute window. - Under ZIP 244 a txid is computed *from* the transaction, so no chain observer, - mempool watcher, or the operator can derive one before publication; realistically - only the user and parties the user tells hold it. That makes this a **targeted** - attack against a user whose txid leaked, not a mass one — it fails the "many - users" bar for High. -- The marginal value of the disclosed *body* over the txid alone is roughly 25 - minutes of earliness: the same transaction is public on chain shortly after, and - `hub/REVIEW.md:177` already concedes batch membership is publicly enumerable. -- The freely-available leg (the residency/flush oracle) measures something the - design already treats as public: the cadence is deliberately deterministic and - publicly known (`batcher.rs:8-17`), and the batch is publicly identifiable - (`REVIEW.md:177`). It is a real violation of `server.rs:22-25`'s stated - invariant, but its independent user harm is small. - -**Why not Low, and why not invalid:** - -- Leg 3 is a genuine, silent defeat of a claimed protection for the user it hits: - the transaction leaves the batch, the timing decorrelation is gone, and every - detection channel reports success. That is a serious outcome available in - realistic (if particular) circumstances, which is the definition of Medium. -- The exposure is gratuitous: there is no legitimate clearnet caller in the - deployed topology, since `HubTransport` is clearnet **xor** mixnet - (`shim/src/hub.rs:219-227`) and `deploy.env.example:18` configures the mixnet. - -**Related issues, deliberately not restated here:** the unbounded concurrency / -upstream-amplification consequence of the same ungated endpoint is filed as -`hub-http-lookup-path-has-no-concurrency-bound.md`; the false rationale in the -test that certifies the path open is filed as -`hub-endpoints-test-certifies-the-ungated-clearnet-lookup-path-on-a-rationale-that-is-false-for-the-deployed-transport.md`; -the runbook omission is filed as -`hub-operators-runbook-endpoint-inventory-omits-the-unauthenticated-lookup-endpoint.md`. -Where the lookup *misses* the queue and is forwarded upstream, see -`hub-lookup-fall-through-hands-every-wallets-txid-to-whichever-indexer-the-hub-is-pointed-at-which-the-shipped-config-makes-the-operator.md`. - -DO NOT DEVIATE FROM THIS TEMPLATE (`docs/SECURITY-ISSUE-TEMPLATE.md`). -LEAVE THIS MESSAGE IN PLACE SO THAT ALL AGENTS KNOW NOT TO DEVIATE FROM THIS TEMPLATE (`docs/SECURITY-ISSUE-TEMPLATE.md`). diff --git a/zeronym-22aa9851caf68-high-medium/medium/indexer-chooses-which-batch-members-reach-the-chain-so-the-on-chain-batch-size-is-adversary-selected-regardless-of-k.md b/zeronym-22aa9851caf68-high-medium/medium/indexer-chooses-which-batch-members-reach-the-chain-so-the-on-chain-batch-size-is-adversary-selected-regardless-of-k.md deleted file mode 100644 index 575504cb..00000000 --- a/zeronym-22aa9851caf68-high-medium/medium/indexer-chooses-which-batch-members-reach-the-chain-so-the-on-chain-batch-size-is-adversary-selected-regardless-of-k.md +++ /dev/null @@ -1,364 +0,0 @@ -# The indexer decides which members of a flushed batch reach the chain, so the on-chain batch size is chosen by that party regardless of how large the hub's batch is - -**Severity**: Medium -**Validation Status**: Confirmed -**Location**: `audit-target/zeronym/hub/src/batcher.rs:341-390` (`flush`: drain, `broadcast_batch`, the achieved/rejected/requeue split), `:396-420` (the telemetry); `audit-target/zeronym/hub/src/chain.rs:176-210` (`broadcast`, `broadcast_batch`), `:269-337` (`unary`/`unary_inner`: one fresh connection and one 10 s budget per call), `:459-474` (`best_of`), `:491-501` (`classify_publish_failure`); `audit-target/zeronym/hub/src/queue.rs:279-295` (`requeue`), `:328-348` (`find_by_txid`); `audit-target/zeronym/hub/src/server.rs:296-303` (`Hub::lookup` answers a queued entry with the mempool sentinel). Deployed endpoint count: `audit-target/zeronym/deploy.env.example:22`. Claims contradicted: `audit-target/zeronym/hub/src/batcher.rs:8-22`, `audit-target/zeronym/README.md:27` and `:34`, `audit-target/zeronym/hub/REVIEW.md` #2 and #8. -**Found by agent:** Global (focus areas G2 "the flush clock as an attack surface" / G14 "admission→flush→publish as one system" / G26 "isolate a target's migration into a batch of one") -**In scope of audit?** Yes - -## Description - -Every defence in the hub's batching design protects the **decision** to publish. -`batcher.rs:8-22` and `hub/REVIEW.md` #2 and #8 are entirely about making the -moment of the flush unreachable by an attacker: no count trigger, no expiry -trigger, no stale-tip trigger, max-over-nodes tip acquisition, admission control -instead of an early-expiry flush. - -Nothing protects the **act** of publishing. Between `flush()` deciding to publish -and the transactions entering a mempool sits the indexer named in `ZIH_INDEXERS`, -and that party's **per-transaction** answer decides, individually and -independently for each batch member, whether that member reaches the network at -this flush or is put back in the queue for a later one. - -The mechanism is `Publish::Retryable`. `flush` splits the batch three ways -(`batcher.rs:365-378`): `Accepted`/`AlreadyKnown` are dropped as delivered, -`Rejected` is dropped as destroyed, and **`Retryable` is requeued and offered -again at the next flush**. `classify_publish_failure` (`chain.rs:491-501`) maps -every gRPC status except `INVALID_ARGUMENT` (3) and `FAILED_PRECONDITION` (9) to -`Retryable` — including `UNAVAILABLE` (14), the status a healthy, -correctly-implemented gRPC service returns when momentarily overloaded, and -therefore the least suspicious answer available. - -So an indexer that wants exactly one transaction from a batch of `k` to appear on -chain answers `SendResponse { error_code: 0, error_message: "" }` for that -one and relays it, and gRPC `UNAVAILABLE` for the other `k-1` and relays nothing. -The hub publishes a batch of `k`; the chain receives a batch of **1**. The other -`k-1` come back at the next flush, where the choice can be repeated. - -This is the isolation attack `REVIEW.md` #2 was written to prevent -("Count-based flushing lets an attacker submit 99 of their own migrations to -isolate a target's", `batcher.rs:9-11`), reached by a route neither #2 nor #8 -considers, because both reason about the queue and the clock and neither reasons -about the hop after them. - -**It is adoption-proof.** `README.md:34` states the residual as "the modal batch -is zero or one … The lever is adoption, not code." Raising adoption raises `k`; -this sets the on-chain batch to 1 for any `k`. - -## Attack Scenario and Steps - -Attacker: the operator of the indexer the hub broadcasts through, or anyone who -compromises or compels it. The audit's threat model names the capability -explicitly — *"hub → indexer: sees the whole batch seconds before it is public; -can lie about the tip and **about publish verdicts**"*. With the shipped -`INDEXERS=66.241.124.200:443` (`deploy.env.example:22`) there is exactly one such -party and `best_of` has one verdict to fold. - -1. The cadence fires. `flush` drains the queue and calls - `chain.broadcast_batch(&payloads)` (`batcher.rs:342`, `:355`). -2. `broadcast_batch` issues one `broadcast` per transaction concurrently - (`chain.rs:208-210`), and each `broadcast` issues one `unary` per endpoint - (`chain.rs:181-197`). `unary_inner` dials a **fresh TCP + TLS + h2 connection - per call** (`chain.rs:300-334`), so the indexer receives `k` separate, - simultaneous `SendTransaction` requests, each carrying one raw transaction in - the clear. -3. Each call has a 10 s budget (`chain.rs:48`, `:280-291`). The indexer therefore - holds the entire batch, with up to ten seconds to decide, before it has to - answer any of them. It selects a victim by content — value balance, action - count, length, `anchorOrchard`, `nExpiryHeight` — or by a txid supplied out of - band. -4. It relays the victim's transaction to its node and answers - `SendResponse { error_code: 0, error_message: "" }`. - `classify_send_response` (`chain.rs:438-445`) makes that `Publish::Accepted`, - and `flush` counts it achieved and drops the entry. -5. For every other member it answers a trailers-only `grpc-status: 14` and relays - nothing. `round_trip` turns that into `GrpcStatusError { code: "14" }` - (`chain.rs:391-401`), `classify_publish_failure` makes it `Publish::Retryable` - (`chain.rs:491-501`), and `flush` pushes the entry into `unplaced` and calls - `queue.requeue` (`batcher.rs:369-372`, `:390`). -6. On chain, at the height of this flush, **exactly one** Orchard-touching - transaction appears from this hub. Its anonymity set is itself, and it is - visible as such to **anyone watching the chain or the mempool**, not only to - the indexer performing the selection. -7. At the next flush the attacker repeats with a different victim — publishing - every migration alone, forever — or releases the remainder. - -**Nothing anywhere observes the difference.** Three detection channels all fail: - -- **The hub's own telemetry.** `batcher.rs:396-410` logs - `flush_size = k, achieved_batch_size = 1, rejected = 0, requeued = k-1` plus a - `warn!` whose `reason` field is *a string the attacker wrote*. That is - byte-for-byte what a genuine ten-second indexer hiccup looks like, and the - project's own test `a_transport_flavoured_grpc_status_is_held_but_invalid_argument_is_not` - (`batcher.rs:763-775`) pins `UNAVAILABLE → hold` as correct behaviour. The one - log line that names the harm — `"batch provides no batching anonymity at this - size"` (`batcher.rs:412-420`) — is *expected to fire today* at current adoption, - so it is pre-normalised as noise. -- **The wallet's confirmation lookup.** The shim routes every `GetTransaction` to - the hub, and `Hub::lookup` answers a queued entry from `Queue::find_by_txid` - with `height: 0` — lightwalletd's mempool sentinel (`server.rs:296-303`, - `server.rs:90-91`, `queue.rs:328-348`). A held-back migration therefore answers - the wallet *"pending in the mempool"*, indefinitely and indistinguishably from a - genuinely unmined transaction. -- **The hub itself.** There is no confirmation tracking (`REVIEW.md` #7 is a - documented designed-not-built item), and `Entry::received_height` — declared at - `queue.rs:135` as *"Drives the confirmation deadline"* — is written once - (`queue.rs:235`) and **read nowhere in the crate**. - -**Attack Requirements and Assumptions:** -- **Control of, compromise of, or compulsion over the indexer(s) the hub - publishes through.** This is not an internet-reachable attack and not available - to a stranger: it requires the configured, semi-trusted endpoint to misbehave. -- Cost: zero. It is a status code per request. -- Reliability: **deterministic** while `node_count() == 1`, which is the shipped - configuration. -- **What makes it not work:** with two or more genuinely independent endpoints, an - honest one answers `Accepted` and relays the transaction, and `best_of` ranks - `Accepted` (3) above `Retryable` (0) (`chain.rs:459-474`), so the hold-back - fails. This attack needs **all** endpoints hostile. Multi-endpoint deployment is - a real, code-supported mitigation here — and it is the *opposite* of the - direction the `max()`-based tip rule scales (see - `tip-and-verdict-aggregation-scale-in-opposite-directions-so-adding-indexers-fixes-one-lever-and-aggravates-three.md`). - -## Impact on Users - -The batch is the entire anonymity mechanism the hub provides; `queue.rs:1-6` says so -("The batch IS the anonymity set"). This reduces the on-chain batch to one -transaction at will, for a victim of the attacker's choosing, while every -component reports normal operation and the wallet is shown "pending". - -The reason this matters despite the operator already holding stronger levers is -specific, and the report should carry it in this frame rather than as additional -independent linkage: - -1. **It survives the fix for the shape-based selection channel.** The filed - core-linkage chain selects within a batch by `(length, anchor, expiry)`, which - works *because* diverted transactions differ in shape. ZIP 318 conformance plus - wallet-side padding makes them uniform and closes that channel; this one uses no - shape at all and is untouched. The two are anti-correlated in time: the shape - channel dominates today, this one dominates the moment the wallet-side condition - the project is designing toward is met. -2. **It reaches adversaries the operator-side attack cannot.** Publishing the - victim **alone on the public chain** hands the result to a mempool watcher, a - chain-analysis firm, or a party who can compel the indexer but not the enclave — - all from public data. Anyone additionally holding the "IP C submitted an - Orchard-touching transaction at time T" half (which `README.md:33` and - `hub/REVIEW.md:181` concede the shim's operator holds) completes IP → - transaction → balance. That is the linkage `README.md:27` says does not survive, - "volume-independent". - -Two secondary harms follow from the same mechanism: - -- **Destruction of held-back transactions in the traffic class the hub is sized - for.** A librustzcash-default wallet transaction carries `expiry = build_height + - 40` (`batcher.rs:49-55`). It was admitted only because it survived - `next_flush_height(tip, 20) + 4` (`queue.rs:380-392`). Holding it back one full - 20-block interval pushes publication to the edge of or past its expiry, after - which the node refuses it, `flush` classifies it `Rejected` and **drops it - permanently** (`batcher.rs:368`) — for a migration the wallet was told at mixnet - hand-off had been sent (`shim/src/hub.rs:238-240`). ZIP 318 migrations, with - 34,561–69,120 blocks of slack, survive indefinitely, which is what makes the - *isolation* variant sustainable against exactly the population the product exists - for: the attacker can serialise a whole batch one transaction per flush with - nothing expiring. -- **Self-escalation into a fleet-wide admission outage.** Requeued entries are - recharged to the byte budget but `requeue` never checks it (`queue.rs:279-295`, - filed as `hub-queue-requeue-ignores-byte-budget-unbounded-growth.md`), so - sustained hold-back grows the queue monotonically past `MAX_QUEUE_BYTES` - (`queue.rs:65`), after which `admit` refuses every new submission `Full` - (`queue.rs:224`) — which, on the deployed dispatch-only mixnet transport, is - silent (`nym-submit-acks-are-never-read-so-every-hub-refusal-is-invisible.md`). - -## Technical Details / Code Analysis - -The verdict split, in full (`hub/src/batcher.rs:361-390`): - -```rust - let mut achieved = 0usize; - let mut rejected = 0usize; - let mut sample_failure: Option = None; - let mut unplaced = Vec::new(); - for (i, entry) in batch.into_iter().enumerate() { - match outcomes.get(i) { - Some(Publish::Accepted { .. }) | Some(Publish::AlreadyKnown) => achieved += 1, - Some(Publish::Rejected { .. }) => rejected += 1, - Some(Publish::Retryable { reason }) => { - sample_failure.get_or_insert_with(|| reason.clone()); - unplaced.push(entry); // <-- held back, offered again next flush - } - None => unplaced.push(entry), - } - } - ... - let requeued = queue.requeue(unplaced); -``` - -The verdicts are **positional and per transaction** — `broadcast_batch` returns one -`Publish` per input (`hub/src/chain.rs:208-210`): - -```rust - pub async fn broadcast_batch(&self, txs: &[Vec]) -> Vec { - join_all(txs.iter().map(|tx| self.broadcast(tx))).await - } -``` - -and each `broadcast` is a separate gRPC call over a separate connection carrying -one transaction (`hub/src/chain.rs:176-198`, `:300-334`), so the indexer has a -distinct, individually answerable request per batch member and can discriminate on -the plaintext it is holding. - -**The crux — the status classification that makes `UNAVAILABLE` a hold rather than -a verdict** (`hub/src/chain.rs:491-501`, with the constants at `:120-121`): - -```rust -const GRPC_INVALID_ARGUMENT: &str = "3"; -const GRPC_FAILED_PRECONDITION: &str = "9"; -... -fn classify_publish_failure(err: &BoxError) -> Publish { - let reason = err.to_string(); - match err.downcast_ref::() { - Some(status) - if status.code == GRPC_INVALID_ARGUMENT || status.code == GRPC_FAILED_PRECONDITION => - { - Publish::Rejected { reason } - } - _ => Publish::Retryable { reason }, - } -} -``` - -`chain.rs:485-490` explains *why* everything else is retryable, and the reasoning -is sound in the direction it considers — dropping a valid migration on a misread -error is unrecoverable. The direction it does not consider is that the same rule -hands the party writing the status code a per-transaction publication switch. - -`round_trip` produces that typed error from either the trailers or a trailers-only -HEADERS frame (`hub/src/chain.rs:360-401`), so a single `grpc-status: 14` header on -an otherwise well-formed HTTP/2 response is enough; no body is needed. - -With one endpoint, `best_of` (`hub/src/chain.rs:459-474`) is the identity function -on that endpoint's answer, so there is no second opinion to override it. - -The lookup path that hides the hold-back from the wallet -(`hub/src/server.rs:296-303`): - -```rust - pub async fn lookup(&self, wire_hash: &[u8]) -> LookupOutcome { - if let Some(bytes) = self.queue.find_by_txid(wire_hash) { - tracing::debug!(source = "queue", "transaction lookup answered"); - return LookupOutcome::Found { data: bytes, height: 0 }; - } -``` - -`height: 0` is documented at `server.rs:90-91` as *"`0` means mempool (a queued, -unflushed transaction), matching lightwalletd's sentinel"*. A wallet polling for -its migration therefore receives "in the mempool" for a transaction that is in no -mempool at all. - -And the module claim this refutes (`hub/src/batcher.rs:8-22`): - -> **Why the cadence is unconditional (REVIEW #2, #8).** Every conditional trigger -> is a lever someone else can pull. […] A deterministic clock nobody can influence -> is the only shape with no lever on it […] - -The clock is only half the mechanism. A deterministic clock firing into a -publication path a single untrusted party gates delivers a deterministic -*schedule* and an adversary-selected *content*. - -## Recommendations - -- **Alarm on the shape that has no benign explanation.** For a single endpoint - answering `k` simultaneous requests, a flush with `achieved >= 1 && requeued >= 1` - is either impossible (a real outage yields `achieved == 0`) or a selective - hold-back. That is a one-line check in `flush` and is the cheapest fix in the hub. -- **Require, and validate, at least two operationally independent indexers before - the hub claims batching anonymity.** `best_of`'s ranking already defeats this at - `n >= 2` with one honest endpoint; the shipped `n = 1` is what makes it - deterministic. Note this is the opposite of what the `max()` tip rule wants, so - the two must be fixed together. -- **Implement the confirmation tracking `REVIEW.md` #7 specifies, or stop claiming - what it would prove.** Re-querying the chain for a published txid one or two - cadences later is the only mechanism that distinguishes "the indexer relayed it" - from "the indexer said it did". Until then, `batcher.rs:333-340`'s description of - `achieved_batch_size` as "the honest measure of the privacy the flush actually - delivered" should say it measures what the indexer reported. `Entry::received_height` - already exists and is unread; it is the natural place to hang this. -- **Distinguish "queued, never offered" from "offered and held back" in the lookup - answer.** One boolean on `Entry` is the difference between a wallet that can - eventually surface the failure and one that cannot. -- **State the residual.** `REVIEW.md`'s inherent-limits section records that a party - degrading the *shim→hub* path chooses when a migration is published. It does not - record that the party on the *hub→indexer* path chooses **whether a migration is - published in a batch at all**, which is a strictly stronger capability over the - same anonymity property. - -## Validation Information - -**Verdict: CONFIRMED at Medium** (downgraded from the filed High). - -**The crux was verified directly, as required:** - -- `hub/src/chain.rs:120-121` defines `GRPC_INVALID_ARGUMENT = "3"` and - `GRPC_FAILED_PRECONDITION = "9"`; `classify_publish_failure` at `:491-501` - returns `Publish::Rejected` **only** for those two codes and - `Publish::Retryable` for everything else, including `UNAVAILABLE` (14). The - doc comment at `:485-490` states this intent explicitly and names `UNAVAILABLE`. -- `round_trip` (`chain.rs:360-401`) reads `grpc-status` from trailers **or** - headers and returns the typed `GrpcStatusError`, so a trailers-only `14` reply - is sufficient — no body required. -- `flush` (`batcher.rs:361-390`) requeues exactly `Retryable` and `None`; the - requeued entries return at the next cadence (`queue.rs:279-295`). -- `broadcast_batch` → `broadcast` → `unary` → `unary_inner` issues one - `TcpStream::connect` per (transaction × endpoint) with a single 10 s - `tokio::time::timeout` around the whole call (`chain.rs:208-210`, `:181-197`, - `:280-291`, `:300-311`). The indexer therefore sees `k` independent, concurrent, - individually answerable requests and has the whole budget to decide. -- `best_of` (`chain.rs:459-474`) ranks `Accepted` 3 > `AlreadyKnown` 2 > - `Rejected` 1 > `Retryable` 0, so with `n = 1` it is the identity and with an - honest endpoint present the hold-back fails — the filed "what makes it not work" - bound is correct. -- The project's own test `a_transport_flavoured_grpc_status_is_held_but_invalid_argument_is_not` - (`batcher.rs:763-775`) asserts `Status("14")` leaves the entry in the queue and - `Status("3")` removes it, pinning the behaviour as intended. -- Detection claims checked: `batcher.rs:396-420` logs only aggregates plus the - attacker-authored `reason` string; `Entry::received_height` appears only at - `queue.rs:135` (declaration) and `:235` (write) and is read nowhere in `hub/src` - or `hub/tests`; `Hub::lookup` answers a queued entry `height: 0` - (`server.rs:296-303`). All three detection failures are as filed. -- `deploy.env.example:22` ships a single endpoint, so `node_count() == 1` in the - shipped configuration. - -**Why Medium and not High.** The coordinator's standing bound for indexer-dependent -findings applies here (PROGRESS.md item 6p: *"the `Rejected` drop and all tip -manipulation require control of a configured indexer endpoint — they are -hub-trust/robustness defects, not internet-reachable weapons"*), and three further -considerations cap the severity: - -1. **The attacker is a configured, semi-trusted endpoint, not a stranger.** No - internet-reachable path exists to this behaviour; it needs the indexer the hub - publishes through to be hostile, compromised or compelled. -2. **Severities must not be stacked (PROGRESS.md item 6v).** Against the *shim's - operator*, forcing `k = 1` adds little today: the already-filed core-linkage - chain lets them select a target *within* a batch by `(length, anchor, expiry)`, - and `hub/REVIEW.md:175` concedes the modal batch is already 0 or 1 at current - adoption. The incremental harm to users **today** is near zero. -3. **A code-supported mitigation exists and is one configuration line:** two or - more independent endpoints defeat it via `best_of`. - -**Why it is nonetheless a real finding and not invalid or Info.** It earns its place -on the two grounds the global pass identified and this validation upholds: it is the -*successor* risk — it survives the wallet-side ZIP 318 + padding fix that closes the -shape-based channel, because it uses no shape — and it reaches adversaries that -channel cannot, because the victim lands **alone on the public chain**, so a mempool -watcher or chain-analysis firm gets the result from public data without ever -touching the enclave or the wallet leg. It also destroys wallet-acknowledged -migrations in the expiry-bounded traffic class as a side effect, with no signal to -the wallet, the operator or the hub. - -**Nothing in the filed issue was found to be factually wrong.** The changes made are -framing only: the "adoption-proof isolation destroys anonymity" headline is now -stated together with the honest bound that today's batch is already 1, the attacker -must be the configured indexer, and `n >= 2` fixes it. - -DO NOT DEVIATE FROM THIS TEMPLATE (`docs/SECURITY-ISSUE-TEMPLATE.md`). -LEAVE THIS MESSAGE IN PLACE SO THAT ALL AGENTS KNOW NOT TO DEVIATE FROM THIS TEMPLATE (`docs/SECURITY-ISSUE-TEMPLATE.md`). diff --git a/zeronym-22aa9851caf68-high-medium/medium/nym-submit-acks-are-never-read-so-every-hub-refusal-is-invisible.md b/zeronym-22aa9851caf68-high-medium/medium/nym-submit-acks-are-never-read-so-every-hub-refusal-is-invisible.md deleted file mode 100644 index 8f7bb3e6..00000000 --- a/zeronym-22aa9851caf68-high-medium/medium/nym-submit-acks-are-never-read-so-every-hub-refusal-is-invisible.md +++ /dev/null @@ -1,446 +0,0 @@ -# Every hub refusal ack is decoded and then discarded, so on the deployed mixnet transport nothing in the shim can distinguish a hub that refuses 100% of migrations from a healthy one - -**Severity**: Medium -**Validation Status**: Confirmed -**Location**: `audit-target/zeronym/shim/src/nym.rs:595-690` (`NymHandle::submit`), `:652` (the dropped receiver), `:835-911` (`correlate`), `:877`, `:903`, `:1026-1062` (`deliver`), `:1031-1033`, `:1038`; payload definition at `audit-target/zeronym/shim/src/wire.rs:203-267`; sole consumer at `audit-target/zeronym/shim/src/hub.rs:228-250`; status surface at `audit-target/zeronym/shim/src/nym.rs:133-177`, `:277-296`; hub side at `audit-target/zeronym/hub/src/nym.rs:313-334` and `audit-target/zeronym/hub/src/server.rs:248-277` -**Found by agent:** Local (file audit of `shim/src/nym.rs`) -**In scope of audit?** Yes — priority area #4 (mixnet transport), #5 (fail-closed discipline), #6 (log/telemetry discipline) - -## Description - -The shim→hub mixnet protocol carries a typed acknowledgement frame, `AckV1`, whose -whole purpose is to tell the shim what the hub did with a diverted migration. It -can say `Accepted`, or `Refused` with one of five reasons — `ExpiryTooTight`, -`TooLarge`, `QueueFull`, `TipStale`, `BadFrame` (`shim/src/wire.rs:217-229`). The -type's own doc comment states the contract (`shim/src/wire.rs:209`): - -> `/// The hub declined the submission. Every refusal fails closed at the shim.` - -**No refusal fails closed at the shim, because no refusal is ever observed by -anything.** `NymHandle::submit` creates the ack waiter and drops its receiving -end in the same statement (`nym.rs:652`), so the verdict is discarded whichever -arm of `deliver` fires — `let _ = waiter.send(kind)` into a dead receiver -(`nym.rs:1031-1033`) if the waiter is still in the map, or `None => {}` -(`nym.rs:1038`) once the correlator's sweep has removed it (`nym.rs:903`). Both -arms are silent. `AckKind` and `AckRefusal` are referenced nowhere in -`shim/src/` outside `wire.rs` and the `Waiter::Ack` variant itself. - -Not blocking the wallet on the round trip is a documented and defensible design -(`nym.rs:563-594`), and this issue does **not** dispute it. The defect is -narrower and is not a consequence of that design: **recording the verdict does -not require blocking the wallet**, and nothing records it. A hub whose queue is -full, whose tip has gone stale, or which is answering `bad_frame` to every frame -it receives produces exactly the same observable behaviour at every deployed shim -as a hub that is accepting and publishing everything: - -* no `tracing` line (both discard arms are silent), -* no counter — `MixnetStatusInner` (`nym.rs:133-177`) has `deaths`, - `consecutive_failures`, `replies_received`, `empty_inbound`, - `sends_dispatched`, `last_reply_unix` and no refusal counter of any kind, -* `/healthz` still answers 200 (`proxy.rs:649` reads only - `MixnetStatus::is_healthy`, which is client connectivity — `nym.rs:307-309`), -* `/nym-status` still reports `mixnet_connected: true`, -* `/nym-diag`, when opened, reports `replies_received` climbing — it is - incremented for *any* non-empty inbound message (`nym.rs:242`), so it rises - identically whether every ack says `Accepted` or every ack says `QueueFull`. - -The system pays the full cost of the feedback channel and then throws the signal -away: 13 reply SURBs are minted and emitted per submit **per configured hub -address** (`nym.rs:96`, `:656`), and the hub spends one of its strictly -serialised outbound send slots building and emitting each ack -(`hub/src/nym.rs:313-334`). - -## Attack Scenario and Steps - -This is a detectability defect. It creates no new attacker capability by itself; -its severity comes from the fact that it is the reason the destruction attacks -this audit has already confirmed can be sustained indefinitely without anyone -being in a position to notice. - -1. An attacker fills the hub's queue to `MAX_QUEUE_BYTES` (64 MiB) over the - unauthenticated, no-ACL Nym ingress — the confirmed High - `issues/confirmed/hub-queue-unauthenticated-fill-silently-destroys-migrations.md`. - From that moment `Hub::admit` returns `Refusal::Full` for every genuine - migration (`hub/src/server.rs:256-276`). -2. A wallet behind any shim sends an Orchard-touching `SendTransaction`. The shim - classifies it, diverts it, dispatches the `SubmitV1` frame and immediately - answers the wallet `SendResponse { error_code: 0, error_message: }` (`intercept.rs:180-203`). -3. The hub refuses it, logs it correctly on its own side, and sends back - `AckV1(Refused(QueueFull))` over one of the 13 SURBs the shim attached. -4. The shim reassembles it, `wire::decode_ack` parses it correctly, and `deliver` - discards the decoded verdict without a log or a counter. -5. The migration is gone. The wallet believes it was sent. **Every readable - surface on the shim is green.** -6. The attacker can sustain this for as long as they like. Nothing raises an - alarm, because the mechanism that would raise one is unwired. - -The same invisibility applies with no attacker at all: `TipStale` (the hub's -indexer is down or lagging), `ExpiryTooTight`, and `BadFrame` (a wire-format skew -between a shim and a hub built from different commits). - -**A second-order effect worth stating: the adversary has better visibility into -the hub's admission state than the honest operator does.** Anyone can run a Nym -client, submit a frame with their own reply SURBs, and *read* the returned -refusal code — that is filed separately as -`hub-ackv1-refusal-codes-are-an-anonymous-real-time-admission-state-oracle.md`. -The one party structurally unable to read it is the shim whose users' migrations -are being destroyed. - -**Attack Requirements and Assumptions:** - -- Causing the *invisibility* requires no access at all; it is unconditional and - affects the shipped transport (`deploy.env.example` selects `HUB_NYM`). -- Causing the *refusals* requires only the ability to send Nym frames to the - hub's published address, which is public by design and has no submitter ACL. -- This issue does not claim the wallet could be told. Under dispatch-only the - wallet has already been answered before the ack exists. The claim is that the - **operator and the project** have no signal, and that a signal is available for - free on a channel already paid for. - -## Impact on Users - -A user's mandatory Orchard→Ironwood migration is destroyed after their wallet was -told it succeeded, and no party — the user, the shim operator, or Shielded Labs — -is in a position to learn that a class of migrations is being destroyed at all. -The user's own recovery path is to notice, unaided, that a transaction they were -told succeeded never confirmed, and to wait out the expiry — which for a ZIP 318 -migration is 30 to 60 days (owned by -`zip318-canonical-expiry-is-the-only-recovery-clock-…`). - -> **CORRECTION 2026-08-18 (validation of the cited file — SUPERSEDES the sentence above).** -> The wallet does **not** wait for expiry. Both official Zcash light-wallet SDKs -> automatically resubmit a sent-but-unmined transaction for as long as it remains -> unexpired — the Android SDK at the head of every ~20 s sync loop and after every -> processed block batch (`CompactBlockProcessor.kt:573,615,723`; selection -> `mined_height IS NULL AND expiry_height > ?`), the iOS SDK at most once per 300 s -> (`TxResubmitter.swift:8-15`, `TransactionDao.swift:218-228`) — and the hub's -> payload-hash dedup makes the resend free. The wallet's non-confirmation signal comes -> from compact-block scanning, which the shim does not intercept (`proxy.rs:1068-1074`). -> Expiry is therefore the **retry horizon**, not the wait: ~50 minutes for the ZIP 203 -> default traffic the shim also diverts, 30–60 days for a ZIP 318 migration. A -> *transient* loss self-heals within minutes; only a loss condition that **outlives the -> horizon** destroys the submission permanently — which is exactly what this issue's -> condition does, so this issue's severity is unaffected. Do not write "the user waits -> 30 to 60 days" in the report. Full refutation and the replacement paragraph: -> `issues/invalid/zip318-canonical-expiry-is-the-only-recovery-clock-and-a-lost-migration-freezes-the-users-notes-for-30-to-60-days.md`. - - -The absence of the signal is what converts a bounded outage into an unbounded -one. Every other loss path in the system has the same property, which is why -`globals/G7-loss-of-a-wallet-acknowledged-migration.md` §5(b) concludes that -**reading the ack is the single change that converts almost every row of the loss -taxonomy from invisible to observable**, at no cost in wallet latency and no new -information for the parent host (the five reason strings are fixed, carry no -per-entry data, and `AckRefusal::as_str` was written and documented as -*"safe to log"*, `wire.rs:257-266`). - -## Technical Details / Code Analysis - -**1. The receiver is dropped at construction.** `shim/src/nym.rs:642-661`: - -```rust - for target in 0..targets { - // A FRESH nonce per address: two hubs answering the same nonce would be - // indistinguishable to the correlator, and the ack is unread anyway. - let nonce = fresh_nonce(); - ... - let frame = wire::encode_submit(&nonce, tx_bytes).map_err(NymError::Encode)?; - let (ack_tx, _drop_receiver) = oneshot::channel(); - let request = Request { - nonce, - frame, - reply_surbs: SUBMIT_REPLY_SURBS, - waiter: Waiter::Ack(ack_tx), - target, - }; -``` - -`_drop_receiver` is a real binding (not `_`), so it lives to the end of the loop -iteration and is dropped as soon as `send()` returns — before `submit` returns to -its caller, and long before the frame reaches the mixnet. - -**2. The verdict is discarded on both arms, so timing is immaterial.** -`shim/src/nym.rs:1026-1044`: - -```rust -fn deliver(pending: &mut HashMap, bytes: &[u8]) { - match bytes.len() { - wire::ACK_BYTES => match wire::decode_ack(bytes) { - Ok((nonce, kind)) => match pending.remove(&nonce) { - Some(Waiter::Ack(waiter)) => { - let _ = waiter.send(kind); - } - Some(other) => { - pending.insert(nonce, other); - tracing::warn!("an ack arrived for a lookup's nonce; ignoring it"); - } - None => {} - }, - Err(err) => { - tracing::warn!(reason = %err, "inbound message could not be decoded as an ack") - } - }, -``` - -`kind` is the fully decoded `AckKind::Refused(AckRefusal::QueueFull)`. If the -waiter is still in `pending`, `let _ = waiter.send(kind)` sends it into a receiver -that no longer exists and the `Err` is discarded by the `let _`. If the sweep has -already removed it, `None => {}` fires. **Neither logs and neither counts.** - -`deliver` has exactly three non-logging arms — the ack-delivery arm above and the -two `None => {}` misses (`:1038`, `:1055`). Every arm that *does* log is a decode -failure or a kind mismatch. Stated as a property: **`deliver` speaks only when it -cannot understand a reply, and is silent exactly when it understood one and threw -it away.** - -**3. In practice it is the `None` arm that fires, because the sweep runs first.** -`shim/src/nym.rs:869-905`: - -```rust - request = requests.recv(), if permit.is_some() && requests_open => match request { - Some(Request { nonce, frame, reply_surbs, waiter, target }) => { - permit - .take() - .expect("the arm is guarded on holding a permit") - .send(OutFrame { frame, reply_surbs, target }); - pending.insert(nonce, waiter); - } - None => requests_open = false, - }, - ... - } - // Callers that timed out (or were cancelled) have dropped their - // receivers; ... - pending.retain(|_, waiter| !waiter.is_abandoned()); -``` - -`pending.retain` runs unconditionally after **every** turn of the select, and -`Waiter::is_abandoned` is `tx.is_closed()` (`nym.rs:320-329`), which is already -true. So the waiter is inserted at `:877` and removed at `:903` on the same turn, -or on the next one at the latest (`SWEEP_INTERVAL = 1 s`, `nym.rs:501`). The ack -cannot arrive inside that window: by the crate's own throughput arithmetic -(`nym.rs:1104-1115`) a single `SubmitV1` is `packets(FRAME_BYTES) + -SUBMIT_REPLY_SURBS` = 32 + 13 = 45 Sphinx packets at -`THROTTLED_PACKETS_PER_SEC ≈ 8.33`, i.e. **≈5.4 s of pure outbound emission -before any mix delay**, and that is only the outbound leg. - -**4. Nothing else consumes an `AckKind`.** A grep over `shim/src/` finds -`AckKind`/`AckRefusal` only in `wire.rs` (codec + unit tests), `nym.rs:316` (the -`Waiter::Ack` variant), and `nym.rs:1169-1290` (tests). `hub.rs:228-250` — the -only caller of `NymHandle::submit` — says so itself: - -```rust - HubTransport::Nym(handle) => match handle.submit(tx_bytes).await { - // ... There is no Refused arm: the hub's verdict is a full round trip - // away and is deliberately not waited for (see `NymHandle::submit`), so - // a refusal is never surfaced here. - Ok(()) => Ok(Submit::Accepted { - txid: crate::nym::local_txid(tx_bytes), - }), -``` - -**5. The hub does its half correctly, and writes it where nobody can read it.** -`hub/src/nym.rs:313-334` builds the ack from the real admission verdict, and -`hub/src/server.rs:273` logs `info!(reason = refusal.as_str(), "submission -refused at admission")`. So it is not true that *nothing* in the system records a -refusal — the hub does. What is true, and is the point, is that (a) **nothing on -the shim side records anything**, and (b) the hub's record goes to `tracing`, -which in a correctly attested enclave (`debug { enabled = false }`) reaches no -console at all — see `globals/G7-…` §5(a) and the separately filed -`hub-health-surface-blind-to-the-states-that-destroy-migrations.md`. The two -readable surfaces the hub exposes, `/healthz` and `/nym-status`, are blind to -`Full` and `TipStale`. - -**6. The dead channel inflates the cost of every submit by 40%, on the exact -budget the confirmed flood attack exhausts.** `nym.rs:88-96` already concedes the -waste: - -```rust -/// NOTE (dispatch-only submit): the shim no longer awaits the ack, so most of -/// these SURBs are now unused send-path overhead — the hub spends them replying -/// into a dropped receiver. ... Trimming it toward the -/// anonymity minimum is a throughput follow-up, ... and low priority, since submits -/// are rare (a migration is ~0.77 per block) next to the continuous cover traffic. -``` - -13 of the 45 packets a submit costs are SURBs for a reply nobody reads. The -stated reason for deprioritising the trim — *"submits are rare"* — is exactly the -premise that the confirmed High -`junk-sendtransaction-flood-consumes-the-shims-whole-mixnet-egress-…` destroys: -an unauthenticated stranger makes submits arbitrarily frequent for about a byte a -second. Trimming the SURB count would reduce that attack's per-frame amplification -by ~29% (45 → 33 packets); it does **not** close it, and it must not be done by -setting the count to zero (see Recommendations). - -## Recommendations - -Keep dispatch-only. Stop discarding the verdict. - -1. **Do not drop the ack receiver in `NymHandle::submit` (`nym.rs:652`).** Either - hold it in a detached task that records the outcome, or replace the - `Waiter::Ack` oneshot with a sink owned by the transport so `deliver` can - record every ack it decodes. -2. **This is the load-bearing fix: add refusal counters to `MixnetStatusInner`, - keyed by `AckRefusal::as_str()`** (five fixed strings, no per-entry data), and - publish them on `/nym-status`. A non-zero `queue_full` or `tip_stale` count is - precisely the "this shim's migrations are being destroyed" signal that has no - representation anywhere today. Do this even if nothing else on this list is - done. -3. A rate-limited `tracing::warn!(reason = refusal.as_str(), …)` is worth adding - for the debug/non-attested case, but **on its own it fixes nothing in the - deployment that matters**: under `debug { enabled = false }` the enclave has no - console, which is why item 2 and not this one is the fix. -4. At minimum, count the `None` arm of `deliver` (`:1038`) so "acks are arriving - and matching nothing" is distinguishable from "no acks are arriving". -5. **Do NOT delete the ack from the `SubmitV1` flow, and do NOT set - `SUBMIT_REPLY_SURBS` to zero.** An earlier draft of this issue recommended - exactly that; it is wrong and would cause harm on two counts. (a) The ack's - inbound packet increments `replies_received` (`nym.rs:242`), which feeds - `inbound_total` (`nym.rs:251-253`), which is the **sole** input to the driver's - liveness probe (`nym_driver.rs:285`, `:312`); `globals/G21-…` records that this - backflow is exactly why the shim is not exposed to the hub's confirmed liveness - fleet-kill. (b) `nym.rs:93-95` states that a zero SURB count would push the - driver off the anonymous-send path (M6/D3), which is an anonymity regression. - The correct change is to *trim* the count toward the measured minimum while - keeping it non-zero, and to *read* what comes back. -6. Fix `shim/src/wire.rs:209`, which currently asserts *"Every refusal fails - closed at the shim"* — false on the deployed transport. (Owned by - `wire-ack-refusal-documented-as-fail-closed-is-discarded-on-the-mixnet-path.md`; - listed here only so the recommendations are complete.) - -**Cross-references (distinct findings; do not merge):** -`hub-queue-unauthenticated-fill-silently-destroys-migrations.md` (confirmed High — -a way to make the hub refuse), -`junk-sendtransaction-flood-…md` (confirmed High — the shim-side denial this -blindness compounds), `hub-tip-overshoot-latches-hub-permanently-stale.md`, -`hub-ackv1-refusal-codes-are-an-anonymous-real-time-admission-state-oracle.md` -(the same codes, readable by an attacker), -`readme-says-a-failed-migration-can-fail-silently-…md` (confirmed Low — the -documentation side), `divert-nym-hub-refusal-test-is-vacuous-and-the-only-wallet-facing-refusal-arm-is-untested.md`, -`shim-hub-submit-verdict-type-means-different-things-on-the-two-transports.md`, -`hub-health-surface-blind-to-the-states-that-destroy-migrations.md`, and -`globals/G7-loss-of-a-wallet-acknowledged-migration.md` (the taxonomy). - -## Validation Information - -**Validated 2026-08-18. Verdict: CONFIRMED. Severity held at Medium.** - -### Mechanics — every claim re-derived from the target, not inherited - -| claim | verified | -|---|---| -| the ack receiver is dropped at construction | yes — `shim/src/nym.rs:652`, `let (ack_tx, _drop_receiver) = oneshot::channel();`. `_drop_receiver` is a binding, not `_`, so it is dropped at the end of the loop iteration, immediately after `send()` returns | -| `is_abandoned` is `tx.is_closed()` and is already true when the waiter is inserted | yes — `nym.rs:320-329` | -| `pending.retain` runs after **every** select turn, not only on the sweep tick | yes — `nym.rs:903`, outside the `select!`, unconditional | -| the sweep therefore removes the waiter on the same turn (worst case one turn / 1 s later) | yes — `SWEEP_INTERVAL = 1 s`, `nym.rs:501` | -| a submit is 45 Sphinx packets ≈ 5.4 s of emission at the throttled rate | yes — `FRAME_BYTES = 65536` (`wire.rs:72`), `PACKET_BYTES = 2048`, `SUBMIT_REPLY_SURBS = 13` (`nym.rs:96`), `THROTTLED_PACKETS_PER_SEC = 1000/120` (`nym.rs:1093`). 32 + 13 = 45; 45 / 8.33 = 5.4 s outbound alone | -| so 100% of acks arrive after their waiter is gone | yes, and **the conclusion does not depend on it** — see the correction below | -| `AckKind`/`AckRefusal` appear nowhere in `shim/src/` outside `wire.rs` and the `Waiter::Ack` variant | yes — full grep of `shim/`; the only other hits are in `shim/tests/nym.rs` and `shim/tests/divert_nym.rs` | -| `HubTransport::submit` has no `Refused` arm on the Nym path | yes — `shim/src/hub.rs:228-250`, with the comment quoted | -| `MixnetStatusInner` has no refusal counter | yes — `nym.rs:133-177` enumerated in full | -| `replies_received` counts any non-empty inbound message | yes — `nym.rs:239-242` | -| `/healthz` reads only client connectivity | yes — `proxy.rs:649` → `nym.rs:307-309` (`!configured || connected`) | -| `wire.rs:209` claims every refusal fails closed at the shim | yes, verbatim | -| the hub emits a correct ack from the real verdict | yes — `hub/src/nym.rs:313-334` | - -### Four corrections applied to the filed text - -1. **"Nothing anywhere in the system records that migrations are being refused" - was too strong and has been struck.** The hub *does* record it — - `hub/src/server.rs:273`, `info!(reason = refusal.as_str(), "submission refused - at admission")`. The accurate claim, now in the file, is that nothing on the - **shim** side records anything, and that the hub's record goes to a `tracing` - console that does not exist under `debug { enabled = false }`. This matters: - a report sentence saying the refusal is recorded nowhere would be refutable in - one grep. -2. **"`nym.rs:1031-1033` is the only arm of `deliver` that does not log" is not - literally true and has been restated.** `deliver` has three silent arms: - the ack-delivery arm and the two `None => {}` misses (`:1038`, `:1055`). The - true and stronger statement, which the report should use instead, is that - *`deliver` logs exactly when it cannot understand a reply and is silent exactly - when it understood one and discarded it.* **Note for the report generator: the - confirmed sibling `readme-says-a-failed-migration-can-fail-silently-…md` says - "Every other arm of this function logs." Use the corrected form above; the - point it was making survives intact and in a stronger form.** -3. **The filed emphasis on the sweep race was misplaced and has been - de-emphasised.** Whether the sweep wins or the ack does is irrelevant: if the - waiter is still present, `let _ = waiter.send(kind)` drops the verdict into a - dead receiver and swallows the `Err`. The verdict is discarded on *both* arms. - This makes the finding independent of any timing argument, which is a - strengthening, not a weakening. -4. **Filed recommendation 5 was inverted because it would have caused harm.** - It read: *"If the decision is instead to keep the ack unread, delete the ack - from the `SubmitV1` flow entirely and drop `SUBMIT_REPLY_SURBS` handling for - submits."* Implementing that would (a) remove the inbound backflow that - `replies_received` → `inbound_total` → the driver's liveness probe depends on - (`nym.rs:242`, `:251-253`, `nym_driver.rs:285`, `:312`) — which - `globals/G21-…` identifies as precisely why the shim is **not** exposed to the - hub's confirmed liveness fleet-kill — and (b) per `nym.rs:93-95`, push the - driver off the anonymous-send path if the count reached zero, an anonymity - regression. The recommendation now says the opposite: keep the channel, keep - the count non-zero, trim it, and read the reply. - -### Two things this issue does **not** claim, stated so the report does not overreach - -- It does not claim the wallet could have been told. Under dispatch-only the - wallet is answered before the ack exists; that trade is documented at - `nym.rs:563-594` and is not the finding. -- It does not claim a new attacker capability. The destruction is owned by two - **confirmed High** issues. This issue owns the reason the destruction is - undetectable and therefore unbounded in duration. - -### Consequence found during validation that the filing did not state - -The dead channel costs 13 of every submit's 45 Sphinx packets, on the shim's -throttled mixnet egress — the exact budget the confirmed High -`junk-sendtransaction-flood-…` exhausts. The code's own note deprioritises -trimming it *"since submits are rare"* (`nym.rs:88-96`), which is the premise that -issue removes: an unauthenticated stranger makes submits arbitrarily frequent for -about a byte a second. Trimming the count is therefore a ~29% mitigation of a -confirmed High as well as a throughput change — but it is a mitigation, not a fix, -and it must not be done by zeroing the count. - -### Why Medium, and not higher or lower - -**Not High.** The defect creates no attacker capability on its own; every -migration it fails to report was destroyed by a mechanism that is filed and graded -elsewhere (two confirmed Highs). No user data is exposed and no funds move. - -**Not Low or Info.** Three things keep it at Medium: - -1. It is the **mechanism carrier** for the entire acknowledgement-boundary family. - The sibling documentation, test and type-level findings - (`readme-says-a-failed-migration-…` Low, - `divert-nym-test-certifies-…`, `wire-ack-refusal-documented-as-fail-closed-…`, - `shim-hub-submit-verdict-type-means-different-things-…`) all describe the same - defect from a different surface; this file is the code. The confirmed README - issue was explicitly deflated Medium → Low **on the stated basis that this file - owns the harm** — grading this one below Medium would leave the family with no - carrier and would trigger that file's recorded revisit condition. It is not - being triggered: this issue is graded Medium, so - `readme-says-a-failed-migration-…` correctly stays at Low. -2. An entire wire-protocol safety feature is inert while the system pays 100% of - its cost, and the code documents the opposite contract (`wire.rs:209`). That is - a real, unconditional defect on the shipped transport, affecting every user of - every diverting endpoint on every failure. -3. `globals/G7-…` §5(b) establishes that reading the ack is the single highest- - leverage change available for the whole silent-loss class. A defect whose fix - is the system's best available detection improvement is not a code-quality note. - -### Instruction to the report generator - -This is the **code-side owner** of the acknowledgement-boundary chapter. Carry it -as the mechanism; carry `readme-says-a-failed-migration-…` (Low) beside it as the -disclosure defect; do not count them as two loss findings. The three facts worth -quoting are (a) `wire.rs:209` states *"Every refusal fails closed at the shim"* -and none does; (b) `deliver` logs exactly when it cannot understand a reply and is -silent exactly when it understood one and discarded it; and (c) the transport -asymmetry — identical hub state yields `error_code = -1` on clearnet -(`hub.rs:143-151` → `intercept.rs:188`) and `error_code = 0` on the mixnet, and -`intercept.rs:188` is the only line in the system that can report a hub refusal to -a wallet, is unreachable in the deployed configuration, and is executed by no -test. - -DO NOT DEVIATE FROM THIS TEMPLATE (`docs/SECURITY-ISSUE-TEMPLATE.md`). -LEAVE THIS MESSAGE IN PLACE SO THAT ALL AGENTS KNOW NOT TO DEVIATE FROM THIS TEMPLATE (`docs/SECURITY-ISSUE-TEMPLATE.md`). diff --git a/zeronym-22aa9851caf68-high-medium/medium/operator-controlled-dns-permits-a-layer-4-relay-that-every-documented-verification-step-passes.md b/zeronym-22aa9851caf68-high-medium/medium/operator-controlled-dns-permits-a-layer-4-relay-that-every-documented-verification-step-passes.md deleted file mode 100644 index ef6f9b15..00000000 --- a/zeronym-22aa9851caf68-high-medium/medium/operator-controlled-dns-permits-a-layer-4-relay-that-every-documented-verification-step-passes.md +++ /dev/null @@ -1,395 +0,0 @@ -# The auditor recipe never looks at the network path, so a plain TCP forwarder on the operator's own DNS record hands them the wallet leg while all four documented checks — and the platform's strongest check — pass, correctly and without a Certificate Transparency trace - -**Severity**: Medium -**Validation Status**: Confirmed -**Location**: `audit-target/zeronym/README.md:71` (the auditor procedure, the claim under test); `audit-target/zeronym/shim/deploy/caution/OPERATORS.md:59` ("A DNS name you control"), `:116-131` (the expected CNAME shape), `:139-142` (the only interposition warning, with a rationale that excludes this case), `:178` (the verify invocation), `:340` (redeploy changes app id and IP); `audit-target/zeronym/deploy.sh:65-69` and `:82-97` (`set_dns_record`), `:162-172` and `:182-195` (the record written and rewritten on every deploy), `:206-221` (what is published to `APP_SOURCE`); `audit-target/zeronym/shim/deploy/caution/assemble-caution.sh:585-601` (`PROVENANCE`, which records the domain but not the app id); `audit-target/zeronym/shim/deploy/caution/caution.hcl.tmpl:79-126` (in-enclave TLS termination) -**Found by agent:** Global, focus area G29 — established while settling whether the operator can observe the wallet→shim leg in an attested deployment -**In scope of audit?** Yes — `README.md` and `OPERATORS.md` are in scope as security claims, and `deploy.sh` is in scope as part of the trust chain. - -> **ANTI-DOUBLE-COUNT, APPLIED DURING VALIDATION.** The *capability* this -> interposition yields — the wallet's source IP, connection timing and cleartext -> TLS record lengths, and from them the exact `|tx|` and the IP↔transaction↔amount -> join — is **already graded inside the confirmed High** -> `core-linkage-survives-in-the-attested-deployment-…md`, which enumerates this as -> **route 1C of its step 1** and names this file as its separate filing. **This -> issue is graded on the detection gap only**, and its severity must never be -> added to that High. The filed addendum's recommendation of **Medium → High is -> rejected** for that reason, on the precedent recorded in coordinator item 7a. -> -> **The relay does not decrypt anything.** TLS terminates on the Caddy inside the -> enclave; the forwarder sees ciphertext. Any sentence implying otherwise is -> wrong and must not be written. - -## Description - -`README.md:71` is the project's whole auditor-facing verification contract: - -> - **Auditors** verify an endpoint without trusting its operator: fetch its -> attestation, check the PCRs against the AWS Nitro root, reproduce the build and -> compare hashes, and check Certificate Transparency for a shadow certificate. - -All four checks are aimed at one interposition: an operator who **terminates** the -wallet's TLS in front of the enclave. That needs a certificate for the -wallet-facing name, which is why Certificate Transparency is named as the -defence. - -**A second interposition costs nothing, leaves no record, and every one of those -four checks passes — as does the strongest check the platform itself offers.** The -operator points the wallet-facing DNS record at a host they own that forwards TCP -port 443 to the Caution endpoint without terminating TLS. - -| Documented check | Under a layer-4 TCP forwarder | -|---|---| -| Fetch `https:///attestation` | Reaches the enclave through the forwarder; the document is genuine and AWS-signed. **Passes.** | -| PCRs against the AWS Nitro root | Nothing about the image changed. **Passes.** | -| Reproduce the build, compare hashes | Unrelated to the network path. **Passes.** | -| Certificate Transparency for a shadow certificate | **No certificate is issued.** The forwarder presents none; the enclave's Caddy still serves the only certificate for the name. Nothing appears in any CT log. **Passes vacuously.** | -| `caution verify`'s attested-TLS `certfp` binding (not in the recipe, but the platform's strongest check) | The verifier's TLS session still terminates on the in-enclave Caddy, so the leaf is unchanged and the attested `certfp` matches. **Passes — and passes *correctly*.** | - -The last row is the point, and it is sharper than "the recipe omits a check": the -platform's binding answers *"did my TLS session terminate inside the attested -enclave?"*, to which the answer under a forwarder is genuinely **yes**. The -binding is sound. It is **orthogonal to this attack by construction**, because -the attack does not touch the key, the certificate or the enclave. - -What the forwarder yields is **metadata**: the wallet's source IP, connection -timing, and every TLS record length in both directions (the record length field -is cleartext in every TLS version). That is exactly the input the confirmed High -`core-linkage-…md` needs for its step 1, and exactly the thing the enclave -architecture otherwise denies the operator — `README.md:27` sells the source-IP -property to users, and G30/G4 confirmed from platform source that the wallet's IP -provably never enters the enclave. - -**The prerequisite is documented as the operator's, not as a trust assumption.** -`shim/deploy/caution/OPERATORS.md:59`: *"**A DNS name you control** for wallets."* - -## Attack Scenario and Steps - -Attacker: the shim operator, or anyone who obtains control of the wallet-facing -DNS record — a compromised or compelled registrar or DNS provider, or an insider. - -1. Deploy the shim exactly as `deploy.sh` does, with `DEBUG=0` and an - `APP_SOURCE`. The enclave is genuine, the build reproduces, `caution verify` - prints `Attestation verification PASSED`. -2. Stand up a TCP forwarder on a host the operator owns — a dozen lines of - `socat`, an `nginx stream` block, or one `iptables` DNAT rule — forwarding - `:443` to `.apps.caution.sh:443`. -3. Repoint the wallet-facing record from the CNAME `deploy.sh:169-172` wrote to an - A record for the forwarder. -4. Capture. Every wallet connection now traverses the operator's host: source IP, - timing to the millisecond, and TLS record lengths in both directions. -5. Run the join described in `core-linkage-…md` steps 2–5. - -**Attack Requirements and Assumptions:** - -- **Access needed:** control of the wallet-facing DNS record — which the runbook - states as a prerequisite the operator must hold — and any host to run a TCP - forwarder on. No AWS account, no platform account, no certificate, no - interaction with the enclave. -- **Which party holds the DNS name, per deployment shape:** the operator, in - *both* shapes. `deploy.sh:65-69` requires `TLS_DOMAIN` to be a label under a - Vultr zone the operator holds, and `set_dns_record` (`:82-97`) deletes and - rewrites every record for that name using the operator's own `VULTR_API_KEY`. - Nothing about the DNS record differs between fully-managed and BYOC. -- **Where the capability is *marginal*:** in **BYOC** (`OPERATORS.md:66-69`, - `caution init --byoc`) the parent host is already in the operator's own AWS - account, so they hold the accepting socket without any DNS trickery and the - forwarder adds nothing. `deploy.sh` does not use BYOC. -- **Where the capability is *load-bearing*:** in the shape `deploy.sh` performs — - `caution apps create`, which `OPERATORS.md:64` defines as *"**Fully managed**: - in Caution's AWS account"* — **the parent host belongs to the Caution platform, - not to the operator**, and with `DEBUG=0` no `--ssh-key` is passed - (`deploy.sh:128-134`), so no `debug.ssh_keys` entry is rendered. In that - deployment the forwarder is one of only two routes by which the operator - reaches the wallet leg at all; the other is hand-running - `assemble-caution.sh --ssh-key` without `--debug`, which `deploy.sh` never - does. -- **What makes it durable:** `OPERATORS.md:340` records that a redeploy produces a - *"new app id AND new IP"*, and `deploy.sh:162-195` rewrites the record on every - deploy — so a third party watching resolution has no stable baseline, and a - change is the expected state rather than an alarm. -- **What it costs the attacker:** nothing detectable. It spends no ACME issuance, - so it does not appear in `RESTARTS.md`'s ledger; it produces no CT entry; it - moves no PCR; and it is reversible in one DNS edit. -- **The one partial observable, stated honestly:** in the honest configuration - `` is a CNAME to `.apps.caution.sh` - (`OPERATORS.md:116-122`), and under a forwarder it is not. So a third party who - has read the *operator runbook* can check the **suffix** of the resolution - chain. They cannot check the **value**, because no zeronym artefact publishes - the app id (see the `PROVENANCE` block below), and `README.md:71` — the only - text addressed to auditors — never mentions DNS at all. - -## Impact on Users - -Wallet users are told at `README.md:71` that an endpoint can be verified *"without -trusting its operator"*, and at `README.md:26-27` that the operator cannot see -their broadcast contents and that the on-chain transaction carries no link to -their IP. For the property that matters most to them — that the operator cannot -associate their IP address with the transaction they broadcast — the stated -verification is not achievable by the stated procedure, because the procedure -never looks at the network path and its one anti-interposition check (CT) is -blind to the interposition that costs nothing. - -The practical consequence: an auditor can run all four documented checks, publish -that the endpoint is verified, and be wrong about exactly the property users act -on. Users have no other signal — no wallet checks attestation, and the linkage -that follows is permanent and retrospective, because the chain is public forever. - -## Technical Details / Code Analysis - -**1. The record is the operator's, and the deploy driver rewrites it.** -`shim/deploy/caution/OPERATORS.md:59`: - -``` -- **A DNS name you control** for wallets. -``` - -`deploy.sh:65-69` pins `TLS_DOMAIN` under an operator-held Vultr zone, and -`:169-172` writes the record: - -```sh -REC_TYPE=CNAME -REC_DATA="$APP_ID.apps.caution.sh" -[ "$DNS_CNAME_TRAILING_DOT" = 1 ] && REC_DATA="$REC_DATA." -set_dns_record "$REC_TYPE" "$REC_DATA" -``` - -with `set_dns_record` (`deploy.sh:82-97`) deleting **every** existing record for -that name first. The record is therefore fully under operator control, is -rewritten routinely, and its value is not published anywhere an auditor is told -to look. - -**2. The only interposition warning in the runbook covers only the detectable -case.** `shim/deploy/caution/OPERATORS.md:139-142`: - -> The record must be **DNS-only**: a Cloudflare-proxied (orange cloud) record -> terminates TLS at Cloudflare, which destroys the in-enclave-key property the -> whole attestation argument rests on, and blocks the ACME challenge so no -> certificate ever issues. Both failures are silent. - -Both named consequences are consequences of *terminating* TLS. A layer-4 -forwarder does neither. TLS passes through end to end, so the in-enclave-key -property is intact; and issuance succeeds normally, because the enclave's Caddy -has **no port-80 path at all** — its only vsock listener is 443 -(`/bin/socat VSOCK-LISTEN:443,reuseaddr,fork TCP:127.0.0.1:443` in the platform's -`run.sh` template) — so validation is TLS-ALPN-01 on 443, which a TCP forwarder -relays like any other connection. The rule "must be DNS-only" is stated with a -rationale that does not cover the case that matters. - -**3. The manifest's own property is preserved, which is why nothing detects it.** -`shim/deploy/caution/caution.hcl.tmpl:81-89`: - -``` - # `e2e_encryption { enabled = true }` is Caution's in-enclave TLS - # termination … the private key is generated and held inside the enclave and - # the operator never holds it, which is the property the whole attestation - # argument depends on. -``` - -Correct, and unaffected. The design protects the *key* and therefore the -*content*. It does not protect the *path*, and the source IP is a property of the -path. - -**4. The platform's `certfp` binding passes, and passes correctly.** -`caution verify` computes the SHA-256 of the leaf certificate of *its own* -connection and requires it to equal the value the enclave signed into the Nitro -attestation (`src/cli/src/lib.rs`, `validate_attested_tls`): - -```rust - anyhow::ensure!(user_data.tls.mode == "tls", "attested TLS mode is not tls"); - anyhow::ensure!(user_data.tls.domain == expected.domain, …); - anyhow::ensure!(user_data.tls.certfp == observed_certfp, - "attested TLS certfp does not match the live leaf certificate"); -``` - -The attested value is produced inside the enclave by `caddy-certfp.sh`, which -opens a local TLS connection to the enclave's own Caddy and publishes -`sha256(leaf DER)` into `/metadata.json`, from which `bootproofd` places it in the -COSE-signed `user_data`. Under a forwarder neither side of that equality changes. - -**5. The platform ships a check that WOULD catch it, and the recipe steers away -from it.** `caution verify` has two TLS paths. `tls_connection` returns -`AttestationResponse` when `--attestation-url` is `https:///…`, -and `PinnedIp(ip)` when it names a raw address. Only the second path resolves the -domain and compares: - -```rust -fn dns_contains_deployment_ip(domain: &str, deployment_ip: IpAddr, addresses: &[SocketAddr]) -> Result { - … - anyhow::ensure!( - addresses.iter().any(|address| address.ip() == deployment_ip), - "configured TLS domain {} does not resolve to deployment IP {}", domain, deployment_ip); -``` - -An auditor who learns the deployment's real address **independently** and passes -it as the attestation URL gets a **hard failure** under a forwarder, because -`` resolves to the forwarder and not to that address. That is -recommendation 1 of this issue, already implemented in the tool. - -It is never invoked, for three reasons that are all zeronym's: -`README.md:71` and `OPERATORS.md:178` both use the domain form; **no zeronym -artefact publishes the app id or the managed hostname** — `assemble-caution.sh` -writes `PROVENANCE` with `serves: $TLS_DOMAIN` and the source commit, and the -tree pushed to `APP_SOURCE` (`deploy.sh:206-221`) is the commit assembled -*before* `caution apps create` ran, so it cannot contain the app id; and -`OPERATORS.md:340` conditions readers to expect the address to change on every -redeploy. Note the failure mode of a half-informed attempt: resolving -`` yields the forwarder's address, `dns_contains_deployment_ip` -compares it against itself and matches, and the pinned connection then reaches -the forwarder and gets the enclave's relayed leaf — so the check **passes**. It -only works with an out-of-band source for the deployment address, which is what -recommendation 3 asks the project to publish. - -**6. `network.ingress` is unmeasured, so the forwarder can be made mandatory.** -`ingress { cidr_ipv4 = … }` is consumed by the platform's security-group -generation, not by the enclave image: `EnclaveManifest` -(`src/enclave-builder/src/manifest.rs`) has fields for `app_source`, -`enclave_source`, `framework_source`, `binary`, `run_command`, `metadata` and the -component commits, and **none for `network` or `resources`** — so changing -ingress changes no PCR. An operator who narrows the shim's ingress from -`0.0.0.0/0` to their forwarder's address turns the interposition from a *default* -into an *enforcement*: a wallet that pins the real deployment address, or an -auditor attempting the raw-IP flow above, can no longer connect at all. - -## Recommendations - -1. **Add the network path to the auditor procedure.** `README.md:71` should - instruct auditors to resolve `` and confirm it points at the - Caution-managed target for the app id named in the published artefacts. This - is one `dig` and it converts an invisible interposition into an observable - one. Pair it with the already-recommended check that `ZIS_HUB_NYM` in - `.manifest.run_command` matches `https:///nym-address`. -2. **Publish the expected resolution target.** Record the app id and its managed - hostname in `PROVENANCE` (which `assemble-caution.sh` already writes) or - alongside the PCRs, so recommendation 1 becomes mechanical, scriptable and - continuously monitorable — the same spirit as a CT watch. Without this, the - platform's own `dns_contains_deployment_ip` check cannot be used, and a - half-informed attempt at it passes. -3. **Restate the "DNS-only" rule with the correct rationale.** - `OPERATORS.md:139-142` should say that *any* interposition on the - wallet-facing record — including one that does not terminate TLS — defeats the - source-IP property, and that the rule is not self-enforcing because a - pass-through forwarder breaks nothing an operator would notice and issues no - certificate. -4. **Say plainly in `README.md` that attestation does not cover the network - path.** List the things it does not bind together: the hub the shim diverts to, - the route by which a wallet reaches the enclave, and (per the - `network.ingress` fact above) who may reach it. State that `certfp` proves the - session terminated in the enclave and proves nothing about what the session - traversed. -5. **Have an auditor check `network.ingress` in the published `caution.hcl`.** - `resources` and `network` are the two manifest blocks covered by no - measurement; a narrowed ingress is a sign the operator intends the - interposition to be unavoidable. - -## Validation Information - -**Validated 2026-08-18. VERDICT: CONFIRMED. Severity HELD at Medium; the -addendum's recommended Medium → High is REJECTED. Every platform claim was -re-derived from the Caution source clone (`codeberg.org/caution/platform` @ -`1f8d8cb`) rather than inherited, and every zeronym citation was checked against -the target at HEAD.** - -### 1. What was verified from primary sources - -- **`validate_attested_tls`** (`src/cli/src/lib.rs:354-385`) compares - `user_data.tls.certfp` against `observed_certfp` and additionally pins - `user_data.tls.mode == "tls"` and `user_data.tls.domain == expected.domain`. -- **`observed_certfp` is the leaf of the verifier's own connection.** On the - documented `https:///attestation` form, `tls_connection` - (`:221-241`) returns `AttestationResponse` and `verify_tls_binding` - (`:6892-6912`) hashes the leaf of that very response. **A layer-4 forwarder - changes neither operand.** -- **`caddy-certfp.sh`** (`src/enclave-builder/templates/caddy-certfp.sh`) is a - loop that `openssl s_client`s the enclave's own Caddy on `127.0.0.1:443` with - `-verify_return_error -verify_hostname`, and writes - `{"tls":{"mode":"tls","domain":…,"certfp":…}}` to `/metadata.json`. The attested - value is therefore the *enclave's* leaf, unconditionally. -- **`dns_contains_deployment_ip`** (`:261-278`) is reached only from the - `PinnedIp` branch (`:6902`, `:6930`), i.e. only when `--attestation-url` names a - raw address. Confirmed by reading both call sites. -- **The enclave has no port-80 ingress.** `run.sh.template` starts exactly one - wallet-facing vsock listener, `VSOCK-LISTEN:443 → TCP:127.0.0.1:443`, and the - parent's Caddy binds `:80` only for health, `/attestation` and a 308 redirect - (`terraform/modules/aws/nitro-enclave/user-data.sh`). So ACME must be - TLS-ALPN-01 on 443, and a TCP forwarder relays it unchanged — the filed claim - that issuance still succeeds is **correct**. -- **`EnclaveManifest`** (`src/enclave-builder/src/manifest.rs:14-45`) has no - `network` or `resources` field. The `network.ingress`-is-unmeasured claim is - **correct**. -- **`deploy.sh` publishes no app id.** `assemble-caution.sh` commits the tree and - writes `PROVENANCE` with `serves: $TLS_DOMAIN`, `source commit`, - `expected binary` — and no app id, which does not exist yet. - `deploy.sh:210-215` pushes that same commit to `APP_SOURCE`. Verified by - reading both scripts end to end. - -### 2. Corrections applied to the filed text - -- **The "What limits it further" hedge is withdrawn and replaced by a precise - per-shape statement.** The filed text hedged on whether the operator already - owns the parent host. They do in **BYOC** (where the forwarder is marginal) and - they do **not** in the fully-managed path `deploy.sh` performs (where it is - load-bearing). Both are now stated, with the deciding citations - (`OPERATORS.md:64` vs `:66-69`, `deploy.sh:156`, `deploy.sh:128-134`). The - audit's `THREATMODEL.md` asserts operator ownership of the parent host - unconditionally; that assertion is false in a managed deploy and was not - inherited here. -- **A partial observable was ADDED against the filing, because omitting it would - have overstated the finding.** `OPERATORS.md:116-122` documents that the record - should be a CNAME to `.apps.caution.sh`, so the *suffix* of the - resolution chain is checkable by anyone who has read the operator runbook. This - does not refute the issue — the *value* is uncheckable because no artefact - publishes the app id, and `README.md:71` (the auditor-facing text) never - mentions DNS — but a report that claimed "no observable exists" would be - falsifiable in one `dig`. Recommendation 2 is strengthened accordingly. -- **The "relay" vocabulary is retained only for the operator's own forwarder.** - Per coordinator item 7h, the Nitro **parent host** is described nowhere in this - file as a relay: it is the enclave's router and DNS resolver. The forwarder in - this attack is a distinct thing on a distinct host. -- **No plaintext claim appears anywhere.** The forwarder sees ciphertext; the - metadata channel is record lengths and timing. Stated explicitly in the banner. -- **The Location line numbers were tightened** (`deploy.sh:162-172` and `:182-195` - for the two `set_dns_record` calls, rather than the filed `:166-192`; the - `PROVENANCE` block added; `caution.hcl.tmpl:79-126` for the http block). - -### 3. Why Medium, and why not High - -**Not High**, for one reason and it is decisive: the *capability* is already -graded inside the confirmed High `core-linkage-…md`, whose step 1 enumerates this -as route 1C by name and whose validation already verified the same platform -mechanics. Grading this High would count one loss twice. Coordinator item 7a -records the governing precedent — a validator rejecting a recommended severity to -avoid multi-counting a harm owned elsewhere — and it applies here in the upward -direction. - -**Not Low**, for three reasons. (i) The falsified text is not an incidental -comment: it is the entire auditor-facing verification contract on the front page -of the project, and `README.md:71`'s claim is a completeness claim -(*"without trusting its operator"*) that the four listed checks cannot support. -(ii) The gap is **structural**, not stale: no check in the recipe looks at the -network path, and the platform's strongest check is orthogonal to this attack by -construction rather than by oversight — so a reader cannot repair it by being -more careful. (iii) A usable remedy exists in the platform *today* -(`dns_contains_deployment_ip`) and is unreachable only because zeronym never -publishes the one value it needs; a finding whose fix is "publish a string you -already have" and which is currently blocking a working check is worth more than -Low. - -This grade matches the closest precedent in `issues/confirmed/`: -`attested-tls-binding-is-verified-once-by-hand-if-ever-…md` (Medium), which is the -same shape — a real binding, verified once by hand if ever, with `README.md:26` -stating the resulting protection unconditionally. The two are **siblings, not -duplicates**: that one is about a check that exists and is not repeated; this one -is about an interposition that check cannot see at all. Report them together -under one heading — *attestation says nothing about how a wallet reaches the -enclave* — and count them once each. - -### 4. Nothing was withdrawn - -Every mechanical claim in the filing and its addendum survived re-derivation. The -changes above are a rejected severity escalation, a withdrawn hedge replaced by a -per-shape statement, one added observable that bounds the finding honestly, and -tightened citations. - -DO NOT DEVIATE FROM THIS TEMPLATE (`docs/SECURITY-ISSUE-TEMPLATE.md`). -LEAVE THIS MESSAGE IN PLACE SO THAT ALL AGENTS KNOW NOT TO DEVIATE FROM THIS TEMPLATE (`docs/SECURITY-ISSUE-TEMPLATE.md`). diff --git a/zeronym-22aa9851caf68-high-medium/medium/publish-verdict-strings-are-zcashds-vocabulary-only-so-a-zebrad-backed-indexer-turns-not-judged-into-permanent-destruction.md b/zeronym-22aa9851caf68-high-medium/medium/publish-verdict-strings-are-zcashds-vocabulary-only-so-a-zebrad-backed-indexer-turns-not-judged-into-permanent-destruction.md deleted file mode 100644 index 61c963d9..00000000 --- a/zeronym-22aa9851caf68-high-medium/medium/publish-verdict-strings-are-zcashds-vocabulary-only-so-a-zebrad-backed-indexer-turns-not-judged-into-permanent-destruction.md +++ /dev/null @@ -1,640 +0,0 @@ -# `classify_publish_error` matches only zcashd's reject vocabulary, so behind the shipped example indexer (lightwalletd in front of zebrad) *every* node answer is a `Rejected` verdict — including the two that mean "the node never looked at this transaction", which the batcher then destroys permanently - -**Severity**: Medium -**Validation Status**: Confirmed -**Location**: `audit-target/zeronym/hub/src/chain.rs:503-533` (`classify_publish_error` and its doc comment), reached from `:438-445` (`classify_send_response`) via `:176-198` (`broadcast`) and `:208-210` (`broadcast_batch`); the requeue-vs-drop split it feeds at `audit-target/zeronym/hub/src/batcher.rs:365-390`, and the telemetry at `:396-402`, `:412-420`. Deployed indexer: `audit-target/zeronym/deploy.env.example:22-23`. Node-side strings (out of scope, reasoned about as an interface): `audit-context/zero/zebra/zebrad/src/components/mempool/error.rs:50,55,61,70,76`, `zebrad/src/components/mempool/storage.rs:106`, `zebrad/src/components/mempool/downloads.rs:127-139`, `zebrad/src/components/mempool.rs:330,419-452,1116-1123`, `zebra-rpc/src/methods.rs:1305-1360`, `zebra-rpc/src/server/error.rs:105-107`, `zebra-rpc/src/queue.rs:39,42,206-241`. Relay: `audit-context/zero/lightwalletd/frontend/service.go:548-580`, `frontend/rawrequest.go:150-161`. Contrast: `audit-context/zero/zcashd/src/rpc/rawtransaction.cpp:1318-1339`, `src/main.cpp:2103,2132`. -**Found by agent:** Global (focus area G9, cross-component invariant drift — dedicated re-run) -**In scope of audit?** Yes - -## Description - -`hub/src/chain.rs` decides, for every member of every flushed batch, whether the -indexer's answer is a **verdict** (the entry is dropped forever) or a **transport -failure** (the entry is requeued for the next flush). When the indexer answers -`OK` carrying a `SendResponse` with a non-zero `error_code`, that decision is made -by matching four substrings against the node's error text: - -`hub/src/chain.rs:513-533` - -```rust -fn classify_publish_error(message: &str) -> Publish { - // Hyphens folded to spaces before matching. Bitcoin-derived nodes report - // these as hyphenated reject reasons (`txn-already-known`) while the longer - // prose forms use spaces, and matching only one shape silently misses the - // other. ... - let m = message.to_ascii_lowercase().replace('-', " "); - if m.contains("already in block chain") - || m.contains("already known") - || m.contains("already in mempool") - || m.contains("duplicate") - { - Publish::AlreadyKnown - } else { - Publish::Rejected { - reason: message.to_string(), - } - } -} -``` - -Its own doc comment states why the match is on text at all (`chain.rs:503-512`): - -> Matched on text because the error codes for these cases are not consistent -> **between zebrad and zcashd**, nor through an indexer that relays them. Kept in -> one place and deliberately conservative: anything unrecognised is a rejection, -> never a silent success. - -**All four substrings are zcashd's vocabulary. None of them is zebrad's.** The -four map one-for-one onto `zcashd/src/rpc/rawtransaction.cpp:1336` -(`"transaction already in block chain"`), `zcashd/src/main.cpp:2103` -(`txn-already-in-mempool`), `:2132` (`txn-already-known`), and the family of -`bad-*-duplicate` reject reasons. - -Every string zebra's `sendrawtransaction` can produce was extracted from the -vendored source and run through the exact predicate above -(`.to_ascii_lowercase().replace('-', " ")`, then the four `contains`). **Not one -matches.** Behind zebrad, `classify_publish_error` returns `Publish::Rejected` -for *every* answer the node can give: - -| zebrad string (verbatim) | source | what it means | `classify_publish_error` | -|---|---|---|---| -| `mempool is disabled since synchronization is behind the chain tip` | `mempool/error.rs:76` (`Disabled`) | **not judged** — mempool not yet active | `Rejected` | -| `transaction dropped because the queue is full` | `mempool/error.rs:70` (`FullQueue`) | **not judged** — request ignored | `Rejected` | -| `transaction already exists in mempool` | `mempool/error.rs:55` (`InMempool`) | already known | `Rejected` | -| `transaction dropped because it is already queued for download` | `mempool/error.rs:61` (`AlreadyQueued`) | already known | `Rejected` | -| `… until a chain reset: transaction was committed to the best chain` | `mempool/error.rs:50` + `storage.rs:106` | already mined | `Rejected` | -| `transaction is already in state` | `mempool/downloads.rs:127` | already mined | `Rejected` | -| `error in state service: …` | `mempool/downloads.rs:130` | **not judged** — internal fault | `Rejected` | -| `transaction download / verification was cancelled` | `mempool/downloads.rs:136` | **not judged** — task cancelled | `Rejected` | -| `transaction did not pass consensus validation: …` | `mempool/downloads.rs:139` | genuinely refused | `Rejected` (correct) | -| `transaction is non-standard` | `mempool/error.rs:80` | genuinely refused | `Rejected` (correct) | - -Four of these are the node saying *it never evaluated the transaction*. The hub -converts all four into `Publish::Rejected`, which `batcher.rs:368` counts as -rejected and **drops permanently** — while the wallet was told `error_code 0` -with a txid at mixnet dispatch, minutes earlier. - -This is the exact outcome `chain.rs:485-490` says the classification exists to -prevent: - -> re-offering an entry costs one call per flush and stops at its expiry, while -> **dropping a valid migration on a misread error is unrecoverable, because the -> shim has already told the wallet it was sent**. - -`classify_publish_failure` (`chain.rs:491-501`) applies that principle correctly -to gRPC-level failures — only `INVALID_ARGUMENT` and `FAILED_PRECONDITION` are -verdicts, everything else is `Retryable`. `classify_publish_error` inverts it for -`SendResponse`-level failures: everything unrecognised is a verdict. Against -zcashd that default is nearly harmless, because zcashd pre-checks the mempool and -the chain before `AcceptToMemoryPool` and so only reaches the error path after it -has actually judged the transaction (`rawtransaction.cpp:1318-1339`). Against -zebrad it is wrong for four of ten reachable answers, and the type-level -distinction `Publish` documents at `chain.rs:100-108` — *"`Retryable` means -nothing judged the transaction at all"* — is simply not delivered. - -## Attack Scenario and Steps - -This fires with **no adversary at all**, which is what separates it from -`hub-flush-destroys-migration-on-single-unverifiable-verdict.md` (a hostile -indexer choosing to lie). Routine operation, in the configuration the project -ships: - -1. The hub's configured indexer is `INDEXERS=66.241.124.200:443` / - `INDEXER_TLS=na.zec.rocks` (`deploy.env.example:22-23`) — a **zec.rocks** - endpoint. zec.rocks publishes its own stack - (): the core is - **Zebra + lightwalletd**, with Zaino as an optional add-on, and the operators - stated on the Zcash community forum (thread 50907, checked 2026-08-18) that as - of June 2026 most production endpoints run *"LWD+zebra"* and only - `zaino.unsafe.zec.rocks` test endpoints run Zaino. lightwalletd-in-front-of- - zebrad is also a first-class, documented, CI-tested configuration - (`zebra/book/src/user/lightwalletd.md`), and zcashd is end-of-life by the - monorepo's own README (*"the long-term direction is the Zebra, Zaino and - Zallet (Z3) stack"*, `audit-context/zero/README.md:31,45-46`). -2. lightwalletd relays the node's JSON-RPC error verbatim: it splits - `": "` and puts the message straight into - `SendResponse.error_message` (`frontend/service.go:554-566`, - `frontend/rawrequest.go:158-159`). zebra's `map_misc_error` puts the - `MempoolError` `Display` string in unmodified - (`zebra-rpc/src/server/error.rs:105-107`). -3. **That zebrad is restarted** — an upgrade, a host reboot, an OOM, a crash, a - state-format migration, or a first deployment still doing its initial sync. - Zebra constructs its mempool `Disabled` (`mempool.rs:330`) and only enables it - once the syncer's last three response lengths average under 20 blocks - (`sync/status.rs:27,72-99`), which cannot be true until several sync rounds - have completed. -4. The hub's 20-block cadence fires. `broadcast_batch` publishes every held - migration; every call returns - `SendResponse{error_code: -1, error_message: "mempool is disabled since - synchronization is behind the chain tip"}`. -5. `classify_publish_error` matches none of the four substrings → - `Publish::Rejected` for **every** member. `best_of` (`chain.rs:459-474`) ranks - `Rejected` above `Retryable`, so a second endpoint that is merely unreachable - cannot rescue them; only a second endpoint that actually *accepts* can. -6. `flush` (`batcher.rs:368`) counts them `rejected` and drops them. `requeued` - is 0, so the `warn!` at `:404-411` (which needs `requeued > 0`) does not fire. - Nothing else in the system holds a copy: the shim answered the wallet - `error_code 0` at mixnet dispatch (`shim/src/intercept.rs:181-202`, - `shim/src/nym.rs:595-600`), keeps no per-migration state, and the enclave is - diskless. - -**The node-side mitigation, quantified — this is what bounds the finding.** -`zebra-rpc`'s `send_raw_transaction` pushes the transaction into zebra's *own* -retry queue **before** it asks the mempool (`methods.rs:1322-1324`), so even the -`Disabled` answer leaves a copy behind. That queue re-offers the transaction on -every tip change and drops it after -`NUMBER_OF_BLOCKS_TO_EXPIRE × spacing + 5 s = 5 × 75 + 5 = 380 s` -(`zebra-rpc/src/queue.rs:39,206-241`), with capacity 20 (`:42`). So: - -> **A migration is destroyed only if zebra is still not close to the tip ~380 s -> after the flush.** A restart that reaches the tip inside 6.3 minutes is fully -> absorbed by zebra and nothing is lost. - -That makes the realistic trigger a *long* node-behind-tip window: a cold start, a -node that was down for hours, an initial sync, a state-format upgrade, or a slow -host — not a fast service restart. The same 380 s queue also covers the -`error in state service` and `cancelled` rows, and the hub's own 10 s -`RPC_TIMEOUT` (`chain.rs:48`) means zebra's 73 s verification timeout string is -not reachable through the hub at all. `FullQueue` needs 500 concurrent pending -mempool downloads (`downloads.rs:107`) sustained across the same window, which is -a flood condition rather than an ordinary one. - -**Attack Requirements and Assumptions:** -- Requires the hub's indexer to relay node errors in `SendResponse.error_code` — - lightwalletd's documented convention. **Against Zaino this path is unreachable - entirely**: `node_backed_indexer.rs:1561-1570` returns `error_code: 0` on - success and a `tonic::Status` otherwise, so a node refusal never becomes a - `SendResponse`. That is the separately filed - `hub-chain-zaino-node-rejections-are-never-verdicts.md`, and the two issues - partition the indexer space between them. -- Requires the node behind that relay to be zebrad. Established above for the - shipped example endpoint; and the hub cannot tell either way, since the - endpoint is operator-chosen and can change without the hub noticing. -- **A correction to the original filing:** zebra does *not* disable an - already-active mempool when it falls behind. `Mempool::update_state` returns - early on `(is_caught_up = false, is_enabled = true, _)` - (`mempool.rs:448-450`), with the comment *"Sync status only gates initial - activation. Once the mempool is active, this method does not disable it."* So - `Disabled` is reachable **only from process start**, not from a running node - drifting behind. That narrows the trigger to restarts and first deployments. -- No privileged access, no network position, no mixnet capability. -- **Deliberate variant.** Whoever runs the indexer — a third party here, not the - hub operator — gets a whole-batch kill with total deniability by restarting - their own node near a cadence boundary: the node tells the exact truth, the hub - mistranslates it, and the hub's log records `rejected = N`. This adds - deniability rather than new power over the already-filed hostile-indexer issue. - -## Impact on Users - -A wallet that migrated is told `error_code 0` with a correct txid, and the -transaction is never broadcast. The loss is silent on both sides: the wallet -believes it succeeded, and the hub logs a truthful-looking `rejected` count that -an operator reads as *"the node refused an invalid transaction"*, not *"I -destroyed a valid one"*. - -Because `MempoolError::Disabled` is a property of the node rather than of any -transaction, the failure is **not per-entry**: one flush landing in one node's -post-restart window destroys the entire batch — every user who migrated in that -20-block window, across every shim in the fleet. - -How long the user is stuck depends on which traffic class they are in, and the -two answers are very different: - -- **Today's traffic** (ordinary Orchard spends built by librustzcash-family - wallets, expiry `= tip + 40` per ZIP 203's Blossom default) — the transaction - expires in ~50 minutes, the wallet's notes unlock, and the user can retry. The - harm is a silent failed migration and an hour of confusion. -- **ZIP 318 conforming traffic** — the canonical expiry is a bucketed absolute - height 30 to 60 days out (`SPEC-NOTES.md` §3). The wallet shows the migration - pending and will not reuse those notes until it expires. That is the - separately-filed - `zip318-canonical-expiry-is-the-only-recovery-clock-and-a-lost-migration-freezes-the-users-notes-for-30-to-60-days.md`; - this issue is one of the mechanisms that reaches it. - -> **CORRECTION 2026-08-18 (validation of the cited file — SUPERSEDES the sentence above).** -> The wallet does **not** wait for expiry. Both official Zcash light-wallet SDKs -> automatically resubmit a sent-but-unmined transaction for as long as it remains -> unexpired — the Android SDK at the head of every ~20 s sync loop and after every -> processed block batch (`CompactBlockProcessor.kt:573,615,723`; selection -> `mined_height IS NULL AND expiry_height > ?`), the iOS SDK at most once per 300 s -> (`TxResubmitter.swift:8-15`, `TransactionDao.swift:218-228`) — and the hub's -> payload-hash dedup makes the resend free. The wallet's non-confirmation signal comes -> from compact-block scanning, which the shim does not intercept (`proxy.rs:1068-1074`). -> Expiry is therefore the **retry horizon**, not the wait: ~50 minutes for the ZIP 203 -> default traffic the shim also diverts, 30–60 days for a ZIP 318 migration. A -> *transient* loss self-heals within minutes; only a loss condition that **outlives the -> horizon** destroys the submission permanently — which is exactly what this issue's -> condition does, so this issue's severity is unaffected. Do not write "the user waits -> 30 to 60 days" in the report. Full refutation and the replacement paragraph: -> `issues/invalid/zip318-canonical-expiry-is-the-only-recovery-clock-and-a-lost-migration-freezes-the-users-notes-for-30-to-60-days.md`. - - -At present adoption the modal batch is 0 or 1 (`README.md:34`, `REVIEW.md:175`), -so a single event typically destroys 0–2 migrations; the per-event population -grows exactly as the product succeeds. - -Separately, and with no user harm, `achieved_batch_size` — the number -`batcher.rs:335-337` calls *"the honest measure of the privacy the flush actually -delivered"* and which `hub/REVIEW.md` design change #9 makes the launch gate — is -under-reported whenever an already-known answer is misread as `rejected`. **This -is narrower than the original filing claimed.** On a healthy zebrad a first -publish returns `error_code 0` with the txid, so the normal case counts -correctly. The mis-mapping bites in two places: (i) after a `Retryable` requeue -where the node in fact already has the transaction, the next flush reads -`transaction already exists in mempool` as `rejected`; and (ii) the second of two -*simultaneously live* hubs, which `shim/src/nym.rs:618-635` states is not the -deployment model (*"sending to every address is therefore safe only while the -other addresses are DEAD"*). The original claim that this happens "on every -honest flush" is withdrawn. - -## Technical Details / Code Analysis - -**The full path from node string to dropped entry.** - -`hub/src/chain.rs:176-198` — one `Publish` per endpoint, folded by `best_of`: - -```rust - pub async fn broadcast(&self, tx_bytes: &[u8]) -> Publish { - let calls = self.endpoints.iter().map(|addr| { - let raw = RawTransaction { data: tx_bytes.to_vec(), height: 0 }; - async move { - match self.unary::<_, SendResponse>(*addr, SEND_TRANSACTION, raw).await { - Ok(resp) => classify_send_response(&resp), - Err(err) => classify_publish_failure(&err), - } - } - }); - best_of(join_all(calls).await) - } -``` - -`hub/src/chain.rs:438-445`: - -```rust -fn classify_send_response(resp: &SendResponse) -> Publish { - if resp.error_code == 0 { - return Publish::Accepted { txid: resp.error_message.clone() }; - } - classify_publish_error(&resp.error_message) -} -``` - -`hub/src/batcher.rs:365-377` — the drop: - -```rust - for (i, entry) in batch.into_iter().enumerate() { - match outcomes.get(i) { - Some(Publish::Accepted { .. }) | Some(Publish::AlreadyKnown) => achieved += 1, - Some(Publish::Rejected { .. }) => rejected += 1, - Some(Publish::Retryable { reason }) => { - sample_failure.get_or_insert_with(|| reason.clone()); - unplaced.push(entry); - } - None => unplaced.push(entry), - } - } -``` - -`Rejected` increments a counter; the entry, moved out of `batch` by `into_iter()`, -is never pushed to `unplaced`, so `queue.requeue(unplaced)` at `:390` does not see -it. It is gone. The comment immediately below (`:379-388`) states the intent this -defeats: - -> A transport failure goes back into the queue for the next cadence. This is the -> only place such a failure can be recovered … **Dropping it because the indexer -> restarted during the flush window would lose the migration outright while the -> wallet believes it was sent.** - -The indexer restarting is handled. The *node behind the indexer* restarting is -not, because that failure arrives as an `OK` response carrying text. - -**Where the strings come from, on the zebra side.** - -`zebra/zebrad/src/components/mempool/error.rs:53-78`: - -```rust - /// Transaction rejected because the mempool already contains another - /// transaction with the same hash. - #[error("transaction already exists in mempool")] - InMempool, - - /// The transaction hash is already queued, so this request was ignored. - #[error("transaction dropped because it is already queued for download")] - AlreadyQueued, - - /// The queue is at capacity, so this request was ignored. - #[error("transaction dropped because the queue is full")] - FullQueue, - - /// The mempool is not enabled yet. - #[error("mempool is disabled since synchronization is behind the chain tip")] - Disabled, -``` - -`zebra/zebrad/src/components/mempool.rs:1116-1123` — while the mempool is off, -every queued transaction gets `Disabled`: - -```rust - Request::Queue(gossiped_txs) => Response::Queued( - iter::repeat_n(MempoolError::Disabled, gossiped_txs.len()) - .map(BoxError::from) - .map(Err) - .collect(), - ), -``` - -`zebra/zebrad/src/components/mempool.rs:448-450` — and the reason `Disabled` is a -startup-only state: - -```rust - // TODO: only disable an already-active mempool when validated sync - // state proves Zebra is behind a higher-work chain ... - (false, true, _) => { - return false; - } -``` - -`zebra/zebra-rpc/src/methods.rs:1322-1348` — the transaction is pushed into -zebra's own retry queue first, then the mempool error is surfaced with its -`Display` string intact: - -```rust - // send transaction to the rpc queue, ignore any error. - let unmined_transaction = UnminedTx::from(raw_transaction.clone()); - let _ = queue_sender.send(unmined_transaction); - ... - let queue_result = queue_results - .pop() - .expect("there should be exactly one item in Vec") - .inspect_err(|err| tracing::debug!("sent transaction to mempool: {:?}", &err)) - .map_misc_error()? -``` - -`zebra/zebra-rpc/src/server/error.rs:105-107` — the message is the error's -`to_string()`, unmodified, under `LegacyCode::Misc = -1`: - -```rust - fn map_error(self, code: impl Into) -> Result { - self.map_err(|error| ErrorObject::owned(code.into().code(), error.to_string(), None::<()>)) - } -``` - -`lightwalletd/frontend/service.go:554-566` — the relay splits `": "` -and puts the message in `SendResponse.error_message` verbatim: - -```go - if rpcErr != nil { - errParts := strings.SplitN(rpcErr.Error(), ":", 2) - ... - errMsg = strings.TrimSpace(errParts[1]) - errCode, err = strconv.ParseInt(errParts[0], 10, 32) -``` - -**Why zcashd hides the defect.** `zcashd/src/rpc/rawtransaction.cpp:1318-1339` -pre-checks the mempool and the chain *before* calling `AcceptToMemoryPool`: - -```cpp - bool fHaveMempool = mempool.exists(hashTx); - bool fHaveChain = existingCoins && existingCoins->nHeight < 1000000000; - if (!fHaveMempool && !fHaveChain) { - ... AcceptToMemoryPool ... - } else if (fHaveChain) { - throw JSONRPCError(RPC_TRANSACTION_ALREADY_IN_CHAIN, "transaction already in block chain"); - } - RelayTransaction(tx); - return hashTx.GetHex(); -``` - -An already-in-mempool transaction takes neither branch and returns the txid with -`error_code 0`, i.e. `Publish::Accepted`. So on zcashd the "already known" case is -handled by a path that never reaches `classify_publish_error`, and two of the four -substrings (`already known`, `already in mempool`) are close to dead code there. -The function's coverage of zcashd is better than it looks, and its coverage of -zebrad is nil. - -**Nothing checks the coupling.** The four strings are a hard-coded model of -another component's error vocabulary, in a monorepo that ships that component's -source. `chain.rs:536-580`'s tests exercise only synthetic strings the author -chose; no test, script or CI job compares the list against any node's actual -output, and no CI runs `cargo test` at all (PROGRESS item 6n). - -## Recommendations - -Recommendations 1 and 2 **must land together**; 1 alone makes things worse, for -the reason given under 2. - -1. **Invert the default for `SendResponse` errors, as `classify_publish_failure` - already does for gRPC errors.** Treat an *unrecognised* message as - `Retryable`, and reserve `Rejected` for messages that positively match a - consensus-failure vocabulary (`bad-txns-*`, `bad-*-nullifiers-*`, - `tx-overwinter-expired`, `insufficient priority/fee`, `Missing inputs`, - `did not pass consensus validation`, `is non-standard`, …). The asymmetry is - the one `chain.rs:485-490` already argues. This single change removes the - `Disabled` / `FullQueue` / `state service` / `cancelled` destruction without - needing to enumerate zebra's vocabulary correctly. -2. **Add zebra's already-known strings to the `AlreadyKnown` set at the same - time**: `already exists in mempool`, `already queued for download`, - `committed to the best chain`, `already in state`. Keep the existing zcashd - strings. **This is not cosmetic once (1) is applied.** With (1) alone, an - already-in-mempool or already-mined transaction becomes `Retryable`, and - `Queue::requeue` has no expiry sweep and no GC path of any kind (confirmed in - `hub-queue-requeue-ignores-byte-budget-unbounded-growth.md`; `REVIEW.md:145` - specifies the deadline that was never implemented). Such an entry would - therefore be **re-broadcast in every batch, forever**, holding queue bytes and - emitting exactly the repeated per-transaction timing signal - `chain.rs:517-521` says the `AlreadyKnown` branch exists to prevent. Applying - (1) without (2) converts silent destruction into permanent republication. - Landing the `REVIEW.md:145` deadline GC alongside is the belt-and-braces - version. -3. **Split `already_known` from `achieved` in the counters** (`batcher.rs:367`, - `:396-401`) rather than folding them, so the launch-gate number distinguishes - "this hub put it on the network" from "someone else already had it". -4. **Pin the coupling with a test.** The monorepo already contains zebra's, - zcashd's, lightwalletd's and Zaino's source as sibling subtrees. A - table-driven unit test listing each node's actual error strings with its - expected `Publish` turns a silently-drifting model of another component into - something a change to that component can break loudly. Nothing about this - requires a live node. -5. **Use `error_code` as a coarse pre-filter, with text as the refinement — but - do not replace the text match with it.** The codes do carry signal (zebra: - `-22` deserialization, `-25` verification refusal, `-1` for everything the - mempool queue rejected without judging; zcashd: `-27` already-in-chain, - `-25`/`-26` rejected), and "only `-22`/`-25`/`-26` may be a verdict" is a - correct and cheap outer guard for both nodes. It is not sufficient on its own, - because `-1` covers both "never judged" and "already known", which recommendation - 2 needs to tell apart. - -## Validation Information - -**Status: CONFIRMED. Severity Medium** (filed as Medium; kept, with the reasoning -below). - -### What was verified, and how - -**1. The predicate is exactly as quoted** — `hub/src/chain.rs:513-533`, -`:438-445`, `:459-474`, `:491-501`; `batcher.rs:341-390`. All line numbers in the -original filing check out against the target. - -**2. Every zebra string was extracted from the vendored source and run through -the actual predicate.** The four `contains` tests were applied mechanically to -all 17 error strings reachable from -`zebrad/src/components/mempool/{error.rs,storage.rs,downloads.rs}`. -**Zero matches.** The claim "not one of them contains any of the four substrings" -is exact, not approximate. Verbatim line numbers: `error.rs:50,55,61,70,76,80`; -`storage.rs:100-116`; `downloads.rs:127,130,133,136,139`. - -**3. The relay chain was verified end to end.** zebra's mempool-queue errors are -surfaced through `map_misc_error` (`methods.rs:1345`) → `map_error` -(`server/error.rs:105-107`) → `ErrorObject{code: -1, message: }`; -lightwalletd's `RawRequest` returns `resp.Error` (a `btcjson.RPCError`, whose -`Error()` is `": "`) at `rawrequest.go:158-159`, and -`service.go:554-566` splits it and copies the message into -`SendResponse.error_message` **verbatim**. The hub then hits -`classify_send_response` → `classify_publish_error`. - -**4. The Zaino partition is confirmed, so the two issues do not double-count.** -`zaino/packages/zaino-state/src/indexer/node_backed_indexer.rs:1561-1570` -constructs `SendResponse { error_code: 0, .. }` on success and returns `Err(..)` -on failure, which becomes a gRPC status and is handled by -`classify_publish_failure`. `classify_publish_error` is genuinely unreachable -behind Zaino. Likewise the `"duplicate"` over-match issue is a *zcashd*-only -phenomenon — zebra's double-spend string is `"transaction inputs were spent, or -nullifiers were revealed, in the best chain"`, which contains no `duplicate`. The -three issues partition cleanly: the function is wrong for every backend, in a -different direction for each. - -**5. Reachability is real, not hypothetical — this was the decisive check.** -- `deploy.env.example:22-23` ships `INDEXERS=66.241.124.200:443` / - `INDEXER_TLS=na.zec.rocks`, and `smoke.sh:46` / `smoke-local.sh:40` use the - same endpoint. -- zec.rocks publishes its own stack: `zecrocks/zcash-stack`'s `docker/` compose - set is **Zebra + lightwalletd**, with `compose.zaino.yaml` as an optional - add-on. Zaino is not the default. -- The zec.rocks operators stated on the Zcash community forum (thread - "Zec.rocks Zcashd Deprecation Timeline", read 2026-08-18) that as of June 2026 - most endpoints run *"LWD+zebra"* and only `zaino.unsafe.zec.rocks` / - `zaino.testnet.unsafe.zec.rocks` run Zaino. -- lightwalletd-on-zebrad is documented and CI-tested upstream - (`zebra/book/src/user/lightwalletd.md`), and lightwalletd's own JSON-RPC client - says it is *"a context-aware JSON-RPC function for zcashd **and zebrad**"* - (`rawrequest.go:72-73`). -- zcashd is end-of-life in this very monorepo (*"a supported fork with a - hardcoded end-of-life, as a transition path only"*, `audit-context/zero/README.md`), - so the vocabulary the function speaks belongs to the node that is being retired - and not to the one that is replacing it. - -**Conclusion on reachability: this is the shipped example configuration, and no -operator has to choose anything unusual to be in it.** - -### Corrections made during validation - -- **The trigger was narrowed.** The original filing said zebra disables its - mempool whenever it "falls behind the tip for any reason". It does not: - `mempool.rs:448-450` returns early for `(not caught up, already enabled)` with - the explicit comment that sync status *"is strong enough to delay initial - activation but not to shut down a working mempool"*. `Disabled` is a - **startup-only** state (`mempool.rs:330`). Rewritten accordingly. -- **The node-side mitigation was quantified and promoted from a footnote to a - bound on the finding.** `send_raw_transaction` enqueues the transaction into - zebra's own retry queue *before* asking the mempool (`methods.rs:1322-1324`), - and that queue retries on every tip change for - `5 × 75 s + 5 s = 380 s` (`queue.rs:39,206-241`). So loss requires zebra to - still be behind the tip ~6.3 minutes after the flush. Short restarts lose - nothing. This is the single largest reason the issue is Medium rather than - High. -- **Two more "never judged" rows were added** that the original filing missed — - `error in state service: …` and `transaction download / verification was - cancelled` (`downloads.rs:130,136`) — and one that was implied but is *not* - reachable was excluded: zebra's `"timeout waiting for verification result"` - fires at 73 s (`downloads.rs:496,514`, `crawler.rs:84`), long after the hub's - own `RPC_TIMEOUT = 10 s` (`chain.rs:48`) has already produced a (correct) - `Retryable`. -- **The "milder half" was cut back.** The claim that `achieved_batch_size = 0` - and `rejected = N` "on every honest flush at the second hub" requires two - simultaneously live hubs, which `shim/src/nym.rs:618-635` states is not the - deployment model. On a healthy zebrad the first publish is `error_code 0` and - counts correctly. The two cases where the mis-mapping does bite are named - explicitly in the impact section. -- **Recommendation 1 was found to be dangerous on its own, and this is now stated - in the issue.** Inverting the default without also extending the `AlreadyKnown` - set turns every already-in-mempool and already-mined answer into `Retryable`; - `Queue::requeue` has no expiry sweep and no GC path (`queue.rs:279-295`; - independently confirmed in - `hub-queue-requeue-ignores-byte-budget-unbounded-growth.md`, and - `REVIEW.md:145` specifies the deadline that was never built), so those entries - would be re-broadcast in every batch forever — the exact repeated timing signal - `chain.rs:517-521` invokes as the reason the `AlreadyKnown` branch exists. - Recommendations 1 and 2 are now explicitly coupled. -- **Recommendation 5 was corrected.** "Prefer the code over the text" as - originally written would be wrong: zebra returns `-1` (`LegacyCode::Misc`) for - *both* "never judged" (`Disabled`, `FullQueue`) and "already known" - (`InMempool`, `Mined`), so the code cannot replace the text match. It is a - useful outer guard, not a replacement. - -### Severity: why Medium, not High or Low - -**Not Low.** This is a real correctness defect in the requeue-vs-drop split — the -seam `chain.rs:479-490` calls out as the one that must be *"drawn honestly"* — -and it silently and permanently destroys migrations the wallet was already told -had succeeded, in the configuration the project ships as its example, with no -adversary present. The failure is per-node rather than per-transaction, so one -occurrence takes the whole batch, and the hub's own log (`rejected = N`) actively -misleads the operator about what happened. - -**Not High.** Four things bound it, and all four were checked rather than assumed: -(i) zebra's own 380 s RPC retry queue absorbs any restart that reaches the tip -inside ~6 minutes, which is most of them; (ii) `Disabled` is reachable only from -process start, not from a running node drifting behind, so the trigger is -restarts and first deployments rather than any lag; (iii) at present adoption the -modal batch is 0–1, so a single event destroys 0–2 migrations; (iv) for today's -traffic class the wallet's automatic-retry horizon is the ~50-minute ZIP 203 -expiry, not the 30–60 day ZIP 318 one. **NOTE, corrected 2026-08-18 (PROGRESS item -8a): this cuts the opposite way from how it was originally written.** The short -horizon is the *worse* one, so today's traffic is the harder case, not the easier -one; the deflation to Medium survives on (i)–(iii) alone, because zebra's 380 s -retry queue and the wallet's own ~20 s resubmissions both sit well inside 50 -minutes. The deliberate variant (an indexer -operator restarting their node at a cadence boundary) adds deniability but no -capability beyond the separately-filed hostile-indexer issue, so it must not be -counted twice. - -**Severity would rise to High** if any of these changed: adoption rises so a batch -holds tens of migrations, or the hub is pointed at an indexer whose node is -routinely far from the tip. **One escalation trigger was STRUCK 2026-08-18** — -*"ZIP 318 conforming wallets ship (the recovery clock becomes 30–60 days)"* rested -on a premise refuted during validation of -`issues/invalid/zip318-canonical-expiry-…md` (PROGRESS item 8a). ZIP 318's long -canonical expiry is a long **automatic-retry horizon**, not a long wait, so -conforming wallets shipping makes this issue's outcome *better*, not worse: the -wallet keeps resubmitting for 30–60 days instead of giving up after ~50 minutes. -Do not escalate on it. - -> **CORRECTION 2026-08-18 (validation of the cited file — SUPERSEDES the sentence above).** -> The wallet does **not** wait for expiry. Both official Zcash light-wallet SDKs -> automatically resubmit a sent-but-unmined transaction for as long as it remains -> unexpired — the Android SDK at the head of every ~20 s sync loop and after every -> processed block batch (`CompactBlockProcessor.kt:573,615,723`; selection -> `mined_height IS NULL AND expiry_height > ?`), the iOS SDK at most once per 300 s -> (`TxResubmitter.swift:8-15`, `TransactionDao.swift:218-228`) — and the hub's -> payload-hash dedup makes the resend free. The wallet's non-confirmation signal comes -> from compact-block scanning, which the shim does not intercept (`proxy.rs:1068-1074`). -> Expiry is therefore the **retry horizon**, not the wait: ~50 minutes for the ZIP 203 -> default traffic the shim also diverts, 30–60 days for a ZIP 318 migration. A -> *transient* loss self-heals within minutes; only a loss condition that **outlives the -> horizon** destroys the submission permanently — which is exactly what this issue's -> condition does, so this issue's severity is unaffected. Do not write "the user waits -> 30 to 60 days" in the report. Full refutation and the replacement paragraph: -> `issues/invalid/zip318-canonical-expiry-is-the-only-recovery-clock-and-a-lost-migration-freezes-the-users-notes-for-30-to-60-days.md`. - - -### Scope note - -Kept strictly to the zebrad-backed-indexer direction. The Zaino direction -(`hub-chain-zaino-node-rejections-are-never-verdicts.md`), the `"duplicate"` -over-match (`hub-chain-duplicate-nullifier-rejection-counted-as-published.md`), -and the hostile-indexer verdict question -(`hub-flush-destroys-migration-on-single-unverifiable-verdict.md`) are separate -issues with separate mechanisms; nothing here restates them. - -DO NOT DEVIATE FROM THIS TEMPLATE (`docs/SECURITY-ISSUE-TEMPLATE.md`). -LEAVE THIS MESSAGE IN PLACE SO THAT ALL AGENTS KNOW NOT TO DEVIATE FROM THIS TEMPLATE (`docs/SECURITY-ISSUE-TEMPLATE.md`). diff --git a/zeronym-22aa9851caf68-high-medium/medium/reproduce-gate-has-never-run-on-any-commit-that-published-a-hash.md b/zeronym-22aa9851caf68-high-medium/medium/reproduce-gate-has-never-run-on-any-commit-that-published-a-hash.md deleted file mode 100644 index 24e40b76..00000000 --- a/zeronym-22aa9851caf68-high-medium/medium/reproduce-gate-has-never-run-on-any-commit-that-published-a-hash.md +++ /dev/null @@ -1,530 +0,0 @@ -# The only automated gate on the attested binaries does not run on `push`, so on a branch that receives only direct pushes it runs only when a human dispatches it — neither hash published at HEAD has ever been checked, the last run of each workflow reported DOES NOT REPRODUCE, and the reference document teaches auditors that a red result is expected - -**Severity**: Medium -**Validation Status**: Confirmed -**Location**: `.github/workflows/zeronym-shim-reproduce.yml:24-38` and `.github/workflows/zeronym-hub-reproduce.yml:18-32` (the trigger sets, in the enclosing monorepo); `audit-target/zeronym/shim/deploy/EXPECTED_SHA256:1` and `audit-target/zeronym/hub/deploy/EXPECTED_SHA256:1`; `audit-target/zeronym/shim/deploy/reproduce.sh:80-120` (the verdict); `audit-target/zeronym/shim/deploy/README.md:36-46`, `:260-264`, `:283-285`, `:473-477`, `:1066-1090`; the shipped verification instruction at `audit-target/zeronym/hub/deploy/caution/assemble-caution.sh:503-519` and `audit-target/zeronym/shim/deploy/caution/assemble-caution.sh:585-604` (the `PROVENANCE` blocks); `audit-target/zeronym/deploy.sh` (no verification step anywhere); repository history and the GitHub Actions run history of both workflows -**Found by agent:** Local (file audit of `shim/deploy/README.md`); raised by the `hub/deploy/reproduce.sh` audit as BRAINSTORM §R26-F; owns coordinator open item 6e -**In scope of audit?** Yes. `*/deploy/**` including `EXPECTED_SHA256` and `reproduce.sh` is explicitly in scope, and `audit-context/AUDIT-INSTRUCTIONS.md` states that "the reproducible-build and attestation chain **is** the trust model here, so a break in it is a security finding, not tooling noise." Markdown claims are in scope "as security claims". The two workflow files sit one directory above `zeronym/` but are the enforcement point for the in-scope scripts, and `shim/deploy/assemble.sh:116` instructs maintainers to keep their filter list in sync with an in-scope file. - -> **FILENAME NOTE — read before quoting it.** This file's *name* encodes the -> original claim *"has never run on any commit that published a hash"*, which -> validation **refuted** (see "The claim that had to be corrected"). The name is -> retained only because `PROGRESS.md`, `BRAINSTORM.md` and two other issue files -> reference it. **Use the title, not the filename.** -> -> **SCOPE AND OWNERSHIP.** -> - The **`paths:` filter** on the `pull_request` trigger — the separate false-green -> hazard this same repository diagnosed on 2026-08-07 and removed from its two -> other PR workflows — is owned by -> `zeronym-reproduce-workflows-keep-the-paths-filter-false-green-the-project-diagnosed-and-removed-from-its-two-sibling-workflows.md` -> (Low). This file owns the **missing `push:` trigger**, the **re-baseline lag**, -> and the **document that sanctions the resulting red state**. Report them -> adjacently; do not count the filter twice. -> - The **fail-open verdict** (an empty or unreadable `EXPECTED_SHA256` yields -> `REPRODUCES` and exit 0) is owned by -> `reproduce-reports-reproduces-and-exits-zero-when-the-published-hash-comparison-is-skipped.md`. -> - The **hash appearing with three different values** across `EXPECTED_SHA256` and -> the two deploy READMEs is owned by -> `shim-published-binary-hash-has-three-different-values-across-expected-sha256-and-the-two-deploy-readmes.md`. - -## Description - -`reproduce.sh` is the project's only automated gate of any kind. Both zeronym -workflows are a single `run: sh .../reproduce.sh` step, and no `cargo test`, -`nextest`, `clippy`, `fmt`, `audit` or `deny` job exists anywhere in the -repository. `shim/deploy/README.md:39-46` describes that gate as the -"independent-second-machine half of the claim" and `:37` describes re-baselining -the published hash as "an explicit, reviewable edit". - -Four verified facts, taken together, say what the gate currently delivers: - -1. **Neither workflow runs on `push`.** - `.github/workflows/zeronym-shim-reproduce.yml:24-26` is `workflow_dispatch:` + - `pull_request:` only; the hub's `:18-20` is the same. Across the **entire run - history of both workflows — 51 shim runs and 18 hub runs — there is not one - `push` event.** -2. **`main` receives direct pushes, not pull requests.** The last 60 commits of - `main` contain zero merge commits from PRs and zero PR-numbered subjects; the - newest `Merge pull request` is 61 commits back (`b174d78`, 2026-08-13). The - consequence is visible in the run history: **every run of either workflow since - 2026-08-13 has been a manual `workflow_dispatch`**, with one exception (shim run - 44, a `pull_request` on a side branch, which failed). Before that date the - workflow did fire automatically on PRs and did confirm published hashes — see - the correction in Validation Information. -3. **`EXPECTED_SHA256` is re-baselined in its own commit, long after the source - change it describes.** Windows of up to **22 commits containing 16 successive - compiled-source changes** are on the record (`c68d320`, 2026-08-11 → - `0fda2c2`, 2026-08-12; independently recomputed during validation). During each - such window `sh zeronym/shim/deploy/reproduce.sh` prints - `zero-indexer-shim: DOES NOT REPRODUCE` for an entirely honest committed tree. -4. **Neither hash published at HEAD has ever been the subject of a run.** The last - run of the shim workflow on a `main` commit at or before HEAD was on - `e91170ed` (2026-08-17, runs 48 and 49) and **failed**; the last hub run on a - `main` commit at or before HEAD was on the same `e91170ed` (run 17) and - **failed**. The re-baselines that fixed those failures — `ff1fccd` (hub) and - `0be521a` (shim) — landed afterwards, and no run of either workflow has been - made on them or on any of the 14 commits between `e91170ed` and HEAD. - -**Mitigating fact, to be stated prominently wherever this is reported: at HEAD -neither hash is stale.** `git log 0be521a..HEAD -- ` and -`git log ff1fccd..HEAD -- ` are both empty, so no compiled -input has changed since either re-baseline. The defect is that the gate has not -been run to *confirm* that, not that a wrong hash is currently published. - -**The contradiction the file was assigned to resolve.** `shim/deploy/README.md` -states the same-commit rule four times (`:263-264`, `:285`, `:475-477`, `:1078`) -— *"`EXPECTED_SHA256` **must move in that same commit**"*, *"Those two must never -drift apart"* — and its normative three-state table names only working-tree -states, each resolved by a single commit. **The practice — source first, hash many -commits later — is none of those three states.** It is covered only by the loose -sentence that follows the table at `:1088-1090`: - -> A red `reproduce.sh` after a commit that touched compiled source is the -> tripwire working, and it has already caught one drifted hash on its first live -> test. - -That sentence is a conflation. The incident it cites (`:368-396`) is a case where -the published hash was **wrong** — a value that "never corresponded to any commit" -— not a case where it was merely **late**. Generalised to "red after a compiled -change is fine", it converts the gate's only alarm state into an expected one, in -the document an auditor is told to read. - -## Attack Scenario and Steps - -There is no single-step exploit. This is the degradation of the assurance -mechanism the product's trust model rests on, and it has two realistic -consequences. - -**Scenario A — a compiled change reaches a published hash with no independent -machine ever executing the gate.** - -1. A change to compiled source lands on `main` by direct push, the normal mode - (60 of the last 60 commits). No `pull_request` event fires; the workflow has no - `push` trigger; no run occurs. -2. The author later pushes a hash-only commit re-baselining `EXPECTED_SHA256` to - whatever their own machine produced (`3fcbaa9` and `0be521a` are exactly this - shape — one file changed each). Again no run occurs. -3. `deploy.sh` is then run. It performs no verification: `EXPECTED_SHA256`, - `build.sh`, `reproduce.sh` and `caution verify` appear in it only inside log - strings and a guard on `APP_SOURCE`, never as invocations. The enclave is built - and deployed from that tree. -4. The published hash for the running enclave has therefore been produced by one - machine, on a commit no automated job has ever seen, and re-checked by nobody. - The document that calls this "an explicit, reviewable edit" (`:37`) describes a - review that does not occur: no reviewer (direct push), no second machine (no - run), and for the shim's two most recent hashes no in-tree record either. -5. A modification introduced anywhere in the compiled input set at step 1 — in - particular in `zeronym/vendor/nym-upgrade-mode-check`, which is compiled into - both enclaves and is watched by no workflow, no `paths:` entry, no dirty-tree - check and no test — is absorbed into the next re-baseline as though intended. - -**Scenario B — an auditor cannot distinguish an honest late re-baseline from a -substitution.** - -1. An auditor follows the published provenance of a live enclave. Both - `assemble-caution.sh` scripts write a `PROVENANCE` file into the pushed app - repository pairing `source commit: $SHA` with `expected binary: $EXPECTED` - (read from the deploying tree's `EXPECTED_SHA256`) and instruct - `git checkout $SHA; sh zeronym//deploy/reproduce.sh`. -2. If the deploy happened inside a re-baseline window — a recurring state of - `main` lasting up to 22 commits — `$SHA` compiles a binary that is not - `$EXPECTED`, so the published instruction **fails on an honest deployment**. - The project recorded this as having already happened: commit `175f375` - (2026-08-13) says *"Until now the assembled `PROVENANCE` quoted the pre-change - hub hash and told a reader to verify with `reproduce.sh` — a claim that was - simply false."* -3. The auditor consults the reference document and reads `:1088-1090`: a red run - after a compiled-source commit "is the tripwire working". They are told the - alarm they just triggered is expected. -4. That is the reading an operator running a substituted binary needs them to - make. The gate's single output has been given two meanings that — in the - document's own phrase about a different pair of states (`:1070-1072`) — "look - identical from the outside and mean opposite things". - -**Attack Requirements and Assumptions:** -- Scenario A requires only the project's normal workflow; no attacker is needed - for the gate to be absent. An attacker *exploiting* it needs the ability to land - a compiled change — a compromised maintainer account, a malicious subtree pull, - or a change to the ungated vendored crate. That is a supply-chain position, not - a remote one. -- Honest bound on what the gate would catch even if it ran: `reproduce.sh` detects - **non-determinism and hash staleness**, not malice. A deterministic malicious - commit reproduces cleanly and is re-baselined normally. What the gate protects is - the *binding* between committed source and the published hash that auditors and - `caution verify` rely on — which is exactly the binding the trust model needs and - the only one anybody outside the project can check. -- Scenario B requires no attacker capability at all to *occur*; a malicious - operator merely benefits, because the honest-red state supplies cover. -- **As of HEAD neither hash is stale** (verified above), so no auditor is being - misled *today* by a wrong value. What is true today is that nobody has confirmed - it, and the last recorded evidence for both components is a failure. - -## Impact on Users - -Wallet users never check a hash. Auditors and third-party indexer operators do, -and they are the entire mechanism by which a user gets any assurance that the -enclave in front of their wallet runs reviewed code. `shim/deploy/README.md:19-27` -states this without hedging: "the deliverable here is not a Dockerfile that -builds. It is a hash anyone can independently recompute." - -What that mechanism delivers today: - -- **Both hashes published at HEAD have no confirmation of any kind.** No run has - been performed against either, and the last runs of both workflows failed. -- **The window in which the committed tree contradicts its own published hash is a - normal, recurring state of `main`** — up to 22 commits — with no bound and no - automatic alarm. -- **Both shipped deploy paths write a `PROVENANCE` file whose verification - instruction fails inside that window**, so the one artefact a third party is - handed to check a live enclave against can be wrong for benign reasons. -- **The document an auditor reads to interpret a red result tells them to expect - one.** - -Combined with the separately-filed fail-open verdict, the gate is weaker in both -directions at once: it goes green when it compared nothing, it goes red when -nothing is wrong, and on `main` it does not run unless a human remembers to press -it. Two of the project's own recorded incidents (2026-08-01, 2026-08-13) were each -caught by this gate and by nothing else, and the 2026-08-13 one could only have -been caught by a manual dispatch, since no automatic trigger existed for `main`. - -## Technical Details / Code Analysis - -**1. The trigger set.** `.github/workflows/zeronym-shim-reproduce.yml:24-38`: - -```yaml -on: - workflow_dispatch: - pull_request: - paths: - - ".github/workflows/zeronym-shim-reproduce.yml" - - "zeronym/shim/**" - - "zebra/Cargo.toml" - - "zebra/zebra-chain/**" - - "zebra/zebra-test/**" - - "zaino/Cargo.toml" - - "zaino/packages/zaino-proto/**" -``` - -There is no `push:`. The same file's header asserts the property the trigger set -cannot deliver for `main`: - -``` -# This runner is the INDEPENDENT SECOND MACHINE. … a matching hash on a native -# x86_64 runner is what upgrades the claim from "deterministic on one host" to -# "deterministic across hosts", which is the property the Auditor Role actually -# needs. -``` - -The hub's workflow is the same shape and was created on **2026-08-09**, two days -*after* the project removed the `paths:`-filtered `pull_request` construct from -`z3-smoke.yml` and `z3-regtest.yml` in `30d6852` (2026-08-07) with a long in-file -write-up of why it is unsafe. Neither zeronym workflow's `on:` block has been -modified since it was written. *(That construct is the sibling issue's subject; -noted here only because it compounds the same blind spot.)* - -**2. The run census.** Retrieved from the GitHub Actions API for both workflows, -complete history: - -| workflow | total runs | `push` events | runs since 2026-08-13 | of which manual | -|---|---|---|---|---| -| `zeronym-shim-reproduce` | 51 | **0** | 15 | 14 (1 `pull_request` on a side branch, failed) | -| `zeronym-hub-reproduce` | 18 | **0** | 13 | 13 | - -The most recent runs bearing on the audited HEAD: - -| workflow | run | date | event | head | conclusion | -|---|---|---|---|---|---| -| shim | 49 | 2026-08-17 15:55Z | `workflow_dispatch` | `e91170ed` | **failure** | -| shim | 48 | 2026-08-17 15:48Z | `workflow_dispatch` | `e91170ed` | **failure** | -| hub | 17 | 2026-08-17 15:48Z | `workflow_dispatch` | `e91170ed` | **failure** | - -and there are **14 commits between `e91170ed` and HEAD**, including both -re-baselines, with no run on any of them. - -**3. The verdict a stale window produces.** `shim/deploy/reproduce.sh:88-120`: - -```sh -if [ -z "$EXPECTED" ]; then - echo "NOTE: no published hash to compare against (EXPECTED_SHA256 missing" - echo " or explicitly cleared). Self-consistency only." -elif [ "$h1" = "$EXPECTED" ]; then - echo "MATCHES PUBLISHED: $EXPECTED" -else - echo "FAIL: this host disagrees with the PUBLISHED hash." - … - fail=1 -fi -… -if [ "$fail" = 0 ]; then - echo "zero-indexer-shim: REPRODUCES" -else - echo "zero-indexer-shim: DOES NOT REPRODUCE (see FAIL lines above)" -fi -``` - -At `e91170e` the shim's `EXPECTED_SHA256` still held `2199d281…` while `proxy.rs` -had moved, so this printed `FAIL: this host disagrees with the PUBLISHED hash` and -exited 1 — which is exactly what runs 48 and 49 recorded. The same output is what -an auditor would get from a substituted binary. - -**4. The re-baseline history.** Every transition of both `EXPECTED_SHA256` files -was extracted with `git log --full-history --format='%H' --reverse main -- ` -(plain `git log -- ` prunes side-branch commits and undercounts). The -same-commit rule was kept **five** times — `1c616c3` (file creation), `c161012`, -`d8306d7`, `3357bd1` on the shim and `a746496` on the hub — the last of them on -**2026-08-11**. Since then, **all fifteen later transitions are late** (8 shim + 7 -hub, across 10 distinct commits). The worst window, recomputed independently during -validation: `c68d320` (2026-08-11) → `0fda2c2` (2026-08-12) is **22 commits -containing 16 compiled-source changes to the shim**, and the hub's equivalent is -23 commits. `0be521a`'s four-commit window is one of the *smaller* ones. - -**5. The provenance consequence.** `hub/deploy/caution/assemble-caution.sh:503-519` -(the shim's copy at `:585-604` is the same block): - -```sh -EXPECTED=$(cat "$ZERO_ROOT/zeronym/hub/deploy/EXPECTED_SHA256" 2>/dev/null || echo "unrecorded") -cat > "$DEST/PROVENANCE" < **Not yet verifiable.** The enclaves are attested and running, but an auditor - > cannot yet tie either back to a public commit. PCR2 (the application binary) - > reproduces; PCR0 and PCR1 do not, so `caution verify` reports FAILED on healthy - > enclaves. The reproduce jobs run on pull requests and manual dispatch, **not on - > every push**; they last ran 2026-08-17 on `e91170ed` and both reported DOES NOT - > REPRODUCE, with re-baselines landed since and no run against them. The live - > pair's provenance also fails: the shim's cites a non-public commit, the hub's a - > hash its own cited commit does not produce. - -- `fa17e92` (12:49:17 -0400, *"Protected as bullets; drop the verifiability - paragraph"*) removed it 26 minutes later, together with the two cross-references - pointing at it. **The removal commit's own message states what is being removed:** - *"Note this removes the README's only disclosure that PCR0/PCR1 do not reproduce, - that the reproduce jobs do not run on push, and that the live pair's provenance - does not check out."* -- **One sentence in that paragraph was genuinely stale**, which is a sufficient and - innocent explanation for removing the paragraph rather than editing it: the - PCR0/PCR1 claim is superseded by `shim/deploy/caution/OPERATORS.md:188-193` - (*"**That is fixed**: on the attested pair deployed 2026-08-14 … **all three PCRs - reproduced** on both"*). -- **The two other sentences were accurate and remain accurate at HEAD**, and this - audit verified both independently against the Actions API rather than taking the - project's word: no `push` trigger and no `push` run in 69 runs; and runs 48/49 - (shim) and 17 (hub) on `e91170ed` all reported failure with the re-baselines - landing afterwards and no run against them. -- Nothing at HEAD replaces the statement: a grep of every `.md` under `zeronym/` - for "not on every push", "DOES NOT REPRODUCE", "last ran", "unverified" or "not - yet verifiable" returns nothing, while `zeronym/README.md`'s auditor bullet still - says "reproduce the build and compare hashes" with no caveat. - -**This finding does not depend on the deletion.** Facts 1-4 in the Description are -established from the workflow files, the commit history and the Actions API alone. -The deletion is reported because it is the only place the project has ever -described the gap, and because its absence leaves the public claim sheet asserting -an auditor procedure with no statement of what that procedure currently -establishes. - -## Recommendations - -1. **Add `push:` (restricted to `main`) to both reproduce workflows.** This is the - single change that turns the window from unobserved into observed, and it costs - one line per workflow. Both historical failures of this control were detected by - `reproduce.sh` and by nothing else, and the 2026-08-13 one required a human to - dispatch it. Do this together with the sibling issue's recommendation to move - path filtering into a job, so a dropped event cannot be silent either. -2. **Make the re-baseline atomic again.** The document already derives the correct - procedure at `:508-513`: measure from a working-tree overlay, commit source + - `EXPECTED_SHA256` + the README row together, then re-run `reproduce.sh` against - the commit with a live comparison. `c161012`, `d8306d7`, `3357bd1` and `a746496` - show it is achievable; the last time it was done was 2026-08-11. -3. **Delete or narrow `shim/deploy/README.md:1088-1090`.** As written it tells - auditors that the gate's failure output is expected. Replace it with the precise - statement: *a red `reproduce.sh` means the published hash does not describe this - tree; there is no benign case, and if you see one it is a defect in our release - process, not in your build.* -4. **Restore an accurate status statement to `zeronym/README.md`.** The paragraph - removed by `fa17e92` contained one stale sentence (PCR0/PCR1) and two that - remain true at HEAD (the push-trigger gap; the current hashes having had no run). - Correct the stale sentence and keep the rest, rather than leaving the public - claim sheet with no statement of what its auditor procedure currently - establishes. -5. **Have `assemble-caution.sh` refuse to write a `PROVENANCE` block pairing a - `source commit` with an `expected binary` it has not verified against that - commit** — or, at minimum, have it print a loud warning when the tree has - compiled-input commits newer than the last `EXPECTED_SHA256` change. Having - `deploy.sh` gate on a passing `reproduce.sh` is the stronger version of the same - fix. - -## Validation Information - -**Verdict: CONFIRMED, Medium. The mechanical facts hold and are now much stronger -than as filed, but the headline claim as originally written is FALSE and has been -replaced. Coordinator open item 6e is closed by this file.** - -### The claim that had to be corrected - -The original title and Description asserted that the gate *"has never executed on -any commit that published a hash"*. **That is false, and would have been a false -positive in the report.** Retrieved from the GitHub Actions API: - -- shim run **34**, 2026-08-09, event `pull_request`, head `a77d9f8fb1`, - conclusion **success** — and `a77d9f8fb1`'s `zeronym/shim/deploy/EXPECTED_SHA256` - is `51ccefed3eda14a5…`, the value that then landed on `main` in `d8306d7`. -- shim run **36**, 2026-08-11, event `pull_request`, head `2cf872d76b`, - conclusion **success** — carrying `f498f82240711872…`, the value that landed in - `3357bd1`. - -`shim/deploy/README.md:293` corroborates this from the project's side -(*"the first shim hash since `51ccefed` to be machine-checked at all"*), as does -the `51ccefed` row's *"Cross-machine confirmed: a native x86_64 CI runner and a -local arm64 build under Rosetta agree."* So the gate did work automatically, on -PR heads, through 2026-08-11. - -The corrected and verified claim, which is what the file now says, is narrower and -still serious: **the gate has no `push` trigger; `main` has taken only direct -pushes since 2026-08-13; consequently every run of either workflow since that date -has been manual; and neither hash published at the audited HEAD has ever been the -subject of a run, with the last run of each workflow reporting failure.** - -### Evidence obtained during validation that the filing did not have - -The unauthenticated GitHub Actions API was queried for the complete run history of -both workflows (`/repos/ShieldedLabs/zero/actions/workflows/{324882157,329724684}/runs`). -This replaces the filing's inference with direct evidence: - -- 51 shim runs and 18 hub runs, **zero `push` events** in either. -- Events by era: `pull_request` dominates through 2026-08-11; from 2026-08-13 - onward all but one run is `workflow_dispatch`. -- shim runs 48/49 and hub run 17, all on `e91170ed` (2026-08-17), all - **failure** — an exact, independent confirmation of the sentence the project - published and then removed. -- No run exists on `ff1fccd`, `0be521a`, or any of the 14 commits between - `e91170ed` and the audited HEAD `62baea8`. - -*(Note for reproduction: upstream `main` has advanced past the audited snapshot. -Shim runs 50/51 and hub run 18 are on commits **after** `62baea8` and are outside -this audit's scope; run 51 succeeded on a post-audit commit. The audit target was -confirmed byte-identical to `audit-context/zero` at `62baea8`.)* - -### Facts re-verified against the target and the repository - -- Trigger sets: `zeronym-shim-reproduce.yml:24-38` and - `zeronym-hub-reproduce.yml:18-32` — `workflow_dispatch` + `pull_request` with a - `paths:` list, no `push:`. Neither `on:` block has been edited since it was - written (`git log -S'pull_request'` on the shim file returns only `1c616c3`, - 2026-07-31). -- `git log --format='%h %s' -60 main | grep -cE 'Merge pull request|\(#[0-9]+\)'` - → **0**; the newest PR merge `b174d78` sits at position **61**. One non-PR merge - (`d1dc077`) is inside the window and fires no `pull_request` event either. -- Re-baseline positions in `main`: `0be521a` (11), `ff1fccd` (12), `3fcbaa9` (17), - `bfa9ad7` (30), `68a1f6f` (36), `3c52859` (42), `175f375` (60) — all after - `b174d78`, therefore all direct pushes. -- Worst window recomputed from scratch: `c68d320`→`0fda2c2` = **22 commits, 16 - compiled-source changes**. The addendum's table is sound; "up to 23 commits and - 16 successive compiled-source changes" is the correct statement, not "one to - four". -- `git log 0be521a..62baea8 -- ` → **0 commits**; - `git log ff1fccd..62baea8 -- ` → **0 commits**. - **Neither hash is stale at HEAD.** -- `shim/deploy/README.md`: `:36-46` (the ledger of what each artefact proves), - `:260-264`, `:283-285`, `:473-477`, `:1066-1090` (the three-state table and the - "tripwire working" sentence) — all quoted verbatim and confirmed. -- `shim/deploy/reproduce.sh:80-120` — the verdict logic as quoted. -- `deploy.sh` — `grep -n 'EXPECTED_SHA256\|reproduce.sh\|caution verify\|build.sh'` - returns four hits, all inside log strings or an `APP_SOURCE` guard; none is an - invocation. The "no verification in `deploy.sh`" claim holds. -- `PROVENANCE` blocks at `hub/deploy/caution/assemble-caution.sh:503-519` and - `shim/deploy/caution/assemble-caution.sh:585-604` — as quoted, including the - `unrecorded` fallback. -- The commit texts of `105c21d2`, `aad6523e`, `175f375`, `53bab86` and `fa17e92` - — all read directly from `git log`/`git show` and quoted accurately. -- `OPERATORS.md:188-193` — the 2026-08-14 all-three-PCRs-reproduced measurement - that makes the deleted paragraph's PCR sentence stale. - -### Two further claims in the filed text that were corrected - -1. *"`shim/deploy/README.md:662` still says '**Nothing about the build is - outstanding**'"* — **misleading as cited.** Line 662 sits inside - `### Re-baseline, 2026-08-01 (second): the predicate widened`, a dated - historical write-up of one re-baseline, not a present-tense status claim. The - sentence has been removed from this issue; the same passage's *"no PCR has been - computed from this binary and no attestation document exists"* is stale in a - different way and is owned by - `deploy-readme-says-no-enclave-no-eif-and-no-pcr-exist-which-its-own-current-row-refutes.md`. -2. The `paths:` filter was argued inside this issue's Recommendation 1. It is kept - only as a cross-reference; the finding belongs to the sibling Low. - -### Handling of the deleted paragraph (coordinator item 6h) - -Reported above as a **dated factual sequence with no motive imputed**, and three -facts are stated alongside it so it cannot be read as an accusation: `53bab86` -*softened* an already-stronger claim rather than being the first disclosure; the -paragraph contained one genuinely stale sentence (PCR0/PCR1), which fully explains -removing the paragraph rather than editing it; and `fa17e92`'s own commit message -explicitly enumerates the disclosure being lost. **The finding does not rest on the -deletion** — every load-bearing fact is established from the workflow files, the -commit graph and the Actions API. Recommendation 4 asks for the accurate parts to -be restored, which is a documentation request, not an allegation. - -### Why Medium - -- **Not High.** No user's funds or privacy is directly lost; exploitation of - Scenario A requires an already-privileged supply-chain position; the control - demonstrably works when it is run (it caught both recorded incidents); the gate - detects staleness and non-determinism rather than malice; and at HEAD neither - hash is actually stale. -- **Not Low.** This is the project's only automated gate on the artefacts the whole - trust model names, and `AUDIT-INSTRUCTIONS.md` states that a break in that chain - is a security finding rather than tooling noise. Both published hashes are - currently unconfirmed and the last recorded evidence for each is a failure; the - stale window is a recurring, unbounded state of `main` up to 22 commits long; the - shipped `PROVENANCE` artefact hands third parties an instruction that fails - inside it; and the reference document teaches the reader to dismiss the gate's - only alarm. The affected population — auditors and third-party operators — is the - entire mechanism by which an end user gets any assurance at all. -- **Remediation is cheap and disproportionately effective:** recommendation 1 is - one line per workflow. - -DO NOT DEVIATE FROM THIS TEMPLATE (`docs/SECURITY-ISSUE-TEMPLATE.md`). -LEAVE THIS MESSAGE IN PLACE SO THAT ALL AGENTS KNOW NOT TO DEVIATE FROM THIS TEMPLATE (`docs/SECURITY-ISSUE-TEMPLATE.md`). diff --git a/zeronym-22aa9851caf68-high-medium/medium/shim-config-no-fail-closed-mode-empty-hub-nym-silently-selects-no-privacy.md b/zeronym-22aa9851caf68-high-medium/medium/shim-config-no-fail-closed-mode-empty-hub-nym-silently-selects-no-privacy.md deleted file mode 100644 index 6ad94f18..00000000 --- a/zeronym-22aa9851caf68-high-medium/medium/shim-config-no-fail-closed-mode-empty-hub-nym-silently-selects-no-privacy.md +++ /dev/null @@ -1,439 +0,0 @@ -# The shim's configuration cannot express "diversion is required": an unset, empty or whitespace-only `ZIS_HUB_NYM` resolves to no-privacy forward-only mode and the process serves wallets normally instead of refusing to start - -**Severity**: Medium -**Validation Status**: Confirmed -**Location**: `audit-target/zeronym/shim/src/config.rs:50-77` (the two hub fields), `:199-213` (`HubSelection`), `:250-289` (`Config::hub_selection`), `:420-457` (the tests that pin the behaviour); consumers at `audit-target/zeronym/shim/src/main.rs:112-143`, `audit-target/zeronym/shim/src/nym.rs:135-136`, `:180-183`, `:305-309`, `audit-target/zeronym/shim/src/intercept.rs:117-124`, `audit-target/zeronym/shim/src/proxy.rs:643-653`; the design intent at `audit-target/zeronym/shim/deploy/caution/OPERATORS.md:10-16` and `:242`; the claim it bears on at `audit-target/zeronym/README.md:26-28` -**Found by agent:** Local (file audit of `shim/src/config.rs`); validated 2026-08-18 -**In scope of audit?** Yes — `audit-context/AUDIT-INSTRUCTIONS.md` names operator-supplied configuration and environment (`ZIS_*`) as an explicit trust boundary and the operator as the primary adversary; priority area #5 (fail-closed discipline) - -> **Note on the filename.** This file keeps its original name so the ten -> cross-references elsewhere in `audit-state/` stay valid. The word "silently" in -> the filename is a legacy artefact and is **wrong** — see "What is not true" -> below. The title above is the corrected statement of the finding. - -## Description - -`Config::hub_selection` resolves the transport into a closed three-member set -(`config.rs:205-213`): - -```rust -pub enum HubSelection { - /// No hub: classify and log, forward everything. No privacy. - ForwardOnly, - /// The transitional clearnet path. - Http(SocketAddr), - /// The mixnet path, over one or more gateway-bound hub addresses. - Nym(Vec), -} -``` - -There is no fourth member and no fourth flag. Nothing in the crate lets an -operator say *"this shim is supposed to divert; if it cannot, refuse to serve."* -There is no `--require-diversion`, no `--strict`, no equivalent knob anywhere in -the shim. So `ForwardOnly` — the mode in which every Orchard-touching transaction -is handed straight to the operator's indexer — is not one of three options an -operator selects. It is the value the resolver produces whenever the other two are -not affirmatively and correctly present. - -Two properties compound that: - -1. **Degenerate values are indistinguishable from absence.** `hub_selection` trims - every `--hub-nym` entry and drops the empty ones *before* deciding - (`config.rs:262-267`), so `ZIS_HUB_NYM` unset, `ZIS_HUB_NYM=`, - `ZIS_HUB_NYM=" "` and `ZIS_HUB_NYM=" , ,"` all yield `ForwardOnly`. Two unit - tests pin exactly this as intended behaviour (`config.rs:420-446`). -2. **The process then reports healthy.** `MixnetStatus::is_healthy` - (`nym.rs:305-309`) is `!configured || connected`, and `configured` is set from - exactly one place — `nym_driver.rs:188`, inside the mixnet driver task, which - `main.rs` spawns only on the `HubSelection::Nym` arm. A forward-only shim never - sets it, so `/healthz` answers `200 ok` for the life of the process - (`proxy.rs:643-653` documents this as deliberate). - -Note the asymmetry with the sibling transport, which is the sharpest evidence that -this is an oversight rather than a decision. `ZIS_HUB=""` is a **hard startup -error**: clap 4.6.5 applies an environment variable that is present even when its -value is empty, so `""` reaches the `SocketAddr` parser and fails. The same -templating accident therefore refuses to boot on the clearnet transport and -degrades to no-privacy on the mixnet one — and the mixnet one is what -`deploy.env.example:18` ships and what `OPERATORS.md:20` tells operators to use. - -The project's own adversarial review demanded the opposite discipline for the -analogous case. `hub/REVIEW.md` #11 requires the shim's last-resort direct -broadcast to be *"off by default, behind an explicitly named config flag"*, -reasoning that *"nearly every attack in this set converges on the same final step: -get the shim to broadcast directly. That is only possible because the shim fails -OPEN on privacy. Making it fail closed removes the whole class in one change."* -The shim has no direct-broadcast path at all, so #11 is satisfied by absence — but -`ForwardOnly` is reached the same way #11 was worried about: by omission rather -than by an explicitly named act. - -### What is not true — four corrections that bound this finding - -These were checked during validation and each one refutes part of the original -framing. They are recorded here so the report does not overstate the issue. - -1. **The state is not undetectable.** `/nym-status` is served on the shim's - wallet-facing listener, unauthenticated, to anyone - (`proxy.rs:654`, `nym.rs:207-211`), and reports `diversion_configured: false` - for a forward-only shim, permanently. `shim/deploy/caution/OPERATORS.md:415` - documents that exact field as an alert condition: *"`false` on a shim you built - with `--hub-nym` — it is forward-only and hiding nothing"*, and - `smoke.sh:331-347` already implements the check. The signal is one `curl` away - for any wallet author, monitor or passer-by. (Its one real weakness: the flag is - mixnet-only, so a *clearnet-hop* shim also reports `false`. Since the clearnet - hop is documented as legacy and non-functional against the current hub - (`OPERATORS.md:18-21`), that ambiguity is not load-bearing today.) -2. **The attestation chain is not blind to it.** Coordinator open item 6q - established from the Caution platform source that `unit.env` **is** measured - into PCR0/PCR1, and `ZIS_HUB_NYM` is rendered into the manifest's env block - (`assemble-caution.sh:352`, `caution.hcl.tmpl:161`). So a forward-only shim has - different PCRs from a diverting one, and the whole environment is additionally - served unauthenticated at `.manifest.run_command` of every `/attestation` - response. Anyone can read whether a hub is configured, and which one. -3. **The deploy tooling does not select it silently.** - `assemble-caution.sh:414` prints `"==> forward-only: no --hub, so migrations are - forwarded to the operator's indexer (no privacy)."` on the operator's terminal, - and the generated `PROVENANCE` file records `diversion: OFF (forward-only, no - privacy)`. Moreover the assembler **hard-refuses** the two accidental routes a - `deploy.sh` operator could take, because `deploy.env.example:63` ships - `NYM_EGRESS` non-empty and `deploy.sh:125-127` passes it unconditionally: - `assemble-caution.sh:189-192` exits 2 on `--nym-egress` without `--hub-nym`, and - the structural check at `:226-235` (`case "$addr" in ?*.?*@?*)`) exits 2 on a - whitespace-only address before the binary ever sees it. -4. **Forward-only is an intentional, documented product phase, not an oversight.** - `shim/deploy/caution/OPERATORS.md:10-16` calls it *"Phase 1 … it classifies and - logs but forwards everything"* and `:242` states *"Forward-only stays the - default: no `--hub-nym`, no diversion."* The mode's existence is by design - (`docs/AVOIDING-FALSE-POSITIVES.md` §6). **The defect is the absence of a way to - opt out of it**, not its existence. - -## Attack Scenario and Steps - -**Path A — the deceptive operator.** Cheap, but detectable. - -1. The operator deploys the real, attested, reproducible shim image in front of - their indexer and advertises the endpoint as running zero-indexer. -2. They omit `HUB_NYM` and `NYM_EGRESS` (or invoke `assemble-caution.sh` by hand - without `--hub-nym`, which is the interface `OPERATORS.md:77-87` documents). - `hub_selection` returns `ForwardOnly`. -3. Every wallet pointed at the endpoint has each Orchard-touching transaction - classified, logged, and then **handed to the operator's own indexer** - (`intercept.rs:117-124`). -4. The operator joins the wallet's TCP source IP to the transaction bytes and, - once it lands on chain, to the value moved — the exact linkage the product - exists to prevent, permanent because the chain is permanent. -5. `/healthz` answers 200 throughout, and wallets, which per `README.md:68` - "install nothing and change no setting", do not check anything. - -**But this path is observable.** `/nym-status` reports `diversion_configured: -false` and `/attestation`'s `.manifest.run_command` shows no `ZIS_HUB_NYM`. What -is missing is not the signal but any documented instruction to read it: the -auditor recipe at `README.md:71` lists four steps (attestation, PCRs, reproduce, -Certificate Transparency) and none of them observes this state. So Path A works -only against a counterparty who never looks — which, absent a wallet-side check, -is every ordinary user. - -**Path B — the configuration accident. This is the path the fix addresses.** - -1. An operator runs the shim outside `deploy.sh` — compose, systemd, Kubernetes, a - hand-written `caution.hcl`, or their own orchestration — with - `ZIS_HUB_NYM=${HUB_NYM}` in an env file, and `HUB_NYM` renders empty. This is - not hypothetical: it is the documented motivation for the filtering - (`config.rs:253-261` — *"an existing clearnet deployment that templates the new - variable in as empty would stop booting"*), and the hub's Nym address must be - re-read and re-templated on every hub restart because a diskless enclave mints a - new one (`OPERATORS.md:369-377`). -2. `config.rs:262-267` drops the empty entry; `hub_selection` returns - `ForwardOnly`. **Startup succeeds.** -3. `main.rs:137-140` emits one `tracing::warn!` at process start. In an attested - enclave that line goes to a console the parent host cannot read (coordinator - open item 7), so in the deployment the product recommends, nobody sees it. -4. Every subsequent runtime signal says the component is working: `/healthz` is - 200, wallets get normal `SendResponse`s from the operator's indexer, and the - per-transaction log line asserts `diverted_in_production=true` for a transaction - it forwarded (filed separately as - `forward-only-log-claims-migration-was-diverted.md`). - -The only durable signal is `/nym-status`, which nothing in this path polls. - -**Attack Requirements and Assumptions:** - -- Path A requires only that the operator omit environment variables. No code - modification is involved, so the *binary* is the audited one — but the PCRs and - the disclosed manifest do change, so the deployment is not indistinguishable - from a diverting one to anyone who reads them. -- Path B requires no attacker at all, but does require a deployment that bypasses - `deploy.sh`/`assemble-caution.sh`, because the assembler hard-refuses the two - routes a template-following operator could take. -- Neither path can be triggered *at runtime* by an outsider: `diversion` is decided - once in `main.rs:112-143` and threaded immutably; no hub refusal, mixnet death, - config reload or attacker input flips it (verified by the G1 global audit, - coordinator open item 6r). Every reachable runtime degradation is toward - fail-closed. - -## Impact on Users - -A wallet user of a forward-only endpoint receives none of the product's privacy -properties while the process, the operator's own liveness monitoring, and the -wallet's experience all report normal operation. `README.md:26-28` states the two -headline protections without condition: - -> - **Broadcast contents.** An Orchard-touching transaction is hidden from the operator… -> - **Source IP.** The on-chain transaction carries no link to the wallet's IP… - -Neither holds in `ForwardOnly`. The operator receives the full transaction bytes -and the wallet's source IP together, and because the transaction lands on the -public chain and pool crossings reveal value in cleartext, they can join -IP → transaction → amount, retrospectively and forever. - -Under the ICTM methodology this engagement uses, a user being told they have a -property they do not have is itself the bug — and the same document that states -the protections unconditionally also says, five lines later, that operators "run -the shim in front of their indexer, **and optionally a hub**" (`README.md:70`). -That broader documentation gap is filed separately as -`readme-promises-protection-that-no-user-wallet-or-passerby-can-verify-is-switched-on.md`; -what belongs to *this* issue is narrower and mechanical: **the binary offers no -setting under which a hub-less start is an error**, so an operator who wants to -guarantee the property for their users cannot ask the software to enforce it. - -## Technical Details / Code Analysis - -The resolver, complete (`shim/src/config.rs:250-289`): - -```rust -impl Config { - /// Resolve the configured transport, rejecting anything ambiguous. - pub fn hub_selection(&self) -> Result { - // Empty entries are dropped, not diagnosed. `ZIS_HUB_NYM=` reaches - // clap as one EMPTY value rather than as no value at all, because with - // a delimiter clap splits whatever the variable holds and an unset - // variable is not the same thing as an empty one. Without this an - // existing clearnet deployment that templates the new variable in as - // empty would stop booting, either because both transports look set or - // because "" looks like a malformed address. ... - let addresses: Vec<&str> = self - .hub_nym - .iter() - .map(|addr| addr.trim()) - .filter(|addr| !addr.is_empty()) - .collect(); - - match (self.hub, addresses.is_empty()) { - (Some(_), false) => Err(ConfigError::BothTransports), - (Some(addr), true) => Ok(HubSelection::Http(addr)), - (None, true) => Ok(HubSelection::ForwardOnly), - (None, false) => { /* structural + duplicate checks, then Nym(..) */ } - } - } -} -``` - -Everything the resolver *does* check is correct and fails closed, which is worth -stating so the finding is not read more broadly than it is: - -- both transports set → `Err(BothTransports)` → `main.rs:112` propagates → the - process exits (`assemble-caution.sh:185-188` rejects the same combination one - layer earlier); -- a malformed non-empty Nym entry → `Err(MalformedNymAddress)` → exit - (`config.rs:276-278`), and `main.rs` then re-parses every surviving address - through the SDK's authoritative `Recipient` parser, still at startup; -- a duplicate entry → `Err(DuplicateNymAddress)` → exit (`config.rs:279-281`); -- a malformed `ZIS_HUB` → clap `SocketAddr` parse error → exit. - -The **only** unguarded cell in that matrix is "neither transport", and it is the -cell that turns the product's privacy off. - -The two tests that pin the degenerate inputs as intended -(`shim/src/config.rs:420-446`): - -```rust - #[test] - fn an_empty_hub_nym_is_the_same_as_an_unset_one() { - let empty = parse(&["--hub-nym", ""]); - assert_eq!(empty.hub_nym, vec![String::new()], - "the field really does hold one empty entry"); - assert_eq!(empty.hub_selection().unwrap(), HubSelection::ForwardOnly); - ... - // Whitespace and stray separators are empty too. - assert_eq!(parse(&["--hub-nym", " , ,"]).hub_selection().unwrap(), - HubSelection::ForwardOnly); - } -``` - -What forward-only actually does with a migration -(`shim/src/intercept.rs:117-124`): - -```rust -117 if inspection.treat_as_migration() { -118 if let Some(diversion) = diversion { -119 return divert(&diversion, tx_data).await; -120 } -121 // Forward-only: no hub configured, so behave exactly like the merged -122 // proof of concept and forward the migration to the operator. No -123 // privacy, but no behaviour change until an operator sets `--hub`. -124 } -``` - -Why the health surface cannot show it (`shim/src/nym.rs:305-309`): - -```rust - /// Whether the shim can currently carry a migration: either diversion is not - /// configured at all (forward-only, nothing to be down), or the client is up. - pub fn is_healthy(&self) -> bool { - !self.0.configured.load(Ordering::Relaxed) || self.0.connected.load(Ordering::Relaxed) - } -``` - -`configured` is written from a single site, `shim/src/nym_driver.rs:188`, inside -`run_driver`, which `main.rs` spawns only on the `HubSelection::Nym` arm. -`main.rs:98-100` acknowledges the consequence — "on a forward-only or **clearnet** -shim nothing ever writes, and it honestly reports 'not configured'". - -And the deploy wrapper contributes no guard of its own (`deploy.sh:113`): - -```sh -[ -n "${HUB_NYM:-}" ] && set -- "$@" --hub-nym "$HUB_NYM" -``` - -`HUB_NYM` is the only shim input with no `: "${VAR:?…}"` guard, and this AND-OR -list does **not** abort under `set -eu` (verified by execution under `sh`, `dash` -and `bash`: a failing non-final command of an AND-OR list is exempt from `-e`). -The flag is simply omitted. That gap is filed separately as -`deploy-gates-the-hubs-mixnet-readiness-and-never-the-shims-so-a-no-privacy-shim-deploys-clean.md`; -it is noted here only because it is the mechanism by which an empty `HUB_NYM` -reaches `hub_selection` at all. - -## Recommendations - -1. **Add an explicit fail-closed mode and make it the documented deployment.** A - `--require-diversion` / `ZIS_REQUIRE_DIVERSION=true` that turns - `HubSelection::ForwardOnly` into a startup error is a few lines in - `hub_selection` and removes the entire class. This is the shape - `hub/REVIEW.md` #11 already asked for, applied to the fail-open that actually - exists. Because the flag would live in `unit.env`, it would also be **measured - into the PCRs and disclosed in `.manifest.run_command`** — turning "this - operator promised to divert" into an attestable fact rather than a claim. -2. **Keep the empty-string accommodation, but make it separable.** The rationale - at `config.rs:253-261` is sound for backwards compatibility, but "the operator - set the variable and it was empty" and "the operator never set the variable" - are distinguishable states. Reject the empty value outright when the require - flag is set, and warn differently in the two cases otherwise. -3. **Report the transport, not a mixnet-only boolean.** Replace - `diversion_configured: bool` with `"transport": "forward-only" | "clearnet" | - "mixnet"`, set on every arm of `main.rs:113-143` rather than only inside the - mixnet driver. That makes `FORWARD-ONLY` and `CLEARNET-HOP` distinguishable to - anyone, which they are not today. -4. **Wire the check that already exists into the deploy path.** `smoke.sh:331-347` - implements exactly this assertion and nothing in the repository invokes or even - mentions it — `smoke.sh` is referenced by no runbook, no README and not by - `deploy.sh`. Have `deploy.sh` run it after a shim deploy, and cite it in - `OPERATORS.md`. -5. **Add "check that a hub is configured, and which one" to the auditor recipe at - `README.md:71`**, naming `/nym-status` and `.manifest.run_command`. This is the - cheapest of the five and closes Path A entirely for anyone who follows it. - -Cross-references: `forward-only-log-claims-migration-was-diverted.md` (the -per-transaction log line asserts the opposite outcome); -`deploy-gates-the-hubs-mixnet-readiness-and-never-the-shims-so-a-no-privacy-shim-deploys-clean.md` -(the missing `deploy.sh` guard); -`readme-promises-protection-that-no-user-wallet-or-passerby-can-verify-is-switched-on.md` -(the user-facing verification gap); -`shim-config-hub-identity-is-unattested-unobservable-operator-configuration.md` -(the variant where a hub *is* configured and it is the operator's own); -`THREATMODEL.md` scenario `FORWARD-ONLY`. - -## Validation Information - -**Verdict: CONFIRMED as a real defect. Severity: Medium — downgraded from the -filed High.** - -**What was verified true** (all read directly from the target): - -| Claim | Verified at | -|---|---| -| `HubSelection` has exactly three members and no "require" flag | `config.rs:205-213`; no `require`/`strict` option anywhere in `shim/src/` | -| Unset, `""`, `" "` and `" , ,"` all resolve to `ForwardOnly` | `config.rs:262-272`; tests at `:420-446` | -| Every *other* cell of the matrix fails closed at startup | `config.rs:274-286`; `main.rs:112` | -| `is_healthy` keeps a forward-only shim at HTTP 200 forever | `nym.rs:305-309`; `proxy.rs:649-653` | -| `configured` is written only from the mixnet driver | `nym.rs:182-183`; single call site `nym_driver.rs:188`; spawned only on the `Nym` arm | -| Forward-only forwards the migration to the operator's indexer | `intercept.rs:117-131` | -| `ZIS_HUB=""` is a hard startup error (the asymmetry) | `config.rs:55-56`; clap 4.6.5 (`shim/Cargo.lock:1113-1114`) applies a present-but-empty env var to the `SocketAddr` parser | -| `deploy.sh:113` has no `:?` guard and does not abort under `set -eu` | executed under `sh`, `dash` and `bash`; all three continue with `--hub-nym` absent | -| `smoke.sh` implements the check and nothing invokes or mentions it | `smoke.sh:331-347`; repository-wide grep for `smoke.sh` returns only `smoke.sh` and `smoke-local.sh` themselves | -| `ForwardOnly` is not runtime-forceable | coordinator open item 6r (G1); `main.rs:112-143` decides once and threads immutably | - -**What was verified FALSE, and struck from the finding.** Four claims in the filed -version do not survive; each is now stated as a bounding fact in the Description -rather than as support for the finding: - -1. *"Undetectable by any wallet, user or passer-by."* False. `/nym-status` is - unauthenticated on the wallet-facing listener and reports - `diversion_configured: false`, and `shim/deploy/caution/OPERATORS.md:415` - documents that value as an alert with the words *"it is forward-only and hiding - nothing"*. The related claim that `/nym-status` "appears in the operator - runbooks only as a way to detect a dead mixnet client" is therefore also false. -2. *"The attestation and reproducibility chain is untouched, which is what makes it - attractive."* False, per open item 6q: `ZIS_HUB_NYM` is rendered into - `unit.env` (`caution.hcl.tmpl:161`), `unit.env` is measured into PCR0/PCR1, and - the environment is served unauthenticated at `.manifest.run_command` of every - `/attestation` response. A forward-only shim's PCRs differ and its manifest - visibly lacks the variable. -3. *"Silently selects no privacy."* False on the deploy path: `assemble-caution.sh:414` - announces it on the terminal and `PROVENANCE` records it in the published tree. - The two accidental routes a template-following operator could take are - **hard-refused** with `exit 2` (`assemble-caution.sh:189-192` and `:226-235`), - because `deploy.env.example:63` ships `NYM_EGRESS` non-empty and - `deploy.sh:125-127` always passes it. The addendum filed by the - `deploy.env.example` auditor was correct and is folded into the body. -4. The framing that forward-only is an accident. It is a **documented product - phase**: `OPERATORS.md:10-16` ("Phase 1 is forward-only, so it adds no privacy - yet") and `:242` ("Forward-only stays the default: no `--hub-nym`, no - diversion"). `docs/AVOIDING-FALSE-POSITIVES.md` §6 applies to the mode's - existence. It does **not** apply to the missing opt-out, which is what this - issue is now about. - -**One fact found during validation that strengthens the finding and was not in the -filed version.** The runbook's own worked *Deploy* command -(`shim/deploy/caution/OPERATORS.md:77-87`) contains neither `--hub-nym` nor -`--nym-egress`. An operator who copies that block verbatim — which is the -documented first deploy — brings up a **no-privacy shim serving live wallets -behind an unchanged public URL**, with diversion added later as a separate "Phase -2" step at `:226-242`. That is deliberate and announced to the operator, but it -means the forward-only state is not an exotic corner: it is the state of every -operator's first deployment, and the users of that endpoint are not told. - -**Severity justification — Medium, and why not High or Low.** - -*Impact when the state occurs:* total and permanent for the affected endpoint's -users — the operator gets the full transaction bytes plus the source IP, joinable -to the public chain forever. That is High impact. - -*Likelihood:* this is what pulls the grade down, and `docs/AVOIDING-FALSE-POSITIVES.md` -§7 is the governing test — *"Will insecure configurations really be used in -practice without the user knowing about the insecurity?"* Unlike the sibling -`DEBUG=1` finding, the polarity here is **not** inverted: `deploy.env.example:18` -ships a populated `HUB_NYM`, so the shipped template selects the *secure* state, -and the assembler refuses the two accidental routes out of it. Reaching -`ForwardOnly` unintentionally therefore requires a deployment that bypasses the -project's own tooling. Reaching it intentionally is available to a deceptive -operator, but it is disclosed by two unauthenticated endpoints and by the attested -manifest, so it is not the invisible attack the filed version described. The -severity guide's Medium band is written for exactly this shape: *"a serious -vulnerability that only exists if the user has configured the application in a -specific, uncommon, way."* - -*Why not High:* the three properties that would justify High — insecure by -default, undetectable, and attacker-triggerable — are all absent. The default is -secure, the state is disclosed three ways, and open item 6r establishes that no -outside party can force the mode at runtime. - -*Why not Low:* the defect is real and the consequence is a complete loss of the -product's stated purpose for every user of an affected endpoint. A fail-closed -option is a few lines, it is what the project's own review #11 demanded for the -analogous case, and — because the flag would be measured into the PCRs — it would -convert an operator's privacy promise into an attestable fact. The sibling -transport already fails closed on the identical input, which makes this an -inconsistency in the crate's own discipline rather than a deliberate trade-off. - -DO NOT DEVIATE FROM THIS TEMPLATE (`docs/SECURITY-ISSUE-TEMPLATE.md`). -LEAVE THIS MESSAGE IN PLACE SO THAT ALL AGENTS KNOW NOT TO DEVIATE FROM THIS TEMPLATE (`docs/SECURITY-ISSUE-TEMPLATE.md`). diff --git a/zeronym-22aa9851caf68-high-medium/medium/shim-mixnet-client-has-neither-retransmission-bound-the-hub-has-so-an-unacked-frame-retransmits-forever.md b/zeronym-22aa9851caf68-high-medium/medium/shim-mixnet-client-has-neither-retransmission-bound-the-hub-has-so-an-unacked-frame-retransmits-forever.md deleted file mode 100644 index 39dae9c8..00000000 --- a/zeronym-22aa9851caf68-high-medium/medium/shim-mixnet-client-has-neither-retransmission-bound-the-hub-has-so-an-unacked-frame-retransmits-forever.md +++ /dev/null @@ -1,437 +0,0 @@ -# The shim's mixnet client has neither of the two retransmission bounds the hub has, so every submit and lookup retransmits out of the shim's whole migration-carrying egress budget, and an un-acknowledgeable destination retransmits forever - -**Severity**: Medium -**Validation Status**: Confirmed -**Location**: `audit-target/zeronym/shim/src/nym_driver.rs:104-156` (`build_client`, no `DebugConfig` on any production path) and `:618` (`send_message` with `IncludedSurbs`); `audit-target/zeronym/shim/deploy/caution/assemble-caution.sh` (no ack-wait argument anywhere); `audit-target/zeronym/deploy.sh:110-113` (the shim branch of the only deploy driver). Contrast: `audit-target/zeronym/hub/src/nym_driver.rs:186-216` and `:638` (`send_reply`), `audit-target/zeronym/hub/deploy/caution/assemble-caution.sh:107,353`, `audit-target/zeronym/deploy.sh:121`, `audit-target/zeronym/deploy.env.example:25-41`. Mechanism in the pinned SDK (`nym-binaries-v2026.15-bydgoszcz`, commit `451c2aa`): `common/client-core/src/client/inbound_messages.rs:107-147`, `common/client-core/config-types/src/lib.rs:25,393-395,446`, `common/client-core/src/client/real_messages_control/acknowledgement_control/mod.rs:124-133`, `.../retransmission_request_listener.rs:67-165`, `.../sent_notification_listener.rs:10-38`, `.../message_handler.rs:500-555`, `common/client-core/src/client/transmission_buffer.rs:149-157,201-223`, `common/client-core/src/client/real_messages_control/real_traffic_stream/sending_delay_controller.rs:23`; `nym-node/src/node/mixnet/handler.rs:281-320` (the ack is emitted by the destination gateway). -**Found by agent:** Global, focus area G15 (the Nym transport read as one system, both sides) -**In scope of audit?** Yes - -## Description - -The hub and the shim run the same `nym-sdk` client, in the same enclave, on the -same mixnet. The project measured that this environment's acknowledgement path is -too slow for the SDK's default retransmission timer, and fixed it **on the hub -only**. Reading the two sides together, the shim is missing *both* of the bounds -that keep the hub's retransmission behaviour finite: - -1. **No ack-wait knob.** `hub/src/nym_driver.rs:202-216` reads - `ZIH_ACK_WAIT_ADDITION_MS` and raises `debug.acknowledgements.ack_wait_addition`; - `deploy.env.example:41` ships `HUB_ACK_WAIT_MS=15000` with the measurement that - justifies it. `shim/src/nym_driver.rs:104-156` builds **no** `DebugConfig` on any - production path — its only `debug_config` call (`:126-140`) is behind - `#[cfg(feature = "mixnet-localnet")]`, sets `message_sending_average_delay` rather - than `ack_wait_addition`, and is not compiled into the shipped image - (`shim/deploy/Containerfile:88`: `ARG CARGO_FEATURES="mixnet-driver"`). There is no - `--ack-wait-ms` in `shim/deploy/caution/assemble-caution.sh` and no variable for it - in `deploy.sh`'s shim branch. Every deployed shim therefore runs at the SDK default - of `1.5 x expected_delay + 1500 ms`. - -2. **No retransmission cap.** The hub's replies go out as `InputMessage::new_reply`, - which sets `max_retransmissions: Some(10)` (`inbound_messages.rs:128-147`). The - shim's submits and lookups go out as `InputMessage::new_anonymous`, which sets - `max_retransmissions: None` (`inbound_messages.rs:107-126`), and the SDK's global - cap defaults to `None` — documented in the field as *"None - no limit"* - (`config-types/src/lib.rs:393-395`, default at `:446`). Neither binary overrides - it. `PendingAcknowledgement::reached_max_retransmissions` is therefore - **permanently false** for every packet the shim sends - (`acknowledgement_control/mod.rs:124-133`), and the only other removal path is the - arrival of the acknowledgement itself - (`retransmission_request_listener.rs:81-90`). - -The consequence in normal operation is wasted egress on the one leg that carries -migrations, on a budget the project's own tests describe as barely sufficient. The -consequence in the failure mode the design itself describes — a `ZIS_HUB_NYM` entry -whose gateway is gone — is a permanently retransmitting frame that is never freed -and never gives up. - -## Attack Scenario and Steps - -**Path A — no attacker (the shipped steady state).** - -1. A shim is deployed by `deploy.sh` with `HUB_NYM` set. Nothing sets an ack-wait - value; the SDK default of 1500 ms applies. -2. The enclave's real acknowledgement round trip exceeds - `1.5 x expected_delay + 1500 ms`. This is not hypothetical: `deploy.env.example:30-40` - records that on the hub, *at 6000 ms*, "the same client ... still resent two of four - replies in their ENTIRETY (31 and 35 duplicates of a 32-packet reply, still arriving - 8-11 s after the first)". The shim runs the same client in the same enclave against - the same mixnet, with **larger** messages (45 packets for a submit against the hub's - 32-41 for a reply). -3. Every duplicate consumes one emission slot at the shaped ~8.33 packets/s the shim's - own `throughput_budget` tests pin (`shim/src/nym.rs:1094`; verified against - `sending_delay_controller.rs:23` `MAX_DELAY_MULTIPLIER = 6` and - `config-types/src/lib.rs:25` 20 ms). Roughly doubling the packets per migration - roughly halves the shim's migration throughput, against a budget the project's own - test says has only ~3x headroom inside `REQUEST_TIMEOUT`. - -**Path B — an unacknowledgeable destination (no attacker needed either).** - -4. `NymHandle::submit` sends every migration to **every** configured `ZIS_HUB_NYM` - entry (`shim/src/nym.rs:595-689`), and the module's own comment describes that list - as *"the current address, and the one it just rotated away from ... nothing is - listening at the stale one"* (`:620-626`). -5. The hub takes a fresh identity precisely when its gateway registration is - unrecoverable — 60 consecutive failed connects, or five short-lived clients - (`hub/src/nym_driver.rs:99,124,240-245,286-306`). In the connect-failure case the - old address's **gateway** is the thing that is down. -6. A sphinx packet's acknowledgement is emitted by the *destination gateway* after the - final hop, not by the recipient client — `nym-node/src/node/mixnet/handler.rs:285-317` - forwards the ack after either pushing to the client or storing for it, and the - forwarding call sits **outside** that match, so an unknown or offline *client* is - still acknowledged. Only an unreachable *gateway node* produces the permanent case. -7. Each submit fanned out to such an address produces 45 fragments that will never be - acknowledged. With `max_retransmissions = None` each one is re-prepared and - re-queued once per ack-wait period **for the life of the client**, and the - `PendingAcknowledgement` holding a clone of the fragment - (`message_handler.rs:529-531,545-546`) is never freed. Two sub-cases, both verified: - - **the dead gateway is still in the topology**: preparation succeeds, the packet is - re-queued on `TransmissionLane::Retransmission`, and 45 fragments each demanding a - re-emission every ~1.65 s cannot be served by an 8.33 packets/s emitter, so that - lane is never empty. Because `pick_random_small_lane` prefers any lane holding - fewer than 100 items (`transmission_buffer.rs:149-157`, `:201-223`), a - permanently-short Retransmission lane **preempts** the General lane that carries - real submits and lookups: real traffic gets roughly half the budget while the - General lane is short, and **nothing at all** while it exceeds 100 packets (about - three full frames). This condition never clears. - - **the dead gateway has fallen out of the topology**: preparation fails and the - listener restarts the timer instead of dropping the ack - (`retransmission_request_listener.rs:116-128`, *"we NEED to start timer here - otherwise we will have this guy permanently stuck in memory"*). No emission is - spent, but the fragment clone is retained for the life of the client and the retry - loop runs forever. Every subsequent migration adds another 45. - -**Path C — an attacker makes Path A worse, cheaply.** The confirmed issues -`gettransaction-flood-starves-migration-diversion.md` and -`junk-sendtransaction-flood-consumes-the-shims-whole-mixnet-egress-...md` both work by -buying sphinx packets out of the shim's shaped emission budget from an unauthenticated -wallet-facing endpoint. Doubling the packets each honest migration costs halves the -traffic an attacker must generate to reach the same starvation. - -**Attack Requirements and Assumptions:** -- Paths A and B require **no attacker at all** — A is the shipped configuration, B is - the failover shape the design documents (D10) plus a gateway-node outage. -- Path B additionally requires that the stale `ZIS_HUB_NYM` entry's gateway **node** be - unreachable rather than merely not hosting that client. A gateway that is up but no - longer registers the client still forwards the ack, so that variant costs only the - honest duplicate rate. The permanent case therefore needs a retired or offline - gateway node — common over months on a public mixnet, and specifically likely in the - connect-failure branch of the hub's own fresh-identity fallback. -- Path C requires only what those two confirmed issues already require. -- What makes this realistic: the project has already measured the underlying condition - in production, on the other component, and shipped a mitigation for it there. - -## Impact on Users - -The shim's mixnet client is the sole carrier of every diverted migration. Its emission -budget is a single, serialised, non-elastic resource, and `NymHandle::submit` answers -the wallet **success at dispatch** (`shim/src/hub.rs:228-240`), before anything has -left the enclave. Consequences, in order of severity: - -- **Migrations the wallet was told had succeeded are destroyed.** Anything still inside - the SDK when the client is torn down is discarded — the driver documents this itself - (`shim/src/nym_driver.rs:418-436`). Halving the drain rate roughly doubles the time a - dispatched-and-acknowledged submit spends inside that window, and Path B keeps the - window permanently occupied. -- **`GetTransaction` over the mixnet fails closed for real wallets** while duplicates - occupy the emission slots, which is the user-visible symptom of both confirmed - flood issues. In Path B's first sub-case this becomes permanent rather than - transient, and no redeploy-free action clears it. -- **The enclave's memory grows and is never reclaimed** in Path B (45 retained fragment - clones, ~2 KB each, per stuck submit, against `memory_mb = 2048`), and a shim OOM - destroys the entire in-flight set. This is bounded per submit but unbounded in the - number of submits, because every new migration fanned out to the dead address adds - another 45. -- **Nothing reports it.** `/nym-status` (`shim/src/nym.rs:205-215`) exposes only - `diversion_configured`, `mixnet_connected`, `client_deaths` and - `consecutive_rebuild_failures`. A client that is connected and spending its whole - budget on duplicates reports `mixnet_connected: true, client_deaths: 0`. - -## Technical Details / Code Analysis - -**The hub has the knob (`hub/src/nym_driver.rs:186-216`):** - -```rust - // How long the SDK waits for a packet's ack before it RETRANSMITS. The SDK - // computes an expected ack round trip from the CONFIGURED mix delays -- not - // measured -- and resends after `expected * ack_wait_multiplier + - // ack_wait_addition` (defaults 1.5x + 1.5 s). ... Measured 2026-08-17: a local hub's - // replies reached a shim with ~1 duplicate fragment per lookup; every - // DEPLOYED hub's reached the same shim with 15-25 -- and each duplicate is a - // full send slot at the throttled rate, which is how a 32-packet reply that - // should take ~5 s took 45-90 s from every enclave and timed out. ... - let builder = match std::env::var("ZIH_ACK_WAIT_ADDITION_MS") - .ok() - .and_then(|v| v.parse::().ok()) - { - Some(ms) => { - let mut debug = nym_sdk::DebugConfig::default(); - debug.acknowledgements.ack_wait_addition = Duration::from_millis(ms); - ... - builder.debug_config(debug) - } - None => builder, - }; -``` - -**The shim's equivalent function has no such branch (`shim/src/nym_driver.rs:104-156`).** -Its whole `build_client` is `MixnetClientBuilder::new_ephemeral()`, an optional -`request_gateway`, an optional localnet-only `debug_config`, an optional localnet-only -topology provider, then `build()`/`connect_to_mixnet()`. The localnet branch is -explicitly and correctly excluded from production: - -```rust - // Gated on `mixnet-localnet` so a PRODUCTION binary cannot read it at all: a - // non-default send rate would make this client's traffic distinguishable from - // every other Nym client, which is a fingerprint ... - #[cfg(feature = "mixnet-localnet")] - let builder = match std::env::var("ZIS_LOCALNET_SEND_DELAY_MS") -``` - -and `shim/deploy/Containerfile:88` sets `ARG CARGO_FEATURES="mixnet-driver"`, so -`mixnet-localnet` is off in every shipped image. A repository-wide search for -`DebugConfig|debug_config|ack_wait` over `shim/`, `hub/` and `nymnet/` returns exactly -the hub's branch, the shim's localnet branch, and the hub's deploy plumbing — the shim -has no production `DebugConfig` at all. - -**The deploy chain mirrors the asymmetry.** `deploy.sh:114-121` (hub branch) forwards -`HUB_ACK_WAIT_MS` as `--ack-wait-ms`, which -`hub/deploy/caution/assemble-caution.sh:107,353` turns into -`ZIH_ACK_WAIT_ADDITION_MS` in `unit.env`. `deploy.sh:110-113` (shim branch) forwards -only `--backend`, `--backend-tls` and `--hub-nym`; `grep -c ack-wait` over -`shim/deploy/caution/assemble-caution.sh` returns **0**. - -**The SDK caps the hub's replies and not the shim's sends** -(`common/client-core/src/client/inbound_messages.rs`): - -```rust - pub fn new_anonymous( // :107 <- every shim submit and lookup - ... - let message = InputMessage::Anonymous { - ... - max_retransmissions: None, // :119 - }; - - pub fn new_reply( // :128 <- every hub reply - ... - let message = InputMessage::Reply { - ... - // \/ set it to SOME sane default so that if we run out of surbs and constantly - // fail to request more, we wouldn't be stuck in limbo - max_retransmissions: Some(10), // :140 - }; -``` - -`shim/src/nym_driver.rs:616-620` calls `sender.send_message(recipient, out.frame.to_vec(), -IncludedSurbs::new(out.reply_surbs))`, and `IncludedSurbs::Amount(_)` routes to -`InputMessage::new_anonymous` (`sdk/rust/nym-sdk/src/mixnet/traits.rs:80-96`). -`hub/src/nym_driver.rs:638` calls `sender.send_reply(tag, frame)`, which routes to -`InputMessage::new_reply` (`traits.rs:122-134`). - -The global cap that could have saved the shim is also unset -(`common/client-core/config-types/src/lib.rs:393-395` and `:446`): - -```rust - /// Specify how many times particular packet can be retransmitted - /// None - no limit - pub maximum_number_of_retransmissions: Option, - ... - maximum_number_of_retransmissions: None, -``` - -so `reached_max_retransmissions` is a disjunction of two `is_some_and` tests over two -`None`s (`acknowledgement_control/mod.rs:124-133`) and can never be true for a shim -packet. `retransmission_request_listener.rs:81-90` is the only place that consults it, -and the only other path that removes a pending acknowledgement is the ack itself. - -**The retained data.** `try_split_and_send_non_reply_message` -(`message_handler.rs:500-555`) clones every fragment before preparing it — *"we need to -clone it because we need to keep it in memory in case we had to retransmit it"* -(`:529-531`) — and stores the clone in a `PendingAcknowledgement` -(`:545-546`). Those clones are held until the ack arrives; with no cap and no ack, for -the life of the client. - -**How fast retransmissions can actually be generated, and why the lane matters.** A -retransmission timer is started **only after the packet has been emitted**, by -`SentNotificationListener`, whose module doc states the reason verbatim: *"It is -required because when we send our packet to the `real traffic stream` controlled by a -poisson timer, there's no guarantee the message will be sent immediately, so we might -accidentally fire retransmission way quicker than we should have"* -(`sent_notification_listener.rs:10-13,30-38`). So retransmission demand is -self-limiting: it cannot exceed the emission rate, and the Retransmission lane does not -grow without bound. What it does instead is **occupy the emitter permanently**: 45 -stuck fragments want a slot every ~1.65 s (about 27 packets/s of demand) against an -8.33 packets/s ceiling, so the lane is always non-empty, and -`pop_next_message_at_random` prefers any lane with fewer than 100 items over a longer -one (`transmission_buffer.rs:149-157`, `:201-223`). Real traffic therefore shares the -budget roughly evenly while the General lane is short and is starved outright once it -is long. The hub's `Some(10)` bounds the same behaviour to ~16 s per stuck reply; the -shim has no such bound. - -**CORRECTION, carried forward from `PROGRESS.md` item 7b-REFUTED and re-verified here.** -An earlier claim held that *"`insert_pending_acks` arms the retransmission timers before -`forward_messages`, so ack timers expire on packets still sitting in the queue"*, i.e. -that retransmission **self-amplifies**. **That is not what the SDK does and the premise -is withdrawn.** `ActionController::handle_insert` inserts the pending ack with -`queue_key = None` and starts no timer; the timer is started only by `Action::StartTimer` -from `SentNotificationListener`, after emission, as quoted above. Nothing in this issue -depends on that withdrawn premise: the retransmission problem here is caused by the ack -round trip exceeding a timer the shim cannot tune, and by the absence of any cap on how -many times a packet may be retried. The **memory-retention** half of the old claim is -correct and is restated above. - -## Recommendations - -1. **Give the shim the knob the hub has.** Add a `ZIS_ACK_WAIT_ADDITION_MS` branch to - `shim/src/nym_driver.rs::build_client` identical to `hub/src/nym_driver.rs:202-216`, - a `--ack-wait-ms` argument to `shim/deploy/caution/assemble-caution.sh`, and a - `SHIM_ACK_WAIT_MS` variable to `deploy.sh`'s shim branch and `deploy.env.example`. - The hub's shipped value (15000) is the measured starting point; the shim's frames - are larger, so it needs at least as much. -2. **Cap the shim's retransmissions.** Set - `debug.traffic.maximum_number_of_retransmissions = Some(n)` in the same - `DebugConfig` (the SDK already applies `Some(10)` to the hub's replies, and - `InputMessage::with_max_retransmissions` exists for per-message control). Without a - cap, an unacknowledgeable destination is a permanent, unrecoverable drain on the one - resource that carries migrations. -3. **Make the drain visible.** `/nym-status` cannot currently distinguish a healthy - client from one spending its entire budget on duplicates. The SDK's - `MixnetClient::shared_lane_queue_lengths()` (which exposes the Retransmission lane - separately) and the pending-ack count are the two numbers that would; failing that, - publish `out_frames.len()` and a dispatched-versus-acked delta (the ack nonce and - waiter already exist — see - `nym-submit-acks-are-never-read-so-every-hub-refusal-is-invisible.md`). -4. **Reconsider the unconditional submit fan-out to stale addresses.** A destination - that has not acknowledged anything for a long interval should stop being fanned out - to, or the list should be prunable without a redeploy. - -## Validation Information - -**Verdict: CONFIRMED. Severity confirmed at Medium.** - -### Verified against the pinned SDK tree at `451c2aa` (read directly, not inferred) - -- `inbound_messages.rs:107-126` — `new_anonymous` sets `max_retransmissions: None`. - `:128-147` — `new_reply` sets `Some(10)` with the quoted comment. `new_regular` - (`:88-105`) is also `None`. **The asymmetry is by construction in the SDK, exactly as - claimed.** -- `config-types/src/lib.rs:393-395` and `:446` — `maximum_number_of_retransmissions` - defaults to `None`. Neither zeronym binary sets it (`grep -rn - "maximum_number_of_retransmissions\|DebugConfig\|debug_config" shim/ hub/` returns - only the hub's `ack_wait_addition` branch and the shim's localnet-only - `message_sending_average_delay` branch). -- `acknowledgement_control/mod.rs:124-133` — `reached_max_retransmissions` is - `local.is_some_and(..) || global.is_some_and(..)`, so it is **permanently false** for a - shim packet. -- `retransmission_request_listener.rs:81-90` — that predicate is the only path to - `Action::new_remove`. `:116-128` — when preparation fails the timer is restarted - rather than the entry dropped, with the comment *"we NEED to start timer here - otherwise we will have this guy permanently stuck in memory"*, which is what makes - the removed-from-topology sub-case a permanent retention rather than a permanent - emission cost. -- `sdk/rust/nym-sdk/src/mixnet/traits.rs:72-96` and `:122-134` — the routing from - `send_message`/`send_reply` to the two constructors, confirming which zeronym call - site gets which cap. -- `sent_notification_listener.rs:10-13,30-38` — timers start after emission. This is - the fact that **refutes** the withdrawn self-amplification premise and also bounds - the retransmission rate; the issue text has been rewritten around it. -- `transmission_buffer.rs:149-157` (`is_small()` = fewer than 100 items) and - `:201-223` (`pop_next_message_at_random` prefers a small lane over everything else) - — this is new to the filing and is what makes a stuck Retransmission lane preempt - real traffic. Verified in source; added to Path B. -- `nym-node/src/node/mixnet/handler.rs:281-320` — `forward_ack_packet` is called - **after** the push-or-store match and outside it, so an offline or unknown client is - still acknowledged. The filing's central caveat ("only an unreachable gateway - produces the permanent case") is therefore correct, and this was checked before the - claim rather than after. -- `sending_delay_controller.rs:23` `MAX_DELAY_MULTIPLIER = 6` and - `config-types/src/lib.rs:25` 20 ms give the 8.33 packets/s ceiling the arithmetic - uses; `real_traffic_stream.rs:373-377` shows the multiplier rises only when the - gateway channel is full, which the project reports its deployed enclaves reach - (`shim/src/nym.rs:1071-1077`). - -### Verified against the target - -- `shim/src/nym_driver.rs:104-156` — `build_client` in full; the only `debug_config` - call is `#[cfg(feature = "mixnet-localnet")]` at `:126-140`, and - `shim/deploy/Containerfile:88` is `ARG CARGO_FEATURES="mixnet-driver"`. -- `hub/src/nym_driver.rs:186-216` — the `ZIH_ACK_WAIT_ADDITION_MS` branch, with the - 2026-08-17 measurement in the comment. -- `deploy.sh:110-113` (shim) versus `:114-121` (hub, including - `--ack-wait-ms "$HUB_ACK_WAIT_MS"`); `hub/deploy/caution/assemble-caution.sh:107` and - `:353`; `grep -c ack-wait shim/deploy/caution/assemble-caution.sh` = **0**. -- `deploy.env.example:25-41` — the shipped `HUB_ACK_WAIT_MS=15000` and the measurement - that at 6000 ms two of four replies were still resent in their entirety. This is the - project's own production evidence that the condition is real in this exact - environment. -- `shim/src/nym.rs:595-689` — the fan-out to every configured address, and `:620-626` - the comment stating that the stale address has nothing listening. So Path B's - precondition is a *documented, intended* configuration, not a hypothetical - misconfiguration. -- `shim/src/hub.rs:228-240` — `Submit::Accepted` at hand-off, with *"a refusal is never - surfaced here"*; `shim/src/nym.rs:205-215` — `/nym-status` carries no queue or - duplicate signal. - -### Corrections made during validation - -- **"27 pkt/s of demand against an 8.33 pkt/s ceiling: one stuck submit is enough to - saturate the shim's entire mixnet egress permanently"** overstated the mechanism. - Retransmission timers restart only after emission, so demand cannot exceed supply and - the lane does not grow without bound. The accurate and still-serious statement, now in - Path B, is that the lane is permanently non-empty and, because the SDK prioritises - short lanes, permanently preempts real traffic — roughly half the budget when the - General lane is short and effectively all of it once the General lane exceeds 100 - packets. -- **Path B was split into its two real sub-cases** (dead gateway still in topology = - permanent emission cost; dead gateway out of topology = permanent memory retention - and a forever retry loop with no emission), because the SDK behaves differently in - each and only the first starves the emitter. -- The memory claim was made precise: ~90 KB per stuck submit is bounded, but it is - **unbounded in the number of submits**, since each new migration fanned out to the - dead address adds another 45 retained clones that are never freed. -- The withdrawn self-amplification premise (item 7b-REFUTED) is restated as a marked - correction, and nothing in the issue now depends on it. The words "self-amplif*" do - not appear except inside that correction. - -### Exploitability and real-world impact - -Paths A and B need no attacker. Path A is the shipped configuration and its cost is -measured by the project itself on the sibling component; its user-visible effect is -reduced migration throughput and more `UNAVAILABLE` on `GetTransaction`, on a budget -the shim's own test says has roughly 3x headroom. Path B needs a `ZIS_HUB_NYM` entry -whose gateway node is offline — a state the design explicitly plans for (the failover -list) and which the hub's own connect-failure fallback tends to produce — and its cost -is permanent until the shim is redeployed, which for an immutable attested enclave is a -~25-minute operator action per shim. - -### Severity: why Medium - -- Not High: no confidentiality breach, no attacker required, and the permanent case - needs a specific (though designed-for and realistic) configuration state. The - everyday cost is roughly a doubling of packets, which the system absorbs. -- Not Low or Info: it is a genuine unbounded-resource defect on the single serialised - resource that carries every diverted migration, in a component that has already told - the wallet the migration succeeded, with no monitoring that can see it, and the - project fixed the identical problem on the other binary — so the omission is a - deviation from its own established practice rather than an accepted residual. -- Not double-counted: the two confirmed flood issues own the *attacker-driven* - starvation of the shim's egress; this issue owns the *no-attacker* and *permanent* - cases and the missing cap, and Path C is noted only as an interaction, not as its own - harm. - -### False-positive checks applied - -- *§1 Assumption an attacker cannot violate?* Not applicable; Paths A and B need no - attacker. -- *§4 Test/debug code?* Checked and inverted: the shim's only `debug_config` is - correctly gated out of production — that is precisely why the production path has no - knob. -- *§5 Impractical resource exhaustion?* No resources are demanded of anyone; the drain - is self-inflicted. -- *§6 Intentional design?* No. The hub's branch, the deploy plumbing and the shipped - `HUB_ACK_WAIT_MS` show the project treats this as a defect worth fixing; the shim was - simply not given the same treatment. -- *§9 Obviously broken functionality?* No — Path A degrades rather than breaks, and - Path B needs a stale address plus a dead gateway node. - -DO NOT DEVIATE FROM THIS TEMPLATE (`docs/SECURITY-ISSUE-TEMPLATE.md`). -LEAVE THIS MESSAGE IN PLACE SO THAT ALL AGENTS KNOW NOT TO DEVIATE FROM THIS TEMPLATE (`docs/SECURITY-ISSUE-TEMPLATE.md`). diff --git a/zeronym-22aa9851caf68-high-medium/medium/shim-nym-driver-every-teardown-path-silently-destroys-acknowledged-submits.md b/zeronym-22aa9851caf68-high-medium/medium/shim-nym-driver-every-teardown-path-silently-destroys-acknowledged-submits.md deleted file mode 100644 index c69fc477..00000000 --- a/zeronym-22aa9851caf68-high-medium/medium/shim-nym-driver-every-teardown-path-silently-destroys-acknowledged-submits.md +++ /dev/null @@ -1,414 +0,0 @@ -# Process shutdown destroys migrations the wallet was already told had succeeded: `Step::Stop` disconnects the mixnet client with no drain, no count and no log, `main` never joins the driver, and one exit path never signals it at all - -**Severity**: Medium -**Validation Status**: Confirmed -**Location**: -`audit-target/zeronym/shim/src/nym_driver.rs:410-415` (`Step::Stop` — `client.disconnect().await; return;`), -`audit-target/zeronym/shim/src/nym_driver.rs:353-358` (the two ways `Step::Stop` is produced), `:362-369` (the `out_frames` arm that is never drained), -`audit-target/zeronym/shim/src/nym_driver.rs:416-467` (`Step::Rebuild`, and the residual statement at `:421-436`), `:468-491` (`Step::Died`), -`audit-target/zeronym/shim/src/nym_driver.rs:25-29` (the module's own "a clean rotation must run it to completion"), -`audit-target/zeronym/shim/src/nym.rs:595-690` (`NymHandle::submit`; `Ok(())` at `:660-661`, `:684-689`), `:915-1010` (`run_supervisor`, the shutdown arm at `:953-956`), `:835-905` (`correlate`, the reserved `out_frames` permit at `:853`), `:392-398` (why disconnect is a command), -`audit-target/zeronym/shim/src/hub.rs:236-240` (`Submit::Accepted` at mixnet hand-off, *"a refusal is never surfaced here"*), -`audit-target/zeronym/shim/src/main.rs:166-176` (the only thing `main` awaits), `:335-339` (channel capacities 32 / 8 / 32 / 8 / 8), `:341-356` (the three detached, never-joined tasks), -`audit-target/zeronym/shim/src/proxy.rs:180` (`DRAIN_TIMEOUT` = 10 s), `:545-553` (the drain, which covers only the wallet-facing leg), `:499` (the fatal-accept-error return) -**Found by agent:** Local (file audit of `shim/src/nym_driver.rs`); validated 2026-08-18 -**In scope of audit?** Yes — priority area #4 (the dispatch-only submit's loss window) - -## Description - -Submit is dispatch-only. `NymHandle::submit` returns `Ok(())` as soon as a -`Request` is accepted into the shim's in-process request channel -(`nym.rs:660-661`, `:684-689`), and `HubTransport::submit` maps that straight to -`Submit::Accepted { txid }` (`hub.rs:236-240`), whose own comment says *"a -refusal is never surfaced here"*. The wallet is told `error_code 0` and handed a -txid **before the frame has reached the SDK, let alone the gateway, let alone -the hub**. - -Everything between that acknowledgement and the mixnet is destroyed on process -shutdown, with no drain, no accounting and no log line: - -```rust -// shim/src/nym_driver.rs:410-415 - match step { - Step::Ferried => {} - Step::Stop => { - client.disconnect().await; - return; - } -``` - -That is the whole of it. There is no attempt to pump `out_frames` first, no -`in_flight` accounting, and no count of how many frames were abandoned. -`Step::Stop` is what **every** SIGTERM, enclave stop and redeploy produces -(`nym_driver.rs:353-358`: `ClientCommand::Disconnect`, which `run_supervisor` -sends from its shutdown arm at `nym.rs:953-956`). - -Three further facts make it worse than a missing log line: - -1. **The 10-second drain the shim does have protects the wrong queue.** - `serve_with_shutdown` waits up to `DRAIN_TIMEOUT` = 10 s for in-flight - *wallet connections* (`proxy.rs:180`, `:545-553`). **Zero seconds** protect - the mixnet queue, which is where the acknowledged submits are. The drain that - exists guards the leg on which nothing has been promised yet; the leg on - which success has already been promised has none. -2. **`main` never joins the driver, so the `disconnect()` the module documents as - load-bearing is not run to completion.** `nym_driver.rs:25-29` and - `nym.rs:392-398` both state the reason `Disconnect` is a command rather than a - drop: *"it is not cancel-safe and a dropped LIVE client leaks its background - tasks (D12), so a clean rotation must run it to completion."* But - `build_nym_transport` spawns `run_transport`, `run_supervisor` and - `run_driver` detached (`main.rs:341-356`) and keeps no `JoinHandle`; `main` - awaits only `serve_with_shutdown` (`main.rs:166-176`). Both the proxy and the - supervisor are woken by the *same* signal, and with no wallet connections open - the drain returns essentially immediately — so `main` returns, the - `#[tokio::main]` runtime is dropped, and the driver task is cancelled wherever - it is, very often inside `client.disconnect().await` at `:413`. -3. **There is a third exit that never signals the driver at all.** A non-transient - `accept()` error returns `Err` from `serve_with_shutdown` (`proxy.rs:499`), - which propagates out of `main`. The `shutdown()` future the supervisor is - waiting on never resolves, so `ClientCommand::Disconnect` is never sent: the - driver is cancelled mid-anything, holding a **live** client. This is the exact - case the module's D12 note says must not happen, and it is reachable — see - `shim-accept-loop-treats-documented-transient-errors-as-fatal.md`. - -## Attack Scenario and Steps - -The loss happens on its own during ordinary operation, and both the primary -adversary and an anonymous stranger can widen or aim it. - -**A. Ordinary operation, no attacker.** An operator stops or redeploys the -enclave. Everything the pipeline holds at that instant is destroyed: - -- up to **32** acknowledged submits in the request channel (`main.rs:335`, - `mpsc::channel(32)`) — the wallet has already had `error_code 0` for each; -- up to **8** in `out_frames` (`main.rs:336`), plus the one capacity slot - `correlate` keeps reserved (`nym.rs:853-863`); -- the **1** hand-off in the driver's `in_flight`; -- everything the SDK holds: a one-slot input, an 8-deep batch channel, and an - unbounded transmission buffer drained at the throttled rate. - -That is **41 frames the shim itself is holding**, plus ~9 more inside the SDK. -At the emission rate the crate derives for itself — `MAX_DELAY_MULTIPLIER` = 6 -against the 20 ms default average delay, i.e. ~8.33 packets/s -(`nym.rs:1085-1094`) — one submit is 32 frame packets plus 13 reply SURBs = 45 -packets ≈ 5.4 s. A full ~50-slot pipeline is therefore **~2,250 packets ≈ 270 -seconds of emission**, all of it already answered "success" to somebody. - -**B. Any anonymous party can guarantee the pipeline is full.** The single -mitigating factor the original filing named — "at ~0.77 Orchard-touching -transactions per block the pipeline is usually empty, so a randomly-timed -restart usually destroys nothing" — is removable by an unauthenticated stranger -for about **one byte per second**. That is the confirmed High -`junk-sendtransaction-flood-consumes-the-shims-whole-mixnet-egress-and-converts-acknowledged-migrations-into-silent-loss.md`: -a 5-byte `SendTransaction` body classifies `Unparseable`, fails safe toward -diversion, and buys 45 packets ≈ 5.4 s of the shim's *entire* egress. Held full, -the pipeline is a standing ~270-second queue, and a real wallet's migration -entering it is queued behind the junk with a success already returned. **Any -restart during the flood therefore destroys real, acknowledged migrations**, and -the "usually empty" defence does not apply. - -**C. The operator can time it.** The indexer operator is adversary #1 and owns -the Nitro parent host and the process lifecycle. The README concedes they learn -*that* a client diverted — a diverted `SendTransaction` is the one request that -produces no corresponding upstream connection to their indexer, on a wallet -connection they can see at the TCP layer. So: watch for a divert, SIGTERM the -shim within the next few seconds, and the frame dies inside `Step::Stop` while -the wallet holds a success and a txid. The action is fully deniable ("we -restarted the enclave"). - -**D. Same outcome without a restart.** Blackholing the enclave's egress to the -gateway port until the SDK hits its 20-consecutive-failure hard stop -(`nym.rs:377-381`) yields `Step::Died` (`nym_driver.rs:468-491`), which drops -`in_flight` and the client — and with it the buffer — with no guard and no -accounting. Egress rules are operator-controlled (`deploy.env.example`, -`NYM_EGRESS`). Here the frames were arguably undeliverable anyway; what the code -withholds is any record that they existed. - -**Attack Requirements and Assumptions:** - -- Scenario A needs no attacker at all; it is the normal redeploy path, and it is - the one that runs on every deployment of every operator. -- Scenario B needs only the ability to send unauthenticated gRPC requests to a - public endpoint, at ~1 byte/second. -- Scenarios C and D need the operator, who controls process lifecycle, - configuration and egress by construction. -- No attacker needs to break the mixnet, the enclave, or any cryptography. - -## Impact on Users - -A wallet is told its migration was broadcast, and is handed a txid it will -display and record. The transaction exists nowhere: not in a mempool, not in the -hub's queue, not on the mixnet. Nothing recovers it: - -- the hub never received it, so there is no queue entry, no payload-hash dedup - entry and no requeue; -- the shim keeps no per-migration state by design (`lib.rs:30-35`); -- confirmation tracking is *"designed, not built"* (`AUDIT-INSTRUCTIONS.md`, - self-declared limitations); -- the only feedback channel the design names is the wallet noticing - non-confirmation and resending (`nym_driver.rs:598-607`), which requires the - wallet to distrust the success it was given; -- **nothing anywhere logs or counts the loss**, so no operator can ever learn - that a redeploy destroyed N migrations, and no user can be told. - -For a migration this is worse than an ordinary lost broadcast: the funds are in a -pool the network has closed to new value, the migration is the user's only route -out, and the user has been told it is done. Under ZIP 318's expiry schedule the -notes can then sit unusable for a long time (see -`zip318-canonical-expiry-is-the-only-recovery-clock-and-a-lost-migration-freezes-the-users-notes-for-30-to-60-days.md`). - -> **CORRECTION 2026-08-18 (validation of the cited file — SUPERSEDES the sentence above).** -> The wallet does **not** wait for expiry. Both official Zcash light-wallet SDKs -> automatically resubmit a sent-but-unmined transaction for as long as it remains -> unexpired — the Android SDK at the head of every ~20 s sync loop and after every -> processed block batch (`CompactBlockProcessor.kt:573,615,723`; selection -> `mined_height IS NULL AND expiry_height > ?`), the iOS SDK at most once per 300 s -> (`TxResubmitter.swift:8-15`, `TransactionDao.swift:218-228`) — and the hub's -> payload-hash dedup makes the resend free. The wallet's non-confirmation signal comes -> from compact-block scanning, which the shim does not intercept (`proxy.rs:1068-1074`). -> Expiry is therefore the **retry horizon**, not the wait: ~50 minutes for the ZIP 203 -> default traffic the shim also diverts, 30–60 days for a ZIP 318 migration. A -> *transient* loss self-heals within minutes; only a loss condition that **outlives the -> horizon** destroys the submission permanently — which is exactly what this issue's -> condition does, so this issue's severity is unaffected. Do not write "the user waits -> 30 to 60 days" in the report. Full refutation and the replacement paragraph: -> `issues/invalid/zip318-canonical-expiry-is-the-only-recovery-clock-and-a-lost-migration-freezes-the-users-notes-for-30-to-60-days.md`. - - -## Technical Details / Code Analysis - -**The wallet is answered before anything leaves the process.** -`shim/src/nym.rs:659-689`: - -```rust -659 match tokio::time::timeout_at(deadline, self.requests.send(request)).await { -660 Ok(Ok(())) => dispatched += 1, -... -684 if dispatched > 0 { -685 Ok(()) -686 } else { -687 Err(NymError::TransportGone) -688 } -``` - -`shim/src/hub.rs:236-240`: - -```rust -236 // ... so a refusal is never surfaced here. -237 Ok(()) => Ok(Submit::Accepted { -238 txid: crate::nym::local_txid(tx_bytes), -239 }), -``` - -`self.requests` is the sender end of `mpsc::channel(32)` (`main.rs:335`). So -"success" means "one of 32 slots in an in-process channel accepted a struct". - -**The unguarded shutdown path**, quoted in full in the Description above -(`nym_driver.rs:410-415`). Note also that the driver's `select!` is **not** -`biased`: when the shutdown arrives, the `commands.recv()` arm and the -`out_frames.recv()` arm are both ready and tokio picks pseudo-randomly, so the -driver can take `Disconnect` with all 8 `out_frames` slots occupied. - -**Nothing joins the driver.** `shim/src/main.rs:341-356`: - -```rust -341 tokio::spawn(nym::run_transport(req_rx, out_tx, in_rx, inflight.clone())); -... -345 tokio::spawn(nym::run_supervisor(rotation, evt_rx, cmd_tx, inflight, shutdown())); -346 tokio::spawn(nym_driver::run_driver( -``` - -Three detached tasks, no `JoinHandle` retained. `main` (`:166-176`) awaits only -`serve_with_shutdown`, whose shutdown path is `shim/src/proxy.rs:545-553`: - -```rust -545 drop(live_tx); -546 tracing::info!("shutdown requested, draining in-flight connections"); -547 match tokio::time::timeout(DRAIN_TIMEOUT, live_rx.recv()).await { -548 Ok(_) => tracing::info!("drained, exiting"), -``` - -With no wallet connections open, every `live_tx` clone is already gone, so -`recv()` resolves immediately and `main` returns in the same poll cycle in which -the signal fired. The supervisor's `commands.send(ClientCommand::Disconnect)` -(`nym.rs:953-956`) succeeds into an empty 8-slot channel without waiting for the -driver, and then the runtime is dropped out from under whatever the driver was -doing. - -**The pipeline depth, precisely.** `correlate` reserves capacity on `out_frames` -*before* accepting a request (`nym.rs:853-863`), and hands the frame over -non-blockingly; the driver takes one frame at a time (`nym_driver.rs:362`, -guarded on `in_flight.is_none()`). So under load `out_frames` sits at 8 and -`requests` backs up to 32. 32 + 8 + 1 = **41** frames held by the shim's own -code, all of which have been answered "success", plus the SDK's one-slot input, -its 8-deep batch channel, and its unbounded transmission buffer. - -**The residual statement at `nym_driver.rs:421-436` is not accurate.** It says -the exposure is *"real, bounded by the drain rate, and recorded in PRODUCTION.md -rather than papered over here."* Three problems: the drain **rate** is not a -bound on the **quantity** (the SDK's buffer is described as unbounded in the same -sentence, and the driver's `out_frames` arm keeps feeding it on every loop turn, -`nym_driver.rs:362-369`); the paragraph discusses only `Step::Rebuild` and never -mentions `Step::Stop`, which is the path every deployment takes; and -`PRODUCTION.md` **does not exist in this repository**, so the residual is -recorded nowhere an auditor or operator can read it. - -**A second inaccurate quantity, in the probe guard's rationale.** -`nym_driver.rs:304-311` argues that *"once we stop feeding it it can only drain, -and two silent rounds is 120 s of drain at the throttled rate — far more than any -residual it could be holding."* Both halves are wrong: the driver never stops -feeding the SDK (the `out_frames` arm at `:362` is armed on every turn of the -loop for the whole window), and a full pipeline is ~2,250 packets ≈ **270 s**, -i.e. **2.3x** the stated 120 s. `out_frames.len() == 0` at a tick instant says -nothing about the SDK's own buffer. See the Validation Information for why this -is a comment/reasoning defect rather than a live loss path. - -## Recommendations - -1. **Account for the loss. This is the cheapest change and the one with no - downside.** `Step::Stop`, `Step::Rebuild` and `Step::Died` should each emit a - count of what was abandoned — `out_frames.len()`, `in_flight.is_some()`, and - ideally the request channel's `len()` — as **counts only**, which is #157-legal - and tells an operator that migrations were destroyed. Today all three are - completely silent. -2. **Drain before disconnecting on the shutdown path.** `Step::Stop` can pump - `out_frames` into the SDK under a bounded deadline before calling - `disconnect()`. That does not empty the SDK's own buffer, but it removes the - 41 frames the shim itself is holding, which is the part under this code's - control. -3. **Join the driver at shutdown.** Return a `JoinHandle` for `run_driver` from - `build_nym_transport` and have `main` await it (bounded) after - `serve_with_shutdown` returns, so the documented "run it to completion" is - actually true and so step 2 has time to run. The mixnet leg deserves at least - the same courtesy as the 10 s `DRAIN_TIMEOUT` already given to the wallet leg. -4. **Signal the driver on *every* exit.** The fatal-accept-error return - (`proxy.rs:499`) must also drive the shutdown sequence, or the supervisor must - observe `main`'s exit some other way; today that path cancels a live client - with no `disconnect()` at all. -5. **Stop calling channel hand-off "success", or hold the acknowledgement - longer.** As long as `hub.rs:236-240` answers `Accepted` at hand-off, every - defect downstream converts into "a transaction the user believes is spent and - that exists nowhere". The honest options are to hold the acknowledgement until - the frame has at least been accepted by the SDK, or to surface the uncertainty - to the wallet. -6. **Correct the two inaccurate comments** — the residual at - `nym_driver.rs:421-436` (the window is not bounded by the drain rate, and it - applies to shutdown as much as to rebuild; and `PRODUCTION.md` is not in this - repository) and the drain figure at `nym_driver.rs:304-311` (a full pipeline - is ~270 s, not 120 s, and the driver never stops feeding the SDK). - -## Validation Information - -**Verdict: CONFIRMED, Medium.** Validated 2026-08-18. Every mechanical claim -about the shutdown path holds; one of the filing's three legs is demoted from a -harm to a comment defect, and two facts are added that the filing did not have. - -### Mechanical verification - -| Claim | Verified at | -|---|---| -| `Step::Stop` is `disconnect(); return;` with no guard, drain, count or log | `shim/src/nym_driver.rs:410-415` | -| `Step::Stop` is what a SIGTERM produces | `shim/src/nym_driver.rs:353-358`; `shim/src/nym.rs:953-956` (`run_supervisor`'s shutdown arm sends `Disconnect` and returns) | -| The driver's `select!` is not `biased`, so `Disconnect` can win over a non-empty `out_frames` | `shim/src/nym_driver.rs:352-369` | -| The wallet is answered at channel hand-off | `shim/src/nym.rs:659-661`, `:684-689`; `shim/src/hub.rs:236-240` | -| Pipeline depth 32 + 8 + 1 = 41 frames, plus the SDK's 1 + 8 | `shim/src/main.rs:335-336`; `shim/src/nym.rs:853-863`; `shim/src/nym_driver.rs:362` | -| 45 packets per submit (32 frame + 13 SURBs) at ~8.33 packets/s | `shim/src/nym.rs:96`, `:1085-1094`, `:1109` | -| `main` retains no `JoinHandle` and awaits only the proxy | `shim/src/main.rs:166-176`, `:341-356` | -| The 10 s drain covers the wallet leg only, and returns immediately when no connections are open | `shim/src/proxy.rs:180`, `:545-553` | -| A fatal accept error returns from `main` with the supervisor never signalled | `shim/src/proxy.rs:499`; `shim/src/nym.rs:929-1010` (the only `Disconnect` sender is the shutdown arm) | -| The module documents `disconnect()` as needing to run to completion | `shim/src/nym_driver.rs:25-29`; `shim/src/nym.rs:392-398` | -| `PRODUCTION.md` does not exist anywhere in the repository | `find audit-target/zeronym -name PRODUCTION.md` returns nothing | - -### One leg demoted — this is a correction, not a confirmation - -The filing's second numbered claim was that the probe-triggered rebuild is a live -loss path: that the `out_frames.len() == 0` guard (`nym_driver.rs:312`) "cannot -see what it is guarding", so a spurious `Step::Silent` fires `ClientEvent::Died`, -the supervisor answers `Rebuild`, and `:440` disconnects a **live** client -holding acknowledged submits. - -**The code observation is correct and the harm does not follow.** The guard's -stated reasoning is genuinely wrong on both halves (the driver never stops -feeding the SDK; 120 s is 2.3x too small), and that is worth fixing. But firing -the false `Silent` additionally requires `seen == mark` — **zero inbound -messages across two 60-second rounds** (`nym_driver.rs:312`, `:574`, `:579`). -Every submit carries 13 reply SURBs and the hub answers each with a one-packet -`AckV1`, and that backflow arrives continuously while the backlog drains, -advancing `inbound_total` (`nym_driver.rs:372-384`, -`shim/src/nym.rs:237-251`, `:1027-1062`). A backlog large enough to strand the probe is -therefore also a backlog generating inbound traffic, and the `_` arm at -`nym_driver.rs:343-350` resets `silent_rounds` to 0. The states in which no -backflow arrives — the hub's mixnet client dead, the addresses all stale, the -gateway genuinely not delivering — are states in which those frames were -undeliverable anyway, so the rebuild is not what destroys them. - -This is exactly the asymmetry the dedicated G21 pass established from the pinned -SDK (`globals/G21-…` §6.3): the hub is vulnerable to the equivalent false-Silent -because a SURB reply generates nothing in return, and **the shim is not**. The -harm is claimed only for `Step::Stop` (and, as an accounting defect, for -`Step::Died`); the probe leg is carried as recommendation 6. - -The *scheduled-rotation* trigger for `Step::Rebuild` is a separate, still-open -filing (`nym-rotation-deferral-cannot-protect-submits-so-rotating-destroys-acknowledged-migrations.md`) -and is deliberately not re-argued here. - -### Two things added that the filing did not have - -1. **The loss window is adversary-widenable, and the filing said the opposite.** - Its "what limits the attack" bullet named "the pipeline is usually empty" as - the mitigating factor. The confirmed High - `junk-sendtransaction-flood-…` removes it: ~1 byte/second of unauthenticated - junk holds the ~50-slot pipeline permanently full, so a real migration - entering it sits behind ~270 seconds of junk with a success already returned, - and **any** restart during that window destroys it. The mitigating factor is - now stated as removable, with its price. -2. **The 10-s-versus-0-s asymmetry, and the third exit.** The shim already - contains the mechanism this issue asks for — a bounded drain — and applies it - to the leg where nothing has been promised (`proxy.rs:545-553`) while giving - zero seconds to the leg where success has already been returned. And a fatal - `accept()` error (`proxy.rs:499`) exits without the supervisor ever being told - to disconnect, so that path violates the module's own D12 rule outright. - -### Corrections to the filed text - -- **`hub.rs:231-233` corrected to `hub.rs:236-240`** (the `Ok(()) => Ok(Submit::Accepted { … })` arm). -- **`nym.rs:376-381` corrected to `nym.rs:377-381`** (the 20-consecutive-failure note). -- **`nym_driver.rs:601-607` corrected to `nym_driver.rs:598-607`** (`send_frame`'s doc comment). -- The severity claim that the truncated `disconnect()` matters in itself was kept - deliberately small: at process exit the leaked SDK tasks die with the process. - What matters is that the mechanism the module exists to provide is not - delivered on the one path that always runs, and that this removes any window in - which a drain could happen. -- The title and framing were narrowed from "all three teardown paths" to the - shutdown path plus the accounting defect on the other two, per the demotion - above. - -### Severity: Medium, and why not higher or lower - -- **Not Low.** It fires on every operator redeploy without any attacker; the - quantity lost is adversary-controlled for about one byte per second; the loss - is of transactions the user was explicitly told had succeeded, in a pool the - network has closed; and it is invisible to everyone — there is no counter, no - log line and no metric anywhere in either binary that records it. -- **Not High.** Each event destroys at most a few dozen frames, and at today's - observed traffic an untimed restart usually destroys nothing unless an attacker - is actively holding the pipeline full or the operator is deliberately aiming - it. The two attacker-driven variants each require another party's capability - (the confirmed High flood, or operator control of the process), and the - worst-case volume is bounded by a ~50-slot pipeline rather than being - open-ended. - -### Distinct from, and not double-counting, three neighbours - -- `nym-submit-acks-are-never-read-so-every-hub-refusal-is-invisible.md` - (**Confirmed, Medium**) owns *refusals being invisible* — the hub receives the - frame and says no, and nobody hears it. This issue owns *teardown destroying - in-flight work* — the hub never receives the frame at all. Different mechanism, - different fix (read the ack vs. drain and join). -- `junk-sendtransaction-flood-…` (**Confirmed, High**) owns the flood itself and - is cited here only as the thing that removes this issue's limiting factor. -- `nym-rotation-deferral-cannot-protect-submits-…` (plausible) owns the - scheduled-rotation trigger for `Step::Rebuild`. - -DO NOT DEVIATE FROM THIS TEMPLATE (`docs/SECURITY-ISSUE-TEMPLATE.md`). -LEAVE THIS MESSAGE IN PLACE SO THAT ALL AGENTS KNOW NOT TO DEVIATE FROM THIS TEMPLATE (`docs/SECURITY-ISSUE-TEMPLATE.md`). diff --git a/zeronym-22aa9851caf68-high-medium/medium/shim-proxy-unbounded-inbound-concurrency-enclave-memory-exhaustion.md b/zeronym-22aa9851caf68-high-medium/medium/shim-proxy-unbounded-inbound-concurrency-enclave-memory-exhaustion.md deleted file mode 100644 index 8db0e3aa..00000000 --- a/zeronym-22aa9851caf68-high-medium/medium/shim-proxy-unbounded-inbound-concurrency-enclave-memory-exhaustion.md +++ /dev/null @@ -1,509 +0,0 @@ -# Nothing bounds the shim's *total* inbound buffering: ~4 MiB per in-flight `SendTransaction`, an uncapped number of concurrent requests, and no timeout of any kind on the wallet leg, against a fixed 2 GB enclave that nothing restarts - -**Severity**: Medium -**Validation Status**: Confirmed -**Location**: -`audit-target/zeronym/shim/src/proxy.rs:478-553` (`serve_with_shutdown`'s accept loop — `tokio::spawn` per connection at `:516`, no semaphore, no counter), -`audit-target/zeronym/shim/src/proxy.rs:575-608` (`serve_connection`; the h2 server builder at `:602-606` sets only the two window sizes — no `max_concurrent_streams`, no `.timer()`, no keep-alive), -`audit-target/zeronym/shim/src/proxy.rs:592` (one `UpstreamPool` per inbound connection), -`audit-target/zeronym/shim/src/proxy.rs:174,177` (`STREAM_WINDOW` 2 MiB, `CONNECTION_WINDOW` 8 MiB), -`audit-target/zeronym/shim/src/proxy.rs:271-287` (`upstream_h2_builder`, the *outbound* leg that does get a timer and keep-alives), -`audit-target/zeronym/shim/src/intercept.rs:68-81` (`MAX_SEND_TX_BYTES` = 4 MiB, `MAX_TX_FILTER_BYTES` = 1 KiB and the comment that describes this exact attack), -`audit-target/zeronym/shim/src/intercept.rs:94-131` (`send_transaction`: the `Limited(...).collect()` at `:102`, the `to_bytes()` copy at `:108`), -`audit-target/zeronym/shim/src/intercept.rs:559-568` (`RawTransaction::decode`, the second full copy), -`audit-target/zeronym/shim/src/wire.rs:124,274-277` (`MAX_NYM_TX_BYTES` = 65,503 — the ceiling on anything the divert path can actually carry), -`audit-target/zeronym/shim/deploy/caution/caution.hcl.tmpl:29-36` (`cpu = 2`, `memory_mb = 2048`), `:51-55` (`ingress 0.0.0.0/0` on 8083), `:97-125` (the in-enclave Caddy that maps 443 onto it) -**Found by agent:** Local (file audit of `shim/src/proxy.rs`), merged at validation with the independently-filed `shim/src/intercept.rs` finding `send-transaction-4mib-buffer-memory-exhaustion.md`; validated 2026-08-18 -**In scope of audit?** Yes - -> **MERGE NOTE.** This file is the single owner of the shim's inbound -> memory-exhaustion finding. `issues/invalid/send-transaction-4mib-buffer-memory-exhaustion.md` -> was filed independently from the `intercept.rs` side of the same defect; its -> substance was validated and found **real**, and it now lives in `invalid/` -> **for bookkeeping only** so the harm is not counted twice. Its three unique -> contributions — the ~8 MiB completion peak, the 64x gap between -> `MAX_SEND_TX_BYTES` and `MAX_NYM_TX_BYTES`, and the fact that the project -> already wrote this attack down and fixed only the other half of it — are -> folded in below. - -## Description - -The shim buffers a whole `SendTransaction` body before it can classify it. That -buffer is capped **per stream** at `MAX_SEND_TX_BYTES = 4 MiB` -(`intercept.rs:71`, `:102`). Nothing caps the aggregate: - -- **No connection cap.** The accept loop spawns a task per accepted socket with - no semaphore, no counter and no admission control (`proxy.rs:478-553`, spawn - at `:516`). `grep -rn "Semaphore\|max_concurrent_streams\|MAX_CONN" shim/src/` - returns nothing. -- **No aggregate byte budget.** There is no process-wide accounting of how much - all in-flight requests are holding. The only number in the system is the - per-stream constant. -- **No timeout of any kind on the wallet leg.** The server h2 builder - (`proxy.rs:602-606`) sets only `initial_stream_window_size` and - `initial_connection_window_size`. There is no request deadline, no idle - timeout, no keep-alive and no timer. `Limited` is a *byte* cap, not a - *duration* cap: a peer that sends 4 MiB minus one byte and then stops without - `END_STREAM` leaves `collect()` pending forever, holding every byte it - received. - -Two multipliers make the per-stream figure worse than 4 MiB: - -1. **The completion peak is ~8 MiB, not 4 MiB.** When a body does complete, - `Collected::to_bytes()` (`intercept.rs:108`) allocates a fresh contiguous - buffer of the full length *before* draining the chunk list - (`http-body-util-0.1.4/src/util.rs`, `BufList::copy_to_bytes` → - `BytesMut::with_capacity(len)`), so both live at that instant. Then - `RawTransaction::decode` allocates `raw.data` as a second full copy - (`intercept.rs:559-568`) while `frame` is still held for the pass-through - replay (`intercept.rs:129`) or the divert. An attacker who parks N streams at - 4 MiB and then sends the last byte of all of them at once turns N x 4 MiB - resident into ~N x 8 MiB in one step, with no extra upload. -2. **4 MiB is 64x larger than anything the divert path can carry.** The mixnet - transport refuses any transaction over `MAX_NYM_TX_BYTES` = 65,503 - (`wire.rs:124`, `:274-277`), so an Orchard-touching body above ~64 KiB is - buffered at up to 4 MiB, copied twice, and *then* refused with - `RESOURCE_EXHAUSTED` (`intercept.rs:167-179`). Only the pass-through path - needs headroom at all, and the constant's own doc comment cites the ~2 MB - Zcash transaction limit as the number it is "well above" — i.e. it is - deliberately double what it needs to be. - -**The project already wrote this attack down and closed only the other half of -it.** `MAX_TX_FILTER_BYTES` was introduced at 1 KiB for the `GetTransaction` -path with this comment (`intercept.rs:73-81`): - -> This used to share `MAX_SEND_TX_BYTES`, which was 4000x looser than the request -> can ever legitimately be, and the looseness had a price: hyper allows ~200 -> streams per connection and connections are uncapped, so a hostile client -> trickling near-4 MiB bodies on many streams could pin gigabytes in an enclave -> whose memory is mostly EnclaveOS. A kilobyte refuses nothing a wallet sends and -> takes that lever away. - -The lever was taken away on the path that never needed a large buffer. On -`SendTransaction` — the path that *must* buffer, and therefore the one a -per-stream constant cannot fix — it is untouched, and line 70's claim that "a -hostile client cannot make the shim buffer unbounded memory" is true per stream -and false in aggregate. - -## Attack Scenario and Steps - -1. Connect to the shim's public endpoint. In the shipped deployment that is the - in-enclave Caddy on 443, which terminates TLS and forwards h2c to the shim on - 8083 (`caution.hcl.tmpl:97-125`); `ingress` is `0.0.0.0/0` (`:51-55`). No - credentials, no wallet, no valid transaction. -2. Open many concurrent HTTP/2 streams, each a - `POST /cash.z.wallet.sdk.rpc.CompactTxStreamer/SendTransaction`. `route_for` - matches on path only, so this reaches `intercept::send_transaction` - unauthenticated. -3. On each stream send just under 4 MiB of DATA and then **stop**, without - `END_STREAM` and without a `RST_STREAM`. `Limited` only errors when the limit - is *exceeded*, so nothing errors; `collect()` stays pending and holds every - byte. HTTP/2 flow control does not bound this: hyper releases receive capacity - as soon as it hands each chunk to the body - (`hyper-1.11.0/src/body/incoming.rs:245`, `h2.flow_control().release_capacity(bytes.len())`), - so the 2 MiB stream / 8 MiB connection windows only pace the upload, they do - not cap it. -4. Add streams and connections until the enclave's memory is gone. Optionally, - send the final byte of every parked stream simultaneously to double the peak - (mechanism 1 above) instead of uploading twice as much. -5. Nothing reaps the stalled streams. Once uploaded, the memory is pinned for - free for as long as the attacker leaves the sockets open. - -**Attack Requirements and Assumptions:** - -- **Access needed:** the ability to open TCP connections to a public endpoint. - Unauthenticated and remote. -- **Concurrency is not capped anywhere on the path.** hyper 1.11.0 caps the - *shim* at 200 streams per connection — verified, not assumed: - `Config::default()` sets `max_concurrent_streams: Some(200)` - (`hyper-1.11.0/src/proto/h2/server.rs:69`) and applies it to the h2 builder - (`:143-144`). But nothing caps *connections*, and in the shipped topology the - shim's peer is Caddy, not the attacker: Caddy's Go HTTP/2 transport opens - additional backend connections once a backend's stream limit is reached, and - Caddy's own server accepts an uncapped number of client connections. So the - attacker's aggregate in-flight request count is bounded only by what they - choose to open. -- **Nothing in front of the shim absorbs it, and this was checked rather than - assumed.** The Caution platform renders the in-enclave Caddyfile from - `src/enclave-builder/templates/run.sh.template:107-121`; it is three `handle` - blocks and a bare `reverse_proxy {{CADDY_UPSTREAM}}` with **no `request_body` - size limit, no rate limit, and no timeout directives**. Caddy's - `reverse_proxy` streams request bodies by default (`request_buffers` unset), so - it relays a slow, incomplete body rather than absorbing it. Nor is there a - second way in that bypasses Caddy: under `e2e_encryption { mode = "tls" }` the - builder *excludes* the `http_port` from the per-port vsock relays - (`src/enclave-builder/src/build.rs:421-435`), so 8083 is reachable only through - Caddy — and Caddy does not change the conclusion. -- **Cost ratio.** This is a bandwidth-priced flood, not an amplifier: roughly one - uploaded byte per one to two bytes of enclave memory at the peak. Pinning - ~1.6 GB costs ~0.8-1.6 GB of upload, once, from a single ordinary host — tens - of seconds on a commodity VPS — after which the hold is free and can be - repeated at will. This ratio is the single reason this issue is Medium rather - than High; see the severity note in the Validation Information. -- **What is *not* required:** no wallet, no valid transaction, no position on the - mixnet, no operator privilege, no interaction with the classifier. - -## Impact on Users - -The enclave is `memory_mb = 2048` with no swap and no autoscaling -(`caution.hcl.tmpl:29-36`), and its own comment says "2 GB is almost entirely -EnclaveOS; the process itself sits in single-digit MB". Exhausting it means an -allocation failure (Rust aborts) or an OOM kill of the shim. - -- **The shim does not come back on its own.** The shim's process is the last - command in the enclave's `run.sh`, so when it dies the script exits and the - enclave terminates. On the parent, the `nitro-enclave.service` unit runs - `nitro-cli run-enclave ... && tail -f /dev/null` with `Restart=on-failure` - (Caution platform, `terraform/modules/aws/nitro-enclave/user-data.sh:157-175`): - `run-enclave` returns as soon as the enclave is launched and `tail -f` keeps - the unit alive forever, so systemd never observes the enclave's death and never - restarts it. There is no watchdog elsewhere in the platform tree. **A - successful attack is therefore an outage that lasts until a human notices** — - and what they must then do is redeploy an immutable attested enclave, which on - this project spends a certificate issuance and re-registers a mixnet identity. -- **Users are pushed off the protected path, permanently and on-chain.** A shim - that is down answers nothing. A user who must complete a mandatory - Orchard->Ironwood migration and cannot will point their wallet at a different, - unprotected indexer — which is the exact linkage the product exists to prevent, - and it is written to the public chain forever. This is the "force fallback - behaviour" attacker goal in the threat model, reached with no privileged - position at all. -- **Acknowledged migrations in flight at that instant are destroyed silently.** - `shim/src/hub.rs:236-240` returns `Submit::Accepted` at mixnet hand-off, and - its own comment says a hub refusal "is never surfaced here". A shim that dies - with frames in its pipeline loses them with the wallet already holding - `error_code 0` and a txid. (That harm is owned in detail by - `shim-nym-driver-every-teardown-path-silently-destroys-acknowledged-submits.md`; - it is named here because it is what makes this a privacy/integrity failure - rather than plain downtime.) -- **The adversary chooses the moment**, cheaply and repeatedly — including the - operator, who has a direct incentive to make the private path look unreliable - and who can do this from anywhere without touching their own infrastructure. - -## Technical Details / Code Analysis - -**The buffer, and the two copies** (`shim/src/intercept.rs:94-131`): - -```rust - 94 pub(crate) async fn send_transaction( - 95 req: Request, - 96 pool: Arc, - 97 diversion: Option>, - 98 ) -> Result, BoxError> { - 99 let (parts, body) = req.into_parts(); -100 -101 // The only buffering in the entire shim, and it is bounded. -102 let collected = match Limited::new(body, MAX_SEND_TX_BYTES).collect().await { -103 Ok(collected) => collected, -104 Err(err) => return Ok(body_read_failed(err)), -105 }; -106 -107 let trailers = collected.trailers().cloned(); -108 let frame = collected.to_bytes(); -109 -110 let (inspection, tx_data) = inspect(&parts.headers, &frame); -... -128 let upstream = pool.get().await?; -129 let replay = ReplayBody::new(frame, trailers).boxed(); -``` - -The comment on line 101 is true of one stream and of nothing else. -`Limited::poll_frame` errors only when a frame's payload *exceeds* what is left -(`http-body-util-0.1.4/src/limited.rs:44-59`), so a peer that stops one byte -short never errors and `collect()` never returns. - -`to_bytes()` is the first copy. `Collected::to_bytes` calls -`BufList::copy_to_bytes(remaining)`, which for a multi-chunk body takes the -`_ =>` arm and does `BytesMut::with_capacity(len)` before draining the chunks -(`http-body-util-0.1.4/src/util.rs`) — both allocations live at that moment. - -`RawTransaction::decode` is the second (`shim/src/intercept.rs:559-568`): - -```rust -559 match RawTransaction::decode(message) { -... -562 Ok(raw) => { -563 let evidence = classify_with_evidence(&raw.data); -564 ( -565 Inspection::Classified(evidence), -566 Some(Bytes::from(raw.data)), -567 ) -568 } -``` - -`raw.data` is a fresh `Vec` holding the whole declared transaction while -`frame` is still alive in the caller. A garbage 4 MiB body decodes as a valid -`RawTransaction`, classifies as `Unparseable`, fails safe toward diversion, and -is then refused by `encode_submit` for exceeding `MAX_NYM_TX_BYTES` -(`wire.rs:274-277`) — so the two copies are paid in full to produce a refusal, -and no mixnet traffic is generated. - -**The accept loop, with nothing bounding it** (`shim/src/proxy.rs:478-543`, -abridged): - -```rust -478 loop { -479 tokio::select! { -480 biased; -481 () = &mut shutdown => break, -482 accepted = listener.accept() => { -... -516 tokio::spawn(async move { -517 let _live = live; -518 match tls { -519 None => { -520 serve_connection(stream, peer, backend, diversion, caution, status) -521 .await -522 } -``` - -`live_tx` is a shutdown tracker, not a limiter: nothing is ever sent on it and -nothing acquires anything. Every accepted socket becomes an unbounded task. - -**The server h2 configuration** (`shim/src/proxy.rs:602-608`): - -```rust -602 if let Err(err) = server_h2::Builder::new(TokioExecutor::new()) -603 .initial_stream_window_size(STREAM_WINDOW) -604 .initial_connection_window_size(CONNECTION_WINDOW) -605 .serve_connection(TokioIo::new(stream), service) -606 .await -607 { -608 tracing::debug!(%peer, %err, "client connection ended"); -``` - -Two window sizes and nothing else. Compare the *outbound* builder in the same -file (`proxy.rs:271-287`), which was given a timer, an interval and a timeout, -with a doc comment explaining that `.timer()` is "load-bearing: hyper silently -disables keepalive (and every other timed behaviour) when no timer is -installed". The inbound leg got none of it. - -**Why "just add `.timer()`" is not the fix** (verified against the pinned -crate, and this inverts the recommendation both original filings made): hyper's -HTTP/1 server has a `header_read_timeout` that a missing timer silently -disables — that is the hazard the hub's `server.rs:392-398` records and -correctly closes. **hyper's HTTP/2 server has no equivalent at all.** Its only -timed mechanism is the keep-alive PING, and `Config::default()` sets -`keep_alive_interval: None` (`hyper-1.11.0/src/proto/h2/server.rs:70-73`). So a -timer alone enables nothing; the shim needs a timer *plus* an explicitly -configured keep-alive, and separately an explicit deadline around the -`Limited(...).collect()` in `intercept.rs`, because neither hyper nor h2 will -ever impose one. - -**A second, cheaper mechanism against a different shared resource.** -`serve_connection` constructs **one `UpstreamPool` per inbound connection** -(`proxy.rs:588-592`), so N inbound connections that each issue one trivial -pass-through request cause the shim to hold N open TCP (and, in the shipped -config, TLS) connections to the operator's indexer, 1:1 with attacker -connections and with no pool and no cap. That consumes enclave file descriptors -and indexer front-end slots. (The *privacy* consequence of the same code — -one upstream session per wallet session — is a separate finding, -`one-upstream-connection-per-wallet-connection-makes-the-operators-indexer-a-per-wallet-session-boundary-and-only-caddys-pooling-hides-it.md`; -only the resource consequence is claimed here.) - -**The envelope** (`shim/deploy/caution/caution.hcl.tmpl:29-36`, `:51-55`): - -```hcl - resources { - # The shim is stateless: ... 2 GB is almost entirely EnclaveOS; the process - # itself sits in single-digit MB. - cpu = 2 - memory_mb = 2048 - } -... - ingress { - cidr_ipv4 = "0.0.0.0/0" - port = 8083 - ip_protocol = "tcp" - } -``` - -## Recommendations - -Ordered by how much each one removes, and corrected against the pinned crates. - -1. **Add an explicit per-request deadline around the buffering step.** Wrap - `Limited::new(body, MAX_SEND_TX_BYTES).collect()` in `intercept.rs:102` (and - the `MAX_TX_FILTER_BYTES` one at `:249`) in a `tokio::time::timeout`, and - answer an expiry with the existing `body_read_failed` shape. This is the - single change that closes the hold; a byte cap can never do it, and hyper's - h2 server will not do it for you. -2. **Add an aggregate byte budget.** A process-wide `Arc` (or a - `Semaphore` of buffering slots) charged on entry to `send_transaction` and - released on exit, sized to what 2 GB minus EnclaveOS can actually hold, so - the bound is on *total* buffered bytes rather than on a per-stream constant - multiplied by an unbounded stream count. Refuse over-budget requests with - `RESOURCE_EXHAUSTED`. -3. **Lower `MAX_SEND_TX_BYTES`** to just above the largest transaction the shim - must pass through — the ~2 MB Zcash limit its own doc comment cites — rather - than double it. Nothing divertible exceeds `MAX_NYM_TX_BYTES` = 65,503, so - the extra headroom serves only the pass-through path and buys nothing there - either. -4. **Cap concurrent inbound connections** in the accept loop - (`proxy.rs:478-543`) with a `tokio::sync::Semaphore` whose permit is held by - the spawned connection task; at minimum, count and log them so the failure - mode is visible before it is fatal. -5. **Set `.max_concurrent_streams(...)` explicitly** on the server builder and - pick it together with (2) and (3) so `max_streams x buffer` fits the enclave - with margin. hyper's 200 is a default, not a contract, and the arithmetic - above depends on it. -6. **Install `.timer(TokioTimer::new())` *and* an explicit - `keep_alive_interval`/`keep_alive_timeout`** on the server builder, so a - vanished (as opposed to merely slow) peer is reaped. Note the correction: - `.timer()` alone is inert on hyper's h2 server — it is necessary but not - sufficient, and it is not a substitute for (1). -7. Avoid the second copy where it is free to do so: `inspect` could classify - from a borrowed slice of `frame` and only materialise `tx_data` on the divert - path, which removes ~4 MiB of the completion peak per stream. - -## Validation Information - -**Verdict: CONFIRMED, Medium (top of Medium).** Validated 2026-08-18. This file -now carries the merged finding; `send-transaction-4mib-buffer-memory-exhaustion.md` -has been moved to `invalid/` for bookkeeping only, with a header stating that its -substance is real and is owned here. - -### Why the two filings are one finding - -They describe the same harm (unauthenticated inbound requests pin the shim -enclave's memory) from the two files that jointly cause it, and they share every -step of the attack, every precondition and most of the fix list. `intercept.rs` -supplies the multiplicand (4 MiB per stream, doubled at completion); -`proxy.rs` supplies the multiplier (no connection cap, no aggregate budget, no -deadline). Counting them separately would double-count one outage. The first -filing said so itself in its own opening paragraph and asked to be merged. - -### Every mechanical claim re-verified - -| Claim | Verified at | -|---|---| -| `SendTransaction` is routed on path only, unauthenticated | `shim/src/proxy.rs:744-782` (`route_for`), `:70` (`SEND_TRANSACTION`) | -| The body is buffered whole, capped per stream at 4 MiB | `shim/src/intercept.rs:71`, `:102` | -| `Limited` errors only on *exceeding* the cap, so stopping one byte short never errors | `http-body-util-0.1.4/src/limited.rs:44-59` | -| Flow control does not bound the total: capacity is released as each chunk is handed to the body | `hyper-1.11.0/src/body/incoming.rs:245` | -| `to_bytes()` allocates a full second buffer before draining the chunks | `http-body-util-0.1.4/src/util.rs`, `BufList::copy_to_bytes`, `_ =>` arm | -| `RawTransaction::decode` allocates a second full copy while `frame` is live | `shim/src/intercept.rs:559-568`, `:108`, `:129` | -| hyper 1.11.0 really does cap the shim at 200 streams/connection | `hyper-1.11.0/src/proto/h2/server.rs:69`, `:143-144` (`max_concurrent_streams: Some(200)`) | -| hyper's h2 server has no header/request-read timeout at all; keep-alive defaults to `None` | `hyper-1.11.0/src/proto/h2/server.rs:70-73` | -| No connection cap, no semaphore, no counter in the accept loop | `shim/src/proxy.rs:478-553`; `grep -rn "Semaphore\|max_concurrent_streams\|MAX_CONN" shim/src/` is empty | -| No timer, no keep-alive, no stream cap on the server builder | `shim/src/proxy.rs:602-606`, contrasted with `:275-287` | -| One `UpstreamPool` per inbound connection | `shim/src/proxy.rs:588-592` | -| Anything divertible is <= 65,503 bytes, so 4 MiB is 64x more headroom than the divert path can use | `shim/src/wire.rs:124`, `:274-277`; refusal at `shim/src/intercept.rs:167-179` | -| Enclave is 2 vCPU / 2 GB, internet-reachable, TLS terminated by an in-enclave Caddy | `shim/deploy/caution/caution.hcl.tmpl:29-36`, `:51-55`, `:97-125` | - -Facts established from the Caution platform tree (out of audit scope, used only -to test whether infrastructure absorbs the attack, and marked as such): - -| Claim | Verified at | -|---|---| -| The in-enclave Caddyfile is a bare `reverse_proxy` — no body limit, no rate limit, no timeouts | `src/enclave-builder/templates/run.sh.template:107-121` | -| Port 8083 gets **no** direct vsock relay under `mode = "tls"`, so Caddy is the only way in | `src/enclave-builder/src/build.rs:396-435` | -| An enclave that exits is not restarted: `run-enclave` returns and `tail -f /dev/null` keeps the unit alive, so `Restart=on-failure` never fires | `terraform/modules/aws/nitro-enclave/user-data.sh:157-175` | - -### `AVOIDING-FALSE-POSITIVES.md` §5 applied honestly, in both directions - -§5's canonical false positives are *"unlimited concurrent connections … each -connection uses minimal resources"* and *"large file upload DoS … requires the -attacker to have 1 GB of upload bandwidth per attempt; CDN/proxy usually limits -to much less."* This issue sits between §5's two poles and I am recording which -half applies where rather than picking the flattering reading: - -**Against the false-positive pattern (why this is real):** - -- A connection here does **not** use minimal resources: one connection can hold - 800 MiB, and 200 simultaneous completions on it peak at ~1.6 GB. §5 names - *"single connection consuming unbounded resources"* as the real-issue shape. -- The "CDN/proxy limits it to much less" mitigation was checked against the - actual platform template and **does not exist**: no body cap, no rate limit, - no timeout, and no route that bypasses Caddy either way. -- "OS limits apply first" is not a mitigation here — the limit that applies - first is a hard 2 GB with no swap, and reaching it is the attack. -- The memory is retained after the upload at zero marginal cost, so it is not - "per attempt"; and the target does not restart itself, so one success is a - standing outage. - -**For the false-positive pattern (why this is not High):** - -- The cost ratio is ~1:1 to ~1:2, not §5's "1 KB request causing 1 GB - allocation". The attacker must actually push on the order of a gigabyte. That - is trivially affordable, but it is a bandwidth-priced flood, not an amplifier, - and it is a genuinely different economic class from the two confirmed Highs on - this same listener (`junk-sendtransaction-flood-…` holds the divert pipeline - permanently full at ~1 byte/second; `gettransaction-flood-…` converts ~100 - bytes into 61 sphinx packets). -- Those two confirmed Highs already take migration diversion down for less - money. The incremental harm this issue adds is "hard process death with no - self-recovery" rather than "starvation", which is worse per event but is not a - new class of victim. -- How much of the 2 GB is actually free could not be measured here; the enclave - image (Caddy, socat, bootproofd, busybox, the shim) is loaded into that same - RAM, so the true free figure is smaller than 2 GB — which makes the attack - cheaper, not more expensive, but the exact number is an inference. - -Net: a cheap, unauthenticated, remote, repeatable DoS whose terminal state is a -privacy failure rather than mere downtime, priced in bandwidth rather than in -amplification. **Medium, at the top of Medium.** - -### Corrections applied to the two filings - -1. **The `max_concurrent_streams` caveat is struck.** The original - `intercept.rs` filing said the ~200 figure "should be verified rather than - relied on; if the default advertises no `SETTINGS_MAX_CONCURRENT_STREAMS`, - the per-connection stream count is bounded only by the peer." It is verified: - hyper 1.11.0 sets `Some(200)`. (h2 0.4.15 alone would indeed be `usize::MAX` - — `frame/settings.rs` leaves it `None` and `Counts::new` uses - `unwrap_or(usize::MAX)` — but hyper sets it before handshaking.) The comment - at `intercept.rs:74-80` is correct as written. -2. **Both filings' `.timer()` recommendation is inverted.** Recommendation 3 in - each ("install `.timer()` … so a stalled client is reaped") does not work: - hyper's h2 server has no `header_read_timeout` analogue, and its only timed - mechanism defaults to off. The fix is an explicit request deadline; the timer - is only a prerequisite for the keep-alive half. This is now recommendation 1 - with the timer demoted to 6. -3. **"800 MiB from a single TCP connection" is kept but scoped.** It is exact for - the shim's h2 peer. In the shipped topology that peer is Caddy, not the - attacker, so the attacker's own connection count and the shim's are - decoupled; the accurate statement is that neither hop caps aggregate - concurrency, which the text now says. -4. **The "a validator should confirm the deployed Caddy configuration" caveat is - discharged** with the platform template quoted above, and the "could Caddy - absorb it?" question answered no. The stronger new fact — that 8083 has no - direct vsock relay under `mode = "tls"` — is also recorded, because it - narrows the ingress to one hop rather than two. -5. **The `UpstreamPool` leg is kept as a resource claim only**, with its privacy - half explicitly deferred to the file that owns it, so the two are not - double-counted. -6. **A "1.6 GB momentary" figure in the original `proxy.rs` filing is restated - correctly.** It read as though completing 200 bodies needs 1.6 GB *in addition - to* the 800 MiB; it does not — the 800 MiB is one of the two copies. The peak - is ~1.6 GB total, reached from 800 MiB of upload, which is the sharper claim. -7. **Recommendation 7 (classify from a borrowed slice) is new**, and is the - cheapest single change that removes half the completion peak. - -### What this issue does *not* claim - -- It does not claim an amplification attack. The ratio is stated as ~1:1-1:2 and - the severity rests on that. -- It does not claim the empty-`DATA`-frame trick is a memory attack. - `Collected::push_frame` explicitly skips empty data frames - (`http-body-util-0.1.4/src/collected.rs:41-52`), so zero-length frames hold a - stream (~1 KB of h2 state) without accumulating bytes. That is a duration hole - worth a sentence, already covered by recommendation 1, and it is not part of - the arithmetic here. -- **It does not claim the response direction.** A lead worth a separate pass and - deliberately excluded from this confirmed finding: nothing bounds the - *outbound* side either — hyper's server `max_send_buffer_size` defaults to - 400 KB per stream (`hyper-1.11.0/src/proto/h2/server.rs:39`, `:75`) and the - shim's upstream connection window is 8 MiB, so a ~100-byte `GetBlockRange` - whose response the client never reads may pin hundreds of kilobytes of - *indexer-supplied* bytes per stream at essentially zero cost to the attacker. - That would be a genuine amplifier, but whether an attacker can drive it through - the in-enclave Caddy's own buffering could not be established here, so it is - recorded in `BRAINSTORM.md` as a lead rather than asserted. -- It does not re-argue the acknowledged-migration loss on shim death; that is - owned by - `shim-nym-driver-every-teardown-path-silently-destroys-acknowledged-submits.md` - and referenced only as the reason this is not plain downtime. - -DO NOT DEVIATE FROM THIS TEMPLATE (`docs/SECURITY-ISSUE-TEMPLATE.md`). -LEAVE THIS MESSAGE IN PLACE SO THAT ALL AGENTS KNOW NOT TO DEVIATE FROM THIS TEMPLATE (`docs/SECURITY-ISSUE-TEMPLATE.md`). diff --git a/zeronym-22aa9851caf68-high-medium/medium/shim-route-for-no-tainted-address-routing-readme-amount-overclaim.md b/zeronym-22aa9851caf68-high-medium/medium/shim-route-for-no-tainted-address-routing-readme-amount-overclaim.md deleted file mode 100644 index e3d81dda..00000000 --- a/zeronym-22aa9851caf68-high-medium/medium/shim-route-for-no-tainted-address-routing-readme-amount-overclaim.md +++ /dev/null @@ -1,381 +0,0 @@ -# `README.md:33`'s "not the amount or which transaction" is false: the reference light-client SDK hands the operator's own indexer the wallet's whole transparent address set on every sync, and the operator's indexer then serves back the txid and the value of any diverted Orchard-to-transparent spend - -**Severity**: Medium -**Validation Status**: Confirmed -**Location**: The claim at `audit-target/zeronym/README.md:33` (and the disclosure it contradicts at `:32`); the routing table that lets it happen at `audit-target/zeronym/shim/src/proxy.rs:744-783` (`route_for`) and `:702-733` (`Route`); the relay at `:787-832` (`pass_through`); the project's own statement of the leak at `audit-target/zeronym/shim/ENDPOINTS.md:130-132`; the stateless-shim commitment that makes the designed fix unimplementable as written at `audit-target/zeronym/shim/src/intercept.rs:56-63` -**Found by agent:** Local (file audit of `shim/src/proxy.rs`); mechanism corrected and re-based on primary sources during validation -**In scope of audit?** Yes — `audit-context/AUDIT-INSTRUCTIONS.md` puts `README.md` in scope "as security claims" and asks that an overclaim be treated as an ICTM finding. Extra focus area 11 is exactly this. - -> **FILENAME NOTE.** The filename is retained because `PROGRESS.md`, -> `audit-state/globals/G11-G33-…` (row **N2**, chain **C**) and -> `issues/confirmed/readme-volume-independent-source-ip-claim-overstated.md` -> all reference it. **Use the title, not the filename.** The original title framed -> this as "`route_for` implements none of the `ENDPOINTS.md` INTERCEPT set"; that -> framing was demoted during validation because `ENDPOINTS.md` disclaims being -> implemented and is separately owned (see *Ownership* below). The finding is the -> README sentence and the concrete leak behind it. - -## Description - -`README.md`'s Security section carries two bullets, four lines apart, on opposite -sides of the *Protected* / *Not protected* boundary. The first states a -mechanism; the second denies that mechanism's consequence. - -`README.md:32` (accurate, and the sharpest disclosure in the document): - -> - **Query content.** Address-level queries (`GetTaddressTxids`, `GetTaddressBalance`, -> `GetAddressUtxos`) are not intercepted and still reach the operator. - -`README.md:33` (false for any diverted transaction with a transparent leg the -wallet owns): - -> - **The operator learns *that* a client migrated**, though not the amount or which -> transaction. - -The project already states the join in its own words. `shim/ENDPOINTS.md:132`, in -the `GetTaddressBalance` row of the INTERCEPT table: - -> **The sharpest amount leak:** a balance poll bracketing the flush yields -> post-minus-pre = the exact deshielded amount, turning "operator learns *that* a -> client migrated" into "amount Y". - -and `:130`, for `GetTaddressTransactions` / `GetTaddressTxids`: - -> Names the migration's transparent leg; the operator joins IP C to the on-chain -> batched tx once the hub publishes. - -None of that routing exists. `route_for` (`proxy.rs:744-783`) is the shim's entire -method-routing table and can distinguish exactly two Zcash methods, -`SendTransaction` and `GetTransaction`. Every transparent-address method returns -`Route::PassThrough` and is relayed verbatim to the operator's indexer. - -**The real mechanism is sharper than the "bracket the flush and subtract" the -issue was originally filed with, and it was established during validation from -the reference SDK rather than inferred.** `zcash_client_backend`'s sync loop calls -`refresh_utxos` on **every sync pass** for every account -(`librustzcash/zcash_client_backend/src/sync.rs:117-126`), and that function sends -**the wallet's complete set of transparent receivers in a single -`GetAddressUtxos` request** (`sync.rs:501-516`). The reply carries, per UTXO, -`txid`, `index`, `valueZat` and `height` -(`lightwalletd/walletrpc/service.proto:212-219`). So the operator does not have to -difference two balances: their own indexer *serves the wallet the txid and the -value*, and the request that asked for it is a durable per-wallet fingerprint. - -## Attack Scenario and Steps - -The adversary is the indexer operator — the primary adversary and the reason the -product exists. No privileged position, no configuration change, no active step: -the data arrives in their indexer's ordinary request log. - -1. A wallet holding legacy Orchard notes spends them to a transparent address it - owns (an unshield — the ordinary "move my Orchard funds to my transparent - receiver" flow). `classify::is_orchard_touching` is true, so the shim diverts - the broadcast to the hub and the operator's indexer is never dialled. -2. On its next sync, the wallet calls `GetAddressUtxos` with **every transparent - receiver of the account** in one request (`sync.rs:501-509`). `route_for` - returns `Route::PassThrough`, and `pass_through` (`proxy.rs:787-832`) relays it - to the operator's indexer unchanged. The operator now holds this wallet's - transparent address set — a stable identifier that survives IP changes, - reconnections and the enclave's connection multiplexing. -3. Once the hub publishes on its 20-block cadence and the transaction is mined, - the same recurring request returns the new UTXO. The operator's own indexer - answers it with `txid`, `valueZat` and `height`. -4. The operator now has, for one identifiable client of theirs: **which - transaction** (`txid`, directly, not inferred) and **the amount** (`valueZat`, - directly). Both are the two things `README.md:33` says they do not get. They - can also confirm it was diverted, since an Orchard-touching transaction is by - construction one they never saw broadcast. -5. `GetTaddressTxids` (`service.proto:285`) gives the same answer a second way, - and `GetTaddressBalance` (`:291`) gives the `ENDPOINTS.md:132` differencing - attack as a third. - -**Attack Requirements and Assumptions:** - -- **Access needed: none beyond running the indexer**, which is the deployment - model. The queries arrive unsolicited on the pass-through path. -- **No parent-host position is required** — this is what distinguishes it from the - other two routes to the same sentence (see *Ownership*). It works in the shipped - fully-managed Caution deployment (`deploy.sh:156` → `caution apps create`), where - validation of `readme-volume-independent-source-ip-claim-overstated.md` - established that the operator does **not** hold the accepting socket and cannot - get the wallet's source IP from their own indexer. -- **Attribution without an IP.** Because the enclave has no NIC and Caddy - multiplexes wallets onto one shim connection, the indexer sees neither the - client address nor a reliable per-wallet connection boundary. The address set in - the request body supplies the missing handle by itself: it is per-wallet, - repeated on every sync, and joinable to the public chain forever. -- **Realistic because it is default behaviour of the reference SDK**, not an - attacker-chosen path. `refresh_utxos` is gated only on the `transparent-inputs` - feature and runs before every shielded scan. -- **Bound, stated honestly: this route does not touch the acute use case.** A - conforming ZIP 318 Orchard-to-Ironwood migration is shielded-to-shielded and has - no transparent leg, so nothing here recovers its amount. The affected class is - the *rest* of the diverted population — every Orchard spend with a transparent - output or input the wallet holds — which the shim diverts because - `is_orchard_touching` is a presence predicate, and which `README.md:33` covers - without qualification. -- **A second bound: the amount is public on chain anyway.** A deshield's output - value is cleartext. What this route supplies is the *linkage* — which of the - operator's clients that on-chain transaction belongs to — which is precisely - what the product exists to break and precisely what `:33` claims survives. - -## Impact on Users - -- A user who moves legacy Orchard funds to their own transparent address through a - zero-indexer endpoint is told, in the Security section they are pointed at, that - the operator does not learn the amount or which transaction. The operator learns - both, from that user's own wallet, on the next sync, with no effort and no - detectable action. The join is permanent, because the chain is permanent and - indexer logs can be replayed later. -- The consequence is the linkage `README.md:54` calls "the attack" — client → - on-chain transaction → value — minus only the IP term in the fully-managed - deployment, and including the IP term under BYOC (`caution init --byoc`) or once - the operator takes the single step described in the confirmed - `core-linkage-…` finding. -- An operator deploying the shim is told at `README.md:70` that "Orchard-touching - transactions and `GetTransaction` lookups stop being yours to see" and may - reasonably conclude their users are protected against the amount join. For - transparent-legged spends they are not. - -## Technical Details / Code Analysis - -**1. The complete routing table.** `shim/src/proxy.rs:744-783`: - -```rust -pub fn route_for(path: &str) -> Route { - if path == SEND_TRANSACTION { - return Route::Intercept; - } - if path == GET_TRANSACTION { - return Route::GetTransaction; - } - // Caution's own endpoints, served on our host: never hand them to the indexer. - if path == CAUTION_HEALTH { return Route::CautionHealth; } - if path == CAUTION_ATTESTATION { return Route::CautionAttestation; } - // The shim's own operator endpoints. - if path == SHIM_HEALTH { return Route::ShimHealth; } - if path == SHIM_NYM_STATUS { return Route::ShimNymStatus; } - if path == SHIM_NYM_DIAG { return Route::ShimNymDiag; } - - let trimmed = path.trim_end_matches('/'); - match trimmed.rsplit('/').next() { - Some(last) if last.eq_ignore_ascii_case("sendtransaction") => Route::InterceptNearMiss, - _ => Route::PassThrough, - } -} -``` - -Of the nine `Route` variants (`:702-733`), three are Zcash methods and the rest -are control-plane or status paths. There is no address handling, no `TaintedAddrs` -structure, no body inspection for any method other than the two intercepted ones, -and no list-splitting for `GetTaddressBalance`. `GetTaddressBalance`, -`GetTaddressBalanceStream`, `GetTaddressTxids`, `GetTaddressTransactions`, -`GetAddressUtxos` and `GetAddressUtxosStream` all land in the `PassThrough` default. - -**2. Where they go.** `shim/src/proxy.rs:664-666`: - -```rust - Route::PassThrough | Route::CautionHealth | Route::CautionAttestation => { - pass_through(req, pool).await - } -``` - -`pass_through` (`:787-832`) dials the operator's indexer and relays the request -head and body verbatim; `forward` (`:840-870`) retargets only the origin, so every -header and the whole body reach the operator unchanged. - -**3. What the wallet actually sends.** `librustzcash/zcash_client_backend/src/sync.rs:501-516`: - -```rust - let request = service::GetAddressUtxosArg { - addresses: db_data - .get_transparent_receivers(account_id, true, true) - .map_err(Error::Wallet)? - .into_keys() - .map(|addr| addr.encode(params)) - .collect(), - start_height: start_height.into(), - max_entries: 0, - }; - … - client.get_address_utxos_stream(request) -``` - -called unconditionally per account from the sync entry point -(`sync.rs:117-126`), *before* shielded scanning. The reply type -(`lightwalletd/walletrpc/service.proto:212-219`) is: - -```protobuf -message GetAddressUtxosReply { - string address = 6; - bytes txid = 1; - int32 index = 2; - bytes script = 3; - int64 valueZat = 4; - uint64 height = 5; -} -``` - -so `txid` and `valueZat` are served by the operator's own indexer, for a -transaction the shim went to some lengths never to show them. - -**4. Why the designed fix is not merely "not built yet".** `ENDPOINTS.md`'s -INTERCEPT set is defined over recognition state — `DivertedMigrations`, -`TaintedAddrs`, `DivertedHeights`, `PendingMigration` (`ENDPOINTS.md:142-157`) — -seeded at divert time. The shipped shim forecloses that by design -(`shim/src/intercept.rs:56-63`): - -```rust -/// Deliberately holds NO state about what it diverted. A stateless shim survives -/// a restart and can run as more than one instance without a follow-up query -/// leaking to the operator, because it recognises nothing: every -/// `GetTransaction` goes to the hub regardless. -pub struct Diversion { - pub hub: HubTransport, -} -``` - -That statelessness is a defensible choice and is why `GetTransaction` is routed -unconditionally rather than recognised. But it means the conditional, -`TaintedAddrs`-keyed design in `ENDPOINTS.md` cannot be implemented as written, so -this leak is not on a path to closing by itself. The two options that remain are -unconditional ("broad") routing of the address methods — which `ENDPOINTS.md:194-206` -names as the load-bearing open decision — or an honest README. - -## Recommendations - -1. **Correct `README.md:33` — this is the cheap fix and it removes a - contradiction that sits inside a single six-line list.** Replace "though not - the amount or which transaction" with something the code supports, e.g.: - - > **The operator learns *that* a client migrated.** For a migration that is - > shielded-to-shielded they learn no more than that. For an Orchard spend with - > a transparent leg they can learn more: address-level queries are not - > intercepted (see above), so their indexer serves the wallet the txid and - > value of the new transparent output and receives the wallet's transparent - > address set with the request. - -2. **Say which claims are adversary-scoped.** Every bullet under *Protected* and - *Not protected* is adversary-dependent and none says so. This one is different - against a chain observer, against the operator in the fully-managed deployment, - and against the operator under BYOC. - -3. **If the leak is to be closed rather than disclosed, pick "broad" routing.** - Route the transparent-address methods away unconditionally, the same shape as - the `GetTransaction` handling that already ships. It is the only branch of - `ENDPOINTS.md:194-206` compatible with the stateless-shim commitment at - `intercept.rs:56-63`. **Do not implement the conditional/`TaintedAddrs` branch** - as written: `endpoints-conditional-intercept-leaks-by-omission-…` shows a - conditional route is itself a signal to an adversary who sees the request - sequence. - -4. **Make the routing table state the policy.** Give the address-level methods - explicit `Route` variants rather than leaving them in the `PassThrough` - default, so a future method is a compile-time decision rather than a silent - omission. - -## Validation Information - -**VERDICT: CONFIRMED, Medium** (filed Medium; held). Validated 2026-08-18 against -the target at HEAD plus two primary sources the filing did not use: the reference -light-client SDK (`audit-context/zero/librustzcash/zcash_client_backend/src/sync.rs`) -and the `CompactTxStreamer` proto (`audit-context/zero/lightwalletd/walletrpc/service.proto`). - -**Every mechanical claim re-verified.** - -| claim | verified | -|---|---| -| `route_for` distinguishes only `SendTransaction` and `GetTransaction` among Zcash methods | yes — `proxy.rs:744-783`, read end to end | -| no `TaintedAddrs` / address state anywhere in the shim | yes — no occurrence in `shim/src/`; `intercept.rs:56-63` forecloses it by design | -| address methods land in `PassThrough` and are relayed verbatim | yes — `proxy.rs:664-666`, `:787-832`, `:840-870` | -| `README.md:32` and `:33` read as quoted | yes, at those exact lines | -| `ENDPOINTS.md:130`, `:131`, `:132` read as quoted | yes, at those exact lines | -| the six address methods exist on the wire surface | yes — `service.proto:285`, `:291-292`, `:323-324` | - -**Three corrections applied against the filing.** - -1. **The mechanism was upgraded and the filed attack step was wrong in its - central detail.** The filing had the wallet "polling the deshield destination - address", which is only the wallet's own address in the self-unshield case and - is *never* the wallet's address when paying a third party — as written it would - have overstated the reach. The real and stronger mechanism is - `refresh_utxos` (`sync.rs:485-540`), which sends the account's **entire** - transparent receiver set on **every** sync and gets back `txid` + `valueZat` - per UTXO. No differencing is needed and no attacker timing is needed. -2. **The headline was re-based.** "`route_for` implements none of the - `ENDPOINTS.md` INTERCEPT set" is a weak lead: `ENDPOINTS.md:5-6` explicitly - disclaims being implemented ("This is the design for the FULL shim; the current - PoC only classifies + logs `SendTransaction`"), and `README.md:32` concedes the - gap. A disclosed gap is not the finding; the sentence that denies its - consequence is. -3. **The IP term was struck from the primary claim.** The filing asserted the - operator obtains "IP address → on-chain transaction → amount". In the shipped - fully-managed deployment they do not obtain the IP from this route — the enclave - has no NIC, Caddy's peer is `127.0.0.1`, and `proxy::forward` adds no - forwarding header (verified: no `X-Forwarded-For` handling anywhere in - `shim/src/`). The claim that survives, and that is sufficient to falsify - `README.md:33`, is **amount + txid attributed to an identifiable client**, - where the identifier is the transparent address set rather than the IP. Under - BYOC, or after the confirmed `core-linkage-…` step 0, the IP term returns. - -**Ownership — read this before the report allocates severity.** - -`README.md:33` is falsified three independent ways, and the report must present -the *sentence* once with three routes rather than three findings that each -re-argue it: - -| route | owner | precondition | closed by | -|---|---|---|---| -| the `log_verdict` INFO line carries `orchard_vb` and `tx_len` | `log-verdict-logs-migration-value-balance-at-info.md` (confirmed **High**) | `DEBUG=1`, i.e. an open enclave console on the parent host | remove the fields **and** default `DEBUG=0` | -| exact `\|tx\|` off the unpadded wallet leg, then select out of the batch and read `valueBalanceOrchard` | `core-linkage-…-self-timestamping.md` (confirmed **High**) | a wallet-leg observation post ("step 0") | wallet-side padding plus ZIP 318 uniformity | -| **unintercepted address queries** | **this issue** | **none** | broad routing, or an honest README | - -**What this issue uniquely owns, and why it is not a triple count.** It is the -only one of the three that needs **no privileged position at all** — no parent -host, no debug deployment, no traffic capture — and therefore the only one that is -live for an operator who holds nothing but their own indexer, which is exactly the -shipped fully-managed posture. It is also the only one that **survives the complete -remediation of both Highs**: fixing the log fields and padding the wallet leg -leaves it untouched, because it runs on ordinary pass-through traffic. A prior -validator already recorded this allocation inside -`readme-volume-independent-source-ip-claim-overstated.md`, naming the -"unintercepted `GetTaddressBalance` bracketing" as one of only two routes to the -amount in an attested deployment and pointing at this file. - -**Why Medium and not Low**, against the item-7n precedent that deflated three -README ICTM issues: those three were deflated because a confirmed code-side -finding fully owned their harm. **Nothing owns this one.** Likelihood is not -merely high but certain — the disclosure is passive, arrives by default from the -reference SDK, and needs no attacker action — and the impact is the product's own -headline harm minus the IP term. - -**Why not High.** Three deliberate limits. (a) The acute use case is untouched: a -conforming ZIP 318 migration has no transparent leg. (b) The amount and the txid -are already public chain data; what leaks is the linkage to a client. (c) In the -shipped fully-managed deployment the linkage terminates in a pseudonymous address -set, not the IP the product exists to protect — the IP term requires BYOC or the -separately-owned step 0. - -**Not covered here, and correctly filed elsewhere — do not merge:** -- `ENDPOINTS.md`'s residual section declaring these leaks *closed* by a "Zeronym - indexer" that exists nowhere in the tree is owned by - `endpoints-residual-section-declares-three-leaks-closed-by-a-component-that-exists-nowhere.md` - (plausible, Low). The coordinator asked whether a confirmed issue already owns - it: none does, but that file does, and it should be validated on its own terms. -- The conditional-route-is-itself-a-signal analysis is owned by - `endpoints-conditional-intercept-leaks-by-omission-and-this-residual-is-never-stated.md` - (plausible, Info). Recommendation 3 above depends on its conclusion. -- `GetLatestTreeState` / per-wallet anchor service is **not** claimed here. The - filing raised it as "a second, independent instance"; it is a different harm - (anchor correlation), it is the wallet-side requirement `README.md:69` does - state, and it is owned by the `core-linkage-…` High and by - `readme-tells-wallet-developers-there-is-exactly-one-requirement-…`. Struck from - this issue to keep the boundary clean. - -**One filed claim checked and left standing but re-scoped:** the near-miss gap on -`GetTransaction` (exact-string match, no case/slash normalisation) is referenced in -the filing's table as "see the separate near-miss finding". It is not part of this -issue's harm and is owned by `shim-route-for-gettransaction-no-nearmiss-arm.md`. - -DO NOT DEVIATE FROM THIS TEMPLATE (`docs/SECURITY-ISSUE-TEMPLATE.md`). -LEAVE THIS MESSAGE IN PLACE SO THAT ALL AGENTS KNOW NOT TO DEVIATE FROM THIS TEMPLATE (`docs/SECURITY-ISSUE-TEMPLATE.md`). diff --git a/zeronym-22aa9851caf68-high-medium/medium/widening-the-flush-window-cannot-raise-the-delivered-anonymity-set-for-todays-traffic-and-achieved-batch-size-over-counts-it.md b/zeronym-22aa9851caf68-high-medium/medium/widening-the-flush-window-cannot-raise-the-delivered-anonymity-set-for-todays-traffic-and-achieved-batch-size-over-counts-it.md deleted file mode 100644 index 6f9e2ca6..00000000 --- a/zeronym-22aa9851caf68-high-medium/medium/widening-the-flush-window-cannot-raise-the-delivered-anonymity-set-for-todays-traffic-and-achieved-batch-size-over-counts-it.md +++ /dev/null @@ -1,484 +0,0 @@ -# Widening the flush window cannot raise the delivered anonymity set for the traffic that exists today, because the batch and the dominant on-chain selector grow at the same rate; the two remedies the project names are null and unreachable, and `achieved_batch_size` is the metric that would have shown it - -**Severity**: Medium -**Validation Status**: Confirmed -**Location**: `audit-target/zeronym/hub/REVIEW.md:52-77` (design change #3's superseding note, "it doubles the batch for free" / "the cheapest available improvement to the k=1 problem"), `:195` ("Decisions for humans", first bullet: "The single highest-value lever is widening the migration expiry"), `:175` (the stated k=1 residual, quoted only for context), `audit-target/zeronym/README.md:34` ("The lever is adoption, not code"); the metric at `audit-target/zeronym/hub/src/batcher.rs:335-343`, `:361-378`, `:396-403` and `:412-420`; the constant at `hub/src/batcher.rs:39-41` -**Found by agent:** Global (focus area G3, anonymity-set arithmetic) -**In scope of audit?** Yes — `hub/REVIEW.md` and `README.md` are in scope as security claims, and `audit-context/AUDIT-INSTRUCTIONS.md` directs auditors to report "where the code or the public README claims more than the residual allows". - -## Description - -This issue is **not** a re-report of the acknowledged k=1 residual -(`hub/REVIEW.md:175`, `README.md:34`). That residual is stipulated. This issue is -about the two remedies the project names for it, both of which are stated as -quantities the code and the documents can act on, and neither of which delivers -what it is said to deliver: - -1. `hub/REVIEW.md:69-77` — raising `N` from 10 to 20 *"doubles the batch for - free"* and is *"the cheapest available improvement to the k=1 problem"*. -2. `hub/REVIEW.md:195` — *"The single highest-value lever is widening the - migration expiry … Widening it would let N grow past 10 and **make batch size a - real function of adoption instead of a fixed loss**."* - -Both statements are about **batch size `k`**. The quantity that matters is not -`k` but the number of batch members that share the target's on-chain -distinguisher — which for today's non-ZIP-318 Orchard traffic is dominated by -`nExpiryHeight`. - -A wallet built on librustzcash sets `expiry_height = target_height + -DEFAULT_TX_EXPIRY_DELTA` with the delta at 40 (`hub/src/batcher.rs:49-52` cites -this default itself), and a light wallet builds a transaction immediately before -broadcasting it. `nExpiryHeight` is therefore **a published, one-block-resolution -timestamp of when the wallet built the transaction**, committed under ZIP 244 into -the txid and readable by anyone off the chain forever. - -The consequence is arithmetic and does not depend on any attacker capability: - -* a `W`-block flush window admits `k ≈ λW` entries, where `λ` is the per-block - arrival rate of diverted transactions across the whole fleet; -* those entries were built over `W` consecutive block heights, so the batch - contains **about `W` distinct expiry values**; -* the expected number of batch members sharing the target's expiry is therefore - `1 + (k−1)/W = 1 + λ − 1/W ≈ 1 + λ`. - -**`W` cancels.** Against an adversary who reads `nExpiryHeight`, a 1152-block -flush window delivers exactly the same anonymity as a 1-block window. The -batching mechanism is not weakened by an attacker; it is neutralised by an -identity. - -The same cancellation makes the second remedy unreachable rather than null. -Adoption raises `λ`, and `1 + λ` does grow with `λ` — but to deliver an anonymity -set of 8 requires `λ ≈ 7` diverted Orchard-touching transactions **per block**. -Blockchair's mainnet stats on 2026-08-18 report `transactions_24h = 5238` over -`blocks_24h = 1135`, i.e. **4.61 transactions per block of every kind on the whole -Zcash network**. The required rate is **1.5× the entire network's throughput**, and -**9.1×** the 0.77 Orchard-touching tx/block the design is sized against. Even the -absolute ceiling — every Zcash transaction being Orchard-touching and diverted -through the hub — yields a delivered set of **5.6**. - -`README.md:34`'s *"The lever is adoption, not code"* is therefore not merely -wrong in the way the already-filed -`core-linkage-survives-in-the-attested-deployment-…` finding establishes -(selection is untouched by `k`); it is wrong in a way that has a hard numeric -ceiling no adoption level can pass, for as long as wallets stamp a per-block -expiry. - -`REVIEW.md:195` reaches the right ask by the wrong mechanism, and the wrong -mechanism corrupts the ask. The mechanism it states is *"let N grow"*, which the -cancellation above shows is null. The mechanism that actually works is the other -half of the same sentence — *"an EPOCH-CANONICAL change adopted by all wallets in -the epoch, in the ZIP 318 sense"* — which works because ZIP 318 makes the expiry -value **identical for every wallet in a 30-day period**, deleting the timestamp -rather than lengthening it. Those are different asks with different outcomes: a -consortium acting on the literal words *"widening the migration expiry"* could -raise every wallet's delta from 40 to 1000 blocks, satisfy `MIN_WALLET_EXPIRY` -with enormous margin, permit any `N` — and change the delivered anonymity set by -exactly zero, because `tip + 1000` is still a one-block-resolution timestamp of -`tip`. - -Finally, the one number the design uses to check itself measures `k`: -`batcher::flush` returns `achieved`, logs it as `achieved_batch_size`, and warns -only when `achieved <= 1` (`hub/src/batcher.rs:361-378`, `:396-403`, `:412-420`). -Its doc comment calls it *"the honest measure of the privacy the flush actually -delivered"* (`:335-336`). It is not: it counts published entries, not distinct -`(length, anchor, expiry)` classes. At `λ = 0.77` and `W = 20` it would report -`achieved_batch_size = 15` on a batch whose delivered anonymity set is about -**1.7**, and it would report a healthy number every time the window was widened. -The instrument the project would use to decide whether the anonymity claim holds — -and which `REVIEW.md:109-113` (design change #9) makes the launch gate — cannot -observe the failure. - -## Attack Scenario and Steps - -No attacker capability is required beyond the one already in the threat model, and -the point of this issue is that the *defence* does not scale, not that a new -attack exists. Concretely, for the operator of the indexer the shim fronts: - -1. Client `C` sends one `SendTransaction` that the shim does not forward. The - operator sees the absence (a conceded residual, `hub/REVIEW.md:181`) and - timestamps it. -2. At the next height `≡ 0 (mod 20)` the hub publishes a batch of `k` - transactions simultaneously. The batch is publicly enumerable (conceded, - `hub/REVIEW.md:177`). -3. The operator reads `nExpiryHeight` off each member. For today's wallets that is - `build_height + 40`, so it names the block in which each member was built. -4. The operator keeps only the members whose implied build height matches the - block in which `C`'s request arrived. **On average `1 + λ` members survive**, - and at any adoption level the network can physically sustain, that is under 6. -5. Steps 3–4 are unaffected if the project ships `FLUSH_INTERVAL_BLOCKS = 144` or - `1152`: the surviving count is `λ`, not `λW`. - -**Attack Requirements and Assumptions:** -- The operator's capabilities are exactly those already stipulated: they see which - client's `SendTransaction` was not forwarded, and they can read the public chain. -- The arithmetic assumes wallets set a per-transaction expiry as a fixed delta - from the chain tip at build time and broadcast promptly. That is the shipped - behaviour of librustzcash-derived wallets, and `hub/src/batcher.rs:49-52` and - `hub/REVIEW.md:52-57` both rest on the same assumption when they choose 40. -- **The estimate is sensitive to one unmeasured quantity.** `λ` — the per-block - arrival rate of Orchard-touching transactions actually reaching hubs — has not - been re-measured in this audit; `audit-context/EXTERNAL-CONTEXT.md` §7 flags the - 0.77/block figure as a 144-block sample taken four days after activation and - never revisited. Every `λ`-dependent number here is **arithmetic, not - observation, and is labelled as such**. The `W`-cancellation itself does not - depend on `λ`. -- Offline signing or any delay between build and broadcast **widens** the spread - of expiry values in a batch, so `1 + λ` is an upper bound on the delivered set, - not a lower one. -- `anchorOrchard` for a latest-anchor wallet has the same one-block resolution and - is derived from the same tip, so it is nearly perfectly correlated with expiry. - It does not multiply the reduction; it is a second, independent confirmation of - the same estimate. Transaction length (a function of action count) *is* - independent and cuts further. - -## Impact on Users - -- Users of a hub that widens its window get a longer wait and no more privacy, for - the traffic that exists on mainnet today. -- The project, its reviewers and any wallet consortium acting on `REVIEW.md:195` - can spend the one expensive, cross-organisational ask they have — a coordinated - wallet change — on the version of it that does nothing (a longer delta) instead - of the version that works (a shared constant). -- The telemetry that is supposed to say "the anonymity claim does not hold at this - adoption level" will say the opposite as soon as either remedy is applied. -- **The adoption lever also has a deadline, which nothing in the documents - acknowledges.** The public tracker at , fetched - 2026-08-18 at height 3,451,561, reports `orchard_at_activation` = 3,660,833 ZEC, - `crossed_zec` = 2,913,084 and `pct_crossed` = 79.57 — i.e. **79.6 % of the - migrating value crossed in the 23,418 blocks (20.3 days) since NU6.3**, most of - it before zeronym's first attested enclave (2026-08-01) and all of it before the - Nym transport went live (2026-08-14). Two honest caveats, both of which must - travel with this number: value is not transaction count and not user count, so a - small number of large holders could account for most of it; and the Sprout - precedent the same tracker cites ("~25,409 ZEC still sits in Sprout, never having - crossed" after eight years) says the *user* tail is long even when the value tail - is short. Still, the population whose members hide each other is measurably - draining, and "the lever is adoption" is a remedy that has to arrive before the - thing it operates on is gone. - -## Technical Details / Code Analysis - -**The claim, verbatim** — `hub/REVIEW.md:69-77`: - -> This is not merely more comfortable, it is the cheapest available improvement to -> the k=1 problem below, and it costs no wallet any change. A 20-block window -> accumulates twice what a 10-block window does: roughly 15 Orchard-touching -> transactions network-wide per window instead of 8, at the measured 0.77 per -> block. It does not solve k=1 at low adoption, because the limiting factor is the -> participating fraction rather than the window, but it doubles the batch for -> free. - -and `hub/REVIEW.md:195`: - -> The single highest-value lever is widening the migration expiry, and it is not a -> hub decision. 20 blocks is Brave's wallet default, not a consensus constant -> (librustzcash uses 40, Zingo 100). Widening it would let N grow past 10 and make -> batch size a real function of adoption instead of a fixed loss. - -**The cancellation.** With `FLUSH_INTERVAL_BLOCKS = W` (`hub/src/batcher.rs:39-41`): - -| quantity | scales as | -|---|---| -| batch size `k` | `λW` | -| distinct `nExpiryHeight` values in the batch | `W` | -| members sharing the target's expiry | `λW / W = λ` | - -Worked out, with `λ` a free parameter because it is unmeasured: - -| `λ` (diverted tx/block, fleet-wide) | `W = 20` | `W = 24` | `W = 144` | `W = 1152` | -|---|---|---|---|---| -| 0.05 (≈ a modal batch of 1 at `W=20`) | 1.00 | 1.01 | 1.04 | 1.05 | -| 0.25 | 1.20 | 1.21 | 1.24 | 1.25 | -| 0.77 (the design's assumed rate, at 100 % participation) | 1.72 | 1.73 | 1.76 | 1.77 | -| 4.61 (**every** Zcash transaction, measured 2026-08-18) | 5.56 | 5.57 | 5.60 | 5.61 | - -Reading down a column shows adoption is a lever; reading across a row shows the -window is not one. The rightmost column is a 24-hour flush interval. - -**The `nExpiryHeight` identity.** `hub/src/batcher.rs:49-56` is where the project -records the wallet behaviour this rests on: - -```rust -/// The tightest wallet expiry the design commits to supporting, in blocks. -/// -/// This is librustzcash's default (40), NOT Brave's 20. Brave is out of scope -/// for v1 and the ask to them is to raise their default to 40. If any wallet -/// with an expiry below 40 comes into scope, `FLUSH_INTERVAL_BLOCKS` must come -/// back down and the batch shrinks with it. -pub const MIN_WALLET_EXPIRY: u32 = 40; -``` - -A wallet honouring that ask still stamps `build_height + 40`. The number 40 is -irrelevant to the leak; the *variability across wallets within one window* is the -leak, and it is exactly `W`. - -**The metric.** `hub/src/batcher.rs:335-343`: - -```rust -/// Returns the achieved batch size, which is the honest measure of the privacy -/// the flush actually delivered. -pub async fn flush(queue: &Arc, chain: &Arc) -> usize { - let batch = queue.drain_shuffled(); - let size = batch.len(); -``` - -and `:361-367`, `:412-420`: - -```rust - let mut achieved = 0usize; - ... - Some(Publish::Accepted { .. }) | Some(Publish::AlreadyKnown) => achieved += 1, - ... - if achieved <= 1 { - // Honest telemetry, not an error. At batch size 1 the anonymity set is - // the transaction itself and the shuffle, the simultaneous publish, Nym - // and the enclave are all irrelevant to it. - tracing::warn!( - achieved_batch_size = achieved, - "batch provides no batching anonymity at this size" - ); - } -``` - -`achieved` is a count of publications. The comment's own reasoning — "at batch -size 1 the anonymity set is the transaction itself" — is the right test applied to -the wrong variable: a batch of 15 transactions with 15 distinct expiry heights is -15 anonymity sets of size 1, and this branch stays silent for all of them. - -**Why the fix direction is not "widen" but "make identical".** `SPEC-NOTES.md` §3, -quoting ZIP 318 directly: - -> The bucketed rule makes the committed value **identical for every migration -> transaction — from any wallet** — whose scheduled broadcast falls within the same -> 30-day period, so an expiry height reveals only the coarse period in which its -> broadcast was scheduled. - -Under that rule the batch contains **one** expiry value regardless of `W`, the -denominator collapses to 1, and the window becomes a real lever for the first time -(see the companion issue on what still partitions a conforming batch). The -distinction matters because `REVIEW.md:195` states the ask as a magnitude -("widening") when it is a *coordination* property, and only the parenthetical -"in the ZIP 318 sense" carries the part that works. - -**Relationship to other filed issues — do not double-count.** -- `core-linkage-survives-in-the-attested-deployment-because-the-wallet-leg-is-unpadded-and-the-published-transaction-is-self-timestamping.md` (High) establishes *that* selection on `(length, anchor, expiry)` defeats batching. This issue supplies the quantity and shows that both stated remedies fail for different reasons. The user harm is the same harm; this is not a second instance of it. -- `hub-flush-interval-pinned-by-a-ceiling-that-does-not-bind-zip318.md` (Low) recommends widening the window. That recommendation is **correct only after wallets adopt ZIP 318's bucketed expiry**, and is null before it. An addendum has been appended to that issue pointing here so the report does not ship the ordering backwards. -- `review9-launch-gate-on-measured-batch-size-has-no-egress-path-so-the-accepted-k1-residual-can-never-be-measured.md` establishes that the launch-gate metric cannot leave the enclave. This issue establishes that, if it could, it would measure the wrong thing. - -## Recommendations - -**Ordering rule, before any of the numbered items: do NOT widen `FLUSH_INTERVAL_BLOCKS` first.** Other findings in this audit recommend widening it. Applied before wallets emit a network-wide canonical expiry, widening changes the delivered anonymity set by zero — and `hub/src/batcher.rs`'s startup assertion and `queue.rs`'s admission rule both refuse the wider window anyway while wallets stamp a 40-block delta (see Validation Information §3). Sequence the wallet-side change first. - -1. Correct `hub/REVIEW.md:69-77`: raising `N` multiplies `k` but not the delivered - anonymity set, because the number of distinct `nExpiryHeight` values in a batch - is proportional to `N`. Say that the window is a lever only once expiry is - epoch-canonical. -2. Restate `hub/REVIEW.md:195`'s ask as what it needs to be: **not "widen the - migration expiry" but "adopt one shared expiry value per epoch"**, i.e. ZIP - 318's `floor(h / EXPIRY_MODULUS) * EXPIRY_MODULUS + 2 * EXPIRY_MODULUS`. A - longer per-wallet delta satisfies the sentence as written and buys nothing. - `zcash_protocol::zip318::expiry_height()` is already compiled into both attested - binaries (PROGRESS.md open item 6l), so both sides can compute the same value - from a crate they already link. -3. Change the metric at `hub/src/batcher.rs:361-421` from a count of published - entries to a count of **distinct `(tx_len, anchorOrchard, nExpiryHeight)` - classes in the batch, and the size of the smallest class**. The hub already - holds every batch member's bytes and already re-parses them as telemetry - (`REVIEW.md` design change #5), so this needs no new information. Keep - `achieved_batch_size` for publication accounting; do not keep calling it the - measure of delivered privacy. -4. Widen the `achieved <= 1` warning to fire on `min_class_size <= 1`, which is the - condition the surrounding comment actually describes. -5. State in `README.md:34` the ceiling as well as the direction: at the measured - whole-network transaction rate, no adoption level delivers an anonymity set - above ~6 while wallets stamp a per-block expiry. -6. Re-measure `λ` before the report ships, and re-derive the table above against - the measurement rather than against 0.77. - -## Validation Information - -**VERDICT: CONFIRMED. Severity: Medium (unchanged).** Validated 2026-08-18. The -central result — **the flush window `W` cancels out of the delivered anonymity -set** — was re-derived from scratch rather than inherited, every cited line in the -target was re-read, the wallet behaviour the derivation rests on was checked -against librustzcash's actual source rather than against the project's summary of -it, and the one external measurement was re-taken. **The derivation holds, the -arithmetic in every table cell is correct, and two independent code facts make the -result stronger than filed.** Four corrections and additions are recorded below; -none of them touches the conclusion. - -**READ THIS FIRST IF YOU ARE FIXING SOMETHING ELSE IN THIS AUDIT.** Several other -findings recommend *widening the flush window* (`FLUSH_INTERVAL_BLOCKS`) as their -remedy — most directly -`hub-flush-interval-pinned-by-a-ceiling-that-does-not-bind-zip318.md`, which -carries a marked addendum pointing here. **Do not do that first.** Against the -adversary this system is built to stop, and for the wallet population that exists -on mainnet today, widening the window multiplies the batch and multiplies the -number of on-chain selector values by the same factor, so the delivered anonymity -set does not move. Widening buys latency and nothing else until wallets emit a -network-wide canonical expiry. The correct order is: (1) wallet-side ZIP 318 -bucketed expiry (and request padding), then (2) widen the window. Reversed, the -engineering effort in (2) is spent for nothing and — because of the metric defect -below — the telemetry will report that it worked. - -### 1. The `W`-cancellation: re-derived independently, and correct - -Two derivations, both giving the same answer. - -*Conditional-on-`k` (the issue's own form).* A batch published at a flush height -holds the entries admitted during the preceding `W` blocks. Under the wallet model -below, each entry's `nExpiryHeight` is a deterministic function of the block in -which it was built, and a light wallet builds immediately before broadcasting, so -a batch drawn from `W` consecutive heights carries about `W` distinct expiry -values, roughly uniformly. Conditioned on a batch of `k`, the expected number of -members sharing the target's value is `1 + (k-1)/W`. Substituting `k = λW` gives -`1 + λ - 1/W`. - -*Unconditional (Poisson).* If diverted arrivals are Poisson at `λ` per block, the -number of *other* entries built in the target's own block is `Poisson(λ)` and is -independent of `W` outright. Expected delivered set `= 1 + λ`, exactly, for every -`W`. - -Both forms agree, and neither has `W` in the answer. Spot-checked table cells -against `1 + λ - 1/W`: (0.05, W=20) = 1.00; (0.05, W=1152) = 1.049; (0.77, W=20) -= 1.72; (4.61, W=20) = 5.56; (4.61, W=1152) = 5.609. **Every published cell is -right to two decimals.** `7/0.77 = 9.09` ("9.1x") and `7/4.61 = 1.52` ("1.5x") -also check out, as does `0.77 x 20 = 15.4` behind "`achieved_batch_size = 15` on a -batch whose delivered set is about 1.7". - -### 2. The wallet behaviour it rests on, checked at the source rather than quoted - -The derivation is only as good as the claim that `nExpiryHeight` is a -one-block-resolution timestamp of build time. Verified directly against the -vendored upstream in `audit-context/zero/librustzcash`, not against the project's -paraphrase: - -- `zcash_primitives/src/transaction/builder.rs:54` — `pub const DEFAULT_TX_EXPIRY_DELTA: u32 = 40;` -- `zcash_primitives/src/transaction/builder.rs:636-640` — for every non-coinbase - transaction, `expiry_height = target_height + DEFAULT_TX_EXPIRY_DELTA`, where - `target_height` is the builder's view of the chain tip. - -So the committed value is `tip_at_build + 41`-ish, it is serialized in the clear -on chain, and it is committed under ZIP 244 into the txid. The delta being 40 -rather than 20 or 1000 is irrelevant to the leak, exactly as the issue says: what -leaks is the *variance across wallets inside one window*, and that is `W`. - -The operator's independent measurement of the target's value is better than the -issue claims, not worse: `GetLatestBlock` and the whole sync surface are -`Route::PassThrough`, so the operator served the victim the very tip the victim -then stamped. The `±1 block` in the issue is conservative. - -### 3. TWO CODE FACTS THAT STRENGTHEN THE FINDING AND ARE NOT IN THE FILING - -**(a) The window cannot be widened today at all — the code refuses to boot.** -`BatchParams::validate` (`hub/src/batcher.rs:96-115`), called unconditionally at -`hub/src/main.rs:45`, enforces -`flush_interval + mining_margin + delivery_lag <= min_wallet_expiry`. With the -shipped constants that is `W + 4 + 6 <= 40`, i.e. **`W <= 30`**. The `W = 144` and -`W = 1152` columns of the issue's table are therefore not merely useless, they are -unshippable while wallets stamp a 40-block delta. - -That is exactly why `hub/REVIEW.md:195` asks for a *wider wallet expiry* — it is -the precondition that would let `N` grow. The finding's force is therefore -sharper than filed: **the project has correctly identified the gate, and the thing -behind the gate is worth zero.** A consortium that raised every wallet's delta from -40 to 1000 would unblock `validate()`, permit `W = 990`, produce batches ~50x -larger, and deliver `1 + λ` — the same number as today. - -**(b) Admission control refuses the wide-window batch independently.** -`survives_next_flush` (`hub/src/queue.rs:380-392`) admits an entry only if -`expiry >= next_flush_height(tip, W) + mining_margin`. At `W = 144` an entry -carrying `tip + 40` fails that test for all but the last few blocks of each -window, so even with the startup assertion relaxed the queue would refuse most of -the traffic rather than accumulate it. Two independent mechanisms, both pointing -at the same wallet-side precondition. - -### 4. The λ measurement: re-taken, correctly characterised, and slightly harsher than filed - -`https://api.blockchair.com/zcash/stats`, re-fetched during this validation: -`transactions_24h = 4,846` over `blocks_24h = 1,146` at height 3,452,485 → -**4.23 transactions per block of every kind, network-wide**. The filed figure -(5,238 / 1,135 = 4.61) is a different sample of the same rolling 24-hour counter -on the same day; both are the same quantity and the drift is ordinary. - -The characterisation in the issue is **correct and correctly bounded**: this is -every Zcash transaction of every kind, used *only* as a hard ceiling on `λ`, and -the issue already labels every λ-dependent number as arithmetic rather than -observation. Re-measured, the 100 %-diversion ceiling on the delivered set is -`1 + 4.23 = 5.23` rather than 5.6 — the finding gets marginally stronger, and the -"~6" in recommendation 5 remains the right thing to write. The -**Orchard-touching** rate was not re-measured here either (no reachable endpoint -exposes per-block pool composition); `EXTERNAL-CONTEXT.md` §7 stays open, and that -is stated in the issue rather than papered over. - -One honest bound on the ceiling argument, which the report should carry: 4.23 -tx/block is *today's network throughput*, not a consensus limit. The ceiling -statement is therefore "no adoption level reachable at the network's current -throughput", not "no adoption level ever". That qualification is already in -recommendation 5's wording ("at the measured whole-network transaction rate") and -should not be dropped. - -### 5. The metric defect is real, and the launch-gate coupling is verified - -Re-read at the source. `hub/src/batcher.rs:335-336` — *"Returns the achieved batch -size, which is the honest measure of the privacy the flush actually delivered."* -`:361-378` — `achieved` is incremented once per `Publish::Accepted | AlreadyKnown` -verdict, i.e. it is a **count of publications**. `:396-403` logs it as -`achieved_batch_size`; `:412-420` warns only on `achieved <= 1`. And -`hub/REVIEW.md:109` is the launch gate in as many words: *"Compute -`achieved_batch_size` per flush, export the distribution to the hub operator, and -gate LAUNCH on a measured distribution rather than gating each batch."* - -The issue's characterisation is exact: a batch of 15 transactions with 15 distinct -expiry heights is fifteen anonymity sets of size 1, and this branch stays silent -for every one of them. The comment beside the warning (*"at batch size 1 the -anonymity set is the transaction itself"*) states the correct test and applies it -to the wrong variable. Recommendation 3 (count distinct -`(tx_len, anchorOrchard, nExpiryHeight)` classes and the smallest class) needs no -new information — the hub already holds every member's bytes and already re-parses -them as telemetry per `REVIEW.md` #5. - -### 6. Not double-counted, and not a re-report of an accepted residual - -- The *privacy failure itself* (selection on `(length, anchor, expiry)` defeats - batching) belongs to the confirmed High - `core-linkage-survives-in-the-attested-deployment-...`. **This issue does not - re-report it** and says so. -- `hub/REVIEW.md:175` and `README.md:34` concede `k = 1` at today's adoption. - `AUDIT-INSTRUCTIONS.md` forbids re-reporting a stated residual but explicitly - invites reporting "where the code or the public README claims more than the - residual allows". The claims tested here are the *remedies* offered alongside the - residual — `REVIEW.md:69-77` ("the cheapest available improvement to the k=1 - problem", "doubles the batch for free"), `REVIEW.md:195` ("make batch size a real - function of adoption instead of a fixed loss") and `README.md:34` ("The lever is - adoption, not code") — none of which is a residual, and all three of which are - wrong in the specific way documented. That is squarely inside the instruction. - -### 7. Why Medium and not higher or lower - -**Not High.** No new leak and no new attacker capability: the harm this issue -adds, on top of the already-confirmed High, is misdirected remediation plus an -instrument that will read green when the property fails. - -**Not Low.** It is not theoretical. `REVIEW.md` #9 makes `achieved_batch_size` the -launch decision, so the defective metric is on the path to a real -ship/don't-ship judgement about a real privacy property; `README.md:34` is a -public, user-facing statement about what would fix the one limitation the project -concedes, and it names a lever that cannot reach its target; and the widening -remedy is recommended by at least one other finding in this same audit, so a -developer working the issue list without this one in front of them would spend the -effort and get nothing. Each of those lands on real users through a decision -someone is going to make. - -**Two things a report author must not overstate.** (i) The `1 + λ` figure is -arithmetic on an unmeasured `λ`; the `W`-cancellation is what does not depend on -it, and that is the load-bearing half. (ii) The selector requires an adversary who -knows *when the target's diverted request arrived* — that is the shim's operator -(a conceded, stipulated capability), not an arbitrary chain observer. For a pure -chain observer with no target to align against, a bigger `k` is still a bigger -haystack. The finding is about the adversary the product exists to stop, and -should be written that way. - - -DO NOT DEVIATE FROM THIS TEMPLATE (`docs/SECURITY-ISSUE-TEMPLATE.md`). -LEAVE THIS MESSAGE IN PLACE SO THAT ALL AGENTS KNOW NOT TO DEVIATE FROM THIS TEMPLATE (`docs/SECURITY-ISSUE-TEMPLATE.md`).