Skip to content

feat: Run Roach as a remote proxy on GCP with GCS recordings - #1

Merged
dcramer merged 11 commits into
mainfrom
poc/remote-service
Oct 10, 2026
Merged

dcramer merged 11 commits into
mainfrom
poc/remote-service

Conversation

@sentry-junior

@sentry-junior sentry-junior Bot commented Oct 9, 2026 •

Copy link
Copy Markdown
Contributor

Run Roach as a hosted man-in-the-middle proxy. Junior and other projects point HTTPS_PROXY at it and trust its CA certificate. They install nothing else.

Cloudflare Workers can't accept an incoming CONNECT, so the service runs on one GCP VM behind an SSL proxy load balancer. Recordings go to GCS. This replaces the R2 Worker store from earlier commits on this branch.

How it works

  • A CI job calls POST /__roach/runs and gets back a run with its own id and token. Each proxied request names its run in Proxy-Authorization, so many CI jobs can share the service at once.
  • Reads are public: anyone can start a replay run, so fork PRs can replay. Every other mode writes, so it needs the tenant's token. The service checks this, not the CI workflow. The config stores only the SHA-256 of each token.
  • Each tenant's recordings sit under <tenant>/ in the bucket. A GCS lifecycle rule deletes each one 30 days after it was written.
  • Each read and write sends a roach.recording count to Sentry, with tenant, rule, key, run and result. Group by key and count run to find recordings that only one run uses.
  • Runs can only use the built-in value patterns plus the ones the operator lists, because one slow regex would block every run.
  • A request body over 64 MiB gets HTTP 413. Upstream responses have no limit. The control API answers a body that isn't JSON with 400.

Production setup

deploy/gcp/ is Terraform for everything: the bucket and its expiry rule, the CA, one token per tenant, the service config in Secret Manager, a Container-Optimized OS VM, and the load balancer with a Google-managed certificate. roach.tfvars.example already lists Junior's origins and value patterns. The README "Deploy" section has the steps. The new Image workflow builds the image on PRs and pushes ghcr.io/getsentry/roach on main. It also runs terraform validate.

Tests

  • tests/service.test.ts: tenants, public replay, concurrent runs, redaction, the body limit on requests but not on responses, and 400/413 from the control API.
  • tests/deployed.test.ts: runs the service the way production does. It starts from the CLI with a config file and a fixed CA, behind a TLS front end like the load balancer. A fake GCS server, metadata server and Sentry server stand in for the real ones. Child processes set only the run's proxy variables. It checks record with a token, replay for a fork, and the Sentry counts.
  • I built the Docker image and ran it by hand: record, public replay, and a refused write without the token.

Known limits

  • Runs live in memory, so a restart or a new image ends the CI runs in progress.
  • Expiry counts from the write, not the last use. In strict replay mode, a request for an expired recording fails.
  • Anyone who can form a request can read its recorded response. Don't record responses that must stay private.
  • The Terraform state holds the CA key and the tenant tokens, so it needs a private backend.
  • Nothing is deployed, and I haven't run terraform apply against a real project.

via David Cramer.

--

View Junior Session [Sentry]

@dcramer
dcramer marked this pull request as ready for review October 9, 2026 18:31
Comment thread src/service.ts Outdated
Comment thread src/service.ts
@sentry-junior
sentry-junior Bot force-pushed the poc/remote-service branch 2 times, most recently from 5811a2c to c694a42 Compare October 9, 2026 18:46
Replace the shared Node proxy service with a remote store. Each CI run
keeps its own local proxy, and only the recordings move to a Worker. This
needs no raw TCP ingress and no shared HTTPS interception.

The recorder now uses a RecordingStore: the file store (as before) or the
remote store, which calls the Worker. The Worker keeps each recording in R2
and one D1 row with its last use and the hash of each request part.

Miss diagnosis no longer reads every recording on the first miss. The
Worker finds the closest recording with one indexed D1 query on the part
hashes, and returns the parts that differ.

Recordings expire by last use, not by upload: a replay refreshes the row
at most once a day, and a daily cron deletes rows and objects that nobody
used for RECORDING_TTL_DAYS. Tenants have hashed tokens and see only their
own recordings.

tests/worker.test.ts runs the real wrangler.jsonc with local R2 and D1
through createTestHarness: record and replay through a proxy, misses,
tenant isolation, size limits, and expiry.

Co-Authored-By: David Cramer <david@sentry.io>
@sentry-junior
sentry-junior Bot force-pushed the poc/remote-service branch from c694a42 to 693377f Compare October 9, 2026 19:03
Comment thread src/recordings.ts Outdated
@sentry-junior sentry-junior Bot changed the title feat: Add a shared Roach service for many projects (POC) feat: Keep recordings in a Cloudflare Worker with R2 and D1 Oct 9, 2026
A failed session lists the keys that it recorded before. keysOf read
every entry of the recordings directory as a rule directory, so a file
such as .DS_Store failed with ENOTDIR and ended the session with HTTP 409.
It now reads only directories.

@cursor cursor Bot left a comment •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

Comment thread src/recordings.ts
The Worker now keeps recordings only in R2. A 30-day R2 lifecycle rule
expires them, so the D1 index, the closest route, and the cron trigger
are gone. The closest recording was only a hint to debug a miss; the
file store still gives it.

Each read and write sends the Sentry metric roach.recording with the
tenant, rule, key, run, and result. The proxy sends its run in the
X-Roach-Run header from the new store.run config. Grouping by key and
run shows recordings that only one run uses.

Co-Authored-By: David Cramer <david@sentry.io>
Comment thread worker/index.ts Outdated
sentry-junior Bot and others added 2 commits October 9, 2026 19:20
Used files list keys with '/', but readdir returns paths with the
platform separator. On Windows every recording looked unused and prune
deleted all of them.
Move the part comparison back into request-key.ts, since the Worker no
longer imports it. RequestParts moves to store.ts so the Worker can
still typecheck store.ts without Node types. Drop the atomic file
writes, which no reader needed. Fix comments and docs that still
described the D1 index and expiry by last use.

Co-Authored-By: David Cramer <david@sentry.io>
@sentry-junior sentry-junior Bot changed the title feat: Keep recordings in a Cloudflare Worker with R2 and D1 feat: Keep recordings in a Cloudflare Worker with R2 Oct 9, 2026
Roach is now a hosted man-in-the-middle proxy. Clients point HTTPS_PROXY
at it and trust its CA; they install nothing else. Cloudflare Workers
cannot accept CONNECT, so the Worker store is gone.

The service holds many runs at once. Each proxied request names its run
in Proxy-Authorization. Replay runs are public, so fork CI can replay;
any mode that writes needs the tenant token. Recordings live in a GCS
bucket per tenant prefix and expire 30 days after they are written. Each
read and write sends a roach.recording count to Sentry with the tenant,
key, run, and result.

deploy/gcp is the production setup: a GCS bucket with the TTL rule, the
CA and tenant tokens, the service config in Secret Manager, one COS VM,
and an SSL proxy load balancer with a managed certificate. The Image
workflow builds the container and pushes it to ghcr on main, and runs
terraform validate.

A request body over 64 MiB now gets a 413 instead of a reset socket.

Co-Authored-By: David Cramer <david@sentry.io>
Comment thread src/server.ts
The 64 MiB limit also applied to recorded upstream responses, so a
large response failed with a 413 that blamed the client. Only client
request bodies now have the limit.
Comment thread src/server.ts
The control API sent 500 for every error, also for a body over the limit
or a body that is not JSON. It now sends the status of the error, with
its message. Any other error is still a 500.
@sentry-junior sentry-junior Bot changed the title feat: Keep recordings in a Cloudflare Worker with R2 feat: Run Roach as a remote proxy on GCP with GCS recordings Oct 9, 2026

@cursor cursor Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes and found 1 potential issue.

Fix All in Cursor

❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.

Reviewed by Cursor Bugbot for commit 5e48d71. Configure here.

Comment thread src/service.ts
Anyone can start a replay run without a tenant token. A request that no
rule matched still went live, so such a run was an open proxy to every
allowed origin. A run without the token now refuses those requests with
403, so it never sends anything live.
Comment thread src/service.ts
Anyone can start a replay run without a token, and each run stays in
memory for up to 6 hours. So repeated calls could fill the memory of the
shared process. A tenant can now have 50 such runs open; another one gets
HTTP 429 until one ends. Runs with the tenant token have no cap.
Comment thread src/service.ts
A body such as `null` parsed as JSON, and the caller then read a field
of it. That threw a TypeError, so the control API answered 500. A body
that is not a JSON object now gets HTTP 400.
@dcramer
dcramer merged commit c7dc81d into main Oct 10, 2026
15 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant