Skip to content

fix(service): hash-tag job cancellation keys for Redis Cluster - #7

Merged
TomasPalsson merged 1 commit into
fix/nsjail-no-cgroup-clonefrom
fix/job-cancellation-cluster-slots
Sep 30, 2026
Merged

TomasPalsson merged 1 commit into
fix/nsjail-no-cgroup-clonefrom
fix/job-cancellation-cluster-slots

Conversation

@TomasPalsson

Copy link
Copy Markdown

What breaks

On the sandbox (Redis in cluster mode), every programmatic run (run_tools_with_bash, and any PTC job) fails. The worker errors right away with:

CROSSSLOT Keys in request don't hash to the same slot

The API never learns that the job failed, so it waits the full JOB_TIMEOUT (about 5 min). Then it returns a bare Code execution failed. Plain /exec runs like echo hi still work.

Cause

The claim/commit/cancel Lua scripts in job-cancellation.ts touch three keys for one job: <key>, <key>:result and <key>:execution. The key had no hash tag, so on a cluster those keys hash to different slots. reconcile() also sent one MGET across many jobs, which crosses slots as well.

Fix

-  return `${JOB_CANCELLATION_PREFIX}:${encodeURIComponent(queueName)}:${encodeURIComponent(jobId)}`;
+  return `${JOB_CANCELLATION_PREFIX}:${hashTag(`${encodeURIComponent(queueName)}:${encodeURIComponent(jobId)}`)}`;

reconcile() now does one GET per job instead of a cross-slot MGET.

Verification

  • A new unit test asserts that a job's keys share one slot. It fails before the fix and passes after; bun test src/job-cancellation.test.ts passes 24/24.
  • Run against a real single-node cluster-mode Redis (--cluster-enabled yes), the old code fails with the exact production CROSSSLOT error. The new code completes claim, commit and read.
  • tsc --noEmit reports the same 7 existing errors before and after; none are in the changed files.
  • Not run: job-cancellation-commit.test.ts, because it needs a local redis-server.

Rollout

Deploy the API and the worker together, because the key names change.

Programmatic jobs failed with CROSSSLOT on cluster-mode Redis: the
claim/commit/cancel scripts touch a job's marker, :result and :execution
keys, which hashed to different slots. The worker failed the job but the
API waited out JOB_TIMEOUT (~5 min) before returning a bare error.

Wrap the per-job part of the key in a hash tag, and reconcile active jobs
with one GET per job instead of a cross-slot MGET.
@TomasPalsson
TomasPalsson merged commit ad213ba into fix/nsjail-no-cgroup-clone Sep 30, 2026
10 checks passed
@TomasPalsson
TomasPalsson deleted the fix/job-cancellation-cluster-slots branch September 30, 2026 11:15
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant