Skip to content

Design: make automatic stale-job recovery honor worker queue scope #223

Description

@obckelbley

Context

This follows the proposal for queue-scoped worker instances: #222.

Sidequest's automatic stale-job recovery currently calls:

backend.staleJobs(maxStaleMs, maxClaimedMs)

without a queue constraint, then transitions every returned stale job:

https://github.com/sidequestjs/sidequest/blob/master/packages/engine/src/routines/release-stale-jobs.ts

That behavior is consistent with the current model in which every engine processes every queue. If engines can process disjoint queue sets, however, automatic maintenance also needs defined ownership semantics.

Problem

Suppose two worker groups share a backend:

  • worker A processes queue-a;
  • worker B processes queue-b;
  • the groups use different stale-job thresholds.

Worker A's maintenance routine can discover and transition stale jobs from queue-b. One worker group's maintenance schedule and timeout configuration can therefore affect work owned by another group.

Even when each transition is concurrency-safe:

  • for claimed jobs, and running jobs without a job-specific timeout, a shorter threshold on one engine may recycle another engine's jobs;
  • an engine performs state transitions for queues it cannot process;
  • each engine may query stale rows belonging to every queue;
  • execution ownership and maintenance ownership become inconsistent.

Existing workaround

A custom backend can override staleJobs(), call the underlying implementation, and filter the returned jobs by queue.

That prevents cross-queue transitions, but it still retrieves stale rows from unrelated queues. Performing the filtering in the database query instead requires backend-specific query implementations.

A first-class queue filter in the backend contract would allow each built-in backend to apply the constraint efficiently.

Proposed semantic rule

For automatic engine maintenance:

An engine should recover stale jobs only from queues that it is eligible to process.

An engine without an explicit queue scope would retain the current backend-wide behavior.

This need not restrict administrative operations. A manual operator query could continue listing stale jobs across every queue, optionally accepting its own queue filter.

Possible API direction

The exact API should follow the queue-scoping design. One possible backward-compatible backend extension would be:

backend.staleJobs(maxStaleMs, maxClaimedMs, {
  queues: ["queue-a"],
});

Applying the filter in each backend query would avoid retrieving and discarding unrelated rows. Filtering only after the query would produce the correct state-transition behavior but retain unnecessary cross-queue reads.

Acceptance cases

Given stale jobs in queue-a and queue-b:

  1. A sweep from an engine scoped to queue-a transitions only the queue-a jobs.
  2. Jobs in queue-b remain unchanged.
  3. An unscoped engine continues recovering stale jobs from both queues.
  4. Queue filtering behaves consistently across every supported backend.
  5. Existing callers that provide no queue scope remain backward-compatible.

Should automatic stale-job recovery follow an engine's queue-processing scope, or is recovery intended to remain backend-global?

I would be happy to implement the agreed behavior after the queue-scoping semantics are settled.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions