Skip to content

[BUG] DistWorker route GC can permanently skip part of the route table when stepHint > 1 #300

Description

@ImDanXie

Environment

  • BiFroMQ 4.0.0 / main @ eef5e3af
  • Module: bifromq-dist/bifromq-dist-worker
  • Files: DistWorkerCoProc.java (gc handler), DistWorkerCleaner.java (cursor bookkeeping)

Summary

When the cleaner sweeps with stepHint > 1, the co-proc inspects every step-th key beneath a per-range cursor. The cleaner resumes each session at the key where the previous one stopped, and when a session wraps it stops at its own start key (the wrap-termination compares against the session start). The cursor is therefore pinned: every session starts at the same key and, with a fixed step, samples the same modulo class of the route table forever. Routes in the remaining classes are never inspected, so dead routes in them are never collected.

Reproduction

Unit-level (deterministic): drive successive GCRequests with stepHint = 2 against an in-memory key set, feeding each reply's nextStartKey back into the next request. The set of inspected keys stops growing after the first session: 6 of 12 keys are inspected, the other 6 never.

Observed on a live broker under subscription churn: the same ~50% of routes remained un-inspected across 400/900 consecutive GC rounds (the skipped set is stable, not random).

Impact

Dead routes in the skipped classes are never collected. The route table (and the memory/cache working set) grows monotonically with subscribe/unsubscribe churn — the exact scenario the GC is meant to bound.

Suggested fix

Rotate the sampling phase per session before the first inspection, so consecutive sessions visit every modulo class:

// DistWorkerCoProc, gc handler: after the cursor seek
int phase = stepUsed > 1 ? Math.floorMod(gcSessionSeq.getAndIncrement(), stepUsed) : 0;
while (phase-- > 0) {
    itr.next();
    if (!itr.isValid()) {
        break; // tail reached; the scan loop below wraps to the first key
    }
}

No wire or protocol change (the phase is derived from a per-range session counter inside the co-proc). A key in the worst case waits at most step sessions to be inspected — acceptable for a sampling sweep.

Test

New regression test (DistWorkerCoProcGCTest): 12 keys, stepHint = 2, successive sessions fed by nextStartKey; asserts every key is inspected within a few sessions. Fails on the unfixed code (6/12 inspected forever), passes with the fix (12/12 within 8 sessions).

Production validation

A 6-node BifroMQ 4.0.0 cluster (standalone.sh, systemd, mTLS clientAuth=REQUIRE), fronted by a TCP load balancer with 6 backends. Live clients: 3 bridge replicas (MQTT5, $oshare subscribers) and 10 backend service instances (persistent sessions). RocksDB engine, default configuration otherwise.

The fix (plus other changes) was rolled across all six nodes one node at a time (wholesale lib/ replacement, then restart), with a per-node gate (service active, all four ports listening, cluster mesh re-established). Six of six nodes succeeded, zero rollbacks.

Post-deploy checks (all measured after the rollout):

Check Result
Cluster mesh (inter-node 8898/8899 connections, per node) 408–417 established on every node
Live client connections (3 bridge replicas + 10 backend services over mTLS) 13/13 reconnected and stable
Cross-node delivery end-to-end publish on one node → bridge (connected elsewhere) consumed it → forwarded (counter 0 → 1)
Client-visible protocol smoke connect / subscribe / publish with QoS 1 all CONNACK 0 / PUBACK 0

Route-table size and dist.gc.* metrics are being tracked against a captured baseline after the deployment (the leak manifests as monotonic growth under churn).

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions