Environment
- BiFroMQ
4.0.0 / main @ eef5e3af
- Module:
bifromq-dist/bifromq-dist-worker
- Files:
DistWorkerCoProc.java (gc handler), DistWorkerCleaner.java (cursor bookkeeping)
Summary
When the cleaner sweeps with stepHint > 1, the co-proc inspects every step-th key beneath a per-range cursor. The cleaner resumes each session at the key where the previous one stopped, and when a session wraps it stops at its own start key (the wrap-termination compares against the session start). The cursor is therefore pinned: every session starts at the same key and, with a fixed step, samples the same modulo class of the route table forever. Routes in the remaining classes are never inspected, so dead routes in them are never collected.
Reproduction
Unit-level (deterministic): drive successive GCRequests with stepHint = 2 against an in-memory key set, feeding each reply's nextStartKey back into the next request. The set of inspected keys stops growing after the first session: 6 of 12 keys are inspected, the other 6 never.
Observed on a live broker under subscription churn: the same ~50% of routes remained un-inspected across 400/900 consecutive GC rounds (the skipped set is stable, not random).
Impact
Dead routes in the skipped classes are never collected. The route table (and the memory/cache working set) grows monotonically with subscribe/unsubscribe churn — the exact scenario the GC is meant to bound.
Suggested fix
Rotate the sampling phase per session before the first inspection, so consecutive sessions visit every modulo class:
// DistWorkerCoProc, gc handler: after the cursor seek
int phase = stepUsed > 1 ? Math.floorMod(gcSessionSeq.getAndIncrement(), stepUsed) : 0;
while (phase-- > 0) {
itr.next();
if (!itr.isValid()) {
break; // tail reached; the scan loop below wraps to the first key
}
}
No wire or protocol change (the phase is derived from a per-range session counter inside the co-proc). A key in the worst case waits at most step sessions to be inspected — acceptable for a sampling sweep.
Test
New regression test (DistWorkerCoProcGCTest): 12 keys, stepHint = 2, successive sessions fed by nextStartKey; asserts every key is inspected within a few sessions. Fails on the unfixed code (6/12 inspected forever), passes with the fix (12/12 within 8 sessions).
Production validation
A 6-node BifroMQ 4.0.0 cluster (standalone.sh, systemd, mTLS clientAuth=REQUIRE), fronted by a TCP load balancer with 6 backends. Live clients: 3 bridge replicas (MQTT5, $oshare subscribers) and 10 backend service instances (persistent sessions). RocksDB engine, default configuration otherwise.
The fix (plus other changes) was rolled across all six nodes one node at a time (wholesale lib/ replacement, then restart), with a per-node gate (service active, all four ports listening, cluster mesh re-established). Six of six nodes succeeded, zero rollbacks.
Post-deploy checks (all measured after the rollout):
| Check |
Result |
| Cluster mesh (inter-node 8898/8899 connections, per node) |
408–417 established on every node |
| Live client connections (3 bridge replicas + 10 backend services over mTLS) |
13/13 reconnected and stable |
| Cross-node delivery end-to-end |
publish on one node → bridge (connected elsewhere) consumed it → forwarded (counter 0 → 1) |
| Client-visible protocol smoke |
connect / subscribe / publish with QoS 1 all CONNACK 0 / PUBACK 0 |
Route-table size and dist.gc.* metrics are being tracked against a captured baseline after the deployment (the leak manifests as monotonic growth under churn).
Environment
4.0.0/main @ eef5e3afbifromq-dist/bifromq-dist-workerDistWorkerCoProc.java(gc handler),DistWorkerCleaner.java(cursor bookkeeping)Summary
When the cleaner sweeps with
stepHint > 1, the co-proc inspects every step-th key beneath a per-range cursor. The cleaner resumes each session at the key where the previous one stopped, and when a session wraps it stops at its own start key (the wrap-termination compares against the session start). The cursor is therefore pinned: every session starts at the same key and, with a fixed step, samples the same modulo class of the route table forever. Routes in the remaining classes are never inspected, so dead routes in them are never collected.Reproduction
Unit-level (deterministic): drive successive
GCRequests withstepHint = 2against an in-memory key set, feeding each reply'snextStartKeyback into the next request. The set of inspected keys stops growing after the first session: 6 of 12 keys are inspected, the other 6 never.Observed on a live broker under subscription churn: the same ~50% of routes remained un-inspected across 400/900 consecutive GC rounds (the skipped set is stable, not random).
Impact
Dead routes in the skipped classes are never collected. The route table (and the memory/cache working set) grows monotonically with subscribe/unsubscribe churn — the exact scenario the GC is meant to bound.
Suggested fix
Rotate the sampling phase per session before the first inspection, so consecutive sessions visit every modulo class:
No wire or protocol change (the phase is derived from a per-range session counter inside the co-proc). A key in the worst case waits at most
stepsessions to be inspected — acceptable for a sampling sweep.Test
New regression test (
DistWorkerCoProcGCTest): 12 keys,stepHint = 2, successive sessions fed bynextStartKey; asserts every key is inspected within a few sessions. Fails on the unfixed code (6/12 inspected forever), passes with the fix (12/12 within 8 sessions).Production validation
A 6-node BifroMQ 4.0.0 cluster (
standalone.sh, systemd, mTLSclientAuth=REQUIRE), fronted by a TCP load balancer with 6 backends. Live clients: 3 bridge replicas (MQTT5,$osharesubscribers) and 10 backend service instances (persistent sessions). RocksDB engine, default configuration otherwise.The fix (plus other changes) was rolled across all six nodes one node at a time (wholesale
lib/replacement, then restart), with a per-node gate (service active, all four ports listening, cluster mesh re-established). Six of six nodes succeeded, zero rollbacks.Post-deploy checks (all measured after the rollout):
Route-table size and
dist.gc.*metrics are being tracked against a captured baseline after the deployment (the leak manifests as monotonic growth under churn).