Environment
- BiFroMQ
4.0.0 / main @ eef5e3af
- Module:
base-kv/base-kv-store-server
- File:
base-kv/base-kv-store-server/src/main/java/org/apache/bifromq/basekv/store/range/KVRangeRestorer.java (restoreFrom)
Summary
restoreFrom() wraps snapshot-install setup in try { ... } catch (Throwable t) { log.error(...) } and returns the session's doneFuture. If startRestore() (or the subsequent messenger.send()) throws synchronously, the future is never completed — and there is no timeout on this path (unlike the receiver-driven path, which has an idle timeout). Worse, the session stays cached in currentSession: a retry with the same snapshot from the same leader hits the reuse branch and returns the same dead future, so the range is permanently stuck — it never applies anything and never retries.
Snapshot install is mandatory on expansion/rebalance, so any transient synchronous failure here is a hard wedge.
Suggested fix
Complete the future exceptionally in the catch block; the caller observes the failure and a retry starts a fresh session:
} catch (Throwable t) {
log.error("Unexpected error", t);
onDone.completeExceptionally(new KVRangeStoreException("Snapshot restore failed to start", t));
}
Test
New regression test: startRestore throws → the future must be completed exceptionally, and a retry with the same snapshot must start a fresh session (not reuse the dead one). Fails on the unfixed code.
Production validation
A 6-node BifroMQ 4.0.0 cluster (standalone.sh, systemd, mTLS clientAuth=REQUIRE), fronted by a TCP load balancer with 6 backends. Live clients: 3 bridge replicas (MQTT5, $oshare subscribers) and 10 backend service instances (persistent sessions). RocksDB engine, default configuration otherwise.
The fix (plus other changes) was rolled across all six nodes one node at a time (wholesale lib/ replacement, then restart), with a per-node gate (service active, all four ports listening, cluster mesh re-established). Six of six nodes succeeded, zero rollbacks.
Post-deploy checks (all measured after the rollout):
| Check |
Result |
| Cluster mesh (inter-node 8898/8899 connections, per node) |
408–417 established on every node |
| Live client connections (3 bridge replicas + 10 backend services over mTLS) |
13/13 reconnected and stable |
| Cross-node delivery end-to-end |
publish on one node → bridge (connected elsewhere) consumed it → forwarded (counter 0 → 1) |
| Client-visible protocol smoke |
connect / subscribe / publish with QoS 1 all CONNACK 0 / PUBACK 0 |
No wedged range (a range stuck in snapshot restore) was observed before or after the deployment; the fix closes the permanent-wedge path should a synchronous start failure ever occur.
Environment
4.0.0/main @ eef5e3afbase-kv/base-kv-store-serverbase-kv/base-kv-store-server/src/main/java/org/apache/bifromq/basekv/store/range/KVRangeRestorer.java(restoreFrom)Summary
restoreFrom()wraps snapshot-install setup intry { ... } catch (Throwable t) { log.error(...) }and returns the session'sdoneFuture. IfstartRestore()(or the subsequentmessenger.send()) throws synchronously, the future is never completed — and there is no timeout on this path (unlike the receiver-driven path, which has an idle timeout). Worse, the session stays cached incurrentSession: a retry with the same snapshot from the same leader hits the reuse branch and returns the same dead future, so the range is permanently stuck — it never applies anything and never retries.Snapshot install is mandatory on expansion/rebalance, so any transient synchronous failure here is a hard wedge.
Suggested fix
Complete the future exceptionally in the catch block; the caller observes the failure and a retry starts a fresh session:
Test
New regression test:
startRestorethrows → the future must be completed exceptionally, and a retry with the same snapshot must start a fresh session (not reuse the dead one). Fails on the unfixed code.Production validation
A 6-node BifroMQ 4.0.0 cluster (
standalone.sh, systemd, mTLSclientAuth=REQUIRE), fronted by a TCP load balancer with 6 backends. Live clients: 3 bridge replicas (MQTT5,$osharesubscribers) and 10 backend service instances (persistent sessions). RocksDB engine, default configuration otherwise.The fix (plus other changes) was rolled across all six nodes one node at a time (wholesale
lib/replacement, then restart), with a per-node gate (service active, all four ports listening, cluster mesh re-established). Six of six nodes succeeded, zero rollbacks.Post-deploy checks (all measured after the rollout):
No wedged range (a range stuck in snapshot restore) was observed before or after the deployment; the fix closes the permanent-wedge path should a synchronous start failure ever occur.