Skip to content

[BUG] KVRangeRestorer leaves the restore future uncompleted when startRestore() throws, permanently wedging the range #303

Description

@ImDanXie

Environment

  • BiFroMQ 4.0.0 / main @ eef5e3af
  • Module: base-kv/base-kv-store-server
  • File: base-kv/base-kv-store-server/src/main/java/org/apache/bifromq/basekv/store/range/KVRangeRestorer.java (restoreFrom)

Summary

restoreFrom() wraps snapshot-install setup in try { ... } catch (Throwable t) { log.error(...) } and returns the session's doneFuture. If startRestore() (or the subsequent messenger.send()) throws synchronously, the future is never completed — and there is no timeout on this path (unlike the receiver-driven path, which has an idle timeout). Worse, the session stays cached in currentSession: a retry with the same snapshot from the same leader hits the reuse branch and returns the same dead future, so the range is permanently stuck — it never applies anything and never retries.

Snapshot install is mandatory on expansion/rebalance, so any transient synchronous failure here is a hard wedge.

Suggested fix

Complete the future exceptionally in the catch block; the caller observes the failure and a retry starts a fresh session:

} catch (Throwable t) {
    log.error("Unexpected error", t);
    onDone.completeExceptionally(new KVRangeStoreException("Snapshot restore failed to start", t));
}

Test

New regression test: startRestore throws → the future must be completed exceptionally, and a retry with the same snapshot must start a fresh session (not reuse the dead one). Fails on the unfixed code.

Production validation

A 6-node BifroMQ 4.0.0 cluster (standalone.sh, systemd, mTLS clientAuth=REQUIRE), fronted by a TCP load balancer with 6 backends. Live clients: 3 bridge replicas (MQTT5, $oshare subscribers) and 10 backend service instances (persistent sessions). RocksDB engine, default configuration otherwise.

The fix (plus other changes) was rolled across all six nodes one node at a time (wholesale lib/ replacement, then restart), with a per-node gate (service active, all four ports listening, cluster mesh re-established). Six of six nodes succeeded, zero rollbacks.

Post-deploy checks (all measured after the rollout):

Check Result
Cluster mesh (inter-node 8898/8899 connections, per node) 408–417 established on every node
Live client connections (3 bridge replicas + 10 backend services over mTLS) 13/13 reconnected and stable
Cross-node delivery end-to-end publish on one node → bridge (connected elsewhere) consumed it → forwarded (counter 0 → 1)
Client-visible protocol smoke connect / subscribe / publish with QoS 1 all CONNACK 0 / PUBACK 0

No wedged range (a range stuck in snapshot restore) was observed before or after the deployment; the fix closes the permanent-wedge path should a synchronous start failure ever occur.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions