Skip to content

[BUG] SessionRegistry kick re-enters the registry compute block and a delayed removal can drop a re-registered session #304

Description

@ImDanXie

Environment

  • BiFroMQ 4.0.0 / main @ eef5e3af
  • Module: bifromq-session-dict
  • File: bifromq-session-dict/src/main/java/org/apache/bifromq/sessiondict/server/SessionRegistry.java

Summary

SessionRegistry.add() invokes the kicked register's callback inside the tenantSessions.compute() block. The real register answers kick() by unregistering itself synchronously on the same thread, re-entering compute() for the same key; on the same-owner takeover branch that nested remove() deletes the mapping add() has just installed — the new session is silently lost from the dictionary, and a subsequent kill cannot find it.

Related: when a taken-over register is torn down late, its remove() can delete the mapping of the replacement session registered by the same clientId (session-dictionary counts drift).

Trigger: same-clientId rapid reconnect — routine for diskless clients that reboot (and thus reconnect) frequently.

Suggested fix

  • collect the kick and perform it after leaving the compute block (the notification is best-effort; the previous stream may already be closed);
  • remove() only drops the dictionary entry when the removed register is still the live one for that owner, so a delayed teardown cannot delete the new session (or its counters).

Test

Two regression tests (re-entrant kick; takeover + delayed removal) fail on the unfixed code; the full session-dict suite (18/18) passes with the fix.

Production validation

A 6-node BifroMQ 4.0.0 cluster (standalone.sh, systemd, mTLS clientAuth=REQUIRE), fronted by a TCP load balancer with 6 backends. Live clients: 3 bridge replicas (MQTT5, $oshare subscribers) and 10 backend service instances (persistent sessions). RocksDB engine, default configuration otherwise.

The fix (plus other changes) was rolled across all six nodes one node at a time (wholesale lib/ replacement, then restart), with a per-node gate (service active, all four ports listening, cluster mesh re-established). Six of six nodes succeeded, zero rollbacks.

Post-deploy checks (all measured after the rollout):

Check Result
Cluster mesh (inter-node 8898/8899 connections, per node) 408–417 established on every node
Live client connections (3 bridge replicas + 10 backend services over mTLS) 13/13 reconnected and stable
Cross-node delivery end-to-end publish on one node → bridge (connected elsewhere) consumed it → forwarded (counter 0 → 1)
Client-visible protocol smoke connect / subscribe / publish with QoS 1 all CONNACK 0 / PUBACK 0

The rolling restarts forced all 13 clients (including the bridge replicas) through same-clientId rapid reconnects; reconnects landed cleanly with no session-dictionary inconsistency observed.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions