Skip to content

Modbus.Tcp bridge tasks not reliably re-attached on a second config-enable of device components within one JVM lifetime #3875

Description

@newman85366-commits

Description

Affected version: 2026.7.0 (edge), Docker deployment, Felix config admin,
Bridge.Modbus.Tcp + custom device components following the
AbstractOpenemsModbusComponent classic-binding pattern.

Summary

Toggling a Modbus device component's enabled configuration property
true -> false -> true at runtime (via updateComponentConfig JSON-RPC)
re-binds cleanly the FIRST time after a JVM boot. A SECOND
disable -> enable cycle within the same JVM lifetime frequently leaves the
device component in a state where enabled=true and the component is
ACTIVE per the config readback, but its Modbus tasks are dead: reads
return null, writes are silently lost. No error is logged.

Reproduction (as observed in our deployment)

  1. Edge with a Bridge.Modbus.Tcp (modbus0) and a device component
    bound to it; both enabled=true, telemetry flowing.
  2. Via JSON-RPC updateComponentConfig, set the device component AND its
    bridge enabled=false. Wait for deactivation.
  3. Re-enable both (bridge first, then device). Telemetry resumes -
    binding works after the FIRST cycle.
  4. Repeat steps 2-3 a second time in the same JVM. With high probability
    the device component reports enabled=true but its channels stay
    null; FC read tasks are not queued to the bridge.

Expected: every disable -> enable cycle re-attaches the protocol
tasks, or a failure is surfaced (component FAULT state / log).

Observed: silent dead binding; only detectable by watching a live
telemetry channel for non-null values (config readback is misleading).

Context / impact

We use runtime enable/disable of the Modbus masters as the single-writer
fence in a hot-standby redundancy scheme (two edges, one set of devices).
Failover promotes the standby by enabling its bridges+devices; failback
demotes and later re-promotes - the second enable in one JVM lifetime
then hits this defect. SCR rebind of only the BRIDGE does not recover the
device; the device component's own activate() is what re-establishes the
binding, and only reliably on the FIRST activation per JVM.

Workarounds we ship (for other users hitting this)

  1. Functional health confirmation: after any enable, poll one live
    telemetry channel per device for a non-null value - never trust
    enabled=true alone.
  2. Re-cycle (disable -> enable) the dead components once on detection.
  3. Last resort: restart the container (fresh JVM) - the first enable
    after boot always binds; a boot-time lease guard keeps the restarted
    instance fenced until it should be active.

We are happy to provide logs, our reproduction harness, or test a patch.

Screenshots

No response

Operating System

Docker, OpenEMS Edge 2026.7.0, Felix config admin

How to reproduce the Error?

reproducibly on the second disable→enable cycle within one JVM

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions