From e93f2e37f87ca8c2e1f7f7229526e85d36031cc9 Mon Sep 17 00:00:00 2001 From: giokur Date: Wed, 23 Sep 2026 18:40:56 +0200 Subject: [PATCH] fix(redis): ask INFO, not CLUSTER INFO, whether this is a cluster MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit cluster_enabled lives in INFO's Cluster section. CLUSTER INFO reports the state of a cluster — cluster_state, cluster_slots_*, cluster_known_nodes — and never carries cluster_enabled, so `clusterInfo().contains("cluster_enabled:1")` is false against every server there is. The failure is one-sided and silent. On dev and qa, single-node Memorystore, the wrong answer is the right one and everything passed. On production's six-node cluster the probe reported "cluster mode disabled", getResetToken took the standalone branch, and the first MGET came back JedisMovedDataException: MOVED 4478 redis-cluster-4.internal.openframe.ai:6379 46 times over the 90s poll, which the rotator reports as a missing token. That took the production tenant report down on the 16:00 run of 2026-09-23 — the first run after 6.36.10 rolled out — and it would have taken every later run with it. Behaviour where the probe cannot connect is unchanged: assume a cluster for that one lookup, remember nothing. RedisConfig.getConfiguredCluster() still wins, and configs/prod pins cluster: true so production does not wait on this release. Co-Authored-By: Claude Opus 5 (1M context) --- .../com/openframe/test/data/redis/Redis.java | 16 +++++++++++++--- 1 file changed, 13 insertions(+), 3 deletions(-) diff --git a/openframe-test-service-core/src/main/java/com/openframe/test/data/redis/Redis.java b/openframe-test-service-core/src/main/java/com/openframe/test/data/redis/Redis.java index ee8f8805ae..a54139ac0c 100644 --- a/openframe-test-service-core/src/main/java/com/openframe/test/data/redis/Redis.java +++ b/openframe-test-service-core/src/main/java/com/openframe/test/data/redis/Redis.java @@ -146,15 +146,25 @@ private static boolean clusterMode(JedisClientConfig config) { } } - /** What the server says about itself, or {@code null} when it did not answer. */ + /** + * What the server says about itself, or {@code null} when it did not answer. + * + *

{@code INFO cluster}, not {@code CLUSTER INFO}. The two are easy to confuse and only one of + * them answers this question: {@code cluster_enabled} lives in INFO's Cluster section, while + * CLUSTER INFO reports a cluster's state - {@code cluster_state}, {@code cluster_slots_*}, + * {@code cluster_known_nodes} - and never carries {@code cluster_enabled} at all. Matching on it + * there is therefore false for every server alive, including a real cluster, which then takes the + * standalone path and answers MOVED to the first MGET. That is exactly how this broke the + * production tenant report on 2026-09-23. + */ private static Boolean probeCluster(JedisClientConfig config) { HostAndPort node = RedisConfig.getNode(); try (Jedis jedis = new Jedis(node, config)) { - boolean enabled = jedis.clusterInfo().contains("cluster_enabled:1"); + boolean enabled = jedis.info("cluster").contains("cluster_enabled:1"); log.info("Redis at {} reports cluster mode {}", node, enabled ? "enabled" : "disabled"); return enabled; } catch (Exception e) { - log.warn("Could not read CLUSTER INFO from {}; assuming a cluster for this lookup", node, e); + log.warn("Could not read INFO from {}; assuming a cluster for this lookup", node, e); return null; } }