Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
25 commits
Select commit Hold shift + click to select a range
4e7071e
ci: add an end-to-end test for operator-managed mode (#176)
usiegj00 Sep 20, 2026
d29e685
Fix cluster_ok metric leak by handling RedisFailover deletion via a f…
Saremox Sep 20, 2026
d350940
Cut CI wall-clock time and fix the broken e2e pod-readiness gate (#177)
Saremox Sep 20, 2026
c64fdea
test: give integration-test operators an explicit short SyncInterval …
Saremox Sep 20, 2026
fee5935
chore(deps): bump actions/checkout from 6 to 7 (#183)
dependabot[bot] Sep 21, 2026
674cf3a
Potential fix for code scanning alert no. 28: Workflow does not conta…
Saremox Sep 24, 2026
b3b24e4
Disconnect a demoted master's clients once it leaves the master Servi…
Saremox Sep 25, 2026
08a63be
Add a skill for running kind in the Claude Code cloud sandbox (#187)
Saremox Sep 25, 2026
8f84d29
Fall back to a mirror for every image the kind skill pulls (#188)
Saremox Sep 25, 2026
8314bf5
Pass watch bookmarks and errors through the namespace filter (#186)
Saremox Sep 25, 2026
d075fc5
Reconcile a RedisFailover when its pods change (#181)
Saremox Sep 25, 2026
b519090
Drop the kooper dependency (#185)
Saremox Sep 25, 2026
6163d62
fix: SentinelCheckQuorum NOQUORUM dead code + MakeSlaveOfWithPort wro…
Saremox Sep 25, 2026
de80035
feat(exporter): make the exporter metrics port configurable (dnse #30…
Saremox Sep 25, 2026
b234930
feat(heal): optionally protect the redis master from autoscaler evict…
Saremox Sep 26, 2026
480be21
feat(sentinel): configurable deployment strategy and PDB minAvailable…
Saremox Sep 26, 2026
6f877ec
feat(env): allow custom env vars on redis and sentinel main container…
Saremox Sep 26, 2026
f43873c
Update dependabot.yml to remove Kubernetes ignore rule
Saremox Sep 26, 2026
616fcfd
feat(redis): operator-managed maxmemory and maxmemory-policy (#190)
Saremox Sep 26, 2026
a011213
feat(redis): resize redis pods in place (#191)
Saremox Sep 26, 2026
3494105
fix(failover): don't promote while a ready Redis pod doesn't answer (…
Saremox Sep 27, 2026
905cfd8
fix(chart): fit the CRD in the upgrade hook's ConfigMap (#195)
Saremox Sep 27, 2026
759b67b
fix(chart): set fsGroup on the operator pod, not its container (#196)
Saremox Sep 27, 2026
861a011
fix(auth): apply a changed Redis password in place (#194)
Saremox Sep 27, 2026
7da3fa2
Merge remote-tracking branch 'upstream/main' into sync/upstream-2026-…
usiegj00 Sep 29, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
88 changes: 88 additions & 0 deletions .claude/skills/kind-cluster/SKILL.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,88 @@
---
name: kind-cluster
description: Create a kind Kubernetes cluster inside the Claude Code cloud sandbox and run this repo's integration tests or a Helm-installed operator against it. Use when a change needs e2e testing on a real API server.
---

# kind cluster in the cloud sandbox

A plain `kind create cluster` fails in the sandbox, and a working cluster
still can't pull images. The scripts here handle all of this.

## Create

```sh
.claude/skills/kind-cluster/kind-up.sh e2e v1.35.0 244
export KUBECONFIG=/tmp/kind-e2e/kubeconfig
```

The script:

1. Starts `dockerd` if it isn't running.
2. Writes a kind config with two sandbox fixes:
- `failCgroupV1: false`: the sandbox is cgroup v1, and kubelet >= 1.35
refuses to start on it.
- containerd `restrict_oom_score_adj = true`: the kernel rejects negative
`oom_score_adj`, so every pod sandbox would fail with `can't get final
child's PID from pipe: EOF`. This affects every node version, 1.34 too.
3. Makes a local registry, `kind-registry`, the node's mirror for docker.io
and quay.io (`registry.sh`).
- The node can't reach any registry, because its `HTTPS_PROXY` points at
the sandbox's 127.0.0.1 proxy.
- `kind load` isn't enough: the operator defaults to `imagePullPolicy:
Always`, so preloaded images still fail with `ImagePullBackOff`.
4. Copies the images from `api/redisfailover/v1/defaults.go`, plus any in
`EXTRA_IMAGES`, into the registry.
- quay.io is blocked from the host too, so `registry.sh` falls back to the
Docker Hub and `mirror.gcr.io` copies.
- Add another image any time with `registry.sh push IMAGE`.
5. Adds a host route to the pod subnet: the integration tests connect to
Redis pod IPs.
6. Applies the RedisFailover CRD.

## Several clusters at once

Give each cluster its own name and subnet octet, e.g. `a … 181` and
`b … 185`. The pod subnets must differ, or the host routes collide. The
registry is shared. The machine has 4 CPUs, so two single-node clusters is
the practical limit.

## Integration tests

```sh
go test ./test/integration/... -tags integration -v -timeout 30m
```

They run the operator in-process against `$KUBECONFIG`.

## Operator image and Helm

```sh
.claude/skills/kind-cluster/build-image.sh pr # redis-operator:pr
helm upgrade --install redis-operator ./charts/redisoperator \
--set image.repository=redis-operator --set image.tag=pr --wait
```

- `docker/app/Dockerfile` fails here, because `apk add` can't reach the
package mirrors. `build-image.sh` builds the binary on the host instead.
- Install helm with `GOBIN=/usr/local/bin go install helm.sh/helm/v3/cmd/helm@v3.19.0`
if it's missing.
- `.github/workflows/e2e.yml` has a RedisFailover manifest and checks you
can reuse.
- Don't run the integration tests while a Helm-installed operator is
running: both would reconcile the test's RedisFailovers.

## Debugging a failed create

Add `--retain` to keep the node. Then read
`docker exec NAME-control-plane journalctl -u kubelet --no-pager` and
`journalctl -u containerd`.

## Clean up

```sh
kind delete cluster --name e2e
docker rm -f kind-registry # once no cluster needs it
```

This leaves the host route behind. It is harmless, and a new cluster
replaces it.
29 changes: 29 additions & 0 deletions .claude/skills/kind-cluster/build-image.sh
Original file line number Diff line number Diff line change
@@ -0,0 +1,29 @@
#!/usr/bin/env bash
# Builds the operator from the current checkout as redis-operator:TAG and
# serves it to the kind nodes through the local registry.
#
# Usage: build-image.sh [TAG] (default TAG: dev)
#
# docker/app/Dockerfile can't be used in the sandbox: its `apk add` steps have
# no route to the package mirrors. Build the binary on the host instead and
# copy it into the same alpine base with the same non-root user.
set -euo pipefail

tag=${1:-dev}
repo=$(git rev-parse --show-toplevel)
here=$(cd "$(dirname "$0")" && pwd)
ctx=$(mktemp -d)
trap 'rm -rf "$ctx"' EXIT

(cd "$repo" && CGO_ENABLED=0 go build -o "$ctx/redis-operator" -ldflags "-w" ./cmd/redisoperator)
cat >"$ctx/Dockerfile" <<'EOF'
FROM alpine:latest
COPY redis-operator /usr/local/bin/redis-operator
RUN addgroup -g 1000 rf && adduser -D -u 1000 -G rf rf
USER rf
ENTRYPOINT ["/usr/local/bin/redis-operator"]
EOF
# docker build pulls a missing base image straight from Docker Hub.
"$here/registry.sh" pull alpine:latest
docker build -q -t "redis-operator:$tag" "$ctx" >/dev/null
"$here/registry.sh" push "redis-operator:$tag"
73 changes: 73 additions & 0 deletions .claude/skills/kind-cluster/kind-up.sh
Original file line number Diff line number Diff line change
@@ -0,0 +1,73 @@
#!/usr/bin/env bash
# Creates a kind cluster that works inside the Claude Code cloud sandbox and
# prepares it for this repo's integration tests.
#
# Usage: kind-up.sh NAME [NODE_VERSION] [SUBNET]
# NAME cluster name; state goes to /tmp/kind-NAME
# NODE_VERSION kindest/node tag (default v1.35.0)
# SUBNET second octet of the pod subnet 10.SUBNET.0.0/16 (default 244).
# Use a different one per cluster when running several.
set -euo pipefail

name=${1:?usage: kind-up.sh NAME [NODE_VERSION] [SUBNET]}
version=${2:-v1.35.0}
subnet=${3:-244}
dir=/tmp/kind-$name
repo=$(git rev-parse --show-toplevel)
here=$(cd "$(dirname "$0")" && pwd)
pod_cidr=10.$subnet.0.0/16
node=$name-control-plane
mkdir -p "$dir"

if ! docker info >/dev/null 2>&1; then
(dockerd >/tmp/dockerd.log 2>&1 &)
for _ in $(seq 1 30); do docker info >/dev/null 2>&1 && break; sleep 1; done
fi

# The sandbox runs cgroup v1, which kubelet >= 1.35 refuses by default, and
# its kernel rejects negative oom_score_adj values, which containerd sets
# on every pod sandbox unless restricted.
cat >"$dir/kind.yaml" <<EOF
kind: Cluster
apiVersion: kind.x-k8s.io/v1alpha4
networking:
podSubnet: $pod_cidr
serviceSubnet: 10.$((subnet + 1)).0.0/16
containerdConfigPatches:
- |-
[plugins."io.containerd.grpc.v1.cri"]
restrict_oom_score_adj = true
[plugins."io.containerd.grpc.v1.cri".registry]
config_path = "/etc/containerd/certs.d"
nodes:
- role: control-plane
kubeadmConfigPatches:
- |
kind: KubeletConfiguration
failCgroupV1: false
EOF

"$here/registry.sh" pull "kindest/node:$version"
kind create cluster --name "$name" --image "kindest/node:$version" \
--config "$dir/kind.yaml" --kubeconfig "$dir/kubeconfig" --wait 180s
export KUBECONFIG=$dir/kubeconfig

# The node can't reach any registry: its inherited HTTPS_PROXY is the
# sandbox's 127.0.0.1 proxy. Serve images from a local registry that
# mirrors docker.io and quay.io, so pods with pullPolicy Always work too.
"$here/registry.sh" connect "$name"
images=$(grep -oE '"[^"]+:[^"]+"' "$repo/api/redisfailover/v1/defaults.go" | tr -d '"' | sort -u)
for img in $images ${EXTRA_IMAGES:-}; do
"$here/registry.sh" push "$img"
done

# The integration tests talk to Redis pod IPs directly. The sandbox has no
# `ip` binary, so borrow the node image's.
node_ip=$(docker inspect -f '{{range .NetworkSettings.Networks}}{{.IPAddress}}{{end}}' "$node")
docker run --rm --net=host --privileged --entrypoint ip "kindest/node:$version" \
route replace "$pod_cidr" via "$node_ip"

kubectl apply --server-side -f "$repo/manifests/databases.spotahome.com_redisfailovers.yaml"

echo
echo "export KUBECONFIG=$dir/kubeconfig"
73 changes: 73 additions & 0 deletions .claude/skills/kind-cluster/registry.sh
Original file line number Diff line number Diff line change
@@ -0,0 +1,73 @@
#!/usr/bin/env bash
# Local registry that the kind nodes use as their docker.io and quay.io
# mirror.
#
# Usage: registry.sh connect CLUSTER start the registry, make it CLUSTER's mirror
# registry.sh push IMAGE copy IMAGE from the host into it
# registry.sh pull IMAGE pull IMAGE to the host, via a mirror if needed
set -euo pipefail

cmd=${1:-}
arg=${2:?usage: registry.sh connect CLUSTER | push IMAGE | pull IMAGE}
reg=kind-registry

# Repository path without the registry host, as containerd asks the mirror
# for it: redis:7 -> library/redis:7, quay.io/a/b:1 -> a/b:1.
repo_path() {
local first=${1%%/*}
if [[ $1 != */* ]]; then
echo "library/$1"
elif [[ $first == *.* || $first == *:* ]]; then
echo "${1#*/}"
else
echo "$1"
fi
}

# Pulls IMAGE unless the host has it. quay.io is blocked and Docker Hub may
# rate limit, so fall back to Docker Hub's copy and mirror.gcr.io.
ensure_local() {
local img=$1 path src
docker image inspect "$img" >/dev/null 2>&1 && return
path=$(repo_path "$img")
for src in "$img" "$path" "mirror.gcr.io/$path"; do
if docker pull -q "$src" >/dev/null 2>&1; then
[[ $src == "$img" ]] || docker tag "$src" "$img"
return
fi
done
echo "failed to pull $img" >&2
return 1
}

case $cmd in
connect)
if ! docker inspect "$reg" >/dev/null 2>&1; then
ensure_local registry:2
docker run -d --restart=always --name "$reg" -p 127.0.0.1:5001:5000 registry:2 >/dev/null
fi
docker network connect kind "$reg" 2>/dev/null || true
# Plain HTTP, so containerd doesn't send it through the unreachable
# HTTPS proxy.
for host in docker.io quay.io; do
docker exec "$arg-control-plane" mkdir -p "/etc/containerd/certs.d/$host"
printf '[host."http://%s:5000"]\n capabilities = ["pull", "resolve"]\n' "$reg" |
docker exec -i "$arg-control-plane" cp /dev/stdin "/etc/containerd/certs.d/$host/hosts.toml"
done
;;
pull)
ensure_local "$arg"
;;
push)
img=$arg
path=$(repo_path "$img")
ensure_local "$img"
docker tag "$img" "localhost:5001/$path"
docker push -q --platform linux/amd64 "localhost:5001/$path" >/dev/null
echo "$img -> $reg/$path"
;;
*)
echo "usage: registry.sh connect CLUSTER | push IMAGE | pull IMAGE" >&2
exit 1
;;
esac
2 changes: 1 addition & 1 deletion .github/copilot-instructions.md
Original file line number Diff line number Diff line change
Expand Up @@ -63,7 +63,7 @@ make helm-test

## Kubernetes Operator Patterns

- The operator uses the `kooper` framework (`github.com/spotahome/kooper/v2`) for controller/reconciler wiring
- The controller is built directly on client-go informers and a workqueue (`operator/redisfailover/controller.go`)
- The reconciliation loop is in `operator/redisfailover/`
- All Kubernetes resources created by the operator carry owner references pointing to the `RedisFailover` CR
- Redis Statefulsets use the prefix `rfr-<name>`; Sentinel Deployments use `rfs-<name>`
Expand Down
3 changes: 0 additions & 3 deletions .github/dependabot.yml
Original file line number Diff line number Diff line change
Expand Up @@ -4,9 +4,6 @@ updates:
directory: "/"
schedule:
interval: "daily"
ignore:
# Ignore Kubernetes dependencies to have full control on them.
- dependency-name: "k8s.io/*"
- package-ecosystem: "github-actions"
directory: "/"
schedule:
Expand Down
22 changes: 21 additions & 1 deletion .github/workflows/ci.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -71,7 +71,10 @@ jobs:
integration-test:
name: Integration test
runs-on: ubuntu-24.04
needs: [check, unit-test]
# No `needs` gate: this matrix doesn't depend on check/unit-test's
# output, so gating it behind them only serializes two independent
# wall-clock costs. Letting it start immediately in parallel is strictly
# faster; a failure here fails the run exactly the same either way.
strategy:
fail-fast: false
matrix:
Expand All @@ -81,6 +84,16 @@ jobs:
- uses: actions/setup-go@v7
with:
go-version-file: go.mod
- name: Pre-pull redis image in the background
# driver=none means minikube schedules pods straight onto this
# runner's own docker daemon, so this image is already the exact one
# the in-process operator will ask the cluster to run. Kicking the
# pull off now and only waiting on it right before the Go tests start
# hides its wall time behind conntrack/minikube/CRD setup below,
# instead of paying for it inline while waitForPodsReady polls.
run: |
docker pull redis:7.2.12-alpine &
echo "REDIS_PULL_PID=$!" >> "$GITHUB_ENV"
- name: Install conntrack
run: sudo apt-get install -y conntrack
- name: Prepare CNI config directory
Expand All @@ -102,6 +115,13 @@ jobs:
run: docker build -t ghcr.io/buildio/redis-operator:test -f docker/app/Dockerfile .
- name: Add redisfailover CRD
run: kubectl create -f manifests/databases.spotahome.com_redisfailovers.yaml
- name: Wait for redis image pre-pull to finish
# $REDIS_PULL_PID isn't a child of this step's shell (each step is
# its own process), so `wait` can't be used on it - poll instead.
run: |
while kill -0 "$REDIS_PULL_PID" 2>/dev/null; do
sleep 1
done
- run: make ci-integration-test

chart-test:
Expand Down
Loading
Loading