What is collected, where it goes, and how to change it.
This document describes the estate's stack, on prometheus (10.0.99.20).
There is a second one. stacks/lab is the lab's own
Prometheus, Loki, Grafana and Alloy, and it is deliberately not part of any of
what follows: no series it holds reaches this Prometheus, no log line reaches
this Loki, and none of the alert rules or dashboards below can see it. That is
ADR-0007's decision — lab telemetry stays in the lab, so that deliberately
hostile data never lands in the store the estate is actually run from — and
ADR-0020
settles its shape. It runs on alexander, the guest
#262 built on 2026-09-05. Its
first client from off that guest is phoenix, the deployment host
(ADR-0043,
built 2026-09-20): its native Alloy pushes over the two ports the lab stack
published for it (#436), and
it too reaches nothing here.
There is a third, and it reports to the second. stacks/soc
is Wazuh and Velociraptor on odin, another guest on Saruman
(ADR-0030),
built 2026-09-27. Its Alloy pushes that guest's telemetry and the Wazuh
indexer's health to the lab's stores on alexander; nothing of it reaches
here either.
The one path that does cross belongs to the hypervisor and not to any guest:
Saruman's own agent remote-writes here over a single unlogged pass
(#88). A DL360 with an ageing
mirrored pair is estate hardware, and its health belongs with the rest of the
estate's.
The consequence worth carrying into everything below: nothing here can tell a
quiet lab from a dead one. RemoteWriteJobStale keys on jobs that arrive on
this Prometheus, so by construction it can never cover a stack that never
arrives. That gap is #257, and
it is not closed by anything in this document.
| Source | Via | Interval | Examples |
|---|---|---|---|
| Linux hosts | Alloy → node_exporter |
60s | CPU, memory, filesystem, network, load, clock offset |
smaug, the NAS |
Prometheus → node_exporter |
60s | The same, minus network — the one host that is SCRAPED and not pushed to, and the one that ships no logs (ADR-0016, #255) |
| Docker containers | Alloy → cAdvisor | 60s | Per-container CPU, memory, network, restarts, OOM |
| Container logs | Alloy → Docker socket | stream | stdout/stderr per container |
| systemd journal | Alloy | stream | unit, boot ID, transport, priority. Delivery is watched by JournalSourceStopped |
/var/log/auth.log |
Alloy | 60s poll | sshd, sudo, PAM |
syslog, /var/log/*.log |
Alloy | 60s poll | Everything else. The *.log source skips auth.log and the facility files rsyslog also writes to syslog (kern, user, mail, cron, daemon), so a line is not stored once per copy (#909) |
| pfSense | snmp-exporter | 60s | pf state table, counters, interface stats |
| pfSense logs | syslog → Alloy on 1514 | stream | filterlog decisions, suricata alerts, kea-dhcp4 leases |
| MokerLink switch | snmp-exporter | 60s | Interface status and 64-bit octet counters |
| APC UPS | snmp-exporter | 60s | Charge, runtime, load, voltage, alarms |
| ProLiant iLO | snmp-exporter | 60s | Temperature, PSU, drive and battery health |
| The stack itself | Prometheus | 15s | Every component scrapes itself |
| Each Alloy agent | Alloy → itself | 60s | Remote-write throughput and lag, component health, lines forwarded |
Retention is 30 days for both metrics (PROMETHEUS_RETENTION in .env) and
logs (retention_period in loki/loki-config.yaml). Change both together or
dashboards will show metrics with no matching logs at the far end of the range.
Metrics carry a second bound: PROMETHEUS_RETENTION_SIZE caps the store at
12 GiB, and whichever limit is reached first wins. It is sized from the write
rate rather than the store's current size — 72.9 MiB/day gives 2.28 GiB over
30 days, 2.73 GiB with Saruman reporting — so it is roughly 4.4x the
planned estate and should never bind. In normal operation the 30 days above is
the limit that applies; if the size cap ever binds, that 30 days stops being
true and PrometheusSizeRetentionActive is what says so. Loki has no
equivalent size bound.
Logs arrive spelling severity about twenty different ways — ERROR, err,
eror, crit, fatal, panic, dbug. config.alloy maps all of them onto
five canonical values before they reach Loki:
emerg, panic, corrupt, fatal, alert, crit, critical → critical
err, eror, error → error
warn, warning → warning
info, information, informational, notice → info
dbug, debug, dbg → debug
This is what makes {level="error"} a useful query across a FreeBSD firewall, a
Ubuntu host and a Go container at the same time.
Three mechanisms do it, because the sources carry severity in three different places. They share the table above and must keep sharing it:
| Source | Mechanism | Reads |
|---|---|---|
/var/log files, container stdout |
loki.process "log_processor" |
the line body, by regex |
| pfSense and other network syslog | loki.relabel "network_syslog" |
the syslog PRI severity |
| the systemd journal | discovery.relabel "journal" |
journald's priority keyword |
Where the protocol carries a severity, it is mapped rather than guessed at —
regex-sniffing a line whose PRI already says err is strictly worse. Where it
does not, the body is all there is. Each mechanism defaults to info for
anything it cannot classify, which is what keeps the vocabulary closed: a
severity nobody anticipated lands inside the five values instead of becoming a
sixth.
scripts/check_dashboards.py asserts that the values config.alloy can emit
and the values the Logs dashboard's Level picker offers are the same set, and
that no relabel rule copies a severity into level unmapped. That check exists
because both halves of #83 were invisible in a diff.
This has been silently broken twice. First the extracting regex used
\binside a double-quoted Alloy string, where\bis a backspace escape rather than a word boundary. The regex never matched, so the template's{{ else }}info{{ end }}fallback labelled every lineinfo. It is now a backtick string.Then (#83) the syslog and journal paths copied their severity to
levelinstead of mapping it, and the Docker path set nolevelat all — so pfSense sat atinformational, the journal atnoticeandalert, and containers at nothing. All of it was outside a picker whose setting said "All", which matched about a seventh of what was being ingested.If you edit any of the three mechanisms, verify afterwards that the vocabulary is still closed and still complete:
sum by (level) (count_over_time({host=~".+"}[10m]))At most five series, every name in the table above. Only
infomeans a regex is not matching. Then check that nothing escaped labelling entirely — these two must return the same number:sum(count_over_time({host=~".+"}[10m])) sum(count_over_time({host=~".+", level=~".+"}[10m]))A gap is a source reaching
loki.writewithout passing through one of the three. Queries spanning an edit will show the old and new spellings side by side until retention ages the old ones out; that is not a regression.
Eight dashboards are provisioned from grafana/dashboards/ into a HomeLab
folder:
| Dashboard | UID | Covers |
|---|---|---|
| Host Overview | homelab-host-overview |
CPU, memory, storage, network per host |
| Docker Containers | homelab-docker |
Per-container resources, restarts, OOM kills |
| Network & Firewall | homelab-network |
pf state table, switch interfaces, iLO health |
| UPS & Power | homelab-ups |
Battery, runtime, load, input voltage |
| Logs | homelab-logs |
Volume by level and source, error and auth streams |
| Observability Stack | homelab-stack |
Scrape health for every target, and Prometheus, Loki, Alertmanager and Alloy watching themselves |
| Internet | homelab-internet |
Speed tests beside WAN utilisation, WAN receive errors, round-trip time, latency under load and the firewall's uplink loss — what to read when the internet is slow (#914) |
| Security | homelab-security |
Firewall blocks by interface and direction, top blocked sources, Suricata classification and priority, terminal-segment violations |
The JSON in git is the source of truth: Grafana re-provisions over its own copy whenever a file changes, so a UI edit never outlives the next commit that touches its dashboard. To change one, edit it in the UI, save, and run:
make dashboards-export # ARGS=--check to report drift and write nothingThat pulls every dashboard back by uid and writes it over the file, so the loop
is edit → one command → git diff. It replaced a hand copy out of Dashboard
settings → JSON Model (#100),
and making it work meant setting allowUiUpdates: true: with false, Grafana
refuses to store a UI edit at all, so the export could only ever hand back the
file it started from — a silent no-op over the edit it was meant to capture.
What that flag used to guarantee is now bought explicitly, by the daily
dashboards-drift job and by a CI check that the save path still works;
grafana/dashboards/README.md
has the whole trade, and what the export drops and refuses to overwrite.
CI checks the result parses, that every datasource UID resolves, that panels fit
the grid and do not overlap, and that every panel expression is syntactically
valid — PromQL through promtool, LogQL through a real Loki, since nothing else
parses it. It also boots the pinned Grafana image and asserts the committed JSON
is what Grafana produces from it, so an export never arrives as a diff nobody
can read.
homelab-stack exists because the stack watched four devices and two hosts
attentively and did not watch itself at all
(#81). Two of the three faults
found while verifying #12 would have been visible on it within a minute:
remote_write failing to a stale address, and cAdvisor reporting one series where
it should report hundreds. Both were invisible for hours because the only view
of the collection path was up{job="alloy"}, which stayed 1 throughout.
up is a poor liveness signal for half of what this stack collects, and the
dashboard says so rather than papering over it. Prometheus scrapes twelve jobs
directly; the rest arrive by remote_write — one <host>-metrics and one
<host>-alloy per agent, plus integrations/cadvisor wherever there is
Docker. A directly scraped target that dies sets up to 0. A remote-writing
agent that dies just stops pushing, so its up goes stale and ages out instead
of falling — and InstanceDown is up == 0, so it cannot see that at all.
Eleven of those twelve jobs are containers on the compose network or devices
behind an exporter. The twelfth is node, and it is a MACHINE — smaug, which
ADR-0016 put
on a segment that may not initiate upward, so Prometheus reaches in and scrapes
it rather than being pushed to (#256).
It is the estate's first scraped host, it needs no new down-detection because
InstanceDown has no job matcher, and it is deliberately not called
smaug-metrics: that name would enrol a pulled job in RemoteWriteJobStale,
whose notification says an Alloy agent has stopped pushing. This host has no
Alloy agent, by decision.
The Sample staleness by job panel is what covers the pushed jobs on the
dashboard, and the Every target table puts Staleness next to Up for the
same reason.
RemoteWriteJobStale in stack.rules.yaml is the alert for it, and it is
deliberately not written as a threshold on that staleness panel. time() - timestamp(up) > 300 reads correctly and cannot fire for any input: an instant
selector stops returning a sample once the lookback delta passes, so the
difference is bounded below any threshold worth alerting on. Measured over 24
hours of real data the largest value any job reached was 79 seconds. The rule
therefore asks the question the other way round — which jobs were reporting in
the last 24 hours and are not reporting now — because count_over_time reads a
range and sees through staleness where an instant selector cannot.
Two consequences worth knowing. It matches on the job-name convention
config.alloy builds (<hostname>-metrics, <hostname>-alloy,
integrations/cadvisor) rather than a list, so a new agent is covered the day
it is deployed — Saruman (#88)
needed nothing added. And the 24-hour window is a real bound: an agent that
comes back inside a day resolves the alert truthfully, one that stays away
longer resolves it falsely once the window no longer remembers it, having
notified at least twice by then.
Alongside it, samples returned per scrape is the panel that catches a
collector which is still answering but has stopped exporting most of what it
used to. That is precisely the cAdvisor fault: up at 1, scrape succeeding,
one series where there should be hundreds.
Since #81 each Alloy agent also scrapes itself and remote-writes the result
under <hostname>-alloy, so prometheus_remote_storage_* and
alloy_component_* exist for every agent rather than only the one on this host.
oracle publishes Alloy's port on loopback (ADR-0012) and there is no address
Prometheus could be pointed at; pushing down the pipe that is already open costs
nothing and needs no new exposure. The monitoring host's agent is consequently
collected twice — job="alloy" by direct scrape and job="prometheus-alloy" by
push — which is deliberate: the first is the only one whose up can reach 0.
homelab-security exists because syslog.alloy went to real trouble to extract
interface, action and direction from pfSense filterlog, and
classification and priority from Suricata; six Loki rules fire on them; and
nothing charted any of it (#82).
"Network & Firewall" is SNMP — the pf state table and interface counters — and
says nothing about what the firewall decided. "Logs" counts lines by level. The
dimensions the parsing exists to produce had no view, so the only way to see a
scan, a misconfigured device or a segmentation failure was as an alert that had
already fired.
Two things about it are worth stating here rather than only in a panel description.
Addresses are parsed at query time, not indexed. Top blocked source
addresses runs | regexp over the line body, because ADR-0003 keeps addresses
out of the labels and a label per source address is the textbook way to detonate
Loki's cardinality. The regex anchors on the adjacent src,dst pair rather than
counting CSV fields, because pfSense's IPv6 filterlog layout puts src at a
different index — v6 lines therefore do not appear in that table, which is a
stated limit rather than an oversight.
Its panels cannot tell you the IDS is alive. Suricata watches Skids (VLAN
20) and Degens (VLAN 10), one process each, and a quiet IDS and a stopped one
produce identical log output, so empty Suricata panels are not evidence of
anything on their own. What is: SuricataStopped in
prometheus/rules/ids.rules.yaml, which reads the firewall's process table over
SNMP and fires per declared interface
(#90), and SuricataLogsStopped
in security.rules.yaml, which fires only once both interfaces have been silent
for nine hours — a window read from the stream rather than chosen
(#441); the dashboard's alert
table lists both alongside the six Loki rules that fire on the parsed labels.
The text panel at the top still says so, rather than letting a flat line be
read as calm.
And it found that the firewall logged blocks only. Building the panels
turned up something the alerts had not: across the full 30-day retention the
action label had exactly one value, block, at roughly 85,000 lines a day and
not one pass. TerminalSegmentReachedInternalNetwork matches
{app="filterlog", action="pass"}, so it could not fire for any input — the
same shape of defect as #63,
where ContainerHighMemory divided by a limit no service set and showed as
loaded and healthy throughout.
#223 armed it, and the way it
was armed is the point. Logging the inter-VLAN pass rules would not have
worked: every one of them is sourced from an internal segment, so no packet they
pass can have a terminal VLAN as its source. Logging the terminal VLANs' own
→ any egress rules would have worked and would have cost about 12.6M lines a
day — roughly 142× the existing volume, against a 30-day retention on one disk.
Instead there are four tripwire rules: one per terminal interface
(igc0.10, igc0.20, igc0.40), each a pass + log for
<terminal net> → Internal_Segments, and one on the lab interface (igc0.30),
a pass + log for <lab net> → House_Segments — every segment but its own,
because Internal_Segments includes the lab and would turn its gateway DNS into
a logged crossing (#234). Each is placed below the block rules that already
stop that path and above the → any egress rule. While segmentation holds,
the blocks match first and the tripwire logs nothing, so it adds no volume. It
can only match if those blocks are removed or reordered — the exact failure the
alert exists for — and in that case the → any rule would have passed the
packet anyway, so nothing is weakened by it being there. The terminal three feed
TerminalSegmentReachedInternalNetwork; the lab's feeds
LabSegmentReachedInternalNetwork, which reads VLAN 30 as a source rather than
a destination.
That alert reads two source subnets, not one.
ADR-0042 routes
the WireGuard peers on 172.31.0.0/24 rather than masquerading them, so a
peer's own address reaches igc0.30 and appears in filterlog. That is the
point of routing rather than translating — a peer is nameable in a rule and in
an alert — and the cost is that a source class written as 10.0.30.0/24 no
longer covers the interface. The alert's regex carries both; the interface
carries a second set of blocks and a second tripwire, sourced from the peers
rather than from the segment. A 10.0.30.x source is the lab, a 172.31.x
source is a remote peer, and a 172.30.x source is neither — that is ifrit's
range bridge escaping, which is a different incident.
A tripwire that never fires is indistinguishable from a broken one, which is
this whole family of defect, so the logging path was proven rather than assumed:
logging was briefly enabled on the lowest-volume terminal egress rule, and 49 of
49 resulting pass lines carried the source address in exactly the position the
alert's regex reads. The destination half of the same regex already matched 2058
block lines. Both halves are therefore verified against real traffic; only the
combination is absent, which is what "the segmentation is holding" looks like.
Passed (1h) still reads not logged, because ordinary egress genuinely is
not. Terminal→internal passes reads none when the query matches nothing —
deliberately not 0. Grafana's noValue fires on an empty result, not on a
measured zero, so rendering a number there would claim a measurement in exactly
the case where the stream is broken, absent or relabelled. None of these stats
render a number they did not get; Firewall log arrival rate next door is what
separates a quiet stream from a stopped one.
174 rules in total: 155 metric-based in prometheus/rules/, and 19 log-based in
loki/rules/.
Some conditions only exist in logs. A metric confirms sshd is running; only the
log shows it rejecting forty passwords in five minutes. loki/rules/security.rules.yaml
covers SSH brute force, SSH accepted from outside VLAN 50/99, repeated sudo
failures, user/group creation, kernel OOM kills, read-only remounts and disk I/O
errors.
The five authentication rules read a three-branch union — authlog, then
journal, then syslog constrained to the sshd/sudo apps — joined with
or, rather than the single {log_type="authlog"} selector they all used until
#261. Only two hosts produce that
label, and the two that do not are saruman (journald-only, so there is no
/var/log/auth.log to tail) and morpheus (pfSense, which arrives over syslog).
Those are the two hosts where root can be reached with a password, so a
password-guessing run against either produced no alert at any volume — 54
accepted logins over the week to 2026-09-05, none of them visible to any of the
five rules. Nothing was missing from the store; the rules could not see it.
or rather than a wider selector because it deduplicates: oracle and
prometheus ship the same sshd events twice, once via auth.log and once via
the journal, and summing a combined selector would double-count them and halve
every threshold on the two hosts that already worked. The dashboard's two
auth panels in logs-explorer.json carry the same union for the same reason.
It also covers a device taking its first DHCP lease on a segment — one rule
for Hicks and one for Winterfell, plus the absent_over_time rule that says the
lease stream itself has stopped. That is
ADR-0019: device joins
come from Kea on morpheus rather than from the eero cloud, because the
firewall sees the join on the wire and the cloud sees it two minutes later over
the WAN. "First" is expressed as the last ten minutes unless the seven days
before it, which needs no state anywhere and no list of known devices in the
repository.
They use the same severity and category labels as the Prometheus rules and
are sent to the same Alertmanager, so routing and inhibition are shared.
Loki's local ruler reads <directory>/<tenant>/, and with auth_enabled: false
the tenant is literally fake — hence the loki/rules:/etc/loki/rules/fake
mount in compose.yaml. Getting that path wrong produces no error, just a ruler
that silently evaluates nothing.
promtool cannot validate these; it parses PromQL and rejects every LogQL
stream selector. scripts/check_loki_rules.sh boots the pinned Loki image with
the rules mounted and fails on a parse error, then asserts the ruler actually
evaluated them. Note that loki -verify-config is not sufficient on its own —
it validates the config file and never opens the rule files. A file containing
count_over_time({{{BROKEN passes -verify-config and is caught only by the
boot check.
Parsing is not matching, so scripts/test_loki_rules.py behaviour-tests them
(#843). It pushes the fixture
lines in loki/tests/*.test.yaml into the same pinned Loki, each in its
source's real format and with the labels Alloy attaches. Then it runs each
rule's own expr as an instant query and compares the series returned with
the ones the case expects. Every rule must have a case that fires and one that
does not, or the run refuses to start. The exceptions are named, each with its
reason, in the runner: today only LokiRulerWatchdog, which is vector(1) and
cannot be quiet. A rule that looks back further than the spacing between cases
(the absence rules, the 7-day new-device rules) must have its cases declare a
window or say isolated_by: how they stay apart, and the runner checks that
against the rule's own ranges. It does not exercise for:, because an instant
query is the expression, not the pending period. Its first run found
DiskIoErrors could never match smartd's failure lines.
A rule that parses and can see every host is still only as good as what reaches
Loki, and two rules in stack.rules.yaml watch that:
JournalSourceStoppedfires when an agent that is up and publishing metrics has read no journal entries for two hours. Zero is a safe assertion rather than a tuned threshold because the quietest host in the estate,Saruman, still reads about three entries an hour — measured, not assumed.LogEntriesDroppedfires when Loki rejects what an agent sends, and only when it keeps doing so for an hour. There is no retry behind a rejection, so those lines are gone — but an agent restart produces a burst of rejections that are not loss at all, andfor: 1his what separates the two.
Both came out of #194, which
reported the journal arriving at 1.5%. That turned out to be a measurement
artifact — the query named job="/var/log/journal" while the stream carries
job="loki.source.journal.journal", so it counted one label set and missed the
other. Compared like with like, delivery was 98.8% on the day of the report and
is 100% now. What the search did find is that Loki had been discarding around
185,000 entries a week and nothing said so, which is
#341.
That check answers whether the rules parse. It cannot answer whether they can
see, and those are different failures with the same symptom — a green run.
#261 was the second kind: five
authentication rules that were valid, loaded, evaluated, and matched nothing on
two of the four monitored hosts, because they selected log_type="authlog"
while Saruman ships a journal and morpheus ships syslog.
scripts/check_loki_coverage.py (make check-loki-coverage) asks the second
question, and it has to run against the live Loki on the monitoring host —
CI has no log store, so this one is deliberately outside make validate. For
each host-scoped rule it compares the hosts the rule's own stream selectors
reach against the hosts shipping logs at all, and for a host it cannot reach it
asks whether lines matching what the rule hunts exist there anyway. A rule blind
to a host that is producing exactly those lines fails; one blind to a host with
nothing to see warns.
There is no table of which rule should see which host, deliberately — a table
like that drifts, and a drifting table is the same class of defect. The
expectation comes out of the data instead, which is also what makes useradd
never matching on FreeBSD morpheus a non-event rather than an exception
somebody has to write down.
It runs daily at 07:45 as homelab-loki-coverage.timer, over a 24-hour window
(#335). The window is short
because it sets detection lag, not sensitivity — see
runbooks/schedule-maintenance.md for the
argument and for what to do when it exits 1.
155 rules across twelve files in prometheus/rules/:
| File | Covers |
|---|---|
host.rules.yaml |
Instance down, predictive disk fill, memory, load, clock skew and a clock with no time source at all — HostClockUnsynchronised reads node_timex_sync_status, because the offset reads zero once timesyncd has restored a clock that is wrong but stable, which is how a cleared RTC wrote fourteen minutes of samples three hours in the past and nothing noticed (#519) — reboots, and, for the two laptops, whether the shelf is on mains, whether the cell that carries them through a cut is still worth relying on (#454, and runbooks/replace-the-laptop-cell.md for the swap), how hot that cell is, whether its temperature is being measured at all — on 2026-09-19 it was not, on either host — and how many minutes the host has left once the cut arrives (#532); whether the wiki's drift check on oracle is still running (#470); and, for the NAS, whether every ZFS pool is online — ZpoolNotOnline reads node_zfs_zpool_state, added after erebor lost a disk on 2026-09-19 with the exporter hung ahead of it and InstanceDown the only page (#558, runbooks/replace-the-nas-disk.md); and whether any drive's unsafe-shutdown count has grown in a day — SmartDriveUnsafeShutdownsGrowing, whose page says when smaug recorded a clean power-off the same day, because mains removed after a clean shutdown ticks the S3520 and a clean shutdown alone does not (#746), the measure of a cut the shutdown sequence ADR-0049 decides did not catch, the rule ADR-0047 left to that issue when it carried smaug's counter under the scrape (#574); and, for Saruman, how full its LVM-thin pools are — which node_filesystem_* cannot see — and whether its Proxmox firewall is still on with an inbound policy that drops, from two agent collectors (#538, #576); and whether the tc mirror that feeds Zeek on fenrir is still delivering packets from every port of the lab bridge — ZeekMirrorInactive, Zeek's absence detection answered on the hypervisor because the lab's Loki has no ruler (#437, ADR-0068); and whether a guest the hypervisor tags disposable has outlived the fortnight it was built for — DisposableGuestOutlived, read from each guest's own qm config rather than from any list of which guests should exist, which is also why HypervisorGuestStopped stays quiet for such a guest, with GuestConfigUnreadable covering the read itself failing (#438, ADR-0071); and how many days the four management consoles' certificates have left — the pfSense GUI, the iLO, Saruman's :8006 and the TrueNAS UI, none of which the blackbox exporter can reach — ManagementCertificateExpiringSoon at 30 days and ManagementCertificateExpiryImminent at 7, read from a handshake made inside each console's own segment, with ManagementCertificateUnchecked for a console that would not answer and CertExpiryStateStale for a collector that stopped (#857, ADR-0084) |
network.rules.yaml |
SNMP reachability, pf not running, state table, switch links, iLO hardware and Smart Array cache. shiva's Smart Storage Battery read failed from 2026-08-18 until it was replaced on 2026-09-02, with the array in write-through as a result, so stored metrics before that date show the failed pack — IloBatteryCondition names the spare part to order, and the controller rollups are deliberately read at failed rather than degraded (#76). Also whether the remote path's dynamic DNS name still resolves to the WAN address — DdnsRecordStale and DdnsRecordUnchecked, from the comparison scripts/collect-gateway-state.sh makes on the firewall every fifteen minutes (#604, ADR-0044); see What the dynamic DNS series disclose below |
ups.rules.yaml |
On battery, low battery, runtime, load, temperature. A pack was fitted on 2026-08-28 and passed its self-test, so these read real hardware; stored metrics older than that date are the card's fabricated values — see runbooks/fit-the-ups-battery.md |
containers.rules.yaml |
Restart loops, OOM kills, memory, throttling |
stack.rules.yaml |
The stack watching itself: config reloads, rule evaluation, notification delivery, log ingestion, and the two cases up == 0 structurally cannot see — a remote-writing agent that stops pushing, and a scraped target that stops being a target at all. The second is ScrapeTargetDisappeared, added with the first scraped host (#256): an emptied or unparseable targets/node.yaml makes the series vanish rather than fall to 0, so InstanceDown stays silent and RemoteWriteJobStale excludes scraped jobs by design. AlloyDown (2026-10-07) is the one alert per host for the agent itself stopping. It keys on the <host>-alloy self-scrape and fires five minutes before RemoteWriteJobStale would, and an inhibit rule in alertmanager.yaml then holds back that host's per-job alerts. The target for smaug was written into targets/node.yaml disabled on 2026-09-17 and enabled on 2026-09-19, once the exporter answered from the pool. Split off containers.rules.yaml onto component: stack in #81 so a Prometheus that cannot reload its config stops being filed as a container fault. Since #575 also whether an Alertmanager silence is about to lapse or names no owning issue, read from the per-silence series scripts/collect_silences.py writes every fifteen minutes — alertmanager_silences is a count per state and cannot say which alert, when, or whose |
watchdog.rules.yaml |
One rule that always fires, so that its absence is detectable |
blackbox.rules.yaml |
Whether an endpoint can actually be reached, from outside the service, and how many days its certificate has left — the sensitive tier's seven-day ACME leaves excepted, which an hours pair watches for a stalled renewal instead — TlsAcmeRenewalLate at 48h, TlsAcmeRenewalStalled at 36h (#426) — Grafana and the wiki verified against the lab CA (#847), the APC card's self-signed one read but not trusted, Prometheus, Loki and Alertmanager over plain http (the switch UI's probes were removed on 2026-09-06, targets/blackbox.yaml), each of the household sites behind trinity's Caddy, verified against the tier's root, paging as SiteDown after ten minutes so a make up restart does not (#855), and — the other way round — that the ingest proxy on 10.0.99.20:9090 and :3100 still refuses a request with no token (IngestAuthNotEnforced, #182). The iLO and pfSense UIs are written into targets/blackbox.yaml and left disabled: each needs a firewall pass from 10.0.99.20 that is a segmentation decision, not a monitoring one (#91). Their certificates' expiry is watched without that pass, from host.rules.yaml (#857) |
dns.rules.yaml |
Whether the house is still filtering DNS, asked directly at AdGuard Home on port 53 rather than through pfSense. Since ADR-0055 AdGuard is the only forwarder, so AdGuardNotAnswering is critical at five minutes: the house cannot resolve outside names. AdGuardNotFiltering stays a warning, because a filter that fails open is a convenience lost, not an outage. The targets in targets/blackbox-dns.yaml are live since 2026-09-28, against AdGuard on trinity (#126, #404) |
backup.rules.yaml |
Whether the scheduled maintenance jobs are still being run at all — staleness, failure, never-ran, whether the age-key proof record exists to be held to its deadline, whether the CA key's offline copy has been proved lately (#496), whether the newest backup sets have been carried onto the second recipient's medium within ninety days (ADR-0048), and whether the household's copy has been carried to the holder's drive within ninety days and proved by the holder within a year (ADR-0073) |
deploy.rules.yaml |
Whether each host that pulls — prometheus and trinity (#533) — is running what the repository says: an uncommitted edit made on the host, a revision that did not verify, and how far behind main it is, and a host whose record stopped arriving. Reads the record scripts/converge.sh writes hourly (#99, ADR-0021). Also whether convergence had to start a stack service that something stopped — DeployServiceRevived — and whether one has been left stopped for two hours, held or under a backup — DeployServicesStopped (ADR-0087) |
internet.rules.yaml |
Why the internet is slow (#914): three scheduled speed tests in a row under 400 Mbit/s down or 100 up, the WAN held above 85% of what it can carry, frames arriving corrupted on the WAN port — WanReceiveErrors, the fault found the day these were written, when a cable between the XB7 and morpheus cut downloads to ~2 Mbit/s — and speedtest-tracker itself going quiet or unreachable. Also records homelab_wan_bits_per_second from pf's own em0 counters. Download thresholds are against the ~940 Mbit/s the 1000baseT WAN port can carry, not the plan's 2000. speedtest-tracker is left out of InstanceDown, so its outage is a warning and not a page |
ids.rules.yaml |
Whether Suricata is running on each interface it is declared for, read from the firewall's process table over SNMP — the fast, per-interface half; SuricataLogsStopped in loki/rules/security.rules.yaml is the slow, aggregate half (#90, #441) |
Every critical alert links to a runbook (#842).
Its runbook_url annotation is a GitHub link to a file in docs/runbooks/,
usually to the section for that alert. A single-alert ntfy page carries it on
its own line, marked 📖 (stacks/sensitive/ntfy/templates/homelab.yml). Security alerts go to
respond-to-a-security-alert.md,
alerts with a runbook of their own go there, and the rest go to
triage-a-critical-alert.md.
check_docs.py fails a critical rule with no runbook_url, a link to a file
that does not exist, or an anchor that is not a heading in that file. A new
critical rule therefore arrives with its runbook section, or CI says which is
missing.
promtool check rules validates that these parse. It does not — and cannot —
tell you whether a rule can ever be true: ContainerHighMemory passed it for
months while dividing by a memory limit no service set at the time, so it showed
as loaded and healthy and could not fire for any input (#63).
prometheus/tests/*.test.yaml holds promtool test rules unit tests, which
feed a rule synthetic series and assert it fires — paired with a case asserting
it stays quiet, because a test that only ever expects silence would have passed
against the broken rule too. Coverage is 155 rules of 155 so far (#843) — all eleven
in blackbox.rules.yaml, all three in dns.rules.yaml, GatewayFilesystemCritical, ContainerHighMemory,
ContainerNearMemoryLimit, ContainerRestartLoop, ContainerCpuThrottled,
the three container-state rules from
#838 and
PrometheusSizeRetentionActive, Watchdog, the three iLO rules from
#76, all seven in
backup.test.yaml, all eight in deploy.test.yaml, the three Loki ruler
rules from #837, the five self-monitoring
rules from #840, RemoteWriteJobStale,
ScrapeTargetDisappeared,
SuricataStopped, the two gateway rules from
#353, the two dynamic DNS rules from
#604, and all forty-one in
host.rules.yaml —
HostDiskWillFillIn24h from #189,
six more from #320, the four
SMART rules from #351,
PatchStateStopped from #360,
SystemUpdateAvailable from #378,
DriftCheckStopped from #470,
the two guest rules from #257,
the three laptop-battery rules from
#454,
HostClockUnsynchronised from #519,
the cell-temperature and runtime rules from
#532, SmartStateStale
from #483, the three ZFS leaf rules from
#744, the three thin-pool rules from
#538, the three firewall rules from
#576, the four guest-disk rules from
#778, the two Zeek mirror rules from
#437, and the two silence rules from
#575, and the last twenty from
#843: the nine UPS rules, seven
network rules, three stack rules and ContainerOomKilled.
That leaves 0 rules without a unit test. The first test of SwitchInterfaceDown
showed it had been unable to fire since it was written: it required the port's
hourly maximum to be 1 while the port read 2. scripts/check_rule_tests.py
now fails CI on any alert, in any stack, that no test selects. It checks per
rule, where the older per-stack guard only refused a stack with no tests at
all. Both numbers here are checked by scripts/check_docs.py — the sentence
they replaced claimed six and named two, and had been wrong for weeks.
ContainerCpuThrottled is the odd one in that list: it is
inert in production and cannot fire against anything cAdvisor
currently reports, because no service sets a CPU quota. Its tests are what make
the rule's correctness checkable anyway, which is the #63 lesson applied before
rather than after the fact (#185).
Saruman's four drives sit behind a P440ar Smart Array, and no single reading
covers all of them.
- The two SAS spindles come through the iLO. SNMP walks
cpqDaPhyDrvSmartStatus, andIloDrivePredictiveFailureis armed on it (#151). - The two SM863a SSDs come through
hpsa. The iLO reports every wear column blank for them, and the controller reportsSSD Smart Trip Wearout: Not Supported.collect-smart-state.shtherefore reads them directly. When a disk's SCSI host ishpsa, it probessmartctl -d cciss,Nthrough the logical drive and keeps only the drives that report no rotation. Each one gets a device label such as/dev/sda:cciss,2. ItsWear_Leveling_Countbecomeshomelab_smart_percentage_used, soSmartDriveWearHighand the otherSmartDrive*rules cover them (#529).
The spindles are left out of the second path on purpose. Reading them there too would page twice for one disk.
homelab_cert_expiry_timestamp_seconds and homelab_cert_checked carry an
endpoint name (pfsense-ui, pve-ui, ilo-ui, truenas-ui) and the
host the console belongs to, and nothing else. No address, subject or
fingerprint leaves the host that made the handshake. A console's expiry date
says when a self-signed certificate was made, which is not on
security.md's
withheld list, and the consoles themselves are already named in
network.md (#857).
homelab_ddns_record_checked and homelab_ddns_record_matches_wan are a
boolean each, labelled with host alone. The dynamic DNS name, its provider,
the address a resolver returned and the WAN address are all on
security.md's
withheld list, and none of them is in a series, a label or a log line. They do
not reach the monitoring host either: the collector sends a script to
morpheus over ssh on stdin, and the script reads the name from config.xml
and the address from the interface, asks 1.1.1.1, and prints one word —
match, stale or unchecked. What the series do tell a reader is that the
estate has a dynamic DNS record, which ADR-0044 already publishes.
unchecked is kept separate from stale on purpose. If the resolver does
not answer, that says nothing about the record, so matches_wan is left out
rather than set to 0 and only DdnsRecordUnchecked can fire. A record that is
deleted (NXDOMAIN, or an empty answer) is stale. So is a record that has one
right address and one wrong one, because a peer picks either.
Disk alerting is predictive rather than a fixed threshold — predict_linear over
a 6-hour window, firing when the extrapolation reaches zero within a day and
free space is already under 30%. A disk sitting at 86% and stable is not an
emergency; one climbing fast at 60% is.
Disk alerts on Saruman come in two kinds. Filesystems — /, /boot/efi,
/etc/pve — are node_filesystem_* from the agent, like every other host. The
guests live on LVM-thin pools, pve/data and large_data/large_data, and a
thin pool is not a filesystem, so none of the rules above can see one. They are
read by scripts/collect-thin-pool-state.sh from lvs every ten minutes:
ThinPoolNearlyFull at 80% data or 60% metadata, and ThinPoolWillFillIn24h
— critical, not a warning, because a pool that fills turns every guest on it
read-only at once — on the same six-hour extrapolation, once the pool is half
used (#538).
The ISO store is checked on Saruman, daily. Packer builds the lab's
templates from installers on smaug-iso, an NFS export that trusts an
address (ADR-0072). scripts/collect-iso-store-state.sh hashes every file on
it against a list kept in the script and writes homelab_iso_state per file
and homelab_iso_store_mounted. IsoChecksumMismatch is critical: a listed
installer whose hash changed. IsoStoreUnexpected warns on a file the list
does not name or a listed one that is gone, IsoStoreNotMounted on the share
being absent, and IsoStoreStateStale after two missed days
(#440).
Disk alerts on Saruman's guests are read through the hypervisor. The lab
guests push to the lab's Prometheus, not here (ADR-0007), so node_filesystem_*
never arrives for them. scripts/collect-guest-disk-state.sh asks each running
VM's qemu-guest-agent for get-fsinfo every ten minutes and writes
homelab_guest_filesystem_size_bytes and _used_bytes per guest and mountpoint.
GuestDiskCritical (below 10% free for 15 minutes) and GuestDiskWillFillIn24h
(the six-hour extrapolation, under 30% free) are both critical: a lab guest
is not somewhere anyone looks, and odin's root reached 98% on 2026-10-01 with
nothing to say so (#778).
ADR-0070
records why this crosses when ADR-0028 kept guest metrics in the lab, and why
the agent's answer is treated as hostile input. GuestAgentSilent warns when a
guest's agent stops answering, and GuestDiskStateStale when the collector
stops writing.
A dead lab service inside a running guest is read the same way. The lab
Prometheus sends nothing (ADR-0020), so a crashed lab Prometheus, a stopped
Zeek or a Wazuh manager with its listener down looked healthy here while the
guest ran. scripts/collect-guest-service-state.sh runs one fixed
docker inspect through each guest's agent every five minutes, for
lab-prometheus on alexander, sensor-zeek on fenrir, and
soc-wazuh-manager and soc-velociraptor on odin. It writes
homelab_guest_service_healthy, which is 1 only for running healthy, so each
container's own healthcheck is the real test. GuestServiceUnhealthy warns
after fifteen minutes. A blind SOC is serious, and not a 2 a.m. page.
GuestServiceUnchecked warns when the agent has not run the check for an hour,
and GuestServiceStateStale when the collector stops writing
(#858,
ADR-0088).
Four receivers, four separate destinations (see
alertmanager/alertmanager.yaml). Three of them were names for one webhook URL
until #66, which meant urgent
and default differed only in how often they repeated — a UPS on battery and a
slow scrape landed in the same place. The routing tree decides which alert is
urgent; only a distinct destination makes that difference audible, because
per-topic sound and do-not-disturb settings live on the receiving end.
The table below is the alert routing, and covers three of the four. The fourth,
heartbeat, carries no alerts at all — it is the dead man's switch, and it is
described in its own section below.
| Matches | Receiver | First notification | Repeats |
|---|---|---|---|
critical + category=power |
urgent |
immediately | 30m |
critical + category=security |
security |
immediately | 1h |
warning + category=security |
security |
30s | 4h |
critical (anything else) |
urgent |
10s | 4h |
warning (anything else) |
default |
30s | 12h |
info |
null |
never | — |
Security has its own destination at both severities because eleven rules carry
category: security — SSH brute force, a terminal segment reaching the internal
network, IoT lateral movement, priority-1 Suricata, Suricata not running — and routed on severity
alone, the warning-severity half of that list arrived in the default channel on
a 12-hour repeat, indistinguishable from a disk filling up. category=power
was the precedent.
First match wins and nothing sets continue, so the order of those rows is
the design. A category route moved below the bare severity rows silently
stops matching, and amtool check-config still reports SUCCESS — that mutation
was tried. scripts/validate.sh and CI therefore assert the table itself with
amtool config routes test --verify.receivers, one assertion per row.
Where the three real channels deliver changed on
#136. They now go to the
sensitive tier's own ntfy on trinity, deny-all, with Alertmanager publishing
by bearer token and verifying the leaf against the tier's root. urgent and
security also keep a second webhook to their ntfy.sh topic, because a phone
off the home network cannot reach the tier. Alertmanager counts each delivery
separately, so either half can fail without taking the other with it. If the
in-house ntfy is the thing that has failed, the page still goes out:
EndpointUnreachable (from the http_2xx_tier_ca probe of
ntfy.matrix.elysium) and AlertmanagerNotificationsFailing are both
critical, so both reach urgent and its ntfy.sh copy. The runbook has the
table and the cutover:
verify-the-alert-path.md.
Inhibit rules stop cascades: a down host suppresses its own disk warnings, a
dead snmp-exporter suppresses the "every device is unreachable" storm that
would otherwise follow, and a certificate inside seven days of expiry suppresses
its own thirty-day warning rather than resolving it — a "resolved" for a
certificate three days from expiry would be a lie.
A silence is how monitoring gets switched off, and the record of the act lives in the system being switched off (ADR-0012). Nothing read that record until #575: both silences active on 2026-09-20 had been found by a person reading the list during a triage pass, one stood over an issue closed by accident, and the other had no issue behind it until that day — the third time that shape had been found by hand (#76, #531, #572). Three times is a pattern, and the answer to a pattern is to watch the control, not to remember harder.
The convention. A silence's comment begins with the number of the open
issue that owns its expiry — #531 … — and the expiry is a date written on
that issue. The issue is where the work that ends the silence is tracked; the
comment is where a person, or a program, looking at the silence finds it.
Whatever else the comment says comes after the number, and createdBy is
free text. A silence is deleted when the work is done, not left to lapse
(runbooks/fit-the-ups-battery.md §3):
one left standing over a freshly fitted part suppresses precisely the thing
you most want to hear about, and both battery runbooks record deleting late as
their one regret.
What watches it. The silence-state job runs scripts/collect_silences.py
every fifteen minutes on the monitoring host — the one host that can reach
Alertmanager, which binds to loopback — and writes two gauges per active or
pending silence:
| Series | Value |
|---|---|
homelab_silence_expires_timestamp_seconds{alert, issue, id} |
when it ends |
homelab_silence_starts_timestamp_seconds{alert, issue, id} |
when it began, or will |
alert is the silence's alertname matcher, verbatim. It is named alert
because Prometheus writes alertname on every alert it raises, so a series
label of that name would be overwritten in the notification by the rule's own.
issue is the number from the comment, present only when the comment begins
with one — a parser that took the first #NNN anywhere would have credited
both 2026-09-20 silences to closed issues cited in their prose and hidden the
finding. id is the UUID, the handle every runbook here deletes by. A label
with nothing to say is omitted rather than written empty, which is what
Prometheus stores either way.
Two rules in prometheus/rules/stack.rules.yaml read them. SilenceWithoutIssue
fires within minutes on a silence whose comment names no owner: a silence
nobody owns is the finding. SilenceExpiresSoon gives a week's notice, and
only to a silence that has already stood a week — the silence this repository
recommends is a five-hour one
(runbooks/fit-the-saruman-ssds.md §1),
and a warning that fires the moment you create one is a warning you learn to
ignore. A thirty-day silence warns on day 23, a fourteen-day one on day 7, a
short one never. Both are warnings and route to default. The collector
stopping is ScheduledJobStale's to report, like every other job in
runbooks/schedule-maintenance.md.
An expired silence stays visible without a second window. Alertmanager
keeps one for five days after it ends, and endsAt is the difference between
deleted and lapsed — the moment of deletion, or the mark. Prometheus keeps the
series' history for thirty days, so
last_over_time(homelab_silence_expires_timestamp_seconds[7d]) matches a page
that has returned to the silence that ended. The collector emits only live
silences for that reason: emitting expired ones would add nothing the history
does not hold, and would make the expiry rule count backwards from every
lapsed silence.
To list them from the monitoring host:
curl -sS http://localhost:9093/api/v2/silences |
python3 -c 'import json,sys; [print(s["id"], s["status"]["state"], s["endsAt"], s["comment"][:60]) for s in json.load(sys.stdin)]'AlertmanagerNotificationsFailing catches delivery errors. It cannot catch a
webhook URL that is well-formed, reachable, and pointed at nothing — a 200 into a
deleted ntfy topic is a successful notification by every measure Alertmanager
has. That is not hypothetical: the webhook was the ntfy.example.invalid
placeholder for the entire life of the stack and nothing noticed, because the
only symptom is that alerts stop arriving, which is also what a healthy week
looks like (#67).
prometheus/rules/watchdog.rules.yaml holds one rule, Watchdog, whose
expression is vector(1). It fires unconditionally and forever. Its firing
carries no information; its absence is the entire signal. One continue: true
— the only one in the tree — sends it to two places:
| Route | Destination | Cadence | Catches |
|---|---|---|---|
heartbeat |
external cron-monitor ping | 5m | Prometheus stopped evaluating, Alertmanager died, no outbound network |
default |
the real alert channel | 24h | the alert channel itself is a 200 into nothing |
Neither half substitutes for the other. The heartbeat proves delivery to a different URL than real alerts use, so it cannot see a deleted topic; the daily notification travels the identical URL your warnings travel, but nothing machine-checks its absence.
The Loki ruler has its own Watchdog, and Prometheus watches it
(#837). Every security alert is
evaluated by Loki's ruler, not Prometheus, and the table above proves only
Prometheus's path. loki/rules/watchdog.rules.yaml holds LokiRulerWatchdog,
also vector(1) and firing forever. Alertmanager routes it to null, because
its delivery is not the point: its sending is. The ruler re-sends a firing
alert every minute, so loki_prometheus_notifications_sent_total climbs while
the ruler is evaluating and has an Alertmanager to reach. Three Prometheus
rules read the ruler from there:
| Rule | Fires when |
|---|---|
LokiRuleEvaluationFailures |
a rule group fails to evaluate (warning) |
LokiRulerNotificationsFailing |
sends error or are dropped (critical) |
LokiRulerSilent |
nothing is sent for a 15-minute window, held for 5 more, so it pages at about 20 minutes; or the counter is absent (critical) |
Those are Prometheus rules, so the heartbeat above already proves the path that would report them. One external check covers both evaluators.
Why not a second external heartbeat for the ruler. It was the first
design in #837. It would prove the ruler's path without depending on
Prometheus scraping Loki, but it needs a second check at the external service
and a second secret URL. As built, a dead Loki or a failed scrape is still
reported, by InstanceDown for the loki job and by LokiRulerSilent's
absent(). So the only case a dedicated heartbeat would add is Prometheus
and the ruler failing at the same moment, and the Prometheus heartbeat already
pages for that. If that trade is ever wrong, the change is a continue: true
route from LokiRulerWatchdog to a second heartbeat receiver.
The heartbeat half became a dead man's switch on 2026-09-09. Until then all
four receivers pointed at ntfy.sh, the heartbeat included, and ntfy is a push
service: it delivers what it is sent and has no notion of an expected interval,
so it cannot notice a ping that never arrived — and absence is the entire
signal. The pings were being delivered to a topic nobody was waiting on
(#359), and they were also
spending almost all of ntfy.sh's free daily budget, so real alerts were refused
at the end of every day (#407).
The heartbeat now pings a healthchecks.io check, period 5m and grace 15m, which
emails when a ping does not arrive; the three real channels stay on ntfy and
have the budget to themselves. (Since #136 only two of them use ntfy.sh at all,
as the off-network copy of what the in-house ntfy receives; see Routing.)
So now, if Prometheus stops evaluating, Alertmanager dies, or this host loses outbound network, something external notices — in principle. That is the failure #214 lived through from the other direction, and the heartbeat is cited as the answer to it in #214's own resolution. It is armed and proven: on 2026-09-09 Alertmanager was stopped for 18 minutes, the check went red and emailed, and the first ping after the restart landed within two minutes; the daily route was confirmed the same sitting (#288, times in the runbook).
check_alert_channels.py --live reports the destination on every deploy,
classifying the heartbeat's host as a watcher or a push service. A push service
is a warning rather than a failure on purpose: it cannot be fixed from this
repository — it needs an account on a watcher service and a decision about where
its notification goes — and a deploy-time check that is permanently red for a
known reason stops being read, which this repository has already written down
about .gitleaksignore. It read ntfy.sh and warned from 2026-09-07 to
2026-09-09; it reads hc-ping.com and passes since. That was the four steps in
runbooks/verify-the-alert-path.md, and
#288's drill is runnable now.
The watcher lives off this host by necessity — a watcher here fails at the same
moment as the thing it is watching. Setting it up, the coupling between
repeat_interval and the external check's period and grace, and how to read
which half went quiet are in
runbooks/verify-the-alert-path.md.
This is the same reasoning loki/rules/security.rules.yaml already applies to
the firewall with FirewallLogsStopped, and the reason SuricataLogsStopped
waits nine hours where the firewall's rule waits thirty minutes: absence of
alerts is indistinguishable from absence of the service for an hour at a time,
so the fast answer is a heartbeat — which prometheus/rules/ids.rules.yaml
reads from the firewall's process table — and the log rule is the slow one. The
notification path was the one place that argument had not been turned on
itself.
The same inversion, applied to maintenance. make backup,
make backup-firewall, make snmp-verify and make secrets-verify-backup were
all commands someone had to remember, and nothing ran any of them
(#77). A job that fails is loud;
a job that stops being run is silent, and silence is also what a healthy week
looks like.
Four of them are now systemd timers on the monitoring host. Every run — timer or
human — goes through scripts/run-scheduled.sh, which records four gauges into
a textfile the node exporter already scrapes:
| Metric | Says |
|---|---|
homelab_job_last_success_timestamp_seconds |
when this job last exited 0. A failed run carries the previous value forward rather than clobbering it |
homelab_job_last_run_timestamp_seconds |
when it last finished, whatever the outcome |
homelab_job_last_exit_code |
0, or 75 for "never started, another job held the lock" |
homelab_job_duration_seconds |
how long it took |
The threshold each job is held to is a fifth series,
homelab_job_max_age_seconds, written by scripts/install-timers.sh from the
same table that decides the cadence. That is what lets the rules in
prometheus/rules/backup.rules.yaml cover every job without naming any of them
— the three that do name one concern the two human proofs, verify-key-backup
and verify-ca-key-backup, below — and what makes make check-timers able to assert that a threshold is
at least twice its timer's real period.
Two things are deliberate and easy to undo by accident:
- The label is
homelab_job, notjob.config.alloy'sdiscovery.relabel "metrics"setsjobon every target from that exporter, and scrapes default tohonor_labels: false— so ajoblabel in the file would arrive asexported_joband every rule would match nothing while still showing as loaded and healthy.backup.test.yamlhas a case that pins this. - The staleness rules compare against a stored timestamp rather than using a
long
for:. A longformeasures continuous pending time in Prometheus's own memory, and one of the jobs being measured is the one that stops Prometheus.UpsBatteryUnprovenrecords the same reasoning.
What this does not prove. Every one of these jobs runs on the machine it is
checking, with the key that is on that machine, against the disk that is in it.
verify-backups proves an archive still decrypts; it says nothing about a dead
disk or a fire. The jobs that prove anything beyond this shelf are the ones
that cannot be automated, because each needs a human to mount removable media:
secrets-verify-backup for the key, and since
ADR-0048 backup-offsite
for the sets themselves, which carries the newest of each kind onto the
second recipient's medium and hashes it there. SecretsKeyBackupUnproven and
OffsiteCopyStale nag at ninety days instead. On trinity, household-copy
and household-proof are the household's equivalents
(ADR-0073),
and HouseholdCopyStale nags at ninety days and at a year. SecretsKeyBackupUnproven is the one rule in backup.rules.yaml not keyed
on homelab_job:
ADR-0024
allows the secrets to be encrypted to more than one age recipient, so it fires
per recipient off homelab_key_recipient_last_proof_timestamp_seconds rather
than off the job. One timestamp for every copy would mean proving either one
vouched for the other, which is backwards when the whole point of the second
copy is that it fails independently. With a single recipient it behaves exactly
as it always has. Three outputs leave: backup-firewall copies each export
to oracle and fails if it cannot, so its failure alert doubles as "the config
has stopped leaving this host"; since
#535 backup-volumes does the
same with each weekly set; and since
#484 backup-nas — the one
job that first fetches from another host, the media tier's state off smaug —
copies its set the same way. verify-backups hashes the far side of both
set directories every morning. No rule names any of the three; the generic
pair carries them all.
That series has to exist for the nag to mean anything, and for four days it did
not (#400): it was written only
by a proof run or by adding a recipient, and this host had proved its key before
the series was invented, so the rule went quiet the day it was deployed while
the fallback it named — ScheduledJobNeverRan — was satisfied by the old proof.
The recipient-state timer now writes the recipient list daily, carrying proofs
forward and setting none, and SecretsKeyRecipientsUnrecorded fires when the
deadline row exists and the recipient series does not — an unless against the
declaration row, not an absent(), so it carries labels like every other rule
in the file.
Installing, tuning and troubleshooting all of it:
runbooks/schedule-maintenance.md.
See runbooks/add-monitored-device.md. In
short:
- A Linux host: run Alloy with
LOKI_URLandPROMETHEUS_REMOTE_WRITE_URLpointed athttps://10.0.99.20, andINGEST_CA_FILEat the estate CA;deploy-agent.shsets all three. The monitoring host needs the host's token added (the runbook). - A Linux host that may not push: a firewall pass first, then
node_exporterin that host's own compose stack, then a target inprometheus/targets/node.yamlwithinstanceset to the hostname. The direction reverses when the segment demands it, and the tool reverses with it —smaugis the only one today. - An SNMP device: append a target to
prometheus/targets/snmp.yamland a module plus auth tosnmp-exporter/generator.yaml. file_sd picks the target up within five minutes without a restart.
The stack is small, but two things will bite if ignored:
- The
iloSNMP module exposes ~1,600 metrics. The HP Insight tree is enormous. It is scraped once a minute from one device, which is fine — but do not add a second module that broad without trimming the OID list ingenerator.yaml. All four modules together are ~1,800 metrics per scrape cycle. - Loki labels must stay low-cardinality.
host,level,log_type,service_nameandunitare bounded. Never promote a request ID, IP address or timestamp to a label; use|=line filters instead. A MAC address is the same class of mistake, which is why the ADR-0019 rules extractmacwith a query-time| regexpand it exists nowhere in the index.
module and auth are deliberately dropped by labeldrop in prometheus.yaml
after being converted to query parameters, so they never become metric labels.
make up # render secrets, start everything
make ps # container status
make logs SERVICE=grafana # tail one service
make reload # hot-reload Prometheus, Alertmanager, snmp-exporter
make validate # everything CI runs
make backup # quiesce, archive, encrypt and verify the data volumes
make restore ARGS=--list # the backup sets that exist
make down # stop, keep data
make nuke # stop, destroy data (prompts) — recoverable, see
# docs/runbooks/restore-the-stack.mdPrometheus and Alertmanager are started with lifecycle endpoints enabled, so
rule and route changes apply via make reload without dropping the TSDB head
block. snmp-exporter serves POST /-/reload unconditionally, with no lifecycle
flag to enable.
make up runs that same reload as its last step. It has to: docker compose up -d keys off the service definition, not the contents of the files it mounts,
so without the reload a re-rendered config would sit on disk while the
container served the copy it parsed at startup.