Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion docs/architecture.md
Original file line number Diff line number Diff line change
Expand Up @@ -191,7 +191,7 @@ hole from the monitoring VLAN into the monitored one.
| `phoenix` (10.0.30.70) | 🟢 30 | *(none — a toolchain host, no Docker)* | **Not built yet** — the deployment host: a Proxmox API token, an SSH key and a checkout, so that the Packer, OpenTofu and Ansible work after [#436](https://github.com/Gerrrt/HomeLab/issues/436) has somewhere to run from. A guest on `Saruman`, placed by [ADR-0043](adr/0043-keep-the-ca-on-prometheus-and-build-phoenix-as-the-deployment-host.md), which also decides that the estate's CA key stays on `prometheus` and does not follow the toolchain here. It holds no age key and converges nothing; it is the WireGuard endpoint of [ADR-0042](adr/0042-terminate-the-remote-path-on-the-lab-and-route-it.md) once that is built; the one thing it reaches that no other guest does is `8006` on `Saruman`, by a single host-firewall rule. Its Alloy pushes to `alexander`, never to VLAN 99; the build is [`build-the-jumpbox.md`](runbooks/build-the-jumpbox.md). Alloy agent (native package, `scripts/deploy-agent.sh`) |
| `oracle` (10.0.99.30) | 🔴 99 | *(none — hand-run containers)* | The Lemmiwinks wiki and its Postgres, since 2025-11-12 ([ADR-0011](adr/0011-keep-the-wiki-internal.md)); Alloy agent (Docker, `scripts/deploy-agent.sh`); the off-host copy of the firewall export (`make backup-firewall`). The estate's host for small off-host jobs — [ADR-0015](adr/0015-give-oracle-the-off-host-jobs.md) |
| `trinity` (10.0.99.40) | 🔴 99 | [`stacks/sensitive`](../stacks/sensitive) | **Not built yet** — ADR-0008's sensitive tier on the ProDesk 600 G4 of [ADR-0034](adr/0034-run-the-sensitive-tier-on-the-prodesk-and-make-it-the-spare-hardware.md), after the firewall restore is rehearsed on it ([#404](https://github.com/Gerrrt/HomeLab/issues/404)). The foundation is authored: Caddy as the published HTTPS port and step-ca issuing beneath the tier's own root rather than the estate's, which is left untouched ([#129](https://github.com/Gerrrt/HomeLab/issues/129), [#130](https://github.com/Gerrrt/HomeLab/issues/130), [ADR-0037](adr/0037-give-the-sensitive-tier-its-own-root-and-issue-beneath-it-over-acme.md)), with AdGuard Home behind Caddy and publishing 53 to the firewall's forwarder alone ([#135](https://github.com/Gerrrt/HomeLab/issues/135), [ADR-0010](adr/0010-keep-the-resolver-on-the-gateway.md)); Home Assistant ([#134](https://github.com/Gerrrt/HomeLab/issues/134)), Immich — four containers behind Caddy with a memory limit on each ([#132](https://github.com/Gerrrt/HomeLab/issues/132)) — Paperless-ngx with a Postgres and a Valkey of its own ([#133](https://github.com/Gerrrt/HomeLab/issues/133)) and Vaultwarden ([#131](https://github.com/Gerrrt/HomeLab/issues/131)) are authored as well; the rest follow. Alloy agent (Docker, `scripts/deploy-agent.sh`), pushing to `prometheus` like `oracle`'s |
| `smaug` (10.0.40.30) | 🟡 40 | [`stacks/media`](../stacks/media) | [ADR-0008](adr/0008-place-services-by-data-trust.md)'s media tier on the ThinkServer TS150 of [#413](https://github.com/Gerrrt/HomeLab/issues/413), placed and addressed by [ADR-0016](adr/0016-open-casabonita-inward-and-keep-it-terminal-outward.md) and running TrueNAS rather than Ubuntu Server by [ADR-0040](adr/0040-run-truenas-on-smaug-and-keep-the-media-stack-in-this-repository.md). **The host is built and the pool is not**: TrueNAS is installed on its boot SSD, it holds the static above since 2026-09-16, and the four inbound rules are created and verified in position — but the ZFS mirror waits on two Exos X20 drives and nothing is deployed on it. The stack is authored and CI-validated ahead of the storage, the way `stacks/sensitive` was ahead of `trinity`: Jellyfin alone, publishing 8096 to the segment because the televisions reach it natively and no firewall rule is involved at all ([#138](https://github.com/Gerrrt/HomeLab/issues/138)). Scraped by `prometheus` on `9100`; it pushes nothing, and runs no Alloy — the estate's first scraped host, and the reason [#256](https://github.com/Gerrrt/HomeLab/issues/256) was more than a line of YAML. That issue settled the fork TrueNAS opened in it: `node_exporter`, as a digest-pinned container in this stack rather than TrueNAS's own endpoint, so the existing `99 → 40:9100` pass, the `host-overview` dashboard and seven rules in `host.rules.yaml` all keep working unchanged. The `node` job and `prometheus/targets/node.yaml` are live; the target itself stays commented until there is a pool to run the exporter from |
| `smaug` (10.0.40.30) | 🟡 40 | [`stacks/media`](../stacks/media) | [ADR-0008](adr/0008-place-services-by-data-trust.md)'s media tier on the ThinkServer TS150 of [#413](https://github.com/Gerrrt/HomeLab/issues/413), placed and addressed by [ADR-0016](adr/0016-open-casabonita-inward-and-keep-it-terminal-outward.md) and running TrueNAS rather than Ubuntu Server by [ADR-0040](adr/0040-run-truenas-on-smaug-and-keep-the-media-stack-in-this-repository.md). **Built, pooled and deployed**: TrueNAS on its boot SSD, the static above since 2026-09-16, the four inbound rules verified in position, the mirror `erebor` since 2026-09-18 and this stack running on it since 2026-09-19 — from a copy of the compose file on the pool, brought up with `docker compose` under TrueNAS's own Docker ([`build-the-nas.md`](runbooks/build-the-nas.md) §6). Jellyfin alone, publishing 8096 to the segment because the televisions reach it natively and no firewall rule is involved at all ([#138](https://github.com/Gerrrt/HomeLab/issues/138)). Scraped by `prometheus` on `9100`; it pushes nothing, and runs no Alloy — the estate's first scraped host, and the reason [#256](https://github.com/Gerrrt/HomeLab/issues/256) was more than a line of YAML. That issue settled the fork TrueNAS opened in it: `node_exporter`, as a digest-pinned container in this stack rather than TrueNAS's own endpoint, so the existing `99 → 40:9100` pass, the `host-overview` dashboard and seven rules in `host.rules.yaml` all keep working unchanged. The `node` job and `prometheus/targets/node.yaml` are live, and the target with them since 2026-09-19, once the exporter answered from the pool |
| `bahamut` (10.0.30.50) | 🟢 30 | *(none — Windows)* | **Not built yet** — Windows Server 2025 domain controller, PDC emulator and DNS for `ad.matrix.elysium` — Tier 0. Static, because every member finds a DC through DNS and the DCs *are* the DNS. Scraped by `alexander` on `9182`; it pushes nothing, and runs no Alloy ([ADR-0029](adr/0029-size-the-lab-domain-and-separate-its-namespace-and-clock.md)) |
| `leviathan` (10.0.30.51) | 🟢 30 | *(none — Windows)* | **Not built yet** — Windows Server 2025 second domain controller and DNS — Tier 0. Static, for the same reason. Scraped by `alexander` on `9182`; it pushes nothing, and runs no Alloy ([ADR-0029](adr/0029-size-the-lab-domain-and-separate-its-namespace-and-clock.md)) |
| `titan` (10.0.30.52) | 🟢 30 | *(none — Windows)* | **Not built yet** — Windows Server 2025 file and member server — the shares, and the NTLM relay target that only exists because 2025 requires outbound SMB signing and not inbound — Tier 1. Scraped by `alexander` on `9182`; it pushes nothing, and runs no Alloy ([ADR-0029](adr/0029-size-the-lab-domain-and-separate-its-namespace-and-clock.md)) |
Expand Down
43 changes: 31 additions & 12 deletions docs/hardware.md
Original file line number Diff line number Diff line change
Expand Up @@ -43,7 +43,7 @@ quietly swapped.
| `Saruman` | HPE ProLiant DL360 Gen9 | 2× Xeon E5-2680 v3 (48 threads) | 128 GB | 2× 1 TB SAS HDD, RAID 1; 2× 960 GB SATA SSD, unassigned | Proxmox VE 9 |
| `prometheus` | Apple MacBook Pro (2012, Retina 13") | i5/i7 | 8 GB | 256 GB SSD | Ubuntu Server 24.04 LTS |
| `oracle` | Dell Inspiron 15-3565 | AMD A6-9200 (2 cores) | 4 GB | 500 GB HDD | Ubuntu Server 24.04 LTS |
| `smaug` | Lenovo ThinkServer TS150 | Xeon E3-1225 v6 (4 cores) | 8 GB ECC | 240 GB SATA SSD (boot) | TrueNAS 25.10 |
| `smaug` | Lenovo ThinkServer TS150 | Xeon E3-1225 v6 (4 cores) | 8 GB ECC | 240 GB SATA SSD (boot) + 2× 18 TB ZFS mirror `erebor` | TrueNAS 25.10 |

The observability stack runs on a thirteen-year-old MacBook. It handles four
SNMP devices at a 60-second interval, four Alloy agents, and 30 days of metric
Expand Down Expand Up @@ -194,10 +194,11 @@ revisions of this repository treated `shiva` as the hypervisor itself.
has created would be the kind of claim this table exists to not make. The boot disk the TrueNAS install wants
([ADR-0040](adr/0040-run-truenas-on-smaug-and-keep-the-media-stack-in-this-repository.md))
and the bracket that carries it in the optical bay are the entries below and
landed with it; the two drives for the mirror, bought 2026-09-11, have not. Nothing
landed with it; the two drives for the mirror, bought 2026-09-11, landed on
2026-09-18 and were in the bays that evening. Nothing
for this machine is outstanding on the roadmap's
[list](roadmap.md#everything-still-to-buy) any more; what is left is those
two drives landing, and a build.
[list](roadmap.md#everything-still-to-buy) any more, and the Storage column
reads the mirror since 2026-09-19.
**Read off the machine on 2026-09-15**, where everything above it came off a
listing: `ThinkServer TS150`, machine type-model `70UB000AUX`, serial
`MJ05N4NK`. Xeon E3-1225 v6 at 3.30 GHz, four cores, and `Active Video: IGD`
Expand Down Expand Up @@ -248,7 +249,8 @@ revisions of this repository treated `shiva` as the hypervisor itself.
access ZFS wants and the thing
[#418](https://github.com/Gerrrt/HomeLab/issues/418) is the cautionary tale
for, and `CSM [Disabled]`, so it boots UEFI as TrueNAS wants.
Two 3.5" trays, both empty — exactly the mirror and no spare.
Two 3.5" trays, filled by the Exos pair on 2026-09-18 — exactly the mirror
and no spare.
**The 5.25" bay was not empty**: a PLDS `DVD-RW DU8AESH` answered on SATA5.
A photograph of the open case had been read here as an empty cage and was
wrong; the BIOS summary is what caught it. The optical drive came out on
Expand All @@ -257,15 +259,32 @@ revisions of this repository treated `shiva` as the hypervisor itself.
that fan is **not optional**: it is the airflow over the drive bays, and two
7200 rpm Exos under a scrub will want it. Reconnected after the swap and
reading `Aux Fan: Operating`.
- 2× Seagate Exos X20 18 TB (`ST18000NM003D`), 3.5" SATA — bought
2026-09-11, in transit. `smaug`'s ZFS mirror
- 2× Seagate Exos X20 18 TB (`ST18000NM003D`, firmware `SN03`), 3.5" SATA —
bought 2026-09-11, **in hand since 2026-09-18**, a day after the carrier's
window lapsed. `smaug`'s ZFS mirror
([ADR-0016](adr/0016-open-casabonita-inward-and-keep-it-terminal-outward.md),
[#413](https://github.com/Gerrrt/HomeLab/issues/413)). A mirror of two is
one drive's capacity, so this is 18 TB usable, not 36. The listing's
**zero power-on hours is a claim, not a fact** — read it back with
`smartctl -a` on arrival and record what the drives actually report here,
before the mirror is built on them. Serials go here when they land. They
enter the Compute table with the NAS, which is not built.
one drive's capacity, so this is 18 TB usable, not 36.
**Read off both drives with `smartctl -a` at the console on 2026-09-18,
before the pool existed**: serials `ZVTBS4NL` and `ZVTBSDL3`, both `PASSED`,
both at **0 power-on hours**, 0 reallocated, 0 pending, 0 uncorrectable, no
errors logged, 26 °C in the bays against a lifetime range of 24–26. **The
zero hours is a fact, not a claim, and it was checked the one way that
settles it.** SMART's hours counter can be reset, and used Exos drives with
it zeroed were sold as new through 2025; Seagate's FARM log keeps a second
counter that the reset does not touch, and `smartctl -l farm` on this host
read **0 power-on hours and 0 spindle hours on both drives**. What the
resettable counters add is consistent with new drives bench-checked by the
seller and with nothing more: `ZVTBSDL3` arrived carrying a Windows quick
format — a 16 MB Microsoft reserved partition and an NTFS volume labelled
`New Volume` filling the rest — with 718,258 LBAs written, about 370 MB
and the size of that format, three power cycles, and one short self-test
logged at lifetime hour 0; `ZVTBS4NL` arrived blank, two power cycles,
nothing written. Both spun up and enumerated on the first power-up. **The
TS150's SATA power lead has four wires and no orange one**, read on
2026-09-18, so this supply puts nothing on pin 3 and the Power Disable trap
`build-the-nas.md` §1 records does not apply on this box — it would on a
supply that does. They entered the Compute table with the pool.
- Intel DC S3520 240 GB, 2.5" SATA 6 Gb/s enterprise SSD with power-loss
protection — bought 2026-09-11, **in hand since 2026-09-15**. `smaug`'s boot
disk, carrying TrueNAS and the media stack it launches
Expand Down
6 changes: 3 additions & 3 deletions docs/network.md
Original file line number Diff line number Diff line change
Expand Up @@ -299,9 +299,9 @@ Televisions and consoles. Internet only.
[ADR-0016](adr/0016-open-casabonita-inward-and-keep-it-terminal-outward.md)
and running TrueNAS by
[ADR-0040](adr/0040-run-truenas-on-smaug-and-keep-the-media-stack-in-this-repository.md)
([#413](https://github.com/Gerrrt/HomeLab/issues/413)). Its ZFS mirror does
not exist yet, and `stacks/media` is authored and undeployed against that
absence; what exists is a host on its address. **It does not change the
([#413](https://github.com/Gerrrt/HomeLab/issues/413)). Its ZFS mirror
`erebor` exists since 2026-09-18 and `stacks/media` runs on it since
2026-09-19, publishing `8096` to this segment. **It does not change the
*Reaches* column**, and that is the point ADR-0016 made in advance: nothing
on this segment initiates anywhere, and the rules created that day all let a
more trusted segment reach **in**. That is the direction this row records,
Expand Down
2 changes: 1 addition & 1 deletion docs/observability.md
Original file line number Diff line number Diff line change
Expand Up @@ -453,7 +453,7 @@ argument and for what to do when it exits 1.
| `network.rules.yaml` | SNMP reachability, pf not running, state table, switch links, iLO hardware and Smart Array cache. `shiva`'s Smart Storage Battery read failed from 2026-08-18 until it was replaced on 2026-09-02, with the array in write-through as a result, so stored metrics before that date show the failed pack — `IloBatteryCondition` names the spare part to order, and the controller rollups are deliberately read at *failed* rather than *degraded* ([#76](https://github.com/Gerrrt/HomeLab/issues/76)) |
| `ups.rules.yaml` | On battery, low battery, runtime, load, temperature. A pack was fitted on 2026-08-28 and passed its self-test, so these read real hardware; stored metrics older than that date are the card's fabricated values — see [`runbooks/fit-the-ups-battery.md`](runbooks/fit-the-ups-battery.md) |
| `containers.rules.yaml` | Restart loops, OOM kills, memory, throttling |
| `stack.rules.yaml` | The stack watching itself: config reloads, rule evaluation, notification delivery, log ingestion, and the two cases `up == 0` structurally cannot see — a remote-writing agent that stops pushing, and a scraped target that stops being a target at all. The second is `ScrapeTargetDisappeared`, added with the first scraped host ([#256](https://github.com/Gerrrt/HomeLab/issues/256)): an emptied or unparseable `targets/node.yaml` makes the series vanish rather than fall to 0, so `InstanceDown` stays silent and `RemoteWriteJobStale` excludes scraped jobs by design. The target for `smaug` is written into `targets/node.yaml` and left disabled until the NAS has a pool to run the exporter from. Split off `containers.rules.yaml` onto `component: stack` in [#81](https://github.com/Gerrrt/HomeLab/issues/81) so a Prometheus that cannot reload its config stops being filed as a container fault |
| `stack.rules.yaml` | The stack watching itself: config reloads, rule evaluation, notification delivery, log ingestion, and the two cases `up == 0` structurally cannot see — a remote-writing agent that stops pushing, and a scraped target that stops being a target at all. The second is `ScrapeTargetDisappeared`, added with the first scraped host ([#256](https://github.com/Gerrrt/HomeLab/issues/256)): an emptied or unparseable `targets/node.yaml` makes the series vanish rather than fall to 0, so `InstanceDown` stays silent and `RemoteWriteJobStale` excludes scraped jobs by design. The target for `smaug` was written into `targets/node.yaml` disabled on 2026-09-17 and enabled on 2026-09-19, once the exporter answered from the pool. Split off `containers.rules.yaml` onto `component: stack` in [#81](https://github.com/Gerrrt/HomeLab/issues/81) so a Prometheus that cannot reload its config stops being filed as a container fault |
| `watchdog.rules.yaml` | One rule that always fires, so that its absence is detectable |
| `blackbox.rules.yaml` | Whether an endpoint can actually be reached, from outside the service, and how many days its certificate has left — Grafana verified against the lab CA, the APC card's self-signed one read but not trusted, the wiki, Prometheus, Loki, Alertmanager and the switch UI over plain http. The iLO and pfSense UIs are written into `targets/blackbox.yaml` and left disabled: each needs a firewall pass from `10.0.99.20` that is a segmentation decision, not a monitoring one ([#91](https://github.com/Gerrrt/HomeLab/issues/91)) |
| `dns.rules.yaml` | Whether the house is still filtering DNS, asked directly at AdGuard Home on port 53 rather than through pfSense — a probe sent down the normal resolver path always passes, because Unbound's fallback is doing its job. [ADR-0010](adr/0010-keep-the-resolver-on-the-gateway.md) made losing the filter silent on purpose, and these two rules are what distinguishes "this site was never on a list" from "AdGuard has been dead for three weeks". Warning, not critical: nothing is down and nobody is blocked. The targets are written into `targets/blackbox-dns.yaml` and left disabled until [#102](https://github.com/Gerrrt/HomeLab/issues/102) builds the mini PC ([#126](https://github.com/Gerrrt/HomeLab/issues/126)) |
Expand Down
12 changes: 7 additions & 5 deletions docs/roadmap.md
Original file line number Diff line number Diff line change
Expand Up @@ -799,11 +799,13 @@ what left this one unfireable for months.
the BIOS flashed, AMT found on Intel's factory-default credential and
unprovisioned, the optical drive swapped for the boot SSD and its SMART read
**before** the install, TrueNAS 25.10 installed, the static and the Kea
reservation both set, and the inbound rules created and verified. What is
left is the two Exos drives and everything downstream of them: the mirror
`erebor`, its two datasets, the household share, the stack, and the one test
that decides whether the stack stays here at all — whether Quick Sync reaches
a container, which is ADR-0040's reopen condition and is still unrun.
reservation both set, and the inbound rules created and verified. **The
drives landed 2026-09-18**, both at zero hours by the FARM log and not only
by SMART, and the rest followed: the mirror `erebor`, its two datasets, the
household share `media`, and the stack running under TrueNAS's Docker with
the scrape live on 2026-09-19. Of ADR-0040's reopen condition, the render
node reaches the container and the process carries the render group; the
transcode itself is the half still unrun.
Reading the enforced ruleset first changed two of the answers, and both were
borne out when the rules were created. **50→40 is not simply a rule to add**:
Hicks and Winterfell each carry an explicit *Block access to
Expand Down
Loading
Loading