You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
At 20:55 PDT on 2026-09-19 (03:55 UTC 2026-09-20) TrueNAS on smaug raised:
Pool erebor state is ONLINE: One or more devices are faulted in response to persistent errors. Sufficient replicas exist for the pool to continue functioning in a degraded state.
Disk ST18000NM003D-3DL103 ZVTBSDL3 is FAULTED
ZVTBSDL3 is sdb, one half of the two-disk Exos X20 mirror build-the-nas.md §3
built on 2026-09-18. It was at about lifetime hour 27, one day after arrival, and its
extended self-test had completed without error at hour 26 that same evening
(hardware.md, Exos entry).
The TS150 has two trays and no spare, so the pool runs on ZVTBS4NL alone until a
replacement is fitted.
What the monitoring host saw, read the same night:
When (PDT, 2026-09-19)
What
20:46
up{job="node",instance="smaug"} went 1 → 0 and stayed there. TCP 9100 still accepts a connection; GET /metrics never returns (curl gives up at 30 s, exit 28). A collector is blocked on the faulted device.
20:50
InstanceDown firing, severity=critical, routed to urgent. Still firing.
20:55
The TrueNAS alert above.
Three things this exposed
No rule reads node_zfs_zpool_state.smaug has exported it since the scrape went live (feat(nas): the drives are in, erebor exists, the stack runs and the scrape is on #522) — one 0/1 series per pool per state, for erebor and boot-pool — and nothing in stacks/observability/prometheus/rules/ reads any node_zfs_* series. The only page was InstanceDown, and only because the exporter hung. The rule that should have fired is in the PR on this issue; it can only catch the case where the exporter survives the fault, and that limit is written into it.
No runbook covered a NAS disk swap, and build-the-nas.md §7 and hardware.md both still read that the pool is ONLINE with no errors.
The clock
The pair is eBay item 237056026029, delivered 2026-09-18. The eBay Money Back Guarantee runs
30 days from delivery, so a "not as described / defective" return must be opened by 2026-10-18. Whether Seagate's own warranty applies to an "0HR" lot is unknown until the
serial is checked.
What has to happen, in this order
Triage at the console: zpool status -v erebor, dmesg for sdb, smartctl -a, -l selftest and -l farm on sdb; readings into the runbook's status block; media-node-exporter restarted, or the hang recorded if it does not come back
Return opened under the eBay guarantee before 2026-10-18; Seagate warranty check by serial recorded
Replacement decided: the seller's replacement, or a purchase and a row in the roadmap's buy list
Disk offlined, swapped, replaced in the UI; resilver complete; scrub clean
hardware.md carries the new serial and its §2 readings; build-the-nas.md §7's zpool status line is true again
The faulted drive wiped, or the decision recorded, before it ships back
homelab-backup-nas.timer installed (§6.2 step 7), so the copy stops being a one-off — done 2026-09-20, next run Sat 2026-09-26 03:32 UTC
Runbook: docs/runbooks/replace-the-nas-disk.md. Refs #483 (why nothing on smaug can
push more than node_exporter offers), #413 (the build this happened one day after).
Corrected 2026-09-20
Point 2 under Three things this exposed is answered: erebor/apps has an off-host copy as of 20260920T060234Z, frodo and SSH exist on smaug, §6.2 reads Done 2026-09-20, and the timer is installed — #484 closed on it. Two boxes above are ticked with the evidence. The tray can be touched.
Milestone NAS, seq/1: this swap is now what that milestone closes on, beside #523. The vdev-level signal (a faulted leaf under an ONLINE pool) stays an open box here; #483, which owns how anything host-level reaches smaug, is re-sequenced to seq/1 in its own milestone so the two are worked together rather than decided twice.
What happened
At 20:55 PDT on 2026-09-19 (03:55 UTC 2026-09-20) TrueNAS on
smaugraised:ZVTBSDL3issdb, one half of the two-disk Exos X20 mirrorbuild-the-nas.md§3built on 2026-09-18. It was at about lifetime hour 27, one day after arrival, and its
extended self-test had completed without error at hour 26 that same evening
(
hardware.md, Exos entry).The TS150 has two trays and no spare, so the pool runs on
ZVTBS4NLalone until areplacement is fitted.
What the monitoring host saw, read the same night:
up{job="node",instance="smaug"}went 1 → 0 and stayed there. TCP 9100 still accepts a connection;GET /metricsnever returns (curl gives up at 30 s, exit 28). A collector is blocked on the faulted device.InstanceDownfiring,severity=critical, routed tourgent. Still firing.Three things this exposed
node_zfs_zpool_state.smaughas exported it since the scrape went live (feat(nas): the drives are in, erebor exists, the stack runs and the scrape is on #522) — one 0/1 series per pool per state, forereborandboot-pool— and nothing instacks/observability/prometheus/rules/reads anynode_zfs_*series. The only page wasInstanceDown, and only because the exporter hung. The rule that should have fired is in the PR on this issue; it can only catch the case where the exporter survives the fault, and that limit is written into it.erebor/appshas zero off-host copies. feat(backup): pull Jellyfin's state off smaug from a ZFS snapshot, encrypt it here and copy it to oracle (#484) #554 mergedscripts/backup-nas.shtoday and closed erebor/apps is declared backed up in two files and backed up by nothing #484, but on the monitoring hostbackups/nasdoes not exist,homelab-backup-nas.timeris not installed,ssh frodo@10.0.40.30is refused, andbuild-the-nas.md§6.2 is marked Not yet done. ADR-0008 calls that dataset the irreplaceable half of the media tier, and it is now on one spinning disk and nowhere else.build-the-nas.md§7 andhardware.mdboth still read that the pool isONLINEwith no errors.The clock
The pair is eBay item 237056026029, delivered 2026-09-18. The eBay Money Back Guarantee runs
30 days from delivery, so a "not as described / defective" return must be opened by
2026-10-18. Whether Seagate's own warranty applies to an "0HR" lot is unknown until the
serial is checked.
What has to happen, in this order
zpool status -v erebor,dmesgforsdb,smartctl -a,-l selftestand-l farmonsdb; readings into the runbook's status block;media-node-exporterrestarted, or the hang recorded if it does not come backerebor/appscopied off and verified — §6.2 andmake backup-nas, or the console tar in the runbook — before the tray is touched — done 2026-09-20: set20260920T060234Z,jellyfin.dbpresent, hashed onoracle(docs(nas): snapshot task and backup pull done; port 22 stops being inert everywhere it was said (#484) #563, erebor/apps is declared backed up in two files and backed up by nothing #484 closed on it)ZpoolNotOnlinerule and its tests merged (feat(nas): erebor lost a disk one day in — a rule that reads pool state, the runbook for the swap, and the record of what fired first #559) — and, read at the console, blind to this fault: the pool kstat saysonlinewith a faulted leaf. A vdev-level signal is now a line of its own belowONLINEpool: TrueNAS alert service to a phone channel, or a textfile metric onsmaug(smaug has no SMART and no patch state, and the mechanism every other host uses is a push it cannot make #483)hardware.mdcarries the new serial and its §2 readings;build-the-nas.md§7'szpool statusline is true againhomelab-backup-nas.timerinstalled (§6.2 step 7), so the copy stops being a one-off — done 2026-09-20, next run Sat 2026-09-26 03:32 UTCRunbook:
docs/runbooks/replace-the-nas-disk.md. Refs #483 (why nothing onsmaugcanpush more than node_exporter offers), #413 (the build this happened one day after).
Corrected 2026-09-20
Point 2 under Three things this exposed is answered:
erebor/appshas an off-host copy as of20260920T060234Z,frodoand SSH exist onsmaug, §6.2 reads Done 2026-09-20, and the timer is installed — #484 closed on it. Two boxes above are ticked with the evidence. The tray can be touched.Milestone NAS,
seq/1: this swap is now what that milestone closes on, beside #523. The vdev-level signal (a faulted leaf under anONLINEpool) stays an open box here; #483, which owns how anything host-level reachessmaug, is re-sequenced toseq/1in its own milestone so the two are worked together rather than decided twice.