Files
homelab/CLAUDE.md
T
darmanandClaude Sonnet 5 2a27d2cf4b terra: kexec-local hangs hard on real hardware, switch docs to USB installer
Confirmed on real hardware: kexec's device_shutdown() pass runs (SCSI disks
sync fine in the log) then the machine goes dark for good — journalctl
--list-boots showed a ~15min gap before the next boot, a genuine hang needing
a manual power cycle, not a slow jump. Near-certainly amdgpu (RX 6800 XT):
discrete AMD GPUs are known to hang during kexec's device-shutdown pass with
no clean handoff before the jump, same class of issue as jupiter's
reboot=pci warm-reboot workaround, just fatal here instead of slow.

README's terra install section now leads with the USB installer path instead
(build ISO, dd to USB, rsync the repo over, disko + nixos-install locally).
CLAUDE.md's gotchas list gets the same warning. installer-iso is renamed from
jupiter-installer to homelab-installer since it's genuinely host-agnostic,
and now ships git.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-24 01:11:52 +02:00

135 lines
8.2 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# CLAUDE.md
This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.
Flake-based NixOS config for a homelab. Hosts: **jupiter** (ZimaBlade NAS, x86_64),
**neptun** (netcup public reverse proxy + tailnet node, x86_64), **mercury** (Raspberry
Pi 3B+ DNS/DHCP, aarch64). See `README.md` for the full install/deploy walkthrough.
## Layout
```
flake.nix # nixosConfigurations: real hosts + test/util targets
common.nix # shared base: user darman (key-only ssh), nix settings, firewall :22, tz
services/<cat>/*.nix # one reusable NixOS module per service, grouped by category
# (media, network, vpn, identity, dev, desktop); each opens
# ITS OWN firewall ports. services/containers.nix (podman
# backend) stays at the top level, shared across categories.
hosts/<h>/ # configuration.nix + disk-config.nix (disko) + hardware-configuration.nix + secrets.nix
secrets/<h>.yaml # sops-nix, age-encrypted per host
scripts/deploy # config-agnostic deploy wrapper (all args mandatory)
scripts/edit_secrets
.sops.yaml # per-host encryption rules (admin key + each host's key)
```
A host = `common.nix` + the `services/**` modules it imports + its `hosts/<h>/configuration.nix`.
`services/` modules are engine-agnostic and shared across hosts (e.g. `services/vpn/tailscale.nix`,
`services/network/caddy.nix` used by jupiter and neptun).
## Commands
Eval/verify a config before deploying (eval only checks the module tree, not
freeform config like pihole's TOML or a container's runtime):
```
nix eval --raw .#nixosConfigurations.<host>.config.system.build.toplevel.drvPath
```
Deploy (from a non-NixOS laptop too — runs nixos-rebuild/nixos-anywhere via `nix run`):
```
./scripts/deploy switch <config> <host> # daily rebuild + activate
./scripts/deploy install <config> <host> # first install (nixos-anywhere, wipes OS disk)
./scripts/deploy kexec <config> <host> # RO-root box (ZimaOS): kexec into a RAM installer first
sudo ./scripts/deploy kexec-local [--yes] # kexec THIS box (no ssh); confirm prompt unless --yes
./scripts/deploy image mercury # build the aarch64 SD image
./scripts/deploy flash mercury /dev/sdX # build + write SD + drop the sops age key
```
Both password prompts are auto-filled from the "HomeLab" Proton Pass vault, keyed by
**`<config>`, not `<host>`** — items `darman@<config>` (sudo) and `root@<config>` (ssh).
A *trashed* Proton Pass item with the same title shadows the active one and yields an
empty password, so the script resolves the title among `--filter-state active` items
first; a plain `pass-cli item view --item-title` silently returns the trashed copy and
you get an interactive prompt with no explanation.
`HOMELAB_KEXEC_TARBALL` (+ `_CPIO` / `_GZIP`) makes `kexec`/`kexec-local` reuse a
prebuilt installer instead of rebuilding ~500MB. The VM test below uses this.
Secrets (needs the admin age key at `~/.config/sops/age/keys.txt`):
```
./scripts/edit_secrets secrets/<host>.yaml
```
Test a service config BEFORE touching hardware — always do this for nontrivial changes:
```
# x86 QEMU VM of mercury's DNS/DHCP stack (fast; validates pihole/unbound at runtime)
nix build .#nixosConfigurations.mercury-vm.config.system.build.vm -o result
./result/bin/run-mercury-vm-vm # ssh -p 2223 darman@localhost (pw: test)
# jupiter services as a VirtualBox OVA
nix build .#nixosConfigurations.jupiter-vbox.config.system.build.virtualBoxOVA
# end-to-end VM test of `deploy kexec-local` (~45s once the tarball is built)
nix build .#checks.x86_64-linux.kexec-local -L
```
`checks.kexec-local` is the only way to exercise `kexec-local` at all: it jumps the
machine you are typing at, so it cannot be rehearsed on real hardware and a failure
looks exactly like a slow boot. It asserts the box actually left the old kernel
(SSH drops then returns), came back as `nixos-installer`, lost its old `/run`, and
kept its ssh host key. Run it after ANY change to the kexec paths.
## Secrets (sops-nix)
- Each `secrets/<host>.yaml` is encrypted to the **admin** key (edit) + that **host's**
key (runtime decrypt); rules in `.sops.yaml`. Private keys live OFF-repo:
`~/.config/sops/age/keys.txt` (admin), `~/.config/homelab/<host>/` (host keys).
- jupiter/neptun decrypt with their **ssh host key** (`ssh-to-age` recipient), shipped at
install via `nixos-anywhere --extra-files`.
- mercury (SD image, no `--extra-files`) uses a **dedicated age key** at
`/var/lib/sops-nix/age.txt``./scripts/deploy flash` writes it to the ext4 root partition.
- A service password that must come from sops but whose module has no `passwordFile`
hook (pihole, adguard) is injected via `sops.templates` → an env file → the service
(`FTLCONF_*` for pihole). See `hosts/mercury/configuration.nix`.
## Non-obvious gotchas (all learned the hard way)
- **aarch64 (mercury)**: the x86 laptop needs `extra-platforms = aarch64-linux` in
`/etc/nix/nix.custom.conf` (NOT `/etc/nix/nix.conf` — Determinate Nix regenerates that)
+ `qemu-user-static-binfmt`, else emulated builds fail with "platform mismatch". Or
build on the Pi with `--build-host darman@<ip>`.
- **pihole on mercury is a CONTAINER** (`services/network/pihole.nix`, official image, host
networking, caps NET_ADMIN/NET_RAW/SYS_NICE/CHOWN, `FTLCONF_*` env config). The native
`services.pihole-ftl` module **segfaults on the Pi 3B+ aarch64** — do not switch back.
- **`services.unbound.resolveLocalQueries = false`** is required: unbound listens on
:5335, so leaving it true points the host's resolv.conf at 127.0.0.1:**53** with nothing
there → boot-time DNS deadlock (starves image pulls / list downloads). Host resolves via
upstream `networking.nameservers`; pihole forwards to unbound explicitly at `127.0.0.1#5335`.
- **Remote deploy pushes unsigned closures**: hosts set `nix.settings.trusted-users =
[ "root" "@wheel" ]` (in common.nix) so a laptop-built closure is accepted by the target.
- **jupiter**: `boot.kernelParams = [ "reboot=pci" ]` (warm reboot hangs on that board);
eMMC initrd modules pinned in `configuration.nix` (generate-config misses them); the
16TB×2 **RAID0** data lives on `/mnt/data` with `nofail`, kept OUT of disko (never wiped).
- **terra: `./scripts/deploy kexec-local` hangs hard — do not use it there.** Confirmed
on real hardware: kexec's `device_shutdown()` pass runs (SCSI disks sync fine in the
log), then the machine goes dark and never comes back — `journalctl --list-boots`
showed a ~15min gap before the next boot, i.e. a genuine hang needing a manual power
cycle, not a slow jump. Near-certainly amdgpu (RX 6800 XT): discrete AMD GPUs are known
to hang during kexec's device-shutdown pass with no clean handoff before the jump —
same class of issue as jupiter's `reboot=pci` workaround, just fatal here instead of
slow. Use the USB installer path instead (README's "First install on terra" section).
- **disko wipes only the OS disk** named in `hosts/<h>/disk-config.nix`; data disks are
plain `fileSystems` in `configuration.nix`.
- `nixos-anywhere`/kexec needs a writable root; **ZimaOS root is read-only**, hence the
`./scripts/deploy kexec` step that streams a RAM installer (with static cpio/gzip since
ZimaOS lacks them).
- **`kexec/run` jumps ~6s AFTER it returns**: nixos-images' `kexec-run.sh` ends with
`nohup sh -c "sleep 6 && $SCRIPT_DIR/kexec -e" &`. So the staging dir must OUTLIVE the
script — an `rm -rf` in an EXIT trap deletes the binary that performs the jump and the
box silently stays on the old kernel. `kexec-local` clears its trap before jumping and
then sleeps 60s on purpose. Covered by `checks.kexec-local`.
- **The kexec installer KEEPS the box's ssh host key**: `kexec-run.sh` copies
`/etc/ssh/ssh_host_*` into the appended initrd and `restore-remote-access.nix` installs
them back. So do NOT `ssh-keygen -R` after a kexec — the key does not change, and
clearing it just throws away the known_hosts record.
- **`kexec-local` stages on `/var/tmp`, not `/tmp`**: `kexec-run.sh` appends a fresh cpio
to `kexec/initrd` in place and execs binaries from that dir, so a size-capped or
`noexec` tmpfs gives a half-written initrd or a bare "Permission denied".