# CLAUDE.md This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository. Flake-based NixOS config for a homelab. Hosts: **jupiter** (ZimaBlade NAS, x86_64), **neptun** (netcup public reverse proxy + tailnet node, x86_64), **mercury** (Raspberry Pi 3B+ DNS/DHCP, aarch64). See `README.md` for the full install/deploy walkthrough. ## Layout ``` flake.nix # nixosConfigurations: real hosts + test/util targets common.nix # shared base: user darman (key-only ssh), nix settings, firewall :22, tz services//*.nix # one reusable NixOS module per service, grouped by category # (media, network, vpn, identity, dev, desktop); each opens # ITS OWN firewall ports. services/containers.nix (podman # backend) stays at the top level, shared across categories. hosts// # configuration.nix + disk-config.nix (disko) + hardware-configuration.nix + secrets.nix secrets/.yaml # sops-nix, age-encrypted per host scripts/deploy # config-agnostic deploy wrapper (all args mandatory) scripts/edit_secrets .sops.yaml # per-host encryption rules (admin key + each host's key) ``` A host = `common.nix` + the `services/**` modules it imports + its `hosts//configuration.nix`. `services/` modules are engine-agnostic and shared across hosts (e.g. `services/vpn/tailscale.nix`, `services/network/caddy.nix` used by jupiter and neptun). ## Commands Eval/verify a config before deploying (eval only checks the module tree, not freeform config like pihole's TOML or a container's runtime): ``` nix eval --raw .#nixosConfigurations..config.system.build.toplevel.drvPath ``` Deploy (from a non-NixOS laptop too — runs nixos-rebuild/nixos-anywhere via `nix run`): ``` ./scripts/deploy switch # daily rebuild + activate ./scripts/deploy install # first install (nixos-anywhere, wipes OS disk) ./scripts/deploy kexec # RO-root box (ZimaOS): kexec into a RAM installer first sudo ./scripts/deploy kexec-local [--yes] # kexec THIS box (no ssh); confirm prompt unless --yes ./scripts/deploy image mercury # build the aarch64 SD image ./scripts/deploy flash mercury /dev/sdX # build + write SD + drop the sops age key ``` Both password prompts are auto-filled from the "HomeLab" Proton Pass vault, keyed by **``, not ``** — items `darman@` (sudo) and `root@` (ssh). A *trashed* Proton Pass item with the same title shadows the active one and yields an empty password, so the script resolves the title among `--filter-state active` items first; a plain `pass-cli item view --item-title` silently returns the trashed copy and you get an interactive prompt with no explanation. `HOMELAB_KEXEC_TARBALL` (+ `_CPIO` / `_GZIP`) makes `kexec`/`kexec-local` reuse a prebuilt installer instead of rebuilding ~500MB. The VM test below uses this. Secrets (needs the admin age key at `~/.config/sops/age/keys.txt`): ``` ./scripts/edit_secrets secrets/.yaml ``` Test a service config BEFORE touching hardware — always do this for nontrivial changes: ``` # x86 QEMU VM of mercury's DNS/DHCP stack (fast; validates pihole/unbound at runtime) nix build .#nixosConfigurations.mercury-vm.config.system.build.vm -o result ./result/bin/run-mercury-vm-vm # ssh -p 2223 darman@localhost (pw: test) # jupiter services as a VirtualBox OVA nix build .#nixosConfigurations.jupiter-vbox.config.system.build.virtualBoxOVA # end-to-end VM test of `deploy kexec-local` (~45s once the tarball is built) nix build .#checks.x86_64-linux.kexec-local -L ``` `checks.kexec-local` is the only way to exercise `kexec-local` at all: it jumps the machine you are typing at, so it cannot be rehearsed on real hardware and a failure looks exactly like a slow boot. It asserts the box actually left the old kernel (SSH drops then returns), came back as `nixos-installer`, lost its old `/run`, and kept its ssh host key. Run it after ANY change to the kexec paths. ## Secrets (sops-nix) - Each `secrets/.yaml` is encrypted to the **admin** key (edit) + that **host's** key (runtime decrypt); rules in `.sops.yaml`. Private keys live OFF-repo: `~/.config/sops/age/keys.txt` (admin), `~/.config/homelab//` (host keys). - jupiter/neptun decrypt with their **ssh host key** (`ssh-to-age` recipient), shipped at install via `nixos-anywhere --extra-files`. - mercury (SD image, no `--extra-files`) uses a **dedicated age key** at `/var/lib/sops-nix/age.txt` — `./scripts/deploy flash` writes it to the ext4 root partition. - A service password that must come from sops but whose module has no `passwordFile` hook (pihole, adguard) is injected via `sops.templates` → an env file → the service (`FTLCONF_*` for pihole). See `hosts/mercury/configuration.nix`. ## Non-obvious gotchas (all learned the hard way) - **aarch64 (mercury)**: the x86 laptop needs `extra-platforms = aarch64-linux` in `/etc/nix/nix.custom.conf` (NOT `/etc/nix/nix.conf` — Determinate Nix regenerates that) + `qemu-user-static-binfmt`, else emulated builds fail with "platform mismatch". Or build on the Pi with `--build-host darman@`. - **pihole on mercury is a CONTAINER** (`services/network/pihole.nix`, official image, host networking, caps NET_ADMIN/NET_RAW/SYS_NICE/CHOWN, `FTLCONF_*` env config). The native `services.pihole-ftl` module **segfaults on the Pi 3B+ aarch64** — do not switch back. - **`services.unbound.resolveLocalQueries = false`** is required: unbound listens on :5335, so leaving it true points the host's resolv.conf at 127.0.0.1:**53** with nothing there → boot-time DNS deadlock (starves image pulls / list downloads). Host resolves via upstream `networking.nameservers`; pihole forwards to unbound explicitly at `127.0.0.1#5335`. - **Remote deploy pushes unsigned closures**: hosts set `nix.settings.trusted-users = [ "root" "@wheel" ]` (in common.nix) so a laptop-built closure is accepted by the target. - **jupiter**: `boot.kernelParams = [ "reboot=pci" ]` (warm reboot hangs on that board); eMMC initrd modules pinned in `configuration.nix` (generate-config misses them); the 16TB×2 **RAID0** data lives on `/mnt/data` with `nofail`, kept OUT of disko (never wiped). - **terra: `./scripts/deploy kexec-local` hangs hard — do not use it there.** Confirmed on real hardware: kexec's `device_shutdown()` pass runs (SCSI disks sync fine in the log), then the machine goes dark and never comes back — `journalctl --list-boots` showed a ~15min gap before the next boot, i.e. a genuine hang needing a manual power cycle, not a slow jump. Near-certainly amdgpu (RX 6800 XT): discrete AMD GPUs are known to hang during kexec's device-shutdown pass with no clean handoff before the jump — same class of issue as jupiter's `reboot=pci` workaround, just fatal here instead of slow. Use the USB installer path instead (README's "First install on terra" section). - **disko wipes only the OS disk** named in `hosts//disk-config.nix`; data disks are plain `fileSystems` in `configuration.nix`. - `nixos-anywhere`/kexec needs a writable root; **ZimaOS root is read-only**, hence the `./scripts/deploy kexec` step that streams a RAM installer (with static cpio/gzip since ZimaOS lacks them). - **`kexec/run` jumps ~6s AFTER it returns**: nixos-images' `kexec-run.sh` ends with `nohup sh -c "sleep 6 && $SCRIPT_DIR/kexec -e" &`. So the staging dir must OUTLIVE the script — an `rm -rf` in an EXIT trap deletes the binary that performs the jump and the box silently stays on the old kernel. `kexec-local` clears its trap before jumping and then sleeps 60s on purpose. Covered by `checks.kexec-local`. - **The kexec installer KEEPS the box's ssh host key**: `kexec-run.sh` copies `/etc/ssh/ssh_host_*` into the appended initrd and `restore-remote-access.nix` installs them back. So do NOT `ssh-keygen -R` after a kexec — the key does not change, and clearing it just throws away the known_hosts record. - **`kexec-local` stages on `/var/tmp`, not `/tmp`**: `kexec-run.sh` appends a fresh cpio to `kexec/initrd` in place and execs binaries from that dir, so a size-capped or `noexec` tmpfs gives a half-written initrd or a bare "Permission denied".