Files
homelab/CLAUDE.md
T
darmanandClaude Sonnet 5 0ea90200b4 deploy: automate a full local reinstall, self-elevating and interactive-safe
./scripts/deploy install <config> localhost now branches on is_live_installer()
(checks uname -n): outside a live installer it builds installer-iso, stages
its kernel/initrd on the ESP and the iso file on a disk the caller picks
(never auto-picked — the wrong disk here is destroyed mid-install), writes a
systemd-boot one-shot findiso= entry with homelab.install=<config> on the
kernel cmdline, and does a real systemctl reboot (not kexec — terra's
kexec-local hang is specifically in kexec's device-shutdown pass, a real ACPI
reboot never runs that code at all).

installer-iso gains homelab-auto-install.service: once homelab-checkout.service
clones the repo, it reads homelab.install= back off /proc/cmdline and re-runs
the identical deploy command itself, now genuinely inside the installer, so
it takes the disko+nixos-install branch instead of preparing again. The whole
reinstall is one command and unattended after the first reboot.

Also: every root-requiring path (kexec-local, the new prepare-and-reboot
branch, the disko+nixos-install branch) self-elevates via a require_root()
helper that re-execs the original invocation under sudo -E, instead of dying
and asking the caller to prefix sudo themselves. Uses an absolute script path
captured before the script's own cd, so the re-exec is correct regardless of
how it was invoked.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-24 01:53:45 +02:00

150 lines
9.4 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# CLAUDE.md
This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.
Flake-based NixOS config for a homelab. Hosts: **jupiter** (ZimaBlade NAS, x86_64),
**neptun** (netcup public reverse proxy + tailnet node, x86_64), **mercury** (Raspberry
Pi 3B+ DNS/DHCP, aarch64). See `README.md` for the full install/deploy walkthrough.
## Layout
```
flake.nix # nixosConfigurations: real hosts + test/util targets
common.nix # shared base: user darman (key-only ssh), nix settings, firewall :22, tz
services/<cat>/*.nix # one reusable NixOS module per service, grouped by category
# (media, network, vpn, identity, dev, desktop); each opens
# ITS OWN firewall ports. services/containers.nix (podman
# backend) stays at the top level, shared across categories.
hosts/<h>/ # configuration.nix + disk-config.nix (disko) + hardware-configuration.nix + secrets.nix
secrets/<h>.yaml # sops-nix, age-encrypted per host
scripts/deploy # config-agnostic deploy wrapper (all args mandatory)
scripts/edit_secrets
.sops.yaml # per-host encryption rules (admin key + each host's key)
```
A host = `common.nix` + the `services/**` modules it imports + its `hosts/<h>/configuration.nix`.
`services/` modules are engine-agnostic and shared across hosts (e.g. `services/vpn/tailscale.nix`,
`services/network/caddy.nix` used by jupiter and neptun).
## Commands
Eval/verify a config before deploying (eval only checks the module tree, not
freeform config like pihole's TOML or a container's runtime):
```
nix eval --raw .#nixosConfigurations.<host>.config.system.build.toplevel.drvPath
```
Deploy (from a non-NixOS laptop too — runs nixos-rebuild/nixos-anywhere via `nix run`):
```
./scripts/deploy switch <config> <host> # daily rebuild + activate
./scripts/deploy install <config> <host> # first install (nixos-anywhere, wipes OS disk)
./scripts/deploy kexec <config> <host> # RO-root box (ZimaOS): kexec into a RAM installer first
sudo ./scripts/deploy kexec-local [--yes] # kexec THIS box (no ssh); confirm prompt unless --yes
./scripts/deploy image mercury # build the aarch64 SD image
./scripts/deploy flash mercury /dev/sdX # build + write SD + drop the sops age key
```
Both password prompts are auto-filled from the "HomeLab" Proton Pass vault, keyed by
**`<config>`, not `<host>`** — items `darman@<config>` (sudo) and `root@<config>` (ssh).
A *trashed* Proton Pass item with the same title shadows the active one and yields an
empty password, so the script resolves the title among `--filter-state active` items
first; a plain `pass-cli item view --item-title` silently returns the trashed copy and
you get an interactive prompt with no explanation.
`HOMELAB_KEXEC_TARBALL` (+ `_CPIO` / `_GZIP`) makes `kexec`/`kexec-local` reuse a
prebuilt installer instead of rebuilding ~500MB. The VM test below uses this.
Secrets (needs the admin age key at `~/.config/sops/age/keys.txt`):
```
./scripts/edit_secrets secrets/<host>.yaml
```
Test a service config BEFORE touching hardware — always do this for nontrivial changes:
```
# x86 QEMU VM of mercury's DNS/DHCP stack (fast; validates pihole/unbound at runtime)
nix build .#nixosConfigurations.mercury-vm.config.system.build.vm -o result
./result/bin/run-mercury-vm-vm # ssh -p 2223 darman@localhost (pw: test)
# jupiter services as a VirtualBox OVA
nix build .#nixosConfigurations.jupiter-vbox.config.system.build.virtualBoxOVA
# end-to-end VM test of `deploy kexec-local` (~45s once the tarball is built)
nix build .#checks.x86_64-linux.kexec-local -L
```
`checks.kexec-local` is the only way to exercise `kexec-local` at all: it jumps the
machine you are typing at, so it cannot be rehearsed on real hardware and a failure
looks exactly like a slow boot. It asserts the box actually left the old kernel
(SSH drops then returns), came back as `nixos-installer`, lost its old `/run`, and
kept its ssh host key. Run it after ANY change to the kexec paths.
## Secrets (sops-nix)
- Each `secrets/<host>.yaml` is encrypted to the **admin** key (edit) + that **host's**
key (runtime decrypt); rules in `.sops.yaml`. Private keys live OFF-repo:
`~/.config/sops/age/keys.txt` (admin), `~/.config/homelab/<host>/` (host keys).
- jupiter/neptun decrypt with their **ssh host key** (`ssh-to-age` recipient), shipped at
install via `nixos-anywhere --extra-files`.
- mercury (SD image, no `--extra-files`) uses a **dedicated age key** at
`/var/lib/sops-nix/age.txt``./scripts/deploy flash` writes it to the ext4 root partition.
- A service password that must come from sops but whose module has no `passwordFile`
hook (pihole, adguard) is injected via `sops.templates` → an env file → the service
(`FTLCONF_*` for pihole). See `hosts/mercury/configuration.nix`.
## Non-obvious gotchas (all learned the hard way)
- **aarch64 (mercury)**: the x86 laptop needs `extra-platforms = aarch64-linux` in
`/etc/nix/nix.custom.conf` (NOT `/etc/nix/nix.conf` — Determinate Nix regenerates that)
+ `qemu-user-static-binfmt`, else emulated builds fail with "platform mismatch". Or
build on the Pi with `--build-host darman@<ip>`.
- **pihole on mercury is a CONTAINER** (`services/network/pihole.nix`, official image, host
networking, caps NET_ADMIN/NET_RAW/SYS_NICE/CHOWN, `FTLCONF_*` env config). The native
`services.pihole-ftl` module **segfaults on the Pi 3B+ aarch64** — do not switch back.
- **`services.unbound.resolveLocalQueries = false`** is required: unbound listens on
:5335, so leaving it true points the host's resolv.conf at 127.0.0.1:**53** with nothing
there → boot-time DNS deadlock (starves image pulls / list downloads). Host resolves via
upstream `networking.nameservers`; pihole forwards to unbound explicitly at `127.0.0.1#5335`.
- **Remote deploy pushes unsigned closures**: hosts set `nix.settings.trusted-users =
[ "root" "@wheel" ]` (in common.nix) so a laptop-built closure is accepted by the target.
- **jupiter**: `boot.kernelParams = [ "reboot=pci" ]` (warm reboot hangs on that board);
eMMC initrd modules pinned in `configuration.nix` (generate-config misses them); the
16TB×2 **RAID0** data lives on `/mnt/data` with `nofail`, kept OUT of disko (never wiped).
- **terra: `./scripts/deploy kexec-local` hangs hard — do not use it there.** Confirmed
on real hardware: kexec's `device_shutdown()` pass runs (SCSI disks sync fine in the
log), then the machine goes dark and never comes back — `journalctl --list-boots`
showed a ~15min gap before the next boot, i.e. a genuine hang needing a manual power
cycle, not a slow jump. Near-certainly amdgpu (RX 6800 XT): discrete AMD GPUs are known
to hang during kexec's device-shutdown pass with no clean handoff before the jump —
same class of issue as jupiter's `reboot=pci` workaround, just fatal here instead of
slow. Use `./scripts/deploy install terra localhost` instead (README's "First install
on terra" section) — it detects it isn't inside a live installer yet and reboots via
a real `systemctl reboot` + systemd-boot one-shot `findiso=` entry, not kexec.
- **`./scripts/deploy install <config> localhost`'s behavior depends on `uname -n`**
(`is_live_installer()`): on a real running OS it builds `installer-iso`, stages it
locally, and reboots into it (`local_install_prepare_and_reboot()`); only inside
`nixos-installer` (kexec) or `homelab-installer` (installer-iso) does it actually run
disko + `nixos-install`. `installer-iso`'s `homelab-auto-install.service` closes the
loop: it reads `homelab.install=<config>` back off `/proc/cmdline` (set by the prepare
step) and re-runs the identical command itself once `homelab-checkout.service` has
cloned the repo — the whole reinstall is one command and unattended after the first
reboot. It always ASKS where to stage the iso file (never auto-picks — the wrong disk
here is destroyed mid-install) and refuses if that turns out to be the disk
`disk-config.nix` is about to wipe; `HOMELAB_INSTALLER_STAGE_DIR` skips the prompt for
scripted use. Both this and `kexec-local` self-elevate via `sudo` (`require_root()`)
rather than requiring you to prefix the command yourself.
- **disko wipes only the OS disk** named in `hosts/<h>/disk-config.nix`; data disks are
plain `fileSystems` in `configuration.nix`.
- `nixos-anywhere`/kexec needs a writable root; **ZimaOS root is read-only**, hence the
`./scripts/deploy kexec` step that streams a RAM installer (with static cpio/gzip since
ZimaOS lacks them).
- **`kexec/run` jumps ~6s AFTER it returns**: nixos-images' `kexec-run.sh` ends with
`nohup sh -c "sleep 6 && $SCRIPT_DIR/kexec -e" &`. So the staging dir must OUTLIVE the
script — an `rm -rf` in an EXIT trap deletes the binary that performs the jump and the
box silently stays on the old kernel. `kexec-local` clears its trap before jumping and
then sleeps 60s on purpose. Covered by `checks.kexec-local`.
- **The kexec installer KEEPS the box's ssh host key**: `kexec-run.sh` copies
`/etc/ssh/ssh_host_*` into the appended initrd and `restore-remote-access.nix` installs
them back. So do NOT `ssh-keygen -R` after a kexec — the key does not change, and
clearing it just throws away the known_hosts record.
- **`kexec-local` stages on `/var/tmp`, not `/tmp`**: `kexec-run.sh` appends a fresh cpio
to `kexec/initrd` in place and execs binaries from that dir, so a size-capped or
`noexec` tmpfs gives a half-written initrd or a bare "Permission denied".