Files
homelab/README.md
T
darmanandClaude Opus 4.8 0995a5fe2f headscale: move the tailnet to orbit.sol, route all DNS through pihole
Three connected changes, all triggered by the same outage.

base_domain leaves mgaction.town. That zone has a wildcard A+AAAA pointing
at neptun, and DNS wildcards match multi-label names, so
jupiter.hosts.mgaction.town resolved publicly to NEPTUN and Caddy proxied
to itself -- a silent loop rather than a lookup failure. Nesting the
tailnet inside the LAN domain as orbit.sol keeps the theme and resolves
unambiguously, since tailscale matches routes by longest suffix.

override_local_dns = true with pihole as the only global nameserver, so
roaming devices get ad blocking and .sol names off-LAN. With it false,
globalResolvers land in the netmap's FallbackResolvers, which a phone
with carrier DNS never consults. No public fallback is listed on purpose:
tailscale treats the list as a set, so a second entry would let queries
slip past the filter whenever mercury is slow. The cost is that mercury
is now a single point of failure for tailnet DNS.

neptun and mercury opt out individually. mercury would otherwise resolve
through itself. neptun must not depend on a Pi behind a domestic line to
renew the certificates for the control server every other node needs --
and it is circular besides, since tailscaled has to resolve
vpn.mgaction.town to connect at all. Instead neptun runs a dnsmasq stub
forwarding just orbit.sol to MagicDNS on 100.100.100.100, which tailscaled
answers whenever it is running regardless of --accept-dns. That resolves
jupiter live, so the hardcoded /etc/hosts pin is gone.

Also sets dns.nameservers.split explicitly: nixpkgs renders its own
dns.split option one level too high, but headscale reads
dns.nameservers.split (hscontrol/types/config.go:722) and so does
headplane, whose DNS page dies on the missing key with "Cannot convert
undefined or null to object". The module's option is dead as written.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-20 23:01:47 +02:00

272 lines
14 KiB
Markdown

# homelab
Flake-based NixOS config. Hosts: `jupiter` (ZimaBlade, NAS + services),
`neptun` (netcup VPS: public reverse proxy, Authentik, headscale),
`mercury` (Raspberry Pi 3B+, DNS/DHCP), `terra` (desktop).
## Structure
```
flake.nix # inputs + nixosConfigurations (jupiter, neptun, kexec, ...)
common.nix # shared base: user, ssh, nix, firewall, timezone
services/ # one reusable module per service, by category
media/ jellyfin, audiobookshelf, the *arrs, sabnzbd, seerr, ...
network/ caddy, samba, avahi, pihole, unbound
vpn/ tailscale, headscale (control server), headplane (its web UI)
identity/ authentik (OIDC provider, from the authentik-nix flake)
dev/ gitea
desktop/ hyprland
containers.nix # podman backend, shared across categories
hosts/
jupiter/ # ZimaBlade NAS
configuration.nix # host bits + imports common + the services it runs
disk-config.nix # disko: eMMC partitions
hardware-configuration.nix
secrets.nix # sops-nix wiring
vm.nix # VirtualBox test image (jupiter-vbox)
neptun/ # netcup public reverse proxy + tailnet node
configuration.nix disk-config.nix hardware-configuration.nix secrets.nix
secrets/ # age-encrypted sops files, one per host
scripts/ # deploy, edit_secrets
```
Hosts compose by importing `common.nix` + whichever `services/*` modules they
run. Each service module opens its own firewall ports.
## Test in VirtualBox (no hardware needed)
```
nix build .#nixosConfigurations.jupiter-vbox.config.system.build.virtualBoxOVA
VBoxManage import result/*.ova --vsys 0 --vmname jupiter-vbox
VBoxManage startvm jupiter-vbox --type headless
```
Login `darman` / `test`. Forward ports with `VBoxManage modifyvm ... --natpf1`.
## First install on the ZimaBlade — nixos-anywhere + disko
Wipes the OS disk and installs the flake over SSH. No USB needed if the box
already runs Linux (ZimaOS) reachable by root SSH — nixos-anywhere kexecs into
an installer, partitions via disko, installs.
> ⚠️ The OS disk in `disk-config.nix` is WIPED. Set `device` to the OS disk
> ONLY (by-id). Back up / physically identify the NAS data disk first — it must
> NOT appear in disko. `lsblk -o NAME,SERIAL,SIZE,MODEL` to identify.
1. Set the real OS disk id in `hosts/jupiter/disk-config.nix`
(`ls -l /dev/disk/by-id`), and the data-disk mount in `configuration.nix`.
2. Add your login SSH pubkey to `users.users.darman.openssh.authorizedKeys.keys`.
3. Set the real samba password:
```
export SOPS_AGE_KEY_FILE=~/.config/sops/age/keys.txt
nix shell nixpkgs#sops -c sops secrets/jupiter.yaml # edit, commit
```
4. Stage the pre-generated host key so sops can decrypt on boot #1
(private key lives off-repo in `~/.config/homelab/jupiter/`):
```
install -Dm600 ~/.config/homelab/jupiter/ssh_host_ed25519_key \
/tmp/extra/etc/ssh/ssh_host_ed25519_key
install -Dm644 ~/.config/homelab/jupiter/ssh_host_ed25519_key.pub \
/tmp/extra/etc/ssh/ssh_host_ed25519_key.pub
```
5. Run from your laptop:
```
nix run github:nix-community/nixos-anywhere -- \
--flake .#jupiter \
--extra-files /tmp/extra \
--generate-hardware-config nixos-generate-config ./hosts/jupiter/hardware-configuration.nix \
--target-host root@<zimablade-ip>
```
`--extra-files` plants the host key before first boot (its age identity is
already a recipient in `.sops.yaml`, so `/run/secrets/samba_password`
decrypts on boot #1). `--generate-hardware-config` pulls the target's real
kernel modules into the placeholder. Commit the result. Reboot into NixOS.
Manual alternative (USB ISO): boot installer, `disko` the disk, then
`nixos-install --flake .#jupiter`.
## Deploy (the `./deploy` wrapper)
All arguments mandatory — no default host, no default config.
```
./deploy kexec <host> # headless kexec into a RAM installer (RO-root box)
./deploy install <config> <host> # first install; wipes OS disk, ships host key
./deploy switch <config> <host> # rebuild + activate on a running host
./deploy boot|test <config> <host> # stage for next boot / activate without boot entry
./deploy image <config> # build an SD-card image (mercury)
./deploy flash <config> <dev> # build SD image, write it, drop the sops age key
```
`switch`/`boot`/`test` prompt for darman's password (`wheelNeedsPassword`).
`<config>` is a `nixosConfigurations` name (`jupiter`, `neptun`). Its pre-generated
SSH host key lives at `~/.config/homelab/<config>/ssh_host_ed25519_key`.
Examples:
```
./deploy switch jupiter jupiter.sol
./deploy install neptun 159.195.64.117
```
Rollback: `nixos-rebuild switch --rollback` on the host, or pick a prior
generation at boot.
## Post-deploy steps (per host)
Things the flake cannot do for you. Skipping these leaves a host that builds
and boots but doesn't work.
### Every host, immediately after a first install
```
ssh darman@<host> sudo -v # DO NOT SKIP
```
`users.mutableUsers` is `true`, so `/etc/shadow` is written **once**, when the
user is created. If the sops secret wasn't readable at that moment the account
gets `!` (locked) permanently — `deploy switch` will never fix it, because the
activation script only sets a password for users not already in `/etc/shadow`.
Combined with `wheelNeedsPassword = true` and `PermitRootLogin = "no"` that
means no way to escalate, and recovery is physical: netcup's rescue system for
neptun, or pulling the SD card for mercury. Verify sudo while you still have
another way in.
### neptun (netcup VPS)
1. **Edge firewall.** In netcup's panel, inbound `ACCEPT` for TCP 22/80/443/2222
**and a rule accepting inbound UDP**. The firewall is stateless: without the
UDP rule every DNS and NTP *reply* is dropped, and nothing on the box reports
an error — it looks like headscale crash-looping on its DERP fetch and Caddy
failing ACME. `grep -A1 '^Udp:' /proc/net/snmp` showing `InDatagrams 0` is the
tell. Rules apply on VM restart, not on save. This is safe: `nixos-fw` is
stateful and default-deny, so it remains the real policy.
Also open UDP 3478 (STUN) and 41641 (tailscale direct).
2. **Authentik** creates `akadmin` on first start; log in at
`https://auth.mgaction.town` with `authentik_bootstrap_password` from sops.
The username is hardcoded upstream and the bootstrap runs once — later
changes to the env vars are ignored.
To use your own admin instead: create a user, add it to the **`authentik
Admins`** group (superuser is a *group* flag in Authentik, there is no
per-user one), verify it works in a private window, then **deactivate**
`akadmin` — do not rename or delete it. The bootstrap blueprint keys on
`username: akadmin` with `state: created`, so if no user by that name
exists it simply makes a new one on the next reconcile.
3. **Bootstrap the tailnet** (headscale starts with an empty database):
```
sudo headscale users create darman
sudo headscale preauthkeys create --user darman --reusable --expiration 24h
```
Put that key in **every** host's sops file as `tailscale_authkey` and rebuild.
4. **Headplane API key** — defaults to 90d, after which headplane silently stops
listing nodes:
```
sudo headscale apikeys create --expiration 999d # -> headplane_headscale_api_key
```
5. **Headplane OIDC.** In Authentik create an OAuth2/OpenID provider
(confidential, redirect `https://vpn.mgaction.town/admin/oidc/callback`,
**a signing key must be selected** or discovery exposes no JWKS) and an
application with slug **`headplane`** — the slug is what makes the issuer
`.../application/o/headplane/` in `services/vpn/headplane.nix`. Client ID goes
in that file, client secret into sops.
6. **Headscale OIDC** (optional — pre-auth keys work without it). A *second*
Authentik provider/application, slug **`headscale`**, redirect
`https://vpn.mgaction.town/oidc/callback` (headscale's own, not headplane's
under `/admin`). Client ID in `services/vpn/headscale.nix`, secret into sops
as `headscale_oidc_client_secret`.
⚠️ headscale runs OIDC discovery **at startup and a failure is fatal** —
an issuer pointing at an application that doesn't exist yet means the
control server won't boot, taking the whole tailnet's control plane with
it. Always verify first:
```
curl -s https://auth.mgaction.town/application/o/headscale/.well-known/openid-configuration
```
Users created by OIDC login are distinct from `headscale users create`
ones: headplane matches the OIDC `sub` claim against the user's
`providerId`, CLI-made users have none, and 0.28 dropped both
`map_legacy_users` and node reassignment — so moving an existing node to
an OIDC user means re-enrolling it.
### jupiter
- `chown -R gitea:gitea /mnt/data/AppData/gitea` after the first deploy (the
repos were copied in over CIFS as `darman:users`).
- **Re-enrolling after the headscale database was recreated:** `tailscaled`
keeps its old node key and reports `Running`, and the autoconnect unit exits
early on that state without ever sending the new pre-auth key. Force it:
```
sudo tailscale logout && sudo systemctl restart tailscaled-autoconnect
```
### mercury (Raspberry Pi 3B+)
- `./deploy flash mercury /dev/sdX` writes the dedicated age key to the root
partition. Without `~/.config/homelab/mercury/age.txt` it silently skips that
step and **no secret decrypts on the box** — check `ls /run/secrets` after
first boot.
- It boots from an SD card, so config changes are `./deploy switch mercury <ip>`
(an aarch64 build — needs `extra-platforms` + binfmt on the laptop, see the
gotchas in `CLAUDE.md`) rather than a reflash.
- **Suspect the card first** when binaries crash with `Illegal instruction` or
services fail inexplicably. Failing flash returns corrupt data with no I/O
errors in `dmesg`:
```
sudo nix-store --verify --check-contents # add --repair to fix
```
A card that has corrupted one path will corrupt more. Replace it and reflash.
- **A reflash wipes `/var/lib/pihole`**, taking the gravity database with it.
The blocklists themselves are declared in `services/network/pihole.nix`, and
the `pihole-adlists` unit re-seeds them on boot and rebuilds gravity when it
finds it empty — so this heals itself, but the first boot after a reflash
spends several minutes downloading lists. Query history and dynamic DHCP
leases are genuinely lost (static leases are declarative). Check with:
```
systemctl status pihole-adlists
sudo podman exec pihole pihole-FTL sqlite3 /etc/pihole/gravity.db \
"SELECT address,enabled FROM adlist; SELECT COUNT(*) FROM gravity;"
```
`Blocked DNS queries: 0` in the pihole logs means gravity is empty — DNS
resolves fine, nothing is filtered.
- mercury's own `resolv.conf` is deliberately public resolvers, not its own
pihole (`resolveLocalQueries = false`, see `CLAUDE.md`) — so `.sol` names do
not resolve *on mercury itself*. That is expected, not a fault.
- **mercury is load-bearing for the whole tailnet's DNS.** headscale sets
`override_local_dns = true` with pihole as the only global nameserver, so
every node — including a phone on mobile data — resolves through it and gets
ad blocking and `.sol` names anywhere. The flip side is that mercury (or the
home connection) going down costs name resolution on every device, not just
`.sol`. Recovery on a stranded device is turning Tailscale off.
There is deliberately no public fallback in `nameservers.global`: tailscale
treats that list as a set, so a second entry would let queries slip past the
filter whenever mercury is slow.
neptun and mercury opt out with `--accept-dns=false` — mercury because it
would otherwise resolve through itself, neptun because a public reverse
proxy must not depend on a Pi at home to renew its certificates.
- Tailnet names are `*.orbit.sol`, LAN names are `*.sol`. Both work everywhere
on the tailnet because tailscale matches DNS routes by **longest suffix**, so
`orbit.sol` reaches MagicDNS even though everything else goes to pihole.
Never name a LAN host `orbit`: pihole's `address=/<host>.sol/<ip>` lines match
a name *and everything beneath it*, which would swallow the entire tailnet
zone.
## Adding a service
Copy the `whoami` block in `oci-containers.containers`, swap image/ports/volumes.
Native NixOS module exists for many apps (Nextcloud, Jellyfin, Grafana...) —
prefer `services.<app>` over a container when available. Add a `caddy`
`virtualHosts` block to expose it.
## Notes
- Backend is Podman with `dockerCompat` — `docker` CLI works, no daemon.
- Samba keeps its own password DB. `services.samba` never sets it; a systemd
oneshot (`samba-smbpasswd`) provisions it. Host reads the password from
`/run/secrets/samba_password` (**sops-nix**); the VM falls back to plaintext
`/etc/samba/smb-password`.
- Secrets: `secrets/jupiter.yaml` is age-encrypted (safe to commit) to two
recipients in `.sops.yaml` — the **admin** key (edit on laptop,
`~/.config/sops/age/keys.txt`) and the **jupiter host** key (derived from its
SSH host key via `ssh-to-age`, decrypts at runtime). Private keys live
off-repo and are gitignored. Rotate/add recipients with `sops updatekeys`.
- Data disk: plain `fileSystems."/mnt/data"` in configuration.nix — kept out of
disko so it is never formatted. Reference by `by-id` / `by-uuid`.
- `system.stateVersion` = `26.05`, install-time schema. Do NOT bump on upgrades.
- Terraform is not used: a single bare-metal box has no provider API. disko +
nixos-anywhere cover provisioning natively.