Three connected changes, all triggered by the same outage. base_domain leaves mgaction.town. That zone has a wildcard A+AAAA pointing at neptun, and DNS wildcards match multi-label names, so jupiter.hosts.mgaction.town resolved publicly to NEPTUN and Caddy proxied to itself -- a silent loop rather than a lookup failure. Nesting the tailnet inside the LAN domain as orbit.sol keeps the theme and resolves unambiguously, since tailscale matches routes by longest suffix. override_local_dns = true with pihole as the only global nameserver, so roaming devices get ad blocking and .sol names off-LAN. With it false, globalResolvers land in the netmap's FallbackResolvers, which a phone with carrier DNS never consults. No public fallback is listed on purpose: tailscale treats the list as a set, so a second entry would let queries slip past the filter whenever mercury is slow. The cost is that mercury is now a single point of failure for tailnet DNS. neptun and mercury opt out individually. mercury would otherwise resolve through itself. neptun must not depend on a Pi behind a domestic line to renew the certificates for the control server every other node needs -- and it is circular besides, since tailscaled has to resolve vpn.mgaction.town to connect at all. Instead neptun runs a dnsmasq stub forwarding just orbit.sol to MagicDNS on 100.100.100.100, which tailscaled answers whenever it is running regardless of --accept-dns. That resolves jupiter live, so the hardcoded /etc/hosts pin is gone. Also sets dns.nameservers.split explicitly: nixpkgs renders its own dns.split option one level too high, but headscale reads dns.nameservers.split (hscontrol/types/config.go:722) and so does headplane, whose DNS page dies on the missing key with "Cannot convert undefined or null to object". The module's option is dead as written. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
272 lines
14 KiB
Markdown
272 lines
14 KiB
Markdown
# homelab
|
|
|
|
Flake-based NixOS config. Hosts: `jupiter` (ZimaBlade, NAS + services),
|
|
`neptun` (netcup VPS: public reverse proxy, Authentik, headscale),
|
|
`mercury` (Raspberry Pi 3B+, DNS/DHCP), `terra` (desktop).
|
|
|
|
## Structure
|
|
|
|
```
|
|
flake.nix # inputs + nixosConfigurations (jupiter, neptun, kexec, ...)
|
|
common.nix # shared base: user, ssh, nix, firewall, timezone
|
|
services/ # one reusable module per service, by category
|
|
media/ jellyfin, audiobookshelf, the *arrs, sabnzbd, seerr, ...
|
|
network/ caddy, samba, avahi, pihole, unbound
|
|
vpn/ tailscale, headscale (control server), headplane (its web UI)
|
|
identity/ authentik (OIDC provider, from the authentik-nix flake)
|
|
dev/ gitea
|
|
desktop/ hyprland
|
|
containers.nix # podman backend, shared across categories
|
|
hosts/
|
|
jupiter/ # ZimaBlade NAS
|
|
configuration.nix # host bits + imports common + the services it runs
|
|
disk-config.nix # disko: eMMC partitions
|
|
hardware-configuration.nix
|
|
secrets.nix # sops-nix wiring
|
|
vm.nix # VirtualBox test image (jupiter-vbox)
|
|
neptun/ # netcup public reverse proxy + tailnet node
|
|
configuration.nix disk-config.nix hardware-configuration.nix secrets.nix
|
|
secrets/ # age-encrypted sops files, one per host
|
|
scripts/ # deploy, edit_secrets
|
|
```
|
|
|
|
Hosts compose by importing `common.nix` + whichever `services/*` modules they
|
|
run. Each service module opens its own firewall ports.
|
|
|
|
## Test in VirtualBox (no hardware needed)
|
|
|
|
```
|
|
nix build .#nixosConfigurations.jupiter-vbox.config.system.build.virtualBoxOVA
|
|
VBoxManage import result/*.ova --vsys 0 --vmname jupiter-vbox
|
|
VBoxManage startvm jupiter-vbox --type headless
|
|
```
|
|
Login `darman` / `test`. Forward ports with `VBoxManage modifyvm ... --natpf1`.
|
|
|
|
## First install on the ZimaBlade — nixos-anywhere + disko
|
|
|
|
Wipes the OS disk and installs the flake over SSH. No USB needed if the box
|
|
already runs Linux (ZimaOS) reachable by root SSH — nixos-anywhere kexecs into
|
|
an installer, partitions via disko, installs.
|
|
|
|
> ⚠️ The OS disk in `disk-config.nix` is WIPED. Set `device` to the OS disk
|
|
> ONLY (by-id). Back up / physically identify the NAS data disk first — it must
|
|
> NOT appear in disko. `lsblk -o NAME,SERIAL,SIZE,MODEL` to identify.
|
|
|
|
1. Set the real OS disk id in `hosts/jupiter/disk-config.nix`
|
|
(`ls -l /dev/disk/by-id`), and the data-disk mount in `configuration.nix`.
|
|
2. Add your login SSH pubkey to `users.users.darman.openssh.authorizedKeys.keys`.
|
|
3. Set the real samba password:
|
|
```
|
|
export SOPS_AGE_KEY_FILE=~/.config/sops/age/keys.txt
|
|
nix shell nixpkgs#sops -c sops secrets/jupiter.yaml # edit, commit
|
|
```
|
|
4. Stage the pre-generated host key so sops can decrypt on boot #1
|
|
(private key lives off-repo in `~/.config/homelab/jupiter/`):
|
|
```
|
|
install -Dm600 ~/.config/homelab/jupiter/ssh_host_ed25519_key \
|
|
/tmp/extra/etc/ssh/ssh_host_ed25519_key
|
|
install -Dm644 ~/.config/homelab/jupiter/ssh_host_ed25519_key.pub \
|
|
/tmp/extra/etc/ssh/ssh_host_ed25519_key.pub
|
|
```
|
|
5. Run from your laptop:
|
|
```
|
|
nix run github:nix-community/nixos-anywhere -- \
|
|
--flake .#jupiter \
|
|
--extra-files /tmp/extra \
|
|
--generate-hardware-config nixos-generate-config ./hosts/jupiter/hardware-configuration.nix \
|
|
--target-host root@<zimablade-ip>
|
|
```
|
|
`--extra-files` plants the host key before first boot (its age identity is
|
|
already a recipient in `.sops.yaml`, so `/run/secrets/samba_password`
|
|
decrypts on boot #1). `--generate-hardware-config` pulls the target's real
|
|
kernel modules into the placeholder. Commit the result. Reboot into NixOS.
|
|
|
|
Manual alternative (USB ISO): boot installer, `disko` the disk, then
|
|
`nixos-install --flake .#jupiter`.
|
|
|
|
## Deploy (the `./deploy` wrapper)
|
|
|
|
All arguments mandatory — no default host, no default config.
|
|
|
|
```
|
|
./deploy kexec <host> # headless kexec into a RAM installer (RO-root box)
|
|
./deploy install <config> <host> # first install; wipes OS disk, ships host key
|
|
./deploy switch <config> <host> # rebuild + activate on a running host
|
|
./deploy boot|test <config> <host> # stage for next boot / activate without boot entry
|
|
./deploy image <config> # build an SD-card image (mercury)
|
|
./deploy flash <config> <dev> # build SD image, write it, drop the sops age key
|
|
```
|
|
`switch`/`boot`/`test` prompt for darman's password (`wheelNeedsPassword`).
|
|
`<config>` is a `nixosConfigurations` name (`jupiter`, `neptun`). Its pre-generated
|
|
SSH host key lives at `~/.config/homelab/<config>/ssh_host_ed25519_key`.
|
|
|
|
Examples:
|
|
```
|
|
./deploy switch jupiter jupiter.sol
|
|
./deploy install neptun 159.195.64.117
|
|
```
|
|
Rollback: `nixos-rebuild switch --rollback` on the host, or pick a prior
|
|
generation at boot.
|
|
|
|
## Post-deploy steps (per host)
|
|
|
|
Things the flake cannot do for you. Skipping these leaves a host that builds
|
|
and boots but doesn't work.
|
|
|
|
### Every host, immediately after a first install
|
|
|
|
```
|
|
ssh darman@<host> sudo -v # DO NOT SKIP
|
|
```
|
|
|
|
`users.mutableUsers` is `true`, so `/etc/shadow` is written **once**, when the
|
|
user is created. If the sops secret wasn't readable at that moment the account
|
|
gets `!` (locked) permanently — `deploy switch` will never fix it, because the
|
|
activation script only sets a password for users not already in `/etc/shadow`.
|
|
Combined with `wheelNeedsPassword = true` and `PermitRootLogin = "no"` that
|
|
means no way to escalate, and recovery is physical: netcup's rescue system for
|
|
neptun, or pulling the SD card for mercury. Verify sudo while you still have
|
|
another way in.
|
|
|
|
### neptun (netcup VPS)
|
|
|
|
1. **Edge firewall.** In netcup's panel, inbound `ACCEPT` for TCP 22/80/443/2222
|
|
**and a rule accepting inbound UDP**. The firewall is stateless: without the
|
|
UDP rule every DNS and NTP *reply* is dropped, and nothing on the box reports
|
|
an error — it looks like headscale crash-looping on its DERP fetch and Caddy
|
|
failing ACME. `grep -A1 '^Udp:' /proc/net/snmp` showing `InDatagrams 0` is the
|
|
tell. Rules apply on VM restart, not on save. This is safe: `nixos-fw` is
|
|
stateful and default-deny, so it remains the real policy.
|
|
Also open UDP 3478 (STUN) and 41641 (tailscale direct).
|
|
2. **Authentik** creates `akadmin` on first start; log in at
|
|
`https://auth.mgaction.town` with `authentik_bootstrap_password` from sops.
|
|
The username is hardcoded upstream and the bootstrap runs once — later
|
|
changes to the env vars are ignored.
|
|
To use your own admin instead: create a user, add it to the **`authentik
|
|
Admins`** group (superuser is a *group* flag in Authentik, there is no
|
|
per-user one), verify it works in a private window, then **deactivate**
|
|
`akadmin` — do not rename or delete it. The bootstrap blueprint keys on
|
|
`username: akadmin` with `state: created`, so if no user by that name
|
|
exists it simply makes a new one on the next reconcile.
|
|
3. **Bootstrap the tailnet** (headscale starts with an empty database):
|
|
```
|
|
sudo headscale users create darman
|
|
sudo headscale preauthkeys create --user darman --reusable --expiration 24h
|
|
```
|
|
Put that key in **every** host's sops file as `tailscale_authkey` and rebuild.
|
|
4. **Headplane API key** — defaults to 90d, after which headplane silently stops
|
|
listing nodes:
|
|
```
|
|
sudo headscale apikeys create --expiration 999d # -> headplane_headscale_api_key
|
|
```
|
|
5. **Headplane OIDC.** In Authentik create an OAuth2/OpenID provider
|
|
(confidential, redirect `https://vpn.mgaction.town/admin/oidc/callback`,
|
|
**a signing key must be selected** or discovery exposes no JWKS) and an
|
|
application with slug **`headplane`** — the slug is what makes the issuer
|
|
`.../application/o/headplane/` in `services/vpn/headplane.nix`. Client ID goes
|
|
in that file, client secret into sops.
|
|
6. **Headscale OIDC** (optional — pre-auth keys work without it). A *second*
|
|
Authentik provider/application, slug **`headscale`**, redirect
|
|
`https://vpn.mgaction.town/oidc/callback` (headscale's own, not headplane's
|
|
under `/admin`). Client ID in `services/vpn/headscale.nix`, secret into sops
|
|
as `headscale_oidc_client_secret`.
|
|
⚠️ headscale runs OIDC discovery **at startup and a failure is fatal** —
|
|
an issuer pointing at an application that doesn't exist yet means the
|
|
control server won't boot, taking the whole tailnet's control plane with
|
|
it. Always verify first:
|
|
```
|
|
curl -s https://auth.mgaction.town/application/o/headscale/.well-known/openid-configuration
|
|
```
|
|
Users created by OIDC login are distinct from `headscale users create`
|
|
ones: headplane matches the OIDC `sub` claim against the user's
|
|
`providerId`, CLI-made users have none, and 0.28 dropped both
|
|
`map_legacy_users` and node reassignment — so moving an existing node to
|
|
an OIDC user means re-enrolling it.
|
|
|
|
### jupiter
|
|
|
|
- `chown -R gitea:gitea /mnt/data/AppData/gitea` after the first deploy (the
|
|
repos were copied in over CIFS as `darman:users`).
|
|
- **Re-enrolling after the headscale database was recreated:** `tailscaled`
|
|
keeps its old node key and reports `Running`, and the autoconnect unit exits
|
|
early on that state without ever sending the new pre-auth key. Force it:
|
|
```
|
|
sudo tailscale logout && sudo systemctl restart tailscaled-autoconnect
|
|
```
|
|
|
|
### mercury (Raspberry Pi 3B+)
|
|
|
|
- `./deploy flash mercury /dev/sdX` writes the dedicated age key to the root
|
|
partition. Without `~/.config/homelab/mercury/age.txt` it silently skips that
|
|
step and **no secret decrypts on the box** — check `ls /run/secrets` after
|
|
first boot.
|
|
- It boots from an SD card, so config changes are `./deploy switch mercury <ip>`
|
|
(an aarch64 build — needs `extra-platforms` + binfmt on the laptop, see the
|
|
gotchas in `CLAUDE.md`) rather than a reflash.
|
|
- **Suspect the card first** when binaries crash with `Illegal instruction` or
|
|
services fail inexplicably. Failing flash returns corrupt data with no I/O
|
|
errors in `dmesg`:
|
|
```
|
|
sudo nix-store --verify --check-contents # add --repair to fix
|
|
```
|
|
A card that has corrupted one path will corrupt more. Replace it and reflash.
|
|
- **A reflash wipes `/var/lib/pihole`**, taking the gravity database with it.
|
|
The blocklists themselves are declared in `services/network/pihole.nix`, and
|
|
the `pihole-adlists` unit re-seeds them on boot and rebuilds gravity when it
|
|
finds it empty — so this heals itself, but the first boot after a reflash
|
|
spends several minutes downloading lists. Query history and dynamic DHCP
|
|
leases are genuinely lost (static leases are declarative). Check with:
|
|
```
|
|
systemctl status pihole-adlists
|
|
sudo podman exec pihole pihole-FTL sqlite3 /etc/pihole/gravity.db \
|
|
"SELECT address,enabled FROM adlist; SELECT COUNT(*) FROM gravity;"
|
|
```
|
|
`Blocked DNS queries: 0` in the pihole logs means gravity is empty — DNS
|
|
resolves fine, nothing is filtered.
|
|
- mercury's own `resolv.conf` is deliberately public resolvers, not its own
|
|
pihole (`resolveLocalQueries = false`, see `CLAUDE.md`) — so `.sol` names do
|
|
not resolve *on mercury itself*. That is expected, not a fault.
|
|
- **mercury is load-bearing for the whole tailnet's DNS.** headscale sets
|
|
`override_local_dns = true` with pihole as the only global nameserver, so
|
|
every node — including a phone on mobile data — resolves through it and gets
|
|
ad blocking and `.sol` names anywhere. The flip side is that mercury (or the
|
|
home connection) going down costs name resolution on every device, not just
|
|
`.sol`. Recovery on a stranded device is turning Tailscale off.
|
|
There is deliberately no public fallback in `nameservers.global`: tailscale
|
|
treats that list as a set, so a second entry would let queries slip past the
|
|
filter whenever mercury is slow.
|
|
neptun and mercury opt out with `--accept-dns=false` — mercury because it
|
|
would otherwise resolve through itself, neptun because a public reverse
|
|
proxy must not depend on a Pi at home to renew its certificates.
|
|
- Tailnet names are `*.orbit.sol`, LAN names are `*.sol`. Both work everywhere
|
|
on the tailnet because tailscale matches DNS routes by **longest suffix**, so
|
|
`orbit.sol` reaches MagicDNS even though everything else goes to pihole.
|
|
Never name a LAN host `orbit`: pihole's `address=/<host>.sol/<ip>` lines match
|
|
a name *and everything beneath it*, which would swallow the entire tailnet
|
|
zone.
|
|
|
|
## Adding a service
|
|
|
|
Copy the `whoami` block in `oci-containers.containers`, swap image/ports/volumes.
|
|
Native NixOS module exists for many apps (Nextcloud, Jellyfin, Grafana...) —
|
|
prefer `services.<app>` over a container when available. Add a `caddy`
|
|
`virtualHosts` block to expose it.
|
|
|
|
## Notes
|
|
|
|
- Backend is Podman with `dockerCompat` — `docker` CLI works, no daemon.
|
|
- Samba keeps its own password DB. `services.samba` never sets it; a systemd
|
|
oneshot (`samba-smbpasswd`) provisions it. Host reads the password from
|
|
`/run/secrets/samba_password` (**sops-nix**); the VM falls back to plaintext
|
|
`/etc/samba/smb-password`.
|
|
- Secrets: `secrets/jupiter.yaml` is age-encrypted (safe to commit) to two
|
|
recipients in `.sops.yaml` — the **admin** key (edit on laptop,
|
|
`~/.config/sops/age/keys.txt`) and the **jupiter host** key (derived from its
|
|
SSH host key via `ssh-to-age`, decrypts at runtime). Private keys live
|
|
off-repo and are gitignored. Rotate/add recipients with `sops updatekeys`.
|
|
- Data disk: plain `fileSystems."/mnt/data"` in configuration.nix — kept out of
|
|
disko so it is never formatted. Reference by `by-id` / `by-uuid`.
|
|
- `system.stateVersion` = `26.05`, install-time schema. Do NOT bump on upgrades.
|
|
- Terraform is not used: a single bare-metal box has no provider API. disko +
|
|
nixos-anywhere cover provisioning natively.
|