Everything here is something the flake cannot do for you, and all of it
was learned by hitting it: a host that builds and boots cleanly is not
necessarily a host that works.
The sudo check applies to every host and is the one that cost the most.
mutableUsers is true, so /etc/shadow is written once at user creation --
if the sops secret wasn't readable at that moment the account is locked
forever and no rebuild will fix it. That happened twice, and recovery was
netcup's rescue system for neptun and pulling the SD card for mercury.
neptun's netcup firewall is stateless and denies inbound UDP by default,
which drops every DNS and NTP reply while reporting nothing anywhere.
Also covers the Authentik/headscale/headplane bootstrap, which is a
chain of manual steps producing values the config needs.
jupiter gets the tailscaled stale-state trap: after the headscale
database is recreated the daemon still reports Running, and the
autoconnect unit exits early without sending the new pre-auth key.
mercury gets the SD-card failure mode, since silent flash corruption
surfaces as SIGILL from random binaries with a clean dmesg.
Also drops a stray code fence that had been dangling at EOF.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The Authentik provider and application now exist (slug "headplane", which
is what makes the issuer .../application/o/headplane/), so the client ID
is a real value rather than a placeholder, and the client secret and
headscale API key are in sops.
The tailscale pre-auth keys for neptun and jupiter are rotated because
the tailnet was recreated from scratch: the old headscale database went
with the VPS's OS disk, so every key issued against it is meaningless to
the new control server.
Note the headscale API key defaults to a 90d expiry. When it lapses
headplane stops listing nodes with no obvious cause -- `headscale apikeys
list` shows the date.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
mercury was the only host with no tailscale at all -- no module import,
no secret, no key in its sops file. It had been enrolled before the NixOS
migration and silently dropped off the tailnet when it was reflashed with
a config that omitted it.
--accept-dns=false, as on neptun and for a sharper reason: headscale
pushes override_local_dns, so accepting MagicDNS would repoint the LAN's
own DNS server at 100.100.100.100 and make house-wide name resolution
depend on tailscaled being up. This host has already deadlocked once on
boot-time DNS (see CLAUDE.md).
darman_password is also rotated: the account had "!" in /etc/shadow,
because on mercury's first boot the secret wasn't readable yet and
update-users-groups.pl falls back to a locked account. mutableUsers is
true, so no later rebuild ever revisited it and the lock was permanent.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
common.nix now sets security.sudo.wheelNeedsPassword = true, but
--use-remote-sudo is deprecated and only prefixes commands with sudo --
it never prompts, so every remote rebuild failed. --ask-sudo-password is
the alias for --elevate=sudo --ask-elevate-password, which asks once and
feeds it via sudo --stdin.
This should have gone in with the wheelNeedsPassword change itself.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
hardware-configuration.nix as regenerated by nixos-anywhere during the
install, replacing the placeholder. The detected initrd modules differ
from what the placeholder guessed (ata_piix, uhci_hcd), but the virtio
modules pinned in configuration.nix merge in regardless, so root mounts
either way.
darman_password is rotated because the previous hash's plaintext was not
recorded anywhere. Combined with wheelNeedsPassword = true and
PermitRootLogin = "no" that left no way to escalate on the box, and
recovery needed netcup's rescue system to edit /etc/shadow directly.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
By default headscale fetches https://controlplane.tailscale.com/derpmap/default
at startup and treats failure as fatal, so it cannot boot when that URL is
unreachable. A self-hosted control plane that will not start without
Tailscale's infrastructure rather misses the point of self-hosting -- and
it crash-looped for exactly that reason while neptun had no DNS.
Enable the embedded DERP server on region 999 and drop the upstream map.
The relay rides Caddy on :443, which is why that vhost already sets
flush_interval -1; only STUN needs a port of its own.
Verified against headscale 0.28.0 before committing: it starts clean with
urls = [], registers "DERP region: {RegionID:999 ...}" pointing at
vpn.mgaction.town with DERPPort 443, and brings up STUN.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Probing the live Debian VPS turned up three mismatches between what it
serves and what this config declares:
- git.mgaction.town had no vhost at all. Gitea's web UI and HTTPS clones
are public today; only its SSH side (the :2222 socat forward) had been
ported, so a deploy would have taken the web side offline.
- Audiobookshelf is served as abs.mgaction.town, not the longer
audiobookshelf.mgaction.town this config used. The mobile app is
configured with the short name.
- The apex returns 200 from Caddy. Left unserved deliberately, so it now
gets Caddy's default 404; noted in a comment so it doesn't look like an
oversight next time.
Gitea's ROOT_URL was http:// while Caddy terminates TLS for that name.
Gitea builds absolute URLs from it, so clone buttons, redirects and
webhooks were handing out downgraded links.
Also record that defaultGateway6 is confirmed rather than assumed --
`ip -6 route show default` on the VPS gives "default via fe80::1 dev
eth0 metric 1024 onlink".
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
nixpkgs only carries Zitadel 2.71, which predates the login-v2 split and
cannot take a v3/v4 database (its migrations are forward-only), so the
instance running on the old Debian VPS could never have moved onto it.
authentik-nix ships 2026.5.4 and tracks upstream closely.
The authentik-nix input deliberately does not follow our nixpkgs, per
upstream's warning that overriding it breaks their pinned python
dependency set. That costs a second nixpkgs in the lock, so add
nix-community's Cachix to common.nix -- without it the closure is ~400
local derivations (npm, rust, python). The laptop that runs
scripts/deploy needs the same two lines in /etc/nix/nix.custom.conf.
Authentik's own module creates the database and orders its units against
postgresql.target, and recent versions need no redis, so the wiring is
just the module plus a secret. Pin postgresql explicitly so that editing
system.stateVersion can never silently demand a pg_upgrade of the
identity store.
Secret ownership is not uniform and the difference matters: authentik
and caddy take a systemd EnvironmentFile, which PID 1 reads as root
before dropping privileges, so root:root 0400 is correct. Headplane
opens its secret paths itself while already running as the headscale
user, so those three need an explicit owner or they fail to start.
Also on neptun:
- Pass Caddy's ACME account email through the same EnvironmentFile
mechanism and reference it with the Caddyfile {$VAR} placeholder.
services.caddy.email would render the address into the world-readable
store.
- Stop accepting MagicDNS from our own control server. headscale pushes
override_local_dns, so joining the tailnet would point neptun's
resolv.conf at a MagicDNS served by the tailscaled neptun itself hosts
-- a tailscaled failure would then also take out DNS, ACME renewal and
finally the certs for the control server every other node needs in
order to recover.
- Give headplane a writable DNS extra-records file. Its view of
headscale's config stays read-only, which is the right outcome for a
declarative box; records are data rather than config.
- Require a password for sudo. Deploys become interactive, but darman's
key is otherwise the only thing between the public internet and root.
- Enable zram (8 GB, and disko leaves no room for a swap device), and let
tailscaled-autoconnect retry instead of failing permanently when the
control server isn't up yet on a first boot.
networking.hosts still carries a PLACEHOLDER address for jupiter --
replace it from `headscale nodes list` once jupiter first enrols.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Group service modules by category (media, network, vpn, identity,
dev, desktop) to make the growing services/ dir easier to navigate.
containers.nix stays at the top level since it's a shared backend,
not a single-category service.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Shorten verbose multi-paragraph comments to essentials, and drop a
stale claim in common.nix that jupiter kept its own copy of the base
config (it now imports common.nix directly).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Path-route Headplane under /admin on the same vhost as headscale instead of
its own subdomain - Caddy handle blocks split on the prefix, headscale gets
everything else. base_url drops to the site root since Headplane appends
/admin (and the OIDC callback path) itself.
Wire Zitadel as the OIDC provider. client_id/client_secret/the headscale
API key can't be real until Zitadel and headscale are actually deployed and
an application/key exist, so those are REPLACE_ME placeholders for now
(documented in services/headplane.nix) - direct API-key login stays enabled
as a fallback so this can't lock anyone out in the meantime.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Headscale is the tailnet control server every host's services/tailscale.nix
already points at (--login-server=https://vpn.mgaction.town). MagicDNS
base_domain "hosts.mgaction.town" matches the "jupiter.hosts.mgaction.town"
names already used in this repo's Caddy vhosts.
Headplane is its web UI, running as headscale's own user (native process
integration, no container). No OIDC wired up - log in with a headscale API
key generated on the box. Both proxied through Caddy; headscale's vhost
needs flush_interval -1 since its node-update endpoint is a long-poll.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Local Postgres, peer-authed over the unix socket (the "zitadel" role is
granted createdb+createrole and doubles as both the runtime and bootstrap
DB user - no password anywhere). TLS terminates at Caddy; Zitadel listens
on localhost:8080 and is proxied at auth.mgaction.town.
Master key and admin bootstrap password come from sops - the admin
password specifically needs the sops.templates -> rendered-file route
(services.zitadel.steps would leak it into the world-readable Nix store),
same pattern as mercury's pihole.env.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Caddy only proxies HTTP; git-over-ssh to gitea needs a raw TCP forward
since gitea's built-in SSH server (jupiter:2222) isn't otherwise reachable
from the public internet. socat forwards the VPS's public :2222 over the
tailnet. Matches what's now live on the (still-Debian) VPS - ready to drop
in once neptun gets migrated to this NixOS config.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
- sabnzbd, prowlarr, sonarr, radarr, clonarr, seerr, cinephage, mediamanager
services, wired into jupiter with LAN Caddy vhosts.
- Gitea: migrated the old ZimaOS docker instance's data (sqlite db, 4 repos,
no LFS objects) into the NixOS module's default stateDir layout. HTTP via
Caddy; git SSH on its own built-in server at :2222 (not :222 - the unpriv
gitea user can't bind <1024).
- mediamanager-nix flake input for the mediamanager service.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
- services/pihole.nix: official pihole/pihole:2026.07.2 via podman, host net,
caps NET_ADMIN/NET_RAW/SYS_NICE/CHOWN; FTLCONF_* env config (upstream unbound,
DHCP 50-200, static lease jupiter, .sol domain, local records)
- unbound: resolveLocalQueries=false (was hijacking resolv.conf to :53 -> boot
DNS deadlock; the real root cause of the earlier failures too)
- password via sops FTLCONF env file; /var/lib/pihole created via tmpfiles
- VM-verified: mercury.sol/jupiter.sol/external all resolve, 0 restarts
mercury is the DHCP server (no lease of its own) so its name wasn't resolvable.
Add explicit dns.hosts A records. VM-verified: both resolve, external unaffected.
Pi's vfat partition isn't mounted at runtime (u-boot reads it pre-boot), so
/boot/firmware doesn't exist -> keyFile moved to the always-mounted root fs.
deploy flash now drops it on the ext4 root partition.
pihole.toml is nix-managed read-only so 'pihole setpassword' fails. pihole-FTL
reads FTLCONF_* env vars (override the toml) — render an env file from the sops
secret pihole_webpassword and feed it via EnvironmentFile. Password stays out of
repo/store. VERIFIED in VM: FTLCONF_webserver_api_password -> API login works.
- after dd, if ~/.config/homelab/<config>/age.txt exists, mount the FAT boot
partition and drop it as sops-age.txt (mercury). Key stays off-repo + out of
the store + out of the image; no manual mount step.
- per-host darman_password (distinct hash) in secrets/{jupiter,vps,mercury}.yaml
-> hashedPasswordFile; different console password per host (ssh still key-only)
- mercury: dedicated age key (on boot partition post-flash), sops-nix wired
- AdGuard: module has no secret hook + writable config -> mutableSettings=true,
admin password set via web setup on first boot (never in repo/store)
- nixosConfigurations.mercury: aarch64, sd-image-aarch64, imports common.nix
- static net placeholders (CHANGE-ME), hostname mercury
- DNS service left undecided: commented pihole (services.pihole-ftl) + adguardhome
- no sops yet (add with the service if it needs a secret)
- services/{samba,avahi,audiobookshelf,containers,caddy,tailscale}.nix
- common.nix grows firewall base + timezone; hosts import what they need
- jupiter/vm/vps import service modules; drop the jupiter/services.nix monolith
- each module opens its own firewall ports; caddy/tailscale shared by hosts
- verified: jupiter/vps/vbox eval + jupiter builds, config equivalent
- services.tailscale auto-registers with headscale using a sops pre-auth key
- trust tailscale0 so LAN services are reachable over the tailnet
- fix missing semicolon on audiobookshelf extraGroups
Installed initrd lacked mmc_block -> can't mount root on eMMC -> no boot.
generate-config runs in the RAM installer and misses mmc modules, and it
overwrites hardware-configuration.nix, so pin them where they merge instead.
kexec/run rebuilds an initrd via cpio|gzip from PATH — absent on ZimaOS, so run
aborted before kexec and the box stayed on ZimaOS. Ship GNU static cpio + busybox
(as gzip) into /tmp/bin, prepend PATH. SSH ControlMaster = one password prompt.
- add nixos-images input; nixosConfigurations.kexec bakes in the ssh login key
- build via config.system.build.kexecInstallerTarball
- deploy: ./deploy kexec <host> streams the installer to /tmp and kexecs
- works around ZimaOS RO root where nixos-anywhere ssh-copy-id fails
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>