Two quick-reference documents, written to be scanned rather than read:
- USER-GUIDE.md for people new to IRC -- connecting, claiming a nickname,
the dozen commands that matter, scrollback, and a plain-language privacy
section (cloaked IP, channels logged for the configured retention, how to
turn off DM storage).
- ADMIN-CHEATSHEET.md split by where you work: from IRC as an operator
(UBAN, KILL, DEFCON, ChanServ, NickServ) and on the host via ergoctl, plus
mode tables and a "when things go wrong" section.
Both use the same placeholder convention as ergo.motd and are rendered by
deploy.sh into $STACK_DIR/docs/ with the network's real name, domain and
retention, so they can be handed straight to users and moderators. Refreshed
on every run, like the other installed files.
Command and mode references were checked against the v2.19.1 sources rather
than written from memory: irc/modes/modes.go for every mode letter, and
irc/chanserv.go and irc/nickserv.go for the service subcommands and which
require an oper capability.
build.sh's embed guard caught the docs being added to FILES without being
added to deploy.sh's EMBEDDED manifest -- the exact failure it was added to
prevent.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Replaces the placeholder MOTD with a usable one: a coloured header, five house
rules, the NickServ/ChanServ/SASL commands people actually need, and a short
"worth knowing" section.
The retention line is generated rather than hard-coded, so it matches the
deployment instead of drifting from it: HISTORY=off says messages are not
stored, otherwise it names the real HISTORY_EXPIRE. Users are told plainly that
channels are logged and for how long, which is the honest counterpart to
enabling persistent history by default.
Rendered and checked: valid UTF-8 (an invalid byte makes the whole config fail
to load), ASCII-only for old clients, no stray $-escapes, only documented colour
names, and 72 columns at the widest.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Seven confirmed findings from an observed-vs-expected review of a real deploy
transcript against the code (four other proposed findings were refuted and no
change was made for them).
- Caddyfile reformatted so `caddy fmt` is clean, removing the warning Caddy
printed on every validate and every start. Two causes, not one: `caddy fmt`
indents with TABS, and it deletes a blank line whose following line begins
with `{` -- which is why the warning pointed at line 17, the blank before the
global options block. The check is a whole-file byte comparison, so the line
number was only the first difference and the entire file had to be
reformatted. Verified whitespace-only outside the heredoc, whose body is left
byte-identical: Caddy strips padding derived from the closing marker's
indentation, so re-indenting it would change what the page serves.
- deploy.sh no longer upgrades Caddy behind the operator's back. A bare
`docker compose pull` refreshed the floating `caddy:2-alpine` on every re-run
and `up -d` then recreated it, doing exactly what CADDY_AUTOUPDATE=0 promises
not to, with none of update.sh's health check or rollback. Pulls are now
per service: the pinned Ergo tag always, Caddy only when absent or opted in.
- deploy.sh seeds the ACME_EMAIL and NETWORK_NAME prompts from .env, so pressing
Enter through a re-run no longer renames the network to the hostname in the
summary while .env keeps the real one.
- The Caddy restart guard compares the container's identity across `up -d`. A
container compose created or recreated has already read the new Caddyfile;
only one left running still holds the old config, and `svc_state` cannot tell
those apart.
- Container logs are now rotated (json-file, 10m x 3) on every service. Alpine's
docker package ships no daemon.json, so the default is unbounded, and the
60-second health probe alone writes a log line per run.
- ergoctl notes that LUSERS counts its own probe connection, which is why a
server with nobody on it reports one invisible user.
Verified: seven local suites pass, the embedded archive round-trips, line
endings are LF, and the reformatted Caddyfile is a verified fixed point of the
formatter (semantically identical token-for-token to the previous one).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Adds client-certificate authentication for operators, so an oper can be
identified by the SHA-256 fingerprint of their TLS client certificate instead of
a typed password:
ergoctl oper certfp <name> <fingerprint|nick> [--auto] [--clear]
The fingerprint can be given directly or read off a connected user over IRC
(276 RPL_WHOISCERTFP). Without --auto Ergo requires BOTH the certificate and the
password -- a second factor; with --auto the password is removed and the
certificate alone grants oper on connect.
It refuses to do this to the 'admin' oper, which matters: ergoctl authenticates
as admin over the loopback PLAINTEXT listener, which presents no client
certificate, and Ergo requires a configured certfp to match. Setting one there
would lock ergoctl -- and the scheduled jobs that use it -- out of the server
permanently.
Polish from the first real deploy on irc-1.srvno.de:
- deploy.sh no longer restarts Caddy on a first deploy. The Caddyfile compare
treated "destination does not exist" as a change, so Caddy was recreated
seconds after starting, while the initial ACME order was in flight.
- ergoctl status no longer prints the raw leading parameters of the 265/266
LUSERS numerics ("1 1 Current local users 1, max 1"), which duplicate the
counts already in the text. 252/254 keep theirs, where it is the only count.
- ergoctl oper list now shows which opers have a certfp and which are auto.
Verified: seven local suites pass, including new coverage for fingerprint
normalisation (colons and uppercase), insertion into the correct oper block
without touching siblings, replace-not-duplicate, the --auto and --clear paths,
276 parsing, and the admin guard.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Ergo keeps history in RAM by default, so it is lost on every restart --
including the ones this stack's updater performs. HISTORY now selects a backend
at first deploy: sqlite (default, a file beside ircd.db, no extra container),
postgres (a pinned container via docker-compose.postgres.yml, loopback-only,
POSTGRES_PASSWORD_FILE so no secret lands in .env), or off. Both SQL backends
need Ergo 2.18.0+, which deploy.sh enforces. HISTORY_EXPIRE (default 30d) sets
retention, because upstream's 1w expire-time DELETES from persistent storage --
persistence with the shipped default would buy only a week.
Ergo opens the history backend only at startup and, unlike MySQL, has no
"after launching the server" guard for sqlite/postgresql: a rehash that enables
one reports success and then silently discards every message. So `ergoctl
history` restarts, and `ergoctl edit` detects a backend change and refuses to
apply it by rehash. The config validator runs with no network by design, so it
neutralises postgresql in its copy and reachability is checked separately.
Audit fixes (six lenses over first-deploy, runtime, lifecycle, security, docs):
- SECURITY (blocker): ircd/ is writable by the container uid while the 15-minute
cert sync and every ergoctl config edit run as root. cp/install/> follow a
symlink, so code execution in Ergo could redirect a root write onto any host
file. All such writes now stage under $STACK_DIR (0700 root) and land via mv
(rename(2) replaces a symlink instead of following it); root reads use cp -P
or refuse. Regression-tested with real symlinks.
- deploy.sh no longer swaps the image when ERGO_TAG changes on a deployed stack
(that bypassed pre-flight, snapshot, user warning and rollback); it points at
`TARGET_VERSION=... ergoctl update update`.
- restore was lossy and could lock you out: it now snapshots the current config
as well as the database, restores message history, re-points the postgres
password and re-hashes the local admin password into the restored config, and
validates before starting.
- `install -d -m` re-modes existing directories: `ergoctl backup /var/backups`
no longer chmods it 0700 root, and the updater no longer re-modes /var/log
(0775 root:syslog on Debian, which rsyslog needs).
- The admin oper password is printed only on a first run at a TTY, so it stays
out of cloud-init serial-console logs.
- A failed update is remembered, so the daily job stops repeating a disruptive
warn/stop/swap/roll-back cycle every night; postgres readiness gates an update
that would otherwise stop a healthy server it cannot restart.
- certsync no longer sends "TLS recovered" for outcomes that synced nothing.
- ergoctl history writes .env only after the restart is healthy, and returns 0.
- CR/LF is stripped before IRC framing, so a multi-line argument cannot inject a
second command; the ntfy token moves out of curl's argv.
- ufw/firewalld are additive, so 6667 is now explicitly revoked when PLAINTEXT=0.
- build.sh refuses to build a deploy.sh whose archive is missing a file the
script reads -- the failure mode that would have shipped a stack aborting on
every host.
Docs corrected against the code throughout, including retention, the pre-connect
account-registration default, encrypted-restore (AGE_IDENTITY), what a re-run
really does to .env, and what the update log does and does not contain.
Verified locally: all six suites pass (config render for each backend against the
real 2.19.1 template, yaml/oper/version/env helpers, the IRC client against a
fake server, and the audit fixes including the symlink escalation). Still not
exercised on a Docker host: the containers, ACME issuance, cert sync and
PostgreSQL itself.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
New deployments/ergo/: the Ergo IRC server behind Caddy for Let's Encrypt TLS
and the IRC-over-WebSocket endpoint. Docker rather than a native OpenRC/systemd
service, because Alpine's apk ergo trails upstream (3.24 ships 2.18.0 against a
2.19.1 security release) and Debian/Alma package it at all -- so native would
mean three install paths plus a per-distro ACME client.
Both containers run with network_mode: host. IRC bans, throttling and cloaking
key on the client's address, and Docker's userland proxy would hide every IPv6
client behind the bridge gateway; host mode also makes the repo's INPUT firewall
genuinely govern 80/443/6697. Caddy reaches Ergo over loopback, which is what
lets Ergo honour X-Forwarded-For (proxy-allowed-from defaults to localhost) and
mark web sessions secure.
- deploy.sh generates ircd.yaml ONCE from the pulled image's own default.yaml
(version-matched), rewriting the listeners/websockets blocks wholesale rather
than patching lines, then asserts hard post-conditions and validates with
`ergo run --smoke` in a throwaway container before anything starts.
- update.sh: pinned vX.Y.Z tags, GHSA + "### Security" release-note policies,
pre-flight against the new image, user NOTICE + grace, stop-consistent DB
snapshot, health check (IRC-level, not a bare TCP connect) and rollback that
restores the DB only when the schema actually moved. Compatibility-break
releases are held for review. certsync copies Caddy's cert pairwise-atomically
and verifies the fingerprint served on 6697 after the rehash.
- ergoctl: status/users/logs, validated edit+rehash, oper add/passwd/rm,
moderation, backup/restore, cert and update passthrough. Talks IRC to the
loopback listener over bash /dev/tcp and strips control characters from
replies.
- Ergo runs as a non-root system user, read-only rootfs, all caps dropped;
Caddy keeps only NET_BIND_SERVICE, with admin API and HTTP/3 off.
Reviewed adversarially across six lenses; 20 confirmed findings fixed, notably
a dead SIGHUP fallback (`rc=$?` after an `if` is always 0), several `set -e`
aborts from non-total pipelines, a release-list cache that only ever populated
in a subshell, and re-runs that used shell defaults instead of the deployed
.env. Verified locally: bash -n, LF endings, the ircd.yaml render against the
real 2.19.1 template in both PLAINTEXT modes, the yaml/oper/version/env helpers,
and the IRC client against a fake server (registration, oper, rehash success and
400-failure, control-character stripping, server-down paths). Not yet exercised
on a Docker host: the containers themselves, ACME issuance and cert sync.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Two Alpine-only failures from an irc-1 run.
1. harden-firewall.sh died with
line 161: /usr/local/sbin/firewall-apply: No such file or directory
and left INPUT unfiltered. Alpine's baselayout ships /usr/local/{bin,lib,
share} but no sbin, and nothing created it. Create it before writing the
engine; same guard in oslib's install_boot_hook / install_daily_job (which
`install` into /usr/local/sbin, and `install` does not make parent dirs) and
for /usr/local/bin in harden-jumphost.
2. The ntfy login notifier never fired despite reporting ACTIVE. Alpine keeps
PAM support in a SEPARATE binary: openssh-server gives /usr/sbin/sshd (no
PAM), openssh-server-pam gives /usr/sbin/sshd.pam. The OpenRC init only
picks the PAM one in start_pre (checkconfig -> update_command), and its
reload/stop match the process with `start-stop-daemon --exec "$command"`.
So a host already running /usr/sbin/sshd never swaps: reload signals a
non-matching process, stop matches nothing, a later start hits "address
already in use". sshd keeps serving without PAM, so /etc/pam.d/sshd -- and
the pam_exec hook in it -- is never consulted.
Add oslib sshd_wanted_binary / sshd_running_binary / sshd_apply_config, and
end the harden scripts with sshd_apply_config instead of svc_reload: reload
as before, but on Alpine stop by pidfile and start again when the running
binary isn't the one the config calls for. Established sessions are separate
processes and survive; only the listener blinks.
install_login_notifier now checks that PAM can actually be reached on Alpine
(sshd.pam present, UsePAM yes) and warns with the fix instead of reporting a
hook that can never run as ACTIVE -- the same "silent notifier" trap the
Alma run hit.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
automations.sh was #!/usr/bin/env bash and written in bash, but a stock Alpine
box has busybox ash and no bash at all -- so the launcher died before it could
install anything:
./automations.sh -> env: 'bash': No such file or directory
curl ... | bash -> bash: not found
sh automations.sh -> syntax errors
cloud-init/base.yml already installs bash/git/curl before touching the repo;
the launcher never got the same treatment.
Give it a #!/bin/sh shebang and a strictly-POSIX prologue that ash can parse:
self-locate via ${0%/*} (busybox dirname takes `--` as the filename, so
`dirname -- "$0"` would answer "."), clone on the piped path as before, install
git/bash via apk/apt-get/dnf/yum, then exec bash on this same file. The bash
pass skips the prologue via BASH_VERSION, and an exported _AUTOMATIONS_REEXEC
guard rules out an exec loop. Everything below the handoff is the unchanged
bash launcher.
The generated bundle stub had the identical bug -- it's the other fresh-host
entry point -- so make it POSIX sh too and have it install bash before
extracting.
Docs: the one-liner now pipes to `sh`, bundle examples use `sh`.
Still bash-only on a bare Alpine host: the scp'd deployments/*/deploy.sh
artifacts.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
An adversarial re-verification of 0812f34 found a couple of the fixes were
incomplete or regressed; addressed here.
- DR was an INCOMPLETE fix: `raft snapshot save`/`restore` are token-gated
(sys/storage/raft/snapshot is sudo-capable) and the container carries no
ambient token, so the previously-"fixed" backup returned "missing client
token" and streamed a zero-byte snapshot. Both the deploy.sh runbook and the
README DR flow now pass `-e BAO_TOKEN=<token>` on save and restore.
- IPv6 bind REGRESSION (introduced by the bind-narrowing): a bare IPv6 literal
in the compose port map (`fd00::10:8200:8200`) is invalid and aborts at
`docker compose pull`. Now IPv4 vs IPv6 are classified separately and IPv6 is
bracketed (`[fd00::10]:8200:8200`).
- Docker readiness race (now reachable since install_docker is actually
called): openrc backgrounds dockerd and returns before the socket is up, so
the next `docker compose pull` raced it under set -e. install_docker now polls
`docker info` for up to 30s.
- UID detection hardened: added `-T` to the one-off `docker compose run`, and an
empty result is now a loud warning (with the manual-chown remedy) instead of a
silent fall-through to root that would re-create the crash-loop on a non-root
image.
- The raft-volume chown is gated to first run (captured before any compose-run
instantiates the volume), so idempotent re-deploys don't recursively re-chown
a live raft dir.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
A multi-agent sanity audit of the freshly-merged deployment found four
end-to-end blockers (and several smaller issues); all fixed here.
HIGH (were blocking):
- deploy.sh never called install_docker(), so `docker compose pull` hit
command-not-found on any host without Docker. Now called before the
compose steps.
- The container's server process runs as the image's own (often non-root)
user but the mounted config/TLS were root-owned 0640/0600 and the raft
volume root-owned -> vault crash-looped, never binding :8200. deploy.sh
now detects the image UID after pull and aligns ownership of config.hcl,
./tls and the data volume (a no-op when the image runs as root);
config.hcl is installed 0644 (holds no secrets).
- Docs told operators to set the daemon key `openbao_ca_cert`, but the
Kanrisha daemon's key is `ca_cert` (config.go, mapstructure:"ca_cert").
The wrong key is fatal on strict unmarshal / leaves TLS unverified.
Renamed in all 5 places (config.hcl, gen-tls.sh, deploy.sh x2, README).
- DR backup used `docker compose cp openbao:… -`, which emits a TAR stream,
so the age-encrypted snapshot was tar-wrapped and would not restore.
Switched to `docker compose exec -T openbao cat` for the raw bytes, wrote
the snapshot to a scratch path (not the live raft dir), and documented the
matching restore.
MEDIUM:
- Swap detection used `swapon --show` (absent on BusyBox) and `\s` (GNU-only)
-> silently no-op on Alpine, leaving swap on. Now uses /proc/swaps and
[[:space:]] so mlock hardening actually holds on musl.
- A Docker-published port bypasses the host INPUT firewall, so the source
rule was illusory. deploy.sh now narrows OPENBAO_BIND to OPENBAO_ADDR when
it is an IP, the compose/README/.env comments state the reality, and a new
Exposure section + an init-immediately warning were added.
- Fixed broken ../kanrisha/ and deployments/kanrisha/ links (separate repo).
LOW:
- OPENBAO_TLS_SANS is now honored (folded into the SAN list from the env).
- .gitignore excludes *.snap / *.snap.age.
- Bootstrap note clarifies bootstrap.sh needs the `bao` CLI (run it from the
Kanrisha host/workstation, not this Docker-only vault host).
- README multi-OS count corrected (eight stacks) + automations.sh header
lists openbao.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
A dedicated OpenBao deployment, kept OFF the Kanrisha tape host so a compromise
of the tape node can't reach the vault.
- Native TLS on the listener (self-signed by default via gen-tls.sh, or a
CA-signed cert from a Smallstep CA over ACME) — no Caddy/Let's Encrypt;
reached over the LAN, not the public internet.
- Integrated raft storage (clean snapshot-based DR).
- mlock on (cap_add IPC_LOCK + memlock unlimited + host swapoff in deploy.sh).
- Manual unseal by default; optional PKCS#11 HSM auto-unseal.
- deploy.sh: Docker install (Alpine/Debian/Alma), self-signed cert, .env seed,
swapoff, firewall 8200/tcp, compose up; then prints init/unseal + the KV-v2 +
AppRole bootstrap for Kanrisha + the raft-snapshot DR flow. Self-contained
(config payload embedded by build.sh).
- Registered in automations.sh + the README deployment table.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
New deployments/copyparty/: copyparty (copyparty/ac) behind Caddy/LE for the
web UI/WebDAV, plus its own SFTP (password auth) and FTPS listeners published
directly. Ships update.sh, which drives container updates off copyparty's
security-advisories API (api.copyparty.eu/advisories) -- policies latest|security|off.
- Real client IP end-to-end: Caddy XFF/X-Real-IP + copyparty xff-src: lan.
- SFTP host key + self-signed FTPS cert generated/persisted in /cfg; admin
password generated on first deploy; conf auto-included via the image's % /cfg.
- Firewall opens 80/443 + SFTP/FTPS + passive range (colon form for ports.d).
- Wired into automations.sh, README, .gitignore; cloud-init for fresh VMs.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
New deployments/squid/: an explicit forward proxy with SSL-bump TLS
interception (local CA, generated on first deploy) and hostname-targeted
static-content caching. Unlike the other stacks it is a forward proxy, not
a Caddy/Let's-Encrypt inbound site.
- Self-built minimal Alpine image (apk squid ships ssl-bump); entrypoint
renders squid.conf and generates the cache policy from the domain lists.
- Wildcard hostname caching (cache-domains.txt leading-dot + optional
cache-domains.regex); boost vs strict-allowlist toggle (CACHE_ONLY_LISTED).
- Storage gate never caches HTML or dynamic content; query strings exempt on
boosted domains so versioned static assets still cache.
- splice-domains.txt passthrough for pinned/banking domains.
- Deny-by-default http_access (TRUSTED_CIDR) + BIND_ADDR pinning; CA key 0600
on host, never embedded, git-ignored.
- Wired into automations.sh, README, .gitignore; cloud-init for fresh VMs.
A fresh Alma box has firewalld active, and the iptables-based harden-firewall.sh
refused to run there (caught by harden-ssh's '|| warn', so the host firewall was
silently skipped). Use firewalld natively on the rhel family instead of fighting it.
- harden-firewall.sh: family-aware backend. On rhel, apply/allow/deny/list/disable
drive firewall-cmd (deny-by-default zone, SSH + registered ports, ping policy,
source-restricted rich rules); Alpine/Debian keep the iptables engine unchanged.
FW_BACKEND=iptables|firewalld overrides.
- oslib: install_firewalld(); sshguard_backend() prefers sshg-fw-firewalld on rhel
so brute-force blocks land in firewalld (no INPUT->sshguard jump needed).
- Deployments already fall through to a firewall-cmd branch when the iptables
engine is absent, so they need no changes.
- README + script header document the per-family backend.
harden-ssh / harden-jumphost are unchanged -- they call harden-firewall.sh apply
and read sshguard_backend(), so the switch happens underneath them.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
install_openssh ended with '[[ "$OS_FAMILY" == alpine ]] && pkg_install ...'.
As the function's LAST statement, that trailing test returns 1 on every
non-Alpine OS (a false '[[ ]]' exits 1), so the function reported failure even
when the packages installed fine. Harmless while the call was bare under set -e
(a short-circuited && is exempt), but the new 'install_openssh || die' guard
read it as a real failure and aborted harden-ssh on Alma right after
'Installing OpenSSH server...'.
Fix: convert the Alpine-only linux-pam step to an if-block, and add '|| return 1'
to the main install so a genuine package failure still propagates honestly.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
On a fresh AlmaLinux 9.8 box, install_bruteforce_protection ran unguarded under
'set -euo pipefail'. When sshguard (from EPEL) wasn't installable at that moment,
the single failed dnf aborted the ENTIRE harden run before it wrote sshd_config
or installed the pam_exec login notifier -- leaving a stock, unhardened box and a
silently-missing ntfy hook.
- oslib: install the iptables backend best-effort first, then sshguard, and
return sshguard's status so callers can treat it as non-fatal.
- harden-ssh/harden-jumphost: install_openssh now dies with a clear message on
failure; sshguard is '|| warn' so sshd hardening and the notifier still apply.
- install_login_notifier verifies the script + pam hook landed and logs
'Login notifier ACTIVE' (or a loud warning) instead of failing silently.
- ntfy-ssh-login.sh: NTFY_DEBUG=1 logs delivery attempts + curl errors to
/var/log/ssh-notify.log so the next silent failure leaves a trace.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The one-liner bootstrap only auto-installed git on Alpine (apk), so piping it
onto a fresh Debian/Alma host with no git fell straight into 'git clone' and
died with 'git: command not found'. oslib's pkg_install can't help here -- the
repo isn't on disk yet. Install git inline via apk/apt-get/dnf/yum, and fail
with a clear message if it still can't.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
base.yml set AUTO_UPDATE=1 but never forwarded it to harden-ssh.sh (which
defaults to 0), so base hosts never actually scheduled daily updates despite
the config implying they did. Auto-update is intentionally bastion-only for
now — jumphost.yml / harden-jumphost.sh enable it by default, and base
(container) hosts will get a dedicated docker-image updater later — so remove
the knob entirely rather than wire it through. No behavior change.
Also drop cloud-init/base.yml from the README "Daily updates" line that
claimed it takes AUTO_UPDATE=1.
Add a reusable iptables baseline that hardens hosts with ICMP + SSH
defaults and lets deployments register the ports they need. INPUT is
deny-by-default (loopback, established, ICMP, SSH on the configured port,
plus registered ports); OUTPUT stays open and FORWARD is left untouched so
Docker container networking is unaffected.
Persistence is native -- no boot hook. Rules are saved and restored by the
distro's own package (iptables/ip6tables on Alpine, iptables-persistent on
Debian, iptables-services on Alma) via the new oslib helpers
install_iptables / fw_save_cmd / fw_enable_restore. The saved ruleset
carries the INPUT->sshguard jump, so brute-force protection survives reboot
without the old sshguard-iptables hook.
A self-contained /usr/local/sbin/firewall-apply rebuilds INPUT from
declarative drop-ins under /etc/firewall/ports.d and runs the native save,
so deployments add a port without needing the repo present:
printf '80/tcp\n443/tcp\n' > /etc/firewall/ports.d/mystack.rule
/usr/local/sbin/firewall-apply
- SSH port read live from sshd_config (custom bastion ports just work);
FW_SSH_SOURCE restricts the source CIDR; FW_ALLOW_PING gates echo
- harden-ssh.sh / harden-jumphost.sh install it when ENABLE_FIREWALL=1
(default) and skip the sshguard-only hook; ENABLE_FIREWALL=0 keeps it
- cloud-init base.yml / jumphost.yml forward the toggle
- the four stack deploy.sh open_web_ports() register 80/443 via the
firewall (ufw/firewalld kept as fallback); Docker-published ports bypass
INPUT, so this is belt-and-braces and self-documenting
- README + cloud-init/README document the mechanism, Docker caveat, and the
`disable` recovery path
Ship allow-everything as the active default (first acl rule), with the
Tailscale default policy translated to headscale's acls format included but
commented for when you lock down. Pre-fix the headscale gotchas: tag:shared
owned by group:admins (not an autogroup), autogroup:self/ssh-check flagged
experimental. Rebuild embedded archive.
headscale rejects autogroups as tagOwners (only user/group:/tag:), which made
the shipped policy fatal. Replace with a valid allow-all default plus correct
commented examples for tightening. Document gating /admin to a pocket-id
superuser group via the headplane client's Allowed User Groups. Rebuild archive.
Integrate headplane (ghcr.io/tale/headplane) into the headscale stack, served
by Caddy at /admin. API-only (no Docker socket); deploy.sh mints a headscale
API key on first run, generates headplane.yaml, and wires optional OIDC login
via pocket-id (second client, /admin/oidc/callback). Adds HEADPLANE_* env,
compose service, Caddy routing; rebuild embedded archive.
Ship policy.hujson (mounted + installed on first deploy, edits preserved) and
wire policy.mode=file / policy.path in config.yaml. Translate the Tailscale
"grants" default into headscale's legacy "acls" format (self-access, tag:shared,
Tailscale SSH), since headscale 0.28 doesn't support grants. Embed in deploy.sh
and document `headscale policy check`.
Make pkce.method: S256 explicit in config.yaml (alongside pkce.enabled: true)
and note the pocket-id client must have PKCE enabled too. Rebuild embedded
archive.
Add /usr/local/bin/headscale (generated by deploy.sh) that runs the headscale
CLI inside the container with the compose file path baked in, so `headscale
users list` etc. work from any directory instead of erroring with
"no configuration file provided". Update the post-deploy hints accordingly.
Explain the project's origin — years of personal infra notes and scripts,
polished with Claude into a consistent, multi-distro toolkit — and note that
more is still being added from the collection. Framed as a starting point for
friends to use and extend. PRs and ideas welcome.
Restructure around a single entry point (automations.sh) with a Gum wizard and
a self-extracting bundle for repo-less installs. Add scripts/oslib.sh so the
provisioning scripts (setup-host, harden-ssh, harden-jumphost, sshuser) run on
Alpine/Debian/Alma; seed root keys from globals/.
- ntfy SSH-login alerts (user, source IP, key, region, jump target) via pam_exec
- daily auto-updates: AUTO_REBOOT=idle reboots only when no SSH active; opt-in
Alpine stable-branch upgrades (ALLOW_RELEASE_UPGRADE)
- cloud-init: generic base/jumphost + per-deployment, which harden SSH by
default on fresh VMs
- pocket-id: optional WebFinger block (BASE_DOMAIN), tag v2.8.0
- headscale: fix oidc.expiry schema for 0.28 so the container starts
- Gitea release workflow on tag (TOKEN_GITEA); repo URLs -> Gitea
- README/LICENSE/.gitignore/.gitattributes (force LF)
Restructure around a single entry point (automations.sh) with a Gum wizard and
a self-extracting bundle for repo-less installs. Add scripts/oslib.sh so the
provisioning scripts (setup-host, harden-ssh, harden-jumphost, sshuser) run on
Alpine/Debian/Alma; seed root keys from globals/.
- ntfy SSH-login alerts (user, source IP, key, region, jump target) via pam_exec
- daily auto-updates with AUTO_REBOOT=idle (reboots only when no SSH active) and
opt-in Alpine stable-branch upgrades
- generic + per-deployment cloud-init; Gitea release workflow on tag
- README/LICENSE/.gitignore/.gitattributes (force LF); repo URLs -> Gitea