Files
automations/README.md
T
57_WolveandClaude Opus 5 85d945ccde feat(ergo): persistent message history, plus deployment-readiness audit fixes
Ergo keeps history in RAM by default, so it is lost on every restart --
including the ones this stack's updater performs. HISTORY now selects a backend
at first deploy: sqlite (default, a file beside ircd.db, no extra container),
postgres (a pinned container via docker-compose.postgres.yml, loopback-only,
POSTGRES_PASSWORD_FILE so no secret lands in .env), or off. Both SQL backends
need Ergo 2.18.0+, which deploy.sh enforces. HISTORY_EXPIRE (default 30d) sets
retention, because upstream's 1w expire-time DELETES from persistent storage --
persistence with the shipped default would buy only a week.

Ergo opens the history backend only at startup and, unlike MySQL, has no
"after launching the server" guard for sqlite/postgresql: a rehash that enables
one reports success and then silently discards every message. So `ergoctl
history` restarts, and `ergoctl edit` detects a backend change and refuses to
apply it by rehash. The config validator runs with no network by design, so it
neutralises postgresql in its copy and reachability is checked separately.

Audit fixes (six lenses over first-deploy, runtime, lifecycle, security, docs):

- SECURITY (blocker): ircd/ is writable by the container uid while the 15-minute
  cert sync and every ergoctl config edit run as root. cp/install/> follow a
  symlink, so code execution in Ergo could redirect a root write onto any host
  file. All such writes now stage under $STACK_DIR (0700 root) and land via mv
  (rename(2) replaces a symlink instead of following it); root reads use cp -P
  or refuse. Regression-tested with real symlinks.
- deploy.sh no longer swaps the image when ERGO_TAG changes on a deployed stack
  (that bypassed pre-flight, snapshot, user warning and rollback); it points at
  `TARGET_VERSION=... ergoctl update update`.
- restore was lossy and could lock you out: it now snapshots the current config
  as well as the database, restores message history, re-points the postgres
  password and re-hashes the local admin password into the restored config, and
  validates before starting.
- `install -d -m` re-modes existing directories: `ergoctl backup /var/backups`
  no longer chmods it 0700 root, and the updater no longer re-modes /var/log
  (0775 root:syslog on Debian, which rsyslog needs).
- The admin oper password is printed only on a first run at a TTY, so it stays
  out of cloud-init serial-console logs.
- A failed update is remembered, so the daily job stops repeating a disruptive
  warn/stop/swap/roll-back cycle every night; postgres readiness gates an update
  that would otherwise stop a healthy server it cannot restart.
- certsync no longer sends "TLS recovered" for outcomes that synced nothing.
- ergoctl history writes .env only after the restart is healthy, and returns 0.
- CR/LF is stripped before IRC framing, so a multi-line argument cannot inject a
  second command; the ntfy token moves out of curl's argv.
- ufw/firewalld are additive, so 6667 is now explicitly revoked when PLAINTEXT=0.
- build.sh refuses to build a deploy.sh whose archive is missing a file the
  script reads -- the failure mode that would have shipped a stack aborting on
  every host.

Docs corrected against the code throughout, including retention, the pre-connect
account-registration default, encrypted-restore (AGE_IDENTITY), what a re-run
really does to .env, and what the update log does and does not contain.

Verified locally: all six suites pass (config render for each backend against the
real 2.19.1 template, yaml/oper/version/env helpers, the IRC client against a
fake server, and the audit fixes including the symlink escalation). Still not
exercised on a Docker host: the containers, ACME issuance, cert sync and
PostgreSQL itself.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-02 14:32:37 -05:00

272 lines
16 KiB
Markdown

# automations
Deployment and automation scripts for self-hosted, security-hardened
infrastructure. The host-provisioning scripts run on **Alpine, Debian, and
Alma Linux** (every distro difference is gated in
[`scripts/oslib.sh`](scripts/oslib.sh)); each Docker stack runs behind Caddy
with automatic Let's Encrypt TLS, orchestrated with Docker Compose.
> **About this repo.** This started as years of personal notes and one-off
> scripts I'd accumulated running my own infrastructure. I worked with Claude
> to clean them up, make them consistent and multi-distro, and add some
> quality-of-life options — so friends can use them too, and build on them.
> I'm still working through my collection and adding more as I go, so expect
> this to keep growing. Treat it as a starting point: fork it, wire in your own
> domains/keys, and send improvements back. PRs and ideas welcome.
## One command to run anything
```bash
curl -fsSL https://git.anomalous.dev/57_Wolve/automations/raw/branch/main/automations.sh \
| REPO_URL=https://git.anomalous.dev/57_Wolve/automations.git sh
```
Or, from a clone:
```bash
./automations.sh
```
Pipe it to `sh`, not `bash` — a stock Alpine box has busybox `ash` and no bash
at all. The launcher's prologue is POSIX sh: it installs `git`/`bash` for the
distro it finds itself on (apk / apt-get / dnf / yum), then re-execs itself
under bash. `./automations.sh` from a clone does the same, so a bare Alpine
install needs nothing prepared by hand.
[`automations.sh`](automations.sh) opens a **Gum** wizard (auto-installed) that
lets you:
- **Deploy on this host** — pick any deployment or generic script and run it,
prompting for the values it needs.
- **Build artifacts locally** — regenerate a deployment's self-contained
`deploy.sh` and optionally `scp` it to a target host.
Shared defaults (email, repo URL, SSH key source, backup recipient) come from
[`globals/`](globals/) so you set them once.
## Self-contained bundle (no repo access)
For hosts that shouldn't have repo/git access, package the whole repo into one
self-extracting script and serve it from a webserver or a **public Gitea
release** (the repo itself can stay private):
```bash
./build-bundle.sh # -> dist/automations-bundle.sh (HEAD)
./build-bundle.sh v1.2.0 # a specific tag
./build-bundle.sh worktree # current working tree
```
On the target host, download then run it (it extracts itself — it can't read a
pipe, so download first):
```bash
curl -fsSLO https://your-host/automations-bundle.sh
sh automations-bundle.sh # launcher wizard
sh automations-bundle.sh bash scripts/setup-host.sh # run one script
SSH_PORT=2222 sh automations-bundle.sh bash scripts/harden-jumphost.sh
```
Like the launcher, the bundle's stub is POSIX sh and installs `bash` if the
host lacks it, so it runs on a bare Alpine box.
It extracts to `INSTALL_DIR` (default `/opt/automations`) and runs the launcher
or the command you pass. The payload excludes ignored files, so no secrets are
embedded.
**Releases via Gitea Actions:** pushing a `vX.Y.Z` tag runs
[`.gitea/workflows/release.yml`](.gitea/workflows/release.yml), which builds the
bundle and attaches it to a release. Create a secret named **`TOKEN_GITEA`**
(a Gitea access token with `write:repository`) — note Gitea reserves the
`GITEA_` prefix, so the secret can't be called `GITEA_TOKEN`. Adjust `runs-on`
to match your runner. Mark the release public and `automations-bundle.sh` is
fetchable without repo access.
## Layout
The rule: **generic, run-anywhere scripts live in `scripts/`; deployment-specific
files live under `deployments/<name>/`.** Shared values live in `globals/`.
```
automations.sh # the launcher (one-liner entry point)
cloud-init/ # generic base + jumphost cloud-init (any of the 3 distros)
globals/ # shared assets: age-pubkey, authorized_keys, motd, globals.env
scripts/ # generic scripts (Alpine/Debian/Alma)
deployments/<name>/ # one folder per stack
```
### `scripts/` — generic (Alpine / Debian / Alma)
| Script | What it does |
|--------|--------------|
| [`setup-host.sh`](scripts/setup-host.sh) | Set hostname per the naming schema (derives FQDN + Node ID) and render the shared MOTD with auto-computed border spacing. |
| [`harden-ssh.sh`](scripts/harden-ssh.sh) | SSH hardening: post-quantum hybrid KEX, fresh Ed25519 host keys, key-only auth, external SFTP subsystem, sshguard. |
| [`harden-jumphost.sh`](scripts/harden-jumphost.sh) | Bastion hardening on top of `harden-ssh`: `ssh-admins` (shell) vs `ssh-jumpers` (ProxyJump-only) with a PermitOpen allow-list. |
| [`harden-firewall.sh`](scripts/harden-firewall.sh) | Deny-by-default host firewall: **iptables** on Alpine/Debian, **firewalld** on Alma/RHEL (set `FW_BACKEND` to override). Loopback, established, ICMP, SSH (configurable port) + registered ports; persisted natively (no boot hook). Same `allow`/`deny`/`list`/`disable` sub-commands on both. |
| [`sshuser.sh`](scripts/sshuser.sh) | Add/edit/remove SSH users on a hardened jump host (Gum TUI or CLI flags). Installed standalone as `sshuser`. |
| [`ntfy-ssh-login.sh`](scripts/ntfy-ssh-login.sh) | `pam_exec` hook that posts SSH logins to ntfy (user, source IP, key used, best-effort jump target), gated by group. Config: [`ssh-notify.conf.example`](scripts/ssh-notify.conf.example). |
| [`auto-update.sh`](scripts/auto-update.sh) | Daily unattended package updates; reports (doesn't auto-jump) a new Alpine branch; reboot detection; ntfy summary. `install`/`run`/`uninstall`. |
| [`oslib.sh`](scripts/oslib.sh) | OS-abstraction layer (detection, package manager, init system, SFTP path, hostname, boot hooks, native firewall persistence, sshguard, login notifier). Sourced, not run. **One file holds every distro difference.** |
| [`lib.sh`](scripts/lib.sh) | Launcher helpers (`ensure_gum`, `load_globals`, `resolve_ssh_keys`); sources `oslib.sh`. Not run directly. |
### `deployments/` — per stack
| Deployment | What it is | Depends on |
|------------|------------|------------|
| [`pocket-id`](deployments/pocket-id/) | OIDC provider (Caddy + Anubis PoW gate + Pocket-ID). | — |
| [`beszel`](deployments/beszel/) | Server monitoring hub. OIDC via pocket-id (post-deploy). | pocket-id (OIDC, optional) |
| [`headscale`](deployments/headscale/) | Self-hosted Tailscale control server, OIDC login. | pocket-id (OIDC) |
| [`webfinger`](deployments/webfinger/) | Serves `/.well-known/webfinger` for OIDC discovery; redirects the rest. | pocket-id (issuer) |
| [`squid`](deployments/squid/) | SSL-bump caching forward proxy — static-content cache + TLS interception via a local CA. **The exception: a forward proxy, not a Caddy/LE site.** | — |
| [`copyparty`](deployments/copyparty/) | Portable file server — web UI/WebDAV behind Caddy, plus direct **SFTP** + **FTPS**. Ships a security-notices-aware updater. | — |
| [`simplex`](deployments/simplex/) | SimpleX SMP + XFTP relay with Tor hidden services + encrypted backups. | globals/age-pubkey.txt |
| [`openbao`](deployments/openbao/) | Hardened tape-encryption key store for **Kanrisha** (the LTO tape-archive system — separate repo/host). **Exception: native TLS on the LAN, no Caddy/LE.** raft storage, mlock, self-signed or Smallstep-ACME cert, manual/HSM unseal. | globals/age-pubkey.txt (backups) |
| [`ergo`](deployments/ergo/) | IRC server ([Ergo](https://ergo.chat)) with Caddy for Let's Encrypt + the websocket endpoint. **Both containers use host networking** (real client IPs for bans/cloaking; the host firewall applies). Persistent message history via SQLite (default) or an optional PostgreSQL container. Ships `ergoctl` (admin CLI) and a health-checked updater that also syncs Caddy's cert into Ergo. | — |
## Conventions
- **Alpine + Docker Compose + Caddy/Let's Encrypt** across every stack, with two
exceptions: `squid` (a forward proxy with a local TLS-interception CA) and
`openbao` (a same-LAN secrets store with native TLS) — neither uses Caddy/LE.
`ergo` keeps Caddy/LE but runs both containers in the **host network
namespace** (IRC needs real client IPs), so it has no published ports.
- **`build.sh``deploy.sh`**: each stack's `build.sh` embeds its
`docker-compose.yml` / `Caddyfile` / `.env.example` into a single
self-contained `deploy.sh` (base64 tar.gz). That one file can be `scp`'d to a
host and run on its own. Rebuild after editing the loose files.
(simplex is the exception — it deploys via `install-simplex.sh`.)
- **Secrets** live in `.env` (generated on first deploy) and `globals/globals.env`,
both git-ignored. Only `*.example` templates and public keys are committed.
- **Non-interactive**: every `deploy.sh` honors `SKIP_PROMPTS=1` with values
supplied via the environment — which is how the per-deployment
`cloud-init.yml` templates stand a stack up unattended.
## Cloud-init
- **Provision a host**: [`cloud-init/base.yml`](cloud-init/base.yml) (hostname +
MOTD + SSH hardening) or [`cloud-init/jumphost.yml`](cloud-init/jumphost.yml)
(bastion). Distro-agnostic — they install prerequisites for whatever the image
is, then run the scripts.
- **Stand up a stack**: each deployment ships its own `cloud-init.yml` (e.g.
[`deployments/pocket-id/cloud-init.yml`](deployments/pocket-id/cloud-init.yml)).
These assume a fresh VM, so they **harden SSH first** (`harden-ssh.sh`,
`HARDEN_SSH=1` by default) and then deploy the stack.
Fill in `REPO_URL` and the values at the top of the `runcmd` block, paste as
instance user-data, and the host configures itself on first boot.
## Multi-OS notes
The host-provisioning scripts (`setup-host`, `harden-ssh`, `harden-jumphost`,
`sshuser`) and the nine Docker stacks (pocket-id, beszel, headscale, webfinger,
squid, copyparty, simplex, openbao, ergo) run on Alpine, Debian, and Alma. Distro differences live in
[`scripts/oslib.sh`](scripts/oslib.sh) — package manager (`apk`/`apt`/`dnf`),
init system (OpenRC/systemd), sshd service name, the per-distro `sftp-server`
path, hostname, boot hooks, and the sshguard log source/backend.
**simplex** remains **Alpine-targeted** — it depends on `awall` and Tor hidden
services with Alpine-specific wiring, so it isn't part of the tri-distro set.
## Host firewall
[`scripts/harden-firewall.sh`](scripts/harden-firewall.sh) installs a
**deny-by-default** baseline, with the backend chosen per family (override with
`FW_BACKEND=iptables|firewalld`):
- **Alpine / Debian → iptables.** `INPUT` drops everything except loopback,
established/related, ICMP, and SSH on the configured port — plus any ports a
deployment registers.
- **Alma / RHEL → firewalld** (its native firewall). The default zone is already
deny-by-default; we strip the stock `ssh`/`cockpit` services, open SSH +
registered ports, and let sshguard block via the `sshguard-firewalld` backend
(no `INPUT → sshguard` jump needed).
`OUTPUT`/egress stays open and `FORWARD` is left untouched, so Docker container
networking is unaffected. The harden scripts and `cloud-init/base.yml` /
`jumphost.yml` install it automatically (`ENABLE_FIREWALL=1` by default).
- **Configurable SSH port** — read live from `sshd_config`, so a bastion on
`2222` is firewalled correctly with no extra flags. Restrict the source with
`FW_SSH_SOURCE=<cidr>`; drop ping with `FW_ALLOW_PING=0`.
- **Native persistence, no boot hook** — on iptables hosts the ruleset is saved
and restored by the distro's own package: `iptables` + `ip6tables`
(Alpine/OpenRC) or `iptables-persistent` (Debian); the saved ruleset carries
the `INPUT → sshguard` jump. On firewalld hosts every change is `--permanent`,
so it persists across reboot natively and sshguard manages its own blocks.
- **Scripted additions** — deployments drop a rule file and re-apply:
```sh
printf '80/tcp\n443/tcp\n' > /etc/firewall/ports.d/mystack.rule
/usr/local/sbin/firewall-apply
```
or interactively: `harden-firewall.sh allow 443/tcp 51820/udp` /
`allow web` / `deny 51820/udp` / `list`.
- **Docker caveat** — containers published with `-p` (e.g. Caddy's 80/443)
reach the host through nat/`FORWARD` and **bypass `INPUT`**, so the firewall
neither blocks nor needs to open them; the per-stack rule files are
belt-and-braces for any host-bound bind and self-documentation. The one
stack that runs with `network_mode: host` (`ergo`) binds on the host
directly, so there the registered ports are the real gate.
- **Recovery** — `harden-firewall.sh disable` un-locks you: on iptables it
flushes the rules and sets `INPUT` back to `ACCEPT` (persisted); on firewalld it
re-opens SSH (the `ssh` service + the configured port) and leaves firewalld
running. A re-apply never drops the live SSH session — on iptables the
established-connection accept is added before the policy flips to `DROP`, and
firewalld reloads preserve established connections.
## SSH login notifications
The harden scripts can install a `pam_exec` hook
([`scripts/ntfy-ssh-login.sh`](scripts/ntfy-ssh-login.sh)) that posts every SSH
login to an [ntfy](https://ntfy.sh) topic. Enable it by passing `NTFY_URL` (and
optionally `NTFY_TOKEN`) when running `harden-ssh.sh` / `harden-jumphost.sh`, or
via the launcher / jumphost cloud-init. Each alert reports:
- the **user** and **source IP**,
- the **SSH key** they authenticated with (fingerprint, via `ExposeAuthInfo`),
- a **region tag** so you know which bastion fired it (derived from the host's
FQDN, e.g. `us-evi-1`),
- best-effort the **next hop** of a ProxyJump (see caveat below).
Filtering is by group: `NOTIFY_GROUPS` limits alerts to certain
groups/security levels, and `NOTIFY_PRIORITY_MAP` sets a per-group ntfy
priority. Config lives at `/etc/ssh-notify.conf` (mode 0600;
[`scripts/ssh-notify.conf.example`](scripts/ssh-notify.conf.example) documents
every key). A publish token is optional — leave it empty for a read-gated topic.
> **Jump-target caveat:** a ProxyJump opens a *direct-tcpip* channel (no
> session), so the destination never reaches `pam_exec`. The bastion only logs
> it at `LogLevel VERBOSE`/`DEBUG`; `harden-jumphost.sh` sets `VERBOSE` and the
> notifier parses the log best-effort. If the target isn't in the log it is
> simply omitted.
## Daily updates
[`scripts/auto-update.sh`](scripts/auto-update.sh) keeps a host patched
unattended — ideal for an SSH-only bastion, where a routine upgrade can barely
break anything. `harden-jumphost.sh` schedules it **by default** (set
`AUTO_UPDATE=0` to skip); `harden-ssh.sh` takes `AUTO_UPDATE=1`. It runs
daily via busybox `crond` (`/etc/periodic/daily`) on Alpine or a systemd
timer on Debian/Alma.
Each run:
- applies all **in-branch** package upgrades (`apk`/`apt`/`dnf`);
- **reports** a new Alpine *branch* (e.g. 3.21 → 3.22) by default — that
rewrites the repo branch, so it's opt-in. Set `ALLOW_RELEASE_UPGRADE=1` to
also **apply** it: it repoints `/etc/apk/repositories` to the newest **stable**
`vX.Y` (never `edge`), runs `apk upgrade --available`, and forces a reboot
flag. Debian/Alma stay report-only;
- detects when a **reboot** is needed (kernel/libc/openssl). `AUTO_REBOOT`
controls it: `0` = never (just flag), `1` = always, **`idle` = only when no
SSH connections are active** — so a bastion reboots itself once the admins and
ProxyJump tunnels have cleared, never mid-session. `harden-jumphost.sh`
defaults the bastion to `idle`. A deferred reboot is tracked in `/run` and
retried each day until it happens;
- sends an ntfy summary (reusing `/etc/ssh-notify.conf`).
Schedule it standalone with `auto-update.sh install` (or via the launcher), run
a pass now with `auto-update.sh run`, and preview safely with
`DRY_RUN=1 auto-update.sh run`.
## License
[MIT](LICENSE).