Author SHA1 Message Date
57_WolveandClaude Opus 5 393223dead feat(openbao): add baoctl, a session wrapper for running bao commands
This host has no `bao` CLI, only Docker, so every authenticated command is a
`docker compose exec` -- and every obvious way to get a token in there leaks it.
`-e BAO_TOKEN=<value>` puts it in the docker process's argv, which
/proc/<pid>/cmdline exposes to every user on the box. An inline
`BAO_TOKEN=<value> cmd` adds shell history on top.

The stdin trick 0eb7f26 uses cannot be the general answer, and the reason is
capability rather than ergonomics: bao's kvbuilder consumes stdin exactly once,
and the operator needs it for `policy write NAME -`, `write PATH -` (a JSON
body) and `key=-` (a single secret value). Spend stdin on the token and an OIDC
client secret has nowhere left to go but argv, reintroducing the leak that was
just closed. update.sh keeps piping because it runs one unattended command that
needs no stdin of its own; interactive work needs something else.

So baoctl is a SESSION wrapper. `baoctl login` prompts once with echo off --
verified against v2.6.2 that bao reads it through termios and requires a TTY --
and afterwards commands are typed verbatim with stdin free.

Three details that are load-bearing, all source-verified at v2.6.2:

- `bao login` prints the token in its success table. Without -no-print the
  interactive path dumps the root token into the exec session's scrollback,
  which is worse than what it replaces.
- The token you type is not what the session keeps. baoctl immediately mints a
  short-lived child and swaps it in via `bao token create -field=token |
  bao login -no-print -`, so the value never reaches an argv or a stdout, the
  session expires on its own, and logout can revoke it without killing the root
  token. A failed mint discards the login rather than leaving the typed token
  sitting in the session.
- logout both revokes AND removes the file. `bao token revoke -self` does not
  delete it and there is no `bao logout` in 2.6.x, so revoking alone leaves a
  stale file that fails with permission errors instead of "not logged in".

The session lives at /dev/shm/.bao-session in the container, pointed at by
BAO_TOKEN_PATH (new in 2.6.0). /dev/shm is already a per-container tmpfs, so the
token never touches disk and dies with the container -- and arranging that
needed no compose change, which matters because recreating this container means
a seal cycle and three unseal keys typed by a human.

It verifies TLS instead of reaching for -tls-skip-verify: ./tls is already
mounted read-only into the container and the generated cert carries
IP:127.0.0.1 in its SANs, so BAO_CACERT validates against the real listener.
-tls-skip-verify would have been the lazy default and is strictly worse.

Also warns when BAO_TOKEN is set in the caller's shell: baoctl never forwards
it, but an operator who set one will assume it is in play and debug the wrong
credential.

Caught while testing: `--ttl` with no value exited SILENTLY, because `shift 2`
with one argument left fails and set -e takes the script down before the
validation ran. Now the argument count is checked first -- another member of
this repo's set -e trap family.

Verified without a live host: help works with no stack present and touches no
docker; the missing-stack path errors cleanly; all four option-validation paths
report rather than exiting silently; the payload carries baoctl.
Not verified: login, the mint-and-swap and logout against a running vault.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-22 10:12:01 -05:00
57_WolveandClaude Opus 5 f8022379e5 fix(openbao): stop documenting unseal keys on the command line
deploy.sh's closing banner and README.md both instructed the operator to run
`bao operator unseal -tls-skip-verify <key>`, three times. That places an
unseal key in the docker process's argv, and /proc/<pid>/cmdline is
world-readable -- so every user on the host can read it. Three of those shares
reconstruct the master key, which makes this a worse instance of the leak class
0eb7f26 fixed for the token.

`bao operator unseal` with NO key argument prompts for it with echo suppressed.
Both places now say to pass nothing, and say why.

These lines predate the updater and this session's work; anyone who followed the
printed instructions has been exposing unseal shares to `ps` on that host.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-22 09:55:11 -05:00
57_WolveandClaude Opus 5 0eb7f26663 fix(openbao): never put the vault token on a command line
`BAO_TOKEN=<root> bash update.sh update` -- which the previous commit and the
pre-existing DR runbook both recommended -- leaks the token twice over. It goes
into the shell history, and `docker compose exec -e BAO_TOKEN=<value>` puts it
in the docker process's argv, where /proc/<pid>/cmdline makes it readable by
every user on the host. The second is the worse one and is not fixed by
anything the operator does at their prompt.

The token now reaches the container over STDIN and is never an argument to
anything: the container's own shell reads one line, exports it, and execs bao.
Verified locally that the child process sees the exact value -- including &, |
and a backslash -- while its argv contains zero occurrences of it.

How a token is supplied, in order: an already-exported BAO_TOKEN; a 0600 file
named by the new BAO_TOKEN_FILE; otherwise an echo-off prompt. With no token
and no terminal it dies telling the operator to use BAO_TOKEN_FILE rather than
an inline assignment, and says why.

Added a `snapshot` subcommand so the DR runbook no longer needs an inline
pipeline at all. That runbook and the deploy banner both carried the -e form
before this deployment had an updater, so both are corrected: taking a backup
is now `bash update.sh snapshot`, and restore keeps the stdin shape rather than
-e. The Kanrisha bootstrap lines say to export the token first instead of
passing it inline.

The conf file still refuses to hold a token, but now points at BAO_TOKEN_FILE
as the unattended answer -- and suggests a snapshot-policy token rather than
the root token for it.

Not verified without a live host: that `docker compose exec -T` forwards stdin
as expected. If it does not, the snapshot fails loudly with a permission error
and the upgrade aborts before touching anything, which is the safe direction.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-22 09:31:47 -05:00
57_WolveandClaude Opus 5 2d3d322b0e feat(openbao): add update.sh, an updater that refuses to strand a sealed vault
Follows the copyparty/ergo updater idiom -- check/update/run/install/uninstall,
a conf file the environment overrides, the version pinned into .env so the
running release is explicit, DRY_RUN -- but inverts its central assumption.

copyparty and ergo come back by themselves after a recreate. OpenBao comes back
SEALED: with the default Shamir seal a restart needs three unseal keys typed in
by a human. A scheduled `latest` update would therefore take the vault offline
at 03:00 and leave it there. So the scheduled path defaults to UPDATE_POLICY=
notify and never changes the running version; `install` schedules a daily CHECK
and says so. UPDATE_POLICY=auto opts in, and is STILL refused unless an
uncommented `seal` stanza is present in config.hcl -- only auto-unseal makes an
unattended update defensible.

`update` preflights before touching anything, because every one of these fails
worse halfway through than up front:
- the container must be running;
- the vault must be UNSEALED, since a sealed vault cannot produce a snapshot and
  there would be no rollback plan;
- BAO_TOKEN must be present, because the snapshot is token-gated on
  sys/storage/raft/snapshot;
- the target must not cross into 2.7.x while a built-in seal "pkcs11" stanza is
  active. That stanza is REMOVED in 2.7.0, not deprecated, so the vault would
  start with no way to unseal at all.

The snapshot is the rollback plan, not a formality: OpenBao's upgrade guide
states that reverting the image alone does not roll back the data store. It is
streamed out with `exec -T ... cat` rather than `compose cp`, which emits a TAR
wrapper that will not restore; written 0600 to /var/backups/openbao; and checked
for being a non-empty valid gzip archive, with a failure treated as fatal.
SKIP_SNAPSHOT=1 exists and warns exactly what it costs.

A failed pull or start rolls the OPENBAO_TAG pin back and restarts the previous
version. BAO_TOKEN is deliberately never written to the conf file: a long-lived
root token sitting next to the vault it opens defeats the vault.

Verified: the 2.7-with-active-pkcs11 refusal and its three negative cases
(2.6.2 active, 3.x active, 2.7 commented); auto-unseal detection distinguishing
a commented stanza from a live one; set_env/env_get; the release-tag parse with
and without a leading v; and all four `run` policy branches, including that
auto refuses without auto-unseal and that an unknown policy dies.

Not verified without a live host: the snapshot, pull and recreate themselves.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-22 09:21:11 -05:00
57_WolveandClaude Opus 5 f8c839a163 feat(openbao): pin 2.6.2, and correct the mlock and recovery-path claims
2.5.5 was two releases behind. Verified against the openbao/openbao source at
tag v2.6.2 that the bump is safe for an existing vault: raft's on-disk format
is unchanged (identical bbolt / hashicorp-raft / raft-boltdb pins,
byte-identical physical/raft/types.proto), snapshots restore in both
directions, and every stanza this config.hcl uses still parses -- the config
parser only moved import paths and the listener diff is additive.

Three corrections that came out of that check.

deploy.sh's printed crash-loop recovery command was already broken by 2.6.0.
That release adds `USER openbao` to the alpine image, so the container no
longer starts as root; the `docker run ... --entrypoint sh ... chown` it tells
the operator to run now executes as uid 100 and fails. Added `--user 0:0`, and
dropped the stale 2.5.5 literal from the same string. The detection logic
itself is unaffected -- `id -u openbao` still reads /etc/passwd and still
returns 100, verified by extracting the passwd layer from both published
images -- so a02524a holds and the volume does not need re-chowning.

The mlock story was wrong, and was already wrong on 2.5.5. OpenBao removed
mlock support: its own config parser carries "OpenBao has dropped support for
mlock. Please remove the line disable_mlock = false from your config and
disable or encrypt swap instead." So `cap_add: [IPC_LOCK]` and
`ulimits: memlock: -1` are inert, and config.hcl's "mlock keeps key material
off swap -- REQUIRED" was describing something that does not happen. The
BEHAVIOUR was right all along, because deploy.sh disables swap, which is the
actual mitigation; only the explanation was wrong. Corrected in config.hcl,
docker-compose.yml and the README. The two compose settings stay: they are
harmless, and removing them would recreate every deployed container for no
gain.

The built-in `seal "pkcs11"` stanza is deprecated in 2.6.0 for removal in
v2.7.0, and the HSM distribution is discontinued by then; PKCS#11 auto-unseal
continues only via the external `plugin "kms" "pkcs11" {}`. Noted where the
commented-out stanza lives, since a stack relying on it has to move before
taking 2.7.x.

Also worth recording and NOT acting on: do not switch to the openbao-distroless
image variant. It ships no shell, which breaks every `docker compose run
--entrypoint sh` probe deploy.sh uses.

Upgrading a LIVE vault is not automatic -- snapshot first, and the restart
comes back sealed. That is the next commit.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-22 09:10:20 -05:00
57_WolveandClaude Opus 5 41d09bc998 docs(openbao): the Proxmox per-VM firewall, and reaching the API over SSH
Two things that cost real debugging time on a live deploy.

A Docker-published port on a Proxmox guest is filtered by the per-VM firewall
on the PVE host, upstream of everything in the guest. The failure is
maximally misleading: inside the VM the DNAT rule, the FORWARD jumps, the
docker-proxy socket and a local curl to the bind address are all correct and
the listener answers 503, while a LAN client times out. SSH working proves
nothing -- it only proves a rule exists for 22. The decisive test is
`tcpdump -ni eth0 'tcp port 8200'` capturing zero packets during a failed
connection, which says the guest never saw them.

And: the publish is pinned to OPENBAO_BIND, so a browser on another subnet
cannot reach the UI. Document the SSH tunnel rather than widening the publish
-- the generated cert already carries DNS:localhost + IP:127.0.0.1, so
https://localhost:8200 validates against it unchanged. That is also the
better way to run the first `operator init`, since the unseal keys are then
shown in a browser instead of a root shell's scrollback.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-16 15:09:35 -05:00
57_WolveandClaude Opus 5 185f404549 fix(openbao): write .env through set_env, not an interpolated sed
947c899 removed this idiom from copyparty, headscale, pocket-id and squid;
openbao was missed. Its first-run seed still interpolated values into
`sed -i -e "s|^KEY=.*|KEY=${VAL}|"`, which silently corrupts any value
containing & (sed expands it to the whole match) and aborts the run under
set -e on one containing the s||| delimiter.

That was survivable while .env held only an address, a bind and an image tag.
It stops being survivable the moment an OIDC client secret or an issuer URL
with a query string goes in there, which is the next commit -- so this lands
first, on its own, with no behaviour change.

Adds the same awk/ENVIRON helper the other four use, replaces the seed's sed
block with set_env calls, and collapses the inlined copy of that awk (added
in dc9761a for the UI write-through) into a call to it.

Verified against plain, sec&ret, pipe|val, back\slash, an Authentik-shaped
discovery URL with an & query string, p@ss&w|rd and the empty string; plus
non-target lines left intact and the 0600 mode preserved across an in-place
update.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-16 15:09:35 -05:00
57_WolveandClaude Opus 5 dc9761a668 feat(openbao): serve the web UI, so init/unseal can happen in a browser
config.hcl shipped `ui = false`, so there was no UI to reach even with working
connectivity. That forces init and unseal through `docker compose exec`, which
puts the unseal keys and root token into a server shell's history and
scrollback -- the one place they should never be.

Make it a substitution point driven by OPENBAO_UI, default on. This adds no
exposure: the UI is served on the same listener, and anything that can reach
:8200 can already do everything through the API. What it changes is where the
unseal material is displayed -- the operator's browser instead of the host.

Rendering config.hcl rather than copying it needed three things to be right:

- Precedence. OPENBAO_UI passed to a run wins; otherwise the value .env already
  deploys applies; otherwise the default. Without that, `OPENBAO_UI=0 bash
  deploy.sh` against an existing node would report success and change nothing.
  Env-presence is captured before the ":=" default, so an explicit 0 is
  distinguishable from "not mentioned" and a re-run cannot silently re-enable
  the UI. A value passed this run is written back to .env.
- Reload. A bind-mounted file's CONTENTS are not part of the compose config
  hash, so `up -d` leaves a changed config.hcl unloaded -- the same trap fixed
  in copyparty (76d2a09) and simplex (0f22735). Restart only on an actual
  change, and say loudly that a restart RE-SEALS the vault, because that is not
  a free action on a running one.
- `cmp -s` stays inside an `if` condition; as `cmp -s A B && CONFIG_CHANGED=1`
  it would trip set -e whenever the files matched.

Verified: default renders `ui = true` and seeds .env; an identical re-run
reports no change; .env's 0 is honoured when nothing is passed; a passed 1 beats
.env's 0 and is written through; a non-boolean value dies.

Note for LAN-isolated hosts: the listener publishes on OPENBAO_BIND only, so a
browser on another subnet still cannot reach it. The generated cert already
carries DNS:localhost + IP:127.0.0.1 in its SANs, so
`ssh -L 8200:<bind-addr>:8200` and https://localhost:8200 works against the
existing cert without widening the publish.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-16 15:01:56 -05:00
57_WolveandClaude Opus 5 a02524ae6e fix(openbao): detect the account the entrypoint drops to, not the probe's
The image's entrypoint starts as root and then runs `su-exec openbao "$@"`
before exec'ing the server. deploy.sh probed with
`--entrypoint id -u`, which BYPASSES the entrypoint, so it reported UID 0 and
chowned config.hcl, tls/ and the raft volume to root -- while the server ran
as uid 100 and could write none of it:

  error initializing storage of type raft: failed to create fsm:
  failed to open bolt file: open /openbao/data/vault.db: permission denied

The container crash-looped on that, and tls.key (0600 root:root) would have
failed the listener straight afterwards. Confirmed on the affected host:
the probe reports 0, `id -u openbao` reports 100, and the entrypoint's line
92 is `set -- su-exec openbao "$@"`.

Ask for the account the entrypoint switches to, falling back to the old probe
when the image has no such account (then the server really does run as
whatever the entrypoint started as). Preferring the service account is also
the safe direction to be wrong in: root ignores file permissions, so chowning
to the unprivileged uid still works if the server turns out to run as root,
whereas the reverse is fatal.

Gate the raft chown on the volume's ACTUAL ownership rather than on a
first-run flag. The flag was false forever after the first deploy, so a volume
left root-owned by an earlier run -- exactly the state this bug created --
could never be repaired by re-running; the operator had to chown it by hand.
Reading the owner costs one container start and still keeps the recursive
chown off a healthy live raft dir. FIRST_RUN is now unused, so it is gone.

Also from the same run, two reporting failures:
- `docker compose ps` printed "Restarting (1)" and the script went on to print
  an unqualified DEPLOYED banner. It now inspects the container state and, when
  it is not `running`, says plainly that this is a crash loop rather than the
  expected sealed-and-unhealthy state -- before and after the banner, with the
  logs command to run.
- Every `docker compose` command in that banner assumed the project directory.
  deploy.sh runs them from $STACK_DIR itself, so the omission only bit the
  operator afterwards, with "no configuration file provided: not found". The
  banner now says to cd there first and quotes that error.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-16 14:49:30 -05:00
57_WolveandClaude Opus 5 0c151ce79b feat(knot-dns): add the missing host-level installer
The deployment shipped its payload but not the thing that installs it.
54a5c09 added README, knot.conf, knsctl, zone.tmpl, secrets.conf.example,
the aliases and the MOTD -- but no deploy.sh, no cloud-init.yml, and no
entry in automations.sh's DEPLOYMENTS. `git log --all` confirms deploy.sh
was never committed and it is not gitignored, yet README.md:96 and :100
tell the operator to run it. So the documented install path did not exist.

Alpine only, native, matching the README: Knot binds :53 directly, needs
real client addresses for RRL and cookies, and keeps its DNSSEC key store on
the host filesystem. The RHEL packaging needs EPEL, which nothing here sets
up, so anything that is not Alpine dies with a clear message rather than
half-installing somewhere untested.

Three decisions worth recording:

The include chain is stubbed. knot.conf include:s seven files the `dns` repo
owns; Knot treats a missing include as a config error, so a node the pipeline
has never delivered to would fail conf-check and never start. deploy.sh
writes a placeholder for each one that is ABSENT -- never over a delivered
file -- so the node comes up healthy serving no zones until the pipeline
lands.

TSIG is generated on a primary and required on a secondary. The keys must
match byte for byte, so a secondary that generated its own would
authenticate nothing; it now refuses to deploy without TSIG_AUTHORITIVE and
TSIG_ADMIN. A primary generates both and prints them once. An existing
secrets.conf is never rewritten, so a re-run cannot rotate a key out from
under a running estate.

PRIMARY_ADDR seeds a minimal remotes.conf on a secondary so it can bootstrap
by AXFR before the pipeline runs -- written only when remotes.conf was
absent, verified by re-running against a delivered file and confirming it is
left untouched.

Re-runs apply changes rather than freezing at first deploy, per the pattern
this repo just adopted elsewhere: knot.conf is re-rendered from .env every
run, env-presence is captured before the ":=" defaults, and values passed to
a re-run are written back to .env with the awk-based set_env from 947c899 --
which matters here because a TSIG secret can contain the characters that
broke the sed-based one.

Two bugs caught while testing this, before it shipped:
- the secrets.conf renderer used `++n` as a gsub argument, which awk
  evaluates on every line, not just matching ones -- both keys would have
  received the SAME secret, making the read-only admin key identical to the
  replication key. Increments on a matching line only now.
- the MOTD is a pre-drawn box, so substituting values of a different width
  than their @TOKEN@ shifted the right border on every login. Values are now
  padded to the token's span, measured over an ASCII-only region so it holds
  under busybox awk in the C locale; an over-long value overflows rather than
  being truncated.

Verified: knot.conf renders identity/NSID/listen and leaves the control
socket alone; all seven stubs are created on a fresh node and skipped on a
re-run; the remotes.conf seed fires only for a fresh secondary; .env seeds
every runtime key; the MOTD renders with no leftover tokens and an aligned
border. Not verified: apk, knotc and the service start, which need an actual
Alpine host.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-16 14:32:45 -05:00
57_WolveandClaude Opus 5 c00fc62b82 fix(ergo): guard a network rename, and persist ACME_EMAIL on a re-run
ergo already defends this bug class deliberately -- it adopts .env settings on
a re-run and dies with a precise remedy for ERGO_DOMAIN, HISTORY, PLAINTEXT
and a floating ERGO_TAG. NETWORK_NAME and ACME_EMAIL were gaps in that
defense, not design choices: both are absent from EXPLICIT and from the
re-run writeback loop, though the header promises it "writes back settings you
passed explicitly on this run".

A plain re-run stays safe, since both are adopted from .env as prompt
defaults. The hole is a value the operator actively supplies. NETWORK_NAME
then splits three ways: the generated docs under $STACK_DIR/docs and (when
exported) the Caddy landing page get the NEW name, ircd.yaml keeps the OLD one
because render_ircd_yaml never rewrites an existing file, and .env is never
updated so the next run reverts the docs too. The summary prints the new name.
Clients keep seeing the old one in NETWORK=.

Add yaml_network_name beside yaml_server_name and die with a remedy pointing
at `ergoctl edit` + `ergoctl rehash`, matching how ERGO_DOMAIN is handled --
a rename ircd.yaml cannot absorb should stop the run, not half-apply. Add
ACME_EMAIL to EXPLICIT and to the writeback loop, since unlike the network
name it genuinely can change on a re-run.

ergolib.sh is embedded, so the payload was regenerated with build.sh; that
rebuild also carries the env_set repair from 947c899.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-16 14:21:36 -05:00
57_WolveandClaude Opus 5 79db4d1e29 fix(beszel): apply BESZEL_DOMAIN passed to a re-run
Same defect as webfinger, and the typo case makes it concrete. A first deploy
with monitroing.example.com never gets an ACME certificate. The operator spots
it and re-runs with the correct hostname -- .env still holds the typo, so
Caddy still has only that site block, still fails ACME, and the endpoint still
does not work, while the script exits 0 printing
"URL: https://monitoring.example.com".

The two sources disagree and nothing compares them: validation greps .env for
non-empty (a typo passes), while the confirmation prompt and the DEPLOYED
banner expand this shell's variable. Interactively the prompted value is not
exported, so compose produces a byte-identical config and does not even
recreate the container -- nothing happens at all.

Fixed with the same idiom as webfinger and d654299.

Verified: a corrected BESZEL_DOMAIN lands in .env and is logged; a re-run
passing nothing leaves .env untouched.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-16 14:21:35 -05:00
57_WolveandClaude Opus 5 237aa913f0 fix(webfinger): apply values passed to a re-run instead of discarding them
.env was written only on the first run. All four values reach Caddy solely
through compose interpolating .env into the container environment -- the
Caddyfile is installed verbatim every run and carries Caddy's own
{$ISSUER_URL} / {$BASE_DOMAIN} / {$REDIRECT_URL} / {$ACME_EMAIL} placeholders,
resolved at config load. So .env is the only thing that decides what is served.

Both paths were wrong. Interactively, the prompts are required and have no
default, so a re-run made the operator retype all four -- and then discarded
every one, because a value read into a nameref is not exported and .env was
not rewritten. Via automations.sh the values ARE exported, so compose
preferred them for that run only and the documented later
`docker compose up -d` reverted to stale .env.

Either way the run exits 0 and the DEPLOYED banner prints this shell's
values, not what was deployed. Moving the IdP from auth.example.com to
id.example.com leaves the endpoint serving "href": the old issuer, pointing
every discovering client at a decommissioned issuer, with the deploy that was
meant to fix it reporting success.

Same shape as d654299: record which keys arrived in the environment before the
":=" defaults, pre-load .env into keys not passed so the prompts stop
demanding a full retype (and SKIP_PROMPTS=1 no longer dies over values .env
already has), then write only the passed keys through with set_env.

Verified: a corrected ISSUER_URL containing a & query string lands in .env
intact (the old escaping would have corrupted it); a re-run passing nothing
leaves .env byte-identical and does not die.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-16 14:21:35 -05:00
57_WolveandClaude Opus 5 947c89931d fix(deployments): stop interpolating .env values into a sed script
The set_env/env_set helper escaped its value with
`esc=${val//\/\\}; esc=${esc//|/\|}; esc=${esc//&/\&}` and then
interpolated it into `sed -i -e "s|^KEY=.*|KEY=${esc}|"`. The escaping does
not do what it looks like. Tested on bash 5.2:

  set_env K 'a&b'   ->  K=aK=seedb        (sed expanded & to the whole match)
  set_env K 'a|b'   ->  sed: unknown option to `s'   (rc!=0, aborts under set -e)

So any value containing & is silently corrupted and any value containing the
s||| delimiter kills the run. That is reachable: headscale writes
OIDC_CLIENT_SECRET through this, pocket-id writes REDIRECT_URL, copyparty
writes DATA_DIR. A generated secret or a URL query string hits both cases.

The copies in headscale and pocket-id were additionally mangled when they
were introduced -- `${val//\/\}` (pattern `\/`, a literal SLASH) and a raw
newline inside `printf '%s=%s\n'`. The mangled form is a no-op rather than a
corrupter, so the practical failure mode was the same as the original.

Replace all of them with an awk rewrite that passes the key and value through
the ENVIRONMENT, so the value is never parsed as part of a script and needs no
escaping at all. ENVIRON and index() are POSIX, so busybox awk handles them.
Output goes to a temp file and is copied back with `cat >`, which preserves the
original mode and owner -- a .env holding secrets stays 0600. If awk fails,
set -e aborts before .env is touched, which `sed -i` could not promise.

Verified against plain, a&b, a|b, a\b, p@ss&w|rd\x, R&D, a URL with a query
string, s/foo/bar/, a trailing space and the empty string; plus the
append-when-key-absent path, the file-does-not-exist path, non-target lines
left intact, no line-count drift, and mode preservation.

copyparty/update.sh and the rebuilt copyparty payload are included because
update.sh is embedded; regenerated with build.sh.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-16 14:21:07 -05:00
57_WolveandClaude Opus 5 0f2273545a fix(simplex): repair the caddy-init YAML fold that broke every deploy
The caddy-init service's `command:` used a YAML FOLDED scalar (`>`). Base
indent is set by `sh -c '` at 6; the closing `}` and the heredoc's `EOF`
were BOTH at 6, so YAML folded them into `} EOF`. The heredoc opened with
`<<EOF` therefore never saw a terminator line.

Confirmed by parsing the generated compose file: the command contains
`} EOF` and no bare EOF line. Feeding the resulting script to sh:

  here-document at line 3 delimited by end-of-file (wanted `EOF')
  syntax error: unexpected end of file          (rc=2)

So caddy-init exited 2 before its `>` redirect ever ran -- no Caddyfile was
ever written, on any deploy. The `if [ ! -f /etc/caddy/Caddyfile ]` guard
never executed at all. caddy gates on
`caddy-init: condition: service_completed_successfully`, and smp-server and
xftp-server gate on caddy being healthy, so nothing in the stack started.
`docker compose up -d` returned non-zero and `set -euo pipefail` aborted the
script before the final report. The same failure hit every entry point:
automations.sh, cloud-init, the OpenRC unit's start() (so every boot), and
restore.sh -- which restores the same broken compose file backup.sh saved.

Generate the Caddyfile from the deploy shell instead and drop caddy-init
entirely, which removes the whole class of problem. Unlike the old
first-run-only guard this also applies a changed DOMAIN, ACME_EMAIL or
KEY_TYPE on a re-run rather than freezing them at the first deploy.

The heredoc here is deliberately unquoted so ${DOMAIN}/${ACME_EMAIL}/
${KEY_TYPE} expand; Caddy's {uri} has no `$` and survives. `cmp -s` stays
inside an `if` condition -- as `cmp -s A B && CADDY_CHANGED=1` it would trip
set -e whenever the files matched. Caddy is restarted only when the file
actually changed AND compose did not already replace the container, since a
bind-mounted file's contents are not part of the compose config hash and
restarting seconds after a first start would interrupt initial ACME issuance.

The Caddyfile is now real on-disk state rather than something a container
regenerates, so it is added to backup.sh's targets and restore.sh's file
list; without that, a restore would start caddy against an empty
./caddy_conf and the servers would never pass their health gate.

Verified: first deploy writes it and reports no change; an identical re-run
reports no change; a corrected DOMAIN rewrites it and flags the restart;
{uri} survives and no .new file is left behind.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-16 14:21:07 -05:00
57_WolveandClaude Opus 5 a785a2bd9e fix(pocket-id): apply BASE_DOMAIN passed to a re-run, not just the first run
The Caddyfile is reassembled every run, but the WebFinger decision reads
BASE_DOMAIN/REDIRECT_URL from .env, and .env was written only on the first
run. So enabling WebFinger on a re-run was a silent no-op: the operator
answers the apex prompt, and the script prints the green
"[+] No BASE_DOMAIN -- pocket-id only" -- contradicting what they just
typed -- then pulls, restarts and prints DEPLOYED with exit 0.

The failure is invisible at the far end too. Caddy has no site block for the
apex, so it never gets a cert for it and discovery fails with a TLS or
connection error rather than a 404. Nothing warns; the one contradicting
line is a green [+] among docker pull output.

Moving the apex had the squid shape: compose reads the shell environment
first, so the new value served for that run only, and the documented later
`docker compose up -d` fell back to stale .env and reverted it.

Record which of the two keys arrived in the environment BEFORE the ":="
defaults, then write those through with a set_env upsert. Only keys actually
passed are touched: automations.sh drops a blank optional answer, so "blank"
cannot be told apart from "not supplied" and must not be read as "disable".
Since that means the prompt cannot retire WebFinger, the enabled path now
says so and points at the .env edit that can.

Verified: enabling on a re-run lands in .env and appends the block; changing
the apex lands and is logged; passing nothing touches nothing; BASE_DOMAIN
without REDIRECT_URL still dies rather than half-enabling.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-16 13:43:59 -05:00
57_WolveandClaude Opus 5 76d2a098cb fix(copyparty): make the documented FTP_NAT re-run remedy actually work
The summary printed at the end of every run says passive FTPS behind NAT is
fixed by re-running with FTP_NAT=... That could never work: the ftp-nat
substitution sat inside `if [[ ! -f "$CONF" ]]`, so on any re-run the conf
existed, the substitution was skipped, and ftp-nat stayed commented out.

cfg/copyparty.conf is the ONLY route FTP_NAT has to the service -- nothing
in docker-compose.yml interpolates it and the copyparty service is given no
environment, so .env's copy is a record, not the live setting. The remedial
run logged "exists; leaving it alone" as a green [+], reported healthy, and
reprinted the same advice. Passive FTPS failed exactly as before, with
nothing to distinguish "you did it wrong" from "it didn't take" -- and
because the template's commented example carries a literal IP, grepping the
conf shows an ftp-nat line with an address in it.

Lift the application out of the creation guard so it runs against an
existing conf too, and match `#?` so an already-set value (ISP change, typo)
is corrected rather than only the commented template line. Warn instead of
guessing if no ftp-nat line exists at all -- appending at EOF would land in
[accounts]. Kept above the chown, since `sed -i` rewrites as root.

`|| true` on the current-value read is load-bearing: on a still-commented
conf the grep matches nothing, and under `set -o pipefail` that would abort
the script inside the assignment -- the exact trap fixed in bf52426.

Also restart copyparty when the conf changed: compose will not recreate a
service whose image and config are unchanged, so `up -d` alone would leave
the edit on disk and the old value in the running process. And record the
value in .env so the two files do not disagree about what is deployed.

Verified: commented template applies; identical value is a no-op; a
different value is corrected; a conf with no ftp-nat line warns without
corruption; empty FTP_NAT skips the block entirely.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-16 13:43:59 -05:00
57_WolveandClaude Opus 5 d6542996fc fix(headscale): let a re-run's values reach .env instead of being reverted
deploy.sh deliberately re-reads .env before rendering config.yaml and
headplane.yaml ("so config.yaml substitution sees what's actually
deployed"). That design is right, but .env was seeded only when absent, so
"what's actually deployed" could never change: `set -a; . "$ENV_FILE"`
overwrote every value passed to the run, and `set -a` re-exported the stale
ones into compose as well.

Worse, the prompts at :192-197 run 56 lines BEFORE ENV_FILE is even defined,
so an interactive re-run asked for all six required values and then threw
every answer away. Via automations.sh the same six arrive exported and meet
the same fate. The validation only greps .env for non-empty values, which
the stale ones satisfy, so the run printed DEPLOYED and exited 0.

The case that matters is a rotated OIDC_CLIENT_SECRET: headscale keeps
presenting the retired secret at pocket-id's token endpoint, and
only_start_if_oidc_is_available probes issuer discovery, not the secret, so
nothing fails at deploy time. Either the leaked credential is still live and
the rotation is fiction, or tailnet OIDC login is broken and surfaces later
at some user's `tailscale up`. A corrected HEADSCALE_DOMAIN leaves the LE
cert hostname and the OIDC redirect URI on the typo; newly-supplied
headplane OIDC creds silently leave /admin on API-key login, so the
IdP-group gating the README recommends is never in force.

Record which keys arrived in the environment BEFORE the ":=" defaults erase
the distinction -- writing a blank over a live OIDC secret is exactly the
wrong move. Pre-load .env into keys NOT passed, before prompting, so the
prompts stop asking questions they will discard (and SKIP_PROMPTS=1 no
longer dies demanding values .env already has). Then write only the passed
keys through with a set_env upsert. Secret values are never echoed; a
changed HEADSCALE_DOMAIN warns about the new cert and the redirect URI.

The `. "$ENV_FILE"` stays: once .env carries this run's values it reads back
what was passed, and .env, config.yaml, headplane.yaml and compose agree.

Verified: rotated secret lands and is not echoed; changed domain lands and
warns; new headplane creds land; a re-run passing nothing touches nothing;
SKIP_PROMPTS=1 with a configured .env and no env vars no longer dies.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-16 13:43:58 -05:00
57_WolveandClaude Opus 5 4643b77083 fix(squid): write explicitly-passed values through to .env on a re-run
.env was seeded only when absent, so a re-run with a corrected value logged
".env exists; leaving it alone" and dropped it. That looked harmless because
compose reads the shell environment before .env: an exported BIND_ADDR did
narrow the bind for that run, `ss -ltn` confirmed it, and the deploy reported
success -- while .env still said 0.0.0.0.

The drift surfaces later. The documented update path is a plain
`docker compose up -d`, which has no such environment, falls back to .env, and
republishes an SSL-bumping intercepting proxy on every interface. Nothing
warned. TRUSTED_CIDR has the same shape: a tightened allow-list silently
reverts to whatever .env kept.

Record which runtime keys actually arrived in the environment BEFORE the ":="
defaults run -- PROXY_PORT especially, whose default is a non-empty 3128, so
afterwards an unset variable is indistinguishable from a supplied one. On a
re-run, write just those keys through with the existing set_env() (a targeted
per-key rewrite, not a file overwrite) and log each change. Keys not passed
that run are untouched, so hand-edits to .env survive.

Verified: narrowing BIND_ADDR updates .env and logs it; a re-run with nothing
exported leaves .env alone; re-passing identical values is a silent no-op; and
a hand-edited PROXY_PORT=8080 survives all three, which is what the
capture-before-defaults ordering exists for.

Found by an adversarial sweep for the openbao bug class (920edc5), then
confirmed by hand.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-16 13:27:42 -05:00
57_WolveandClaude Opus 5 bf52426299 fix(openbao): keep a failing address probe from aborting the deploy
host_addrs() ends in a pipeline, so under `set -o pipefail` a probe that
exits non-zero -- even after printing perfectly usable addresses, or
because awk is missing -- made `addrs="$(host_addrs)"` non-zero, and
`set -e` killed deploy.sh at that line.

Nothing was printed when it did: the 2>/dev/null had already swallowed the
tool's own error and the fail-open guard on the next lines was never
reached, so the operator got a bare exit 1 mid-deploy with nothing to
diagnose. Worst case it aborted a deploy whose bind address was CORRECT --
reproduced with an `ip` stub that prints the matching address, then exits 1.

Capture with `|| true` so the emptiness test actually drives the fail-open
the comment beside it already promised. Neutralising inside host_addrs
instead would not cover a missing awk, since pipefail takes the rightmost
non-zero status.

Re-verified the check is not weakened: a healthy probe with the address
genuinely absent still lists the host's addresses and dies with the full
message, and all seven .env/environment precedence cases are unchanged.

Found by adversarial review of 920edc5.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-16 13:18:40 -05:00
57_WolveandClaude Opus 5 920edc50b3 fix(openbao): validate the publish bind before writing any state
deploy.sh narrows the Docker publish bind to OPENBAO_ADDR when that is an
IP, but never checked that the host actually holds that address. A typo'd
IP therefore failed late, inside `docker compose up`:

  failed to bind host port 192.160.100.50:8200/tcp:
  cannot assign requested address

...by which point .env had been seeded and the TLS cert generated with the
bad address in its SAN. Neither is rewritten on a re-run (.env is never
overwritten, gen-tls.sh never regenerates over an existing pair), so
re-running with a corrected OPENBAO_ADDR silently changed nothing.

Add host_addrs() + check_bind_addr(), run before anything is written:
- lists the host's addresses from plain `ip addr show` -- no -o/scope
  filters, since busybox ip supports neither -- falling back to ifconfig,
  and skipping the check when neither exists rather than blocking;
- skips 0.0.0.0 / :: / *, and unwraps an [IPv6] publish literal;
- SKIP_BIND_CHECK=1 overrides for an address that only comes up later.

Resolve the bind compose will really interpolate, which follows compose's
own precedence -- shell environment before .env:
- exported (automations.sh passes answers via `env VAR=...`, or a
  standalone OPENBAO_BIND=... run): the environment wins, so warn when
  .env disagrees, because a later bare `docker compose up` would not;
- derived here: that assignment is not exported, so .env wins;
- in neither: compose falls back to 0.0.0.0 and publishes the API on
  every interface -- warn, since that is a silent exposure.

Also warn when .env's OPENBAO_ADDR differs from this run's, naming the
cert that has to be deleted for the SAN to be regenerated.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-16 13:09:23 -05:00
57_WolveandClaude Opus 5 a3843d3d85 fix(ssh): build KexAlgorithms from what OpenSSH supports, add classic opt-in
Hardened hosts rejected clients that implement the very same key exchange. The
KEX list was assembled from version arithmetic and emitted only the
standardised spellings:

    KexAlgorithms mlkem768x25519-sha256,sntrup761x25519-sha512

OpenSSH called that hybrid sntrup761x25519-sha512@openssh.com before the method
was standardised (8.5, in the default proposal from 8.9) and
sntrup761x25519-sha512 after (9.9), and KEXINIT matches names byte-exactly with
no alias resolution -- so every client older than the rename was refused with
"no matching key exchange method found" despite implementing the algorithm. The
same arithmetic was a latent server-side bug: on OpenSSH 9.0-9.8 it wrote the
post-standardisation name into sshd_config, which those builds do not know, and
sshd fatals on an unknown KexAlgorithms token rather than starting.

Ask the binary instead of guessing. oslib gains kex_supported(),
ssh_kex_pq_list(), ssh_kex_classic_list(), ssh_kex_list() and ssh_kex_has_pq(),
which filter candidates through `ssh -Q kex` and offer every spelling the host
actually has. Version thresholds are gone, and with them both failure modes --
including on distros whose backports make the version string meaningless.

SSH_ALLOW_CLASSIC_KEX=1 (off by default) additionally offers curve25519-sha256
and its @libssh.org spelling. Some clients have no PQ method at all: notably
Windows' in-box ssh.exe, which is not merely old -- Microsoft's fork compiles
sntrup761 out because it needs C99 VLAs that MSVC lacks, so even a fully patched
9.5p2 reports zero PQ methods. The knob is a real trade and says so in the
warning, the generated sshd_config comment, and the README: such a session is
safe against a classical attacker but has no store-now-decrypt-later protection.
Modern clients still negotiate PQ, since the client's preference order decides.

Three defects found reviewing the above, fixed here:

- the printed pre-reload verification command pinned the server's full list via
  `-o KexAlgorithms=`, which ssh rejects at option-parse time when the client
  lacks any one name. That made the one safety gate before a wholesale
  sshd_config swap a false negative for exactly the clients this commit admits.
  Dropped, matching harden-jumphost.sh.
- the no-PQ branch was unreachable: without the opt-in the classical names are
  never collected, so a host with no PQ hybrid died reporting "no usable key
  exchange method" instead of the actionable message written for it. The branch
  now keys off a separate PQ probe, and the empty-list die is narrowed to a
  genuinely empty `ssh -Q kex`.
- SSH_VER is cosmetic but its grep could abort the whole run under pipefail on
  any banner that does not match (vendor forks, OpenSSH_for_Windows_9.5p2) --
  silently, with no message. Guarded.

Wired through cloud-init/base.yml and jumphost.yml, since harden-ssh.sh rewrites
sshd_config wholesale on every run and a hand edit there does not survive.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-10 22:03:07 -05:00
57_WolveandClaude Opus 5 339c62a1b0 fix(firewall): skip the host firewall on Proxmox
Proxmox VE and Proxmox Mail Gateway are Debian, so os_detect classified them as
debian and harden-firewall.sh installed the iptables backend on top of
pve-firewall.

The conflict is not the obvious one. pve-firewall does NOT delete third-party
rules: it restores with `iptables-restore -n` (--noflush), only ever flushes
chains matching its own patterns (PVEFW-*, tapNiM-*, vethNiM-*, fwbrN-*,
GROUP-*), appends `-A INPUT -j PVEFW-INPUT` only when that hook is missing, and
never sets a built-in chain's policy. Our rules survive it. We are the ones
doing damage:

- `-P INPUT DROP` is ours alone, and PVEFW-HOST-IN RETURNs on accept rather than
  ACCEPTing (it still has to check tap rules), so traffic Proxmox explicitly
  allowed falls out of its chain onto our DROP -- we silently override the
  platform's own accepts.
- netfilter-persistent restores at boot with a full iptables-restore (no
  --noflush), wiping PVE's hook along with everything else until the daemon
  re-appends it ~10s later.
- a deny-by-default chain has to enumerate the whole platform to stay usable:
  8006, 5405-5412/udp corosync, 60000-60050, 5900-5999, 3128, 22, 111/udp, plus
  Ceph when hyperconverged -- and `-i lo`, or pveproxy loses pvedaemon on :85.
- under the nftables backend (PVE 8.2+) an nft DROP beats an iptables ACCEPT, so
  our rules would not even be authoritative.

So don't manage a firewall there at all:

- oslib: is_proxmox() -- matches hosts shipping pve-firewall (VE/PMG), not PBS.
- harden-firewall.sh: a third backend, "pve", that deliberately does nothing.
  apply explains and exits 0, allow/deny refuse loudly rather than fake success
  for a rule they didn't add, list shows pve-firewall status. It overrides an
  explicit FW_BACKEND; FW_IGNORE_PVE=1 is the one escape hatch.
- harden-ssh.sh / harden-jumphost.sh: skip the firewall and install the
  standalone INPUT -> sshguard boot hook instead. That jump is safe alongside
  pve-firewall -- inserted with -I, it sits ahead of the appended PVEFW-INPUT
  hook and keeps first look at NEW connections.

Detection only helps hosts built from here on, so `disable` cleans up one that
was hardened earlier: it detects leftovers (and says which signal fired), sets
INPUT ACCEPT *before* flushing so it can't drop the SSH session it runs over,
deletes /etc/firewall and the engine, disables boot restore and renames the
saved rulesets aside rather than persisting the open state, re-adds the sshguard
jump, and restarts pve-firewall. `apply` points at it when it spots leftovers.

Documented plainly that skipped is NOT protected: Proxmox's firewall is off by
default (cluster-wide enable defaults to 0, and the daemon tears its chains down
every ~10s while it is), so these hosts have no host firewall until someone
enables it -- and the node panel's "Firewall: Yes" is ignored while the
datacenter one says No.

Also: svc_disable + fw_restore_services/fw_saved_files in oslib
(fw_enable_restore now derives from the former), and usage() prints the whole
header block instead of a hardcoded line range.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-10 18:51:50 -05:00
57_Wolve 752385cb4c Merge pull request 'feat(knot-dns): authoritative Knot DNS node deployment' (#8) from feat/knot-dns into main
Reviewed-on: #8
2026-09-10 05:40:18 +00:00
57_WolveandClaude Opus 5 54a5c0931a feat(knot-dns): authoritative Knot DNS node deployment
Native Alpine deployment for the anycast DNS estate -- no Docker, no Caddy,
alongside squid and openbao as an exception to the repo norm. Knot binds :53
directly, needs real client addresses for RRL and DNS cookies, and its DNSSEC
key store must live on the host filesystem.

Deploys a NODE. Zone data lives in the separate dns repo and arrives from its
pipeline. The split is /etc/knot/knot.conf: written here once as a skeleton of
include: lines covering only what belongs to a box (identity, NSID, storage
paths, listen, logging, control socket); everything that is DNS policy --
templates, dnssec policy, remotes/ACLs, modules, the domain inventory and the
zone files -- is delivered by the dns repo.

knsctl replaces adddns.pl and adddnssec.pl, fixing four defects:
- the duplicate check searched for the domain in BIND named.conf double-quote
  syntax (/"$domain"/) against unquoted YAML, so it could never match; only
  the -f zone-file test ever caught anything
- neither script consulted the other class's manifest, so a domain already in
  public.conf could be appended to dnssec.conf and fail the reload AFTER both
  files had been written
- nothing validated before reloading
- the reload was non-blocking, so a rejected config reported success

Its manifest matching is anchored on the YAML key and escapes the dot, so
barsrvno.de and srvnoXde no longer false-positive against srvno.de.

Aliases preserve the existing muscle memory with three corrections: -b on
every triggering knotc command (without it knotc returns OK when the command
was SENT, not when it succeeded); knzr (zone-reload) added alongside knrl
(reload), since reloading one zone's data is the right verb for a record
change and a full reload is only needed when a zone is added or removed; and
serial/NSID helpers that query unicast addresses, because asking the anycast
service address reaches whichever node is nearest and says nothing about
which node is stale.

Break-glass writes (add/remove/edit) warn and audit-log: they are overwritten
by the next pipeline deploy unless the change also lands in git. Removal
refuses to purge DNSSEC keys -- zone-purge +keys is irreversible on Knot
3.5.x, the key trash bin having arrived in 3.6.0 -- and prints the ordering
requirement, since removing a signed zone before the parent DS is withdrawn
is an outage for validating resolvers rather than a graceful shutdown.

deploy.sh, build.sh and cloud-init.yml are deliberately not included yet;
they are blocked on the Knot version decision, which sets the apk pin and
feature availability. See the Status section in the README.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-10 00:36:38 -05:00
57_WolveandClaude Opus 5 6dba4a5f09 docs(ergo): user guide and admin/moderator cheat sheet
Two quick-reference documents, written to be scanned rather than read:

- USER-GUIDE.md for people new to IRC -- connecting, claiming a nickname,
  the dozen commands that matter, scrollback, and a plain-language privacy
  section (cloaked IP, channels logged for the configured retention, how to
  turn off DM storage).
- ADMIN-CHEATSHEET.md split by where you work: from IRC as an operator
  (UBAN, KILL, DEFCON, ChanServ, NickServ) and on the host via ergoctl, plus
  mode tables and a "when things go wrong" section.

Both use the same placeholder convention as ergo.motd and are rendered by
deploy.sh into $STACK_DIR/docs/ with the network's real name, domain and
retention, so they can be handed straight to users and moderators. Refreshed
on every run, like the other installed files.

Command and mode references were checked against the v2.19.1 sources rather
than written from memory: irc/modes/modes.go for every mode letter, and
irc/chanserv.go and irc/nickserv.go for the service subcommands and which
require an oper capability.

build.sh's embed guard caught the docs being added to FILES without being
added to deploy.sh's EMBEDDED manifest -- the exact failure it was added to
prevent.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-02 21:48:01 -05:00
57_WolveandClaude Opus 5 73eeb981b8 feat(ergo): a real MOTD with house rules, and honest retention wording
Replaces the placeholder MOTD with a usable one: a coloured header, five house
rules, the NickServ/ChanServ/SASL commands people actually need, and a short
"worth knowing" section.

The retention line is generated rather than hard-coded, so it matches the
deployment instead of drifting from it: HISTORY=off says messages are not
stored, otherwise it names the real HISTORY_EXPIRE. Users are told plainly that
channels are logged and for how long, which is the honest counterpart to
enabling persistent history by default.

Rendered and checked: valid UTF-8 (an invalid byte makes the whole config fail
to load), ASCII-only for old clients, no stray $-escapes, only documented colour
names, and 72 columns at the widest.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-02 15:53:00 -05:00
57_WolveandClaude Opus 5 134736c082 fix(ergo): findings from reviewing the first live deploy
Seven confirmed findings from an observed-vs-expected review of a real deploy
transcript against the code (four other proposed findings were refuted and no
change was made for them).

- Caddyfile reformatted so `caddy fmt` is clean, removing the warning Caddy
  printed on every validate and every start. Two causes, not one: `caddy fmt`
  indents with TABS, and it deletes a blank line whose following line begins
  with `{` -- which is why the warning pointed at line 17, the blank before the
  global options block. The check is a whole-file byte comparison, so the line
  number was only the first difference and the entire file had to be
  reformatted. Verified whitespace-only outside the heredoc, whose body is left
  byte-identical: Caddy strips padding derived from the closing marker's
  indentation, so re-indenting it would change what the page serves.
- deploy.sh no longer upgrades Caddy behind the operator's back. A bare
  `docker compose pull` refreshed the floating `caddy:2-alpine` on every re-run
  and `up -d` then recreated it, doing exactly what CADDY_AUTOUPDATE=0 promises
  not to, with none of update.sh's health check or rollback. Pulls are now
  per service: the pinned Ergo tag always, Caddy only when absent or opted in.
- deploy.sh seeds the ACME_EMAIL and NETWORK_NAME prompts from .env, so pressing
  Enter through a re-run no longer renames the network to the hostname in the
  summary while .env keeps the real one.
- The Caddy restart guard compares the container's identity across `up -d`. A
  container compose created or recreated has already read the new Caddyfile;
  only one left running still holds the old config, and `svc_state` cannot tell
  those apart.
- Container logs are now rotated (json-file, 10m x 3) on every service. Alpine's
  docker package ships no daemon.json, so the default is unbounded, and the
  60-second health probe alone writes a log line per run.
- ergoctl notes that LUSERS counts its own probe connection, which is why a
  server with nobody on it reports one invisible user.

Verified: seven local suites pass, the embedded archive round-trips, line
endings are LF, and the reformatted Caddyfile is a verified fixed point of the
formatter (semantically identical token-for-token to the previous one).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-02 15:38:13 -05:00
57_WolveandClaude Opus 5 f916eb549a feat(ergo): ergoctl oper certfp, plus polish from the first live deploy
Adds client-certificate authentication for operators, so an oper can be
identified by the SHA-256 fingerprint of their TLS client certificate instead of
a typed password:

  ergoctl oper certfp <name> <fingerprint|nick> [--auto] [--clear]

The fingerprint can be given directly or read off a connected user over IRC
(276 RPL_WHOISCERTFP). Without --auto Ergo requires BOTH the certificate and the
password -- a second factor; with --auto the password is removed and the
certificate alone grants oper on connect.

It refuses to do this to the 'admin' oper, which matters: ergoctl authenticates
as admin over the loopback PLAINTEXT listener, which presents no client
certificate, and Ergo requires a configured certfp to match. Setting one there
would lock ergoctl -- and the scheduled jobs that use it -- out of the server
permanently.

Polish from the first real deploy on irc-1.srvno.de:

- deploy.sh no longer restarts Caddy on a first deploy. The Caddyfile compare
  treated "destination does not exist" as a change, so Caddy was recreated
  seconds after starting, while the initial ACME order was in flight.
- ergoctl status no longer prints the raw leading parameters of the 265/266
  LUSERS numerics ("1 1 Current local users 1, max 1"), which duplicate the
  counts already in the text. 252/254 keep theirs, where it is the only count.
- ergoctl oper list now shows which opers have a certfp and which are auto.

Verified: seven local suites pass, including new coverage for fingerprint
normalisation (colons and uppercase), insertion into the correct oper block
without touching siblings, replace-not-duplicate, the --auto and --clear paths,
276 parsing, and the admin guard.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-02 15:34:32 -05:00
57_WolveandClaude Opus 5 85d945ccde feat(ergo): persistent message history, plus deployment-readiness audit fixes
Ergo keeps history in RAM by default, so it is lost on every restart --
including the ones this stack's updater performs. HISTORY now selects a backend
at first deploy: sqlite (default, a file beside ircd.db, no extra container),
postgres (a pinned container via docker-compose.postgres.yml, loopback-only,
POSTGRES_PASSWORD_FILE so no secret lands in .env), or off. Both SQL backends
need Ergo 2.18.0+, which deploy.sh enforces. HISTORY_EXPIRE (default 30d) sets
retention, because upstream's 1w expire-time DELETES from persistent storage --
persistence with the shipped default would buy only a week.

Ergo opens the history backend only at startup and, unlike MySQL, has no
"after launching the server" guard for sqlite/postgresql: a rehash that enables
one reports success and then silently discards every message. So `ergoctl
history` restarts, and `ergoctl edit` detects a backend change and refuses to
apply it by rehash. The config validator runs with no network by design, so it
neutralises postgresql in its copy and reachability is checked separately.

Audit fixes (six lenses over first-deploy, runtime, lifecycle, security, docs):

- SECURITY (blocker): ircd/ is writable by the container uid while the 15-minute
  cert sync and every ergoctl config edit run as root. cp/install/> follow a
  symlink, so code execution in Ergo could redirect a root write onto any host
  file. All such writes now stage under $STACK_DIR (0700 root) and land via mv
  (rename(2) replaces a symlink instead of following it); root reads use cp -P
  or refuse. Regression-tested with real symlinks.
- deploy.sh no longer swaps the image when ERGO_TAG changes on a deployed stack
  (that bypassed pre-flight, snapshot, user warning and rollback); it points at
  `TARGET_VERSION=... ergoctl update update`.
- restore was lossy and could lock you out: it now snapshots the current config
  as well as the database, restores message history, re-points the postgres
  password and re-hashes the local admin password into the restored config, and
  validates before starting.
- `install -d -m` re-modes existing directories: `ergoctl backup /var/backups`
  no longer chmods it 0700 root, and the updater no longer re-modes /var/log
  (0775 root:syslog on Debian, which rsyslog needs).
- The admin oper password is printed only on a first run at a TTY, so it stays
  out of cloud-init serial-console logs.
- A failed update is remembered, so the daily job stops repeating a disruptive
  warn/stop/swap/roll-back cycle every night; postgres readiness gates an update
  that would otherwise stop a healthy server it cannot restart.
- certsync no longer sends "TLS recovered" for outcomes that synced nothing.
- ergoctl history writes .env only after the restart is healthy, and returns 0.
- CR/LF is stripped before IRC framing, so a multi-line argument cannot inject a
  second command; the ntfy token moves out of curl's argv.
- ufw/firewalld are additive, so 6667 is now explicitly revoked when PLAINTEXT=0.
- build.sh refuses to build a deploy.sh whose archive is missing a file the
  script reads -- the failure mode that would have shipped a stack aborting on
  every host.

Docs corrected against the code throughout, including retention, the pre-connect
account-registration default, encrypted-restore (AGE_IDENTITY), what a re-run
really does to .env, and what the update log does and does not contain.

Verified locally: all six suites pass (config render for each backend against the
real 2.19.1 template, yaml/oper/version/env helpers, the IRC client against a
fake server, and the audit fixes including the symlink escalation). Still not
exercised on a Docker host: the containers, ACME issuance, cert sync and
PostgreSQL itself.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-02 14:32:37 -05:00
57_WolveandClaude Opus 5 4cbd808eb1 feat(ergo): add Ergo IRC server deployment with ergoctl admin CLI
New deployments/ergo/: the Ergo IRC server behind Caddy for Let's Encrypt TLS
and the IRC-over-WebSocket endpoint. Docker rather than a native OpenRC/systemd
service, because Alpine's apk ergo trails upstream (3.24 ships 2.18.0 against a
2.19.1 security release) and Debian/Alma package it at all -- so native would
mean three install paths plus a per-distro ACME client.

Both containers run with network_mode: host. IRC bans, throttling and cloaking
key on the client's address, and Docker's userland proxy would hide every IPv6
client behind the bridge gateway; host mode also makes the repo's INPUT firewall
genuinely govern 80/443/6697. Caddy reaches Ergo over loopback, which is what
lets Ergo honour X-Forwarded-For (proxy-allowed-from defaults to localhost) and
mark web sessions secure.

- deploy.sh generates ircd.yaml ONCE from the pulled image's own default.yaml
  (version-matched), rewriting the listeners/websockets blocks wholesale rather
  than patching lines, then asserts hard post-conditions and validates with
  `ergo run --smoke` in a throwaway container before anything starts.
- update.sh: pinned vX.Y.Z tags, GHSA + "### Security" release-note policies,
  pre-flight against the new image, user NOTICE + grace, stop-consistent DB
  snapshot, health check (IRC-level, not a bare TCP connect) and rollback that
  restores the DB only when the schema actually moved. Compatibility-break
  releases are held for review. certsync copies Caddy's cert pairwise-atomically
  and verifies the fingerprint served on 6697 after the rehash.
- ergoctl: status/users/logs, validated edit+rehash, oper add/passwd/rm,
  moderation, backup/restore, cert and update passthrough. Talks IRC to the
  loopback listener over bash /dev/tcp and strips control characters from
  replies.
- Ergo runs as a non-root system user, read-only rootfs, all caps dropped;
  Caddy keeps only NET_BIND_SERVICE, with admin API and HTTP/3 off.

Reviewed adversarially across six lenses; 20 confirmed findings fixed, notably
a dead SIGHUP fallback (`rc=$?` after an `if` is always 0), several `set -e`
aborts from non-total pipelines, a release-list cache that only ever populated
in a subshell, and re-runs that used shell defaults instead of the deployed
.env. Verified locally: bash -n, LF endings, the ircd.yaml render against the
real 2.19.1 template in both PLAINTEXT modes, the yaml/oper/version/env helpers,
and the IRC client against a fake server (registration, oper, rehash success and
400-failure, control-character stripping, server-down paths). Not yet exercised
on a Docker host: the containers themselves, ACME issuance and cert sync.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-02 13:19:31 -05:00
57_Wolve 1c1c41ac17 Merge pull request 'fix(alpine): create /usr/local/sbin, and run the PAM sshd so pam_exec fires' (#7) from fix/alpine-posix-launcher into main
Reviewed-on: #7
2026-09-02 13:18:59 +00:00
57_WolveandClaude Opus 5 7c13e18ca5 fix(alpine): create /usr/local/sbin, and run the PAM sshd so pam_exec fires
Two Alpine-only failures from an irc-1 run.

1. harden-firewall.sh died with

     line 161: /usr/local/sbin/firewall-apply: No such file or directory

   and left INPUT unfiltered. Alpine's baselayout ships /usr/local/{bin,lib,
   share} but no sbin, and nothing created it. Create it before writing the
   engine; same guard in oslib's install_boot_hook / install_daily_job (which
   `install` into /usr/local/sbin, and `install` does not make parent dirs) and
   for /usr/local/bin in harden-jumphost.

2. The ntfy login notifier never fired despite reporting ACTIVE. Alpine keeps
   PAM support in a SEPARATE binary: openssh-server gives /usr/sbin/sshd (no
   PAM), openssh-server-pam gives /usr/sbin/sshd.pam. The OpenRC init only
   picks the PAM one in start_pre (checkconfig -> update_command), and its
   reload/stop match the process with `start-stop-daemon --exec "$command"`.
   So a host already running /usr/sbin/sshd never swaps: reload signals a
   non-matching process, stop matches nothing, a later start hits "address
   already in use". sshd keeps serving without PAM, so /etc/pam.d/sshd -- and
   the pam_exec hook in it -- is never consulted.

   Add oslib sshd_wanted_binary / sshd_running_binary / sshd_apply_config, and
   end the harden scripts with sshd_apply_config instead of svc_reload: reload
   as before, but on Alpine stop by pidfile and start again when the running
   binary isn't the one the config calls for. Established sessions are separate
   processes and survive; only the listener blinks.

   install_login_notifier now checks that PAM can actually be reached on Alpine
   (sshd.pam present, UsePAM yes) and warns with the fix instead of reporting a
   hook that can never run as ACTIVE -- the same "silent notifier" trap the
   Alma run hit.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-02 08:18:09 -05:00
57_Wolve 844b6b4254 Merge pull request 'fix(launcher): run on a bare Alpine host (POSIX prologue, self-installs bash)' (#6) from fix/alpine-posix-launcher into main
Reviewed-on: #6
2026-08-10 22:18:56 +00:00
57_WolveandClaude Opus 5 f71f8d615a fix(launcher): run on a bare Alpine host (POSIX prologue, self-installs bash)
automations.sh was #!/usr/bin/env bash and written in bash, but a stock Alpine
box has busybox ash and no bash at all -- so the launcher died before it could
install anything:

  ./automations.sh        -> env: 'bash': No such file or directory
  curl ... | bash         -> bash: not found
  sh automations.sh       -> syntax errors

cloud-init/base.yml already installs bash/git/curl before touching the repo;
the launcher never got the same treatment.

Give it a #!/bin/sh shebang and a strictly-POSIX prologue that ash can parse:
self-locate via ${0%/*} (busybox dirname takes `--` as the filename, so
`dirname -- "$0"` would answer "."), clone on the piped path as before, install
git/bash via apk/apt-get/dnf/yum, then exec bash on this same file. The bash
pass skips the prologue via BASH_VERSION, and an exported _AUTOMATIONS_REEXEC
guard rules out an exec loop. Everything below the handoff is the unchanged
bash launcher.

The generated bundle stub had the identical bug -- it's the other fresh-host
entry point -- so make it POSIX sh too and have it install bash before
extracting.

Docs: the one-liner now pipes to `sh`, bundle examples use `sh`.

Still bash-only on a bare Alpine host: the scp'd deployments/*/deploy.sh
artifacts.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-10 17:13:30 -05:00
57_WolveandClaude Opus 4.8 b027de2182 fix(openbao): follow-ups from adversarial fix-verification
An adversarial re-verification of 0812f34 found a couple of the fixes were
incomplete or regressed; addressed here.

- DR was an INCOMPLETE fix: `raft snapshot save`/`restore` are token-gated
  (sys/storage/raft/snapshot is sudo-capable) and the container carries no
  ambient token, so the previously-"fixed" backup returned "missing client
  token" and streamed a zero-byte snapshot. Both the deploy.sh runbook and the
  README DR flow now pass `-e BAO_TOKEN=<token>` on save and restore.
- IPv6 bind REGRESSION (introduced by the bind-narrowing): a bare IPv6 literal
  in the compose port map (`fd00::10:8200:8200`) is invalid and aborts at
  `docker compose pull`. Now IPv4 vs IPv6 are classified separately and IPv6 is
  bracketed (`[fd00::10]:8200:8200`).
- Docker readiness race (now reachable since install_docker is actually
  called): openrc backgrounds dockerd and returns before the socket is up, so
  the next `docker compose pull` raced it under set -e. install_docker now polls
  `docker info` for up to 30s.
- UID detection hardened: added `-T` to the one-off `docker compose run`, and an
  empty result is now a loud warning (with the manual-chown remedy) instead of a
  silent fall-through to root that would re-create the crash-loop on a non-root
  image.
- The raft-volume chown is gated to first run (captured before any compose-run
  instantiates the volume), so idempotent re-deploys don't recursively re-chown
  a live raft dir.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-08 19:33:23 -05:00
57_WolveandClaude Opus 4.8 0812f345a8 fix(openbao): sanity-audit fixes — 4 blockers + hardening
A multi-agent sanity audit of the freshly-merged deployment found four
end-to-end blockers (and several smaller issues); all fixed here.

HIGH (were blocking):
- deploy.sh never called install_docker(), so `docker compose pull` hit
  command-not-found on any host without Docker. Now called before the
  compose steps.
- The container's server process runs as the image's own (often non-root)
  user but the mounted config/TLS were root-owned 0640/0600 and the raft
  volume root-owned -> vault crash-looped, never binding :8200. deploy.sh
  now detects the image UID after pull and aligns ownership of config.hcl,
  ./tls and the data volume (a no-op when the image runs as root);
  config.hcl is installed 0644 (holds no secrets).
- Docs told operators to set the daemon key `openbao_ca_cert`, but the
  Kanrisha daemon's key is `ca_cert` (config.go, mapstructure:"ca_cert").
  The wrong key is fatal on strict unmarshal / leaves TLS unverified.
  Renamed in all 5 places (config.hcl, gen-tls.sh, deploy.sh x2, README).
- DR backup used `docker compose cp openbao:… -`, which emits a TAR stream,
  so the age-encrypted snapshot was tar-wrapped and would not restore.
  Switched to `docker compose exec -T openbao cat` for the raw bytes, wrote
  the snapshot to a scratch path (not the live raft dir), and documented the
  matching restore.

MEDIUM:
- Swap detection used `swapon --show` (absent on BusyBox) and `\s` (GNU-only)
  -> silently no-op on Alpine, leaving swap on. Now uses /proc/swaps and
  [[:space:]] so mlock hardening actually holds on musl.
- A Docker-published port bypasses the host INPUT firewall, so the source
  rule was illusory. deploy.sh now narrows OPENBAO_BIND to OPENBAO_ADDR when
  it is an IP, the compose/README/.env comments state the reality, and a new
  Exposure section + an init-immediately warning were added.
- Fixed broken ../kanrisha/ and deployments/kanrisha/ links (separate repo).

LOW:
- OPENBAO_TLS_SANS is now honored (folded into the SAN list from the env).
- .gitignore excludes *.snap / *.snap.age.
- Bootstrap note clarifies bootstrap.sh needs the `bao` CLI (run it from the
  Kanrisha host/workstation, not this Docker-only vault host).
- README multi-OS count corrected (eight stacks) + automations.sh header
  lists openbao.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-08 19:11:37 -05:00
57_WolveandClaude Opus 4.8 2efc9dbffb feat(openbao): hardened same-LAN tape-encryption key store for Kanrisha
A dedicated OpenBao deployment, kept OFF the Kanrisha tape host so a compromise
of the tape node can't reach the vault.

- Native TLS on the listener (self-signed by default via gen-tls.sh, or a
  CA-signed cert from a Smallstep CA over ACME) — no Caddy/Let's Encrypt;
  reached over the LAN, not the public internet.
- Integrated raft storage (clean snapshot-based DR).
- mlock on (cap_add IPC_LOCK + memlock unlimited + host swapoff in deploy.sh).
- Manual unseal by default; optional PKCS#11 HSM auto-unseal.
- deploy.sh: Docker install (Alpine/Debian/Alma), self-signed cert, .env seed,
  swapoff, firewall 8200/tcp, compose up; then prints init/unseal + the KV-v2 +
  AppRole bootstrap for Kanrisha + the raft-snapshot DR flow. Self-contained
  (config payload embedded by build.sh).
- Registered in automations.sh + the README deployment table.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-08 14:10:22 -05:00
57_Wolve c6b0f7d418 Merge pull request 'feat(copyparty): add file-server deployment with SFTP/FTPS + security-notices updater' (#5) from claude/agitated-bose-f427e5 into main
Reviewed-on: #5
2026-06-29 20:57:39 +00:00
57_WolveandClaude Opus 4.8 c00ca055f2 feat(copyparty): add file-server deployment with SFTP/FTPS + security-notices updater
New deployments/copyparty/: copyparty (copyparty/ac) behind Caddy/LE for the
web UI/WebDAV, plus its own SFTP (password auth) and FTPS listeners published
directly. Ships update.sh, which drives container updates off copyparty's
security-advisories API (api.copyparty.eu/advisories) -- policies latest|security|off.

- Real client IP end-to-end: Caddy XFF/X-Real-IP + copyparty xff-src: lan.
- SFTP host key + self-signed FTPS cert generated/persisted in /cfg; admin
  password generated on first deploy; conf auto-included via the image's % /cfg.
- Firewall opens 80/443 + SFTP/FTPS + passive range (colon form for ports.d).
- Wired into automations.sh, README, .gitignore; cloud-init for fresh VMs.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-29 15:56:24 -05:00
57_Wolve cd402ba79e Merge pull request 'feat(squid): add SSL-bump caching forward-proxy deployment' (#4) from feat/squid-proxy into main
Reviewed-on: #4
2026-06-22 21:33:34 +00:00
57_Wolve e7ba68790b Done — the Squid deployment is fully built, wired in, and statically verified on the feat/squid-proxy branch. ## What landed deployments/squid/ (13 files) plus wiring into [automations.sh](automations.sh:51), [README.md](README.md), and [.gitignore](.gitignore). All six approved design decisions are implemented: - **Self-built minimal Alpine image** — apk add squid (ssl-bump compiled in), openssl/tini. [entrypoint.sh](deployments/squid/entrypoint.sh) renders squid.conf and generates the cache policy from the lists at start. - **SSL-bump** peek→splice→bump with [splice-domains.txt](deployments/squid/splice-domains.txt) passthrough. - **Wildcard hostname caching** — leading-dot [cache-domains.txt](deployments/squid/cache-domains.txt) + optional regex; CACHE_ONLY_LISTED boost/allowlist toggle. - **HTML/dynamic storage gate**, with query strings exempt on boosted domains. - **CA lifecycle** — generated on first deploy via the built image, key 0600 on the host, staged to tmpfs for the container, never embedded, git-ignored. - **Deny-by-default http_access + BIND_ADDR pinning** as the real access gate (Docker bypasses INPUT). ## One bug caught & fixed The test of the entrypoint's policy generation showed it emitting ubuntu.com with an **unescaped dot** — the s/\./\\./g double-backslash is sed-implementation-dependent and silently failed. Switched to a [.] character class (ubuntu[.]com), portable across busybox/git-bash, re-embedded, re-audited. ✅ ## Verified vs. pending **Verified locally:** bash -n on all scripts, LF endings, launcher parses, embedded archive byte-matches the loose files, and the cache-policy generation produces correct Squid directives in both modes. **Pending a Docker host** (this Windows box has none, as the plan anticipated): the container functional smoke test — build, CA mint, a bumped HTTPS fetch, a TCP_HIT, and a splice check. A few version-sensitive specifics get confirmed there: tls-cert=/tls-key= option names, the security_file_certgen path, busybox install -o/-g, ssl::server_name subdomain matching, and the squid -z flag. (All isolated and easy to adjust if any differ.) ## Commit (for you to GPG-sign) Staged and ready. Per your convention I haven't committed — here's the message: `` feat(squid): add SSL-bump caching forward-proxy deployment New deployments/squid/: an explicit forward proxy with SSL-bump TLS interception (local CA, generated on first deploy) and hostname-targeted static-content caching. Unlike the other stacks it is a forward proxy, not a Caddy/Let's-Encrypt inbound site. - Self-built minimal Alpine image (apk squid ships ssl-bump); entrypoint renders squid.conf and generates the cache policy from the domain lists. - Wildcard hostname caching (cache-domains.txt leading-dot + optional cache-domains.regex); boost vs strict-allowlist toggle (CACHE_ONLY_LISTED). - Storage gate never caches HTML or dynamic content; query strings exempt on boosted domains so versioned static assets still cache. - splice-domains.txt passthrough for pinned/banking domains. - Deny-by-default http_access (TRUSTED_CIDR) + BIND_ADDR pinning; CA key 0600 on host, never embedded, git-ignored. - Wired into automations.sh, README, .gitignore; cloud-init for fresh VMs. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> ` `bash git commit -S -F- <<'EOF' <paste message above> EOF ` Want me to **push feat/squid-proxy** once you've committed (you'd open the PR in Gitea), or hold? And if you can point me at a Docker host (or run docker compose build in /srv/squid` after a test deploy), I'll drive the functional smoke test and fix anything the live Squid version flags.
New deployments/squid/: an explicit forward proxy with SSL-bump TLS
interception (local CA, generated on first deploy) and hostname-targeted
static-content caching. Unlike the other stacks it is a forward proxy, not
a Caddy/Let's-Encrypt inbound site.
- Self-built minimal Alpine image (apk squid ships ssl-bump); entrypoint
  renders squid.conf and generates the cache policy from the domain lists.
- Wildcard hostname caching (cache-domains.txt leading-dot + optional
  cache-domains.regex); boost vs strict-allowlist toggle (CACHE_ONLY_LISTED).
- Storage gate never caches HTML or dynamic content; query strings exempt on
  boosted domains so versioned static assets still cache.
- splice-domains.txt passthrough for pinned/banking domains.
- Deny-by-default http_access (TRUSTED_CIDR) + BIND_ADDR pinning; CA key 0600
  on host, never embedded, git-ignored.
- Wired into automations.sh, README, .gitignore; cloud-init for fresh VMs.
2026-06-22 16:32:25 -05:00
57_Wolve 8555f7cbe4 Merge pull request 'feat(firewall): drive firewalld on Alma/RHEL with full CLI parity' (#3) from fix/alma-harden-resilience into main
Reviewed-on: #3
2026-06-14 22:42:55 +00:00
57_WolveandClaude Opus 4.8 fe25f35305 feat(firewall): drive firewalld on Alma/RHEL with full CLI parity
A fresh Alma box has firewalld active, and the iptables-based harden-firewall.sh
refused to run there (caught by harden-ssh's '|| warn', so the host firewall was
silently skipped). Use firewalld natively on the rhel family instead of fighting it.

- harden-firewall.sh: family-aware backend. On rhel, apply/allow/deny/list/disable
  drive firewall-cmd (deny-by-default zone, SSH + registered ports, ping policy,
  source-restricted rich rules); Alpine/Debian keep the iptables engine unchanged.
  FW_BACKEND=iptables|firewalld overrides.
- oslib: install_firewalld(); sshguard_backend() prefers sshg-fw-firewalld on rhel
  so brute-force blocks land in firewalld (no INPUT->sshguard jump needed).
- Deployments already fall through to a firewall-cmd branch when the iptables
  engine is absent, so they need no changes.
- README + script header document the per-family backend.

harden-ssh / harden-jumphost are unchanged -- they call harden-firewall.sh apply
and read sshguard_backend(), so the switch happens underneath them.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-14 17:16:24 -05:00
57_Wolve 4b02cac919 Merge pull request 'fix(oslib): install_openssh must not report failure on non-Alpine' (#2) from fix/alma-harden-resilience into main
Reviewed-on: #2
2026-06-14 22:00:53 +00:00
57_WolveandClaude Opus 4.8 c3e2e9c52b fix(oslib): install_openssh must not report failure on non-Alpine
install_openssh ended with '[[ "$OS_FAMILY" == alpine ]] && pkg_install ...'.
As the function's LAST statement, that trailing test returns 1 on every
non-Alpine OS (a false '[[ ]]' exits 1), so the function reported failure even
when the packages installed fine. Harmless while the call was bare under set -e
(a short-circuited && is exempt), but the new 'install_openssh || die' guard
read it as a real failure and aborted harden-ssh on Alma right after
'Installing OpenSSH server...'.

Fix: convert the Alpine-only linux-pam step to an if-block, and add '|| return 1'
to the main install so a genuine package failure still propagates honestly.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-14 16:59:39 -05:00
57_Wolve a844037cb0 Merge pull request 'fix(launcher): install git across all distros in the curl-pipe bootstrap' (#1) from fix/alma-harden-resilience into main
Reviewed-on: #1
2026-06-14 21:54:47 +00:00
57_WolveandClaude Opus 4.8 60433e4c8d fix(harden): keep hardening and the ntfy notifier alive when sshguard can't install
On a fresh AlmaLinux 9.8 box, install_bruteforce_protection ran unguarded under
'set -euo pipefail'. When sshguard (from EPEL) wasn't installable at that moment,
the single failed dnf aborted the ENTIRE harden run before it wrote sshd_config
or installed the pam_exec login notifier -- leaving a stock, unhardened box and a
silently-missing ntfy hook.

- oslib: install the iptables backend best-effort first, then sshguard, and
  return sshguard's status so callers can treat it as non-fatal.
- harden-ssh/harden-jumphost: install_openssh now dies with a clear message on
  failure; sshguard is '|| warn' so sshd hardening and the notifier still apply.
- install_login_notifier verifies the script + pam hook landed and logs
  'Login notifier ACTIVE' (or a loud warning) instead of failing silently.
- ntfy-ssh-login.sh: NTFY_DEBUG=1 logs delivery attempts + curl errors to
  /var/log/ssh-notify.log so the next silent failure leaves a trace.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-14 16:53:39 -05:00
57_WolveandClaude Opus 4.8 3c02574dd0 fix(launcher): install git across all distros in the curl-pipe bootstrap
The one-liner bootstrap only auto-installed git on Alpine (apk), so piping it
onto a fresh Debian/Alma host with no git fell straight into 'git clone' and
died with 'git: command not found'. oslib's pkg_install can't help here -- the
repo isn't on disk yet. Install git inline via apk/apt-get/dnf/yum, and fail
with a clear message if it still can't.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-14 16:53:39 -05:00
85 changed files with 12231 additions and 800 deletions
+7
View File
@@ -13,3 +13,10 @@ Caddyfile text eol=lf
# Treat the embedded-archive deploy scripts as text too (they're base64+shell).
deployments/*/deploy.sh text eol=lf
# Extension-less / unusual-extension scripts and configs that must stay LF.
deployments/ergo/ergoctl text eol=lf
deployments/knot-dns/knsctl text eol=lf
*.tmpl text eol=lf
*.caddy text eol=lf
*.motd text eol=lf
+18
View File
@@ -19,6 +19,24 @@ id_ed25519
id_ed25519_*
*.pem
# Squid TLS-interception CA -- generated on the host at deploy time, never
# committed (the private key can MITM any client that trusts it).
deployments/squid/ssl/
# copyparty generated config -- cfg/copyparty.conf holds the admin password and
# cfg/ftps.pem the FTPS key; both are generated on the host at deploy time. The
# copyparty.conf.example template stays tracked.
deployments/copyparty/cfg/
# Ergo -- generated on the host at deploy time: ircd/ (config with the oper
# hash, database, TLS key), caddy/ (ACME account + certs), secrets/ (admin oper
# password), backups/, templates/. Only the loose sources stay tracked.
deployments/ergo/ircd/
deployments/ergo/caddy/
deployments/ergo/secrets/
deployments/ergo/backups/
deployments/ergo/templates/
# ── Backups ─────────────────────────────────────────────────────────────────
*.tar.gz.age
*-backup-*.tar.gz*
+184 -30
View File
@@ -18,7 +18,7 @@ with automatic Let's Encrypt TLS, orchestrated with Docker Compose.
```bash
curl -fsSL https://git.anomalous.dev/57_Wolve/automations/raw/branch/main/automations.sh \
| REPO_URL=https://git.anomalous.dev/57_Wolve/automations.git bash
| REPO_URL=https://git.anomalous.dev/57_Wolve/automations.git sh
```
Or, from a clone:
@@ -27,6 +27,12 @@ Or, from a clone:
./automations.sh
```
Pipe it to `sh`, not `bash` — a stock Alpine box has busybox `ash` and no bash
at all. The launcher's prologue is POSIX sh: it installs `git`/`bash` for the
distro it finds itself on (apk / apt-get / dnf / yum), then re-execs itself
under bash. `./automations.sh` from a clone does the same, so a bare Alpine
install needs nothing prepared by hand.
[`automations.sh`](automations.sh) opens a **Gum** wizard (auto-installed) that
lets you:
@@ -55,11 +61,14 @@ pipe, so download first):
```bash
curl -fsSLO https://your-host/automations-bundle.sh
bash automations-bundle.sh # launcher wizard
bash automations-bundle.sh bash scripts/setup-host.sh # run one script
SSH_PORT=2222 bash automations-bundle.sh bash scripts/harden-jumphost.sh
sh automations-bundle.sh # launcher wizard
sh automations-bundle.sh bash scripts/setup-host.sh # run one script
SSH_PORT=2222 sh automations-bundle.sh bash scripts/harden-jumphost.sh
```
Like the launcher, the bundle's stub is POSIX sh and installs `bash` if the
host lacks it, so it runs on a bare Alpine box.
It extracts to `INSTALL_DIR` (default `/opt/automations`) and runs the launcher
or the command you pass. The payload excludes ignored files, so no secrets are
embedded.
@@ -90,9 +99,9 @@ deployments/<name>/ # one folder per stack
| Script | What it does |
|--------|--------------|
| [`setup-host.sh`](scripts/setup-host.sh) | Set hostname per the naming schema (derives FQDN + Node ID) and render the shared MOTD with auto-computed border spacing. |
| [`harden-ssh.sh`](scripts/harden-ssh.sh) | SSH hardening: post-quantum hybrid KEX, fresh Ed25519 host keys, key-only auth, external SFTP subsystem, sshguard. |
| [`harden-ssh.sh`](scripts/harden-ssh.sh) | SSH hardening: post-quantum hybrid KEX (every spelling the host's OpenSSH supports; `SSH_ALLOW_CLASSIC_KEX=1` adds a curve25519 fallback for clients with no PQ method), fresh Ed25519 host keys, key-only auth, external SFTP subsystem, sshguard. |
| [`harden-jumphost.sh`](scripts/harden-jumphost.sh) | Bastion hardening on top of `harden-ssh`: `ssh-admins` (shell) vs `ssh-jumpers` (ProxyJump-only) with a PermitOpen allow-list. |
| [`harden-firewall.sh`](scripts/harden-firewall.sh) | Deny-by-default iptables baseline: loopback, established, ICMP, SSH (configurable port) + registered ports. Persisted **natively** (no boot hook). Deployments add ports via `/etc/firewall/ports.d` + `firewall-apply`. Sub-commands: `allow`/`deny`/`list`/`disable`. |
| [`harden-firewall.sh`](scripts/harden-firewall.sh) | Deny-by-default host firewall: **iptables** on Alpine/Debian, **firewalld** on Alma/RHEL (set `FW_BACKEND` to override), **skipped on Proxmox** (`pve-firewall` owns the ruleset). Loopback, established, ICMP, SSH (configurable port) + registered ports; persisted natively (no boot hook). Same `allow`/`deny`/`list`/`disable` sub-commands on both. |
| [`sshuser.sh`](scripts/sshuser.sh) | Add/edit/remove SSH users on a hardened jump host (Gum TUI or CLI flags). Installed standalone as `sshuser`. |
| [`ntfy-ssh-login.sh`](scripts/ntfy-ssh-login.sh) | `pam_exec` hook that posts SSH logins to ntfy (user, source IP, key used, best-effort jump target), gated by group. Config: [`ssh-notify.conf.example`](scripts/ssh-notify.conf.example). |
| [`auto-update.sh`](scripts/auto-update.sh) | Daily unattended package updates; reports (doesn't auto-jump) a new Alpine branch; reboot detection; ntfy summary. `install`/`run`/`uninstall`. |
@@ -107,12 +116,19 @@ deployments/<name>/ # one folder per stack
| [`beszel`](deployments/beszel/) | Server monitoring hub. OIDC via pocket-id (post-deploy). | pocket-id (OIDC, optional) |
| [`headscale`](deployments/headscale/) | Self-hosted Tailscale control server, OIDC login. | pocket-id (OIDC) |
| [`webfinger`](deployments/webfinger/) | Serves `/.well-known/webfinger` for OIDC discovery; redirects the rest. | pocket-id (issuer) |
| [`squid`](deployments/squid/) | SSL-bump caching forward proxy — static-content cache + TLS interception via a local CA. **The exception: a forward proxy, not a Caddy/LE site.** | — |
| [`copyparty`](deployments/copyparty/) | Portable file server — web UI/WebDAV behind Caddy, plus direct **SFTP** + **FTPS**. Ships a security-notices-aware updater. | — |
| [`simplex`](deployments/simplex/) | SimpleX SMP + XFTP relay with Tor hidden services + encrypted backups. | globals/age-pubkey.txt |
| [`kanrisha`](deployments/kanrisha/) | Custom app on a ScoutFS filesystem; two-phase install (host prep + ScoutFS, then the Docker stack). **Alma/Rocky only.** | ScoutFS (host) |
| [`openbao`](deployments/openbao/) | Hardened tape-encryption key store for **Kanrisha** (the LTO tape-archive system — separate repo/host). **Exception: native TLS on the LAN, no Caddy/LE.** raft storage, mlock, self-signed or Smallstep-ACME cert, manual/HSM unseal. | globals/age-pubkey.txt (backups) |
| [`ergo`](deployments/ergo/) | IRC server ([Ergo](https://ergo.chat)) with Caddy for Let's Encrypt + the websocket endpoint. **Both containers use host networking** (real client IPs for bans/cloaking; the host firewall applies). Persistent message history via SQLite (default) or an optional PostgreSQL container. Ships `ergoctl` (admin CLI) and a health-checked updater that also syncs Caddy's cert into Ergo. | — |
## Conventions
- **Alpine + Docker Compose + Caddy/Let's Encrypt** across every stack.
- **Alpine + Docker Compose + Caddy/Let's Encrypt** across every stack, with two
exceptions: `squid` (a forward proxy with a local TLS-interception CA) and
`openbao` (a same-LAN secrets store with native TLS) — neither uses Caddy/LE.
`ergo` keeps Caddy/LE but runs both containers in the **host network
namespace** (IRC needs real client IPs), so it has no published ports.
- **`build.sh` → `deploy.sh`**: each stack's `build.sh` embeds its
`docker-compose.yml` / `Caddyfile` / `.env.example` into a single
self-contained `deploy.sh` (base64 tar.gz). That one file can be `scp`'d to a
@@ -141,8 +157,8 @@ instance user-data, and the host configures itself on first boot.
## Multi-OS notes
The host-provisioning scripts (`setup-host`, `harden-ssh`, `harden-jumphost`,
`sshuser`) and the four Docker stacks (pocket-id, beszel, headscale, webfinger)
run on Alpine, Debian, and Alma. Distro differences live in
`sshuser`) and the nine Docker stacks (pocket-id, beszel, headscale, webfinger,
squid, copyparty, simplex, openbao, ergo) run on Alpine, Debian, and Alma. Distro differences live in
[`scripts/oslib.sh`](scripts/oslib.sh) — package manager (`apk`/`apt`/`dnf`),
init system (OpenRC/systemd), sshd service name, the per-distro `sftp-server`
path, hostname, boot hooks, and the sshguard log source/backend.
@@ -150,29 +166,83 @@ path, hostname, boot hooks, and the sshguard log source/backend.
**simplex** remains **Alpine-targeted** — it depends on `awall` and Tor hidden
services with Alpine-specific wiring, so it isn't part of the tri-distro set.
**kanrisha** is the inverse — **Alma/Rocky-targeted**. It installs ScoutFS (a
RHEL-only kernel filesystem) and its host package set via `dnf`, so it runs on
the RHEL family only for now.
### Key exchange and old clients
The KEX list is built from what the host's OpenSSH actually supports (`ssh -Q kex`),
not from its version number, and it offers **every spelling** of each algorithm.
That matters: OpenSSH called the same hybrid `sntrup761x25519-sha512@openssh.com`
before the method was standardised and `sntrup761x25519-sha512` after, and SSH
matches algorithm names as exact strings — so a server offering only the new
spelling rejects a client implementing the identical algorithm under the old one,
with `Unable to negotiate ... no matching key exchange method found`.
Client floor for a PQ-only host, from the OpenSSH release notes:
| Client | Result |
|---|---|
| 9.9+ | works — knows both spellings, and ML-KEM |
| 8.9 – 9.8 | works — its default proposal carries `sntrup761x25519-sha512@openssh.com` |
| 8.5 – 8.8 | has the algorithm but does **not offer it** by default; needs `KexAlgorithms +sntrup761x25519-sha512@openssh.com` client-side |
| < 8.5 | no PQ hybrid exists — locked out by design |
| **Windows in-box `ssh.exe`** | **locked out at every version, including a fully-patched 9.5p2** |
The Windows case is not an age problem, which makes it easy to misdiagnose: Microsoft's
fork **compiles `sntrup761` out**, because it needs C99 variable-length arrays that MSVC
does not support (Win32-OpenSSH #2140, #2391). So a fully-patched Windows 11 reports
`OpenSSH_for_Windows_9.5p2` — new enough on paper — and still lists zero PQ methods in
`ssh -Q kex`. Two ways out, in order of preference:
1. **Use a client that has PQ.** Check what is already on the box before installing
anything — Git for Windows bundles genuine upstream OpenSSH:
```powershell
& "$env:ProgramFiles\Git\usr\bin\ssh.exe" -Q kex | Select-String 'mlkem|sntrup'
```
If that prints anything, use that binary and install nothing. Otherwise
`winget install --id Microsoft.OpenSSH.Preview` (10.0.0.0p2 added both algorithms).
**`Add-WindowsCapability` does not help** — it installs the same in-box 9.5p2. Mind
the `PATH` order: the MSI appends `C:\Program Files\OpenSSH`, but
`C:\Windows\System32\OpenSSH` is already ahead of it, so a bare `ssh` still resolves
to the old binary. Confirm with `(Get-Command ssh).Source`.
2. **`SSH_ALLOW_CLASSIC_KEX=1`** — also offer `curve25519-sha256` (and its
`@libssh.org` spelling). Off by default. A session that negotiates it is secure
against a classical attacker but has **no** store-now-decrypt-later protection.
Modern clients still pick a PQ method, because the client's preference order
decides. Set it per host rather than fleet-wide, and drop it once the old client
is gone.
> `harden-ssh.sh` rewrites `/etc/ssh/sshd_config` wholesale on every run, so hand
> edits to that file do not survive a re-run. Use the knob, not `sed`.
## Host firewall
[`scripts/harden-firewall.sh`](scripts/harden-firewall.sh) installs a
**deny-by-default** iptables baseline: `INPUT` drops everything except loopback,
established/related, ICMP, and SSH on the configured port — plus any ports a
deployment registers. `OUTPUT` stays open and `FORWARD` is left untouched, so
Docker container networking is unaffected. The harden scripts and
`cloud-init/base.yml` / `jumphost.yml` install it automatically
(`ENABLE_FIREWALL=1` by default; set `0` to fall back to the minimal
sshguard-only jump).
**deny-by-default** baseline, with the backend chosen per family (override with
`FW_BACKEND=iptables|firewalld`):
- **Alpine / Debian → iptables.** `INPUT` drops everything except loopback,
established/related, ICMP, and SSH on the configured port — plus any ports a
deployment registers.
- **Alma / RHEL → firewalld** (its native firewall). The default zone is already
deny-by-default; we strip the stock `ssh`/`cockpit` services, open SSH +
registered ports, and let sshguard block via the `sshguard-firewalld` backend
(no `INPUT → sshguard` jump needed).
- **Proxmox → nothing. The host is skipped** — and left with no host firewall
until you enable Proxmox's own. See below.
`OUTPUT`/egress stays open and `FORWARD` is left untouched, so Docker container
networking is unaffected. The harden scripts and `cloud-init/base.yml` /
`jumphost.yml` install it automatically (`ENABLE_FIREWALL=1` by default).
- **Configurable SSH port** — read live from `sshd_config`, so a bastion on
`2222` is firewalled correctly with no extra flags. Restrict the source with
`FW_SSH_SOURCE=<cidr>`; drop ping with `FW_ALLOW_PING=0`.
- **Native persistence, no boot hook** — rules are saved and restored by the
distro's own package: `iptables` + `ip6tables` (Alpine/OpenRC),
`iptables-persistent` (Debian), or `iptables-services` (Alma). The saved
ruleset carries the `INPUT → sshguard` jump, so brute-force protection
survives reboot without a custom hook.
- **Native persistence, no boot hook** — on iptables hosts the ruleset is saved
and restored by the distro's own package: `iptables` + `ip6tables`
(Alpine/OpenRC) or `iptables-persistent` (Debian); the saved ruleset carries
the `INPUT → sshguard` jump. On firewalld hosts every change is `--permanent`,
so it persists across reboot natively and sshguard manages its own blocks.
- **Scripted additions** — deployments drop a rule file and re-apply:
```sh
printf '80/tcp\n443/tcp\n' > /etc/firewall/ports.d/mystack.rule
@@ -183,11 +253,95 @@ sshguard-only jump).
- **Docker caveat** — containers published with `-p` (e.g. Caddy's 80/443)
reach the host through nat/`FORWARD` and **bypass `INPUT`**, so the firewall
neither blocks nor needs to open them; the per-stack rule files are
belt-and-braces for any host-bound bind and self-documentation.
- **Recovery** — `harden-firewall.sh disable` flushes the rules and sets `INPUT`
back to `ACCEPT` (persisted), should a rule ever lock you out. A re-apply never
drops the live SSH session: the established-connection accept is added before
the policy flips to `DROP`.
belt-and-braces for any host-bound bind and self-documentation. The one
stack that runs with `network_mode: host` (`ergo`) binds on the host
directly, so there the registered ports are the real gate.
- **Recovery** — `harden-firewall.sh disable` un-locks you: on iptables it
flushes the rules and sets `INPUT` back to `ACCEPT` (persisted); on firewalld it
re-opens SSH (the `ssh` service + the configured port) and leaves firewalld
running. A re-apply never drops the live SSH session — on iptables the
established-connection accept is added before the policy flips to `DROP`, and
firewalld reloads preserve established connections.
### Proxmox hosts are skipped
Proxmox VE and Proxmox Mail Gateway are Debian underneath, so everything else in
this repo treats them as Debian — but they already ship a firewall, and
`pve-firewall` owns the host ruleset. The conflict is not the one you'd expect:
`pve-firewall` does **not** delete third-party rules. It restores with
`iptables-restore -n` (`--noflush`), only ever flushes chains matching its own
patterns (`PVEFW-*`, `tapNiM-*`, `vethNiM-*`, `fwbrN-*`, `GROUP-*`), *appends*
`-A INPUT -j PVEFW-INPUT` only when that hook is missing, and never sets a
built-in chain's policy. Our rules would survive it fine. The damage runs the
other way — **we break Proxmox**:
- `-P INPUT DROP` is ours alone, and `PVEFW-HOST-IN` **returns** on accept rather
than accepting (it still has to check the tap rules), so traffic Proxmox
explicitly allowed falls out of its chain and lands on our `DROP`. We silently
override the platform's own accepts;
- the persistence layer is worse: `netfilter-persistent` restores at boot with a
full `iptables-restore` (no `--noflush`), wiping PVE's `-j PVEFW-INPUT` hook
along with everything else until the daemon re-appends it on its next ~10s pass;
- a deny-by-default chain has to enumerate the whole platform to stay usable:
`8006/tcp` web UI, `5405-5412/udp` corosync (5405 + knet link number, up to 8
links — the bare `5405` is the pre-6.x multicast-era number), `60000-60050/tcp`
migration, `5900-5999/tcp` VNC, `3128/tcp` SPICE, `22/tcp` SSH, `111/udp`
rpcbind with NFS storage, plus Ceph's `6789`/`3300`/`6800-7300` when
hyperconverged. Miss one and you lose the GUI or the cluster; miss `-i lo` and
`pveproxy` can't reach `pvedaemon` on `127.0.0.1:85`, breaking the API locally;
- under the opt-in **nftables** backend (PVE 8.2+), `proxmox-firewall` registers
its own nft input hook, and an nft `DROP` beats an iptables `ACCEPT` — our
rules wouldn't even be authoritative.
So `harden-firewall.sh` detects Proxmox (`is_proxmox` in `oslib.sh`: it looks
for `pve-firewall`/`pveversion`/`/etc/pve/nodes`) and does nothing — `apply`
explains and exits cleanly, `allow`/`deny` refuse loudly rather than pretend,
and `list` shows `pve-firewall status`. `harden-ssh.sh` / `harden-jumphost.sh`
skip the firewall too and install only the `INPUT → sshguard` boot hook, so
brute-force protection still works while the Proxmox firewall is off. Manage the
host firewall where Proxmox expects it — *Datacenter → Firewall* and
*Node → Firewall*, or the `.fw` files directly:
```
[RULES]
IN ACCEPT -p tcp -dport 443
```
`FW_IGNORE_PVE=1` forces our firewall on anyway. It will fight `pve-firewall`
and can lock you out of the GUI and the cluster; Proxmox Backup Server ships no
firewall of its own and is *not* detected, so it hardens as a normal Debian host.
> **Skipped is not the same as protected.** Proxmox's firewall is **off by
> default** — the cluster-wide `enable` in `cluster.fw` defaults to `0`, and while
> it is, the `pve-firewall` daemon actively tears its chains down every ~10 seconds,
> leaving `INPUT` at policy `ACCEPT` with no rules. Until you enable it at
> *Datacenter → Firewall → Options*, a Proxmox host has **no host firewall at all**
> and `8006`, `22`, `3128` and `111` are open on every interface. Don't be reassured
> by the *node* panel reading `Firewall: Yes` — that setting is ignored while the
> datacenter one reads `No`. On these hosts, do your filtering in Proxmox or upstream
> of it.
The `INPUT → sshguard` jump the harden scripts install *is* safe alongside
`pve-firewall`: it is inserted with `-I`, so it sits ahead of the appended
`PVEFW-INPUT` hook and keeps getting first look at new connections either way.
**Already hardened a Proxmox host?** Detection only helps hosts set up from now
on, so `disable` cleans up one that already has our firewall:
```sh
bash scripts/harden-firewall.sh disable
```
On Proxmox that sub-command checks for leftovers (`/etc/firewall`, the engine, or
a `DROP` policy) and, if it finds them, sets `INPUT` back to `ACCEPT` and flushes
it — policy first, so it never drops the SSH session you are running it over —
deletes `/etc/firewall` and `/usr/local/sbin/firewall-apply`, **disables** the
boot-restore service and renames the saved rulesets to `*.bak-harden-firewall`
(saving the open state, as the normal `disable` does, would snapshot
`pve-firewall`'s own `PVEFW-*` chains and restore that stale copy at the next
boot), re-inserts the `INPUT → sshguard` jump, and restarts `pve-firewall`. On a
host that was correctly skipped it finds nothing and says so. `apply` points you
at it when it spots leftovers.
## SSH login notifications
+120 -29
View File
@@ -1,4 +1,4 @@
#!/usr/bin/env bash
#!/bin/sh
#
# automations.sh -- one command to run or deploy anything in this repo.
#
@@ -6,49 +6,119 @@
#
# 1. One-liner on a fresh target host (clones the repo, then launches):
# curl -fsSL https://git.anomalous.dev/57_Wolve/automations/raw/branch/main/automations.sh \
# | REPO_URL=https://git.anomalous.dev/57_Wolve/automations.git bash
# | REPO_URL=https://git.anomalous.dev/57_Wolve/automations.git sh
#
# 2. From a clone:
# ./automations.sh
#
# It opens a Gum wizard (auto-installed) that lets you:
# • Mode: deploy on THIS host, or build deploy.sh artifacts locally.
# • Pick any deployment (pocket-id, beszel, headscale, webfinger, simplex)
# or any generic script (harden-ssh, harden-jumphost, sshuser).
# • Pick any deployment (pocket-id, beszel, headscale, webfinger, squid,
# copyparty, simplex, openbao, ergo) or any generic script (setup-host,
# harden-ssh, harden-jumphost, sshuser, auto-update).
# Shared defaults come from globals/ (see globals/README.md).
#
# Non-interactive: set SKIP_PROMPTS=1 plus the needed vars and pipe the menu
# choices in, or just call the underlying deployments/<name>/deploy.sh
# directly -- they all honor SKIP_PROMPTS=1.
set -euo pipefail
# ============================================================================
# PROLOGUE -- POSIX sh only. Everything below the "exec bash" handoff is bash.
#
# The shebang is /bin/sh, not bash, on purpose: a stock Alpine box has busybox
# ash and NO bash at all, so a `#!/usr/bin/env bash` launcher dies before it
# can install anything ("env: 'bash': No such file or directory"). This part
# therefore has to parse and run under ash: no [[ ]], no arrays, no
# BASH_SOURCE, no printf -v. It locates (or clones) the repo, makes sure bash
# exists, and re-execs this same file under bash -- which then skips the
# prologue via BASH_VERSION and runs the real launcher.
# ============================================================================
set -eu
_boot_log() { printf '\033[1;32m[+]\033[0m %s\n' "$*"; }
_boot_die() { printf '\033[1;31m[x]\033[0m %s\n' "$*" >&2; exit 1; }
# Install packages with whichever manager this distro has -- apk (Alpine),
# apt-get (Debian/Ubuntu), dnf/yum (Alma/RHEL). oslib.sh's pkg_install can't
# help here: it's bash, and on the piped path it isn't even on disk yet.
_boot_install() {
if command -v apk >/dev/null 2>&1; then
apk add -q "$@" 2>/dev/null && return 0
apk update -q >/dev/null 2>&1 || true # stale/absent index on a fresh box
apk add -q "$@"
elif command -v apt-get >/dev/null 2>&1; then
apt-get update -qq && DEBIAN_FRONTEND=noninteractive apt-get install -y -qq "$@"
elif command -v dnf >/dev/null 2>&1; then
dnf install -y -q "$@"
elif command -v yum >/dev/null 2>&1; then
yum install -y -q "$@"
else
return 1
fi
}
_boot_need() { # _boot_need <command> [package] -> 0 if it's available afterwards
if command -v "$1" >/dev/null 2>&1; then return 0; fi
_boot_log "$1 not found; installing it..."
_boot_install "${2:-$1}" || true
command -v "$1" >/dev/null 2>&1
}
# ----------------------------------------------------------------------------
# Self-locate, or bootstrap by cloning the repo (one-liner / piped form).
# ----------------------------------------------------------------------------
_self="${BASH_SOURCE[0]:-}"
if [[ -n "$_self" && -f "$(cd "$(dirname "$_self")" 2>/dev/null && pwd)/scripts/lib.sh" ]]; then
ROOT="$(cd "$(dirname "$_self")" && pwd)"
ROOT=""
# Strip the last path component ourselves rather than calling dirname: busybox
# dirname takes its first argument literally, so `dirname -- "$0"` would answer
# "." on Alpine. `$0` is "sh"/"bash" (no slash) when we're piped from curl.
case "$0" in
*/*) _dir="${0%/*}" ;;
*) _dir="." ;;
esac
_dir="$(CDPATH= cd "$_dir" 2>/dev/null && pwd)" || _dir=""
if [ -n "$_dir" ] && [ -f "$_dir/scripts/lib.sh" ]; then
ROOT="$_dir"
else
# Piped via curl: we don't have the repo on disk. Clone it, then re-exec.
# Piped via curl: we don't have the repo on disk. Clone it, then hand off.
: "${REPO_URL:=}"
: "${REPO_BRANCH:=main}"
[[ -n "$REPO_URL" ]] || {
echo "[x] Running standalone (piped). Set REPO_URL=... so I can clone the repo." >&2
exit 1
}
command -v git >/dev/null 2>&1 || { command -v apk >/dev/null 2>&1 && apk add -q git; }
[ -n "$REPO_URL" ] || _boot_die "Running standalone (piped). Set REPO_URL=... so I can clone the repo."
_boot_need git || _boot_die "git is required to clone the repo, but it isn't installed and I couldn't install it automatically (need root + a supported package manager). Install git, then re-run."
_tmp="$(mktemp -d -t automations.XXXXXX)"
echo "[+] Cloning $REPO_URL ($REPO_BRANCH)..."
_boot_log "Cloning $REPO_URL ($REPO_BRANCH)..."
git clone --depth 1 --branch "$REPO_BRANCH" "$REPO_URL" "$_tmp"
exec bash "$_tmp/automations.sh" "$@"
ROOT="$_tmp"
fi
# ----------------------------------------------------------------------------
# Hand off to bash. Needed when we're running under ash/dash, and when the body
# we want is the freshly cloned copy rather than the piped stdin we came from.
# ----------------------------------------------------------------------------
if [ -z "${BASH_VERSION:-}" ] || [ ! -f "$0" ] || [ "$ROOT" != "$_dir" ]; then
if [ "${_AUTOMATIONS_REEXEC:-0}" = 1 ]; then
# Already handed off once. If we're in bash the handoff worked and only
# the path comparison differs (symlinked checkout) -- just continue.
[ -n "${BASH_VERSION:-}" ] || _boot_die "Re-exec under bash did not take effect. Run it explicitly: bash $ROOT/automations.sh"
else
# The launcher, everything it sources (scripts/lib.sh, scripts/oslib.sh),
# and every deploy.sh it invokes are bash. Alpine images routinely ship
# without it, so install it before going any further.
_boot_need bash || _boot_die "bash is required, but it isn't installed and I couldn't install it automatically (need root + a supported package manager). Install bash, then re-run."
_AUTOMATIONS_REEXEC=1; export _AUTOMATIONS_REEXEC
exec bash "$ROOT/automations.sh" "$@"
fi
fi
# ============================================================================
# Running under bash from here down.
# ============================================================================
set -euo pipefail
# shellcheck source=scripts/lib.sh
. "$ROOT/scripts/lib.sh"
load_globals
DEPLOYMENTS=(pocket-id beszel headscale webfinger simplex kanrisha)
DEPLOYMENTS=(pocket-id beszel headscale webfinger squid copyparty simplex openbao ergo knot-dns)
SCRIPTS=(setup-host harden-ssh harden-jumphost sshuser auto-update)
# ----------------------------------------------------------------------------
@@ -94,25 +164,46 @@ ask_deployment_vars() {
ask ISSUER_URL "OIDC issuer URL (e.g. https://auth.example.com)"
ask REDIRECT_URL "Redirect target for other traffic (e.g. https://example.org)"
ask ACME_EMAIL "Let's Encrypt email" ;;
squid)
ask TRUSTED_CIDR "Trusted client CIDR(s) allowed to use the proxy (e.g. 100.64.0.0/10)"
ask BIND_ADDR "Host IP to bind the proxy on (blank = 0.0.0.0)" optional
ask CACHE_SIZE_MB "On-disk cache size in MB (blank = 5000)" optional
ask CACHE_ONLY_LISTED "Cache ONLY listed domains? (1=yes, blank=boost mode)" optional ;;
copyparty)
ask COPYPARTY_DOMAIN "Public hostname for the web UI (e.g. files.example.com)"
ask ACME_EMAIL "Let's Encrypt email"
ask DATA_DIR "Host data folder shared as the root (blank = /srv/copyparty/data)" optional
ask FTP_NAT "Public IP for passive FTPS via NAT (blank = none)" optional
ask UPDATE_POLICY "Auto-update policy: latest | security | off (blank = latest)" optional ;;
simplex)
ask DOMAIN "Apex domain (creates smp.DOMAIN, xftp.DOMAIN)"
ask ACME_EMAIL "Let's Encrypt email"
ask XFTP_QUOTA "XFTP disk quota" optional
ask SSH_PORT "SSH port" optional
ask ALLOWED_IP "Your IP to whitelist in sshguard" optional ;;
kanrisha)
# License: reuse globals.env value if present, else prompt for it.
if [[ -z "${SCOUTFS_LICENSE:-}" ]]; then
ask SCOUTFS_LICENSE "ScoutFS license key"
else
ENVS+=("SCOUTFS_LICENSE=$SCOUTFS_LICENSE")
openbao)
ask OPENBAO_ADDR "LAN address the Kanrisha tape host reaches the vault at (IP or DNS)"
ask OPENBAO_BIND "Host IP to bind the API on (blank = 0.0.0.0)" optional ;;
knot-dns)
ask NODE_ID "Node ID for server.identity / NSID (e.g. ANYCAST-DNS-3)"
ask ROLE "Role: primary | secondary"
ask LISTEN "Listen addresses (blank = 0.0.0.0@53, ::@53)" optional
# A secondary's TSIG keys must byte-match the primary's, so they are
# copied from it rather than generated here.
if [[ "${ROLE:-}" == secondary ]]; then
ask PRIMARY_ADDR "Address of the primary this node transfers from"
ask TSIG_AUTHORITIVE "TSIG authortive-tsig secret (from the primary)" password
ask TSIG_ADMIN "TSIG admin-tsig secret (from the primary)" password
fi
# Blank path -> phase 1 (host prep). A live path -> phase 2 (deploy).
ask SCOUTFS_PATH "ScoutFS mount path (blank = host prep / phase 1)" optional
if [[ -n "${SCOUTFS_PATH:-}" ]]; then
ask KANRISHA_DOMAIN "Public hostname (e.g. kanrisha.example.com)"
ask ACME_EMAIL "Let's Encrypt email"
fi ;;
ask DATACENTER "Data centre label for the MOTD" optional
ask PEERS "Peer addresses for 'knsctl serials' (space-separated)" optional ;;
ergo)
ask ERGO_DOMAIN "IRC server hostname (e.g. irc.example.com)"
ask ACME_EMAIL "Let's Encrypt email"
ask NETWORK_NAME "IRC network name, no spaces (blank = the hostname)" optional
ask HISTORY "Persistent message history: sqlite | postgres | off (blank = sqlite)" optional
ask PLAINTEXT "Also serve PUBLIC plaintext IRC on 6667? (1 = yes, blank = no)" optional
ask UPDATE_POLICY "Auto-update policy: latest | security | off (blank = latest)" optional ;;
esac
}
+28 -10
View File
@@ -49,34 +49,52 @@ PAYLOAD="$(make_tar | base64 | tr -d '\n')"
# ---------------------------------------------------------------------------
{
cat <<'STUB'
#!/usr/bin/env bash
#!/bin/sh
#
# automations-bundle.sh -- self-extracting bundle of the automations repo.
# Generated by build-bundle.sh. Download, then run (it can't extract from a
# pipe -- it needs to read itself as a file):
#
# curl -fsSLO https://your-host/automations-bundle.sh
# bash automations-bundle.sh # launcher wizard
# bash automations-bundle.sh bash scripts/setup-host.sh # run a script
# SSH_PORT=2222 bash automations-bundle.sh bash scripts/harden-jumphost.sh
# sh automations-bundle.sh # launcher wizard
# sh automations-bundle.sh bash scripts/setup-host.sh # run a script
# SSH_PORT=2222 sh automations-bundle.sh bash scripts/harden-jumphost.sh
#
# POSIX sh, not bash: a stock Alpine host has busybox ash and no bash, so a
# bash stub would fail before it could install anything. It installs bash
# itself (everything inside the bundle is bash) and then runs the launcher.
#
# Env:
# INSTALL_DIR where to extract (default /opt/automations)
# BUNDLE_KEEP 1 to keep the extracted repo (default), 0 to use a temp dir
# and remove it after the command finishes
set -euo pipefail
set -eu
: "${INSTALL_DIR:=/opt/automations}"
: "${BUNDLE_KEEP:=1}"
SELF="${BASH_SOURCE[0]:-$0}"
if [[ ! -f "$SELF" ]]; then
SELF="$0"
if [ ! -f "$SELF" ]; then
echo "[x] Run me as a downloaded file, not via a pipe:" >&2
echo " curl -fsSLO <url>/automations-bundle.sh && bash automations-bundle.sh" >&2
echo " curl -fsSLO <url>/automations-bundle.sh && sh automations-bundle.sh" >&2
exit 1
fi
if [[ "$BUNDLE_KEEP" != "1" ]]; then
# The repo contents -- launcher, scripts/, deployments/ -- are all bash.
if ! command -v bash >/dev/null 2>&1; then
echo "[+] bash not found; installing it..." >&2
if command -v apk >/dev/null 2>&1; then { apk add -q bash 2>/dev/null || { apk update -q && apk add -q bash; }; } || true
elif command -v apt-get >/dev/null 2>&1; then { apt-get update -qq && DEBIAN_FRONTEND=noninteractive apt-get install -y -qq bash; } || true
elif command -v dnf >/dev/null 2>&1; then dnf install -y -q bash || true
elif command -v yum >/dev/null 2>&1; then yum install -y -q bash || true
fi
command -v bash >/dev/null 2>&1 || {
echo "[x] bash is required, but it isn't installed and I couldn't install it automatically (need root + a supported package manager). Install bash, then re-run." >&2
exit 1
}
fi
if [ "$BUNDLE_KEEP" != "1" ]; then
INSTALL_DIR="$(mktemp -d -t automations.XXXXXX)"
trap 'rm -rf "$INSTALL_DIR"' EXIT
fi
@@ -88,7 +106,7 @@ sed -e '1,/^__ARCHIVE_BELOW__$/d' "$SELF" | base64 -d | tar -xz -C "$INSTALL_DIR
cd "$INSTALL_DIR"
chmod +x automations.sh build-bundle.sh scripts/*.sh deployments/*/*.sh 2>/dev/null || true
if [[ "$#" -gt 0 ]]; then
if [ "$#" -gt 0 ]; then
exec "$@"
else
exec bash ./automations.sh
+6 -1
View File
@@ -29,7 +29,11 @@ runcmd:
DATACENTER="Globally Everywhere"
SSH_PORT=22
ALLOWED_IP= # optional: whitelist your client IP in sshguard
ENABLE_FIREWALL=1 # deny-by-default host firewall (0 to skip)
SSH_ALLOW_CLASSIC_KEX=0 # 1 = also offer curve25519-sha256 for clients
# with no post-quantum KEX (old Windows ssh.exe);
# costs store-now-decrypt-later protection
ENABLE_FIREWALL=1 # deny-by-default host firewall (0 to skip;
# always skipped on Proxmox -- pve-firewall owns it)
OPEN_PORTS="" # extra inbound ports, e.g. "80/tcp 443/tcp"
# ==================
@@ -53,5 +57,6 @@ runcmd:
# SSH hardening (key-only, PQ KEX, sshguard) + deny-by-default host firewall.
SSH_PORT="$SSH_PORT" ALLOWED_IP="$ALLOWED_IP" \
SSH_ALLOW_CLASSIC_KEX="$SSH_ALLOW_CLASSIC_KEX" \
ENABLE_FIREWALL="$ENABLE_FIREWALL" OPEN_PORTS="$OPEN_PORTS" \
FORCE=1 bash scripts/harden-ssh.sh
+6 -1
View File
@@ -25,7 +25,11 @@ runcmd:
DATACENTER="Globally Everywhere"
SSH_PORT=22
ALLOWED_IP= # optional: whitelist your client IP
ENABLE_FIREWALL=1 # deny-by-default host firewall (0 to skip)
SSH_ALLOW_CLASSIC_KEX=0 # 1 = also offer curve25519-sha256 for clients
# with no post-quantum KEX (old Windows ssh.exe);
# costs store-now-decrypt-later protection
ENABLE_FIREWALL=1 # deny-by-default host firewall (0 to skip;
# always skipped on Proxmox -- pve-firewall owns it)
JUMP_TARGETS="10.0.0.5:22 10.0.0.6:22" # hosts jumpers may ProxyJump to
# Optional login notifications (pam_exec -> ntfy). Leave NTFY_URL empty to
# skip. NTFY_REGION defaults to the region segment of this host's FQDN.
@@ -56,6 +60,7 @@ runcmd:
# Bastion hardening (admins shell + jumpers ProxyJump whitelist + optional
# login notifications).
SSH_PORT="$SSH_PORT" ALLOWED_IP="$ALLOWED_IP" JUMP_TARGETS="$JUMP_TARGETS" \
SSH_ALLOW_CLASSIC_KEX="$SSH_ALLOW_CLASSIC_KEX" \
ENABLE_FIREWALL="$ENABLE_FIREWALL" \
NTFY_URL="$NTFY_URL" NTFY_TOKEN="$NTFY_TOKEN" NTFY_EMAIL="$NTFY_EMAIL" NTFY_REGION="$NTFY_REGION" \
FORCE=1 bash scripts/harden-jumphost.sh
+65 -2
View File
@@ -7,7 +7,9 @@
# What this does:
# 1. Installs docker + docker-cli-compose if missing.
# 2. Lays down docker-compose.yml, Caddyfile, .env.example in $STACK_DIR.
# 3. Generates .env on first run; existing .env is never overwritten.
# 3. Generates .env on first run. On a re-run an existing .env is kept,
# except for values passed explicitly that run, which are written
# through -- .env is what compose interpolates into the container.
# 4. Prompts for required values not preset (BESZEL_DOMAIN, ACME_EMAIL).
# 5. Opens TCP 80/443 in UFW if active.
# 6. Pulls images, brings the stack up, waits for healthchecks.
@@ -32,6 +34,18 @@ set -euo pipefail
: "${FORCE:=0}"
: "${SKIP_PROMPTS:=0}" # non-interactive: require values via env, no prompts
[[ "$SKIP_PROMPTS" == "1" ]] && FORCE=1
ENV_FILE="$STACK_DIR/.env"
# Both values reach Caddy only by compose interpolating .env, so .env has to absorb
# what this run passed or the run deploys the old values while reporting the
# new ones. Which keys arrived in the ENVIRONMENT must be recorded BEFORE the
# ":=" defaults below, which make an unset variable look like an empty one.
RUNTIME_KEYS=(BESZEL_DOMAIN ACME_EMAIL)
KEYS_FROM_ENV=()
for _k in "${RUNTIME_KEYS[@]}"; do
if [[ -n "${!_k+x}" ]]; then KEYS_FROM_ENV+=("$_k"); fi
done
: "${BESZEL_DOMAIN:=}"
: "${ACME_EMAIL:=}"
@@ -142,6 +156,24 @@ prompt() {
fi
}
# On a re-run the deployed values live in .env. Load them into any key NOT
# passed this run, so the prompts below do not force the operator to retype
# everything (and so SKIP_PROMPTS=1 does not die demanding values .env has).
if [[ -f "$ENV_FILE" ]]; then
_preloaded=()
for _k in "${RUNTIME_KEYS[@]}"; do
if [[ " ${KEYS_FROM_ENV[*]-} " == *" ${_k} "* ]]; then continue; fi
_v=$(sed -n "s/^${_k}=//p" "$ENV_FILE" | tail -n1)
if [[ -z "$_v" ]]; then continue; fi
printf -v "$_k" '%s' "$_v"
_preloaded+=("$_k")
done
if (( ${#_preloaded[@]} > 0 )); then
log "Reusing ${ENV_FILE} values for: ${_preloaded[*]}"
log " (pass VAR=... to change one; it is then written back to .env)"
fi
fi
prompt BESZEL_DOMAIN "Public hostname for beszel (e.g. monitoring.example.com)"
prompt ACME_EMAIL "Let's Encrypt email"
@@ -163,6 +195,28 @@ install -d -m 0750 "$STACK_DIR"
install -m 0640 "$SCRIPT_DIR/docker-compose.yml" "$STACK_DIR/docker-compose.yml"
install -m 0640 "$SCRIPT_DIR/Caddyfile" "$STACK_DIR/Caddyfile"
set_env() { # <KEY> <value>: update KEY in .env, or append if absent
# The value goes through the ENVIRONMENT, never interpolated into a sed
# script. Interpolating it corrupts any value containing & (sed expands it to
# the whole match) and aborts the run on one containing the s||| delimiter --
# which is reachable for an OIDC issuer URL or a redirect query string.
local key="$1" val="$2" tmp
if [[ ! -f "$ENV_FILE" ]]; then
printf '%s=%s
' "$key" "$val" >> "$ENV_FILE"
return 0
fi
tmp="$(mktemp)"
_SE_KEY="$key" _SE_VAL="$val" awk '
BEGIN { k = ENVIRON["_SE_KEY"]; v = ENVIRON["_SE_VAL"]; seen = 0 }
!seen && index($0, k "=") == 1 { print k "=" v; seen = 1; next }
{ print }
END { if (!seen) print k "=" v }
' "$ENV_FILE" > "$tmp"
cat "$tmp" > "$ENV_FILE" # rewrite in place: keeps the original mode/owner
rm -f "$tmp"
}
ENV_FILE="$STACK_DIR/.env"
if [[ ! -f "$ENV_FILE" ]]; then
log "Seeding $ENV_FILE..."
@@ -172,7 +226,16 @@ if [[ ! -f "$ENV_FILE" ]]; then
-e "s|^ACME_EMAIL=.*|ACME_EMAIL=${ACME_EMAIL}|" \
"$ENV_FILE"
else
log ".env exists; leaving it alone."
# Keep the file, but absorb what was passed this run: compose reads .env,
# so letting it drift means deploying the old values while the banner
# prints the new ones. Keys not passed are left untouched.
log ".env exists; keeping it (only values passed this run are updated)."
for _k in ${KEYS_FROM_ENV[@]+"${KEYS_FROM_ENV[@]}"}; do
_cur=$(sed -n "s/^${_k}=//p" "$ENV_FILE" | tail -n1)
if [[ "$_cur" == "${!_k}" ]]; then continue; fi
set_env "$_k" "${!_k}"
log " ${_k}: ${_cur:-<unset>} -> ${!_k:-<empty>}"
done
fi
# Validate
+50
View File
@@ -0,0 +1,50 @@
# Copy to .env and fill in. docker compose picks .env up automatically.
# --- Public hostname (web UI / WebDAV via Caddy) ----------------------------
# Bare hostname (no scheme) where the copyparty web UI is reached.
COPYPARTY_DOMAIN=files.example.com
# Email for Let's Encrypt registration / expiry notifications.
ACME_EMAIL=admin@example.com
# --- Image -----------------------------------------------------------------
# `ac` is the recommended edition (Pillow + FFmpeg thumbnails/transcode).
# Use ghcr.io/9001/copyparty-ac to pull from GitHub instead of Docker Hub.
COPYPARTY_IMAGE=copyparty/ac
# Pinned version, e.g. 1.20.11. deploy.sh resolves "latest" to the newest
# release on first deploy so the running version is always explicit; update.sh
# bumps this tag when it updates.
COPYPARTY_TAG=latest
# --- Data + identity --------------------------------------------------------
# Host directory shared as the data root (mounted at /w). Created by deploy.sh
# and chowned to PUID:PGID. Point this at a big disk / mount for real use.
DATA_DIR=/srv/copyparty/data
PUID=1000
PGID=1000
# --- SFTP / FTPS (direct listeners, bypass Caddy) ---------------------------
# Host interface to publish the direct listeners on (blank = 0.0.0.0 = all).
# A published Docker port bypasses the host firewall, so pin this to a trusted
# interface (e.g. a Tailscale IP) on multi-homed boxes.
BIND_ADDR=
SFTP_PORT=3922
FTPS_PORT=3990
# Passive-FTP TCP port range (hyphen form; published as-is in docker-compose).
FTP_PASV_RANGE=12000-12099
# Public IP for passive FTPS through NAT. Leave blank if clients reach the
# server directly; set it to the server's external IP otherwise.
FTP_NAT=
# --- Updates (scripts: update.sh) -------------------------------------------
# How the scheduled updater behaves:
# latest - update to the newest release whenever one exists (default)
# security - update ONLY when the running version has a known advisory
# off - never change the running version (check/notify only)
UPDATE_POLICY=latest
# Advisory feed used by `security` policy + `update.sh check`:
# advisories-panic (critical only) | advisories (important) | advisories-all
VC_FEED=advisories
# --- Image tags -------------------------------------------------------------
CADDY_TAG=2-alpine
+42
View File
@@ -0,0 +1,42 @@
# Caddyfile for the copyparty stack.
#
# Auto-issues a Let's Encrypt cert for $COPYPARTY_DOMAIN and reverse-proxies the
# web UI / WebDAV to copyparty on :3923. SFTP and FTPS do NOT pass through here
# -- they are published straight from the copyparty container.
{
email {$ACME_EMAIL}
}
{$COPYPARTY_DOMAIN} {
encode zstd gzip
reverse_proxy copyparty:3923 {
# Real client IP: Caddy already appends the genuine remote address to
# X-Forwarded-For (it is the internet edge); copyparty trusts it via
# `xff-src` in copyparty.conf. X-Real-IP is set too for good measure.
header_up X-Real-IP {http.request.remote.host}
# Stream uploads/downloads (no buffering) and allow long-lived large
# transfers -- copyparty resumable uploads can run for a long time.
# Caddy imposes no request-body size limit by default, so big uploads
# stream straight through.
flush_interval -1
transport http {
read_timeout 1h
write_timeout 1h
}
}
header {
Strict-Transport-Security "max-age=31536000; includeSubDomains"
X-Content-Type-Options "nosniff"
Referrer-Policy "strict-origin-when-cross-origin"
-Server
}
log {
output stdout
format console
}
}
+155
View File
@@ -0,0 +1,155 @@
# copyparty
[copyparty](https://github.com/9001/copyparty) — a portable file server with
resumable uploads, dedup, WebDAV, **SFTP**, **FTPS**, media indexing and
thumbnails — behind Caddy with automatic Let's Encrypt TLS. Ships with a
**security-notices-aware updater** (`update.sh`).
The web UI / WebDAV is fronted by Caddy on 443; **SFTP and FTPS are served
directly by copyparty** (Caddy is HTTP-only and not in that path).
## Required `.env` values
| Variable | Notes |
|----------|-------|
| `COPYPARTY_DOMAIN` | Public hostname for the web UI (e.g. `files.example.com`). |
| `ACME_EMAIL` | Let's Encrypt registration email. |
| `DATA_DIR` | Host folder shared as the data root (mounted at `/w`). Point at a data disk. |
| `COPYPARTY_IMAGE` / `COPYPARTY_TAG` | Image edition (`copyparty/ac`) + pinned version. |
| `SFTP_PORT` / `FTPS_PORT` / `FTP_PASV_RANGE` | Direct-listener ports (`3922` / `3990` / `12000-12099`). |
| `FTP_NAT` | This host's public IP, for passive FTPS through NAT (blank if direct). |
| `UPDATE_POLICY` | `latest` (default) / `security` / `off` — see [Updates](#updates--security-notices). |
See [`.env.example`](.env.example) for the rest. Accounts and shares live in
[`copyparty.conf.example`](copyparty.conf.example) (deploy.sh generates the real
`cfg/copyparty.conf` with a random admin password on first run).
## Deploy
```bash
./automations.sh # Deploy on this host → deploy: copyparty
```
Or build + run the self-contained artifact:
```bash
./build.sh
scp deploy.sh root@host:
ssh root@host 'bash deploy.sh'
# non-interactive:
# COPYPARTY_DOMAIN=files.example.com ACME_EMAIL=me@example.com SKIP_PROMPTS=1 bash deploy.sh
```
Unattended provisioning: [`cloud-init.yml`](cloud-init.yml). On first run deploy.sh
pins `COPYPARTY_TAG` to the newest release, generates `cfg/copyparty.conf` (random
admin password, printed once) and a self-signed `cfg/ftps.pem`, opens the ports,
brings the stack up, and schedules `update.sh`.
## Accounts & shares
The default config is **private**: one `admin` account with full access to the
whole data root, nothing public. Edit `cfg/copyparty.conf` (commented examples
included) to add users or open read-only / anonymous-upload shares, then hot-reload:
```bash
docker compose exec copyparty kill -s USR1 1
```
## Real client IP (behind Caddy)
For the web UI, copyparty must see the genuine client IP (for its logs and abuse
bans), which means **both** sides are configured:
- **Caddy** is the internet edge, so it appends the real address to
`X-Forwarded-For` and also sets `X-Real-IP` (see [`Caddyfile`](Caddyfile)).
- **copyparty** trusts that upstream via `xff-src: lan` in
[`copyparty.conf.example`](copyparty.conf.example) — without it copyparty
ignores the header and logs/bans the Docker-network address.
`lan` (trust private ranges) is safe here because copyparty's HTTP port is never
published — only Caddy can reach it. If copyparty prints an `xff-src` warning at
startup, narrow it to the exact subnet it names. SFTP/FTPS are direct TCP, so
they already see the real client IP (no XFF involved).
## SFTP & FTPS
Both authenticate against the same `[accounts]` as the web UI.
```bash
sftp -P 3922 admin@files.example.com # SFTP (password auth enabled)
lftp -u admin -e 'set ftp:ssl-force true; set ssl:verify-certificate no' \
ftp://files.example.com:3990 # explicit-TLS FTPS
```
- **SFTP host key** is generated on first start and persisted under `/cfg`, so
clients don't get "host key changed" warnings across restarts.
- **FTPS cert** defaults to the self-signed `cfg/ftps.pem` (clients must accept
it). To serve a publicly-trusted cert, reuse Caddy's Let's Encrypt cert: mount
the `caddy-data` volume read-only into the copyparty service and point
copyparty at it, e.g.
```yaml
# docker-compose.yml, copyparty service:
volumes:
- caddy-data:/caddy:ro
```
```ini
# cfg/copyparty.conf [global]:
cert: /caddy/caddy/certificates/acme-v02.api.letsencrypt.org-directory/files.example.com/files.example.com.crt
certkey: /caddy/caddy/certificates/acme-v02.api.letsencrypt.org-directory/files.example.com/files.example.com.key
```
(copyparty re-reads the cert on `kill -s USR1`; refresh after a renewal.)
- **Passive FTPS** needs the `FTP_PASV_RANGE` ports published (they are) **and**
`FTP_NAT` set to the server's public IP when clients connect through NAT.
- A published Docker port bypasses the host `INPUT` firewall, so set `BIND_ADDR`
to pin the listeners to a trusted interface on multi-homed hosts.
## Updates & security notices
copyparty doesn't self-update, but it publishes a machine-readable
**security-advisories** feed (the same one its built-in `--vc-url` check uses).
This deployment uses it on **two** levels:
1. **In-app** — `vc-url` + `vc-age` in `cfg/copyparty.conf` make copyparty itself
check whether the running version has a known advisory and warn in the log /
control panel. (`vc-exit`, opt-in, makes it shut down if vulnerable.)
2. **[`update.sh`](update.sh)** — reads the same feed to *act*, pinning the new
version in `.env` and recreating the container (with health-checked rollback).
```bash
./update.sh check # report current vs latest + any advisory; change nothing
./update.sh update # update now (TARGET_VERSION=1.20.11 to pin a version)
./update.sh install # (re)schedule the daily run / uninstall to stop
```
`UPDATE_POLICY` (in `.env` or `/etc/copyparty-update.conf`) controls the
scheduled `run`:
| Policy | Behaviour |
|--------|-----------|
| `latest` *(default)* | Update to the newest release whenever one exists. |
| `security` | Update **only** when the running version has a known advisory, to the patched release named in the feed. |
| `off` | Never change the running version (check/notify only). |
`VC_FEED` selects the feed (`advisories-panic` / `advisories` / `advisories-all`).
Update results reuse the ntfy config at `/etc/ssh-notify.conf` (same as the SSH
login notifier / host auto-update), when present.
## Files
| File | Purpose |
|---|---|
| `docker-compose.yml` | caddy + copyparty; publishes SFTP/FTPS/passive ports. |
| `Caddyfile` | TLS + reverse proxy for the web UI/WebDAV (real-IP headers). |
| `copyparty.conf.example` | Config template (globals, security check, SFTP/FTPS, accounts, volumes). |
| `update.sh` | Security-notices-aware container updater. |
| `.env.example` | Stack tunables. |
| `deploy.sh` / `build.sh` | Self-contained installer + archive embedder. |
| `cloud-init.yml` | Fresh-VM bootstrap (harden SSH, then deploy). |
## Notes
- No Anubis PoW gate — it would break WebDAV / API / upload clients.
- The web UI/WebDAV go through Caddy; SFTP/FTPS bypass it by design.
- DNS for `COPYPARTY_DOMAIN` must resolve to the host and 80/443 be reachable
before deploy for the LE cert to issue.
+47
View File
@@ -0,0 +1,47 @@
#!/usr/bin/env bash
#
# build.sh -- (re)embed the loose deployment files into deploy.sh as a base64
# tar.gz payload after __ARCHIVE_BELOW__. Idempotent: strips any existing
# payload first.
#
# Run this after editing ANY embedded file below, then re-stage deploy.sh --
# the deployed stack uses the EMBEDDED copies, not the loose files.
set -euo pipefail
DIR=$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)
SCRIPT="$DIR/deploy.sh"
MARKER="__ARCHIVE_BELOW__"
# README, cloud-init and the build/deploy scripts are NOT embedded. The generated
# cfg/ (copyparty.conf + ftps.pem) is produced on the host at deploy time.
FILES=(
docker-compose.yml
Caddyfile
copyparty.conf.example
.env.example
update.sh
)
[[ -f "$SCRIPT" ]] || { echo "deploy.sh not found at $SCRIPT" >&2; exit 1; }
for f in "${FILES[@]}"; do
[[ -f "$DIR/$f" ]] || { echo "Missing $DIR/$f" >&2; exit 1; }
done
PAYLOAD=$(tar -czf - -C "$DIR" "${FILES[@]}" | base64)
TMP=$(mktemp)
trap 'rm -f "$TMP"' EXIT
sed "/^${MARKER}\$/,\$d" "$SCRIPT" > "$TMP"
{
echo "$MARKER"
echo "$PAYLOAD"
} >> "$TMP"
mv "$TMP" "$SCRIPT"
chmod +x "$SCRIPT"
trap - EXIT
size=$(wc -c < "$SCRIPT")
echo "Built $SCRIPT (${size} bytes)"
+41
View File
@@ -0,0 +1,41 @@
#cloud-config
#
# copyparty (file server) — harden SSH, then deploy, on a fresh host.
#
# Fill in REPO_URL and the values in the runcmd block, then paste this as the
# instance user-data. DNS for COPYPARTY_DOMAIN must point at this host and ports
# 80/443 must be reachable before boot, or the Let's Encrypt cert request fails.
# SFTP/FTPS (3922/3990 + the passive range) are opened too.
packages:
- git
runcmd:
- hostnamectl set-hostname copyparty || true
- |
set -e
REPO_URL=https://git.anomalous.dev/57_Wolve/automations.git
REPO_BRANCH=main
HARDEN_SSH=1 # harden SSH on this fresh VM (set 0 to skip)
SSH_PORT=22
ALLOWED_IP= # optional: whitelist your client IP in sshguard
git clone --depth 1 --branch "$REPO_BRANCH" "$REPO_URL" /opt/automations
cd /opt/automations
# Harden SSH first (PQ KEX, key-only auth, sshguard + deny-by-default
# firewall). The firewall it installs is what deploy.sh registers the
# copyparty ports with.
if [ "$HARDEN_SSH" = 1 ]; then
SSH_PORT="$SSH_PORT" ALLOWED_IP="$ALLOWED_IP" SKIP_PROMPTS=1 FORCE=1 \
bash scripts/harden-ssh.sh
fi
# Deploy copyparty. DATA_DIR is the shared folder (point it at a data disk);
# FTP_NAT should be this host's public IP if you want passive FTPS via NAT.
COPYPARTY_DOMAIN=files.example.com \
ACME_EMAIL=admin@example.com \
DATA_DIR=/srv/copyparty/data \
UPDATE_POLICY=latest \
FTP_NAT= \
SKIP_PROMPTS=1 \
bash deployments/copyparty/deploy.sh
@@ -0,0 +1,62 @@
# copyparty configuration (TEMPLATE).
#
# deploy.sh copies this to cfg/copyparty.conf on first deploy, replacing
# __ADMIN_PW__ with a generated password (and uncommenting ftp-nat if FTP_NAT is
# set). The image auto-includes every *.conf in /cfg (its baked-in config ends
# with `% /cfg`), so this file is loaded automatically -- no -c flag needed.
#
# Reload after editing, no restart: docker compose exec copyparty kill -s USR1 1
# Permission letters: r=read w=write/upload m=move/rename d=delete a=admin g=get-only
[global]
# -- Behind Caddy (real client IP) --
# Trust the reverse proxy's X-Forwarded-For. Only Caddy can reach :3923 (it is
# not published), so `lan` (trust private nets) is safe and robust; narrow it
# to the exact docker subnet copyparty names in a startup warning if you like.
xff-src: lan
# -- Security notices (copyparty's built-in vulnerability check) --
# Periodically checks whether THIS version has a known advisory and warns in
# the log / control panel. The companion update.sh acts on the same feed.
vc-url: https://api.copyparty.eu/advisories
vc-age: 3
# vc-exit # opt-in: shut down if running a vulnerable version -- only
# enable together with the security updater, or it will loop.
# -- Indexing / media --
e2dsa # up2k database + filesystem scan (dedup, search)
e2ts # parse media tags for the indexer
# -- SFTP + FTPS (direct listeners; NOT proxied by Caddy) --
sftp: 3922 # SFTP server
sftp-pw # allow password auth (not just ssh keys)
ftps: 3990 # explicit-TLS FTPS (plaintext `ftp` intentionally omitted)
ftp-pr: 12000-12099 # passive TCP port range (also published in docker-compose.yml)
cert: /cfg/ftps.pem # self-signed by deploy.sh; swap in Caddy's LE cert to taste
# ftp-nat: 203.0.113.10 # set to the server's public IP for passive FTPS via NAT
[accounts]
# deploy.sh replaces __ADMIN_PW__ with a generated password on first run.
admin: __ADMIN_PW__
# Default volume: the whole data root, private to admin. Open it up below.
[/]
/w
accs:
rwmda: admin
# --- examples (uncomment + adjust) -----------------------------------------
# Public read-only share of /w/pub for anyone (no login):
# [/pub]
# /w/pub
# accs:
# r: *
#
# Anonymous write-only dropbox (uploads, no listing or download):
# [/inbox]
# /w/inbox
# accs:
# w: *
#
# Add another account: put `bob: s3cret` under [accounts] above, then grant it
# access in a volume, e.g. `rw: bob` inside that volume's accs: block.
+638
View File
@@ -0,0 +1,638 @@
#!/usr/bin/env bash
#
# deploy.sh -- deploy copyparty (caddy + copyparty file server) on Alpine,
# Debian or Alma. Single-node; runs as root.
#
# What this does:
# 1. Installs docker + compose if missing.
# 2. Lays down the stack files in $STACK_DIR.
# 3. Generates cfg/copyparty.conf on first run (random admin password) and a
# self-signed cfg/ftps.pem for FTPS; neither is overwritten on re-run.
# 4. Generates .env on first run, pinning COPYPARTY_TAG to the newest release.
# 5. Prompts for required values not preset (COPYPARTY_DOMAIN, ACME_EMAIL).
# 6. Opens 80/443 + SFTP/FTPS/passive ports on the host firewall if present.
# 7. Pulls images, brings the stack up, waits for health.
# 8. Installs + schedules update.sh (security-notices-aware) unless disabled.
#
# Caddy fronts the web UI / WebDAV over TLS; SFTP and FTPS are published directly
# by copyparty (not proxied). copyparty trusts Caddy's X-Forwarded-For so logs
# and bans use the real client IP (see cfg/copyparty.conf `xff-src`).
#
# Idempotent: re-run to apply config changes / pull new images.
#
# Self-contained: docker-compose.yml, Caddyfile, copyparty.conf.example,
# .env.example and update.sh are embedded as a base64 tar.gz at the bottom of
# this file. Rebuild with build.sh after editing the loose source files.
#
# Usage:
# bash deploy.sh # interactive prompts
# COPYPARTY_DOMAIN=files.example.com ACME_EMAIL=me@x.com \
# SKIP_PROMPTS=1 bash deploy.sh # non-interactive
# DATA_DIR=/mnt/disk/files FTP_NAT=203.0.113.10 bash deploy.sh
# SKIP_DOCKER_INSTALL=1 bash deploy.sh
set -euo pipefail
: "${STACK_DIR:=/srv/copyparty}"
: "${SKIP_DOCKER_INSTALL:=0}"
: "${FORCE:=0}"
: "${SKIP_PROMPTS:=0}"
[[ "$SKIP_PROMPTS" == "1" ]] && FORCE=1
: "${COPYPARTY_DOMAIN:=}"
: "${ACME_EMAIL:=}"
: "${COPYPARTY_IMAGE:=copyparty/ac}"
: "${COPYPARTY_TAG:=latest}"
: "${PUID:=1000}"
: "${PGID:=1000}"
: "${DATA_DIR:=/srv/copyparty/data}"
: "${BIND_ADDR:=}"
: "${SFTP_PORT:=3922}"
: "${FTPS_PORT:=3990}"
: "${FTP_PASV_RANGE:=12000-12099}"
: "${FTP_NAT:=}"
: "${UPDATE_POLICY:=latest}"
: "${VC_FEED:=advisories}"
: "${COPYPARTY_AUTOUPDATE:=1}" # 0 = install update.sh but don't schedule it
: "${CADDY_TAG:=2-alpine}"
log() { printf '\033[1;32m[+]\033[0m %s\n' "$*"; }
warn() { printf '\033[1;33m[!]\033[0m %s\n' "$*" >&2; }
die() { printf '\033[1;31m[x]\033[0m %s\n' "$*" >&2; exit 1; }
[[ $EUID -eq 0 ]] || die "Run as root."
# ---------------------------------------------------------------------------
# OS detection + Docker install (Alpine / Debian / Alma). Inlined (this deploy.sh
# is self-contained / scp'd standalone) rather than sourced from oslib.sh.
# ---------------------------------------------------------------------------
osfam() {
local id="" like=""
if [[ -r /etc/os-release ]]; then
id="$(. /etc/os-release 2>/dev/null && echo "${ID:-}")"
like="$(. /etc/os-release 2>/dev/null && echo "${ID_LIKE:-}")"
fi
case " $id $like " in
*" alpine "*) echo alpine ;;
*" debian "*|*" ubuntu "*) echo debian ;;
*" rhel "*|*" fedora "*|*" centos "*) echo rhel ;;
*) echo "${id:-unknown}" ;;
esac
}
pkg_install() { # best-effort install of a package across the three families
case "$(osfam)" in
alpine) apk add -q "$@" || true ;;
debian) DEBIAN_FRONTEND=noninteractive apt-get install -y -qq "$@" || true ;;
rhel) dnf install -y -q "$@" || true ;;
esac
}
# fetch a URL to stdout with whatever's available (curl or wget).
fetch() {
if command -v curl >/dev/null 2>&1; then curl -fsSL "$1" 2>/dev/null
elif command -v wget >/dev/null 2>&1; then wget -qO- "$1" 2>/dev/null
fi
}
install_docker() {
if command -v docker >/dev/null 2>&1; then
log "Docker already installed: $(docker --version)"
else
log "Installing Docker (OS: $(osfam))..."
case "$(osfam)" in
alpine) apk add -q docker docker-cli-compose openrc ;;
debian|rhel) command -v curl >/dev/null 2>&1 || pkg_install curl
curl -fsSL https://get.docker.com | sh ;;
*) die "Unsupported OS for auto Docker install. Set SKIP_DOCKER_INSTALL=1 and install Docker yourself." ;;
esac
fi
if command -v rc-update >/dev/null 2>&1; then
rc-update add docker default >/dev/null 2>&1 || true
rc-service docker status >/dev/null 2>&1 || rc-service docker start
elif command -v systemctl >/dev/null 2>&1; then
systemctl enable --now docker >/dev/null 2>&1 || systemctl start docker || true
fi
}
open_web_ports() {
# Register web (80/443) + SFTP/FTPS + the passive-FTP range. Prefer the host
# firewall (harden-firewall.sh); else ufw/firewalld if active.
#
# NOTE: published Docker ports reach the host via nat/FORWARD and BYPASS the
# INPUT firewall, so this is belt-and-braces + self-documentation. Pin the
# real exposure with BIND_ADDR and copyparty's own access control.
local pasv_colon="${FTP_PASV_RANGE/-/:}" # 12000-12099 -> 12000:12099 (iptables/ufw/ports.d)
local pasv_hyphen="${FTP_PASV_RANGE}" # 12000-12099 (firewalld)
if [[ -d /etc/firewall/ports.d && -x /usr/local/sbin/firewall-apply ]]; then
log "Registering 80,443,${SFTP_PORT},${FTPS_PORT},${pasv_colon}/tcp with host firewall..."
printf '80/tcp\n443/tcp\n%s/tcp\n%s/tcp\n%s/tcp\n' \
"$SFTP_PORT" "$FTPS_PORT" "$pasv_colon" > /etc/firewall/ports.d/copyparty.rule
/usr/local/sbin/firewall-apply
elif command -v ufw >/dev/null 2>&1 && ufw status 2>/dev/null | grep -q '^Status: active'; then
log "ufw active -- allowing web + SFTP/FTPS..."
local p
for p in 80/tcp 443/tcp "${SFTP_PORT}/tcp" "${FTPS_PORT}/tcp" "${pasv_colon}/tcp"; do
ufw allow "$p" >/dev/null
done
elif command -v firewall-cmd >/dev/null 2>&1 && firewall-cmd --state >/dev/null 2>&1; then
log "firewalld active -- allowing web + SFTP/FTPS..."
firewall-cmd -q --add-service=http --permanent
firewall-cmd -q --add-service=https --permanent
firewall-cmd -q --add-port="${SFTP_PORT}/tcp" --permanent
firewall-cmd -q --add-port="${FTPS_PORT}/tcp" --permanent
firewall-cmd -q --add-port="${pasv_hyphen}/tcp" --permanent
firewall-cmd -q --reload
fi
}
# ----------------------------------------------------------------------------
# Extract embedded archive
# ----------------------------------------------------------------------------
SCRIPT_DIR=$(mktemp -d -t copyparty-deploy.XXXXXX)
trap 'rm -rf "$SCRIPT_DIR"' EXIT
extract_archive() {
grep -a -A 9999999 '^__ARCHIVE_BELOW__$' "$0" \
| tail -n +2 \
| base64 -d \
| tar -xz -C "$SCRIPT_DIR"
}
if grep -q -a '^__ARCHIVE_BELOW__$' "$0"; then
log "Extracting embedded deployment files..."
extract_archive
else
die "No embedded archive found. Run build.sh to embed deployment files."
fi
EMBEDDED=(docker-compose.yml Caddyfile copyparty.conf.example .env.example update.sh)
for f in "${EMBEDDED[@]}"; do
[[ -f "$SCRIPT_DIR/$f" ]] || die "Embedded archive missing $f"
done
# ----------------------------------------------------------------------------
# Prompt for required vars
# ----------------------------------------------------------------------------
prompt() {
local varname="$1" message="$2"
local -n ref="$varname"
if [[ -z "${ref:-}" ]]; then
[[ "$SKIP_PROMPTS" == "1" ]] && die "$varname required (set it in the environment; running with SKIP_PROMPTS=1)."
read -r -p "$message: " ref
[[ -n "$ref" ]] || die "$varname required."
fi
}
prompt COPYPARTY_DOMAIN "Public hostname for the web UI (e.g. files.example.com)"
prompt ACME_EMAIL "Let's Encrypt email"
# ----------------------------------------------------------------------------
# Docker + firewall
# ----------------------------------------------------------------------------
if [[ "$SKIP_DOCKER_INSTALL" != "1" ]]; then
install_docker
fi
open_web_ports
# ----------------------------------------------------------------------------
# Stack directory + files
# ----------------------------------------------------------------------------
log "Setting up $STACK_DIR..."
install -d -m 0750 "$STACK_DIR"
install -m 0644 "$SCRIPT_DIR/docker-compose.yml" "$STACK_DIR/docker-compose.yml"
install -m 0644 "$SCRIPT_DIR/Caddyfile" "$STACK_DIR/Caddyfile"
install -m 0755 "$SCRIPT_DIR/update.sh" "$STACK_DIR/update.sh"
install -d -m 0750 "$STACK_DIR/cfg"
# Data root (bind-mounted at /w), owned by the runtime UID/GID.
install -d -m 0750 "$DATA_DIR"
chown "$PUID:$PGID" "$DATA_DIR"
ENV_FILE="$STACK_DIR/.env"
set_env() { # <KEY> <value>: update KEY in .env, or append if absent
# The value goes through the ENVIRONMENT, never interpolated into a sed
# script. Interpolating it corrupts any value containing & (sed expands it to
# the whole match) and aborts the run on one containing the s||| delimiter --
# which is reachable for an OIDC secret, a password or a URL query string.
local key="$1" val="$2" tmp
if [[ ! -f "$ENV_FILE" ]]; then
printf '%s=%s
' "$key" "$val" >> "$ENV_FILE"
return 0
fi
tmp="$(mktemp)"
_SE_KEY="$key" _SE_VAL="$val" awk '
BEGIN { k = ENVIRON["_SE_KEY"]; v = ENVIRON["_SE_VAL"]; seen = 0 }
!seen && index($0, k "=") == 1 { print k "=" v; seen = 1; next }
{ print }
END { if (!seen) print k "=" v }
' "$ENV_FILE" > "$tmp"
cat "$tmp" > "$ENV_FILE" # rewrite in place: keeps the original mode/owner
rm -f "$tmp"
}
# Pin COPYPARTY_TAG to the newest release so the running version is explicit.
if [[ "$COPYPARTY_TAG" == "latest" || -z "$COPYPARTY_TAG" ]]; then
_v="$(fetch "https://api.github.com/repos/9001/copyparty/releases/latest" \
| grep -oE '"tag_name"[[:space:]]*:[[:space:]]*"[^"]+"' | head -n1 \
| sed -E 's/.*"v?([0-9][^"]*)".*/\1/')"
if [[ -n "${_v:-}" ]]; then COPYPARTY_TAG="$_v"; log "Pinned COPYPARTY_TAG=${_v} (newest release)."
else warn "Could not resolve the newest release; using tag 'latest'."; COPYPARTY_TAG="latest"; fi
fi
# ----------------------------------------------------------------------------
# Generate copyparty.conf (admin password) + self-signed FTPS cert -- once.
# ----------------------------------------------------------------------------
CONF="$STACK_DIR/cfg/copyparty.conf"
if [[ ! -f "$CONF" ]]; then
ADMIN_PW="$(head -c 32 /dev/urandom | base64 | tr -dc 'A-Za-z0-9' | head -c 24)"
log "Generating $CONF (admin account)..."
install -m 0640 "$SCRIPT_DIR/copyparty.conf.example" "$CONF"
sed -i "s|__ADMIN_PW__|${ADMIN_PW}|" "$CONF"
else
log "$CONF exists; keeping it."
fi
# FTP_NAT has exactly one route to the running service: this conf file. Nothing
# in docker-compose.yml interpolates it and the copyparty service is given no
# environment, so .env's copy is a record, not the live setting. The summary at
# the end tells operators to fix passive FTPS by re-running with FTP_NAT=... --
# so that has to apply to an EXISTING conf too, not only a freshly created one.
# Must stay above the chown further down, since `sed -i` rewrites as root.
CONF_CHANGED=0
if [[ -n "$FTP_NAT" ]]; then
# `|| true` is load-bearing: on a still-commented conf the grep matches
# nothing, and under `set -o pipefail` that would abort this assignment.
cur_nat="$(grep -E '^[[:space:]]*ftp-nat:' "$CONF" | head -n1 | sed -E 's/^[[:space:]]*ftp-nat:[[:space:]]*//; s/[[:space:]]*#.*$//; s/[[:space:]]*$//' || true)"
if [[ "$cur_nat" == "$FTP_NAT" ]]; then
: # already applied
elif grep -qE '^[[:space:]]*#?[[:space:]]*ftp-nat:' "$CONF"; then
log "Setting 'ftp-nat: ${FTP_NAT}' in $CONF (was: ${cur_nat:-unset})."
# `#?` so an already-set value is corrected, not just the commented
# template line.
sed -i -E "s|^[[:space:]]*#?[[:space:]]*ftp-nat:.*| ftp-nat: ${FTP_NAT}|" "$CONF"
CONF_CHANGED=1
else
warn "No ftp-nat line in $CONF to set. Add ' ftp-nat: ${FTP_NAT}' under [global] by hand."
fi
fi
# Read back the admin password for the summary (works on first run + re-runs).
# Strip the leading 'admin:', any trailing inline #comment, and surrounding space.
ADMIN_PW="$(grep -E '^[[:space:]]*admin:' "$CONF" | head -n1 \
| sed -E 's/^[[:space:]]*admin:[[:space:]]*//; s/[[:space:]]*#.*$//; s/[[:space:]]*$//')"
FTPS_PEM="$STACK_DIR/cfg/ftps.pem"
if [[ ! -f "$FTPS_PEM" ]]; then
command -v openssl >/dev/null 2>&1 || { log "Installing openssl for the FTPS cert..."; pkg_install openssl; }
command -v openssl >/dev/null 2>&1 || die "openssl is required to mint the FTPS cert; install it and re-run."
log "Generating self-signed FTPS certificate (CN=${COPYPARTY_DOMAIN})..."
_k="$(mktemp)"; _c="$(mktemp)"
if ! openssl req -x509 -newkey rsa:2048 -sha256 -days 3650 -nodes \
-keyout "$_k" -out "$_c" -subj "/CN=${COPYPARTY_DOMAIN}" \
-addext "subjectAltName=DNS:${COPYPARTY_DOMAIN}" 2>/dev/null; then
# older openssl without -addext
openssl req -x509 -newkey rsa:2048 -sha256 -days 3650 -nodes \
-keyout "$_k" -out "$_c" -subj "/CN=${COPYPARTY_DOMAIN}"
fi
cat "$_k" "$_c" > "$FTPS_PEM" # copyparty --cert wants key + chain concatenated
rm -f "$_k" "$_c"
chmod 0640 "$FTPS_PEM"
else
log "$FTPS_PEM exists; leaving it."
fi
# Everything copyparty (running as PUID) must read/write under /cfg.
chown -R "$PUID:$PGID" "$STACK_DIR/cfg"
# ----------------------------------------------------------------------------
# Seed .env (first run only)
# ----------------------------------------------------------------------------
if [[ ! -f "$ENV_FILE" ]]; then
log "Seeding $ENV_FILE..."
install -m 0600 "$SCRIPT_DIR/.env.example" "$ENV_FILE"
set_env COPYPARTY_DOMAIN "$COPYPARTY_DOMAIN"
set_env ACME_EMAIL "$ACME_EMAIL"
set_env COPYPARTY_IMAGE "$COPYPARTY_IMAGE"
set_env COPYPARTY_TAG "$COPYPARTY_TAG"
set_env PUID "$PUID"
set_env PGID "$PGID"
set_env DATA_DIR "$DATA_DIR"
set_env BIND_ADDR "$BIND_ADDR"
set_env SFTP_PORT "$SFTP_PORT"
set_env FTPS_PORT "$FTPS_PORT"
set_env FTP_PASV_RANGE "$FTP_PASV_RANGE"
set_env FTP_NAT "$FTP_NAT"
set_env UPDATE_POLICY "$UPDATE_POLICY"
set_env VC_FEED "$VC_FEED"
set_env CADDY_TAG "$CADDY_TAG"
else
log ".env exists; leaving it alone."
fi
# .env's FTP_NAT is a record rather than the live setting, but let it disagree
# with the conf and the next reader cannot tell which one is deployed.
if [[ -n "$FTP_NAT" ]]; then
cur_env_nat=$(sed -n 's/^FTP_NAT=//p' "$ENV_FILE" | tail -n1)
if [[ "$cur_env_nat" != "$FTP_NAT" ]]; then
set_env FTP_NAT "$FTP_NAT"
log "Recorded FTP_NAT=${FTP_NAT} in $ENV_FILE."
fi
fi
# Validate required values landed.
missing=()
for var in COPYPARTY_DOMAIN ACME_EMAIL; do
grep -E "^${var}=.+$" "$ENV_FILE" >/dev/null || missing+=("$var")
done
(( ${#missing[@]} == 0 )) || die "Missing values in $ENV_FILE: ${missing[*]}"
# ----------------------------------------------------------------------------
# Confirm
# ----------------------------------------------------------------------------
if [[ "$FORCE" != "1" ]]; then
cat <<EOF
About to pull images and start the copyparty stack from $STACK_DIR.
Caddy will request a Let's Encrypt cert for ${COPYPARTY_DOMAIN}. DNS for that
name must already point at this host, and ports 80/443 must be reachable from
the internet. SFTP (${SFTP_PORT}) and FTPS (${FTPS_PORT}) will also be opened.
Continue? [y/N]
EOF
read -r ans
[[ "${ans,,}" == "y" || "${ans,,}" == "yes" ]] || { warn "Aborted."; exit 0; }
fi
cd "$STACK_DIR"
log "Pulling images..."
docker compose pull
log "Starting stack..."
docker compose up -d --remove-orphans
# compose does not recreate a service whose image and compose config are
# unchanged, so a conf edit on a re-run would sit on disk while the running
# process keeps the old ftp-nat. Restart it explicitly.
if [[ "$CONF_CHANGED" == "1" ]]; then
log "copyparty.conf changed; restarting copyparty to apply it..."
docker compose restart copyparty
fi
# ----------------------------------------------------------------------------
# Wait for health
# ----------------------------------------------------------------------------
log "Waiting for services to become healthy (up to 120s)..."
deadline=$(( $(date +%s) + 120 ))
while (( $(date +%s) < deadline )); do
status=$(docker compose ps --format '{{.Service}} {{.Health}}' 2>/dev/null || true)
unhealthy=$(echo "$status" | awk '$2 != "healthy" && $2 != "" {print $1}')
if [[ -z "$unhealthy" && -n "$status" ]]; then
log "All services healthy."
break
fi
sleep 5
done
# ----------------------------------------------------------------------------
# Install + schedule the updater
# ----------------------------------------------------------------------------
if [[ "$COPYPARTY_AUTOUPDATE" != "0" ]]; then
log "Scheduling the copyparty updater (policy=${UPDATE_POLICY})..."
STACK_DIR="$STACK_DIR" UPDATE_POLICY="$UPDATE_POLICY" VC_FEED="$VC_FEED" \
bash "$STACK_DIR/update.sh" install || warn "Updater scheduling failed (non-fatal)."
else
log "COPYPARTY_AUTOUPDATE=0 -- update.sh installed but not scheduled."
fi
echo
log "Stack status:"
docker compose ps
echo
cat <<EOF
================================================================
DEPLOYED -- copyparty file server
Web UI: https://${COPYPARTY_DOMAIN}
SFTP: sftp -P ${SFTP_PORT} admin@${COPYPARTY_DOMAIN}
FTPS: explicit TLS on port ${FTPS_PORT}
lftp -u admin -e 'set ftp:ssl-force true; set ssl:verify-certificate no' \\
ftp://${COPYPARTY_DOMAIN}:${FTPS_PORT}
Stack dir: ${STACK_DIR}
Login: user 'admin' / password '${ADMIN_PW}'
(stored in ${STACK_DIR}/cfg/copyparty.conf)
Files: ${DATA_DIR} (mounted at /w; shared at the site root)
Notes:
* FTPS uses a self-signed cert (${STACK_DIR}/cfg/ftps.pem). Clients must accept
it, or swap in Caddy's Let's Encrypt cert -- see deployments/copyparty/README.md.
* Passive FTPS through NAT needs FTP_NAT set to this host's public IP (re-run
with FTP_NAT=..., or set ftp-nat in cfg/copyparty.conf).
* Real client IP comes from Caddy via X-Forwarded-For (trusted with xff-src).
Add users / open shares:
edit ${STACK_DIR}/cfg/copyparty.conf, then:
docker compose exec copyparty kill -s USR1 1 # hot-reload, no restart
Updates (security-notices aware):
${STACK_DIR}/update.sh check # current vs latest + any advisory match
${STACK_DIR}/update.sh update # update now (TARGET_VERSION=x.y.z to pin)
Scheduled daily; change UPDATE_POLICY (latest|security|off) in
${STACK_DIR}/.env or /etc/copyparty-update.conf, or: update.sh uninstall
Manage:
docker compose logs -f
docker compose pull && docker compose up -d # manual update
docker compose down # stop, keep data
docker compose down -v # stop, keep host data dir
Or just re-run this script -- it's idempotent.
================================================================
EOF
# IMPORTANT: do not put any code below this exit. Everything after the
# __ARCHIVE_BELOW__ marker is the embedded tar.gz payload (base64), appended by
# build.sh.
exit 0
__ARCHIVE_BELOW__
H4sIAAAAAAAAA+xb+3PbRpLOz/wrJpQck7bAh2jZMXRMrSLRtmplSSXRSXyKlwKBIYkVCGAxgCja
0f3t93XP4EFKzmvt7NWVoSqRAGZ6enr68XXP0IvcK5lYbjSPIyVby3nw1Se/OriePnnyVedpp/ts
p8ufnU6Xn9Or7s6Tr7o72zudp73esx7adXudTu8r0fn0rNy9MpU6iRBf7Twb/RgF1/KvGPP/0LUh
3Chexk6SLgUk4V4JyxKu43lL0RgenTfF40qDRhwlqTMOpJj4+Kdkci2TZqu2UdsQwyiOgmi6tPFd
iO+TaIHX7R/l+GDvB9D8ThO1nzzp6bucqN17vt1Dj8ar4fB0S/hhKpPQCZpM5vzF8FS0Bf6f426t
4wf03N7qPX/e2YodpfxreUt0PD+Rbroljk+G4tp39LhN5nGf54UB5n7opFIJzFA4oScSiYkoacVJ
dOPjeTqTYiHHOftvDlvluA8VswVy1JNZC3yVyhAUhJNI6uwnIlqEYrh/ykLyXam2hIro1ZLbxNkY
nWbSg9ATx5/OUtCbJNGchy4l7kZh6vggTcui2feVIFFZURgsmYUwSiE2dHRSETvpzKwH6PhzZyrB
79i5kp6FNiA38afCizDHS3cGUYn24hJrfPlAtN3J9FI0/NANMk8KkshSPGpRF5AHRWrQ5GlEWSJa
dNsuODXtlIh9OBRPZLFwsjSaO6nvOgE4Bf9hhEGj+RxM25cgGGGIxMdYoZSe9Frip4OXo/2T4xeH
L0evTl4P+jSEnpqZC48+JVFj+Txxfv5KzCKViiu5VCAYQAdEFnoQF3cl6cRYFiyPcNwkUgorTfae
KtEgomM/9Kx5lIUpsZ5Ix7MWiZ9K6HQtdObSLpeiVruOgmwulV0TWqcsz0md8k7L1q7V8hWnVxuY
9ye7mJ5WAkiTdDdXZT8KsYhGiwVp8VJMouQ+PSa9INVhYoZVPBDfdtqwzZbYw6pZvlIZVMQRRzKF
/gxCN1nGqXBlkjLdzQ/7J6dvT/fOhm9HByev9w6Pb1ufYbbaZdTI9Hn5bfMEw+8dHLwdDfde2ta2
5QQxTOSW2xUGMzLrRx34jVl5GwoSSKUslUZxLD1+R35N6YHgZUT92479bade3kMy5LnuPmlnXlzn
Zxtsle0eN6moim7favO6kdu02zJ128xX5WESFW0rytWm/2svjJ619Se/lOG1n0ThXIZpPuT6+tii
fs+i5RPa2389GA3w6IjblbemhSdjGXpqFIXllErLoPuZdIJ05s6ke5U3gYOFtC/q+68P6luivpjK
lD6tf51Y9DlL09hut7vbz1od/HXt7U73uZlVu/7OEOF4cO0Etuh1VE7Yn8soA+2d/Eki0wR+G43M
A17qEWzfjzxbdNH1sxhj6ahhkGRslbjYYoUQHN7AzlIVsU00tBHD34fwRH7a3GVq7NGJig5hHGba
9wSYIni0xFkW4pkSp28OD+zTl4cHcJBMixhRZNdEj31kxdfNHPhJ5YSSohSozvyYzZd77lWC0wGj
QzYP8T1U5/x8cF5SPDw+fTPESIlcwMWzb1YSDQ+PD0awzzMml0YICcwGPFM5Dzx2RJoAfWEYFszE
gSPyJ0x+HN2QQ55nQepbs2iOqX4O91JAiaqLqZrJ4eu9lwPbKhq2HffWrjZgDxQQlkg/4n9WrOTX
fFAGtWHr47W0gJA7NBYtqrmr3+urNj8UAretDtsSdyTtGZ2enA1ti4DSLQGtbe2rjMNiGPObdEgB
CzrPO0TneWeFDqtoQ97Ege/6KQWm5u8iO4JC/TA62zsmGXe3MUX6//z5b72uF4Ma3He/z938cLA3
3BsdHGLgtkquS7DCbvXWbi8q/hl4waZ/H/NmGysIkNBdvExnMDANbOQNcBz5SbUCWBBPhcpcoAE1
yQIChBVyYUhWnjIs1M4itw8xl06oYah2JmQNGRvpPZ5VM9Jj5+rSf3/OFqvIetNd89lygW9SOTLj
AjI0Gg8L7/twixxVc6vXbLkBMsFG89O64G10/U/nOl+uu1eBQD7jGL+R/28/edap5P895P/dLpp9
yf//gsskFIxb8oxhrSKg88nflxesI8yPZdcgiLwECQmye5OcwA+WA8OvcmGgpUsARabtRZzXk98H
mSTKpjM464ToaQT2sfT6V5NrZHsfNJCeO34gPmxWAHDtFi/vzOtWmA6hGyGBfa9ST0zf+3HNhHme
8EgnYmuljg+1MnKeIc4IN/AB3MXhqW2SOyegTBSfMaNu5hpJbwZWQXoepVKgHbAEwagKtZ+sF1Gy
cBLk0vQNuTyntRyMGHoCnElvKpu7FSEwBiMQSsWSCrHLm8nEUol7KbhuUE3zkahbxLl1eEr0CfOl
UcQKMI0ij4KXyhLZKqghnCInH2VxpeMHwv+tRP4LCpW29LRaBCtvaxUuzlOIYo7QF0SOp9oe4i5/
Ew3E3HE2mSC8hNMmKwgwaLQQQRROLSoEeCJwkqmsEIMqhGpC6NOqpC8EyrI5V7XMMEi0QpFkIc/I
YYoc8FoVWqYWw1VTRQDAzMQaR3iu/PcEdueQ6ngJVDBxgGQZII/9aT5MhZrS0yyU1ah2OeAkyNRs
lEdiYXWLNzwpDvck0Ip2aUV0vJGJ1UJ0ZysvudJRvK281EjWrINeuwpdLInvptYwH9c6l24GUktR
nzs3FkBPv9fd6T2FZ98Vppp0no0PIthWqOoFnZ+sfdgfNN8aLmNpncQESRS9qYeRCv3JpGx7JrFs
iUys0wgYc1mZRF1pdqLEn/qhtZjJ0OJKj3lS0gCfhKSqcwuiaWVikEIMQcCU8a2UfJTMnZR8hYoC
aXrf/r/BMmuWLW+cefypwcBvxP+dzpNn6/X/Z50v8f8vuTZW4+HEn2aJric2hoPXp0d7w4Ep8CO3
CaJlS82ohw7jPufx99SB0X3iJyo1nbbgiOLAceGqQWiEZPD14fHo9MfRSCz8dAYfW9Z0KbQvosQT
DfLpWUgFYzgJdBWTNLZC2KI/ITAwOt4bIvzUqIqZNltlwZsLz5bxPGq9kK0rwwiNd8viFGwJmBBL
eUG8aer2mCpjJHyS8waj95W3LRee2pnmBW2W25mkDsKZwHcj/Po0lS0dMUwxAKiGyyxmEw65pHQr
y3LlB/D4Srw5P+uKLiieUtUXiS+kHMgUZJF1JX3y9XDqfXbrbR1jhJj359G1bCeSyhEYqe9J9ME3
p+94c8xdTPtTmfJuQq12MQ2isRO8M8UW8b2c+VgFHewayQpaaQpTBBsSgmCUsVKABj5cAyQtcUIV
Lk2NY6x03JkGegasMEHa0CgQnF6Ay8AJL0WDwQrI+9fQFUg5VU1GIM5EaqQZjdFgV4ROkgAL+Gle
giLu4NuQbhtZq2xMcKiUMsmHc3dHZ64AK+A8JL2Dvi2jDPH8iiGAAUY28EWYlxVFEQTBPBX/RaNa
LhhnfpCSpl1nASn62A+oLRcYCjmecqps9IlfKYF4BtYTMXx1eC5IuLToM4cg+FVIRQjHu/ZVlOid
IGKY5qBnjSlTgGszzk2iAKYVykAbCqmaExKxLPYgSzJrSKeoGCrSlonUdbdr18oSpP8EMJTdbjux
3yoNXmZtwwScgm7NdbQec4E7eYOlLa4NEcUkCluoWUbLgUlAwIBbLGunEBGMLZ8w5Gu2LO65NmC3
3DqNplpYbMC6dmIWRU8y2RIAdGBmQRYVRFHcKtbvMPTAZ0jimsNIHb0octtTzvpwWbx9JaiCNHag
6491qXWpUgkAR1rdgLpnMdRWOok7azKZVN3hGtJDdz1Y6kxVkX75xAqASqFalAE9NlU2UxouCqm7
Oh/izMojsMnWZXRKwWPaVIHerozL5FSOhaiJFS9WWdNIuvDE8HQzAtyp+CeZn4Ku0IYbzWxCGiGo
IlhdD1MJtGiLSrMN70/g9SYVl+hyyflISGGGlT0CUIbvNwSRJ9qiUu8z0uJKH2+pMtoF/ISrbzgB
vEOZ7sHEvDsnGogupag2u/Q28dyKsVp6ByyANfvTUEuvCHK7Qi2cmOixQGHDRwOd55I3cSB9XWbX
IYnqWz0qo3V7rW7HEC5Ke1rYIMF8uvCdvNb5lFhCtFONgAYX7LguFemVdsJl1NUhFJ7ld8bPIgbD
tMiK2d3bK71rCCYHOjkx5VNbbxnOAHVZxeFRIyQuucelqj2RgSdHakqmBDc5RoBbtGoXbeKYa6qY
gqnDJou559i6U42TdEsYjAkXWUR3qLfjkXI1f3+JnwKhlifv2/LmiZpR9h9NwEYbwtYJXLiMkDhT
wgh36IdNOqFwQa/f8RkD3ZS/ar43WI2hg4904SOMwuU8ypROmPRAXhLFtEnRMLkcB3QySnIhGDRP
VM1gfojGxXB8d2fARTGgB5ODtZErM8pgC8pMLsfRGG6z5yYyvTT73KW6CGeMSL9FCxiKKcwjpfi3
QSSoUsCRTS/ylpCtaUtcJhgSJMkaFW3D8wkC3QTKyryJcRBRBeg/DVO/XJ/pasnw+vNkfeX1G/nf
k97d+u/O06df8r+/4toQ+0By5NhJExhFTggdkY9fS0voZI/SzdZP97Ry3248MtXxOOForFd5Kc4V
EOVXnfv35MpLQvCvCqB4LpsEivUxqwp8N+PoUzwu7Y3X1mu2fQZqubIDwc6J6wGXfClQrNa0Ezn1
qR7HuXCbMI2faHQ/wZy5VtWqlXVinU39bY04ieSQc9J/b7NaS+TSoWqsMqmWDp2UiXJSSRn7qc/A
DVjxxTyWUzTM5uMQ81NtLhNSrZqyefEGyzmduUnLj9rPYYpl/m45Lm/ZZ9ABLpm/9NNXiKQIESkl
mIit5lAAnlZlzFvl/epOOQVoPyRgZXC8iTvd1jYhpdYKtlFke0rU9V56PYdOoVzgFpQSpK0Et9cq
C+ZQXZE+5BkDpOQECzp0kYPR3TLTAblxNo/zCgbydSobajRDTVR1XsO9l33NVL6gB4SLHgtETCBY
Pvrxp1f0FR2j0JCecjiGLx6d5+BTIDn+Eg0+tkFvUuCHZkvs81buKmI1BxLdGZ3p8EiAxaGQljiN
gLj1fEHD4TK056srSnf4SAgZAOf3GTBzLd8479+zb14jsn06kFAj2vqbEU31vOadVGUL7PLGzW/b
fy6a8mBIWqD86gmZ8jwJqf8YCfmV6AtzygDf4JtY3z92pkUzZA588pmWldMs+uCKv3paBeRKthqs
0Y4YkpHBFUoujICZytEVOs1COlWcg+jXirMZfUrOasURiz5lUmQ3OjWwSJzrGc9sGZO2Ul16tzIt
R1m+upv/YPqrRyn6lcyqhND3pST5/hrSkhacI50Z0hJGuq6rQMbZmg09c1JBL02w3OUUyL+bBSEL
1MegMCij3IVPOmdKev1cl95oWxQN5SZ+nAKLFgb8B7IEo0sLzQIFhizg46lcEEDuQmehDAbXZk5b
Beb9qhcqfBB5C6p0Ccor5I1P+2cNs8mjzy0XhYeC0snx0VvtZu5zV/dVdJhQNJmY3Quhh3RnrAX3
EWlwzajNIWrJBZNm7c0pTHkA3To63H+b+zFKMUzViEo8ZPLsSi5zti+hcLzP8lhclvUhJn+pRVVW
fCyqI7kYHB0JDOiBxS+VJqKhD6QgJVl9YcHQaj/sj14MBgf9ShFpJXJybeTP+litAcVp0X5+WPRL
QmOuYnk/4xgE8p/t7HwE/z/p9Toa/2/vPN158nSHz38gJfiC//+Ca+PrdqaS9tgP2wTsxw4BidpG
pSxsld7w/tMTW1TRzUt5XoL4EZI3qdS/OTho12Ll9fG900PgmKIMDU9bqb0R2gQ5Cmu0+26qE7J6
HqVkhH/NQFVQruVpXoE1Mg4+ByeD8wI5OGKOcAWeLSoZ6R+ykAOMJhUOVxxXUQsnV897Rnkl/9Ky
dF38UjtG8qKKt6H0nhQHLS5NKV1Y4ZF4mi5NR8eCiKruWwQ0VB5pSgwbcrK1ZYAdoiqjPmXWofLD
kN8NgokQkcM6e0U4o5r1GL0jFviYfgOE5TM7VcyClvh5Nja/3jDxUs9bX0X8ozoloRUIM6HK3rXK
4+pjKsVhBVJagmkR53apdKbDmmKydPIiv5w4hlLFmOZKJDNVLhPqGozRfNotcdKVQE+0lA7JRofz
PX2XNozyh4i7IBIZRnmXYLh39nIwHP0wODs/PDnu37SWrfeakNFNJlSMo/G6D14vMeSlHjEsm9L5
lmu5whvL1JxnaKzMrrkGR0ru/208UpL644Bkywxv6pXlFdOSylKhKGH38mOwpParUAbXn8IyJK8B
fJSu+fN2J2eo+vcURQJrPBft6e5WfxkhFlgOI9rz4d7+3+9JcPDqPsSEx/egFCIEt3P44m2/SzTP
X430Lf+Gqc9sKTWz9CT0/rMQB2dvR2dvjvsd7r+mZlWpvnw1OhucnvRX83P6eVEqLJlRdhJjjf2g
VjsfHL3o1zcbrifwHwCcSyZ09HoPTJ2fvDnbH1x03t3Wm3XxzTciXnjN9maDtq8+3pBA2BFtIRtF
UtUTdJiJKa2La0RP2hrKD60x6NdLXxE+iDVieDCJpNmV9sq6fKcJflf4VuQsaz9XMYtCcrX7H1/v
23oNucnFhbAS9L+3e128e7fL3oP3JxYz2tQ/fHHe1wqFjqMAbnUXUaXYbgTB+iY/rov+/4h/XFzY
Kkb2Z79792hD/PKLsN4Tu9yi3a68bd/ScCRyctd+mJWn0UZX/bzLgwf9R7f1XTG6Lh5tPOrzE93o
ap3oGmdXhq09678d6/3onfnSsZ6P3j3aJA7A4x0OSEwh8f316Mq2foVTzdb1g5/rFS6vN+i2aBQn
yIknwro2/Dx8oB7S12vz+yHyTP/10TWpTfyaXvLCMO01y8RY3GDFPG1jn/lLY6R2xUrzV9oy7X43
f7Bmrva99po3NkZr9zv5k1XDtfv5c2O19prZ5q+PTl5iJNhMO4imd5UYD9Gyho9GU4gPuVgf/tzp
9S66u73t+cXjd3zTmYsH6ueQhPwIq3Jbo71/dLrbpze/+PqePuK7b7apn+fL+8fqzi9uPtqP9/S7
1B+eW/fXllffNLKCTvZFvVvam5DuLBJ1vLYF8ywDOutCxxnrm3/D/cQnenTSaKXuvIWIGcUaOxUY
kKoMooEwUPlh3Zb+FUdeTDRYh+vVhIToUJA+ugP/4rksq4Zgl1loHbvHtbI3MSeazFqkJs6cOprT
g5Tu+l6/XuezIfjUvyPKXRBrVKSsPC6ueB5u6ZHXbt1puP1d25PX7ZCKsGBIC27zA/0QiF1z3l+P
+ocojI4O/z6okIHl0YdLnepi0/fEJlHFd7/kE6uu02ZRf9TUxMz97m61kSfHPlBV/dEvuMnGWZhm
ZQ/zcrVHMpOBaT+RXpQ45sZFYIhU2ZnbVbs2izn5nm1lIUMVODHTRirHrd2aQsKnunJtAi/0q+JP
THxw/MPoxeHRoF/RxzaNV6/h32gqU1ZZUq0J5p03z/32FLhbWANR/8dmt19fafALH+L93/autqmN
Iwl/168Yr8Wxa6Pl5V7qDlhcjiGJyz5wAXbiYEIJtIBiSatoBQ4H+u/XT3fP7OxKMiSxK192ymVJ
uzM9Mz09Pf06ELdfpa9npBe1OgkBWWvhXBiPrlIQN/rItQ/z2Gy+2nm/ZTZpd16lW9b0ZuihU0sg
JUuEOii+fZrjmAfuJfecW5qLLC0C9bF7aVQv9/d2/7uze7ikYiBbU4dZr62pgLC15poUZxWp2Lx0
tTgkC6Fbo9HVkOQSkULQmypFqPA34g5wjfw2hL4iZkiFWEQ4sC6iweOnSKuzsigsuDiyPIDMeu4I
YZ2UA7tZ81KIJEWcXTrfk+iVLB6ZvZfbLyB9j1JSK9pecAYCy9/uvza/XiE6EUHMg4vYYysf0xsi
BGKfNDP6shaYcX/osZdH01RQYS6Wny/kyULeAPMmmCAMgkhcfMtv7RrROK9GA7Pi8wbqGEym/3Gc
9ofKNk4Odk6IGBIFip/vnr9OFHj700ez6GB+s/Pdy12iKpjmdfmPAgUQ0Kivq88JEp7nKR0aiVnR
eHSUR/yMqJ1jpMLmyhJBDZIgwmGzag8xeWauHQg6qAaIPSoA2ZrFk53dbTnEQu4kKoPSiosllAOF
hJxAOehYf5oSaoVCRikHj2DzcBTPOi1wOhR6k0B1WvN+1kmXOSNXYmf6ssbcBfMykMkQGQm0nVlb
EtId3IxFo+YQU4D8MX4f/4QdkjYG2ahPu6wkGohwtgrOIGwj2zGLRyQvHj/9EMtnaL9EzxY9BgJO
8dg0V81mYppr5lmDQJ/0UsuZGCjOfRCsyJ38NCx6/jCwssRqIPXuTA6DQeud10+kYFR+mHA/F8Kd
HhnptAChcsOX5MaPzTtVRTtdkj8QQ/zFe3ghJpLeTaui/q7TSiHjQWI44ZDxjG3nsCiwlUYFoiE7
WQkePJnKneOG2l9OFGRFarm2wgqk+bBzJjHHrUPPrKZZpabVL6W5K7zTkogxn47s52LpFJJjJyrJ
S9Chrj1xkUem1AsS0kPQlFyzEckxG5ZRlcn7WjfNbslY4pzRhTqrHm6twJZJExIuOYNbfdbX6rQm
0VFUjim8qmkMKhChvmc87Kxt/Y3nfMsBuibg9xzBPbjYgPkN+rOVWtXioKOJaXaWJ28oE+LmrfP8
4LVcomBDci+648urU4QeLMMGly87bWSybCe3bP3r/up9cEzQX8fAoiDw9d11/0dw9HNw/DQora0P
DOgmkWQxX46fPATc9bOQIUZB/GT5w+ryoqyg81YhvHbdtPOP83KxQxfJVrGQDjJx7/7y6zIB1Npy
q0U6OLuJXKA1sTY6x3nDi1MM0oULsqbd1VQlMzZvQG5qCt17xR87+/tms59fbPGvd29f79LP9m8t
tZDZzbP1YbDJ9NLKU5oFCz5FH518q2F/nWDK1b2bjkQ0gGktCeZGZTt92F7iUezyVmpevHmH4zod
BfjFtq2gCYjBLB5gNjcX37xfLG/6O5XBMWtADhqaeP5LjmCPLEfeB6lvN/SFaLbXPbX5fo2z4bWh
gznLY7UExZBvAxoU8teDqMHW8ukKGKfW6KTn5nocwqSH2SGHLTFHkrTOvmys1iiN82GvOw5Hi0fx
h9bxIg0n5mM0jOKefAmu3wVRtO7otm9gDIpZOKR2HzpPqdUwcu+pp1gEXyK3cdiPL0i8HYYrUQRW
1he9diVqePLU+GpIhyQ1lGGf9Ydhe8mcaqdj+j4+pV5pOu1oCR+n0hzCC9FP2EsHxJHoFX85jSJt
Z54mJlxZiswTEw5olbQeaasEb9Y721THRZVJWqGn9B7fN/Gdh5iTnJ2fd9M8vMahMxpc6GAlTGDE
Rj56qugNlgIPg/w+4Q+L7ELnPZf8DmS+lAL6pwxPlXUIN5O7reQuoX93m3db0bMP+RMwDBwv+E+M
XljkJxEWjDqY6rVf7lKx8G2bVqxY3iGQrou6GkFUJxHQvT6jt1g/RoutthZ55AE5Nwwzpr+QmhK5
JhBROToHYs0KA0UNyDibgb7ZlBeVLAdXLXH1qhC27Jut6gvXZgttptCRfbwHH/rgkA7rhrahnbte
nLYhb/5BZjlbQKRH+z2GfYh2RGM80stVRu1PhJYyF4jpZ0b7KMRmX7J3SiSrK1EMY6wSDYdBJcxT
Yo61DglW1Eh/O0uHY7PDHzjP27lJp0YWmKeQmsM0Olr/98pxZXRd3FjGTprBWRqioyUS987GSsra
Mz6E+xRmxYDxXLzBN3l2hK/HjcvuOF9ynhEwpiUwJ1YLO9ccp6O3h8mW4kchvRJ4fp6Q6+7o2Ntj
bLkHoWkLOkQ/ti9SqXk7ieQpH7hCwwExvOwTNIESGQAM3yKHL4+IaBz3D+7Zn7T7wbPKA+6lVjI6
4Xgl27nfJxqiy4LHEFH5PMYWINGyWoeai8u8fdLtCGD39Ow6dQ/tOt0EUQnc8LoYr66MHaxiWF9y
hOFJpcrUTHQ2HgVhFh79TPU9VPgSOkgzH+lZNlUT59bVKXG9n49+BoN7QhxtEYcX0fLwOoqqg7Cb
sgRIxm8RSM3sHgZeS1tl7xUNAueWPD2FBJrYubrDVAGuV9kIV6dKYIvDJf4ZgReVh6RAh36/LCFh
j3Ib/81SEP+SdQchVMO0E8IyhVFHVBpv3n8N096uOpNH6VWe2uy57kDcAV2OXCPtwmbJLrHxVa1T
+TJbl4v4tC88uJzW70QCJ5x9btwd99Its0n4yuDapa+4emCLkchKt3g3PDs8JLaSgcd6yiq+kJl1
H5v8Mu31xOdPOjG2enLwYnXlP/I+ngHHdcJ+pt1Dev52/7V1NlV7eIAGVaov0jDjQeRhoEJsZcAE
fft7wFI/18YXaJQuorx1XpVm3atZUq6amFVnZSdsi7UwKMj/NK3vTfBj6xCDwRViPCqapjx+o2uE
NxjkZBZmDvde7ex6jjjtgt0PAPP8imTxUfd/HHu+br5J2yPoK17jmWDlVr/5YH9scQT8uvHrKyTF
cOKdCRsV8Ps738H/VcAHkm/Hk6XS+0kJa17fh+2LnPE18W0B851IjqN4ziQmFt6aYRPYjQggCEAg
pr5sx1Wle/h9blFtYo6at3zxyDGsSZZKg1n0B6uFbxp19PgV+NH3fAeZ+dTujk3IMUj5ZXuYQs5x
kVJfmtGgsxO5/KyidHZIIFPHuH5LmmFomiH7BZ4u5LgpeHVtxejZJI71So1N19hAR/Fc7NIJHaPj
q5y2mgyhSPiVF2qqGiIcVK/lWLy9jQ/kEtPJxNAPwdpkMqWqegYnFNcHgKorSXqBQZIN2M01loe0
HrsG9VFgbsU63FydLEYldzybsRxwbsS7xcKWTVLiYzzBXpoOzT+ds9ynrtWvQV3PObDKhUCt20sa
YURc4vSPJRdxtgSTA52NGxwjJrZHGNCIbfA1P192bBzydSLDcqcdjHV0wo2zLXNkpbtWt5MfeySK
OnIQjPUYoJogmtu/gz9pzQsTcGQ5zK2FoaN5i9brrWcT09oCP8omcRz/cZ6UVw2VCcPcmMqpAnW6
YUy9vhoSo/q8d+aMqrRLDgeasP0xiU/bH2Pdgq0r2oXvF/oLnYXvF/67cBAFs/aIME1xBpanAFdE
5lDyCPak6vBL6BBjp2QSEamknQ2+8iKDq2sKOUD/JC6mOm8EqFfU8gQjT3Gx3spv6STb2SaEXHYv
LgMdSnYuy+vGhLtZOWt3oFQwmUL4qo9wmjYvCxG9RPW1stHwsj3ILWY8JlrBCFOfT3QYRzfX6yZv
vOkLRZ8mgUylM4tAvbMYhG55yyko/nTCyiTvE5yvqDCJHoi4DuFKo75ATaefJ0C1aVdm1el2VDcg
Wk7tDIWDANvWgVFe+ftXfS72S/T7++hCRuyY9kYpFLYYY/DVmfK2hAYLGx7B2eGcFfZqPo2qlfyM
2Byk49zIpS25BhMiXlijk15uy/X2By+e70Iw3zGhvRVfbrmCNNamIzrCfSgs35tP2VWvY3AFKDrs
R1+avcvQkoCotxgl/3SjhG9K3SEnggd3Duj8nSHdY//0irk/P3pgLwV7521UDuqaweEtWN8xVWkU
lfvQdXRX3dvIsYnn2Cmxcw7EaZbC3EqxONn5eVn79zoTqkgQaRsWQdWR15UfQ2ODgsvgVBJDULhe
H5Hm41INvAMKSj4KGjNB84OTZFIj0brKrjk0CEre3ooZ8nON2Kk0MItrw8VKO8Udd1rCmS0wNzyJ
ph6XV7bcqwJTkdCKfWuTat9lOEJtTUxkdjVv0TConf3n37ym3dm8RRNweRKPrfEQTielGryZDfCx
vTqbHSdymRGRgIXhMnDZe3wutzIZF/I+E6SVZqVre7zM2gJh2ScKp6xPaLbsvYrKUx9khduLzfw0
6OYt0dFkZvsnUQV1rrFmYAza13Soc/wNoZKXbjI9FMSG+c9k8HcVwnjgPO/Zitbd+pXD0rzUjC8M
G4EIPdxJ6fRBIovmztuX26aV/koMRQw5nS7tvP0rtsIjZTkuTBHnfoDl8vTtPIEPYjfT+E7S7YpW
iOUaF0H7kUL3DEYqN89Q2RmsvoZAcp5dDToxxwZ0shMmHDc1N9nysWJ98uBA0jEfNWElxEKJQSo/
jHBYINyvBn/wFiiCGon5lIMCUEOelCoJyeFl6fSwtpf7ubaqUcz1SbfaeDj73kD938OypSsx2PW7
VImWxgXQzuXjwr+1bsE3MWXqn0Rey+508z9m+/IE45OuWODbUJG1CTgl7iAMSuETc6pesXZ+nuKO
NF2fZzjDeQWnIQmrUkgPY1Ozh8UcA188MQVL5gwLlsMId5YwKe+x1J6lhzznNKKRy9pBjLQdl9Kg
6houyN62Z3UYGBVdmBdQ+hc0W+WhcFYxHpq3+JxsVDp2HRT9l2KT5wjzjMzADcRCtztb4tJn7es5
yrxE6NiULWgYfYO01XJKzeu974LZivOG6v8JQawG9htfBf/2cOHwJ0QsJhqGCZg2uMcxnfsZTVlE
1pUu2Io2Shy10kZs3rojqmRwrZzzs8iFOHORhMY2mvjzemGVYIVAHhUkavubT6KHMjMnM4NGwcOF
fsYk3iqdIodQYiCRYJndMzLfwmRH5w+nkOEsMRXWKPWFcAbhus2NKgQp/cMelbyuZU2nz5GL6a48
+NyRc//qTy1cobYQmqc0CJA56QVefN2sRLeNwrA9j7bKulLp8H7BqiMWqJPKX6ZKcW1GWZMNfRqM
4ikd7CG0gpfzJj+bep9rdlpmKcbokVwQV/RVieix+YEDgOcnr9ks3uKSCgSidwdELV2JTMctGXJp
N+oXkYftMYccz8l329zEZ3HTX3HRZNFT1UQyQv4ofAxI0Mu9C2AKF2cs2Zgui++6PeqCdeckwOEC
Kafxyl/+afM9hHKLepF96Wd8EUMqk2Q124sq2CzMwI9t0xxMl+RFj3jGjJvLftYx//rHP+bnnjka
+WGEm/XnVRNGoMfDHxKGi6XzpZyQk4qikpwjyTUVJ//0wcS0NNRbY5c5/biiD4MwZtSaIkAikp09
opFHfBFAftngQEFcBgDhfec1kRGtXgN1Sh0wdp/+9qA+yi0L0X10ZgdRld5pl996r9sdIlQiN5eS
Os9Dt4FG9k/bSQu+0nd+/UlpcEwNB25zSF43jGYPQmWoql+VeonXlSXEeWsld8l29HOaWdiZyZod
vR10x8eN7VRCEYjFJp6M5GJgdV83niNHIBF9yEJqNI7Uh3bcwF38STZI88ts3NghKjgA5pKpSyIM
k8Vsqnj4VMAPRvMnss2Yn7LfsuRHgz5E6+PG3uBFu0cyYpvmxXvgjfzJRwg/bBHeJ0rL+t3/pZ3t
tNe+OUjPktVLav9SttRx44c2btf65ibh8eSxHFlTE5NZnI17RBJpPxsg2y5rd+bU0VuJWy3cLjBn
5p9zMd9LjYpcI6Dm9PAAarRGCcg5gz/D4CRh5cG8ANdIFWygQFwFKZXj3FsECUL5cxj2xvwwYv09
W3TWoH3KeUiIgZgI2NXRue/UFrOGHiy3q+st3ikTd7bofeMo1vZhiYA2cqSDFeXJvhDQkb7QHajv
lFgieWdPKNtw4F77lGVfk4bMZPQ259vCi/tmjnhcdzSEO3l4py3vHIxjpl2m27/6/py61KUudalL
XepSl7rUpS51qUtd6lKXutSlLnWpS13qUpe61KUudalLXepSl7rUpS51+avK/wFMvaeYAKAAAA==
+78
View File
@@ -0,0 +1,78 @@
# copyparty stack -- caddy (TLS) + copyparty (portable file server).
#
# Topology:
# Browser/WebDAV --> caddy:443 --> copyparty:3923 (HTTP, internal)
# SFTP / FTPS --> copyparty:{3922,3990,passive} (direct, NOT via caddy)
#
# Caddy terminates TLS and reverse-proxies the web/WebDAV UI. copyparty's SFTP
# and FTPS listeners are their own TCP services, so they are published straight
# from the copyparty container -- Caddy is HTTP-only and not in that path.
#
# The image's baked-in config does `chdir /w` + `% /cfg` (include every *.conf in
# /cfg), so our ./cfg/copyparty.conf is picked up automatically -- no `command:`
# override needed. XDG_CONFIG_HOME=/cfg in the image, so generated SSH host keys
# live under /cfg and persist across restarts (the bind-mount is read-write).
name: copyparty
volumes:
caddy-data:
caddy-config:
services:
# ---------------------------------------------------------------------------
# Caddy -- TLS termination + reverse proxy for the web/WebDAV UI. The only
# service on 80/443. Auto-issues a Let's Encrypt cert for ${COPYPARTY_DOMAIN}.
# ---------------------------------------------------------------------------
caddy:
image: caddy:${CADDY_TAG:-2-alpine}
container_name: caddy
restart: unless-stopped
ports:
- "80:80"
- "443:443"
- "443:443/udp" # HTTP/3
volumes:
- ./Caddyfile:/etc/caddy/Caddyfile:ro
- caddy-data:/data
- caddy-config:/config
environment:
COPYPARTY_DOMAIN: "${COPYPARTY_DOMAIN}"
ACME_EMAIL: "${ACME_EMAIL}"
depends_on:
- copyparty
healthcheck:
test: ["CMD", "wget", "-qO-", "http://127.0.0.1:2019/config/"]
interval: 30s
timeout: 5s
retries: 3
start_period: 10s
# ---------------------------------------------------------------------------
# copyparty -- the file server. HTTP 3923 stays internal (Caddy fronts it);
# only the direct SFTP/FTPS listeners are published. Runs as PUID:PGID so
# files on the host bind-mount have sane ownership.
#
# A published Docker port BYPASSES the host INPUT firewall, so set BIND_ADDR
# to pin these listeners to a trusted interface if the box is multi-homed.
# ---------------------------------------------------------------------------
copyparty:
image: ${COPYPARTY_IMAGE:-copyparty/ac}:${COPYPARTY_TAG:-latest}
container_name: copyparty
restart: unless-stopped
user: "${PUID:-1000}:${PGID:-1000}"
ports:
- "${BIND_ADDR:-0.0.0.0}:${SFTP_PORT:-3922}:3922" # SFTP
- "${BIND_ADDR:-0.0.0.0}:${FTPS_PORT:-3990}:3990" # FTPS (explicit TLS)
- "${BIND_ADDR:-0.0.0.0}:${FTP_PASV_RANGE:-12000-12099}:${FTP_PASV_RANGE:-12000-12099}" # FTPS passive
volumes:
- ${DATA_DIR:-/srv/copyparty/data}:/w
- ./cfg:/cfg
healthcheck:
# copyparty's own python -- no extra deps in the image. A successful TCP
# connect to the HTTP listener means the server is up.
test: ["CMD", "python3", "-c", "import socket; socket.create_connection(('127.0.0.1',3923),3).close()"]
interval: 30s
timeout: 5s
retries: 3
start_period: 20s
+451
View File
@@ -0,0 +1,451 @@
#!/usr/bin/env bash
#
# update.sh -- update the copyparty container, optionally driven by copyparty's
# security-notices API. Companion to deploy.sh; installed alongside the stack.
#
# copyparty does NOT self-update, but it DOES publish a machine-readable feed of
# security advisories (the same one its built-in `--vc-url` check uses). This
# script reads that feed to decide when to act, pins the new version in .env,
# and recreates the container -- so the running version is always explicit and
# a bad release can be rolled back by editing .env.
#
# Subcommands:
# check (default) report current vs latest + any matching advisory; no changes
# run apply per UPDATE_POLICY, then notify (this is what the schedule runs)
# update force an update now (to latest, or TARGET_VERSION=x.y.z)
# install schedule the daily `run`
# uninstall remove the schedule
#
# Policy (UPDATE_POLICY):
# latest update to the newest release whenever one exists (default)
# security update ONLY when the running version has a known advisory, to the
# patched release named in the feed
# off never change the running version (check/notify only)
#
# Env (also read from /etc/copyparty-update.conf; environment wins):
# STACK_DIR=/srv/copyparty UPDATE_POLICY=latest VC_FEED=advisories
# NOTIFY=1 SSH_NOTIFY_CONF=/etc/ssh-notify.conf DRY_RUN=0
# TARGET_VERSION= GH_REPO=9001/copyparty
set -euo pipefail
SELF="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)/$(basename "${BASH_SOURCE[0]}")"
# Load defaults from the conf for any var not already set in the environment
# (precedence: environment > conf > built-in).
: "${COPYPARTY_UPDATE_CONF:=/etc/copyparty-update.conf}"
if [[ -r "$COPYPARTY_UPDATE_CONF" ]]; then
while IFS= read -r _line; do
[[ "$_line" =~ ^[[:space:]]*# || -z "${_line//[[:space:]]/}" ]] && continue
_k="${_line%%=*}"; _v="${_line#*=}"; _k="${_k//[[:space:]]/}"
[[ "$_k" =~ ^[A-Za-z_][A-Za-z0-9_]*$ ]] || continue
[[ -n "${!_k:-}" ]] && continue
_v="${_v%\"}"; _v="${_v#\"}"
printf -v "$_k" '%s' "$_v"
done < "$COPYPARTY_UPDATE_CONF"
fi
: "${STACK_DIR:=/srv/copyparty}"
: "${UPDATE_POLICY:=latest}"
: "${VC_FEED:=advisories}"
: "${NOTIFY:=1}"
: "${SSH_NOTIFY_CONF:=/etc/ssh-notify.conf}"
: "${DRY_RUN:=0}"
: "${TARGET_VERSION:=}"
: "${GH_REPO:=9001/copyparty}"
: "${LOG:=/var/log/copyparty-update.log}"
log() { printf '\033[1;32m[+]\033[0m %s\n' "$*"; }
warn() { printf '\033[1;33m[!]\033[0m %s\n' "$*" >&2; }
die() { printf '\033[1;31m[x]\033[0m %s\n' "$*" >&2; exit 1; }
run() { if [[ "$DRY_RUN" == "1" ]]; then echo "DRY: $*"; else eval "$@"; fi; }
# docker compose, scoped to the stack dir (so ./Caddyfile, ./cfg resolve and
# .env is auto-loaded).
dc() { ( cd "$STACK_DIR" && docker compose "$@" ); }
osfam() {
local id="" like=""
if [[ -r /etc/os-release ]]; then
id="$(. /etc/os-release 2>/dev/null && echo "${ID:-}")"
like="$(. /etc/os-release 2>/dev/null && echo "${ID_LIKE:-}")"
fi
case " $id $like " in
*" alpine "*) echo alpine ;;
*" debian "*|*" ubuntu "*) echo debian ;;
*" rhel "*|*" fedora "*|*" centos "*) echo rhel ;;
*) echo "${id:-unknown}" ;;
esac
}
# ---------------------------------------------------------------------------
# .env helpers
# ---------------------------------------------------------------------------
ENV_FILE="$STACK_DIR/.env"
env_get() { [[ -f "$ENV_FILE" ]] && grep -E "^$1=" "$ENV_FILE" | head -n1 | cut -d= -f2- || true; }
env_set() { # <KEY> <value>: update KEY in .env, or append if absent
# The value goes through the ENVIRONMENT, never interpolated into a sed
# script. Interpolating it corrupts any value containing & (sed expands it to
# the whole match) and aborts the run on one containing the s||| delimiter --
# which is reachable for an OIDC secret, a password or a URL query string.
local key="$1" val="$2" tmp
if [[ ! -f "$ENV_FILE" ]]; then
printf '%s=%s
' "$key" "$val" >> "$ENV_FILE"
return 0
fi
tmp="$(mktemp)"
_SE_KEY="$key" _SE_VAL="$val" awk '
BEGIN { k = ENVIRON["_SE_KEY"]; v = ENVIRON["_SE_VAL"]; seen = 0 }
!seen && index($0, k "=") == 1 { print k "=" v; seen = 1; next }
{ print }
END { if (!seen) print k "=" v }
' "$ENV_FILE" > "$tmp"
cat "$tmp" > "$ENV_FILE" # rewrite in place: keeps the original mode/owner
rm -f "$tmp"
}
# strip a leading v and anything after the X.Y.Z core
normver() { printf '%s' "$1" | grep -oE '[0-9]+\.[0-9]+(\.[0-9]+)?' | head -n1; }
# $1 <= $2 ?
ver_le() { [[ "$1" == "$2" ]] || [[ "$(printf '%s\n%s\n' "$1" "$2" | sort -V | head -n1)" == "$1" ]]; }
ver_gt() { ! ver_le "$1" "$2"; }
# ---------------------------------------------------------------------------
# Version discovery
# ---------------------------------------------------------------------------
# Currently-running version: prefer the live container, fall back to the pinned
# tag in .env.
current_version() {
local v=""
v="$(dc exec -T copyparty python3 -m copyparty --versionb 2>/dev/null | grep -oE '[0-9]+\.[0-9]+\.[0-9]+' | head -n1 || true)"
if [[ -z "$v" ]]; then v="$(normver "$(env_get COPYPARTY_TAG)")"; fi
printf '%s' "$v"
}
# Newest released version from the GitHub releases API (tag_name, e.g. v1.20.11).
latest_version() {
command -v curl >/dev/null 2>&1 || { warn "curl missing; cannot resolve latest release."; return 0; }
curl -fsSL "https://api.github.com/repos/${GH_REPO}/releases/latest" 2>/dev/null \
| grep -oE '"tag_name"[[:space:]]*:[[:space:]]*"[^"]+"' | head -n1 \
| sed -E 's/.*"tag_name"[[:space:]]*:[[:space:]]*"v?([^"]+)".*/\1/'
}
# Advisory scan: ask copyparty's own python (inside the container -- no host jq/
# python dependency) whether $1 is covered by any advisory in $VC_FEED. Prints:
# OK
# ERR <msg>
# VULN <max-patched-version>\n<comma-separated advisory ids>
advisory_scan() {
local ver="$1" feed="https://api.copyparty.eu/${VC_FEED}"
dc exec -T -e CPV="$ver" -e FEED="$feed" copyparty python3 - <<'PY' 2>/dev/null || echo "ERR exec"
import json, os, re, sys, urllib.request
cpv = os.environ.get("CPV", "")
feed = os.environ.get("FEED", "")
def vt(s):
out = []
for p in re.split(r'[.\-]', s.strip().lstrip("vV")):
m = re.match(r'\d+', p)
out.append(int(m.group(0)) if m else 0)
return tuple(out)
def cmp(a, b):
ta, tb = vt(a), vt(b)
n = max(len(ta), len(tb))
ta += (0,) * (n - len(ta)); tb += (0,) * (n - len(tb))
return (ta > tb) - (ta < tb)
def satisfies(ver, rng):
for part in rng.split(","):
part = part.strip()
if not part:
continue
m = re.match(r'(<=|>=|==|=|<|>)?\s*v?([0-9][0-9A-Za-z.\-]*)', part)
if not m:
return False
op = m.group(1) or "="
c = cmp(ver, m.group(2))
ok = ((op in ("=", "==") and c == 0) or (op == "<" and c < 0) or
(op == "<=" and c <= 0) or (op == ">" and c > 0) or (op == ">=" and c >= 0))
if not ok:
return False
return True
if not cpv:
print("ERR no-version"); sys.exit(0)
try:
raw = urllib.request.urlopen(feed, timeout=10).read()
data = json.loads(raw)
except Exception as e:
print("ERR " + str(e)[:80]); sys.exit(0)
if isinstance(data, dict):
data = data.get("advisories") or data.get("data") or [data]
hits, patched = [], []
for adv in data:
for v in (adv.get("vulnerabilities") or []):
name = ((v.get("package") or {}).get("name") or "").lower()
if name and name != "copyparty":
continue
rng = v.get("vulnerable_version_range") or ""
if rng and satisfies(cpv, rng):
hits.append(adv.get("ghsa_id") or adv.get("cve_id") or "advisory")
pv = v.get("patched_versions") or v.get("first_patched_version") or ""
if isinstance(pv, dict):
pv = pv.get("identifier", "")
pv = re.sub(r'^[^0-9]*', '', str(pv))
if pv:
patched.append(pv)
if not hits:
print("OK")
else:
best = ""
for p in patched:
if not best or cmp(p, best) > 0:
best = p
print("VULN " + best)
print(",".join(sorted(set(hits))))
PY
}
# ---------------------------------------------------------------------------
# Notify (reuse the login-notifier's ntfy config, like scripts/auto-update.sh)
# ---------------------------------------------------------------------------
send_notice() { # <title> <priority> <body>
[[ "$NOTIFY" == "1" ]] || return 0
[[ -r "$SSH_NOTIFY_CONF" ]] || return 0
# shellcheck disable=SC1090
. "$SSH_NOTIFY_CONF"
[[ -n "${NTFY_URL:-}" ]] || return 0
command -v curl >/dev/null 2>&1 || return 0
local title="$1" prio="$2" body="$3" host
host="$(hostname -f 2>/dev/null || hostname 2>/dev/null || echo unknown)"
set -- -fsS -m 5 -H "X-Title: ${title}" -H "X-Priority: ${prio}"
[[ -n "${NTFY_TOKEN:-}" ]] && set -- "$@" -H "Authorization: Bearer ${NTFY_TOKEN}"
[[ -n "${NTFY_EMAIL:-}" ]] && set -- "$@" -H "X-Email: ${NTFY_EMAIL}"
local t="copyparty"; [[ -n "${NTFY_REGION:-}" ]] && t="${t},${NTFY_REGION}"
set -- "$@" -H "X-Tags: ${t}"
if [[ "$DRY_RUN" == "1" ]]; then
echo "DRY: curl ntfy ($prio): $body"
else
curl "$@" -d "${body} [${host}]" "$NTFY_URL" >/dev/null 2>&1 || true
fi
return 0
}
# ---------------------------------------------------------------------------
# Health wait (same shape as deploy.sh)
# ---------------------------------------------------------------------------
wait_health() {
local deadline; deadline=$(( $(date +%s) + 120 ))
while (( $(date +%s) < deadline )); do
local status unhealthy
status="$(dc ps --format '{{.Service}} {{.Health}}' 2>/dev/null || true)"
unhealthy="$(echo "$status" | awk '$2 != "healthy" && $2 != "" {print $1}')"
[[ -z "$unhealthy" && -n "$status" ]] && return 0
sleep 5
done
return 1
}
# ---------------------------------------------------------------------------
# Apply an update: pin the tag, pull, recreate, verify; roll back on failure.
# ---------------------------------------------------------------------------
apply_update() { # <from> <to> [advisory-ids]
local from="$1" to="$2" advs="${3:-}"
log "Updating copyparty ${from:-?} -> ${to}..."
if [[ "$DRY_RUN" == "1" ]]; then
echo "DRY: set COPYPARTY_TAG=${to}; docker compose pull copyparty; docker compose up -d"
return 0
fi
cp -a "$ENV_FILE" "${ENV_FILE}.bak.$(date -u +%Y%m%d%H%M%S)" 2>/dev/null || true
env_set COPYPARTY_TAG "$to"
if ! dc pull copyparty; then
warn "pull failed; restoring COPYPARTY_TAG=${from}."
env_set COPYPARTY_TAG "$from"
send_notice "copyparty update FAILED" "high" "pull of ${to} failed; staying on ${from}"
return 1
fi
dc up -d --remove-orphans
if wait_health; then
log "copyparty ${to} is healthy."
local b="updated ${from:-?} -> ${to}"; [[ -n "$advs" ]] && b="${b} (advisory: ${advs})"
send_notice "copyparty updated" "default" "$b"
return 0
fi
warn "copyparty ${to} did not become healthy; rolling back to ${from}."
env_set COPYPARTY_TAG "$from"
dc up -d --remove-orphans || true
send_notice "copyparty update FAILED" "high" "${to} unhealthy; rolled back to ${from}"
return 1
}
# ---------------------------------------------------------------------------
# Decide the target version for the current policy. Sets globals TARGET,
# TARGET_IDS and SCAN_NOTE (NOT via stdout -- a $() subshell would lose them).
# ---------------------------------------------------------------------------
TARGET=""; TARGET_IDS=""; SCAN_NOTE=""
resolve_target() { # <current-version>
local cur="$1"
TARGET=""; TARGET_IDS=""; SCAN_NOTE=""
if [[ -n "$TARGET_VERSION" ]]; then
TARGET="$(normver "$TARGET_VERSION")"; SCAN_NOTE="target override ${TARGET}"; return 0
fi
case "$UPDATE_POLICY" in
off)
SCAN_NOTE="policy=off (no changes)"; return 0 ;;
security)
local scan first rest
scan="$(advisory_scan "$cur")"
first="$(printf '%s' "$scan" | head -n1)"
rest="$(printf '%s' "$scan" | sed -n '2p')"
case "$first" in
VULN*)
TARGET="$(printf '%s' "$first" | awk '{print $2}')"
TARGET_IDS="$rest"
SCAN_NOTE="VULNERABLE (${rest:-?}); patched in ${TARGET:-?}"
# if the feed names no patched version, fall forward to latest
[[ -z "$TARGET" ]] && TARGET="$(normver "$(latest_version)")" ;;
OK) SCAN_NOTE="no advisory matches ${cur}" ;;
*) SCAN_NOTE="advisory check unavailable (${first})" ;;
esac ;;
latest|*)
TARGET="$(normver "$(latest_version)")"
SCAN_NOTE="policy=latest" ;;
esac
}
# ---------------------------------------------------------------------------
# Subcommands
# ---------------------------------------------------------------------------
preflight() {
[[ $EUID -eq 0 ]] || die "Run as root."
[[ -f "$STACK_DIR/docker-compose.yml" ]] || die "No stack at $STACK_DIR (set STACK_DIR)."
command -v docker >/dev/null 2>&1 || die "docker not found."
}
do_check() {
preflight
local cur latest scan
cur="$(current_version)"
latest="$(normver "$(latest_version)")"
log "Running version: ${cur:-unknown} | latest release: ${latest:-unknown} | policy: ${UPDATE_POLICY}"
scan="$(advisory_scan "$cur")"
local first ids; first="$(printf '%s' "$scan" | head -n1)"; ids="$(printf '%s' "$scan" | sed -n '2p')"
local prio="min" note=""
case "$first" in
VULN*) note="VULNERABLE: ${ids} (patched ${first#VULN })"; prio="high"; warn "$note" ;;
OK) note="no known advisory affects ${cur:-?}"; log "$note" ;;
*) note="advisory check unavailable (${first})"; warn "$note" ;;
esac
if [[ -n "$cur" && -n "$latest" ]] && ver_gt "$latest" "$cur"; then
log "A newer release is available: ${cur} -> ${latest}"
[[ "$prio" == "min" ]] && prio="default"
note="${note}; newer release ${latest} available"
fi
send_notice "copyparty check" "$prio" "${note}"
}
do_run() {
preflight
[[ "$DRY_RUN" == "1" ]] || { install -d -m 0755 "$(dirname "$LOG")" 2>/dev/null || true; echo "=== copyparty-update $(date -u +%FT%TZ) ===" >> "$LOG"; }
local cur
cur="$(current_version)"
resolve_target "$cur"
log "current=${cur:-?} | ${SCAN_NOTE}"
if [[ -z "$TARGET" ]]; then
log "No update to apply."
return 0
fi
if [[ -n "$cur" ]] && ! ver_gt "$TARGET" "$cur"; then
log "Target ${TARGET} is not newer than ${cur}; nothing to do."
return 0
fi
apply_update "$cur" "$TARGET" "$TARGET_IDS"
}
do_update() {
# force: default to latest unless TARGET_VERSION/policy says otherwise
preflight
local cur
cur="$(current_version)"
if [[ -z "$TARGET_VERSION" && "$UPDATE_POLICY" == "off" ]]; then UPDATE_POLICY=latest; fi
resolve_target "$cur"
[[ -n "$TARGET" ]] || die "Could not determine a target version (${SCAN_NOTE})."
if [[ -n "$cur" ]] && ! ver_gt "$TARGET" "$cur" && [[ -z "$TARGET_VERSION" ]]; then
log "Already on ${cur} (latest ${TARGET}); nothing to do."
return 0
fi
apply_update "$cur" "$TARGET" "$TARGET_IDS"
}
# Write /etc/copyparty-update.conf so the scheduled run inherits these.
write_conf() {
cat > "$COPYPARTY_UPDATE_CONF" <<CONF
# Defaults for the scheduled copyparty updater (deployments/copyparty/update.sh).
# Environment variables still override these at runtime.
STACK_DIR="${STACK_DIR}"
UPDATE_POLICY="${UPDATE_POLICY}"
VC_FEED="${VC_FEED}"
NOTIFY="${NOTIFY}"
CONF
chmod 644 "$COPYPARTY_UPDATE_CONF"
log "Wrote $COPYPARTY_UPDATE_CONF"
}
do_install() {
[[ $EUID -eq 0 ]] || die "Run as root."
write_conf
case "$(osfam)" in
alpine)
install -d -m 0755 /etc/periodic/daily
cat > /etc/periodic/daily/copyparty-update <<EOF
#!/bin/sh
exec bash "$SELF" run
EOF
chmod +x /etc/periodic/daily/copyparty-update
command -v rc-update >/dev/null 2>&1 && { rc-update add crond default >/dev/null 2>&1 || true; rc-service crond start >/dev/null 2>&1 || true; }
log "Scheduled daily via /etc/periodic/daily/copyparty-update (policy=${UPDATE_POLICY})." ;;
*)
cat > /etc/systemd/system/copyparty-update.service <<EOF
[Unit]
Description=copyparty container updater
After=docker.service
[Service]
Type=oneshot
ExecStart=/usr/bin/env bash $SELF run
EOF
cat > /etc/systemd/system/copyparty-update.timer <<EOF
[Unit]
Description=Daily copyparty update check
[Timer]
OnCalendar=daily
Persistent=true
RandomizedDelaySec=1h
[Install]
WantedBy=timers.target
EOF
systemctl daemon-reload
systemctl enable --now copyparty-update.timer >/dev/null 2>&1 || true
log "Scheduled daily via systemd timer copyparty-update.timer (policy=${UPDATE_POLICY})." ;;
esac
}
do_uninstall() {
[[ $EUID -eq 0 ]] || die "Run as root."
rm -f /etc/periodic/daily/copyparty-update
if command -v systemctl >/dev/null 2>&1; then
systemctl disable --now copyparty-update.timer >/dev/null 2>&1 || true
rm -f /etc/systemd/system/copyparty-update.timer /etc/systemd/system/copyparty-update.service
systemctl daemon-reload >/dev/null 2>&1 || true
fi
log "Removed the scheduled copyparty updater."
}
case "${1:-check}" in
check) do_check ;;
run) do_run ;;
update) do_update ;;
install) do_install ;;
uninstall) do_uninstall ;;
*) die "Usage: update.sh [check|run|update|install|uninstall]" ;;
esac
+90
View File
@@ -0,0 +1,90 @@
# Copy to .env and fill in. docker compose picks .env up automatically.
# deploy.sh seeds this file on first run and never rewrites it afterwards.
#
# NOTE: values are NOT quoted and must not contain a dollar sign (Compose
# interpolates .env). Secrets never live here: the admin oper password is in
# secrets/admin.pass (0600) and only its bcrypt hash is in ircd/ircd.yaml.
# --- Identity -----------------------------------------------------------------
# Public hostname (lower-case). Becomes Ergo's server.name, the Let's Encrypt
# cert subject, the Caddy site and the allowed websocket origin. Immutable after
# the first deploy (Ergo refuses to rehash a changed server name).
ERGO_DOMAIN=irc.example.com
# Email for Let's Encrypt registration / expiry notices.
ACME_EMAIL=admin@example.com
# IRC network name (ISUPPORT NETWORK=, welcome line). Letters, digits, . _ -
# only -- no spaces. deploy.sh defaults it to ERGO_DOMAIN.
NETWORK_NAME=irc.example.com
# --- Image --------------------------------------------------------------------
# ghcr.io tags carry the leading 'v' (v2.19.1). deploy.sh pins the newest
# release on first deploy; update.sh bumps this when it updates. Never pin
# 'stable'/'latest' here -- those are separate builds and defeat rollback.
ERGO_IMAGE=ghcr.io/ergochat/ergo
ERGO_TAG=v2.19.1
# uid:gid Ergo runs as inside the container = the host's `ergo` system user
# (created by deploy.sh; owns ./ircd). Filled in by deploy.sh.
ERGO_UID=
ERGO_GID=
# --- Listeners ----------------------------------------------------------------
# 0 (default): plaintext 6667 on 127.0.0.1 only (healthcheck + ergoctl).
# 1: ALSO serve plaintext IRC on :6667 publicly, with STS advertised so capable
# clients pin TLS. Credentials sent over 6667 cross the internet in clear.
# Chosen at the FIRST deploy: it selects a listener in ircd/ircd.yaml, which
# deploy.sh never rewrites. To change it later, edit that listener with
# 'ergoctl edit', set the value here, 'ergoctl restart', then re-run deploy.sh
# so the host firewall matches.
PLAINTEXT=0
# --- Message history ----------------------------------------------------------
# Ergo keeps history in RAM by default, so it is lost on every restart --
# including the ones the updater performs. Persist it with:
# sqlite (default) a file beside ircd.db; no extra container
# postgres a pinned PostgreSQL container (docker-compose.postgres.yml)
# off RAM only, upstream behaviour
# Needs Ergo >= 2.18.0. How long messages are kept is governed by
# history.restrictions.expire-time in ircd.yaml (1 week by default), NOT here.
# Chosen at the first deploy; switch later with 'ergoctl history <backend>'
# (which does not migrate existing messages).
HISTORY=sqlite
# Set by deploy.sh so every `docker compose` command in this stack sees the same
# services. It gains ':docker-compose.postgres.yml' when HISTORY=postgres.
COMPOSE_FILE=docker-compose.yml
# --- PostgreSQL (only when HISTORY=postgres) ----------------------------------
# The MAJOR version is pinned on purpose: PostgreSQL refuses to start on a data
# directory written by a different major, so update.sh never touches this image.
# Changing it is a deliberate dump-and-restore -- see the README.
POSTGRES_TAG=17-alpine
POSTGRES_USER=ergo
POSTGRES_DB=ergo_history
# Published on 127.0.0.1 only. The password lives in secrets/postgres.pass
# (0600), never here -- compose interpolates .env and a dollar sign would break it.
POSTGRES_PORT=5432
# --- Updates (update.sh) ------------------------------------------------------
# This file is where the scheduled jobs read these four knobs from; edit them
# here. (An explicit environment variable, or a value pinned in
# /etc/ergo-update.conf, overrides what is set here for that run.)
# latest - update to the newest release whenever one exists (default)
# security - update ONLY when a published GitHub security advisory covers the
# running version, or a release in range has a "### Security"
# section in its notes
# off - never change the running version (check/notify only)
UPDATE_POLICY=latest
# Releases whose notes announce "Compatibility breaks" are held for review and
# only notified, unless FORCE_UPDATE=1.
FORCE_UPDATE=0
# Seconds of warning (NOTICE to all users) before an update restarts Ergo.
UPDATE_GRACE=60
# 1 = the daily run also pulls a newer Caddy image and recreates it (drops every
# web-client websocket at that moment). 0 = Caddy updates via `ergoctl update caddy`.
CADDY_AUTOUPDATE=0
# --- Image tags ---------------------------------------------------------------
CADDY_TAG=2-alpine
+150
View File
@@ -0,0 +1,150 @@
# __NETWORK_NAME__ — admin & moderator cheat sheet
Two places to work:
- **From IRC**, as an operator — moderation, accounts, channels. Any moderator.
- **On the host**, via `ergoctl` — the server itself. Needs root/SSH.
---
## Become an operator
```
/OPER <name> <password>
```
Per connection: do it again each time you reconnect. Operator status is hidden
from ordinary users. Passwordless login by TLS certificate is available — see
"Operators" in the deployment README.
---
## Moderation — from IRC
`UBAN` is the one command for bans; it works out whether you gave it an IP, a
CIDR, a `nick!user@host` mask or an account name.
| Task | Command |
|---|---|
| Look someone up first | `UBAN INFO <nick>` (shows their real IP) |
| Ban an address | `UBAN ADD 203.0.113.9 DURATION 1d spamming` |
| Ban a range | `UBAN ADD 203.0.113.0/24 DURATION 7d` |
| Suspend an account | `UBAN ADD <account> DURATION 30d` |
| Lift a ban | `UBAN DEL <target>` |
| List bans | `UBAN LIST` |
| Disconnect someone | `KILL <nick> :reason` |
| Require accounts, network-wide | `DEFCON 4` (`DEFCON 5` = normal) |
> A **bare name** is treated as an *account*, not a connected nick. For someone
> who isn't registered, `UBAN INFO` them and ban the IP.
Durations: `1y 12mo 31d 10h 8m 13s`. Omit `DURATION` for permanent.
**Always-on users** stay reachable while disconnected — banning the account
(`UBAN ADD <account>`) is what removes them, not `KILL`.
### Channels
| Task | Command |
|---|---|
| Join a channel you can't enter | `SAJOIN #channel` |
| Force a mode change | `SAMODE #channel +m` |
| Kick | `/kick #channel nick reason` |
| Ban from a channel | `/mode #channel +b nick!*@*` |
| Ask the server how best to ban | `/msg ChanServ HOWTOBAN #channel <nick>` |
| Blacklist a channel entirely | `/msg ChanServ PURGE ADD #channel reason` |
| Undo that | `/msg ChanServ PURGE DEL #channel` |
### Accounts
| Task | Command |
|---|---|
| Register for someone | `/msg NickServ SAREGISTER <user> <password>` |
| Reset a password | `/msg NickServ PASSWD <user> <newpassword>` |
| Suspend | `/msg NickServ SUSPEND ADD <nick> DURATION 30d reason` |
| Unsuspend / list | `/msg NickServ SUSPEND DEL <nick>` · `SUSPEND LIST` |
| Erase permanently | `/msg NickServ ERASE <user>` (two-step, returns a code) |
Two-step commands (`ERASE`, `CS UNREGISTER`, `TRANSFER`, `PURGE`) reply with a
confirmation code — repeat the command with the code appended.
---
## The server — on the host
```bash
ergoctl status # health, versions, TLS, users, scheduled jobs
ergoctl users # who's on
ergoctl logs -f # live server log
ergoctl announce <text> # notice to everyone
ergoctl cmd <raw IRC> # run any of the above as the admin oper
```
| Task | Command |
|---|---|
| Edit config (validates, then applies) | `ergoctl edit` |
| Edit the rules shown on connect | `ergoctl motd` |
| Apply config without disconnecting anyone | `ergoctl rehash` |
| Restart (drops every user) | `ergoctl restart` |
| Add a moderator | `ergoctl oper add <name> chat-moderator` |
| Add an admin | `ergoctl oper add <name> server-admin` |
| Certificate login for an oper | `ergoctl oper certfp <name> <nick> --auto` |
| Back up config + database + history | `ergoctl backup` |
| Restore | `ergoctl restore <file>` |
| Check / renew TLS | `ergoctl cert` · `ergoctl cert sync` |
| Check for updates | `ergoctl update check` |
| Update now (snapshot + rollback) | `ergoctl update update` |
| Message history backend | `ergoctl history [sqlite\|postgres\|off]` |
`ergoctl help` lists everything.
---
## Mode reference
**Channel modes** — `/mode #channel +X`
| | | | |
|---|---|---|---|
| `i` | invite only | `m` | moderated (voiced+ may speak) |
| `k` | requires a key | `n` | no messages from outside |
| `l` | user limit | `s` | secret (hidden from `/list`) |
| `R` | registered users only | `t` | ops set the topic |
| `M` | registered users may speak | `C` | no CTCP |
| `U` | op-moderated (unvoiced msgs go to ops) | `u` | auditorium |
| `b` | ban mask | `e` | ban exception |
| `I` | invite exception | `f` | forward to another channel |
**Member prefixes** — `/mode #channel +o nick`
| | | | | |
|---|---|---|---|---|
| `q` `~` | founder | `a` `&` | admin | `o` `@` | op |
| `h` `%` | half-op | `v` `+` | voice | | |
Make any of these stick across reconnects with
`/msg ChanServ AMODE #channel +o account`.
**User modes** — `/mode yournick +X`
| | | | |
|---|---|---|---|
| `i` | invisible | `R` | only registered users may PM you |
| `T` | block CTCP | `B` | mark as a bot |
| `Z` | secure connection (automatic) | `s` | receive server notices (opers) |
---
## When things go wrong
| Symptom | Do this |
|---|---|
| Clients reject the certificate | `ergoctl cert` — if it says SELF-SIGNED, check DNS and ports 80/443, then `ergoctl cert sync` |
| Config change didn't take | Some settings only apply at startup — `ergoctl restart` |
| Rehash refused a change | The file is kept and reverted; the server's own error is printed |
| Server won't start after an edit | `ergoctl logs` shows the failing line; restore `ircd.yaml.last-good` |
| Update broke something | It rolls back automatically; `ergoctl status` confirms, and `backups/` holds a database snapshot |
| Locked out of `/OPER` | The password is in `secrets/admin.pass` on the host — `ergoctl passwd show` |
| Spam flood in progress | `DEFCON 4` requires accounts network-wide until you set `DEFCON 5` |
Full detail lives in [`README.md`](README.md).
+75
View File
@@ -0,0 +1,75 @@
# Caddyfile for the Ergo stack. Caddy runs in the HOST network namespace.
#
# Roles:
# 1. Obtain and renew the Let's Encrypt certificate for {$ERGO_DOMAIN}. Ergo's
# :6697 listener reuses that cert -- `update.sh certsync` copies it into
# ./ircd and rehashes Ergo (runs every 15 minutes, silent when unchanged).
# 2. Terminate HTTPS for IRC-over-WebSocket: wss://{$ERGO_DOMAIN}/webirc ->
# Ergo's plaintext websocket listener on 127.0.0.1:8097. Caddy adds
# X-Forwarded-For / X-Forwarded-Proto and Ergo trusts them from loopback
# (proxy-allowed-from: localhost), so web users keep their real IP and are
# marked secure (+Z).
# 3. A plain-text landing page with connection details at /.
#
# Add your own site config (e.g. a Gamja web client, see README) as
# conf.d/*.caddy -- deploy.sh installs this file but never touches conf.d/.
# Apply changes with: docker compose restart caddy
{
email {$ACME_EMAIL}
# Host networking would put the admin API on the HOST's 127.0.0.1:2019,
# reachable by every local process and container. The config is static, so
# turn it off; changes are applied by restarting the container.
admin off
# No HTTP/3: it would bind udp/443 on the host (not in the firewall's port
# list) and browsers do not run WebSockets over h3 anyway.
servers {
protocols h1 h2
}
}
# Loopback-only health endpoint for the compose healthcheck (no admin API to ask).
http://127.0.0.1 {
respond "ok" 200
}
{$ERGO_DOMAIN} {
encode zstd gzip
# IRC over WebSocket. Ergo ignores the request path; Caddy proxies the
# Upgrade transparently and supplies X-Forwarded-For/-Proto itself.
handle_path /webirc* {
reverse_proxy 127.0.0.1:8097
}
# Operator additions (web client, redirects, ...). See conf.d/00-readme.caddy.
import conf.d/*.caddy
handle {
header Content-Type "text/plain; charset=utf-8"
respond <<TXT
{$NETWORK_NAME} -- IRC server
Connect with any IRC client:
server: {$ERGO_DOMAIN}
port: 6697 (TLS)
web: wss://{$ERGO_DOMAIN}/webirc (IRC-over-WebSocket endpoint)
Register a nickname: /msg NickServ REGISTER <password>
Powered by Ergo (https://ergo.chat).
TXT 200
}
header {
Strict-Transport-Security "max-age=31536000; includeSubDomains"
X-Content-Type-Options "nosniff"
Referrer-Policy "strict-origin-when-cross-origin"
-Server
}
log {
output stdout
format console
}
}
+416
View File
@@ -0,0 +1,416 @@
# ergo
[Ergo](https://ergo.chat) — a modern IRC server with built-in services
(NickServ/ChanServ/HostServ), message history, always-on "bouncer" clients and
IRCv3 — behind Caddy for automatic Let's Encrypt TLS. Ships a host CLI
(**`ergoctl`**) for day-to-day administration and a health-checked updater
(**`update.sh`**) that also keeps Ergo's TLS certificate in sync with Caddy.
## Docker, not an OpenRC/systemd service
Ergo is a single static Go binary, so a native install is perfectly possible.
This stack still uses Docker Compose because it matches the repo's goals better:
| | Docker (this stack) | Native OpenRC / systemd |
|---|---|---|
| Alpine, Debian, Alma | one compose file, one image | Alpine packages `ergo` in community, but a stable branch keeps the version it shipped with (3.24 → 2.18.0, 3.23/3.22 → 2.16.0) while upstream is at 2.19.1, a *security* release; only edge tracks upstream. Debian/Alma have no package at all → hand-rolled tarball install + a unit per init system |
| Upgrade / rollback | `ERGO_TAG` in `.env`, health-checked, DB snapshot, rollback | replace a binary + hand-restore the DB |
| ACME TLS | stock Caddy (already used by every other stack) | certbot/acme.sh per distro + deploy hook |
| Real client IPs | yes — **host networking** (see below) | yes |
The usual objection to containerised IRC — bans, throttling and IP cloaking
need the client's real address, and Docker's userland proxy hides IPv6 clients
behind the bridge gateway — is solved by running **both** containers with
`network_mode: host`. That also means the host's deny-by-default firewall
genuinely governs the ports (the "Docker bypasses INPUT" caveat from the root
README does not apply here).
## Topology
| Port | Who listens | Purpose |
|---|---|---|
| `6697/tcp` | Ergo | IRC over TLS (Let's Encrypt cert, copied from Caddy) |
| `443/tcp` | Caddy | `wss://DOMAIN/webirc` → Ergo's websocket listener on `127.0.0.1:8097`; text landing page at `/` |
| `80/tcp` | Caddy | ACME HTTP-01 + redirect to HTTPS |
| `127.0.0.1:6667` | Ergo | loopback plaintext: compose healthcheck + `ergoctl` (exempt from ip-limits/bans, treated as secure) |
| `6667/tcp` public | Ergo | only with `PLAINTEXT=1` (STS is then advertised) — **not recommended** |
No HTTP/3 (would bind udp/443), Caddy's admin API is **off** (it would sit on
the host's loopback), and there is no `ports:` section — host mode.
## Required `.env` values
| Variable | Notes |
|---|---|
| `ERGO_DOMAIN` | Public hostname (lower-case). Server name, cert subject, Caddy site and websocket origin. **Immutable** after the first deploy. |
| `ACME_EMAIL` | Let's Encrypt registration email. |
| `NETWORK_NAME` | IRC network name (letters, digits, `. _ -`; no spaces). Defaults to `ERGO_DOMAIN`. |
| `ERGO_TAG` | Pinned image tag **with the `v`** (`v2.19.1`). Set by deploy.sh to the newest release; bumped by update.sh. Never `stable`/`latest` — if a first deploy could not reach GitHub it falls back to `stable`, and the next update run pins it. |
| `PLAINTEXT` | `0` (default) or `1` — public plaintext 6667. **Fixed at the first deploy** (it selects a listener in `ircd.yaml`, which deploy.sh never rewrites); see below to change it. |
| `HISTORY` | `sqlite` (default) / `postgres` / `off` — see [Message history](#message-history). |
| `UPDATE_POLICY` | `latest` (default) / `security` / `off` — see [Updates](#updates). |
See [`.env.example`](.env.example) for the rest (`UPDATE_GRACE`,
`FORCE_UPDATE`, `CADDY_AUTOUPDATE`, `CADDY_TAG`). `.env` must not contain `$`
(Compose interpolates it) — secrets never live there. The four update knobs are
read from `.env` by the scheduled jobs; an environment variable, or a value
pinned in `/etc/ergo-update.conf`, overrides it for that run.
Deploy-time-only options (not stored in `.env`): `ERGO_AUTOUPDATE=0` installs
`update.sh` without scheduling the daily update (cert sync is still scheduled),
and `CERT_WAIT` (default 180s) bounds the wait for the first Let's Encrypt cert.
To switch `PLAINTEXT` after the first deploy: change the plaintext listener with
`ergoctl edit` (`"127.0.0.1:6667":` ↔ `":6667":`, and `server.sts.enabled` with
it), set `PLAINTEXT` in `.env`, run `ergoctl restart`, then re-run `deploy.sh`
so the host firewall matches. Passing a conflicting `PLAINTEXT=` to a re-run
stops with those instructions rather than half-applying the change.
## Deploy
```bash
./automations.sh # Deploy on this host → deploy: ergo
```
Or build + run the self-contained artifact:
```bash
./build.sh
scp deploy.sh root@host:
ssh root@host 'bash deploy.sh'
# non-interactive:
# ERGO_DOMAIN=irc.example.com ACME_EMAIL=me@example.com NETWORK_NAME=MyNet SKIP_PROMPTS=1 bash deploy.sh
```
Unattended provisioning: [`cloud-init.yml`](cloud-init.yml). On first run
deploy.sh:
1. installs Docker (+ `openssl`, `jq`, `curl`) and creates the `ergo` system
user the container runs as;
2. registers `80,443,6697/tcp` with the host firewall
(`/etc/firewall/ports.d/ergo.rule`, or ufw/firewalld);
3. pins `ERGO_TAG` to the newest release and seeds `.env`;
4. generates **`ircd/ircd.yaml` once**, from the pulled image's *own*
`default.yaml` (so it always matches the running version): server/network
name, the listeners above, the websocket origin, the cloak suffix, and an
`admin` oper whose random password is saved to `secrets/admin.pass` (0600;
only the bcrypt hash goes into the yaml). Hard post-checks refuse to
continue if the upstream template layout ever changes under the edits;
5. validates the config with `ergo run --smoke` in a throwaway container
**before** anything starts, validates the Caddyfile, then starts the stack
and waits for both healthchecks (Ergo's is IRC-level: a `QUIT` must be
answered with `ERROR`);
6. waits for Caddy's Let's Encrypt cert, copies it into `ircd/` and rehashes
Ergo (until then Ergo serves a self-signed cert from its own `mkcerts`);
7. schedules `update.sh certsync` every 15 min and `update.sh run` daily
(`ERGO_AUTOUPDATE=0` schedules only the cert sync).
Re-runs are idempotent: they never touch `ircd/ircd.yaml`, `secrets/` or
`caddy/etc/conf.d/`, and they adopt the deployed settings from `.env` rather
than this shell's defaults, so a plain `bash deploy.sh` cannot silently close a
`PLAINTEXT=1` firewall port or reset your update policy. A setting you do pass
explicitly is written back to `.env`. A changed `Caddyfile` triggers a Caddy
restart, since Caddy has no admin API here.
## Administer: `ergoctl`
Installed at `/usr/local/bin/ergoctl`. It talks to Ergo over the loopback
listener as the `admin` oper.
| Command | What it does |
|---|---|
| `ergoctl status` / `users` / `logs [-f]` / `version` | health, versions, TLS state, user count; server log with health-probe noise filtered |
| `ergoctl edit` / `motd` | `$EDITOR` on `ircd.yaml` / `ergo.motd`, then **validate** (throwaway container) and **REHASH**; the previous file is restored if either step fails |
| `ergoctl rehash` | reload config, MOTD and TLS certs without disconnecting anyone |
| `ergoctl restart` / `stop` / `start` | restart the container (drops every user; needed for settings Ergo fixes at startup, such as `server.name`, `datastore.path`, `casemapping`, `enforce-utf8`, `max-line-len`, `idle-timeouts`). Changes to `docker-compose.yml` or `.env` need `docker compose up -d` instead |
| `ergoctl oper list\|add\|passwd\|certfp\|rm` | manage operators in `ircd.yaml`; passwords are generated (24 chars) and printed once. See [Operators](#operators) |
| `ergoctl passwd [show\|rotate]` | the `admin` oper password |
| `ergoctl announce <text>` | NOTICE to everyone |
| `ergoctl kill <nick> [reason]` · `ban add\|del\|list\|info` · `defcon [1-5]` | moderation (`UBAN` under the hood). `ban add` takes an **IP, CIDR, `nick!user@host` mask, or account name** — a bare name is an *account* to suspend, not a connected nick, so run `ban info <nick>` first to get their IP |
| `ergoctl cmd <raw IRC line>` | anything else as the admin oper, e.g. `ergoctl cmd NS SAREGISTER alice hunter2`, `ergoctl cmd CS PURGE ADD #spam` (`CS PURGE DEL` to undo); replies are printed with control characters stripped |
| `ergoctl cert [show\|sync]` | cert on disk vs the one served on 6697; force a sync |
| `ergoctl backup [--live] [dir]` / `restore <file>` | see [Backups](#backups) |
| `ergoctl update …` | passthrough to `update.sh` |
| `ergoctl history [backend]` | show or switch the persistent-history backend |
| `ergoctl debug on\|off` · `caddy-restart` · `shell` | log level; apply Caddyfile/conf.d changes; a shell in the container |
Editing by hand works too: change `ircd/ircd.yaml`, then `ergoctl rehash` — a
rejected config leaves the old one running (Ergo's rehash is transactional), but
the *file* on disk would then be unbootable, which is why `ergoctl edit`
validates first. Ergo's [manual](https://github.com/ergochat/ergo/blob/master/docs/MANUAL.md)
covers every option; the pristine template for your version is in `templates/`.
## TLS
Caddy owns the ACME account and renews the cert on its own schedule (ARI-driven).
Ergo terminates TLS itself on 6697, so the cert has to reach `ircd/fullchain.pem`
+ `privkey.pem`. `update.sh certsync` (every 15 minutes, via busybox crond or a
systemd timer) does that the way upstream's certbot deploy-hook recipe does,
plus safety checks:
- picks the newest `<domain>.crt` under `caddy/data/caddy/certificates/*/` (Let's
Encrypt or the ZeroSSL fallback), skipping a renewal that is still mid-write;
- verifies the key matches the cert (certmagic writes them as separate files and
generates a new key on every renewal), stages both in `ircd/` and installs
them with two atomic `mv`s — Ergo never sees a half pair, which would be
fatal at its next start;
- `SIGHUP`s Ergo (rehash swaps the TLS config for new connections only) and
confirms the fingerprint served on 6697 changed; if not, the previous pair is
restored and you get an ntfy alert.
Silent when nothing changed. The daily run also warns when the served cert is
still self-signed (DNS/80/443 not right yet) or expires within 14 days
(certsync not running). `ergoctl cert` shows both sides.
## Operators
An oper name is a **credential, not a nick** — you authenticate as one and keep
your own nickname. Deploy creates `admin` (class `server-admin`, `hidden: true`),
whose password is in `secrets/admin.pass` (`ergoctl passwd show`):
```
/OPER admin <password>
```
Add your own rather than sharing that one. Ergo casefolds oper names, so use
lowercase letters, digits, `_` and `-`, starting with a letter:
```bash
ergoctl oper add alice server-admin
```
The class is `chat-moderator` by default; `server-admin` adds rehash, account and
channel administration, defcon and massmessage.
### Client-certificate auth
Instead of typing a password every session, an oper can be identified by the
SHA-256 fingerprint of their TLS client certificate. Find it on the `276` line of
your own `/WHOIS`, or from the certificate:
```bash
openssl x509 -noout -fingerprint -sha256 -in client.pem
```
Then pass that fingerprint, or let ergoctl read it off a connected user:
```bash
ergoctl oper certfp alice 57_Wolve --auto
```
Without `--auto`, Ergo requires **both** the certificate and the password, which
makes the certificate a second factor. With `--auto` the password is removed and
the certificate alone grants operator status the moment they connect — nothing
typed, nothing replayable, but anyone holding that key is an operator on sight.
`ergoctl oper certfp alice --clear` revokes it.
**This is refused for `admin`, deliberately.** `ergoctl` opers up over the
loopback plaintext listener, which presents no client certificate, so a certfp
there would make Ergo reject it — and every ergoctl command needing oper,
including the scheduled jobs, would stop working. Keep `admin` on its password
and put certificates on personal opers.
## Message history
Ergo keeps channel and DM history for `CHATHISTORY`, `/HISTORY`, autoreplay-on-join
and always-on (bouncer) clients. **Upstream keeps it in RAM**, so it is lost on
every restart — including the ones this stack's updater performs. So persistence
is on by default here:
| `HISTORY` | What it does |
|---|---|
| `sqlite` *(default)* | A file at `ircd/ergo_history.db`, next to the account database. No extra container, backed up with everything else. Right for a single-node server. |
| `postgres` | A pinned PostgreSQL container (`docker-compose.postgres.yml`), reachable only on `127.0.0.1`. Choose it if you want a real database to query, or already run Postgres. |
| `off` | Upstream behaviour: RAM only. |
Both SQL backends need **Ergo 2.18.0+** (deploy.sh refuses older pins). How long
messages are kept is `history.restrictions.expire-time` in `ircd.yaml` — **one
week by default** — not a size limit: enabling persistence removes the in-memory
`channel-length`/`client-length` caps. Raise or lower it with `ergoctl edit`, and
note the privacy point upstream flags: persisted messages are personal data, so
check what your jurisdiction expects of you before extending retention.
```bash
ergoctl history # backend, database size / connectivity, retention
ergoctl history postgres # switch (restarts Ergo; does NOT migrate messages)
```
Switching backends leaves the old store on disk and starts the new one empty —
there is no migration path, so pick one at deploy time if you can.
### PostgreSQL
`HISTORY=postgres` adds `docker-compose.postgres.yml` to `COMPOSE_FILE` in `.env`,
which `docker compose` reads by itself — so every command in this stack sees the
same services with no extra flags. The container publishes **only to
`127.0.0.1:5432`**, Ergo reaches it there (the host-networked containers cannot
use compose service DNS), and `depends_on: service_healthy` keeps Ergo from
starting before the database accepts connections. The password is generated into
`secrets/postgres.pass` (0600) and passed via `POSTGRES_PASSWORD_FILE`, so it
never lands in `.env`.
**The major version is pinned and the updater never touches it.** PostgreSQL
refuses to start on a data directory written by a different major version, so
moving from `17-alpine` to `18-alpine` is a deliberate dump-and-restore:
```bash
ergoctl backup # contains a pg_dump of the history database
ergoctl history off && docker compose down -v postgres # drops the old data volume
# set POSTGRES_TAG=18-alpine in .env, then:
ergoctl history postgres && ergoctl restore backups/ergo-backup-<ts>.tar.gz
```
Note that Postgres 18 also changed the image's default data directory layout; the
compose file pins `PGDATA` explicitly so a future image default cannot move it.
## Updates
```bash
ergoctl update check # running vs latest, published security advisories, cert state
# reports only, with one exception: a floating ERGO_TAG
# ('stable') is pinned to the running version in .env
ergoctl update update # update now (TARGET_VERSION=2.19.1 to pin); honours the safety rails
ergoctl update install # (re)schedule / uninstall to stop
```
`UPDATE_POLICY` (in `.env`) drives the daily `run`:
| Policy | Behaviour |
|---|---|
| `latest` *(default)* | Update to the newest release whenever one exists. |
| `security` | Update **only** when a published GitHub security advisory covers the running version, or a release since it has a `### Security` section in its notes (upstream only opened GHSAs from 2.19.1). |
| `off` | Never change the running version (check/notify only). |
Every update, scheduled or manual:
1. pulls the pinned `ghcr.io/ergochat/ergo:vX.Y.Z` and **pre-flights** it against
a copy of `ircd/` (`ergo run --smoke` — also dry-runs a DB schema upgrade);
2. `NOTICE`s connected users and waits `UPDATE_GRACE` seconds (default 60);
3. stops Ergo and snapshots `ircd.db` to `backups/` (consistent copy; buntdb is
append-only, so a live copy could miss the last second), keeps the last 5;
4. flips `ERGO_TAG`, starts, waits for the healthcheck **and** a registration
handshake;
5. on failure: previous tag back; if the old version refuses the upgraded
database ("Database requires update") the snapshot is restored too. Cause
and snapshot path go into the ntfy message;
6. on success: saves the new image's `default.yaml` under `templates/` and, if it
differs from the previous version's, writes `templates/diff-vA-vB.txt` so
new/renamed options are visible.
Releases whose notes announce **compatibility breaks** are held (one ntfy, not
daily) until you run `ergoctl update update` or set `FORCE_UPDATE=1`. Restarting
Ergo disconnects every user (there is no hot restart), so pick `security` if
that matters more than being current. Caddy is only updated by
`ergoctl update caddy` (or daily with `CADDY_AUTOUPDATE=1`); recreating it drops
web-client websockets. Update results reuse the ntfy config at
`/etc/ssh-notify.conf` when present.
## Backups
`ergoctl backup` writes `backups/ergo-backup-<ts>.tar.gz` (0600) with
`ircd.yaml`, `ircd.db`, `ergo.motd`, the message history (the SQLite file, or a
`pg_dump` — a file copy of a live PostgreSQL data directory would not be a valid
backup) and a `meta` file (version, history backend, date). By default
it stops Ergo for a few seconds for a consistent copy (`--live` skips that).
Put an age public key in `$STACK_DIR/age-recipients.txt` (e.g. from
`globals/age-pubkey.txt`) **and install `age`** (`apk add age`, `apt install age`,
`dnf install age` — deploy.sh does not) and backups are encrypted with it; with
the recipients file present but `age` missing, `ergoctl backup` refuses rather
than writing the account and oper hashes out in the clear. TLS files are not
included — certsync regenerates them. `ergoctl restore <file>` stops Ergo, keeps the current config and database as
`backups/{ircd.yaml,ircd.db}.pre-restore.<ts>`, warns if the backup came from a
newer version (older Ergo cannot read a newer schema), restores, validates the
result before starting, and waits for health. Because the archive carries the
*original* host's credentials, restore re-points `datastore.postgresql.password`
at this host's `secrets/postgres.pass` and re-hashes this host's
`secrets/admin.pass` into the config — otherwise a rebuild would come up healthy
but leave you unable to `/OPER`. An encrypted backup needs the private key:
```bash
AGE_IDENTITY=/root/age.key ergoctl restore backups/ergo-backup-<ts>.tar.gz.age
``` Ergo's own `datastore.autoupgrade`
leaves `ircd/ircd.db.v<N>.<ts>.bak` files behind on schema upgrades; prune them
once you are happy with a release.
## Web client (optional)
The `wss://DOMAIN/webirc` endpoint is ready; Ergo only accepts browser
connections whose `Origin` is `https://DOMAIN` (`server.websockets.allowed-origins`).
To host [Gamja](https://codeberg.org/emersion/gamja) on the same origin, unpack a
release tarball into `caddy/www/gamja`, add
`{"server": {"url": "/webirc", "autojoin": "#lobby"}}` as `config.json`, mount
`./caddy/www:/www:ro` in the caddy service and drop a `handle { root * /www/gamja
file_server }` block into `caddy/etc/conf.d/` (see `00-readme.caddy`), then
`ergoctl caddy-restart`.
## Security model
- **Host loopback is Ergo's trust boundary.** Loopback peers are exempt from
bans and ip-limits, are treated as secure, and (because
`proxy-allowed-from: [localhost]` is what lets Caddy's `X-Forwarded-For`
through) may assert a client IP. Every local process and every other
host-network container can reach `127.0.0.1:6667`/`:8097`. This stack is for
a single-purpose container host without untrusted local users; deploy.sh
warns about other host-network containers.
- **Ergo runs unprivileged**: as the host's `ergo` system user, read-only root
filesystem, all capabilities dropped, `no-new-privileges`. `ircd/` (database,
config with the oper hash, TLS key) is `0700 ergo:ergo`; `secrets/`,
`backups/`, `caddy/data` are `0700 root`.
- **Root never follows a symlink into `ircd/`.** That directory is writable by
the container uid while the cert sync and every `ergoctl` config edit run as
root, so a plain `cp`/`install`/`>` there would let code execution inside Ergo
redirect a root write onto any file on the host. Every such write stages under
`$STACK_DIR` (0700 root) and lands with `mv` — `rename(2)` replaces a symlink
instead of following it — and root reads use `cp -P` or refuse outright.
- **Caddy runs as root** (it needs `NET_BIND_SERVICE` for 80/443) with every
other capability dropped, admin API off, HTTP/3 off.
- **Loopback plaintext is the control plane by design**: OPER over loopback is
not a downgrade (Ergo marks it secure), and it lets `ergoctl` work even when a
ban or throttle would lock out a remote client. `PLAINTEXT=1` puts real users'
credentials on the wire in clear — leave it at `0`.
- The admin password is sent to Ergo only over loopback and stored only in
`secrets/admin.pass`. `ergo genpasswd` hashes at bcrypt cost 4, so oper
passwords are always generated (24 random chars), never typed.
## Files
| File | Purpose |
|---|---|
| `docker-compose.yml` | caddy + ergo, both `network_mode: host`, hardened; IRC-level healthcheck. |
| `docker-compose.postgres.yml` | Optional PostgreSQL overlay for `HISTORY=postgres` (loopback-only, pinned major). |
| `Caddyfile` | ACME, `/webirc` websocket proxy, text landing page, `admin off`, no h3; imports `conf.d/*.caddy`. |
| `conf.d-readme.caddy` | Installed once as `caddy/etc/conf.d/00-readme.caddy` (operator drop-ins; Gamja example). |
| `ergo.motd` | MOTD template (`__NETWORK_NAME__`, `__DOMAIN__`, `__HISTORY_NOTE__`). |
| [`USER-GUIDE.md`](USER-GUIDE.md) | Quick-start for people new to IRC. Rendered to `docs/` on the host. |
| [`ADMIN-CHEATSHEET.md`](ADMIN-CHEATSHEET.md) | Moderator and admin quick reference. Rendered to `docs/` on the host. |
| `ergolib.sh` | Shared helpers: `.env`, compose, health/IRC probes, `--smoke` validator, oper edits, certsync, ntfy. |
| `update.sh` | Updater + certsync + Caddy update + scheduling. |
| `ergoctl` | Host admin CLI. |
| `.env.example` | Stack tunables. |
| `deploy.sh` / `build.sh` | Self-contained installer + archive embedder. |
| `cloud-init.yml` | Fresh-VM bootstrap (harden SSH, then deploy). |
On the host: `/srv/ergo/{ircd,caddy/{etc,data,config},secrets,backups,templates}`
plus `.state/` (notification de-duplication) and `.env.bak.<ts>` copies from each
update (last 5 kept); `/etc/ergo-update.conf` (where the jobs find the stack);
`/usr/local/bin/ergoctl`; `/var/log/ergo-update.log` (the **daily update run**
logs there, since busybox crond has nowhere else to put its output — cert-sync
failures instead surface through ntfy and as `certsync: LAST FAILURE` in
`ergoctl status`).
## Notes
- Only one host-networked Caddy fits on a box: don't co-locate this stack with
another Caddy/80/443 stack (deploy.sh checks the ports).
- DNS for `ERGO_DOMAIN` must resolve to the host and 80/443 be reachable before
deploy for the cert to issue; Ergo stays up on a self-signed cert meanwhile.
- Ergo's built-in registration is open by default, and **`allow-before-connect`
is also on**, so accounts can be created by a client that has not finished
connecting, throttled only globally (30 attempts per 10 minutes ≈ 4,300/day).
Each account is a durable row in `ircd.db`. For a private network set
`accounts.registration.enabled: false`, or keep registration but set
`allow-before-connect: false`, with `ergoctl edit`.
- The admin oper password is printed by `deploy.sh` only on the first run and
only to a terminal. Unattended runs (cloud-init) print the path instead, so the
credential does not end up in the provider's serial-console log; read it with
`ergoctl passwd show`.
+119
View File
@@ -0,0 +1,119 @@
# __NETWORK_NAME__ — new to IRC? Start here
IRC is group chat. You pick a nickname, join channels (chat rooms whose names
start with `#`), and talk. Commands start with `/`.
---
## 1. Connect
| | |
|---|---|
| **Server** | `__DOMAIN__` |
| **Port** | `6697` |
| **TLS/SSL** | **on** (required) |
| **Network** | __NETWORK_NAME__ |
Any IRC client works. Good starting points: **HexChat** (Windows/Linux),
**Textual** or **Halloy** (macOS), **Revolution IRC** (Android), **Goguma** (iOS).
> If your client asks for a "network" and a "server", the network is just a
> label — the server address above is what matters.
---
## 2. Claim your nickname
On this network your nickname *is* your account, so register it or someone else
could use it later.
```
/msg NickServ REGISTER your-password
```
Then tell your client to log in automatically: enable **SASL**, method
**PLAIN**, with your nickname as the username and that password. Most clients
have this in the same dialog as the server address.
Not using SASL yet? Log in by hand each time:
```
/msg NickServ IDENTIFY your-password
```
Someone else on your nick? `/msg NickServ GHOST yournick` disconnects them.
---
## 3. Talk
| Do this | Command |
|---|---|
| Join a channel | `/join #channel` |
| Leave | `/part #channel` |
| Private message | `/msg someone hello` |
| See who's here | `/names` |
| Look someone up | `/whois someone` |
| Set yourself away | `/away back in 10` (`/away` to return) |
| Change nick | `/nick newnick` |
| List channels | `/list` |
| Quit | `/quit see you` |
Type normally to speak in the channel you're looking at. `/me waves` writes an
action line.
---
## 4. Scrollback
Messages are stored for **__HISTORY_DAYS__**, so you see what you missed while
away — no need to stay connected. Most modern clients fetch it automatically;
otherwise `/history #channel 50` pulls the last 50 lines.
To have a channel's recent messages replayed every time you join:
```
/msg NickServ SET autoreplay-lines 25
```
---
## 5. Privacy, plainly
- Your IP address is **hidden** from other users (you'll see a scrambled
hostname). Server staff can still see it.
- Channels are **logged for __HISTORY_DAYS__**. Treat anything you type as
recorded — including private messages, which are stored for logged-in users.
- Turn off storage of your own DMs with `/msg NickServ SET dm-history off`.
- Connections are encrypted (TLS). Your client shows `+Z` when secure.
---
## 6. Making your own channel
Just join one that doesn't exist yet — you become its founder:
```
/join #mychannel
/msg ChanServ REGISTER #mychannel
```
Registering keeps it yours. A few useful settings once you own it:
| Want | Command |
|---|---|
| Invite-only | `/mode #mychannel +i` then `/mode #mychannel +I nickname` |
| Registered users only | `/mode #mychannel +R` |
| Set the topic | `/topic #mychannel Welcome!` |
| Give someone ops | `/msg ChanServ OP #mychannel nickname` |
| Make that permanent | `/msg ChanServ AMODE #mychannel +o nickname` |
---
## 7. Stuck?
- `/msg NickServ HELP` and `/msg ChanServ HELP` list everything.
- `/motd` shows the network's rules.
- `/admin` shows who runs the server.
Be kind, and enjoy yourself.
+70
View File
@@ -0,0 +1,70 @@
#!/usr/bin/env bash
#
# build.sh -- (re)embed the loose deployment files into deploy.sh as a base64
# tar.gz payload after __ARCHIVE_BELOW__. Idempotent: strips any existing
# payload first.
#
# Run this after editing ANY embedded file below, then re-stage deploy.sh --
# the deployed stack uses the EMBEDDED copies, not the loose files.
set -euo pipefail
DIR=$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)
SCRIPT="$DIR/deploy.sh"
MARKER="__ARCHIVE_BELOW__"
# README, cloud-init and the build/deploy scripts are NOT embedded. ircd.yaml is
# generated on the host from the pulled image's own default config.
FILES=(
docker-compose.yml
docker-compose.postgres.yml
Caddyfile
conf.d-readme.caddy
ergo.motd
USER-GUIDE.md
ADMIN-CHEATSHEET.md
.env.example
ergolib.sh
update.sh
ergoctl
)
[[ -f "$SCRIPT" ]] || { echo "deploy.sh not found at $SCRIPT" >&2; exit 1; }
for f in "${FILES[@]}"; do
[[ -f "$DIR/$f" ]] || { echo "Missing $DIR/$f" >&2; exit 1; }
done
# Guard: every file deploy.sh reads out of the extracted archive must be in
# FILES, and must be in deploy.sh's own EMBEDDED manifest. Forgetting either is
# a deploy-time abort on every host, so catch it here instead.
missing=0
while IFS= read -r ref; do
case " ${FILES[*]} " in *" $ref "*) ;; *) echo "deploy.sh reads \$SCRIPT_DIR/$ref but build.sh does not embed it" >&2; missing=1 ;; esac
done < <(grep -oE '\$SCRIPT_DIR/[A-Za-z0-9._-]+' "$SCRIPT" | sed 's|^\$SCRIPT_DIR/||' | sort -u)
for f in "${FILES[@]}"; do
grep -q "EMBEDDED=(.*$f" "$SCRIPT" || { echo "$f is embedded but missing from deploy.sh's EMBEDDED manifest" >&2; missing=1; }
done
(( missing == 0 )) || { echo "Refusing to build a deploy.sh that would abort on the target host." >&2; exit 1; }
# Syntax-check the bash payload before embedding it.
for f in ergolib.sh update.sh ergoctl deploy.sh; do
bash -n "$DIR/$f" || { echo "bash -n failed: $f" >&2; exit 1; }
done
PAYLOAD=$(tar -czf - -C "$DIR" "${FILES[@]}" | base64)
TMP=$(mktemp)
trap 'rm -f "$TMP"' EXIT
sed "/^${MARKER}\$/,\$d" "$SCRIPT" > "$TMP"
{
echo "$MARKER"
echo "$PAYLOAD"
} >> "$TMP"
mv "$TMP" "$SCRIPT"
chmod +x "$SCRIPT"
trap - EXIT
size=$(wc -c < "$SCRIPT")
echo "Built $SCRIPT (${size} bytes)"
+43
View File
@@ -0,0 +1,43 @@
#cloud-config
#
# Ergo (IRC server) -- harden SSH, then deploy, on a fresh host.
#
# Fill in REPO_URL and the values in the runcmd block, then paste this as the
# instance user-data. DNS for ERGO_DOMAIN must point at this host and ports
# 80/443 must be reachable before boot, or the Let's Encrypt cert request fails
# (Ergo then keeps serving a self-signed cert until certsync catches up).
# 6697/tcp (IRC over TLS) is opened on the host firewall as well.
packages:
- git
runcmd:
- hostnamectl set-hostname irc || true
- |
set -e
REPO_URL=https://git.anomalous.dev/57_Wolve/automations.git
REPO_BRANCH=main
HARDEN_SSH=1 # harden SSH on this fresh VM (set 0 to skip)
SSH_PORT=22
ALLOWED_IP= # optional: whitelist your client IP in sshguard
git clone --depth 1 --branch "$REPO_BRANCH" "$REPO_URL" /opt/automations
cd /opt/automations
# Harden SSH first (PQ KEX, key-only auth, sshguard + deny-by-default
# firewall). Because the Ergo stack uses host networking, that firewall
# genuinely governs 80/443/6697 -- deploy.sh registers them with it.
if [ "$HARDEN_SSH" = 1 ]; then
SSH_PORT="$SSH_PORT" ALLOWED_IP="$ALLOWED_IP" SKIP_PROMPTS=1 FORCE=1 \
bash scripts/harden-ssh.sh
fi
# Deploy Ergo. The admin oper password is generated and stored in
# /srv/ergo/secrets/admin.pass (also printed to this log once).
ERGO_DOMAIN=irc.example.com \
ACME_EMAIL=admin@example.com \
NETWORK_NAME=ExampleNet \
HISTORY=sqlite \
PLAINTEXT=0 \
UPDATE_POLICY=latest \
SKIP_PROMPTS=1 \
bash deployments/ergo/deploy.sh
+20
View File
@@ -0,0 +1,20 @@
# Operator drop-ins for the Ergo stack's Caddy site block.
#
# Every conf.d/*.caddy file is imported INSIDE the {$ERGO_DOMAIN} site block, after
# the /webirc websocket route and before the text landing page. deploy.sh installs
# this readme once and never touches conf.d/ again. Apply with:
#
# docker compose restart caddy (or: ergoctl caddy-restart)
#
# Example -- serve the Gamja web client (https://codeberg.org/emersion/gamja) from
# ./caddy/www/gamja (unpack a release tarball there and write a config.json with
# {"server": {"url": "/webirc", "autojoin": "#lobby"}}):
#
# handle {
# root * /www/gamja
# file_server
# }
#
# ...and add `- ./caddy/www:/www:ro` to the caddy service volumes in
# docker-compose.yml. Gamja's Origin (https://{$ERGO_DOMAIN}) already matches
# server.websockets.allowed-origins in ircd.yaml.
File diff suppressed because it is too large Load Diff
@@ -0,0 +1,69 @@
# PostgreSQL overlay -- used ONLY when HISTORY=postgres.
#
# deploy.sh activates it by writing
# COMPOSE_FILE=docker-compose.yml:docker-compose.postgres.yml
# into .env, which docker compose reads by itself. So every `docker compose`
# command in this stack (up, ps, stop, logs, exec) sees the same set of services
# with no extra flags -- unlike compose profiles, where a service can be silently
# absent from one command and present in another.
#
# Networking: ergo and caddy run in the HOST namespace, so they cannot use
# compose's service DNS. PostgreSQL therefore stays on a normal bridge network
# and publishes only to the host's loopback; Ergo reaches it at 127.0.0.1:5432
# (datastore.postgresql.host in ircd.yaml). Nothing is exposed off-box.
#
# The major version is PINNED. PostgreSQL will not start on a data directory
# written by a different major version, so `update.sh` never touches this image.
# Upgrading it is a deliberate dump-and-restore -- see the README ("PostgreSQL").
services:
postgres:
image: postgres:${POSTGRES_TAG:-17-alpine}
container_name: ergo-postgres
restart: unless-stopped
# Loopback only. A Docker-published port bypasses the host INPUT firewall,
# so the bind address is the real restriction here.
ports:
- "127.0.0.1:${POSTGRES_PORT:-5432}:5432"
environment:
POSTGRES_USER: "${POSTGRES_USER:-ergo}"
POSTGRES_DB: "${POSTGRES_DB:-ergo_history}"
# The password is read from a file so it never has to live in .env
# (compose interpolates .env, and a '$' in a password would break it).
POSTGRES_PASSWORD_FILE: /run/secrets/postgres-password
# 17 and below keep the classic layout; set it explicitly so a future
# image default cannot move the data directory under us.
PGDATA: /var/lib/postgresql/data/pgdata
volumes:
- postgres-data:/var/lib/postgresql/data
- ./secrets/postgres.pass:/run/secrets/postgres-password:ro
cap_drop: [ALL]
cap_add:
- CHOWN # initdb/entrypoint fix ownership of PGDATA
- DAC_READ_SEARCH
- FOWNER
- SETGID # the entrypoint drops from root to the postgres user
- SETUID
security_opt: [no-new-privileges:true]
logging:
driver: json-file
options:
max-size: "10m"
max-file: "3"
healthcheck:
test: ["CMD-SHELL", "pg_isready -U \"${POSTGRES_USER:-ergo}\" -d \"${POSTGRES_DB:-ergo_history}\" -q"]
interval: 15s
timeout: 5s
retries: 5
start_period: 30s
# Ergo must not come up before the database is accepting connections: with
# persistent history enabled it fails to start if the backend is unreachable.
# depends_on is orchestration only, so it works across the host/bridge split.
ergo:
depends_on:
postgres:
condition: service_healthy
volumes:
postgres-data:
+90
View File
@@ -0,0 +1,90 @@
# Ergo stack -- Caddy (Let's Encrypt + WebSocket front) + Ergo IRC server.
#
# Topology -- BOTH services run in the HOST network namespace (no Docker NAT):
#
# IRC clients --> :6697 TLS ----------------------> ergo (sees the real IPv4/IPv6)
# web clients --> :443 caddy --/webirc--> 127.0.0.1:8097 ergo websocket (X-Forwarded-For)
# Let's Encrypt --> :80/:443 caddy (ACME); the issued cert is copied into ./ircd for
# Ergo's :6697 listener by `update.sh certsync` (+ SIGHUP rehash)
# loopback --> 127.0.0.1:6667 plaintext: compose healthcheck + `ergoctl` ONLY
# (loopback is exempt from ip-limits/bans and counts as secure)
#
# Why host networking: IRC bans, throttling and IP cloaking key on the client's
# address, and Docker's userland proxy would hide every IPv6 client behind the
# bridge gateway. Host mode also means the host's deny-by-default INPUT firewall
# really governs these ports (deploy.sh registers 80/443/6697). Consequences:
# no `ports:`/`networks:` here (Compose rejects them in host mode), Caddy reaches
# Ergo over 127.0.0.1, and the host's loopback is Ergo's trust boundary -- see the
# README ("Security model").
name: ergo
services:
# ---------------------------------------------------------------------------
# Caddy -- obtains/renews the Let's Encrypt cert, terminates HTTPS for the
# IRC-over-WebSocket path and serves a text landing page. Static config:
# apply Caddyfile edits with `docker compose restart caddy` (admin API is off).
# ---------------------------------------------------------------------------
caddy:
image: caddy:${CADDY_TAG:-2-alpine}
container_name: ergo-caddy
restart: unless-stopped
network_mode: host
cap_drop: [ALL]
cap_add: [NET_BIND_SERVICE] # :80/:443 (root in-container needs only this)
security_opt: [no-new-privileges:true]
volumes:
- ./caddy/etc:/etc/caddy # Caddyfile + conf.d/ (directory mount, not a file)
- ./caddy/data:/data # ACME account + certs; host-readable for certsync
- ./caddy/config:/config
environment:
ERGO_DOMAIN: "${ERGO_DOMAIN}"
ACME_EMAIL: "${ACME_EMAIL}"
NETWORK_NAME: "${NETWORK_NAME}"
logging:
driver: json-file
options:
max-size: "10m"
max-file: "3"
healthcheck:
# Loopback-only health site defined in the Caddyfile (the admin API is off).
test: ["CMD-SHELL", "wget -qO- http://127.0.0.1/ | grep -q '^ok'"]
interval: 30s
timeout: 5s
retries: 3
start_period: 10s
# ---------------------------------------------------------------------------
# Ergo -- the IRC server. Runs as the host's `ergo` system user (deploy.sh
# creates it and chowns ./ircd), read-only rootfs, no capabilities. Every
# file it needs lives in ./ircd: ircd.yaml, ircd.db, ergo.motd, fullchain.pem,
# privkey.pem. `init: true` puts tini at PID 1 so `docker compose kill -s HUP`
# (rehash: reloads config, MOTD and TLS certs) reaches the ergo process.
# ---------------------------------------------------------------------------
ergo:
image: ${ERGO_IMAGE:-ghcr.io/ergochat/ergo}:${ERGO_TAG:-stable}
container_name: ergo
restart: unless-stopped
network_mode: host
init: true
user: "${ERGO_UID:-1000}:${ERGO_GID:-1000}"
cap_drop: [ALL]
security_opt: [no-new-privileges:true]
read_only: true
tmpfs: [/tmp]
volumes:
- ./ircd:/ircd
logging:
driver: json-file
options:
max-size: "10m"
max-file: "3"
healthcheck:
# IRC-level probe: Ergo answers a pre-registration QUIT with "ERROR :Quit"
# and closes. A bare TCP connect would pass even with a wedged server.
# busybox nc exits 0 after a timeout, hence the grep.
test: ["CMD-SHELL", "printf 'QUIT\r\n' | nc -w 3 127.0.0.1 6667 | grep -q '^ERROR'"]
interval: 60s
timeout: 6s
retries: 3
start_period: 20s
+27
View File
@@ -0,0 +1,27 @@
$c[cyan]__NETWORK_NAME__$r -- __DOMAIN__
$c[cyan]============================================================$r
Welcome! This is a small IRC network running Ergo.
$bHouse rules$r
1. Be decent to one another. No harassment, hate speech, or
personal attacks.
2. No spam, flooding, or unsolicited advertising.
3. Keep it legal, and roughly work-safe unless a channel's
topic says otherwise.
4. Don't evade bans. If you disagree with an operator, take it
to a private message rather than into the channel.
5. Channels may add rules of their own -- read the /TOPIC.
$bGetting set up$r
Register your nick /msg NickServ REGISTER <password>
Stay logged in enable SASL PLAIN in your client
Register a channel /msg ChanServ REGISTER #channel
Help /msg NickServ HELP /msg ChanServ HELP
$bWorth knowing$r
__HISTORY_NOTE__
Your IP address is hidden from other users, but not from staff.
Questions or appeals: message an operator -- see /ADMIN.
Enjoy your stay.
+698
View File
@@ -0,0 +1,698 @@
#!/usr/bin/env bash
#
# ergoctl -- administer the Ergo IRC stack from the host.
#
# Installed by deploy.sh as $STACK_DIR/ergoctl with a wrapper at
# /usr/local/bin/ergoctl. Talks to Ergo over the loopback plaintext listener
# (127.0.0.1:6667) as the 'admin' oper, using the password in secrets/admin.pass.
# Config changes are validated in a throwaway container (ergo run --smoke on a
# copy) and applied with REHASH; on failure the previous file is restored.
#
# Run `ergoctl help` for the command list.
set -euo pipefail
SELF_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
log() { printf '\033[1;32m[+]\033[0m %s\n' "$*"; }
warn() { printf '\033[1;33m[!]\033[0m %s\n' "$*" >&2; }
die() { printf '\033[1;31m[x]\033[0m %s\n' "$*" >&2; exit 1; }
# shellcheck source=ergolib.sh
. "$SELF_DIR/ergolib.sh"
load_conf
: "${STACK_DIR:=$SELF_DIR}"
ergo_set_paths
CONF="$IRCD_DIR/ircd.yaml"
MOTD="$IRCD_DIR/ergo.motd"
usage() {
cat <<EOF
ergoctl -- administer the Ergo IRC server (stack: $STACK_DIR)
Status & logs
status containers, versions, TLS cert, user count, schedule
users LUSERS summary
logs [-f] [N] ergo log (health-probe noise filtered; default last 100)
version running and pinned version, latest release
cert [show|sync] cert on disk vs served on :6697; 'sync' pulls Caddy's now
Configuration (validated, then REHASH; reverted on failure)
edit \$EDITOR ircd/ircd.yaml
motd \$EDITOR ircd/ergo.motd
rehash reload config, MOTD and TLS certs (no disconnects)
restart | stop | start the ergo container (restart drops every user)
caddy-restart apply Caddyfile / conf.d changes
debug on|off switch the log level (info <-> debug)
history [sqlite|postgres|off]
show, or switch, the persistent-history backend.
Switching does NOT migrate existing messages and
needs a restart (Ergo fixes the backend at startup).
Operators
oper list
oper add <name> [class] class: chat-moderator (default) | server-admin; prints the password once
oper passwd <name> new random password for an oper
oper certfp <name> <fingerprint|nick> [--auto]
authenticate that oper by TLS client certificate.
Pass a SHA-256 fingerprint, or the nick of a connected
user to read it from the server. Without --auto the
cert AND the password are both required; with --auto
the password is dropped and they are opered on connect.
--clear removes it. Not allowed for 'admin' (ergoctl
opers over loopback with no certificate).
oper rm <name>
passwd [show|rotate] the 'admin' oper password (secrets/admin.pass)
Moderation (as the admin oper)
announce <text> NOTICE to all users
kill <nick> [reason]
ban add <ip|cidr|nick!user@host|account> [duration] [reason] e.g. ban add 203.0.113.9 1d spam
(a bare name is treated as an ACCOUNT to suspend, not a
connected nick -- use 'ban info <nick>' to get their IP first)
ban del <target> | ban list | ban info <target>
defcon [1-5] show or set the DEFCON level
cmd <raw IRC line> anything else, e.g. cmd NS SAREGISTER alice hunter2
(replies are printed; control characters stripped)
Data
backup [--live] [dir] tar.gz of ircd.yaml, ircd.db, ergo.motd and the message
history (SQLite file, or a pg_dump). Stops Ergo briefly
for a consistent copy; --live = crash-consistent.
Encrypted with age when $STACK_DIR/age-recipients.txt exists.
restore <file> restore a backup: stops Ergo, keeps the current config and
database as backups/*.pre-restore.<ts>, reconciles this
host's credentials, validates, then restarts.
Encrypted backup: AGE_IDENTITY=/path/to/age.key ergoctl restore <file>.age
Updates (update.sh)
update [check|update|run|certsync|caddy|install|uninstall]
shell sh inside the ergo container
EOF
}
require_root() { [[ $EUID -eq 0 ]] || die "Run as root."; }
require_stack() { [[ -f "$STACK_DIR/docker-compose.yml" && -f "$CONF" ]] || die "No deployed stack at $STACK_DIR (set STACK_DIR)."; }
# Validate the current ./ircd with the running image; on failure print the log.
validate_or_fail() { # <what>
if ergo_validate_config; then rm -f "$VALIDATE_LOG"; return 0; fi
warn "$1 did not load in $(ergo_image):"
tail -n 15 "$VALIDATE_LOG" >&2 || true
rm -f "$VALIDATE_LOG"
return 1
}
# Apply a config change: validate, then REHASH; revert the file(s) on failure.
# $1 = description; $2.. = files that were changed (each has a .last-good copy).
apply_change() {
local what="$1"; shift
local f rc=0
if ! validate_or_fail "$what"; then
for f in "$@"; do [[ -f "$f.last-good" ]] && cat "$f.last-good" > "$f"; done
die "$what rejected; previous version restored."
fi
if ! ergo_running; then
log "$what saved (Ergo is not running; it will load on start)."
for f in "$@"; do rm -f "$f.last-good"; done
return 0
fi
# Ergo opens the history backend ONLY at startup. Unlike MySQL, sqlite and
# postgresql have no "after launching the server" guard, so a rehash that
# turns one on reports success and then silently discards every message.
# Catch that here rather than letting the operator believe it worked.
for f in "$@"; do
[[ "$f" == "${CONF:-}" && -f "$f.last-good" ]] || continue
if [[ "$(history_backend "$f")" != "$(history_backend "$f.last-good")" ]]; then
warn "This change switches the message-history backend, which Ergo only reads at startup."
warn "A rehash would report success and silently do nothing, so a restart is required."
if [[ "${ERGOCTL_ASSUME_YES:-0}" != 1 ]]; then
read -r -p "Restart Ergo now (disconnects every user)? [y/N] " ans
[[ "${ans,,}" == y* ]] || { cat "$f.last-good" > "$f"; die "Reverted; nothing changed. Use 'ergoctl history <backend>' to switch it properly."; }
fi
if ! validate_or_fail "$what"; then cat "$f.last-good" > "$f"; die "$what rejected; previous version restored."; fi
irc_raw --oper --quiet 1 -- 'NOTICE $$* :Server restarting to change message-history storage.' >/dev/null 2>&1 || true
dc up -d --no-deps ergo >/dev/null 2>&1 || dc restart ergo >/dev/null 2>&1 || true
if wait_healthy ergo 120 && irc_probe; then
rm -f "$f.last-good"; log "$what applied (Ergo restarted)."; return 0
fi
dc logs --tail 30 ergo >&2 || true
cat "$f.last-good" > "$f"
die "Ergo is not healthy after the restart; config reverted. Run 'ergoctl restart'."
fi
done
# Take the status from the call itself: `$?` after an `if` whose condition
# failed and that has no else branch is 0, which would make the SIGHUP
# fallback below unreachable.
irc_rehash || rc=$?
if (( rc == 0 )); then
log "$what applied (rehash complete)."
for f in "$@"; do rm -f "$f.last-good"; done
return 0
fi
# Ergo refuses to rehash settings that are fixed at startup. It phrases them
# all as "... after launching the server", so match that, not one variant.
if [[ "$REHASH_MSG" == *"after launching the server"* || "$REHASH_MSG" == *"cannot be changed after launching"* ]]; then
warn "Rehash refused: $REHASH_MSG"
warn "The file is valid and saved; apply it with: ergoctl restart (disconnects everyone)"
for f in "$@"; do rm -f "$f.last-good"; done
return 0
fi
if (( rc == 2 )); then
warn "Could not REHASH over IRC ($REHASH_MSG); sending SIGHUP instead."
dc kill -s HUP ergo >/dev/null 2>&1 || true
sleep 2
if dc logs --since 15s ergo 2>/dev/null | grep -qiE 'rehash.*(complete|success)'; then
log "$what applied (rehash via SIGHUP)."; for f in "$@"; do rm -f "$f.last-good"; done; return 0
fi
fi
for f in "$@"; do [[ -f "$f.last-good" ]] && cat "$f.last-good" > "$f"; done
die "Rehash failed: ${REHASH_MSG}. Previous version restored (Ergo kept running on the old config)."
}
snapshot_file() { cp -p "$1" "$1.last-good"; }
edit_file() { # <file> <what>
local f="$1" what="$2" tmp
require_root; require_stack
tmp="$(mktemp)"; cat "$f" > "$tmp"
"${EDITOR:-vi}" "$tmp"
if cmp -s "$f" "$tmp"; then rm -f "$tmp"; log "No changes."; return 0; fi
snapshot_file "$f"
cat "$tmp" > "$f"; rm -f "$tmp"
apply_change "$what" "$f"
}
oper_name_ok() { [[ "$1" =~ ^[a-z][a-z0-9_-]{1,31}$ ]]; }
cmd_status() {
require_stack
echo "Stack: $STACK_DIR (domain $(ergo_domain), network $(env_get NETWORK_NAME))"
dc ps 2>/dev/null || true
echo
local st cur latest
st="$(svc_state ergo || true)"; cur="$(ergo_version_running || true)"
latest="$(normver "$(ergo_latest_tag || true)")"
printf 'Ergo: %s | running %s | pinned %s | latest %s\n' "${st:-not created}" "${cur:-?}" "$(ergo_tag)" "${latest:-?}"
if [[ "$(env_get PLAINTEXT)" == "1" ]]; then echo " PUBLIC PLAINTEXT LISTENER on :6667 (PLAINTEXT=1)"; fi
local hb; hb="$(history_backend)"
case "$hb" in
postgres) pg_check || true; printf 'History: persistent, PostgreSQL -- %s\n' "$PG_CHECK_MSG" ;;
sqlite) printf 'History: persistent, SQLite (%s)\n' "$IRCD_DIR/$(yaml_get_nested "$CONF" datastore sqlite database-path)" ;;
*) printf 'History: RAM only -- lost on every restart\n' ;;
esac
cmd_cert show
if ergo_running; then
echo; cmd_users || true
fi
echo
printf 'Policy: %s | schedule: ' "${UPDATE_POLICY:-$(env_get UPDATE_POLICY)}"
if [[ -x /etc/periodic/15min/ergo-certsync || -f /etc/systemd/system/ergo-certsync.timer ]]; then printf 'certsync 15min'; else printf 'certsync MISSING'; fi
if [[ -x /etc/periodic/daily/ergo-update || -f /etc/systemd/system/ergo-update.timer ]]; then printf ', update daily\n'; else printf ', update NOT scheduled\n'; fi
local s; s="$(state_get certsync)"; [[ -z "$s" ]] || echo "certsync: LAST FAILURE: $s"
}
cmd_users() {
local out
out="$(irc_raw --quiet 1 -- LUSERS)" || die "Could not talk to Ergo on 127.0.0.1:6667 (is it running?)."
# ":server 251 nick :There are N users..." -> "There are N users..."
# 265/266 repeat their counts in the trailing text ("1 1 :Current local users
# 1, max 1"), so drop their leading parameters; 252/254 carry the only count
# there ("0 :IRC Operators online"), so keep it.
printf '%s\n' "$out" | awk '$2=="251"||$2=="252"||$2=="254"||$2=="265"||$2=="266" {
n=$2
sub(/^:[^ ]+ [0-9]+ [^ ]+ /, "")
if (n=="265" || n=="266") sub(/^[^:]*:/, "")
sub(/^:/, ""); sub(/ :/, " ")
print " " $0 }'
# LUSERS is answered while our own throwaway client is still registered, so
# every count above includes it (and Ergo's default +i makes it "invisible").
printf '%s\n' " (counts include ergoctl's own probe connection)"
}
cmd_logs() {
require_stack
local follow="" n=100
while [[ $# -gt 0 ]]; do
case "$1" in -f|--follow) follow="-f" ;; *[0-9]*) n="$1" ;; esac; shift
done
# Each healthcheck probe logs a 'Client connecting: real IP 127.0.0.1' line; hide those.
dc logs --tail "$n" $follow ergo 2>&1 | grep -vE 'real IP 127\.0\.0\.1, proxied IP <nil>|Preexisting TLS cert' || true
}
cmd_cert() {
require_stack
local f="$IRCD_DIR/fullchain.pem" served
case "${1:-show}" in
sync) require_root; bash "$STACK_DIR/update.sh" certsync ;;
*)
if [[ ! -f "$f" ]]; then echo "TLS: no cert in ircd/ yet"; return 0; fi
if cert_is_selfsigned "$f"; then
echo "TLS: SELF-SIGNED (Caddy has not issued a cert yet -- DNS/80/443?) expires $(cert_enddate "$f")"
else
echo "TLS: $(cert_subject_cn "$f") expires $(cert_enddate "$f") (from Caddy)"
fi
if ergo_running; then
served="$(served_fingerprint || true)"
if [[ -z "$served" ]]; then echo " :6697 did not complete a TLS handshake"
elif [[ "$served" == "$(cert_fingerprint "$f")" ]]; then echo " :6697 serves this cert"
else echo " :6697 serves a DIFFERENT cert -- run: ergoctl rehash"; fi
fi ;;
esac
}
cmd_oper() {
require_root; require_stack
local sub="${1:-list}" name="${2:-}" class="${3:-chat-moderator}" pw hash
case "$sub" in
list) yaml_oper_list "$CONF" | awk '{printf " %-20s %s\n", $1, $2}' \
| while read -r line; do
n="${line%% *}"; n="${n# }"
fp="$(yaml_oper_get_field "$CONF" "$n" certfp)"
au="$(yaml_oper_get_field "$CONF" "$n" auto)"
printf '%s%s%s\n' "$line" \
"$( [[ -n "$fp" ]] && printf ' certfp:%s…' "${fp:0:12}" )" \
"$( [[ "$au" == "true" ]] && printf ' auto' )"
done ;;
certfp)
# Authenticate an oper by TLS client certificate instead of (or as
# well as) a typed password.
[[ -n "$name" ]] || die "Usage: ergoctl oper certfp <name> <fingerprint|nick> [--auto] | ergoctl oper certfp <name> --clear"
yaml_oper_has "$CONF" "$name" || die "No oper '$name' in ircd.yaml (see: ergoctl oper list)."
# ergoctl itself opers as 'admin' over the loopback PLAINTEXT listener,
# which by definition presents no client certificate. Ergo requires a
# configured certfp to match, so this would lock the tooling -- and the
# scheduled jobs that use it -- out of the server for good.
[[ "$name" != "admin" ]] || die "Refusing to set a certfp on 'admin': ergoctl opers up over the loopback plaintext listener with no client certificate, so Ergo would reject it and every ergoctl command that needs oper would stop working. Make a personal oper instead: ergoctl oper add <you> server-admin && ergoctl oper certfp <you> <nick> --auto"
local arg="" auto=0 clear=0 a
for a in "${@:3}"; do
case "$a" in
--auto) auto=1 ;;
--clear) clear=1 ;;
--*) die "Unknown option '$a'." ;;
*) arg="$a" ;;
esac
done
snapshot_file "$CONF"
if (( clear )); then
yaml_oper_rm_field "$CONF" "$name" certfp || true
yaml_oper_rm_field "$CONF" "$name" auto || true
[[ -n "$(yaml_oper_get_field "$CONF" "$name" password)" ]] \
|| warn "'$name' now has neither a certfp nor a password; Ergo refuses to load an oper with no way to authenticate. Set one with: ergoctl oper passwd $name"
apply_change "certfp removal for '$name'" "$CONF"
return 0
fi
[[ -n "$arg" ]] || { rm -f "$CONF.last-good"; die "Give a SHA-256 fingerprint, or the nick of a connected user to read it from. Your own is on the 276 line of /WHOIS <yournick>, or: openssl x509 -noout -fingerprint -sha256 -in client.pem"; }
local fp=""
if fp="$(normalize_certfp "$arg")"; then
: # a fingerprint was given directly
else
log "Looking up the certificate fingerprint of '$arg' over IRC..."
local found; found="$(certfp_of_nick "$arg" || true)"
[[ -n "$found" ]] || { rm -f "$CONF.last-good"; die "No certificate fingerprint for '$arg'. They must be connected AND using a TLS client certificate. Have them check the 276 line of their own /WHOIS, then pass the fingerprint here directly."; }
fp="$(normalize_certfp "$found")" || { rm -f "$CONF.last-good"; die "The server returned '$found', which is not a SHA-256 fingerprint."; }
log "Found: $fp"
fi
yaml_oper_set_field "$CONF" "$name" certfp "\"$fp\"" \
|| { cat "$CONF.last-good" > "$CONF"; die "Could not write certfp into the '$name' oper block."; }
if (( auto )); then
# certfp + password means Ergo requires BOTH, so the password has
# to go for a no-typing login.
yaml_oper_rm_field "$CONF" "$name" password || true
yaml_oper_set_field "$CONF" "$name" auto true \
|| { cat "$CONF.last-good" > "$CONF"; die "Could not set auto for '$name'."; }
fi
apply_change "certfp for oper '$name'" "$CONF"
echo
if (( auto )); then
echo " '$name' is now granted operator status automatically on connect with that"
echo " certificate -- no /OPER, no password. The password has been removed."
warn "Anyone holding that client certificate is now an operator on sight. Keep the key safe, and 'ergoctl oper certfp $name --clear' revokes it."
else
echo " '$name' now requires BOTH that certificate and the password: /OPER $name <password>"
echo " For passwordless automatic oper instead: ergoctl oper certfp $name --auto"
fi ;;
add)
[[ -n "$name" ]] || die "Usage: ergoctl oper add <name> [chat-moderator|server-admin]"
oper_name_ok "$name" || die "Oper names: lowercase letters, digits, _ - (2-32 chars)."
yaml_oper_has "$CONF" "$name" && die "Oper '$name' already exists."
grep -qE "^ \"?${class}\"?:" "$CONF" || warn "Class '$class' is not defined under oper-classes in ircd.yaml -- the rehash will fail unless you add it."
pw="$(ergo_random_password 24)"
hash="$(ergo_genpasswd "$(ergo_image)" "$pw")" || die "genpasswd failed."
snapshot_file "$CONF"
yaml_oper_add "$CONF" "$name" "$class" "$hash" || { cat "$CONF.last-good" > "$CONF"; die "Could not insert the oper block."; }
apply_change "oper '$name'" "$CONF"
echo
echo " Oper '$name' ($class) added. Log in with: /OPER $name $pw"
echo " (shown once; rotate with: ergoctl oper passwd $name)" ;;
passwd)
[[ -n "$name" ]] || die "Usage: ergoctl oper passwd <name>"
yaml_oper_has "$CONF" "$name" || die "No oper '$name' in ircd.yaml."
pw="$(ergo_random_password 24)"
hash="$(ergo_genpasswd "$(ergo_image)" "$pw")" || die "genpasswd failed."
snapshot_file "$CONF"
yaml_oper_set_password "$CONF" "$name" "$hash" || { cat "$CONF.last-good" > "$CONF"; die "Oper '$name' has no password line to replace."; }
apply_change "oper '$name' password" "$CONF"
if [[ "$name" == "admin" ]]; then ( umask 077; printf '%s\n' "$pw" > "$ADMIN_PASS_FILE" ); fi
echo
echo " New password for '$name': /OPER $name $pw" ;;
rm|remove|del)
[[ -n "$name" ]] || die "Usage: ergoctl oper rm <name>"
[[ "$name" != "admin" ]] || die "Refusing to remove 'admin' (ergoctl uses it). Rotate its password instead."
yaml_oper_has "$CONF" "$name" || die "No oper '$name' in ircd.yaml."
snapshot_file "$CONF"
yaml_oper_rm "$CONF" "$name" || { cat "$CONF.last-good" > "$CONF"; die "Could not remove the oper block."; }
apply_change "removal of oper '$name'" "$CONF" ;;
*) die "Usage: ergoctl oper list|add|passwd|certfp|rm" ;;
esac
}
# A connected user's TLS certificate fingerprint, via 276 RPL_WHOISCERTFP.
# Empty when they are not connected or not using a client certificate.
certfp_of_nick() {
local out
out="$(irc_raw --oper --quiet 2 -- "WHOIS $1")" || return 1
printf '%s\n' "$out" | awk '$2=="276" { for (i=1;i<=NF;i++) if ($i ~ /^[0-9a-fA-F]{64}$/) { print $i; exit } }'
}
cmd_passwd() {
require_root; require_stack
case "${1:-show}" in
show) [[ -f "$ADMIN_PASS_FILE" ]] || die "No $ADMIN_PASS_FILE -- run: ergoctl passwd rotate"
echo "/OPER admin $(head -n1 "$ADMIN_PASS_FILE")" ;;
rotate) cmd_oper passwd admin
if irc_raw --oper --quiet 1 -- >/dev/null 2>&1; then log "Verified: the new admin password works."; else warn "Could not verify the new password over IRC (is Ergo running?)."; fi ;;
*) die "Usage: ergoctl passwd [show|rotate]" ;;
esac
}
cmd_cmd() { # raw line as oper
require_root; require_stack
[[ $# -gt 0 ]] || die "Usage: ergoctl cmd <raw IRC line>"
local out rc=0
out="$(irc_raw --oper --quiet 2 -- "$*")" || rc=$?
case "$rc" in
0) ;;
3) die "OPER failed (wrong admin password? see secrets/admin.pass; the server logs details under type 'opers')." ;;
4) die "Cannot connect to 127.0.0.1:6667 -- is Ergo running?" ;;
*) die "Registration with the server failed (rc=$rc)." ;;
esac
# Drop the routine registration burst and show only what the command produced.
# The MODE filter is anchored to our own throwaway nick so that MODE replies
# caused by the command itself (e.g. SAMODE #chan +m) still print.
printf '%s\n' "$out" | grep -vE '^:[^ ]+ (00[1-5]|25[0-9]|26[56]|37[256]|422|381) |^:[^ ]+ MODE ergoctl[0-9]+ ' || true
}
cmd_backup() {
require_root; require_stack
local live=0 dest="$BACKUP_DIR" ts stage out was_running=0 users=""
while [[ $# -gt 0 ]]; do case "$1" in --live) live=1 ;; *) dest="$1" ;; esac; shift; done
# `install -d -m` also re-modes an EXISTING directory, so only create-and-mode
# our own default; never change the mode of a directory the operator named
# (`ergoctl backup /var/backups` must not make it 0700 root).
if [[ "$dest" == "$BACKUP_DIR" ]]; then install -d -m 0700 "$dest"
else [[ -d "$dest" ]] || install -d -m 0700 "$dest"; fi
# A recipients file means "encrypt these" -- refuse rather than silently
# writing account hashes and the oper hash out in the clear.
if [[ -f "$STACK_DIR/age-recipients.txt" ]] && ! command -v age >/dev/null 2>&1; then
die "$STACK_DIR/age-recipients.txt exists but 'age' is not installed, so the backup would be UNENCRYPTED. Install it (apk add age / apt install age / dnf install age) or move the recipients file away."
fi
ts="$(date -u +%Y%m%d-%H%M%S)"
stage="$(mktemp -d)"
if ergo_running; then
was_running=1
if [[ "$live" == 0 ]]; then
users="$(irc_raw --quiet 1 -- LUSERS 2>/dev/null | awk '$2=="251" {print}' | sed 's/.*:There are //' || true)"
log "Stopping Ergo for a consistent snapshot (${users:-users unknown})..."
irc_raw --oper --quiet 1 -- 'NOTICE $$* :Brief restart for a backup -- back in a few seconds.' >/dev/null 2>&1 || true
dc stop -t 15 ergo >/dev/null 2>&1 || true
else
warn "Live backup: ircd.db is copied while Ergo writes to it (crash-consistent; may miss the last second)."
fi
fi
# -P everywhere: ircd/ is writable by the container uid, so a symlink there
# must be copied as a link, never followed by this root process.
refuse_symlink "$CONF" "ircd/ircd.yaml" || die "Refusing to back up through a symlink."
cp -Pp "$CONF" "$stage/ircd.yaml"
[[ -f "$IRCD_DIR/ircd.db" ]] && cp -Pp "$IRCD_DIR/ircd.db" "$stage/ircd.db"
[[ -f "$MOTD" ]] && cp -Pp "$MOTD" "$stage/ergo.motd"
# Message history, whichever backend holds it.
local backend hf
backend="$(history_backend)"
case "$backend" in
sqlite)
hf="$IRCD_DIR/$(yaml_get_nested "$CONF" datastore sqlite database-path)"
# SQLite writes -wal/-shm siblings; with Ergo stopped they are already
# checkpointed, and on a --live backup we take them along.
for f in "$hf" "$hf-wal" "$hf-shm"; do
[[ -f "$f" ]] && cp -Pp "$f" "$stage/$(basename "$f")"
done ;;
postgres)
if pg_check; then
log "Dumping the PostgreSQL history database..."
# A file copy of a live PGDATA is not a valid backup; pg_dump is.
dc exec -T postgres pg_dump -U "$(env_get POSTGRES_USER)" -d "$(env_get POSTGRES_DB)" \
> "$stage/history.sql" 2>/dev/null || { rm -f "$stage/history.sql"; warn "pg_dump failed; the backup will NOT contain message history."; }
else
warn "PostgreSQL is not reachable ($PG_CHECK_MSG); the backup will NOT contain message history."
fi ;;
esac
cat > "$stage/meta" <<EOF
ERGO_TAG=$(ergo_tag)
ERGO_DOMAIN=$(ergo_domain)
HISTORY=$backend
DATE=$(date -u +%FT%TZ)
HOST=$(hostname -f 2>/dev/null || hostname)
CONSISTENT=$(( live == 0 ))
EOF
if [[ "$was_running" == 1 && "$live" == 0 ]]; then dc start ergo >/dev/null 2>&1 || true; fi
out="$dest/ergo-backup-${ts}.tar.gz"
( umask 077; tar -czf "$out" -C "$stage" . )
rm -rf "$stage"
if [[ -f "$STACK_DIR/age-recipients.txt" ]]; then
age -R "$STACK_DIR/age-recipients.txt" -o "$out.age" "$out" || { rm -f "$out.age"; die "age encryption failed; the plaintext tar is still at $out."; }
rm -f "$out"; out="$out.age"
fi
chmod 0600 "$out"
log "Backup written: $out ($(wc -c < "$out") bytes). Copy it off this host."
[[ "$out" == *.age ]] || warn "Unencrypted: it holds account hashes and the oper hash. Put an age public key in $STACK_DIR/age-recipients.txt (and install 'age') to encrypt future backups."
}
cmd_restore() {
require_root; require_stack
local file="${1:-}" stage tag
[[ -f "$file" ]] || die "Usage: ergoctl restore <backup.tar.gz[.age]>"
stage="$(mktemp -d)"
if [[ "$file" == *.age ]]; then
command -v age >/dev/null 2>&1 || die "age is required to decrypt $file."
age -d -i "${AGE_IDENTITY:?set AGE_IDENTITY=/path/to/age/key}" "$file" | tar -xzf - -C "$stage"
else
tar -xzf "$file" -C "$stage"
fi
[[ -f "$stage/ircd.yaml" && -f "$stage/ircd.db" ]] || { rm -rf "$stage"; die "Backup lacks ircd.yaml/ircd.db."; }
tag="$(grep '^ERGO_TAG=' "$stage/meta" 2>/dev/null | cut -d= -f2 || true)"
if [[ -n "$tag" ]] && ver_gt "$(normver "$tag")" "$(normver "$(ergo_tag)")"; then
warn "Backup was taken with Ergo $tag but $(ergo_tag) is pinned -- an older Ergo cannot read a newer database schema. Update first (ergoctl update update) or pin ERGO_TAG=$tag."
read -r -p "Continue anyway? [y/N] " ans; [[ "${ans,,}" == y* ]] || { rm -rf "$stage"; die "Aborted."; }
fi
log "Stopping Ergo and restoring from $file..."
dc stop -t 15 ergo >/dev/null 2>&1 || true
install -d -m 0700 "$BACKUP_DIR"
# Keep the CURRENT config and database, not just the database: if the restore
# turns out to be unloadable there has to be a way back.
local ts uid gid
ts="$(date -u +%Y%m%d%H%M%S)"; uid="$(ergo_uid)"; gid="$(ergo_gid)"
[[ -f "$CONF" ]] && cp -Pp "$CONF" "$BACKUP_DIR/ircd.yaml.pre-restore.$ts"
[[ -f "$IRCD_DIR/ircd.db" ]] && cp -Pp "$IRCD_DIR/ircd.db" "$BACKUP_DIR/ircd.db.pre-restore.$ts"
safe_replace "$stage/ircd.yaml" "$CONF" 0600 "$uid" "$gid"
safe_replace "$stage/ircd.db" "$IRCD_DIR/ircd.db" 0600 "$uid" "$gid"
[[ -f "$stage/ergo.motd" ]] && safe_replace "$stage/ergo.motd" "$MOTD" 0644 "$uid" "$gid"
# The backup carries the ORIGINAL host's credentials. Reconcile them with
# this host's, or the restored server is unreachable and unadministrable.
if [[ "$(history_backend)" == "postgres" && -f "$SECRETS_DIR/postgres.pass" ]]; then
log "Re-pointing datastore.postgresql at this host's database password..."
yaml_set_nested "$CONF" datastore postgresql password "\"$(cat "$SECRETS_DIR/postgres.pass")\"" \
|| warn "Could not update the postgresql password; Ergo may fail to reach its history database."
fi
if [[ -f "$ADMIN_PASS_FILE" ]]; then
local h
if h="$(ergo_genpasswd "$(ergo_image)" "$(head -n1 "$ADMIN_PASS_FILE")")"; then
yaml_oper_set_password "$CONF" admin "$h" \
&& log "Re-hashed this host's admin password into the restored config (ergoctl keeps working)." \
|| warn "Could not update the admin oper hash; 'ergoctl passwd rotate' will fix it."
else
warn "Could not hash the local admin password; the restored config keeps the backup's admin hash. If OPER fails, run 'ergoctl passwd rotate'."
fi
fi
# Message history
local f
for f in "$stage"/ergo_history.db*; do
[[ -f "$f" ]] || continue
install -m 0600 -o "$(ergo_uid)" -g "$(ergo_gid)" "$f" "$IRCD_DIR/$(basename "$f")"
done
if [[ -f "$stage/history.sql" ]]; then
log "Starting PostgreSQL and restoring the history dump..."
dc up -d --no-deps postgres >/dev/null 2>&1 || true
if wait_healthy postgres 120; then
dc exec -T postgres psql -U "$(env_get POSTGRES_USER)" -d "$(env_get POSTGRES_DB)" < "$stage/history.sql" >/dev/null 2>&1 \
|| warn "Restoring the history dump failed; Ergo will start with an empty history."
else
warn "PostgreSQL did not become healthy; message history was NOT restored."
fi
fi
rm -rf "$stage"
# Validate before starting, so a bad backup does not cost an outage.
if ! validate_or_fail "the restored config"; then
[[ -f "$BACKUP_DIR/ircd.yaml.pre-restore.$ts" ]] && cp -p "$BACKUP_DIR/ircd.yaml.pre-restore.$ts" "$CONF"
[[ -f "$BACKUP_DIR/ircd.db.pre-restore.$ts" ]] && cp -p "$BACKUP_DIR/ircd.db.pre-restore.$ts" "$IRCD_DIR/ircd.db"
dc up -d --no-deps ergo >/dev/null 2>&1 || true
die "The backup's config does not load; rolled back to what was here before (copies kept as $BACKUP_DIR/*.pre-restore.$ts)."
fi
dc up -d --no-deps ergo >/dev/null 2>&1 || true
if wait_healthy ergo 120 && irc_probe; then
log "Restored; Ergo is healthy."
log "Previous state kept at $BACKUP_DIR/{ircd.yaml,ircd.db}.pre-restore.$ts"
else
dc logs --tail 20 ergo >&2 || true
die "Ergo is not healthy after the restore. Roll back with: cp $BACKUP_DIR/ircd.yaml.pre-restore.$ts $CONF && cp $BACKUP_DIR/ircd.db.pre-restore.$ts $IRCD_DIR/ircd.db && ergoctl start"
fi
}
cmd_history() {
require_stack
local want="${1:-}" cur pw
cur="$(history_backend)"
if [[ -z "$want" ]]; then
echo "History backend: $cur"
case "$cur" in
sqlite) local f="$IRCD_DIR/$(yaml_get_nested "$CONF" datastore sqlite database-path)"
[[ -f "$f" ]] && echo " $f ($(wc -c < "$f") bytes)" || echo " (database not created yet)" ;;
postgres) pg_check && echo " $PG_CHECK_MSG" || echo " $PG_CHECK_MSG" ;;
off) echo " messages are kept in RAM only and lost on every restart" ;;
esac
echo " retention: history.restrictions.expire-time = $(yaml_get_nested "$CONF" history restrictions expire-time)"
return 0
fi
require_root
case "$want" in sqlite|postgres|off) ;; *) die "Usage: ergoctl history [sqlite|postgres|off]" ;; esac
[[ "$want" != "$cur" ]] || { log "History is already using '$cur'."; return 0; }
# .env and ircd.yaml must never disagree: deploy.sh reconciles them from .env
# on its next run, so a half-applied switch would let it tear down the wrong
# thing. Remember both and put them back on every failure path.
local prev_hist prev_compose
prev_hist="$(env_get HISTORY)"; prev_compose="$(env_get COMPOSE_FILE)"
_hist_revert() {
[[ -f "$CONF.last-good" ]] && cat "$CONF.last-good" > "$CONF"
[[ -n "$prev_hist" ]] && env_set HISTORY "$prev_hist"
[[ -n "$prev_compose" ]] && env_set COMPOSE_FILE "$prev_compose"
return 0
}
if [[ "$want" == "postgres" ]]; then
[[ -f "$STACK_DIR/docker-compose.postgres.yml" ]] || die "docker-compose.postgres.yml is missing; re-run deploy.sh to install it."
if [[ ! -f "$SECRETS_DIR/postgres.pass" ]]; then
log "Generating the PostgreSQL password..."
( umask 077; ergo_random_password 32 > "$SECRETS_DIR/postgres.pass" )
fi
chmod 0600 "$SECRETS_DIR/postgres.pass"; pw="$(cat "$SECRETS_DIR/postgres.pass")"
[[ -n "$(env_get POSTGRES_TAG)" ]] || env_set POSTGRES_TAG 17-alpine
[[ -n "$(env_get POSTGRES_USER)" ]] || env_set POSTGRES_USER ergo
[[ -n "$(env_get POSTGRES_DB)" ]] || env_set POSTGRES_DB ergo_history
[[ -n "$(env_get POSTGRES_PORT)" ]] || env_set POSTGRES_PORT 5432
env_set COMPOSE_FILE "docker-compose.yml:docker-compose.postgres.yml"
fi
warn "Switching the history backend does NOT migrate existing messages; the old store is left in place."
snapshot_file "$CONF"
# Turn every backend off first, then enable the requested one.
yaml_set_nested "$CONF" datastore sqlite enabled false || true
yaml_set_nested "$CONF" datastore postgresql enabled false || true
case "$want" in
sqlite) yaml_set_nested "$CONF" datastore sqlite enabled true || { cat "$CONF.last-good" > "$CONF"; die "Could not enable datastore.sqlite."; }
yaml_set_nested "$CONF" history persistent enabled true || true ;;
postgres) yaml_set_nested "$CONF" datastore postgresql enabled true || { cat "$CONF.last-good" > "$CONF"; die "Could not enable datastore.postgresql."; }
yaml_set_nested "$CONF" datastore postgresql host "\"127.0.0.1\"" || true
yaml_set_nested "$CONF" datastore postgresql port "$(env_get POSTGRES_PORT)" || true
yaml_set_nested "$CONF" datastore postgresql user "\"$(env_get POSTGRES_USER)\"" || true
yaml_set_nested "$CONF" datastore postgresql password "\"${pw}\"" || true
yaml_set_nested "$CONF" datastore postgresql history-database "\"$(env_get POSTGRES_DB)\"" || true
yaml_set_nested "$CONF" history persistent enabled true || true
log "Starting PostgreSQL..."
dc up -d --no-deps postgres >/dev/null 2>&1 || true
wait_healthy postgres 120 || { _hist_revert; die "PostgreSQL did not become healthy; config reverted."; } ;;
off) yaml_set_nested "$CONF" history persistent enabled false || true ;;
esac
if ! validate_or_fail "history backend $want"; then _hist_revert; die "Config rejected; reverted."; fi
log "Restarting Ergo (the history backend is fixed at startup; every user is disconnected)..."
irc_raw --oper --quiet 1 -- 'NOTICE $$* :Server restarting to change message-history storage.' >/dev/null 2>&1 || true
dc up -d --no-deps ergo >/dev/null 2>&1 || dc restart ergo >/dev/null 2>&1 || true
if wait_healthy ergo 120 && irc_probe; then
# Only now is the switch real, so only now does .env change.
env_set HISTORY "$want"
[[ "$want" == "postgres" ]] || env_set COMPOSE_FILE "docker-compose.yml"
rm -f "$CONF.last-good"
log "History backend is now '$want'."
if [[ "$want" == "off" ]]; then
warn "Existing persisted messages are still on disk; remove them yourself if that is the point."
fi
else
dc logs --tail 30 ergo >&2 || true
_hist_revert
dc up -d --no-deps ergo >/dev/null 2>&1 || true
die "Ergo is not healthy after the switch; config and .env reverted. Check 'ergoctl logs'."
fi
return 0 # never let a trailing false test become this command's exit status
}
cmd_debug() {
require_root; require_stack
local level
case "${1:-}" in on) level=debug ;; off) level=info ;; *) die "Usage: ergoctl debug on|off" ;; esac
snapshot_file "$CONF"
yaml_set_log_level "$CONF" "$level" || { cat "$CONF.last-good" > "$CONF"; die "No 'level:' line found in the logging block."; }
apply_change "log level $level" "$CONF"
}
case "${1:-help}" in
help|-h|--help) usage ;;
status) cmd_status ;;
users) require_stack; cmd_users ;;
logs) shift; cmd_logs "$@" ;;
version) require_stack; printf 'running %s | pinned %s | latest %s\n' "$(ergo_version_running)" "$(ergo_tag)" "$(ergo_latest_tag || echo '?')" ;;
cert) shift; cmd_cert "$@" ;;
edit) edit_file "$CONF" "ircd.yaml" ;;
motd) edit_file "$MOTD" "MOTD" ;;
rehash) require_root; require_stack
if irc_rehash; then log "Rehash complete."; else
if [[ "$REHASH_MSG" == *"could not reach"* ]]; then warn "$REHASH_MSG -- sending SIGHUP."; dc kill -s HUP ergo; else die "Rehash failed: $REHASH_MSG"; fi
fi ;;
restart) require_root; require_stack; irc_raw --oper --quiet 1 -- 'NOTICE $$* :Server restarting now.' >/dev/null 2>&1 || true
dc restart -t 15 ergo; wait_healthy ergo 120 && log "Ergo is healthy." || die "Ergo did not become healthy (ergoctl logs)." ;;
stop) require_root; require_stack; dc stop -t 15 ergo ;;
start) require_root; require_stack; dc start ergo; wait_healthy ergo 120 && log "Ergo is healthy." || die "Ergo did not become healthy (ergoctl logs)." ;;
caddy-restart) require_root; require_stack
dc run --rm --no-deps -T caddy caddy validate --config /etc/caddy/Caddyfile --adapter caddyfile >/dev/null || die "Caddyfile invalid (see caddy/etc/Caddyfile, conf.d/)."
dc restart caddy; wait_healthy caddy 60 && log "Caddy is healthy." || die "Caddy did not become healthy." ;;
debug) shift; cmd_debug "$@" ;;
history) shift; cmd_history "$@" ;;
oper) shift; cmd_oper "$@" ;;
passwd) shift; cmd_passwd "$@" ;;
cmd) shift; cmd_cmd "$@" ;;
announce) shift; [[ $# -gt 0 ]] || die "Usage: ergoctl announce <text>"; cmd_cmd "NOTICE \$\$* :$*" ;;
kill) shift; [[ $# -ge 1 ]] || die "Usage: ergoctl kill <nick> [reason]"; n="$1"; shift; cmd_cmd "KILL $n :${*:-Killed by an operator}" ;;
ban) shift; sub="${1:-list}"; shift || true
case "$sub" in
add) [[ $# -ge 1 ]] || die "Usage: ergoctl ban add <ip|cidr|nick!user@host|account> [duration e.g. 1d] [reason] -- a bare name means an ACCOUNT; use 'ban info <nick>' for a connected user's IP"
t="$1"; shift; d=""; if [[ "${1:-}" =~ ^[0-9]+[ymwdhs]$|^[0-9]+mo$ ]]; then d="DURATION $1"; shift; fi
cmd_cmd "UBAN ADD $t $d ${*:+:$*}" ;;
del) [[ $# -ge 1 ]] || die "Usage: ergoctl ban del <target>"; cmd_cmd "UBAN DEL $1" ;;
info) [[ $# -ge 1 ]] || die "Usage: ergoctl ban info <target>"; cmd_cmd "UBAN INFO $1" ;;
list|*) cmd_cmd "UBAN LIST" ;;
esac ;;
defcon) shift; cmd_cmd "DEFCON ${1:-}" ;;
backup) shift; cmd_backup "$@" ;;
restore) shift; cmd_restore "$@" ;;
update) shift; require_root; require_stack; exec bash "$STACK_DIR/update.sh" "${1:-check}" ;;
shell) require_root; require_stack; dc exec ergo sh ;;
*) usage; exit 1 ;;
esac
+744
View File
@@ -0,0 +1,744 @@
#!/usr/bin/env bash
#
# ergolib.sh -- shared helpers for the Ergo stack. Sourced by deploy.sh (from the
# extracted archive), update.sh and ergoctl (from $STACK_DIR). Not run directly.
#
# Everything that touches the running stack lives here so the three scripts stay
# consistent: .env access, the compose wrapper, health + IRC probes over the
# loopback listener, the throwaway-container config validator, Caddy->Ergo cert
# sync, oper edits in ircd.yaml, and ntfy notifications.
#
# Conventions: functions return non-zero instead of exiting (callers decide),
# and everything is portable to busybox (Alpine) -- POSIX awk, no GNU date -d,
# no sort -V, no grep -P.
# ---------------------------------------------------------------------------
# Logging (reuse the caller's if defined)
# ---------------------------------------------------------------------------
if ! declare -f log >/dev/null 2>&1; then
log() { printf '\033[1;32m[+]\033[0m %s\n' "$*"; }
warn() { printf '\033[1;33m[!]\033[0m %s\n' "$*" >&2; }
die() { printf '\033[1;31m[x]\033[0m %s\n' "$*" >&2; exit 1; }
fi
# ---------------------------------------------------------------------------
# Config file (/etc/ergo-update.conf): KEY=VALUE defaults for the scheduled
# jobs. Environment wins over the file, the file over built-ins.
# ---------------------------------------------------------------------------
: "${ERGO_UPDATE_CONF:=/etc/ergo-update.conf}"
load_conf() {
[[ -r "$ERGO_UPDATE_CONF" ]] || return 0
local _line _k _v
while IFS= read -r _line || [[ -n "$_line" ]]; do
[[ "$_line" =~ ^[[:space:]]*# || -z "${_line//[[:space:]]/}" ]] && continue
_k="${_line%%=*}"; _v="${_line#*=}"; _k="${_k//[[:space:]]/}"
[[ "$_k" =~ ^[A-Za-z_][A-Za-z0-9_]*$ ]] || continue
[[ -n "${!_k:-}" ]] && continue
_v="${_v%\"}"; _v="${_v#\"}"
printf -v "$_k" '%s' "$_v"
done < "$ERGO_UPDATE_CONF"
return 0
}
# Derived paths. Call after STACK_DIR is final.
ergo_set_paths() {
: "${STACK_DIR:=/srv/ergo}"
ENV_FILE="$STACK_DIR/.env"
IRCD_DIR="$STACK_DIR/ircd"
CADDY_DIR="$STACK_DIR/caddy"
SECRETS_DIR="$STACK_DIR/secrets"
BACKUP_DIR="$STACK_DIR/backups"
TEMPLATES_DIR="$STACK_DIR/templates"
STATE_DIR="$STACK_DIR/.state"
ADMIN_PASS_FILE="$SECRETS_DIR/admin.pass"
: "${GH_REPO:=ergochat/ergo}"
: "${NOTIFY:=1}"
: "${SSH_NOTIFY_CONF:=/etc/ssh-notify.conf}"
: "${DRY_RUN:=0}"
}
osfam() {
local id="" like=""
if [[ -r /etc/os-release ]]; then
id="$(. /etc/os-release 2>/dev/null && echo "${ID:-}")"
like="$(. /etc/os-release 2>/dev/null && echo "${ID_LIKE:-}")"
fi
case " $id $like " in
*" alpine "*) echo alpine ;;
*" debian "*|*" ubuntu "*) echo debian ;;
*" rhel "*|*" fedora "*|*" centos "*) echo rhel ;;
*) echo "${id:-unknown}" ;;
esac
}
# fetch a URL to stdout (curl or wget); rc 1 on failure.
fetch() {
if command -v curl >/dev/null 2>&1; then curl -fsSL -m 20 "$1" 2>/dev/null
elif command -v wget >/dev/null 2>&1; then wget -qO- -T 20 "$1" 2>/dev/null
else return 1
fi
}
# ---------------------------------------------------------------------------
# Writing into ircd/ safely
#
# ircd/ is owned by the unprivileged container uid, but several things that
# write there run as ROOT (the 15-minute cert sync, every ergoctl config edit).
# A plain `cp`, `install` or `>` follows a symlink at the destination, so code
# execution inside Ergo could point one of those writes at any root-owned file
# and win the host. rename(2) is the one operation that REPLACES a symlink
# instead of following it, so every root write into ircd/ goes: create a temp
# under $STACK_DIR (root-owned 0700, same filesystem) -> set mode/owner -> mv.
# ---------------------------------------------------------------------------
# A temp file that an unprivileged uid cannot reach or pre-empt.
stage_tmp() { mktemp "$STACK_DIR/.stage.XXXXXX"; }
# safe_replace <tmpfile> <target> [mode] [uid] [gid]
safe_replace() {
local tmp="$1" target="$2" mode="${3:-0600}" uid="${4:-}" gid="${5:-}"
chmod "$mode" "$tmp" 2>/dev/null || true
[[ -n "$uid" ]] && chown "${uid}:${gid:-$uid}" "$tmp" 2>/dev/null
mv -f "$tmp" "$target"
}
# Refuse to read a path under ircd/ that the container uid has turned into a
# symlink (a root reader would otherwise copy out whatever it points at).
refuse_symlink() { # <path> [what]
[[ -L "$1" ]] || return 0
warn "${2:-$1} is a symlink; refusing to touch it as root (see README: Security model)."
return 1
}
# ---------------------------------------------------------------------------
# .env access
# ---------------------------------------------------------------------------
env_get() { [[ -f "$ENV_FILE" ]] && grep -E "^$1=" "$ENV_FILE" | head -n1 | cut -d= -f2- || true; }
env_set() { # <KEY> <value>: update KEY in .env, or append if absent
# The value goes through the ENVIRONMENT, never interpolated into a sed
# script. Interpolating it corrupts any value containing & (sed expands it to
# the whole match) and aborts the run on one containing the s||| delimiter --
# which is reachable for an OIDC secret, a password or a URL query string.
local key="$1" val="$2" tmp
if [[ ! -f "$ENV_FILE" ]]; then
printf '%s=%s
' "$key" "$val" >> "$ENV_FILE"
return 0
fi
tmp="$(mktemp)"
_SE_KEY="$key" _SE_VAL="$val" awk '
BEGIN { k = ENVIRON["_SE_KEY"]; v = ENVIRON["_SE_VAL"]; seen = 0 }
!seen && index($0, k "=") == 1 { print k "=" v; seen = 1; next }
{ print }
END { if (!seen) print k "=" v }
' "$ENV_FILE" > "$tmp"
cat "$tmp" > "$ENV_FILE" # rewrite in place: keeps the original mode/owner
rm -f "$tmp"
}
ergo_domain() { env_get ERGO_DOMAIN | tr 'A-Z' 'a-z' || true; }
# The compose invocation must include the postgres overlay whenever history runs
# on PostgreSQL. deploy.sh writes COMPOSE_FILE into .env, which docker compose
# reads by itself -- these helpers are for code that needs to know.
pg_enabled() { [[ -f "$STACK_DIR/docker-compose.postgres.yml" ]] && grep -q '^COMPOSE_FILE=.*postgres' "$ENV_FILE" 2>/dev/null; }
# Is PostgreSQL up and accepting our credentials? rc 0 ok, 1 not.
PG_CHECK_MSG=""
pg_check() {
PG_CHECK_MSG=""
if ! pg_enabled; then PG_CHECK_MSG="postgres not in use"; return 0; fi
if [[ "$(svc_state postgres)" != running* ]]; then PG_CHECK_MSG="the postgres container is not running"; return 1; fi
local db user
db="$(env_get POSTGRES_DB)"; user="$(env_get POSTGRES_USER)"
if dc exec -T postgres pg_isready -U "${user:-ergo}" -d "${db:-ergo_history}" >/dev/null 2>&1; then
PG_CHECK_MSG="postgres ready (${db:-ergo_history})"; return 0
fi
PG_CHECK_MSG="postgres is running but not accepting connections yet"
return 1
}
# Fill the named variables from .env when they are not already set in the
# environment. Used by update.sh/ergoctl so the knobs .env documents really are
# the ones the scheduled jobs obey (environment > /etc/ergo-update.conf > .env).
env_defaults() { # env_defaults <VAR>...
local k v
for k in "$@"; do
[[ -n "${!k:-}" ]] && continue
v="$(env_get "$k")"
[[ -n "$v" ]] && printf -v "$k" '%s' "$v"
done
return 0
}
ergo_uid() { local v; v="$(env_get ERGO_UID)"; printf '%s' "${v:-1000}"; }
ergo_gid() { local v; v="$(env_get ERGO_GID)"; printf '%s' "${v:-1000}"; }
ergo_tag() { local v; v="$(env_get ERGO_TAG)"; printf '%s' "${v:-stable}"; }
ergo_image() { local i; i="$(env_get ERGO_IMAGE)"; printf '%s:%s' "${i:-ghcr.io/ergochat/ergo}" "$(ergo_tag)"; }
# docker compose, scoped to the stack dir (so ./ircd, ./caddy resolve and .env
# is auto-loaded).
dc() { ( cd "$STACK_DIR" && docker compose "$@" ); }
# ---------------------------------------------------------------------------
# Version helpers (pure bash; busybox sort has no -V)
# ---------------------------------------------------------------------------
# NOTE the trailing `|| true` here and in the cert/state helpers below: these are
# assigned in `set -euo pipefail` scripts, where a no-match grep would abort the
# caller instead of yielding "". Callers test for an empty result.
normver() { printf '%s' "$1" | grep -oE '[0-9]+\.[0-9]+(\.[0-9]+)?' | head -n1 || true; }
ver_cmp() { # <a> <b> -> prints -1 / 0 / 1
local a b i x y
IFS=. read -r -a a <<< "$(normver "$1")"
IFS=. read -r -a b <<< "$(normver "$2")"
for i in 0 1 2; do
x="${a[$i]:-0}"; y="${b[$i]:-0}"
if (( x > y )); then echo 1; return; fi
if (( x < y )); then echo -1; return; fi
done
echo 0
}
ver_gt() { [[ "$(ver_cmp "$1" "$2")" == "1" ]]; }
ver_ge() { [[ "$(ver_cmp "$1" "$2")" != "-1" ]]; }
ver_le() { [[ "$(ver_cmp "$1" "$2")" != "1" ]]; }
# Running version as reported by the binary ('ergo-v2.19.1' or 'ergo-2.19.1[-hash]');
# falls back to the pinned tag when the container is down.
ergo_version_running() {
local v=""
if ergo_running; then
v="$(dc exec -T ergo /ircd-bin/ergo --version 2>/dev/null | grep -oE '[0-9]+\.[0-9]+\.[0-9]+' | head -n1 || true)"
fi
[[ -n "$v" ]] || v="$(normver "$(ergo_tag)")"
printf '%s' "$v"
}
# Newest non-prerelease tag from GitHub (e.g. v2.19.1). rc 1 if unreachable.
ergo_latest_tag() {
local t
t="$(fetch "https://api.github.com/repos/${GH_REPO}/releases/latest" \
| grep -oE '"tag_name"[[:space:]]*:[[:space:]]*"[^"]+"' | head -n1 \
| sed -E 's/.*"tag_name"[[:space:]]*:[[:space:]]*"([^"]+)".*/\1/')"
[[ -n "$t" ]] || return 1
printf '%s' "$t"
}
# ---------------------------------------------------------------------------
# Container state
# ---------------------------------------------------------------------------
# "<state> <health>" for one compose service, e.g. "running healthy", "exited ".
# Empty when the container does not exist or the docker daemon is unreachable.
svc_state() {
dc ps -a --format '{{.Service}} {{.State}} {{.Health}}' 2>/dev/null | awk -v s="$1" '$1==s {print $2" "$3; exit}' || true
}
ergo_running() { [[ "$(svc_state ergo)" == running* ]]; }
# Wait for a service to be "running healthy"; rc 1 on timeout. Exited/restarting
# containers never count (docker compose ps hides them without -a).
wait_healthy() { # <service> [timeout-seconds]
local svc="$1" deadline st
deadline=$(( $(date +%s) + ${2:-120} ))
while (( $(date +%s) < deadline )); do
st="$(svc_state "$svc")"
[[ "$st" == "running healthy" ]] && return 0
sleep 3
done
return 1
}
# ---------------------------------------------------------------------------
# Talking IRC to the loopback listener (127.0.0.1:6667) with bash /dev/tcp.
#
# irc_raw [--oper] [--quiet SECS] -- <line>...
#
# Registers a throwaway nick, optionally opers up as 'admin' (password from
# $ADMIN_PASS_FILE or $ERGO_OPER_PASS), sends each line, prints the server's
# replies (control characters stripped) until the server has been quiet for
# --quiet seconds, then QUITs. IRC_LAST holds the raw transcript, but only when
# irc_raw is called directly -- inside $(...) it stays in the subshell, so parse
# the printed output instead.
#
# rc: 0 ok | 3 OPER refused | 4 connect failed | 5 registration failed
# Loopback is exempt from ip-limits/bans, but fakelag applies until opered:
# NICK+USER+OPER is 3 of the 5-line burst, so keep pre-oper payloads short.
# ---------------------------------------------------------------------------
IRC_LAST=""
# Strip CR/LF/NUL before framing: an argument containing a newline (e.g.
# `ergoctl announce "$(cat file)"`) would otherwise terminate this line and make
# the rest a second command, executed with the admin oper's full privileges.
_irc_send() {
local l="${1//$'\r'/}"; l="${l//$'\n'/ }"; l="${l//$'\0'/}"
printf '%s\r\n' "$l" >&3 2>/dev/null
}
_irc_cmd() { # echo the command/numeric token of a server line
local l="$1"
if [[ "$l" == :* ]]; then l="${l#* }"; fi
printf '%s' "${l%% *}"
}
irc_raw() {
local oper=0 quiet=2 line cmd nick tries=0 pass="" registered=0 opered=0 out=""
while [[ $# -gt 0 ]]; do
case "$1" in
--oper) oper=1; shift ;;
--quiet) quiet="$2"; shift 2 ;;
--) shift; break ;;
*) break ;;
esac
done
IRC_LAST=""
if [[ "$oper" == 1 ]]; then
pass="${ERGO_OPER_PASS:-}"
[[ -n "$pass" ]] || pass="$(head -n1 "$ADMIN_PASS_FILE" 2>/dev/null || true)"
[[ -n "$pass" ]] || { warn "No admin oper password (expected $ADMIN_PASS_FILE)."; return 3; }
fi
if ! { exec 3<>/dev/tcp/127.0.0.1/6667; } 2>/dev/null; then
return 4
fi
nick="ergoctl$RANDOM"
_irc_send "NICK $nick"
_irc_send "USER ergoctl 0 * :ergoctl"
while IFS= read -r -t 15 -u 3 line; do
line="${line%$'\r'}"
IRC_LAST+="$line"$'\n'
cmd="$(_irc_cmd "$line")"
case "$cmd" in
PING) _irc_send "PONG ${line#PING }" ;;
376|422) registered=1; break ;;
433) # nick in use -- pick another (up to 3 times)
tries=$((tries + 1)); (( tries > 3 )) && break
nick="ergoctl$RANDOM"; _irc_send "NICK $nick" ;;
432|464|465|ERROR|FAIL) break ;;
esac
done
if [[ "$registered" != 1 ]]; then
exec 3>&- 2>/dev/null; return 5
fi
if [[ "$oper" == 1 ]]; then
_irc_send "OPER admin $pass"
while IFS= read -r -t 15 -u 3 line; do
line="${line%$'\r'}"
IRC_LAST+="$line"$'\n'
cmd="$(_irc_cmd "$line")"
case "$cmd" in
PING) _irc_send "PONG ${line#PING }" ;;
381) opered=1; break ;;
491|464|ERROR) break ;;
esac
done
if [[ "$opered" != 1 ]]; then
_irc_send "QUIT :ergoctl"; exec 3>&- 2>/dev/null; return 3
fi
fi
for line in "$@"; do _irc_send "$line"; done
while IFS= read -r -t "$quiet" -u 3 line; do
line="${line%$'\r'}"
IRC_LAST+="$line"$'\n'
cmd="$(_irc_cmd "$line")"
case "$cmd" in
PING) _irc_send "PONG ${line#PING }"; continue ;;
ERROR) out+="$line"$'\n'; break ;;
esac
out+="$line"$'\n'
done
_irc_send "QUIT :ergoctl"
exec 3>&- 2>/dev/null
# Never let user-controlled IRC text drive the terminal: drop control chars
# (this also strips IRC formatting bytes 0x02/0x03/0x1d-0x1f).
[[ -n "$out" ]] && printf '%s' "$out" | tr -d '\000-\010\013\014\016-\037\177'
return 0
}
# Registration-only probe (no oper): proves the config loaded and the server
# completes a handshake. rc 0/1.
irc_probe() { irc_raw --quiet 1 -- >/dev/null 2>&1; }
# REHASH as the admin oper; sets REHASH_MSG. rc 0 = "Rehash complete",
# 1 = server refused (message in REHASH_MSG), 2 = could not connect/oper.
REHASH_MSG=""
irc_rehash() {
local out rc
REHASH_MSG=""
out="$(irc_raw --oper --quiet 3 -- REHASH)"; rc=$?
if (( rc != 0 )); then
REHASH_MSG="could not reach Ergo as oper (rc=$rc)"; return 2
fi
# $out holds every reply after the REHASH was sent (the $() subshell keeps
# IRC_LAST from us, so parse the printed transcript).
if printf '%s' "$out" | grep -qi 'Rehash complete'; then return 0; fi
REHASH_MSG="$(printf '%s\n' "$out" | awk '$2=="400" {sub(/^[^ ]+ 400 [^ ]+ REHASH :?/,""); print; exit}')"
[[ -n "$REHASH_MSG" ]] || REHASH_MSG="no confirmation from the server: $(printf '%s' "$out" | tail -n 3 | tr '\n' ' ')"
return 1
}
# ---------------------------------------------------------------------------
# Helper containers: never on a network, never privileged, run as the ergo uid.
# ---------------------------------------------------------------------------
ergo_helper() { # ergo_helper [docker-run-opts...] -- <image> <args...>
local opts=()
while [[ $# -gt 0 && "$1" != "--" ]]; do opts+=("$1"); shift; done
[[ "${1:-}" == "--" ]] && shift
docker run --rm --network none --cap-drop ALL --security-opt no-new-privileges \
--user "$(ergo_uid):$(ergo_gid)" "${opts[@]}" "$@"
}
# The image's own default config (already patched by Ergo's Dockerfile: public
# ":6667", no [::1]). Writes to <dest>; rc 1 if nothing came out.
ergo_extract_template() { # <image> <dest>
local image="$1" dest="$2"
ergo_helper --entrypoint cat -- "$image" /ircd-bin/default.yaml > "$dest" 2>/dev/null || true
if ! grep -q '^server:' "$dest" 2>/dev/null; then
ergo_helper --entrypoint /ircd-bin/ergo -- "$image" defaultconfig > "$dest" 2>/dev/null || true
fi
grep -q '^server:' "$dest" 2>/dev/null
}
# bcrypt hash for a password (stdin -> stdout, one line). rc 1 if it doesn't
# look like a bcrypt hash.
ergo_genpasswd() { # <image> <password>
local h
h="$(printf '%s\n' "$2" | ergo_helper -i --entrypoint /ircd-bin/ergo -- "$1" genpasswd 2>/dev/null | tr -d '\r\n')"
[[ "$h" == \$2a\$* || "$h" == \$2b\$* ]] || return 1
printf '%s' "$h"
}
# Random password from [A-Za-z0-9_-]: never trips Ergo's passphrase checks.
ergo_random_password() { head -c 256 /dev/urandom | tr -dc '_A-Za-z0-9-' | head -c "${1:-24}"; }
# ---------------------------------------------------------------------------
# Config validation: `ergo run --smoke` in a throwaway container against a COPY
# of ./ircd. Its own network namespace means the wildcard listeners bind fine,
# the copied ircd.db/lock are disposable, and a pending schema upgrade is
# dry-run on the copy. rc 0 = config (and certs) load; output in $VALIDATE_LOG.
# Ergo has no offline config checker and --smoke on the live dir would fight
# the running server for the flock and the ports.
# ---------------------------------------------------------------------------
VALIDATE_LOG=""
ergo_validate_config() { # [image]
local image="${1:-$(ergo_image)}" tmp rc
tmp="$(mktemp -d "$STACK_DIR/.validate.XXXXXX")" || return 1
VALIDATE_LOG="$(mktemp)"
if ! cp -a "$IRCD_DIR"/. "$tmp"/ 2>>"$VALIDATE_LOG"; then rm -rf "$tmp"; return 1; fi
rm -f "$tmp/ircd.lock"
# The validator has no network (so it cannot fight the live server for ports),
# which means it cannot reach a PostgreSQL history backend either. Neutralise
# that one setting in the COPY so everything else -- listeners, opers, MOTD,
# certs, and any pending DB schema upgrade -- is still really validated.
# PostgreSQL reachability is checked separately (see pg_check in update.sh).
if [[ "$(yaml_get_nested "$tmp/ircd.yaml" datastore postgresql enabled)" == "true" ]]; then
yaml_set_nested "$tmp/ircd.yaml" datastore postgresql enabled false || true
yaml_set_nested "$tmp/ircd.yaml" history persistent enabled false || true
fi
chown -R "$(ergo_uid):$(ergo_gid)" "$tmp"
ergo_helper -v "$tmp:/ircd" -w /ircd --entrypoint /bin/sh -- "$image" \
-c '/ircd-bin/ergo mkcerts >/dev/null 2>&1; exec /ircd-bin/ergo run --smoke' >>"$VALIDATE_LOG" 2>&1
rc=$?
rm -rf "$tmp"
return $rc
}
# ---------------------------------------------------------------------------
# ircd.yaml oper block edits (POSIX awk). The opers block is
# opers:
# <name>:
# class: ...
# password: "..."
# Each helper rewrites in place via a temp file and verifies the result.
# ---------------------------------------------------------------------------
yaml_server_name() {
awk '/^[a-z]/ {top=$1} top=="server:" && /^ name:/ {sub(/^ name:[ \t]*/,""); gsub(/["\047]/,""); print; exit}' "$1"
}
yaml_network_name() {
awk '/^[a-z]/ {top=$1} top=="network:" && /^ name:/ {sub(/^ name:[ \t]*/,""); gsub(/["\047]/,""); print; exit}' "$1"
}
yaml_oper_list() { # <file> -> "name class" per oper
awk '
/^[a-z]/ { top=$1 }
top=="opers:" && /^ [A-Za-z0-9_.-]+:/ { if (name!="") print name, cls; name=$1; sub(/:$/,"",name); cls="?" }
top=="opers:" && name!="" && /^ class:/ { cls=$2; gsub(/"/,"",cls) }
END { if (name!="") print name, cls }' "$1"
}
yaml_oper_has() { yaml_oper_list "$1" | awk -v n="$2" '$1==n {f=1} END {exit !f}'; }
yaml_oper_set_password() { # <file> <name> <hash>
local tmp; tmp="$(mktemp)"
NAME="$2" HASH="$3" awk '
/^[a-z]/ { top=$1; cur="" }
top=="opers:" && /^ [A-Za-z0-9_.-]+:/ { cur=$1; sub(/:$/,"",cur) }
top=="opers:" && cur==ENVIRON["NAME"] && /^ password:/ { print " password: \"" ENVIRON["HASH"] "\""; done=1; next }
{ print }
END { exit done ? 0 : 1 }' "$1" > "$tmp" || { rm -f "$tmp"; return 1; }
safe_replace "$tmp" "$1" 0600 "$(ergo_uid)" "$(ergo_gid)"
}
# Read one field from an oper's block ("" when absent). Commented-out template
# lines (#certfp:) do not count -- only a real setting.
yaml_oper_get_field() { # <file> <name> <key>
NAME="$2" KEY="$3" awk '
/^[a-z]/ { top=$1; cur="" }
top=="opers:" && /^ [A-Za-z0-9_.-]+:/ { cur=$1; sub(/:$/,"",cur) }
top=="opers:" && cur==ENVIRON["NAME"] && $0 ~ "^ " ENVIRON["KEY"] ":" {
sub(/^ [a-z0-9-]+:[ \t]*/,""); gsub(/"/,""); print; exit }' "$1"
}
# Set a field inside ONE oper's block (certfp, auto, vhost, modes...). Replaces
# the line when it is already there, otherwise inserts it after that oper's
# `class:` line so it lands inside the right block.
yaml_oper_set_field() { # <file> <name> <key> <value>
local tmp; tmp="$(mktemp)"
if NAME="$2" KEY="$3" VAL="$4" awk '
/^[a-z]/ { top=$1; cur="" }
top=="opers:" && /^ [A-Za-z0-9_.-]+:/ { cur=$1; sub(/:$/,"",cur) }
top=="opers:" && cur==ENVIRON["NAME"] && $0 ~ "^ " ENVIRON["KEY"] ":" {
print " " ENVIRON["KEY"] ": " ENVIRON["VAL"]; done=1; next }
{ print }
END { exit done ? 0 : 1 }' "$1" > "$tmp"; then
safe_replace "$tmp" "$1" 0600 "$(ergo_uid)" "$(ergo_gid)"
return 0
fi
NAME="$2" KEY="$3" VAL="$4" awk '
/^[a-z]/ { top=$1; cur="" }
top=="opers:" && /^ [A-Za-z0-9_.-]+:/ { cur=$1; sub(/:$/,"",cur) }
{ print }
top=="opers:" && cur==ENVIRON["NAME"] && /^ class:/ && !done {
print " " ENVIRON["KEY"] ": " ENVIRON["VAL"]; done=1 }
END { exit done ? 0 : 1 }' "$1" > "$tmp" || { rm -f "$tmp"; return 1; }
safe_replace "$tmp" "$1" 0600 "$(ergo_uid)" "$(ergo_gid)"
}
yaml_oper_rm_field() { # <file> <name> <key>
local tmp; tmp="$(mktemp)"
NAME="$2" KEY="$3" awk '
/^[a-z]/ { top=$1; cur="" }
top=="opers:" && /^ [A-Za-z0-9_.-]+:/ { cur=$1; sub(/:$/,"",cur) }
top=="opers:" && cur==ENVIRON["NAME"] && $0 ~ "^ " ENVIRON["KEY"] ":" { done=1; next }
{ print }
END { exit done ? 0 : 1 }' "$1" > "$tmp" || { rm -f "$tmp"; return 1; }
safe_replace "$tmp" "$1" 0600 "$(ergo_uid)" "$(ergo_gid)"
}
# A SHA-256 certfp as Ergo stores it: 64 lowercase hex digits, no colons.
normalize_certfp() {
local fp; fp="$(printf '%s' "$1" | tr -d ': \t\r\n' | tr 'A-Z' 'a-z')"
[[ "$fp" =~ ^[0-9a-f]{64}$ ]] || return 1
printf '%s' "$fp"
}
yaml_oper_add() { # <file> <name> <class> <hash> [whois-line]
local tmp; tmp="$(mktemp)"
local block
block=" $2:
class: \"$3\"
hidden: true
whois-line: ${5:-is a network operator}
password: \"$4\"
"
BLOCK="$block" awk '
/^[a-z]/ { if (inopers && !done) { printf "%s\n", ENVIRON["BLOCK"]; done=1 } inopers=($1=="opers:") }
{ print }
END { if (inopers && !done) printf "%s", ENVIRON["BLOCK"] }' "$1" > "$tmp" || { rm -f "$tmp"; return 1; }
safe_replace "$tmp" "$1" 0600 "$(ergo_uid)" "$(ergo_gid)"
yaml_oper_has "$1" "$2"
}
yaml_oper_rm() { # <file> <name>
local tmp; tmp="$(mktemp)"
NAME="$2" awk '
/^[a-z]/ { top=$1; skip=0 }
top=="opers:" && /^ [A-Za-z0-9_.-]+:/ { n=$1; sub(/:$/,"",n); skip=(n==ENVIRON["NAME"]); if (skip) next }
top=="opers:" && skip && (/^ / || /^[ \t]*$/) { next }
top=="opers:" && skip { skip=0 }
{ print }' "$1" > "$tmp" || { rm -f "$tmp"; return 1; }
safe_replace "$tmp" "$1" 0600 "$(ergo_uid)" "$(ergo_gid)"
! yaml_oper_has "$1" "$2"
}
# Set a key nested as <top>:/<4-space sub>:/<8-space key>. Only the key inside
# THAT block is rewritten -- datastore's mysql/postgresql/sqlite sub-blocks all
# have an `enabled` key, so a plain sed would hit the wrong one. A key that is
# not present is an error (rc 1), never a silent no-op.
yaml_set_nested() { # <file> <top> <sub> <key> <value>
local tmp; tmp="$(mktemp)"
TOP="$2" SUB="$3" KEY="$4" VAL="$5" awk '
/^[a-z][a-z-]*:/ { t=$0; sub(/:.*/,"",t); cur_top=t; cur_sub="" }
/^ [a-z][a-z0-9-]*:/ { s=$0; sub(/^ /,"",s); sub(/:.*/,"",s); cur_sub=s }
cur_top==ENVIRON["TOP"] && cur_sub==ENVIRON["SUB"] && $0 ~ "^ " ENVIRON["KEY"] ":" {
print " " ENVIRON["KEY"] ": " ENVIRON["VAL"]; done=1; next
}
{ print }
END { exit done ? 0 : 1 }' "$1" > "$tmp" || { rm -f "$tmp"; return 1; }
safe_replace "$tmp" "$1" 0600 "$(ergo_uid)" "$(ergo_gid)"
}
# Read a nested key back (for status/asserts); empty when absent.
yaml_get_nested() { # <file> <top> <sub> <key>
TOP="$2" SUB="$3" KEY="$4" awk '
/^[a-z][a-z-]*:/ { t=$0; sub(/:.*/,"",t); cur_top=t; cur_sub="" }
/^ [a-z][a-z0-9-]*:/ { s=$0; sub(/^ /,"",s); sub(/:.*/,"",s); cur_sub=s }
cur_top==ENVIRON["TOP"] && cur_sub==ENVIRON["SUB"] && $0 ~ "^ " ENVIRON["KEY"] ":" {
sub(/^ [a-z0-9-]+:[ \t]*/,""); gsub(/"/,""); print; exit
}' "$1"
}
# Which persistent-history backend the deployed ircd.yaml actually uses:
# prints sqlite | postgres | off. Reads the config, not .env, so it reports what
# Ergo is really doing after a hand edit.
history_backend() { # [ircd.yaml]
local f="${1:-$IRCD_DIR/ircd.yaml}"
[[ -f "$f" ]] || { printf 'off'; return 0; }
if [[ "$(yaml_get_nested "$f" history persistent enabled)" != "true" ]]; then printf 'off'; return 0; fi
if [[ "$(yaml_get_nested "$f" datastore postgresql enabled)" == "true" ]]; then printf 'postgres'; return 0; fi
if [[ "$(yaml_get_nested "$f" datastore sqlite enabled)" == "true" ]]; then printf 'sqlite'; return 0; fi
printf 'off'
}
# The first logging block's level (info <-> debug), for `ergoctl debug on|off`.
yaml_set_log_level() { # <file> <level>
local tmp; tmp="$(mktemp)"
LEVEL="$2" awk '
/^[a-z]/ { top=$1 }
top=="logging:" && !done && /^ level:/ { print " level: " ENVIRON["LEVEL"]; done=1; next }
{ print }
END { exit done ? 0 : 1 }' "$1" > "$tmp" || { rm -f "$tmp"; return 1; }
safe_replace "$tmp" "$1" 0600 "$(ergo_uid)" "$(ergo_gid)"
}
# ---------------------------------------------------------------------------
# TLS: Caddy's cert -> ./ircd/{fullchain,privkey}.pem -> SIGHUP (rehash)
# ---------------------------------------------------------------------------
# Newest <domain>.crt across issuer dirs (Let's Encrypt, ZeroSSL fallback, ...):
# certmagic writes .key, then .crt, then .json per issuance, each atomically.
ergo_cert_source() {
local d best="" f
d="$(ergo_domain)"
for f in "$CADDY_DIR"/data/caddy/certificates/*/"$d"/"$d".crt; do
[[ -f "$f" ]] || continue
if [[ -z "$best" || "$f" -nt "$best" ]]; then best="$f"; fi
done
[[ -n "$best" ]] || return 1
printf '%s' "$best"
}
cert_fingerprint() { openssl x509 -in "$1" -noout -fingerprint -sha256 2>/dev/null | cut -d= -f2 || true; }
cert_enddate() { openssl x509 -in "$1" -noout -enddate 2>/dev/null | cut -d= -f2 || true; }
cert_subject_cn() { openssl x509 -in "$1" -noout -subject 2>/dev/null | sed -E 's/.*CN ?= ?([^,\/]+).*/\1/' || true; }
cert_expires_within() { ! openssl x509 -in "$1" -noout -checkend "$2" >/dev/null 2>&1; } # <file> <seconds>
cert_is_selfsigned() {
[[ "$(openssl x509 -in "$1" -noout -issuer 2>/dev/null | cut -d= -f2-)" == "$(openssl x509 -in "$1" -noout -subject 2>/dev/null | cut -d= -f2-)" ]]
}
served_fingerprint() { # what :6697 currently presents ("" if the handshake fails)
local d; d="$(ergo_domain)"
timeout 10 openssl s_client -connect 127.0.0.1:6697 -servername "$d" </dev/null 2>/dev/null \
| openssl x509 -noout -fingerprint -sha256 2>/dev/null | cut -d= -f2 || true
}
# Sync Caddy's cert into ./ircd if it changed. Pairwise + atomic: both files are
# verified to match (pubkey) and staged on the same filesystem before two mv's,
# so Ergo can never see a half-renewed pair (which is fatal at the next start).
# rc: 0 synced or nothing to do | 2 no cert from Caddy yet | 1 failed (caller notifies)
CERTSYNC_MSG=""
ergo_certsync() {
local src_crt src_key src_json dst_crt dst_key tmpc tmpk want got i uid gid
CERTSYNC_MSG=""
command -v openssl >/dev/null 2>&1 || { CERTSYNC_MSG="openssl missing on the host"; return 1; }
src_crt="$(ergo_cert_source)" || { CERTSYNC_MSG="no certificate from Caddy yet"; return 2; }
src_key="${src_crt%.crt}.key"; src_json="${src_crt%.crt}.json"
[[ -f "$src_key" ]] || { CERTSYNC_MSG="renewal in progress (key missing)"; return 0; }
# .json is written last: if it is older than .crt/.key a renewal is mid-flight.
if [[ -f "$src_json" ]] && { [[ "$src_crt" -nt "$src_json" ]] || [[ "$src_key" -nt "$src_json" ]]; }; then
CERTSYNC_MSG="renewal in progress; retrying later"; return 0
fi
dst_crt="$IRCD_DIR/fullchain.pem"; dst_key="$IRCD_DIR/privkey.pem"
if [[ -f "$dst_crt" && -f "$dst_key" ]] && cmp -s "$src_crt" "$dst_crt" && cmp -s "$src_key" "$dst_key"; then
CERTSYNC_MSG="up to date"; return 0
fi
# cert and key must belong together
want="$(openssl x509 -in "$src_crt" -noout -pubkey 2>/dev/null | sha256sum | cut -d' ' -f1)"
got="$(openssl pkey -in "$src_key" -pubout 2>/dev/null | sha256sum | cut -d' ' -f1)"
if [[ -z "$want" || "$want" != "$got" ]]; then
CERTSYNC_MSG="cert/key mismatch in Caddy storage (renewal in progress?); retrying later"; return 0
fi
uid="$(ergo_uid)"; gid="$(ergo_gid)"
# Stage and keep the rollback copies OUTSIDE ircd/: that directory is
# writable by the container uid, and `install`/`cp` would follow a symlink
# planted there and write as root wherever it points. Only `mv` (rename(2),
# which replaces a symlink rather than following it) touches ircd/.
local prevdir="$STACK_DIR/.certsync"
install -d -m 0700 "$prevdir" || { CERTSYNC_MSG="could not create $prevdir"; return 1; }
tmpc="$prevdir/fullchain.pem.new"; tmpk="$prevdir/privkey.pem.new"
if ! install -m 0600 -o "$uid" -g "$gid" "$src_crt" "$tmpc" || ! install -m 0600 -o "$uid" -g "$gid" "$src_key" "$tmpk"; then
rm -f "$tmpc" "$tmpk"; CERTSYNC_MSG="could not stage the new pair in $prevdir"; return 1
fi
# source must not have moved under us mid-copy
if ! cmp -s "$src_crt" "$tmpc" || ! cmp -s "$src_key" "$tmpk"; then
rm -f "$tmpc" "$tmpk"; CERTSYNC_MSG="source changed during copy; retrying later"; return 0
fi
# Roll-back copies, also outside ircd/. Read the live pair only if it is a
# real file -- a symlink there is the container trying to exfiltrate.
rm -f "$prevdir/fullchain.pem.prev" "$prevdir/privkey.pem.prev"
if ! refuse_symlink "$dst_crt" "ircd/fullchain.pem" || ! refuse_symlink "$dst_key" "ircd/privkey.pem"; then
rm -f "$tmpc" "$tmpk"; CERTSYNC_MSG="ircd/ TLS path is a symlink; refusing to sync"; return 1
fi
[[ -f "$dst_crt" ]] && cp -p "$dst_crt" "$prevdir/fullchain.pem.prev"
[[ -f "$dst_key" ]] && cp -p "$dst_key" "$prevdir/privkey.pem.prev"
mv -f "$tmpc" "$dst_crt"; mv -f "$tmpk" "$dst_key"
if ergo_running; then
dc kill -s HUP ergo >/dev/null 2>&1 || true
want="$(cert_fingerprint "$dst_crt")"; got=""
for i in 1 2 3 4 5 6 7 8 9 10; do
got="$(served_fingerprint)"
[[ -n "$got" && "$got" == "$want" ]] && break
sleep 1
done
if [[ "$got" != "$want" ]]; then
# Ergo kept the previous tls config (a failed rehash leaves the old one
# active). Put the old pair back so the next start is not fatal.
if [[ -f "$prevdir/fullchain.pem.prev" && -f "$prevdir/privkey.pem.prev" ]]; then
mv -f "$prevdir/fullchain.pem.prev" "$dst_crt"; mv -f "$prevdir/privkey.pem.prev" "$dst_key"
dc kill -s HUP ergo >/dev/null 2>&1 || true
fi
CERTSYNC_MSG="Ergo did not present the new cert after SIGHUP (see: docker compose logs ergo); previous pair restored"
return 1
fi
fi
rm -f "$prevdir/fullchain.pem.prev" "$prevdir/privkey.pem.prev"
CERTSYNC_MSG="installed cert for $(cert_subject_cn "$dst_crt") (expires $(cert_enddate "$dst_crt")) from ${src_crt#"$CADDY_DIR"/data/caddy/certificates/}"
return 0
}
# ---------------------------------------------------------------------------
# Small persistent state for de-duplicating notifications (.state/<key>)
# ---------------------------------------------------------------------------
state_get() { cat "$STATE_DIR/$1" 2>/dev/null || true; }
state_set() { install -d -m 0700 "$STATE_DIR" 2>/dev/null; printf '%s\n' "$2" > "$STATE_DIR/$1"; }
state_rm() { rm -f "$STATE_DIR/$1"; }
# ---------------------------------------------------------------------------
# Notify (reuse the login-notifier's ntfy config, like scripts/auto-update.sh)
# ---------------------------------------------------------------------------
send_notice() { # <title> <priority> <body>
[[ "${NOTIFY:-1}" == "1" ]] || return 0
[[ -r "$SSH_NOTIFY_CONF" ]] || return 0
# shellcheck disable=SC1090
. "$SSH_NOTIFY_CONF"
[[ -n "${NTFY_URL:-}" ]] || return 0
command -v curl >/dev/null 2>&1 || return 0
local title="$1" prio="$2" body="$3" host
host="$(hostname -f 2>/dev/null || hostname 2>/dev/null || echo unknown)"
set -- -fsS -m 5 -H "X-Title: ${title}" -H "X-Priority: ${prio}"
[[ -n "${NTFY_EMAIL:-}" ]] && set -- "$@" -H "X-Email: ${NTFY_EMAIL}"
local t="ergo"; [[ -n "${NTFY_REGION:-}" ]] && t="${t},${NTFY_REGION}"
set -- "$@" -H "X-Tags: ${t}"
if [[ "${DRY_RUN:-0}" == "1" ]]; then
echo "DRY: curl ntfy ($prio): $body"
elif [[ -n "${NTFY_TOKEN:-}" ]]; then
# The token goes in on stdin, never in argv: /proc/<pid>/cmdline is
# world-readable, and every local process is inside this stack's trust
# boundary already (see README: Security model).
printf 'header = "Authorization: Bearer %s"\n' "$NTFY_TOKEN" \
| curl "$@" --config - -d "${body} [${host}]" "$NTFY_URL" >/dev/null 2>&1 || true
else
curl "$@" -d "${body} [${host}]" "$NTFY_URL" >/dev/null 2>&1 || true
fi
return 0
}
+616
View File
@@ -0,0 +1,616 @@
#!/usr/bin/env bash
#
# update.sh -- keep the Ergo stack current: container updates with a DB snapshot
# and health-checked rollback, plus the Caddy -> Ergo TLS certificate sync.
# Companion to deploy.sh; installed in $STACK_DIR next to ergolib.sh.
#
# Subcommands:
# check (default) current vs latest, advisories, cert state; changes nothing
# run what the daily schedule runs: apply UPDATE_POLICY, cert sync, expiry check
# update update now (to latest, or TARGET_VERSION=x.y.z); honours FORCE_UPDATE
# certsync copy Caddy's cert into ircd/ and rehash if it changed (15-min schedule)
# caddy pull a newer Caddy image and recreate it (health-checked)
# install write /etc/ergo-update.conf + schedule the jobs
# uninstall remove the schedule
#
# Policy (UPDATE_POLICY):
# latest update to the newest release whenever one exists (default)
# security update ONLY when a published GitHub security advisory covers the
# running version, or a release between current and latest has a
# "### Security" section in its notes
# off never change the running version (check/notify only)
#
# Safety rails (all policies):
# * releases whose notes announce "Compatibility breaks" are HELD (notify
# only) unless FORCE_UPDATE=1 -- review the notes, then `update update`
# * the new image must load the current ircd/ (ergo run --smoke on a copy,
# which also dry-runs any DB schema upgrade) before the live server is touched
# * users get a NOTICE and UPDATE_GRACE seconds before the restart
# * ircd.db is snapshotted with Ergo stopped (buntdb is append-only; a live
# copy can miss the last second of writes)
# * health = compose healthcheck (IRC-level) + a registration handshake; on
# failure the previous tag is restored and, if the new version bumped the
# DB schema, the snapshot is put back
#
# Env, in precedence order: environment > /etc/ergo-update.conf > the stack's
# .env > the built-in default.
# STACK_DIR=/srv/ergo UPDATE_POLICY=latest FORCE_UPDATE=0 UPDATE_GRACE=60
# CADDY_AUTOUPDATE=0 NOTIFY=1 SSH_NOTIFY_CONF=/etc/ssh-notify.conf DRY_RUN=0
# TARGET_VERSION= GH_REPO=ergochat/ergo LOG=/var/log/ergo-update.log
# SCHEDULE_UPDATES=1 (install: 0 schedules certsync only, no daily update)
# VERBOSE= (certsync: also log "nothing to do" outcomes)
set -euo pipefail
SELF="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)/$(basename "${BASH_SOURCE[0]}")"
SELF_DIR="$(dirname "$SELF")"
log() { printf '\033[1;32m[+]\033[0m %s\n' "$*"; }
warn() { printf '\033[1;33m[!]\033[0m %s\n' "$*" >&2; }
die() { printf '\033[1;31m[x]\033[0m %s\n' "$*" >&2; exit 1; }
# shellcheck source=ergolib.sh
. "$SELF_DIR/ergolib.sh"
load_conf
: "${STACK_DIR:=$SELF_DIR}"
ergo_set_paths
# .env is the stack's own settings file and documents these knobs, so honour it
# for anything the environment and /etc/ergo-update.conf did not already set.
env_defaults UPDATE_POLICY FORCE_UPDATE UPDATE_GRACE CADDY_AUTOUPDATE
: "${UPDATE_POLICY:=latest}"
: "${FORCE_UPDATE:=0}"
: "${UPDATE_GRACE:=60}"
: "${CADDY_AUTOUPDATE:=0}"
: "${TARGET_VERSION:=}"
: "${SCHEDULE_UPDATES:=1}"
: "${LOG:=/var/log/ergo-update.log}"
# ---------------------------------------------------------------------------
# Release metadata (GitHub). Unauthenticated: 60 requests/hour -- a run uses 2-3.
# ---------------------------------------------------------------------------
RELEASES_JSON=""
# Fetch the release list ONCE per process. This must be called as a plain
# statement, never inside a pipeline or $(): the assignment would land in a
# subshell and every use would re-download the list (and burn the 60/hour
# anonymous API budget).
load_releases() {
[[ -n "$RELEASES_JSON" ]] && return 0
RELEASES_JSON="$(fetch "https://api.github.com/repos/${GH_REPO}/releases?per_page=50" || true)"
printf '%s' "$RELEASES_JSON" | jq -e 'type=="array"' >/dev/null 2>&1 || RELEASES_JSON="[]"
return 0
}
# Bodies of stable releases with cur < version <= upto, newest first:
# "=== <tag>" header lines followed by the body.
release_notes_between() { # <cur> <upto>
local cur="$1" upto="$2" tag v
load_releases
while IFS= read -r tag; do
[[ -n "$tag" ]] || continue
v="$(normver "$tag")"; [[ -n "$v" ]] || continue
ver_gt "$v" "$cur" || continue
ver_le "$v" "$upto" || continue
printf '=== %s\n' "$tag"
printf '%s' "$RELEASES_JSON" | jq -r --arg t "$tag" '.[] | select(.tag_name==$t) | .body // ""'
printf '\n'
done < <(printf '%s' "$RELEASES_JSON" | jq -r '.[] | select(.prerelease==false and .draft==false) | .tag_name' 2>/dev/null)
}
# Published GitHub security advisories covering <cur>. Prints "GHSA-... <patched>"
# lines. Ranges look like "<= v2.19.0" or ">= v2.15.0, < v2.19.1".
ghsa_covering() { # <cur>
local cur="$1" json id rng patched part op ver ok
json="$(fetch "https://api.github.com/repos/${GH_REPO}/security-advisories" || true)"
printf '%s' "$json" | jq -e 'type=="array"' >/dev/null 2>&1 || return 0
while IFS=$'\t' read -r id rng patched; do
[[ -n "$rng" ]] || continue
ok=1
IFS=',' read -r -a parts <<< "$rng"
for part in "${parts[@]}"; do
part="$(printf '%s' "$part" | tr -d ' ')"
op="$(printf '%s' "$part" | grep -oE '^(<=|>=|==|<|>|=)' || echo '=')"
ver="$(normver "$part")"; [[ -n "$ver" ]] || { ok=0; break; }
case "$op" in
'<') ver_gt "$ver" "$cur" || ok=0 ;;
'<=') ver_ge "$ver" "$cur" || ok=0 ;;
'>') ver_gt "$cur" "$ver" || ok=0 ;;
'>=') ver_ge "$cur" "$ver" || ok=0 ;;
*) [[ "$(ver_cmp "$ver" "$cur")" == "0" ]] || ok=0 ;;
esac
(( ok )) || break
done
(( ok )) && printf '%s %s\n' "$id" "$(normver "$patched")"
done < <(printf '%s' "$json" | jq -r '.[] | select(.state=="published") | .ghsa_id as $id | (.vulnerabilities // [])[] | [$id, (.vulnerable_version_range // ""), ((.patched_versions // "") | tostring)] | @tsv' 2>/dev/null)
}
# ---------------------------------------------------------------------------
# Decide what to do. Sets globals (not via stdout -- a $() would lose them):
# TARGET (x.y.z or ""), REASON, HOLD (1 = compat break, notify only),
# NOTES (one-line summary for the notification), DBCHANGE (0/1)
# ---------------------------------------------------------------------------
TARGET=""; REASON=""; HOLD=0; NOTES=""; DBCHANGE=0
resolve_target() { # <current x.y.z>
local cur="$1" latest notes ghsa ids
TARGET=""; REASON=""; HOLD=0; NOTES=""; DBCHANGE=0
latest="$(normver "$(ergo_latest_tag || true)")"
if [[ -n "$TARGET_VERSION" ]]; then
TARGET="$(normver "$TARGET_VERSION")"; REASON="target override ${TARGET}"
else
case "$UPDATE_POLICY" in
off) REASON="policy=off (no changes)"; return 0 ;;
security)
ghsa="$(ghsa_covering "$cur")"
if [[ -n "$ghsa" ]]; then
ids="$(printf '%s\n' "$ghsa" | awk '{print $1}' | paste -sd, - || true)"
# NF>=2: an advisory with no patched version yields only an id.
TARGET="$(printf '%s\n' "$ghsa" | awk 'NF>=2 {print $2}' | tail -n1 || true)"
[[ -n "$TARGET" ]] || TARGET="$latest"
REASON="security advisory covers ${cur}: ${ids}"
elif [[ -n "$latest" ]] && ver_gt "$latest" "$cur" \
&& grep -qiE '^#+ *security|security vulnerabilit' <<< "$(release_notes_between "$cur" "$latest")"; then
TARGET="$latest"; REASON="a release since ${cur} carries a Security section"
else
REASON="no security advisory or security release covers ${cur}"; return 0
fi ;;
latest|*)
[[ -n "$latest" ]] || { REASON="could not resolve the latest release"; return 0; }
TARGET="$latest"; REASON="policy=latest" ;;
esac
fi
[[ -n "$TARGET" ]] || return 0
notes="$(release_notes_between "$cur" "$TARGET")"
# Feed grep from a here-string, not a pipe: `printf | grep -q` on notes
# larger than the pipe buffer makes printf die of SIGPIPE, and pipefail
# then reports "no match" for what was in fact a match.
if grep -qiE 'compatibility break|not backwards[- ]compatible|backwards[- ]incompatible' <<< "$notes"; then
HOLD=1
fi
if grep -qiE 'database change|database file format changes|schema' <<< "$notes" \
&& ! grep -qiE 'no changes to the config file format or database file format' <<< "$notes"; then
DBCHANGE=1
fi
NOTES="$(grep -iE 'config file format|database (file )?format|database change|compatibility break' <<< "$notes" \
| head -n 3 | sed 's/^[ *-]*//' | tr '\n' ' ' | cut -c1-300 || true)"
return 0
}
# ---------------------------------------------------------------------------
# Apply an update: pre-flight on a copy, warn users, stop, snapshot, up, verify;
# roll back (tag + DB snapshot when needed) on failure.
# ---------------------------------------------------------------------------
notify_users() { # <text> -- best effort, only when Ergo is up and we can oper
ergo_running || return 0
irc_raw --oper --quiet 1 -- "NOTICE \$\$* :$1" >/dev/null 2>&1 || true
}
apply_update() { # <from x.y.z> <to x.y.z> <reason>
local from="$1" to="$2" reason="$3" image_base new_image snap ts cause prev_tag
image_base="$(env_get ERGO_IMAGE)"; image_base="${image_base:-ghcr.io/ergochat/ergo}"
new_image="${image_base}:v${to}"
# Roll back to the tag that is actually pinned right now. Reconstructing
# "v${from}" would write a bare "v" when the running version is unknown
# (e.g. .env still holds the 'stable' fallback and the container is down).
prev_tag="$(ergo_tag)"
log "Updating Ergo ${from:-?} -> ${to} (${reason})..."
if [[ "$DRY_RUN" == "1" ]]; then
echo "DRY: pull ${new_image}; validate on a copy; NOTICE users; stop; snapshot ircd.db; ERGO_TAG=v${to}; up -d --no-deps ergo"
return 0
fi
cp -a "$ENV_FILE" "${ENV_FILE}.bak.$(date -u +%Y%m%d%H%M%S)" 2>/dev/null || true
log "Pulling ${new_image}..."
if ! docker pull -q "$new_image" >/dev/null 2>&1; then
send_notice "Ergo update FAILED" "high" "pull of ${new_image} failed; staying on ${from}"
return 1
fi
log "Pre-flight: loading the current ircd/ with ${new_image} (throwaway container)..."
if ! ergo_validate_config "$new_image"; then
cause="$(tail -n 5 "$VALIDATE_LOG" 2>/dev/null | tr '\n' ' ' | cut -c1-300)"
warn "Pre-flight failed: $cause"
send_notice "Ergo update HELD" "high" "${to} cannot load the current config/DB: ${cause}. Staying on ${from}; run 'ergoctl update check'."
return 1
fi
rm -f "$VALIDATE_LOG"
if (( UPDATE_GRACE > 0 )) && ergo_running; then
log "Warning connected users (${UPDATE_GRACE}s grace)..."
notify_users "Server restarting for an upgrade to Ergo ${to} in ${UPDATE_GRACE} seconds -- you will be disconnected briefly."
sleep "$UPDATE_GRACE"
fi
log "Stopping Ergo and snapshotting ircd.db..."
dc stop -t 15 ergo >/dev/null 2>&1 || true
ts="$(date -u +%Y%m%d%H%M%S)"
install -d -m 0700 "$BACKUP_DIR"
snap="$BACKUP_DIR/ircd.db.${from:-unknown}.${ts}"
# -P: never follow a symlink in the uid-owned ircd/ (see ergolib's note).
if [[ -f "$IRCD_DIR/ircd.db" ]] && refuse_symlink "$IRCD_DIR/ircd.db" "ircd/ircd.db"; then
cp -Pp "$IRCD_DIR/ircd.db" "$snap" && chmod 0600 "$snap"
else
snap=""
fi
# The message history is a separate store; snapshot it too so a rollback is
# not silently lossy.
case "$(history_backend)" in
sqlite) [[ -f "$IRCD_DIR/ergo_history.db" ]] && cp -Pp "$IRCD_DIR/ergo_history.db" "$BACKUP_DIR/ergo_history.db.${from:-unknown}.${ts}" 2>/dev/null || true ;;
postgres) dc up -d --no-deps postgres >/dev/null 2>&1 || true
if pg_check; then
dc exec -T postgres pg_dump -U "$(env_get POSTGRES_USER)" -d "$(env_get POSTGRES_DB)" \
> "$BACKUP_DIR/history.sql.${from:-unknown}.${ts}" 2>/dev/null \
|| { rm -f "$BACKUP_DIR/history.sql.${from:-unknown}.${ts}"; warn "pg_dump failed; the rollback would not restore message history."; }
fi ;;
esac
env_set ERGO_TAG "v${to}"
dc up -d --no-deps ergo >/dev/null 2>&1 || true
if wait_healthy ergo 120 && irc_probe; then
log "Ergo ${to} is healthy."
state_rm failed
record_template "$new_image" "$from" "$to"
prune_snapshots
local b="updated ${from:-?} -> ${to} (${reason})"
[[ -n "$NOTES" ]] && b="${b}. Notes: ${NOTES}"
[[ -n "$TEMPLATE_NOTE" ]] && b="${b}. ${TEMPLATE_NOTE}"
send_notice "Ergo updated" "default" "$b"
return 0
fi
cause="$(dc logs --tail 60 ergo 2>/dev/null | grep -E 'Config file did not load|Database requires update|Could not load server|failed to|fatal|panic' | tail -n 2 | cut -c1-200 | tr '\n' ' ' || true)"
warn "Ergo ${to} did not become healthy (${cause:-no diagnostic in the log}); rolling back to ${prev_tag}."
dc stop -t 10 ergo >/dev/null 2>&1 || true
env_set ERGO_TAG "$prev_tag"
dc up -d --no-deps ergo >/dev/null 2>&1 || true
if ! wait_healthy ergo 90; then
# The old binary refuses a newer schema ("Database requires update") --
# only then does the snapshot go back (otherwise keep the live DB).
if [[ -n "$snap" ]] && { [[ "$DBCHANGE" == 1 ]] || dc logs --tail 40 ergo 2>/dev/null | grep -q 'Database requires update'; }; then
warn "Old version rejects the upgraded database; restoring the snapshot ${snap}."
dc stop -t 10 ergo >/dev/null 2>&1 || true
install -m 0600 -o "$(ergo_uid)" -g "$(ergo_gid)" "$snap" "$IRCD_DIR/ircd.db"
dc up -d --no-deps ergo >/dev/null 2>&1 || true
wait_healthy ergo 90 || true
fi
fi
prune_snapshots
# Remember the target that failed so the daily run does not repeat this
# whole disruptive cycle (warn users, stop, swap, roll back) every night.
# A manual `ergoctl update update` still forces a retry.
state_set failed "$to"
if [[ "$(svc_state ergo)" == "running healthy" ]]; then
send_notice "Ergo update FAILED" "high" "${to} unhealthy (${cause:-see docker compose logs ergo}); rolled back to ${prev_tag}. Snapshot: ${snap:-none}. Will not retry automatically; run 'ergoctl update update' after investigating."
else
send_notice "Ergo DOWN after failed update" "urgent" "${to} failed (${cause:-?}) and ${prev_tag} did not come back healthy. Snapshot: ${snap:-none}. Run 'ergoctl status' / 'docker compose logs ergo'."
fi
return 1
}
# Keep the image's default.yaml per version and diff it against the previous one
# so template drift (new options, renamed keys) is visible after an upgrade.
TEMPLATE_NOTE=""
record_template() { # <image> <from> <to>
local image="$1" from="$2" to="$3" new old diff added removed
TEMPLATE_NOTE=""
install -d -m 0700 "$TEMPLATES_DIR"
new="$TEMPLATES_DIR/default.yaml.v${to}"
ergo_extract_template "$image" "$new" 2>/dev/null || { rm -f "$new"; return 0; }
chmod 0600 "$new"
old="$TEMPLATES_DIR/default.yaml.v${from}"
[[ -f "$old" ]] || return 0
diff="$TEMPLATES_DIR/diff-v${from}-v${to}.txt"
if diff -u "$old" "$new" > "$diff" 2>/dev/null; then rm -f "$diff"; return 0; fi
chmod 0600 "$diff" 2>/dev/null || true
added="$(grep -cE '^\+[^+]' "$diff" || true)"; removed="$(grep -cE '^-[^-]' "$diff" || true)"
TEMPLATE_NOTE="default config changed upstream (+${added}/-${removed} lines): review ${diff} against ircd/ircd.yaml"
log "$TEMPLATE_NOTE"
}
prune_snapshots() { # keep the last 5 of each snapshot kind, and 5 .env backups
local f n=0 pat
for pat in 'ircd.db.*' 'ergo_history.db.*' 'history.sql.*'; do
n=0
for f in $(ls -1t "$BACKUP_DIR"/$pat 2>/dev/null); do
n=$((n + 1)); (( n > 5 )) && rm -f "$f"
done
done
n=0
for f in $(ls -1t "$STACK_DIR"/.env.bak.* 2>/dev/null); do
n=$((n + 1)); (( n > 5 )) && rm -f "$f"
done
return 0
}
# The pinned tag must be an immutable vX.Y.Z: deploy.sh falls back to 'stable'
# when GitHub is unreachable, and a floating tag silently changes the running
# version on any pull, with no pre-flight, snapshot or rollback. Pin it to
# whatever is running as soon as we can see a version (no restart needed).
pin_floating_tag() { # <current x.y.z>
local cur="$1" tag
tag="$(ergo_tag)"
[[ -n "$cur" ]] || return 0
case "$tag" in v[0-9]*) return 0 ;; esac
env_set ERGO_TAG "v${cur}"
log "Pinned ERGO_TAG=v${cur} (was '${tag}', which floats)."
}
# ---------------------------------------------------------------------------
# Cert sync + expiry watch
# ---------------------------------------------------------------------------
do_certsync() {
local rc=0 prev
ergo_certsync || rc=$?
prev="$(state_get certsync)"
case "$rc" in
0) [[ "$CERTSYNC_MSG" == "up to date" ]] || log "certsync: $CERTSYNC_MSG"
# rc 0 also covers the "renewal in progress, retrying later" cases,
# which synced nothing -- clearing a stored failure on those would
# send a false all-clear while the real problem persists.
if [[ -n "$prev" && ( "$CERTSYNC_MSG" == "up to date" || "$CERTSYNC_MSG" == installed\ cert* ) ]]; then
state_rm certsync; send_notice "Ergo TLS recovered" "default" "cert sync ok: ${CERTSYNC_MSG}"
fi ;;
2) [[ -n "${VERBOSE:-}" ]] && log "certsync: $CERTSYNC_MSG"
return 0 ;;
*) warn "certsync: $CERTSYNC_MSG"
if [[ "$prev" != "$CERTSYNC_MSG" ]]; then state_set certsync "$CERTSYNC_MSG"; send_notice "Ergo TLS sync FAILED" "high" "$CERTSYNC_MSG"; fi
return 1 ;;
esac
return 0
}
cert_expiry_check() { # daily: warn once per day if the served cert is self-signed or expiring
local f="$IRCD_DIR/fullchain.pem" today msg=""
[[ -f "$f" ]] || return 0
today="$(date -u +%F)"
if cert_is_selfsigned "$f"; then
msg="Ergo is serving a SELF-SIGNED certificate on 6697 (Caddy has not issued one yet: check DNS for $(ergo_domain) and ports 80/443)."
elif cert_expires_within "$f" $((14 * 86400)); then
msg="Ergo's TLS cert expires soon ($(cert_enddate "$f")) and certsync has not replaced it -- is the 15-minute job running? ($SELF certsync)"
fi
[[ -n "$msg" ]] || { state_rm certexpiry; return 0; }
warn "$msg"
[[ "$(state_get certexpiry)" == "$today" ]] && return 0
state_set certexpiry "$today"
send_notice "Ergo TLS attention" "high" "$msg"
}
# ---------------------------------------------------------------------------
# Caddy: pull + recreate when the image moved (drops web-client websockets)
# ---------------------------------------------------------------------------
do_caddy() {
preflight
local tag before after
tag="$(env_get CADDY_TAG)"; tag="${tag:-2-alpine}"
before="$(docker inspect --format '{{.Image}}' ergo-caddy 2>/dev/null | head -n1 || true)"
log "Pulling caddy:${tag}..."
docker pull -q "caddy:${tag}" >/dev/null 2>&1 || { warn "pull failed"; return 1; }
after="$(docker image inspect --format '{{.Id}}' "caddy:${tag}" 2>/dev/null | head -n1 || true)"
if [[ -n "$before" && "$before" == "$after" ]]; then log "Caddy is current (${tag})."; return 0; fi
[[ "$DRY_RUN" == "1" ]] && { echo "DRY: docker compose up -d --no-deps caddy"; return 0; }
log "Recreating Caddy (web-client websockets will reconnect)..."
dc up -d --no-deps caddy >/dev/null 2>&1 || true
if wait_healthy caddy 60; then
send_notice "Caddy updated" "min" "caddy:${tag} recreated (image $(printf '%s' "$after" | cut -c8-19))"
return 0
fi
send_notice "Caddy update FAILED" "high" "caddy:${tag} is not healthy after recreate; check docker compose logs caddy"
return 1
}
# ---------------------------------------------------------------------------
# Subcommands
# ---------------------------------------------------------------------------
preflight() {
[[ $EUID -eq 0 ]] || die "Run as root."
[[ -f "$STACK_DIR/docker-compose.yml" ]] || die "No stack at $STACK_DIR (set STACK_DIR)."
command -v docker >/dev/null 2>&1 || die "docker not found."
command -v jq >/dev/null 2>&1 || die "jq is required (release metadata); install it."
}
# With HISTORY=postgres, Ergo will not start if its history backend is down --
# so an update must not stop a healthy server until the database is confirmed up.
# The config validator cannot check this: it runs with no network on purpose.
pg_gate() { # <what we are about to do>
pg_enabled || return 0
pg_check && return 0
warn "PostgreSQL is not ready ($PG_CHECK_MSG); skipping $1 -- Ergo would not restart."
send_notice "Ergo update skipped" "high" "PostgreSQL not ready (${PG_CHECK_MSG}); ${1} was not attempted."
return 1
}
do_check() {
preflight
local cur latest ghsa held="" prio="min" note
cur="$(ergo_version_running)"
pin_floating_tag "$cur"
latest="$(normver "$(ergo_latest_tag || true)")"
log "Running: ${cur:-unknown} (pinned $(ergo_tag)) | latest release: ${latest:-unknown} | policy: ${UPDATE_POLICY}"
ghsa="$(ghsa_covering "$cur" || true)"
if [[ -n "$ghsa" ]]; then
note="VULNERABLE: $(printf '%s\n' "$ghsa" | awk '{print $1}' | paste -sd, -) covers ${cur}"; prio="high"; warn "$note"
else
note="no published advisory covers ${cur}"; log "$note"
fi
if [[ -n "$cur" && -n "$latest" ]] && ver_gt "$latest" "$cur"; then
resolve_target "$cur"
log "A newer release is available: ${cur} -> ${latest}"
[[ "$prio" == "min" ]] && prio="default"
note="${note}; newer release ${latest} available"
if [[ "$HOLD" == 1 ]]; then held=" -- HELD: release notes announce compatibility breaks (review, then FORCE_UPDATE=1 or 'ergoctl update update')"; warn "${held# -- }"; fi
[[ -n "$NOTES" ]] && log "Release notes: $NOTES"
send_notice "Ergo check" "$prio" "${note}${held}"
fi
local f="$IRCD_DIR/fullchain.pem"
if [[ -f "$f" ]]; then
if cert_is_selfsigned "$f"; then warn "TLS: serving a SELF-SIGNED cert (no Let's Encrypt cert from Caddy yet)."
else log "TLS: $(cert_subject_cn "$f"), expires $(cert_enddate "$f") (served: $( [[ "$(served_fingerprint)" == "$(cert_fingerprint "$f")" ]] && echo matches || echo DIFFERS -- rehash needed ))"; fi
fi
}
do_run() {
preflight
if [[ "$DRY_RUN" != "1" ]]; then
# Create the log directory only when it is missing. `install -d -m` also
# re-modes an EXISTING directory, and this resolves to /var/log, which on
# Debian/Ubuntu ships 0775 root:syslog so rsyslog (running as `syslog`)
# can create files there -- re-moding it to 0755 silently breaks logging.
[[ -d "$(dirname "$LOG")" ]] || install -d -m 0755 "$(dirname "$LOG")" 2>/dev/null || true
echo "=== ergo-update $(date -u +%FT%TZ) ===" >> "$LOG"
# Scheduled runs have nowhere to send stdout (busybox crond would try to
# mail it and these hosts have no MTA), so keep the whole run in the log.
[[ -t 1 ]] || exec >>"$LOG" 2>&1
fi
do_certsync || true
cert_expiry_check || true
local cur
cur="$(ergo_version_running)"
pin_floating_tag "$cur"
resolve_target "$cur"
log "current=${cur:-?} | ${REASON}"
if [[ -n "$TARGET" && -n "$cur" ]] && ver_gt "$TARGET" "$cur"; then
if [[ "$(state_get failed)" == "$TARGET" && "$FORCE_UPDATE" != "1" ]]; then
log "Skipping ${TARGET}: it already failed here and was rolled back. Investigate, then run 'ergoctl update update' (or set FORCE_UPDATE=1)."
elif ! pg_gate "the update to ${TARGET}"; then
: # postgres is down; pg_gate already warned and notified
elif [[ "$HOLD" == 1 && "$FORCE_UPDATE" != "1" ]]; then
warn "HELD: ${TARGET} announces compatibility breaks. Review the release notes, then 'ergoctl update update' (or FORCE_UPDATE=1)."
if [[ "$(state_get held)" != "$TARGET" ]]; then
state_set held "$TARGET"
send_notice "Ergo update HELD" "default" "${cur} -> ${TARGET} announces compatibility breaks; not applied automatically. Notes: ${NOTES:-see GitHub release}. Apply with 'ergoctl update update'."
fi
else
state_rm held
apply_update "$cur" "$TARGET" "$REASON" || true
fi
else
log "No update to apply."
fi
if [[ "$CADDY_AUTOUPDATE" == "1" ]]; then do_caddy || true; fi
}
do_update() {
preflight
local cur
cur="$(ergo_version_running)"
[[ -n "$cur" ]] || die "Cannot tell which version is running (container down and ERGO_TAG='$(ergo_tag)' is not a version). Start Ergo, or pin ERGO_TAG=vX.Y.Z in $ENV_FILE first."
pin_floating_tag "$cur"
if [[ -z "$TARGET_VERSION" && "$UPDATE_POLICY" == "off" ]]; then UPDATE_POLICY=latest; fi
resolve_target "$cur"
[[ -n "$TARGET" ]] || die "Could not determine a target version (${REASON})."
if [[ -n "$cur" ]] && ! ver_gt "$TARGET" "$cur" && [[ -z "$TARGET_VERSION" ]]; then
log "Already on ${cur} (latest ${TARGET}); nothing to do."
return 0
fi
if [[ "$HOLD" == 1 && "$FORCE_UPDATE" != "1" ]]; then
warn "Release notes between ${cur} and ${TARGET} announce compatibility breaks: ${NOTES}"
warn "Proceeding because you asked explicitly (this is 'update update'). Set FORCE_UPDATE=1 to also auto-apply such releases."
fi
pg_gate "the update to ${TARGET}" || die "Start PostgreSQL first (docker compose up -d postgres), then retry."
apply_update "$cur" "$TARGET" "${REASON} (manual)"
}
# Only what the scheduled job needs to FIND the stack lives here. The update
# knobs (UPDATE_POLICY, FORCE_UPDATE, UPDATE_GRACE, CADDY_AUTOUPDATE) are read
# from the stack's .env so there is a single place to edit them; setting one
# here (or in the environment) still overrides .env for that run.
write_conf() {
cat > "$ERGO_UPDATE_CONF" <<CONF
# Where the scheduled Ergo jobs find the stack (deployments/ergo/update.sh).
# The update knobs live in \${STACK_DIR}/.env -- edit them there. Anything set
# here overrides .env, so leave it alone unless you mean to pin a value.
STACK_DIR="${STACK_DIR}"
NOTIFY="${NOTIFY}"
CONF
chmod 644 "$ERGO_UPDATE_CONF"
log "Wrote $ERGO_UPDATE_CONF (update knobs come from ${STACK_DIR}/.env)"
}
do_install() {
[[ $EUID -eq 0 ]] || die "Run as root."
write_conf
case "$(osfam)" in
alpine)
install -d -m 0755 /etc/periodic/15min /etc/periodic/daily
printf '#!/bin/sh\nexec bash "%s" certsync\n' "$SELF" > /etc/periodic/15min/ergo-certsync
chmod +x /etc/periodic/15min/ergo-certsync
if [[ "$SCHEDULE_UPDATES" == "1" ]]; then
printf '#!/bin/sh\nexec bash "%s" run\n' "$SELF" > /etc/periodic/daily/ergo-update
chmod +x /etc/periodic/daily/ergo-update
else
rm -f /etc/periodic/daily/ergo-update
fi
if command -v rc-update >/dev/null 2>&1; then
rc-update add crond default >/dev/null 2>&1 || true
rc-service crond status >/dev/null 2>&1 || rc-service crond start >/dev/null 2>&1 || true
fi
log "Scheduled: /etc/periodic/15min/ergo-certsync$( [[ "$SCHEDULE_UPDATES" == "1" ]] && printf ', /etc/periodic/daily/ergo-update (policy=%s)' "$UPDATE_POLICY" )." ;;
*)
cat > /etc/systemd/system/ergo-certsync.service <<EOF
[Unit]
Description=Ergo: sync Caddy's TLS cert into the IRC server
After=docker.service
[Service]
Type=oneshot
ExecStart=/usr/bin/env bash $SELF certsync
EOF
cat > /etc/systemd/system/ergo-certsync.timer <<EOF
[Unit]
Description=Ergo cert sync every 15 minutes
[Timer]
OnCalendar=*:0/15
Persistent=true
[Install]
WantedBy=timers.target
EOF
cat > /etc/systemd/system/ergo-update.service <<EOF
[Unit]
Description=Ergo container updater
After=docker.service
[Service]
Type=oneshot
ExecStart=/usr/bin/env bash $SELF run
EOF
cat > /etc/systemd/system/ergo-update.timer <<EOF
[Unit]
Description=Daily Ergo update check
[Timer]
OnCalendar=daily
Persistent=true
RandomizedDelaySec=1h
[Install]
WantedBy=timers.target
EOF
systemctl daemon-reload
systemctl enable --now ergo-certsync.timer >/dev/null 2>&1 || true
if [[ "$SCHEDULE_UPDATES" == "1" ]]; then systemctl enable --now ergo-update.timer >/dev/null 2>&1 || true
else systemctl disable --now ergo-update.timer >/dev/null 2>&1 || true; fi
log "Scheduled: ergo-certsync.timer (15 min)$( [[ "$SCHEDULE_UPDATES" == "1" ]] && printf ', ergo-update.timer (daily, policy=%s)' "$UPDATE_POLICY" )." ;;
esac
}
do_uninstall() {
[[ $EUID -eq 0 ]] || die "Run as root."
rm -f /etc/periodic/15min/ergo-certsync /etc/periodic/daily/ergo-update
if command -v systemctl >/dev/null 2>&1; then
systemctl disable --now ergo-certsync.timer ergo-update.timer >/dev/null 2>&1 || true
rm -f /etc/systemd/system/ergo-certsync.{timer,service} /etc/systemd/system/ergo-update.{timer,service}
systemctl daemon-reload >/dev/null 2>&1 || true
fi
log "Removed the scheduled Ergo jobs (config kept at $ERGO_UPDATE_CONF)."
}
case "${1:-check}" in
check) do_check ;;
run) do_run ;;
update) do_update ;;
certsync) preflight; do_certsync ;;
caddy) do_caddy ;;
install) do_install ;;
uninstall) do_uninstall ;;
*) die "Usage: update.sh [check|run|update|certsync|caddy|install|uninstall]" ;;
esac
+80 -5
View File
@@ -8,7 +8,9 @@
# 1. Installs docker + docker-cli-compose if missing.
# 2. Lays down docker-compose.yml, Caddyfile, and a substituted
# config.yaml in $STACK_DIR.
# 3. Generates .env on first run; existing .env is never overwritten.
# 3. Generates .env on first run. On a re-run an existing .env is kept,
# except for values passed explicitly that run, which are written
# through (.env is the source of truth for every consumer below).
# 4. Prompts for required values not preset (HEADSCALE_DOMAIN,
# ACME_EMAIL, TAILNET_DOMAIN, POCKETID_DOMAIN, OIDC_CLIENT_ID,
# OIDC_CLIENT_SECRET).
@@ -19,8 +21,9 @@
# 7. Pulls images, brings the stack up, waits for healthchecks.
#
# Idempotent: re-run to apply config changes / pull new images. Re-running
# regenerates config.yaml from the current .env so edits to .env propagate
# (but secrets in .env are never touched once seeded).
# regenerates config.yaml from the current .env, so edits to .env propagate --
# as do values passed to the re-run itself, including a rotated
# OIDC_CLIENT_SECRET. Values not passed are left exactly as they are.
#
# Self-contained: docker-compose.yml, Caddyfile, config.yaml, policy.hujson
# and .env.example are embedded as a base64-encoded tar.gz at the bottom of
@@ -44,6 +47,21 @@ set -euo pipefail
: "${FORCE:=0}"
: "${SKIP_PROMPTS:=0}" # non-interactive: require values via env, no prompts
[[ "$SKIP_PROMPTS" == "1" ]] && FORCE=1
ENV_FILE="$STACK_DIR/.env"
# Everything downstream reads .env (see the `. "$ENV_FILE"` below), so .env has
# to absorb what this run passed or the run silently deploys the old values.
# Which keys actually arrived in the ENVIRONMENT must be recorded BEFORE the
# ":=" defaults below, which make an unset variable look like an empty one --
# and writing a blank over a live OIDC secret is exactly the wrong move.
RUNTIME_KEYS=(HEADSCALE_DOMAIN ACME_EMAIL TAILNET_DOMAIN POCKETID_DOMAIN
OIDC_CLIENT_ID OIDC_CLIENT_SECRET
HEADPLANE_OIDC_CLIENT_ID HEADPLANE_OIDC_CLIENT_SECRET)
KEYS_FROM_ENV=()
for _k in "${RUNTIME_KEYS[@]}"; do
if [[ -n "${!_k+x}" ]]; then KEYS_FROM_ENV+=("$_k"); fi
done
: "${HEADSCALE_DOMAIN:=}"
: "${ACME_EMAIL:=}"
: "${TAILNET_DOMAIN:=}"
@@ -189,6 +207,25 @@ prompt() {
fi
}
# On a re-run the deployed values live in .env. Load them into any key NOT passed
# this run, so the prompts below do not ask six questions whose answers the
# `. "$ENV_FILE"` further down would discard -- and so SKIP_PROMPTS=1 does not
# die demanding values the .env already has.
if [[ -f "$ENV_FILE" ]]; then
_preloaded=()
for _k in "${RUNTIME_KEYS[@]}"; do
if [[ " ${KEYS_FROM_ENV[*]-} " == *" ${_k} "* ]]; then continue; fi
_v=$(sed -n "s/^${_k}=//p" "$ENV_FILE" | tail -n1)
if [[ -z "$_v" ]]; then continue; fi
printf -v "$_k" '%s' "$_v"
_preloaded+=("$_k")
done
if (( ${#_preloaded[@]} > 0 )); then
log "Reusing ${ENV_FILE} values for: ${_preloaded[*]}"
log " (pass VAR=... to change one; it is then written back to .env)"
fi
fi
prompt HEADSCALE_DOMAIN "Public hostname for headscale (e.g. hs.example.com)"
prompt ACME_EMAIL "Let's Encrypt email"
prompt TAILNET_DOMAIN "Tailnet base domain for MagicDNS (e.g. tail.example.com)"
@@ -220,7 +257,28 @@ install -m 0640 "$SCRIPT_DIR/docker-compose.yml" "$STACK_DIR/docker-compose.yml"
install -m 0640 "$SCRIPT_DIR/Caddyfile" "$STACK_DIR/Caddyfile"
# .env
ENV_FILE="$STACK_DIR/.env"
set_env() { # <KEY> <value>: update KEY in .env, or append if absent
# The value goes through the ENVIRONMENT, never interpolated into a sed
# script. Interpolating it corrupts any value containing & (sed expands it to
# the whole match) and aborts the run on one containing the s||| delimiter --
# which is reachable for an OIDC secret, a password or a URL query string.
local key="$1" val="$2" tmp
if [[ ! -f "$ENV_FILE" ]]; then
printf '%s=%s
' "$key" "$val" >> "$ENV_FILE"
return 0
fi
tmp="$(mktemp)"
_SE_KEY="$key" _SE_VAL="$val" awk '
BEGIN { k = ENVIRON["_SE_KEY"]; v = ENVIRON["_SE_VAL"]; seen = 0 }
!seen && index($0, k "=") == 1 { print k "=" v; seen = 1; next }
{ print }
END { if (!seen) print k "=" v }
' "$ENV_FILE" > "$tmp"
cat "$tmp" > "$ENV_FILE" # rewrite in place: keeps the original mode/owner
rm -f "$tmp"
}
if [[ ! -f "$ENV_FILE" ]]; then
log "Seeding $ENV_FILE..."
install -m 0600 "$SCRIPT_DIR/.env.example" "$ENV_FILE"
@@ -236,7 +294,24 @@ if [[ ! -f "$ENV_FILE" ]]; then
-e "s|^HEADPLANE_OIDC_CLIENT_SECRET=.*|HEADPLANE_OIDC_CLIENT_SECRET=${HEADPLANE_OIDC_CLIENT_SECRET}|" \
"$ENV_FILE"
else
log ".env exists; leaving it alone."
# Keep the file, but absorb what was passed this run -- otherwise the
# `. "$ENV_FILE"` below reverts it and the run deploys the old values while
# reporting success. A rotated OIDC_CLIENT_SECRET silently not taking effect
# is the case that matters. Keys not passed are left untouched.
log ".env exists; keeping it (only values passed this run are updated)."
for _k in ${KEYS_FROM_ENV[@]+"${KEYS_FROM_ENV[@]}"}; do
_cur=$(sed -n "s/^${_k}=//p" "$ENV_FILE" | tail -n1)
if [[ "$_cur" == "${!_k}" ]]; then continue; fi
set_env "$_k" "${!_k}"
case "$_k" in
*SECRET*)
log " ${_k}: updated (value changed; not echoed)" ;;
HEADSCALE_DOMAIN)
warn "HEADSCALE_DOMAIN ${_cur} -> ${!_k}: Caddy will request a new Let's Encrypt cert, and the pocket-id redirect URI must become https://${!_k}/oidc/callback" ;;
*)
log " ${_k}: ${_cur:-<unset>} -> ${!_k:-<empty>}" ;;
esac
done
fi
# Validate
-23
View File
@@ -1,23 +0,0 @@
# Copy to .env and fill in. docker compose picks .env up automatically.
# deploy.sh (phase 2) seeds .env from this on first run; an existing .env is
# never overwritten.
# ─── Public hostname ────────────────────────────────────────────────────────
# Bare hostname (no scheme) where Kanrisha will be reached. DNS must point here
# and 80/443 be reachable before deploy, or the Let's Encrypt request fails.
KANRISHA_DOMAIN=kanrisha.example.com
# Email for Let's Encrypt registration / expiry notifications.
ACME_EMAIL=admin@example.com
# ─── ScoutFS ────────────────────────────────────────────────────────────────
# Absolute path to the mounted ScoutFS filesystem Kanrisha operates on. You
# create + mount this between phase 1 and phase 2 (see README). Bind-mounted
# into the Kanrisha container by docker-compose.yml.
SCOUTFS_PATH=/mnt/kanrisha
# ─── Image tags ─────────────────────────────────────────────────────────────
CADDY_TAG=2-alpine
# TODO (part two): set the real Kanrisha image + tag once the compose lands.
# KANRISHA_IMAGE=
# KANRISHA_TAG=latest
-32
View File
@@ -1,32 +0,0 @@
# Caddyfile for the kanrisha stack.
#
# Auto-issues a Let's Encrypt cert for $KANRISHA_DOMAIN and reverse-proxies to
# the Kanrisha app container.
#
# STUB: the upstream below points at the placeholder service on :80. Update the
# reverse_proxy target (service name + port) to match the real Kanrisha service
# in docker-compose.yml when you wire in part two.
{
email {$ACME_EMAIL}
}
{$KANRISHA_DOMAIN} {
encode zstd gzip
reverse_proxy kanrisha:80 { # TODO: real Kanrisha service + port
header_up X-Real-IP {http.request.remote.host}
}
header {
Strict-Transport-Security "max-age=31536000; includeSubDomains"
X-Content-Type-Options "nosniff"
Referrer-Policy "strict-origin-when-cross-origin"
-Server
}
log {
output stdout
format console
}
}
-70
View File
@@ -1,70 +0,0 @@
# kanrisha
Kanrisha (custom application) deployed as a Docker Compose stack behind Caddy,
running on a **ScoutFS** filesystem. **Alma / Rocky Linux only** for now — it
installs ScoutFS (a RHEL-only kernel filesystem) and a host package set via
`dnf`.
> **Status:** the host-prep half (phase 1) is complete; the application stack
> (phase 2) is scaffolded with a **stub** `docker-compose.yml` / `Caddyfile`
> pending the real Kanrisha compose file. See the TODOs in those files.
## Two-phase deploy
Kanrisha runs on a ScoutFS filesystem you create by hand (it depends on your
storage layout), so [`deploy.sh`](deploy.sh) is split into two phases and runs
whichever fits when you invoke it — decided by whether `SCOUTFS_PATH` points at
a live ScoutFS mount.
**Phase 1 — host prep (automated):**
1. `dnf install epel-release`
2. `dnf install git tar sparse mt-st mtx sg3_utils pciutils mbuffer wget curl zstd jq bash-completion sos lsof`
3. Install Docker.
4. Install ScoutFS: `curl -fsSL https://rpm-1.srvno.de/setup-scoutfs.sh | bash -s -- --license <KEY>`
5. Print the steps to create your ScoutFS filesystem, then stop.
**You do (manual):** create + mount a ScoutFS filesystem, e.g. at
`/mnt/kanrisha` (and add it to `/etc/fstab`).
**Phase 2 — deploy the stack (automated):** lay down the compose + Caddyfile,
seed `.env`, open 80/443, pull images, bring the stack up behind Caddy/TLS.
## Required values
| Variable | Phase | Notes |
|----------|-------|-------|
| `SCOUTFS_LICENSE` | 1 | License key for `setup-scoutfs.sh`. Prompted, or set in `globals/globals.env`. |
| `SCOUTFS_PATH` | 2 | Absolute path to the mounted ScoutFS filesystem. Its presence is what triggers phase 2. |
| `KANRISHA_DOMAIN` | 2 | Public hostname (e.g. `kanrisha.example.com`). |
| `ACME_EMAIL` | 2 | Let's Encrypt registration email. |
See [`.env.example`](.env.example) for image tags.
## Deploy
```bash
./automations.sh # Deploy on this host → deploy: kanrisha
```
Or build + run the self-contained artifact:
```bash
./build.sh
scp deploy.sh root@host:
ssh root@host 'bash deploy.sh' # phase 1 (prompts for license)
# ... create + mount your ScoutFS filesystem on the host ...
ssh root@host 'SCOUTFS_PATH=/mnt/kanrisha KANRISHA_DOMAIN=kanrisha.example.com \
ACME_EMAIL=me@example.com bash deploy.sh' # phase 2
```
Unattended provisioning (phase 1): [`cloud-init.yml`](cloud-init.yml).
## Notes
- **Rebuild after editing loose files.** `docker-compose.yml`, `Caddyfile`, and
`.env.example` are embedded into `deploy.sh` by [`build.sh`](build.sh); re-run
it after any edit or the deploy ships the stale copy.
- DNS for `KANRISHA_DOMAIN` must resolve to the host and 80/443 be reachable
before phase 2, or the Let's Encrypt cert request fails.
- The ScoutFS `mkfs`/`mount` recipe in the phase-1 handoff is a placeholder —
fill in your canonical commands.
-44
View File
@@ -1,44 +0,0 @@
#cloud-config
#
# Kanrisha (phase 1) -- harden SSH, then run host prep + ScoutFS install on a
# fresh Alma / Rocky host.
#
# This covers PHASE 1 ONLY. Kanrisha needs a ScoutFS filesystem that you create
# by hand (it depends on your storage layout), so after first boot:
# 1. read the handoff at the end of /var/log/cloud-init-output.log
# 2. create + mount your ScoutFS filesystem (e.g. at /mnt/kanrisha)
# 3. finish with phase 2:
# SCOUTFS_PATH=/mnt/kanrisha KANRISHA_DOMAIN=... ACME_EMAIL=... \
# SKIP_PROMPTS=1 bash /opt/automations/deployments/kanrisha/deploy.sh
#
# Fill in REPO_URL, SCOUTFS_LICENSE, and the values in the runcmd block, then
# paste this as the instance user-data.
packages:
- git
runcmd:
- hostnamectl set-hostname kanrisha || true
- |
set -e
REPO_URL=https://git.anomalous.dev/57_Wolve/automations.git
REPO_BRANCH=main
HARDEN_SSH=1 # harden SSH on this fresh VM (set 0 to skip)
SSH_PORT=22
ALLOWED_IP= # optional: whitelist your client IP in sshguard
SCOUTFS_LICENSE= # <-- your ScoutFS license key (required)
git clone --depth 1 --branch "$REPO_BRANCH" "$REPO_URL" /opt/automations
cd /opt/automations
# Harden SSH on this fresh VM: PQ KEX, key-only auth, sshguard. Seeds root
# from globals/authorized_keys (or SSH_KEYS_URL).
if [ "$HARDEN_SSH" = 1 ]; then
SSH_PORT="$SSH_PORT" ALLOWED_IP="$ALLOWED_IP" SKIP_PROMPTS=1 FORCE=1 \
bash scripts/harden-ssh.sh
fi
# Phase 1: EPEL + host packages + Docker + ScoutFS. With no SCOUTFS_PATH set
# this stops after install and prints the next steps (create the ScoutFS
# filesystem, then run phase 2).
SCOUTFS_LICENSE="$SCOUTFS_LICENSE" SKIP_PROMPTS=1 \
bash deployments/kanrisha/deploy.sh
-374
View File
@@ -1,374 +0,0 @@
#!/usr/bin/env bash
#
# deploy.sh -- deploy Kanrisha (custom app on a ScoutFS filesystem) on
# Alma / Rocky Linux. Two-phase, idempotent, self-contained.
#
# Kanrisha runs on a ScoutFS filesystem that you create by hand (it depends on
# your storage layout), so the deploy splits into two phases and this one
# script runs whichever is appropriate when you invoke it:
#
# PHASE 1 -- host prep (no ScoutFS path yet):
# 1. Installs EPEL + the host package set Kanrisha needs.
# 2. Installs Docker.
# 3. Installs ScoutFS from the vendor repo (setup-scoutfs.sh + license).
# 4. Prints the steps to create your ScoutFS filesystem, then exits.
#
# <you create + mount a ScoutFS filesystem, e.g. at /mnt/kanrisha>
#
# PHASE 2 -- deploy the stack (SCOUTFS_PATH points at that mount):
# 5. Lays down docker-compose.yml + Caddyfile + .env in $STACK_DIR.
# 6. Opens 80/443, pulls images, brings the stack up behind Caddy/TLS.
#
# The phase is chosen automatically: if SCOUTFS_PATH is set and is a live
# ScoutFS mount, it deploys; otherwise it preps the host. Re-run any time.
#
# Self-contained: docker-compose.yml, Caddyfile, .env.example are embedded as a
# base64 tar.gz after __ARCHIVE_BELOW__. Rebuild with build.sh after editing the
# loose source files. (The host-prep phase needs none of them.)
#
# Alma / Rocky (RHEL family) only for now -- it installs ScoutFS (a RHEL-only
# kernel filesystem) and host packages via dnf.
#
# Usage:
# bash deploy.sh # phase 1 (host prep)
# SCOUTFS_LICENSE=... bash deploy.sh # phase 1, license preset
# SCOUTFS_PATH=/mnt/kanrisha KANRISHA_DOMAIN=k.example.com \
# ACME_EMAIL=me@example.com bash deploy.sh # phase 2 (deploy)
# SKIP_PROMPTS=1 ... bash deploy.sh # non-interactive (cloud-init)
set -euo pipefail
: "${STACK_DIR:=/srv/kanrisha}"
: "${SKIP_DOCKER_INSTALL:=0}"
: "${FORCE:=0}"
: "${SKIP_PROMPTS:=0}" # non-interactive: require values via env, no prompts
[[ "$SKIP_PROMPTS" == "1" ]] && FORCE=1
: "${SCOUTFS_LICENSE:=}"
: "${SCOUTFS_PATH:=}"
: "${SCOUTFS_SETUP_URL:=https://rpm-1.srvno.de/setup-scoutfs.sh}"
: "${KANRISHA_DOMAIN:=}"
: "${ACME_EMAIL:=}"
# Host packages Kanrisha needs in place before ScoutFS + the stack.
HOST_PACKAGES=(git tar sparse mt-st mtx sg3_utils pciutils mbuffer wget curl
zstd jq bash-completion sos lsof)
log() { printf '\033[1;32m[+]\033[0m %s\n' "$*"; }
warn() { printf '\033[1;33m[!]\033[0m %s\n' "$*" >&2; }
die() { printf '\033[1;31m[x]\033[0m %s\n' "$*" >&2; exit 1; }
[[ $EUID -eq 0 ]] || die "Run as root."
# ---------------------------------------------------------------------------
# OS detection. This deploy.sh is self-contained (scp'd standalone), so the OS
# logic is inlined here instead of sourced from scripts/oslib.sh.
# ---------------------------------------------------------------------------
osfam() {
local id="" like=""
if [[ -r /etc/os-release ]]; then
id="$(. /etc/os-release 2>/dev/null && echo "${ID:-}")"
like="$(. /etc/os-release 2>/dev/null && echo "${ID_LIKE:-}")"
fi
case " $id $like " in
*" alpine "*) echo alpine ;;
*" debian "*|*" ubuntu "*) echo debian ;;
*" rhel "*|*" fedora "*|*" centos "*) echo rhel ;;
*) echo "${id:-unknown}" ;;
esac
}
require_rhel() {
[[ "$(osfam)" == "rhel" ]] || \
die "Kanrisha supports Alma / Rocky Linux (RHEL family) only for now (detected: $(osfam))."
}
install_docker() {
if command -v docker >/dev/null 2>&1; then
log "Docker already installed: $(docker --version)"
return 0
fi
log "Installing Docker..."
command -v curl >/dev/null 2>&1 || dnf install -y -q curl
curl -fsSL https://get.docker.com | sh
systemctl enable --now docker >/dev/null 2>&1 || systemctl start docker || true
}
# Is the host already prepared (ScoutFS tooling present)?
scoutfs_installed() {
command -v mkfs.scoutfs >/dev/null 2>&1
}
# Is SCOUTFS_PATH a directory sitting on a live ScoutFS mount?
scoutfs_path_ready() {
[[ -n "$SCOUTFS_PATH" && -d "$SCOUTFS_PATH" ]] || return 1
if command -v findmnt >/dev/null 2>&1; then
[[ "$(findmnt -no FSTYPE --target "$SCOUTFS_PATH" 2>/dev/null)" == "scoutfs" ]] || return 1
else
mountpoint -q "$SCOUTFS_PATH" 2>/dev/null || return 1
fi
return 0
}
open_web_ports() {
# Register 80/443 for this stack. Prefer the host firewall
# (harden-firewall.sh) when present; else fall back to firewalld/ufw if
# active (no-op when neither is).
#
# NOTE: Caddy publishes 80/443 via Docker, which reaches the host through
# nat/FORWARD and BYPASSES the INPUT firewall -- so this is belt-and-braces
# for any host-bound bind and self-documents the stack ports.
if [[ -d /etc/firewall/ports.d && -x /usr/local/sbin/firewall-apply ]]; then
log "Registering 80,443/tcp with host firewall..."
printf '80/tcp\n443/tcp\n' > /etc/firewall/ports.d/kanrisha.rule
/usr/local/sbin/firewall-apply
elif command -v firewall-cmd >/dev/null 2>&1 && firewall-cmd --state >/dev/null 2>&1; then
log "firewalld active -- allowing http,https..."
firewall-cmd -q --add-service=http --permanent
firewall-cmd -q --add-service=https --permanent
firewall-cmd -q --reload
elif command -v ufw >/dev/null 2>&1 && ufw status 2>/dev/null | grep -q '^Status: active'; then
log "ufw active -- allowing 80,443/tcp..."
ufw allow 80/tcp >/dev/null; ufw allow 443/tcp >/dev/null
fi
}
prompt() {
local varname="$1" message="$2"
local -n ref="$varname"
if [[ -z "${ref:-}" ]]; then
[[ "$SKIP_PROMPTS" == "1" ]] && die "$varname required (set it in the environment; running with SKIP_PROMPTS=1)."
read -r -p "$message: " ref
[[ -n "$ref" ]] || die "$varname required."
fi
}
# ===========================================================================
# PHASE 1 -- host preparation + ScoutFS
# ===========================================================================
phase1_prep() {
require_rhel
log "Phase 1/2: preparing host (EPEL, packages, Docker, ScoutFS)."
log "Installing EPEL release..."
dnf install -y -q epel-release || die "Failed to install epel-release."
log "Installing ${#HOST_PACKAGES[@]} host packages..."
dnf install -y -q "${HOST_PACKAGES[@]}" || die "Failed to install host packages."
if [[ "$SKIP_DOCKER_INSTALL" != "1" ]]; then
install_docker
fi
if scoutfs_installed; then
log "ScoutFS already installed (mkfs.scoutfs present); skipping vendor setup."
else
if [[ -z "$SCOUTFS_LICENSE" ]]; then
[[ "$SKIP_PROMPTS" == "1" ]] && \
die "SCOUTFS_LICENSE required (set it in the environment or globals.env; running with SKIP_PROMPTS=1)."
read -r -p "ScoutFS license key: " SCOUTFS_LICENSE
[[ -n "$SCOUTFS_LICENSE" ]] || die "SCOUTFS_LICENSE required."
fi
log "Installing ScoutFS from ${SCOUTFS_SETUP_URL}..."
curl -fsSL "$SCOUTFS_SETUP_URL" | bash -s -- --license "$SCOUTFS_LICENSE" \
|| die "ScoutFS setup failed."
scoutfs_installed || \
warn "setup-scoutfs.sh finished but mkfs.scoutfs is not on PATH yet -- check the vendor output."
fi
print_handoff
}
print_handoff() {
cat <<EOF
================================================================
KANRISHA -- HOST PREP COMPLETE (phase 1 of 2)
Installed: EPEL, host packages, Docker, and ScoutFS.
ScoutFS is ready to use.
NEXT -- create the filesystem Kanrisha will run on:
1. Identify the device(s) for ScoutFS:
lsblk
2. Make a ScoutFS filesystem (your metadata + data devices):
mkfs.scoutfs ... # <-- your canonical recipe
3. Mount it, and add it to /etc/fstab so it survives reboot:
mkdir -p /mnt/kanrisha
mount -t scoutfs ... /mnt/kanrisha
THEN -- finish the deploy (phase 2):
SCOUTFS_PATH=/mnt/kanrisha \\
KANRISHA_DOMAIN=kanrisha.example.com \\
ACME_EMAIL=you@example.com \\
bash deploy.sh
(or re-run ./automations.sh, pick "deploy: kanrisha", and give the path)
================================================================
EOF
}
# ===========================================================================
# PHASE 2 -- deploy the stack
# ===========================================================================
phase2_deploy() {
require_rhel
log "Phase 2/2: deploying the Kanrisha stack on ${SCOUTFS_PATH}."
prompt KANRISHA_DOMAIN "Public hostname for Kanrisha (e.g. kanrisha.example.com)"
prompt ACME_EMAIL "Let's Encrypt email"
if [[ "$SKIP_DOCKER_INSTALL" != "1" ]]; then
install_docker
fi
command -v docker >/dev/null 2>&1 || \
die "Docker is not installed (run phase 1 first, or unset SKIP_DOCKER_INSTALL)."
# --- extract embedded archive ---
SCRIPT_DIR=$(mktemp -d -t kanrisha-deploy.XXXXXX)
trap 'rm -rf "$SCRIPT_DIR"' EXIT
if grep -q -a '^__ARCHIVE_BELOW__$' "$0"; then
log "Extracting embedded deployment files..."
grep -a -A 9999999 '^__ARCHIVE_BELOW__$' "$0" | tail -n +2 | base64 -d | tar -xz -C "$SCRIPT_DIR"
else
die "No embedded archive found. Run build.sh to embed deployment files."
fi
for f in docker-compose.yml Caddyfile .env.example; do
[[ -f "$SCRIPT_DIR/$f" ]] || die "Embedded archive missing $f"
done
open_web_ports
log "Setting up $STACK_DIR..."
install -d -m 0750 "$STACK_DIR"
install -m 0640 "$SCRIPT_DIR/docker-compose.yml" "$STACK_DIR/docker-compose.yml"
install -m 0640 "$SCRIPT_DIR/Caddyfile" "$STACK_DIR/Caddyfile"
ENV_FILE="$STACK_DIR/.env"
if [[ ! -f "$ENV_FILE" ]]; then
log "Seeding $ENV_FILE..."
install -m 0600 "$SCRIPT_DIR/.env.example" "$ENV_FILE"
sed -i \
-e "s|^KANRISHA_DOMAIN=.*|KANRISHA_DOMAIN=${KANRISHA_DOMAIN}|" \
-e "s|^ACME_EMAIL=.*|ACME_EMAIL=${ACME_EMAIL}|" \
-e "s|^SCOUTFS_PATH=.*|SCOUTFS_PATH=${SCOUTFS_PATH}|" \
"$ENV_FILE"
else
log ".env exists; leaving it alone."
fi
missing=()
for var in KANRISHA_DOMAIN ACME_EMAIL SCOUTFS_PATH; do
grep -E "^${var}=.+$" "$ENV_FILE" >/dev/null || missing+=("$var")
done
(( ${#missing[@]} == 0 )) || die "Missing values in $ENV_FILE: ${missing[*]}"
if [[ "$FORCE" != "1" ]]; then
cat <<EOF
About to pull images and start the stack from $STACK_DIR.
Caddy will request a Let's Encrypt cert for ${KANRISHA_DOMAIN}. DNS for
that name must already point at this host, and ports 80/443 must be
reachable from the internet.
Continue? [y/N]
EOF
read -r ans
[[ "${ans,,}" == "y" || "${ans,,}" == "yes" ]] || { warn "Aborted."; exit 0; }
fi
cd "$STACK_DIR"
log "Pulling images..."
docker compose pull
log "Starting stack..."
docker compose up -d --remove-orphans
log "Waiting for services to become healthy (up to 120s)..."
deadline=$(( $(date +%s) + 120 ))
while (( $(date +%s) < deadline )); do
status=$(docker compose ps --format '{{.Service}} {{.Health}}' 2>/dev/null || true)
unhealthy=$(echo "$status" | awk '$2 != "healthy" && $2 != "" {print $1}')
if [[ -z "$unhealthy" && -n "$status" ]]; then
log "All services healthy."
break
fi
sleep 5
done
echo
log "Stack status:"
docker compose ps
echo
cat <<EOF
================================================================
KANRISHA DEPLOYED
URL: https://${KANRISHA_DOMAIN}
Stack dir: ${STACK_DIR}
ScoutFS: ${SCOUTFS_PATH}
Manage:
docker compose logs -f
docker compose pull && docker compose up -d # update
docker compose down # stop, keep volumes
docker compose down -v # stop, WIPE caddy volumes
Or just re-run this script -- it's idempotent.
NOTE: the Kanrisha service in docker-compose.yml is a STUB pending the real
compose file. Drop it in, run build.sh, then re-run this deploy.
================================================================
EOF
}
# ===========================================================================
# Dispatch -- phase 2 if a live ScoutFS mount was given, else phase 1.
# ===========================================================================
if scoutfs_path_ready; then
phase2_deploy
else
if [[ -n "$SCOUTFS_PATH" ]]; then
warn "SCOUTFS_PATH='$SCOUTFS_PATH' is not a live ScoutFS mount -- running host prep (phase 1)."
fi
phase1_prep
fi
# IMPORTANT: do not put any code below this exit. Everything after the
# __ARCHIVE_BELOW__ marker is the embedded tar.gz payload (base64).
exit 0
__ARCHIVE_BELOW__
H4sIAAAAAAAAA+1YbW/jNhLOZ/2KgRPgHHQlv+VlT70U503SNuhms4i9uBaHg0tLtE1EIlWSiuMG
Bvoj7hfeL+kMZb3Em9t+uHa7e/AksCS+DMmZ4TMPGavojms/UmmmDA9WabL3u0sX5eToaK970u2d
Hvfcs9vtuXKUQX9wutc77vcGJ8eD/vHJXrc36Hb7e9D9/afyvuTGMg2wd3w6+YdK7vnHGPMTkn24
Y1ILs2CAhojuwPchYnG8gvb49egQvqjr2xHaSqXAsgyUhFGkcvv16DDw9r19GI3fvQrBLjhozhL4
rlLK9b2IOAgDmVb3IuYxCAkZ0xbsUgUwXlBVwiK+UEnMNeqKVZSnXFrj9PGHjEcWu6G6jINRMM1F
EgdmAQtm8DvldiHkHKwCnk6xIZMxZFjHoY/a1AxiniVqVfZgcM8SEcMm5mEmEk6dcf3YMc/gnNYf
wC1306JJoJofSzv8WK1pKeyiXnLMZ0IKK5R8QYUSS32dy2q2hZ3GKlOJmq9CfAd4pdUStYH/VWH0
8OhoQB/lWOHfMqXtV65nZVKVcc0sN+QEGn2hjP2LKf3hlmNWxvL0BUyFjP1U5ZLsh7OM/aUWlpbD
LBw8js5v3o2/Hk3eDsffrtG/2IIaTlewUjlMuV1yXEdhyl7TrOh0T7KUh9VMPe9eJeg1E3pQrMWP
mWX1V6TkTMxDz9tYr25ILwAiZXPUV5QcPJ4PLy5+mIyH34R+32dJJiRfu3aoxzL80pNiAq6Dq9Gc
trINIZdoAeNjsGYZj10dmdEUAwH40HrZDV92W/U32p1sX5Q0VlJUBx0XEmTZsMNt1HGDNgq1qto2
1t6h362KjRk6xdNVcnkvtJIU8eWQ3w3f3F6Nvh1OLm6uh1dvQmgdPG6VrcvZD8+vLyeXWPTaNas/
Ny0w+LmMzUTJekGV1+hzgdFrF9GCR3dlC4wutOM/W+fXF60X0FrOuaWn/9ONT8+FtVnY6fT6p0EX
/3phv9v762ZJnda/NkoERp3GrRbCoGtKxSLlGKYhHJclmlst0Ngw2BQ4J04wxoWKQ+hhV6zAjXNz
cQPtEjcOQ+xYbk8EEGPz6dZ+3IagwKmh1btqF28Ib5bNoV3ZFgOO4ClAnxxW7UklBVABDIhDKk9o
O91zbfgEYe1hRfjRzjN0OIcqLLBQ1VqYXJGroQOFofClYXk3CCGr5Dwm21QQ0NwdjSi4uh5+cxn6
VjPEnbvOcqFYKtZho4XbPAkhhV0DzYIwutD07E56EhQf2kxb+2P/GeiploMI6ICqTSiFsI8emmlM
Is7CQb0Bt8Co2Dst789Oj//3UgXrHzjGb/A/on5b/K/XOzrZ8b+PIfsNuJop7fbtU0ZY8JZhbpUv
jMk5EajXnBjHpYz0KrMQccRG6nywlaMcZdjgpE84iUCPoIjqaJgKoAknKjDappN5ZiwCdop0JFFL
xGFBxBDpi0PlmjdWrAzBBnN7AO8KNC7Y2xZYM40ZDdplF8I/TAWE8YeE5Cmz0QcyCerDHBG/d3KC
JdE+ok5LofkTmut5j0WuT5lI4PGgkaW9NVZuW24Nm/YyUjGHn42NYf6zyLwNODdXU2WKl114hDJX
hv+FiBfL3ACvS/5ovAmy3u/9W+zgX72FR8rvgeY/obMtPlNleUD4XTCwtefVPTfzJBlhIo+sP9ZM
GhrDH/EoR7q5glbKHnxMO2eD3vHgBLf8l2icKMljPsqnFwptIk2r0vO9f47BgGzIH68y7t9kxKkN
1bSkMlLMZnXbWz7jWqMb3qpERCuopWWK6Sgt5kL65Bo/0gpTWVFS68B5arRnc23I0BsLw+SW5RZ3
Q4xvVSkGPMYJBa5RCd/0Xn9uGYvycMAfWJr9cSngN/C/f4Jgv4X/x6eDHf5/DEH8V5mjrxQJDrAx
FyA9lsEG4apzciaiO1M0Q7hgmBAw/kXEkmQV0Km9OmS3y1Migg6S2aKLY32OqyNAz4RGNojc8Esc
EU/4wlg6wbuGwqAySQgHCn/owIpggAi6D//59y/FP7zNp7jf3dHXgXdV84n/4yJeMc3ribelAoMn
gJQfUvrQjbS4JD9MXQ7CBnEAF29GkOZEoykJArWmkzy67GW3QzcHZWM2xWw+5YhQfOOWF7DJ7U8z
9wbjYYZZyQTeVhI6KzNLCRABhgL54dJlMcr42+rm6EnNCLDxeMMfMqFXIJUVM4wTB+OBV6e+Mxan
Qv59S3nt5PJY8ad77X93+nCKSSJHOuJOQLjbyBflxcwzx6fnLnsC+EHlqKu4psFE7voXe+oDdzVE
czjcXg4vri8PA3jVuBJyRGYzmWrEionRPdD7JCfwmge1s04qbae+BGq678qdsPF8bT5XD3rVLdRZ
eQnlPXMTYbh9hiw2LxiULO4RKyxN0EGGUPPpef6sWULDFsf3z41U7GQnO9nJTnayk53sZCc72clO
drKTT0p+BQmmdmwAKAAA
-54
View File
@@ -1,54 +0,0 @@
# kanrisha stack -- caddy (TLS) + kanrisha (custom app on ScoutFS).
#
# STUB: the real Kanrisha service is provided in part two. This placeholder
# documents the expected shape so build.sh has something to embed and phase 2
# of deploy.sh has a valid compose file to stand up Caddy. Replace the
# `kanrisha` service with the real definition, then re-run build.sh.
#
# Topology:
# Browser -> caddy:443 -> kanrisha:<port>
#
# Kanrisha operates on the host's ScoutFS filesystem, bind-mounted read-write
# at ${SCOUTFS_PATH} (created by you between phase 1 and phase 2).
name: kanrisha
volumes:
caddy-data:
caddy-config:
services:
caddy:
image: caddy:${CADDY_TAG:-2-alpine}
container_name: caddy
restart: unless-stopped
ports:
- "80:80"
- "443:443"
volumes:
- ./Caddyfile:/etc/caddy/Caddyfile:ro
- caddy-data:/data
- caddy-config:/config
environment:
KANRISHA_DOMAIN: "${KANRISHA_DOMAIN}"
ACME_EMAIL: "${ACME_EMAIL}"
depends_on:
- kanrisha
healthcheck:
test: ["CMD", "wget", "-qO-", "http://127.0.0.1:2019/config/"]
interval: 30s
timeout: 5s
retries: 3
start_period: 10s
# TODO (part two): replace this stub with the real Kanrisha service.
# - real image + tag (KANRISHA_TAG in .env)
# - the port Caddy should reverse_proxy to (update Caddyfile too)
# - any env / config / healthcheck the app needs
kanrisha:
image: ${KANRISHA_IMAGE:-traefik/whoami}:${KANRISHA_TAG:-latest} # STUB image
container_name: kanrisha
restart: unless-stopped
volumes:
# ScoutFS filesystem the app runs on (host path from .env).
- "${SCOUTFS_PATH}:/data"
+36
View File
@@ -0,0 +1,36 @@
# Copy to /etc/knot/.env (0640 root:knot). deploy.sh seeds it on first run and,
# on a re-run, updates only the keys you passed that run. Never commit it.
#
# knot.conf is re-rendered from these values on every deploy, so correcting one
# here (or passing it to a re-run) actually takes effect.
# ─── Identity ───────────────────────────────────────────────────────────────
# Shown as server.identity and server.nsid, and in the MOTD. Use the estate's
# node naming, e.g. ANYCAST-DNS-3.
NODE_ID=
# primary | secondary
# primary -- holds zone files, signs with automatic DNSSEC, owns the KASP
# database, notifies the secondaries. Exactly one node.
# secondary -- receives zones by AXFR/IXFR with TSIG and signs nothing.
ROLE=secondary
# ─── Listening ──────────────────────────────────────────────────────────────
# Substituted into knot.conf's `listen:`. Comma-separated addr@port entries.
# Narrow this to the node's service address if it should not answer on every
# interface (an anycast node usually should not).
LISTEN=0.0.0.0@53, ::@53
# ─── Replication ────────────────────────────────────────────────────────────
# secondary only: the primary this node transfers from. deploy.sh seeds a
# minimal remotes.conf from it so the node can bootstrap by AXFR before the
# `dns` repo's pipeline has ever run. The pipeline overwrites that file later.
PRIMARY_ADDR=
# ─── MOTD ───────────────────────────────────────────────────────────────────
# Data centre label for the login banner. Cosmetic.
DATACENTER=
# Peers for `knsctl serials` / the knsers alias -- space-separated addresses of
# the other nodes in the estate. Written to /etc/profile.d/knotdns-peers.sh.
PEERS=
+5
View File
@@ -0,0 +1,5 @@
# Never commit the rendered TSIG key material or a populated .env. secrets.conf
# holds the shared secrets every node's AXFR/NOTIFY authenticates with; leaking
# it lets anyone transfer -- or spoof a NOTIFY for -- every zone in the estate.
secrets.conf
.env
+134
View File
@@ -0,0 +1,134 @@
# knot-dns
Authoritative DNS node — [Knot DNS](https://www.knot-dns.cz/) on Alpine, native
(no container), for a global anycast estate with automatic DNSSEC.
**Exception to the repo norm**, alongside `squid` and `openbao`: no Docker, no
Caddy, no Let's Encrypt. Knot binds :53 directly, needs real client addresses
for RRL and DNS cookies, and its DNSSEC key store must live on the host
filesystem. Containerising it buys nothing and costs the host firewall.
## What this deploys, and what it does not
This stands up **a node**. It does not manage zones — that is the
[`dns`](../../../dns) repo, and its pipeline delivers zone data here.
The split point is `/etc/knot/knot.conf`. This deployment writes it once as a
skeleton of `include:` lines covering only what belongs to a *box*: identity,
NSID, storage paths, listen addresses, logging, control socket. Everything that
belongs to *DNS policy* — templates, DNSSEC policy, remotes/ACLs, modules, the
domain inventory, the zone files — arrives from the `dns` repo.
| Deployed here, once | Delivered by the `dns` repo, continuously |
|---|---|
| `knot.conf` skeleton | `templates.conf`, `policy.conf`, `remotes.conf`, `modules.conf` |
| `secrets.conf` (TSIG, rendered locally, never in git) | `dnssec.conf`, `public.conf`, `arpa.conf` |
| `knsctl`, aliases, MOTD, `zone.tmpl` | `zones/**/*.zone` |
| packages, users, directories, firewall | |
## Files
| File | Purpose |
|---|---|
| `deploy.sh` | The installer. Alpine only, native. Self-contained: the files below are embedded: rebuild with `build.sh` after editing them. |
| `.env.example` | Seeds `/etc/knot/.env`. `knot.conf` is re-rendered from it every run. |
| `cloud-init.yml` | Fresh-VM user-data: harden SSH, host firewall, then deploy. |
| `knot.conf` | Bootstrap skeleton. `@NODE_ID@` / `@LISTEN@` substituted from `.env`. |
| `secrets.conf.example` | TSIG `key:` block. Rendered to `/etc/knot/secrets.conf`, `0640 root:knot`. Never committed. |
| `knsctl` | Admin CLI. Replaces `adddns.pl` / `adddnssec.pl` and fixes four defects in them. |
| `knotdns-aliases.sh` | `/etc/profile.d/` — the existing muscle memory, with `-b` added everywhere. |
| `knot-dns.motd` | The cheat-sheet MOTD. |
| `zone.tmpl` | New-zone skeleton (mirrors `dns/skeleton/zone.tmpl`). |
## What deploy.sh does on a fresh node
Two things are worth knowing before the first run.
**The include chain is stubbed.** `knot.conf` `include:`s seven files this
deployment does not own (`remotes`, `policy`, `modules`, `templates`, `arpa`,
`dnssec`, `public`). Knot treats a missing include as a config error, so a node
the `dns` pipeline has never delivered to would fail `conf-check` and never
start. `deploy.sh` writes a placeholder for each one that is absent — and only
for those, so a delivered file is never clobbered. The node therefore comes up
healthy, serving no zones, until the pipeline lands.
**TSIG keys are generated on a primary and copied to a secondary.** They have to
match byte for byte or nothing authenticates, so a secondary cannot invent its
own. `ROLE=primary` generates both keys, writes `secrets.conf` and prints them
once at the end of the run — capture them. `ROLE=secondary` requires
`TSIG_AUTHORITIVE` and `TSIG_ADMIN` and refuses to deploy without them. An
existing `secrets.conf` is never rewritten, so a re-run cannot rotate a key out
from under a running estate.
A secondary also gets a minimal `remotes.conf` seeded from `PRIMARY_ADDR`, so it
can bootstrap by AXFR before the pipeline has ever run. That seed is written
only when `remotes.conf` was absent; once the pipeline owns the file, re-running
leaves it alone.
Re-running is safe and is how you change things: `knot.conf` is re-rendered from
`.env` every time, and values passed to the re-run are written back to `.env`
first, so a corrected `NODE_ID` or `LISTEN` actually takes effect. Values you do
not pass are left as they are.
## Roles
`ROLE` in `.env` selects what the node is:
- **`primary`** — holds zone files, signs with automatic DNSSEC, owns the KASP
database, notifies the secondaries. Exactly one node. Not necessarily
public-facing: a hidden (stealth) primary is the recommended shape.
- **`secondary`** — receives zones by AXFR/IXFR with TSIG, serves queries,
signs nothing. `dnssec-signing` must be **off**; a secondary serves data that
is already signed.
Adding a public node is a `secondary` deploy plus one address in the primary's
`remotes.conf`. With catalog zones configured, the new node self-populates.
## What survives from the previous setup
Everything an operator types. `knrl`, `knsc`, `knzc`, `knpbz`, `knsec`,
`knarpa`, `ozf` all still work, from the same paths, on the same box. Three
corrections were folded in:
1. **`-b` on every triggering `knotc` command.** Without it `knotc` returns OK
when the command was *sent*, not when it succeeded, so a rejected zone file
reports green. `knrl` was a bare `knotc reload`.
2. **`knzr` (`zone-reload`) added** next to `knrl` (`reload`). Reloading one
zone's data is the right verb for a record change — smaller blast radius,
and a parse error in one zone cannot disturb the others. `knrl` reloads
configuration and is only needed when a zone is added or removed.
3. **`knsctl` replaces the Perl scripts**, whose duplicate check searched for
the domain in BIND `named.conf` double-quote syntax against unquoted YAML
and therefore never matched; which never consulted the other class's
manifest, so a domain in `public.conf` could be appended to `dnssec.conf`
and fail the reload after both files were already written; and which
reloaded without validating.
## Key material
`/var/lib/knot/keys` (the KASP LMDB) is node-local to the primary and
replicates nowhere. Losing it is the one unrecoverable failure in this system:
Knot will happily generate fresh KSKs, every published DS will point at keys
that no longer exist, and every signed domain goes bogus until each registrar
is updated by hand.
```sh
knotc -b zone-backup +backupdir /var/backups/knot/$(date -u +%FT%H) +journal
```
**`+journal` is not the default** — the documented default filter set excludes
it, and in this configuration the journal holds the only durable copy of the
signed zone and the last real SOA serial. A default `zone-backup` captures your
keys and loses your zone data. Restore order is config → KASP + journal → *then*
start `knotd`, never the reverse.
## Usage
```sh
# interactive
bash deploy.sh
# non-interactive (this is what cloud-init.yml does)
ROLE=secondary NODE_ID=ANYCAST-DNS-3 PRIMARY_ADDR=10.1.24.64 \
SKIP_PROMPTS=1 bash deploy.sh
```
+45
View File
@@ -0,0 +1,45 @@
#!/usr/bin/env bash
#
# build.sh -- (re)embed knot.conf, secrets.conf.example, knsctl, zone.tmpl,
# knotdns-aliases.sh, knot-dns.motd and .env.example into deploy.sh as a
# base64-encoded tar.gz payload after __ARCHIVE_BELOW__. Idempotent: strips any
# existing payload first.
set -euo pipefail
DIR=$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)
SCRIPT="$DIR/deploy.sh"
MARKER="__ARCHIVE_BELOW__"
FILES=(
knot.conf
secrets.conf.example
knsctl
zone.tmpl
knotdns-aliases.sh
knot-dns.motd
.env.example
)
[[ -f "$SCRIPT" ]] || { echo "deploy.sh not found at $SCRIPT" >&2; exit 1; }
for f in "${FILES[@]}"; do
[[ -f "$DIR/$f" ]] || { echo "Missing $DIR/$f" >&2; exit 1; }
done
PAYLOAD=$(tar -czf - -C "$DIR" "${FILES[@]}" | base64)
TMP=$(mktemp)
trap 'rm -f "$TMP"' EXIT
sed "/^${MARKER}\$/,\$d" "$SCRIPT" > "$TMP"
{
echo "$MARKER"
echo "$PAYLOAD"
} >> "$TMP"
mv "$TMP" "$SCRIPT"
chmod +x "$SCRIPT"
trap - EXIT
size=$(wc -c < "$SCRIPT")
echo "Built $SCRIPT (${size} bytes)"
+55
View File
@@ -0,0 +1,55 @@
#cloud-config
#
# Authoritative Knot DNS node — harden SSH, then deploy, on a fresh Alpine host.
# Native (no container): Knot binds :53 directly. Alpine only.
#
# Fill in REPO_URL and the values in the runcmd block, then paste this as the
# instance user-data.
#
# A SECONDARY needs the primary's TSIG secrets, byte for byte, or nothing will
# ever transfer — take them from the primary's deploy summary (or its
# /etc/knot/secrets.conf) and paste them below. A PRIMARY generates its own and
# prints them once; capture them from the console before they scroll away.
packages:
- git
runcmd:
- hostnamectl set-hostname anycast-dns-3 || true
- |
set -e
REPO_URL=https://git.anomalous.dev/57_Wolve/automations.git
REPO_BRANCH=main
HARDEN_SSH=1 # harden SSH on this fresh VM (set 0 to skip)
SSH_PORT=22
ALLOWED_IP= # optional: whitelist your client IP in sshguard
git clone --depth 1 --branch "$REPO_BRANCH" "$REPO_URL" /opt/automations
cd /opt/automations
# Harden SSH: PQ KEX, key-only auth, sshguard. Seeds root from
# globals/authorized_keys (or SSH_KEYS_URL).
if [ "$HARDEN_SSH" = 1 ]; then
SSH_PORT="$SSH_PORT" ALLOWED_IP="$ALLOWED_IP" SKIP_PROMPTS=1 FORCE=1 \
bash scripts/harden-ssh.sh
fi
# Host firewall, so deploy.sh's 53/tcp+udp drop-in has somewhere to register.
SKIP_PROMPTS=1 FORCE=1 bash scripts/harden-firewall.sh || true
# Deploy the node. ROLE=primary holds and signs the zones; ROLE=secondary
# transfers them in and signs nothing. NODE_ID becomes server.identity and
# NSID, so make it the estate's node name.
ROLE=secondary \
NODE_ID=ANYCAST-DNS-3 \
PRIMARY_ADDR=10.1.24.64 \
LISTEN='0.0.0.0@53, ::@53' \
DATACENTER='Stockholm SE' \
PEERS='10.1.24.64 10.1.24.68' \
TSIG_AUTHORITIVE=REPLACE_WITH_THE_PRIMARYS_SECRET \
TSIG_ADMIN=REPLACE_WITH_THE_PRIMARYS_SECRET \
SKIP_PROMPTS=1 \
bash deployments/knot-dns/deploy.sh
# The node now serves NO zones — DNS policy and zone data are delivered by
# the `dns` repo's pipeline. Add this node there (and to the primary's
# remotes.conf, unless a catalog zone picks it up) to put it in rotation.
+551
View File
@@ -0,0 +1,551 @@
#!/usr/bin/env bash
#
# deploy.sh -- deploy an authoritative Knot DNS node on Alpine, native (no
# container). Single-purpose box, runs as root.
#
# What this does:
# 1. Installs knot + knot-utils and creates the knot user/directories.
# 2. Seeds /etc/knot/.env on first run; on a re-run updates only the values
# passed that run, and re-renders knot.conf from it either way.
# 3. Renders /etc/knot/knot.conf from the skeleton (@NODE_ID@, @LISTEN@).
# 4. TSIG: ROLE=primary generates secrets.conf and prints the keys once;
# ROLE=secondary REQUIRES them to be supplied, because they must byte-match
# the primary's or no transfer will ever authenticate.
# 5. Stubs the seven policy files knot.conf includes, so conf-check passes and
# knotd starts clean on a box the `dns` repo has never delivered to.
# 6. Installs knsctl, zone.tmpl, the login aliases and the MOTD.
# 7. Opens 53/tcp + 53/udp on the host firewall.
# 8. conf-check, then enables and starts knot.
#
# What this does NOT do: manage zones or DNS policy. Those come from the `dns`
# repo's pipeline -- see the README's split table. This script never overwrites
# a policy file the pipeline owns, nor an existing secrets.conf.
#
# Idempotent: re-run to change identity/listen/peers or pull a newer knot.
#
# Self-contained: knot.conf, secrets.conf.example, knsctl, zone.tmpl,
# knotdns-aliases.sh, knot-dns.motd and .env.example are embedded as a base64
# tar.gz at the bottom. Rebuild with build.sh after editing the loose files.
#
# Usage:
# ROLE=primary NODE_ID=ANYCAST-DNS-1 bash deploy.sh
# ROLE=secondary NODE_ID=ANYCAST-DNS-3 PRIMARY_ADDR=10.1.24.64 \
# TSIG_AUTHORITIVE=... TSIG_ADMIN=... SKIP_PROMPTS=1 bash deploy.sh
set -euo pipefail
: "${CONF_DIR:=/etc/knot}"
: "${DATA_DIR:=/var/lib/knot}"
: "${SKIP_PROMPTS:=0}" # non-interactive: require values via env, no prompts
: "${FORCE:=0}"
[[ "$SKIP_PROMPTS" == "1" ]] && FORCE=1
ENV_FILE="$CONF_DIR/.env"
# Which keys arrived in this script's ENVIRONMENT has to be recorded BEFORE the
# ":=" defaults below, which make an unset variable look like an empty one --
# and a blank written over a live TSIG secret would break every transfer.
RUNTIME_KEYS=(NODE_ID ROLE LISTEN PRIMARY_ADDR DATACENTER PEERS)
KEYS_FROM_ENV=()
for _k in "${RUNTIME_KEYS[@]}"; do
if [[ -n "${!_k+x}" ]]; then KEYS_FROM_ENV+=("$_k"); fi
done
: "${NODE_ID:=}"
: "${ROLE:=}"
: "${LISTEN:=}"
: "${PRIMARY_ADDR:=}"
: "${DATACENTER:=}"
: "${PEERS:=}"
: "${TSIG_AUTHORITIVE:=}" # secondary: must match the primary's
: "${TSIG_ADMIN:=}"
LISTEN_DEFAULT='0.0.0.0@53, ::@53'
log() { printf '\033[1;32m[+]\033[0m %s\n' "$*"; }
warn() { printf '\033[1;33m[!]\033[0m %s\n' "$*" >&2; }
die() { printf '\033[1;31m[x]\033[0m %s\n' "$*" >&2; exit 1; }
[[ $EUID -eq 0 ]] || die "Run as root."
# Alpine only, deliberately: Knot binds :53 directly and keeps its DNSSEC key
# store on the host filesystem, and this estate's nodes are Alpine. The RHEL
# packaging needs EPEL, which nothing in this repo sets up, so fail loudly
# rather than half-install somewhere untested.
_osid="$( [[ -r /etc/os-release ]] && . /etc/os-release 2>/dev/null && echo "${ID:-}" )"
[[ "$_osid" == alpine ]] || \
die "This deployment is Alpine-only (detected: ${_osid:-unknown}). Knot on other distros is not packaged or tested here."
# ---------------------------------------------------------------------------
# Extract embedded archive (falls back to the source dir before build.sh runs)
# ---------------------------------------------------------------------------
SCRIPT_DIR=$(mktemp -d -t knot-deploy.XXXXXX)
trap 'rm -rf "$SCRIPT_DIR"' EXIT
grep -a -A 9999999 '^__ARCHIVE_BELOW__$' "$0" \
| tail -n +2 \
| base64 -d \
| tar -xz -C "$SCRIPT_DIR" 2>/dev/null || true
if [[ ! -f "$SCRIPT_DIR/knot.conf" ]]; then
SRC=$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)
for f in knot.conf secrets.conf.example knsctl zone.tmpl knotdns-aliases.sh knot-dns.motd .env.example; do
[[ -f "$SRC/$f" ]] || die "Missing $f (no embedded payload and not in $SRC -- run build.sh)."
cp "$SRC/$f" "$SCRIPT_DIR/$f"
done
fi
# ---------------------------------------------------------------------------
# .env: seed once, then update only what this run passed
# ---------------------------------------------------------------------------
set_env() { # <KEY> <value>: update KEY in .env, or append if absent
# The value goes through the ENVIRONMENT, never interpolated into a sed
# script: a TSIG secret or a listen list containing & or the s||| delimiter
# would otherwise be silently corrupted or abort the run.
local key="$1" val="$2" tmp
if [[ ! -f "$ENV_FILE" ]]; then
printf '%s=%s\n' "$key" "$val" >> "$ENV_FILE"
return 0
fi
tmp="$(mktemp)"
_SE_KEY="$key" _SE_VAL="$val" awk '
BEGIN { k = ENVIRON["_SE_KEY"]; v = ENVIRON["_SE_VAL"]; seen = 0 }
!seen && index($0, k "=") == 1 { print k "=" v; seen = 1; next }
{ print }
END { if (!seen) print k "=" v }
' "$ENV_FILE" > "$tmp"
cat "$tmp" > "$ENV_FILE" # rewrite in place: keeps the original mode/owner
rm -f "$tmp"
}
install -d -m 0755 "$CONF_DIR"
# On a re-run the deployed values live in .env. Load them into any key NOT
# passed this run, so the prompts below do not demand a full retype and
# SKIP_PROMPTS=1 does not die over values the node already has.
if [[ -f "$ENV_FILE" ]]; then
_preloaded=()
for _k in "${RUNTIME_KEYS[@]}"; do
if [[ " ${KEYS_FROM_ENV[*]-} " == *" ${_k} "* ]]; then continue; fi
_v=$(sed -n "s/^${_k}=//p" "$ENV_FILE" | tail -n1)
if [[ -z "$_v" ]]; then continue; fi
printf -v "$_k" '%s' "$_v"
_preloaded+=("$_k")
done
if (( ${#_preloaded[@]} > 0 )); then
log "Reusing ${ENV_FILE} values for: ${_preloaded[*]}"
log " (pass VAR=... to change one; it is then written back to .env)"
fi
fi
prompt() {
local varname="$1" message="$2"
local -n ref="$varname"
if [[ -z "${ref:-}" ]]; then
[[ "$SKIP_PROMPTS" == "1" ]] && die "$varname required (set it in the environment; running with SKIP_PROMPTS=1)."
read -r -p "$message: " ref
[[ -n "$ref" ]] || die "$varname required."
fi
}
prompt NODE_ID "Node ID for server.identity / NSID (e.g. ANYCAST-DNS-3)"
prompt ROLE "Role: primary or secondary"
case "$ROLE" in
primary|secondary) ;;
*) die "ROLE must be 'primary' or 'secondary' (got '$ROLE')." ;;
esac
[[ -n "$LISTEN" ]] || LISTEN="$LISTEN_DEFAULT"
if [[ "$ROLE" == secondary ]]; then
prompt PRIMARY_ADDR "Address of the primary this node transfers from"
fi
# ---------------------------------------------------------------------------
# Packages, user, directories
# ---------------------------------------------------------------------------
log "Installing Knot DNS..."
apk add -q knot knot-utils
# The package normally creates knot:knot; make that certain before any chown.
getent group knot >/dev/null 2>&1 || addgroup -S knot
getent passwd knot >/dev/null 2>&1 || adduser -S -D -H -G knot -s /sbin/nologin knot
# storage/journal/kasp/timer/catalog as declared in knot.conf's database:
# section, plus the zone root knsctl writes into.
for d in database journal keys timer catalog zones; do
install -d -m 0750 -o knot -g knot "$DATA_DIR/$d"
done
install -d -m 0750 -o knot -g knot /run/knot
# ---------------------------------------------------------------------------
# knot.conf -- re-rendered every run, since it is derived purely from .env
# ---------------------------------------------------------------------------
log "Rendering $CONF_DIR/knot.conf (NODE_ID=$NODE_ID, ROLE=$ROLE)..."
_KC_NODE_ID="$NODE_ID" _KC_LISTEN="$LISTEN" awk '
{ gsub(/@NODE_ID@/, ENVIRON["_KC_NODE_ID"]); gsub(/@LISTEN@/, ENVIRON["_KC_LISTEN"]); print }
' "$SCRIPT_DIR/knot.conf" > "$CONF_DIR/knot.conf.new"
install -m 0640 -o root -g knot "$CONF_DIR/knot.conf.new" "$CONF_DIR/knot.conf"
rm -f "$CONF_DIR/knot.conf.new"
# ---------------------------------------------------------------------------
# TSIG. A secondary's keys must byte-match the primary's, so they can never be
# generated here -- a generated key would authenticate nothing.
# ---------------------------------------------------------------------------
gen_tsig() { # gen_tsig <key-id> -> base64 secret on stdout
local id="$1" out=""
if command -v keymgr >/dev/null 2>&1; then
# keymgr prints a whole key: block; take just the secret.
out="$(keymgr -t "$id" hmac-sha256 2>/dev/null \
| awk '/secret:/ { sub(/^[[:space:]]*secret:[[:space:]]*/, ""); print; exit }' || true)"
fi
[[ -n "$out" ]] || out="$(openssl rand -base64 32)"
printf '%s' "$out"
}
SECRETS="$CONF_DIR/secrets.conf"
TSIG_PRINTED=0
if [[ -f "$SECRETS" ]]; then
log "$SECRETS exists; leaving it alone (TSIG keys are never rotated by a re-run)."
elif [[ "$ROLE" == primary ]]; then
log "Generating TSIG keys for $SECRETS..."
[[ -n "$TSIG_AUTHORITIVE" ]] || TSIG_AUTHORITIVE="$(gen_tsig authortive-tsig)"
[[ -n "$TSIG_ADMIN" ]] || TSIG_ADMIN="$(gen_tsig admin-tsig)"
TSIG_PRINTED=1
else
[[ -n "$TSIG_AUTHORITIVE" && -n "$TSIG_ADMIN" ]] || die \
"A secondary's TSIG keys must match the primary's byte for byte, so they cannot be generated here.
Copy them off the primary ($CONF_DIR/secrets.conf) and re-run with:
TSIG_AUTHORITIVE=<secret> TSIG_ADMIN=<secret> bash deploy.sh"
fi
if [[ ! -f "$SECRETS" ]]; then
# Increment only on a MATCHING line: as a gsub argument, `++n` would be
# evaluated on every line of the file, so both keys would get the second
# secret -- making the read-only admin key identical to the replication key.
_TS_AUTH="$TSIG_AUTHORITIVE" _TS_ADMIN="$TSIG_ADMIN" awk '
/REPLACE_ME/ { n++; sub(/REPLACE_ME/, n == 1 ? ENVIRON["_TS_AUTH"] : ENVIRON["_TS_ADMIN"]) }
{ print }
' "$SCRIPT_DIR/secrets.conf.example" > "$SECRETS.new"
install -m 0640 -o root -g knot "$SECRETS.new" "$SECRETS"
rm -f "$SECRETS.new"
fi
# ---------------------------------------------------------------------------
# The seven policy files knot.conf includes. The `dns` repo owns them; a fresh
# node has none, and Knot treats a missing include as a config error -- so
# conf-check would fail and knotd would never start. Stub only what is absent;
# never touch a file the pipeline has already delivered.
# ---------------------------------------------------------------------------
_stubbed=()
for f in remotes policy modules templates arpa dnssec public; do
[[ -f "$CONF_DIR/$f.conf" ]] && continue
cat > "$CONF_DIR/$f.conf" <<EOF
# $f.conf -- PLACEHOLDER written by deploy.sh.
#
# This file is owned by the \`dns\` repo and is overwritten by its pipeline.
# It exists only so knot.conf's include: chain resolves and knotd can start on
# a node the pipeline has not delivered to yet. Do not edit it here.
EOF
chown root:knot "$CONF_DIR/$f.conf"; chmod 0640 "$CONF_DIR/$f.conf"
_stubbed+=("$f.conf")
done
(( ${#_stubbed[@]} == 0 )) || log "Stubbed policy files (the dns pipeline overwrites these): ${_stubbed[*]}"
# A secondary otherwise cannot transfer anything until the pipeline first runs.
# Seed remotes.conf from PRIMARY_ADDR so the node can bootstrap by AXFR -- but
# only when we just stubbed it, so a delivered remotes.conf is never clobbered.
if [[ "$ROLE" == secondary && -n "$PRIMARY_ADDR" ]] \
&& printf '%s\n' "${_stubbed[@]-}" | grep -qx 'remotes.conf'; then
log "Seeding $CONF_DIR/remotes.conf with the primary ($PRIMARY_ADDR)..."
cat > "$CONF_DIR/remotes.conf" <<EOF
# remotes.conf -- BOOTSTRAP seed written by deploy.sh from PRIMARY_ADDR.
#
# Enough for this node to AXFR from the primary before the \`dns\` repo's
# pipeline has ever delivered here. The pipeline owns this file and will
# overwrite it; do not add remotes by hand.
remote:
- id: primary
address: ${PRIMARY_ADDR}
key: authortive-tsig
acl:
- id: acl_primary
address: ${PRIMARY_ADDR}
key: authortive-tsig
action: [ transfer, notify ]
EOF
chown root:knot "$CONF_DIR/remotes.conf"; chmod 0640 "$CONF_DIR/remotes.conf"
fi
# ---------------------------------------------------------------------------
# Operator tooling
# ---------------------------------------------------------------------------
install -m 0755 "$SCRIPT_DIR/knsctl" /usr/local/sbin/knsctl
install -m 0644 -o root -g knot "$SCRIPT_DIR/zone.tmpl" "$CONF_DIR/zone.tmpl"
install -m 0644 "$SCRIPT_DIR/knotdns-aliases.sh" /etc/profile.d/knotdns-aliases.sh
# PEERS drives `knsctl serials` / the knsers alias.
if [[ -n "$PEERS" ]]; then
printf 'export PEERS="%s"\n' "$PEERS" > /etc/profile.d/knotdns-peers.sh
chmod 0644 /etc/profile.d/knotdns-peers.sh
fi
_FQDN="$(hostname -f 2>/dev/null || hostname 2>/dev/null || echo "$NODE_ID")"
# The MOTD is a pre-drawn box, so a value wider or narrower than its @TOKEN@
# would shift the right-hand border on every login. Replace "@TOKEN@ + the run
# of spaces after it" with the value padded to that same span, which keeps the
# border put. That span is pure ASCII, so byte length == display width and this
# stays correct under busybox awk in the C locale (the box-drawing characters
# themselves are never measured). A value longer than its span is left to
# overflow rather than truncated: a clipped hostname in an ops banner is worse
# than a ragged edge.
_MD_NODE_ID="$NODE_ID" _MD_ROLE="$ROLE" _MD_DC="${DATACENTER:-unset}" _MD_FQDN="$_FQDN" awk '
function fit(line, tok, val, n, span, pad) {
while (match(line, tok " *")) {
span = RLENGTH
pad = val
while (length(pad) < span) pad = pad " "
line = substr(line, 1, RSTART - 1) pad substr(line, RSTART + RLENGTH)
}
return line
}
{
$0 = fit($0, "@NODE_ID@", ENVIRON["_MD_NODE_ID"])
$0 = fit($0, "@ROLE@", ENVIRON["_MD_ROLE"])
$0 = fit($0, "@DATACENTER@", ENVIRON["_MD_DC"])
$0 = fit($0, "@FQDN@", ENVIRON["_MD_FQDN"])
print
}
' "$SCRIPT_DIR/knot-dns.motd" > /etc/motd
chmod 0644 /etc/motd
# ---------------------------------------------------------------------------
# .env: seed on first run, else update only what was passed this run
# ---------------------------------------------------------------------------
if [[ ! -f "$ENV_FILE" ]]; then
log "Seeding $ENV_FILE..."
install -m 0640 -o root -g knot "$SCRIPT_DIR/.env.example" "$ENV_FILE"
for _k in "${RUNTIME_KEYS[@]}"; do set_env "$_k" "${!_k}"; done
else
# knot.conf is re-rendered from .env above, so letting .env drift from what
# was passed would silently redeploy the old identity on the next run.
log ".env exists; keeping it (only values passed this run are updated)."
for _k in ${KEYS_FROM_ENV[@]+"${KEYS_FROM_ENV[@]}"}; do
_cur=$(sed -n "s/^${_k}=//p" "$ENV_FILE" | tail -n1)
if [[ "$_cur" == "${!_k}" ]]; then continue; fi
set_env "$_k" "${!_k}"
log " ${_k}: ${_cur:-<unset>} -> ${!_k:-<empty>}"
done
fi
# ---------------------------------------------------------------------------
# Firewall: authoritative DNS answers on 53, TCP and UDP both.
# ---------------------------------------------------------------------------
if [[ -d /etc/firewall/ports.d && -x /usr/local/sbin/firewall-apply ]]; then
log "Registering 53/tcp + 53/udp with the host firewall..."
printf '53/tcp\n53/udp\n' > /etc/firewall/ports.d/knot-dns.rule
/usr/local/sbin/firewall-apply
else
warn "No host firewall found (scripts/harden-firewall.sh not installed); 53/tcp+udp not registered."
fi
# ---------------------------------------------------------------------------
# Validate, then start. -b so a rejected config fails here, not silently.
# ---------------------------------------------------------------------------
log "Validating configuration..."
knotc -b conf-check || die "knotc conf-check failed -- not starting. Fix $CONF_DIR/knot.conf and re-run."
rc-update add knot default >/dev/null 2>&1 || true
if rc-service knot status >/dev/null 2>&1; then
log "Reloading knot..."
knotc -b reload || warn "Reload failed; check 'rc-service knot status' and the syslog."
else
log "Starting knot..."
rc-service knot start || die "knot failed to start -- check the syslog (log target is syslog, zone at info)."
fi
cat <<EOF
================================================================
DEPLOYED ${NODE_ID} (${ROLE})
Config: ${CONF_DIR}/knot.conf (re-rendered from ${ENV_FILE} every run)
Secrets: ${SECRETS} (0640 root:knot, never in git)
Data: ${DATA_DIR}
Admin: knsctl -- 'knsctl list', 'knsctl status'; aliases in a login shell
This node serves NO ZONES yet. Zone data and DNS policy come from the \`dns\`
repo's pipeline; the policy files here are placeholders until it delivers.
EOF
if [[ "$ROLE" == secondary ]]; then
cat <<EOF
Secondary bootstrap: remotes.conf points at ${PRIMARY_ADDR:-<unset>} with
authortive-tsig. Add this node's address to the PRIMARY's remotes.conf (or let
the catalog zone do it) before it can transfer.
EOF
fi
if [[ "$TSIG_PRINTED" == "1" ]]; then
cat <<EOF
>> TSIG KEYS -- SHOWN ONCE. Every secondary needs these byte for byte:
TSIG_AUTHORITIVE=${TSIG_AUTHORITIVE}
TSIG_ADMIN=${TSIG_ADMIN}
Store them out of band. They are in ${SECRETS} on this node.
EOF
fi
cat <<EOF
>> BACK UP THE KASP DATABASE (${DATA_DIR}/keys). It replicates nowhere, and
losing it makes every signed domain go bogus until each registrar is updated
by hand:
knotc -b zone-backup +backupdir /var/backups/knot/\$(date -u +%FT%H) +journal
EOF
__ARCHIVE_BELOW__
H4sIAAAAAAAAA+xcbXfbNrLu1/BXzMpybCUmJdl5aeU4jWsrqU4cKddy2u06rgOJkMQ1RaoEacdJ
vL/9zgwAknpx0+3dpnvuKdsmEgAOgMHgmWcGUC+iOPWGcTT66o97Gvg8evDgq8ajRvPxwyb/3Wg0
uRyfne2HO181H24/bDza2Xm8g+2azeajh19B4w8cU/5kKhUJwFcPH5//GIeX8kv0+V/0rMGFtQBw
Xfiu1zvpnxzvv4b+y/ZR+6TX9eBQzsL4WvoQR0MJg2vwucBTk11IRCLDaxhORDSWynPWnDU4mQQK
RkEoIb6KFL6FDa4mIoWBDONorCCNIZ1I7OvvLQh8GaVBer0FKo0TMZYwE+lEbUEYqFRGKE74fiKV
klQWj8dBNN4CHG6axCGoeHghUw/alzK5TidYh5Lnezrs9uF176hz8BPOD8WlcjoLRUri/AjFDmEW
h8EQB5DIaYzl9f2DI6ycxn4WUisR+TxcP56KIIIgusQRx8m1FodT9WUYYP+ooFEST7ntOxT9DgXO
YtJXkCo4OIRZMMOWkUSNxoBKp5nBhziSLYiT0shagKNKgxh1N0GxWqk/iDCTCnuHZ29ev24fPwPU
PKhsoFB7WYqdlxdGj8ST0aV++xj1IfwW8GIPwR2QBkfucCKHF3D3blGccENeREnfcXo00nhG4xGh
Bz8G6STOUpyTeSmRaZbgUHsvcZFlxNMfxtMpqe1KKBR1T6HC7m2xIG6C76psOJTSx2Gj0QkU8k+c
stTq0LZDyktQceNEyshzHCUTVHLLAXys0bSg8qzbO2yfdw6fVbgmUoG/XJpkkR8kWF7HT3Uaty7P
UGiL59GiP7hMZCkudBoMXTEMcWUiLtXW2IJTeHbU6Z+0u8/gzHGMGbbmmhSd8B8e2ajuLg2mEnXX
gobjoCnTay4g+owllqlrhWXczkwV1ZdEaNJcZruaK9TGE0SjGPJnjUtdWkajVrYkWo6X/ZcuWsw0
UArLsFZlIWrYdrvqwe2Xcge79MH0TWt8zfaHmyFQwSCULOCXLBjOD1BEuEYySdC+HV+kYiCU1Noy
u70F9UuR1MNgoDVmG3Gbf8ZoWSJ0/cFiM1PDrS6Emq1ociGvVa71ZEUDLte6xU5RCSvamBoHrdhF
Qy12Om62hX1ud/cuLp/UKETlG4rtqw5ogLgWGjdcJ4iGYebT9GU61H3hpse9pDQUa90nMvJtd/ne
3gJ0pg1IYmO1W9DtnRAwjIN0lVwDayx3Vb2Gv1urDQ7eWp/j6a0tRDITt1ZqDL59cNkAR6er/2xX
+f/yKVudJ98LXMz/OAX6DP9rNnYW+d/OTuPRX/zvSzxrsMoCwH16OzJtzuNPjRlGt/1D+xgOeq9e
dU7g5Ps2HLe7h+3j9iE87xy1PeikxCUQpTyEqWAcxUxs4IWMZILogRxA9wJXSDCIqMjr6TgBN4Un
+DESU/kUJlMxdNVEbD98xH0WzNUgBxI+4p4DOULxIIkUajJh4HiEYIoclppJJakPpBLIWh38pL0x
sQdkABOkHoj0bqqC8bw7RDrpNuEJ6oc+bRPGI0AJcrKwuf/358f1Dv4B9wmUO89/qmk3GI7jBOc1
bc3NQft6mnUL1fX6aP+gff6q7RQD8adBtDSG3DcI32VuTb0CzhgOOuAnwSjVvI4cNXl9lqJ+xzj+
bNP86/kCz0Wkhmn4x/ZBIP/44cNb8H9751GjsYD/jYd/xf9f5ln7Wz1TSX0QRHWMFgG598RgK9kF
RWcvKW6jKJqBBA6OOh7Gk8j5CEgxhEUG581CqJvPxOZmoQ47nwfvsc0IyTrBVYgcdoQcmN0AsefX
MsEQfpgEMypjvsxSW/gqQNOjIDRIwM80xGJcyfGqkiLBTz5jXikyv0rEbIbF+NGPkTdKDEiI+bI0
8lv1SlU3rdRrnOzodA+BXIuvvYi6jlLxHmNc8hYc41wlKIAcVxaxLN/K+mn/1RFsviOk1iIxzNGu
E0VN39W2QOkkhx7yMM5CHyLySIDB5XDiQQ+h20qjhu5IB27aX0mVsv/C2CQbT1J0j7kWPH5r24Ou
RCjHJlqDFCFSPCd1tqKHPvgYhqFQCkMQjMaDEcrkYQmrMBGSE8lHgSUlum3GPEDXjFqNKFRPYyjR
9V3uZxSHYXxF8Z6R8s6mBSgCfUdtIhgJnBTlj8rz0LE+hf+jFGcxiNMJO2sFE5Hr2QwRhyFpgYMU
Q+wtCKW4NHGoCZYBg1oaVDDOMF7SKtpBFcU6KXQpwgBDS4qkNDfQw8Mq3fQB25oppbQFRHHkDsKY
3ajRWp6k0B2ZDAV+52SGMtmvbu/41f4RHHy/333R7sOLHrKh496bF9/Di86JR5uEI7JLWZc+2jxl
eDiU/u64vf/SfXG03++32PhQFk1YmsgRHT3nzdAYzZRtYiqS79M8/DRxItKoMATsJWEZlF2ayCma
MepXGcOklB3ZtghxelPhS7svi6iW1XLNIb3tjsYlMhy6i4Gx5yikbK7MdPxL6+w4/+h12/3z417v
ZK/6sfjScudja1p+deMc9LrPzw87x9jWfsSWlnveODYPifX2Y8ut2qYsxUtx4904+28OO9Ql/217
i8d1DWUuD9rDghvnJZKz896P3Tb1WnxpuXkq6MZx/EBu1gA+wiwJonQEG1pOC9bV22gDKtV7FXh6
d3sXNz6uY3MXbhzSE77z+VduHB4NtWX2ZduvK+B/detNMllwM7i//vxk/eQftQoWfuy/Oeydv+nT
cKsf9d/ExW9uKrqDt3ky5+nTSpWVUYHtp3VfXtajDK3i0ydIk0w6OMdMibHMR4G7Ep482XjT33/R
3nCMB3hiknlP4VQkY3XmOJ1IzXRaCTZFeCWu0SLFSNaIQVMSbFUeicuLDcp7nzgpOt80U3CqEeks
b8/YoCt30UBD/QIEI4inBAI+vquh9Yl+9WnRlwYgBlPdBDGGzDrJIs5e6e2DqBhfRcAbnlLODufd
AtwMS4NHDcxoi/Lmi2Jfbqi8rRhj5zi31+32cR9lkKgVY9IrzMMoUpym21Iqdu6xoGVSqlp3HGk4
zv5sFnLIYiBrpQJdU2mHswsjWv8cvriypFLnOwTbC3dMTgM2GTjUAujswhQJE3kFxKQ4GjKu48xi
yj6xBVBSO5+/67IHMm7jk3YvPGpCwLwdljAYrtBbPDM55UJtiFLVNlp179hx2tFlkMTRVEYpdc6r
sFdpNrymt/3Ae/Q15B+/qdAqSJko5g3vzAK+c9jcaTOsQVsNxUwWDpIaZhgooj8QaMLoppDEkOOE
171+5++A8a2nl+lcv4FbqXCqCTrcJNB8Z5alpMlT4X5ouN947llOEPyYI2P6yNHcVKYCsTkRQ/KK
HLgOBZ+DoBsWiQebhzFn/LCb4YSHKGDMYXQIA3ztAhFZvp/RmQltUXIwlMbm7YRRLXvqyqlXYdWq
fC5DdOQi5f1xPR3EIeO9Qn1H5HoGZBl6lGi1OLLIqzko+Zxp0RzmrSvCrmYFPqHN+LCh6m+9+lv8
b7xBwIe9/zgJcOQFJ0ETFNH1Vu7sCSxMKK/V+i0c0PZQsH90BPFI+zI+fgnjWAe6qdYezZPc9yK9
3KAAN8YZDTJ0WqMg8s+D6NwOQeUQiB6feCok72HEBf5eFaE9eb9X3cyni9PzKzqyJ/VT+gGLrEuq
lHOac+VldlUuzxOkuzjfHL1PT4kTVkdwdkaQTQcAQZTJvH4N9qPhJE40saIJMymlpAaa1rtRHHuR
TN+R9RCDZ9oJ7wYisTVeLgoXYIzuHtxf2lD5+fS0pWbIw1tnZ/fc8hfDdMtF1Y/J+5tywdsquaFR
hdlhNHeuUBiIdm8fR2tr9+o3lblG+jgJGnnhKNDrgCvrlBo0acfyOhIwnBMwLC3iMF/BfK2GHB2Y
cz9eDiDtzyketaF1X6kW7KVeHdarvkedVXBFfsP0Vr/7+ye7gDN6qjjQahP2/gU/G2g52zQf3LN7
tqj27eZb71fra/erOCnqhbVJMVGxJ0r+1LjrOTYB4uoCNuo/vz2tfxyrbLBZPz07PatvQaWyhYOr
7RrfV23ebFAPw6l/jk5icbU04mp3obEZPoxw4xmFUC2StWbLvdnVrfYqWplqEozSnNRQydWE3ASp
Zg3ccQoNXrDSCg8FgjrqLZhfQeOtakZ+9eM296Y72Ibd3bnW92rIzBELsggVhExCn5G2gKCv1FQq
MZxfVDIudGFmxnp3a0mKT8QM8fq8I9XzL9sF5BFuLjSIuEEeptoGTqEILbqkjLleauXZ2Enb8Vg2
YLYU7jAzNPsST7+0zkiWVWqPB3Gn2e+Ir0uoXMymtrDfeAxWgxzCaM9RYpjEEqxwPVmzuT6M9hY2
J82kbqSVdimtE2rnAyPw3bumU/xqe2PxiuhlIkeZ4og0LoKuQgrCuI1dyuutLmQoUwRvwudRnFGy
Ni2aVrTeOPiqzIWHaBrUmVVAHiCyH8i5mlcpvQ8c4N4eKhpGhiTMxIjXcaYjQ7LDgNjLvDjmLsgf
+PBzk+n5KlXCfRgGdeQnrl1WT01qnpmbvmdgtntg7opw3kMQIfFd5ix0ouq3dCIK5QfjSDt4Oow1
LxtpGZ2XMlSRQ3D1pQc/GJnEvxvFrm7PBxJsvGggOH004ZAO4on1GFnHzw+g+c3X264SkWS6RsPg
8wlUhi+uMQogem5Gj8x8ogkb4rVnz/GxZm8ukPtpfbru1xpNg110AYIsof7ssPdqv9N9ZpVXH1ds
DYZ4nf0jrNHyqKZSGAk8JZukEFiLnF5CXsCfzE6fEEZVqkW0O187jX06V35gCo0O9jn1Yy/s2BXU
y6TTNgvUnHN7nMUzSRNKCeB+NPJ0BozSNxhZ5HGIXh2mcaXjGj+W+vIJbzNvLk4u593I31JdcX1m
XbkUnSS+8cQWEivaPhdD5DyXoKuZif2a0la3NzP8wQZt37Wf947bRaLJRqFzSU9dS2kp3JC+Z0Hx
b6vu6SxAoN6HpdhRJ9mQqyKRp/4Gwt47oWeZwRpkzVuwLQZQqf9fCWB9637TLyuZnl9TM49vqhnX
h1IZg+Rc4AvGm+2iYdhLZxQYhYH0c5C3tgvd9o/aNiPcuBTpLIe/3lygzHatP/PeFsQZjTiR50vd
hTwGR0pXOsMbp+b6B4OZ9EoEKr9dleO/+aoXrjLftJS+mPPpetlLtXkWcpbEg1BOdXibe3lzp4mi
6wohebVMtPaqZb61ZwCmMrfV8CUUvq5gc13VthZROc9VLe2xypw0zalNd3t7li8sEWnb62EfSjeV
DP4kcoxIkIiEcAbBKBNhyx4U4wAXtrs1B0M5tYNbzTpzYGPGqS0kmS4zz3+Hvi2kOAw4m44+w3Zy
kSWKQ6ZVGF7FUhkjKA+CFqRoc8HNUmC0JhrLRkLZv6d3t+HJk3bvuWMzlyXeoaeEWsMI3LxKb/aO
D9vH8Gr/5KR93PfgOSckyEmb7aEvVnK+3ua7TRJeZxMQ+1N0pYd9zgsRbNN0yc36iUD4HYpM4S5H
L/jD8/3OEY9cH+2bBBlJTqSiw7tEoyzdWaPsIu7ZOEtxVfTlQ4FRLqIUogBy+iz1Edx13h87JQaF
Fv1LRjwB325ppTQ9eE2EVk14tMQJvm48/ppugkmEUZqmCLeI95r1plZovPoAqbBYDQTbHvwoUAl0
gbKYPLU/OTmCTcq0MgR9/ehBo0EyxidHh6qmX96hY7e8j26fxzBmWNzFdS6J7bI8/dYDfdKkz2FK
Y2R70NdLHdBXdu2tzZlMUMcIJCqVM3jAqMitDrv9fvuAdpziIwu6Djfg6xvYAaWkZhlSKGRcG4xP
/A3uU/MNZgaoZFwie18Q9zUzuh3vofceNmlQlLtAbaGyB5RlSxIkeMzkd7xHXqOGU0lmCPpYlo8h
lCO6t2lZh8mHEQ9ItkoDRLyTCBgzzyHjZoeDUA4uGswMKifXM1k+SUS8MRjQQqpbhgMGMf2dYKzA
vRIOiAEDstlWv+54/zinW3a4FmMWfEyJPNixL/EJQ/M4PYt7K9C5f/i88M/7OpNnQSz9+GGE6Lrs
DQyX1UzV09bru8t0uqAMhXu0t4hTVB5NQ8FvkWFiROMwzX6pLriJvYU5W6dlhJMv4vvX68qb2zN2
MN6iyyzEGWdlktML3kpnr339zbSBmR4CVaGX4gT4gqfSr2kD1fqhQzudKt8opco9z9vQpq+PO8Qw
iZU+8FAVw4L+JyPg5cyzOQh50+0c7PdP7DV9Tzexp7Miuh4KirIwIg1we5pWRhqnsBHbryghzKEB
CWVvJ0XCuWGuQm4fhjokjcyJLm6yLNW1NrYwryI5CQ3x0okgu9P9pfQqufHb0qtlJeMOvvCRM95X
dDEMnjW3H3sN/KcJ/d4+Z4NXpcQ+mrzXzs1GwbLzLKG780DpTvbW3eY2J819PuAr99xyvy0lSAnc
ZpzZ4DWdm09ZOFpeIXS2SMU1OszNqPKMmv3WyZQYus1plXtH684TXvAEnmwWOcUaGzgnnZhQ2aQT
mkXNiuF02y6YNCEO6FmeTNIbrLbYzGzTcks6QrASi+5trU5jzslZTnKWxTEi1ubGZ6GrqZHrFhY4
f0ZZWVqH1Xy/yCAWeFxbidt2QhoveIQl/LDVdNBotfFvDX/uNHN59EvUs1nkHRc4Z1LOi5I3+V3j
mTslXDGeVXS4WeLTqwa1JGVFwo36VcVdDZviS5FTmV8CcGZNJ9Rukwjl6xcK7bF0B6N8/cJm1bxl
QcYvsRpWDR3RQ5+MttzL4EafxBSbhzxwyeiX3i455OZKf1x+1uYiZz5bQQhFOmtOmcshtanF2OUW
UTaU3xSoyKuCTusIouatfG1+75j+KqvUIkMlf12E+d3TYv0osNrb2PjkTj657kSGs0/0Rw3YNG39
4mmAIdN0HIDcNrnOTZjf5UXh1Pifff8vvz/zB/bxufv/O8u//9xu7vx1//NLPLvOrr4HOY+LNrlq
09LOruNUe8edF51uXoYxYRXjzDt0f9dx8tI7ne4dZBJ3EMbcpkdZtCGigky88dgjbPNEiEEclSUB
XZb0YNO5Y5Pcd3ahr3ktBhsHvaNDt3+yf3yC0X/7EHrdo588HTTyL0sXzwDu3MHXQR9GfCb/X6Pr
k5c2sKF0nX37kn5o2aK79BSrKL60GNlfuNrfMZhEId0VpLha2JeZg3J2TGcwKE1pfqsGNOFrc2cV
RfDGc1h59O6xHCHhndC0RyIMKXnL9yLM7wjYRwQKw0Xnzjf2DZTn3GluN74xMtrvZ0HCVyqbD3wP
Xr1BWo6Dueafw15BgqH32A0x0qCfvqG4KMtnresSM4jN7R2/xlkNYe5NCBw6nyopk6PJ37yaxMok
QtIs4dubuCLFERkNyaeJ6jF2OWWBLQ6I+lOewqktGE+3v9J2VjbaXmzk7KIeaE2/4zlzYuOIfnTr
/OlQ+1/5kAckRYowQFquPDX5z/fxGfx/1NxpLuD/9uPHD/7C/y/x6F9RzVtA/uuvWRITjnp+fbmR
+am/pGtkiUhjulBpz9vpZHWInmSKJC653gK6jEbg4QNC7gD3/5RwM+HjZv0bcYsXScC33SPf5mSl
b/8HAkX74lraVfK/7V1LjxvHEb7zV7SlA7n2cuTIlmQwICKCpGxCK3JNUhZkBI6Gzx0sl0PPkNp4
YQNBzjnkoEMOOfqX+Ze4vqrq7hkuqbUeq42CaRCQluTU9HRXd9fjq4/xkk6VAW2xYy6TZUjYnAzp
9IRDBrR7S6yYTw1sz4EW876/5urrzR3x967hBlKKjxjrOE7PaAMfv74UH4MU4+hKcwX5JzzYDKaX
qnybWhHEwGtq9DmZ620Efw0DHRUYb2v2LSBeeBkMHfTzOR27pBiuI5wGAL4+DUqsVPQMyaJe3rLI
yzkvIZu5q5lVBEjgZgWX4U7WX0idxHSckejd5rKV6CC2eXitu3xtL88EBsolP+G9bps/EihAGJhj
zmZr/g/2BZ6KD31NbainUzPpGR3zPHyjBSJkSTiJNko3ESIzQKeqJs/hGGLkefgVzjehlbZJRoIb
BfIA5Qdyt0rG/zpw1spSplLsFRGFpdRotciyov71209637VbQel0eZEwoHOvf6UwTsZkWlkKf3Fq
sQ2vVtwEUjM6YD6ZVUIRKqkFSBfGAvTUnG4gnah8dUAecnKa5sIklU/Jj2L6iKNes3GkGxHmDyr5
4yZijROsQn6CGZuAvSSGTO4Zb3BgJUAJkcwCmYCCvE02vB4s1DuUVT6dVLH6VCnPdTGSjiwnAmXQ
Mg0kHW6bAUOBtWrjDNkvubAangu0PFzKteECw3mKB+Xb2BAo3Eos3h0BUAd98JMBLhaayPGOifRB
Jp3H971XdcOX0dxOxPsVbhfmanRRL48nZkdFiaJsy34RT8d7vytpaP9dYEH3fhkf8uLvrfwasoh0
aD58iThBvnMhNq5uruNNwnm+LAUER9nJcsaU6tKgL9LSlwMW00sCKtz9P5v4YmbS5OUyDiaIhtCf
PLGXwz1/usWoK53ZQXg2PTQj0hNExeQo4A7zaQrlJLWGB6X7DuJUqSzSkM4KeofdIwDssE9lqgNC
qf9Ya6VcNvy1VXpkt3fqc1gvZ+J45evQPV+bcg3C1TPFbu62Y6X80cIQg7tHs2hMe/tPgWlIdcyP
nDTBMGmyxLMa1eiEPr0iWbI/TYLt4A0SJdspEiiXbvWvy24wlH87CZCD8rNcUokNHpseIz0Ha8lf
THfQafH9pmtklfX+S1+4z32IJpk+4E+kI4C9dR355Va2KwJVj5jrxyq6zIxkrqIUho8kr3hZ4twm
131gs17oQt6+NVpJyoRJfuNIUqexGlFnpW0ylD50ZmhV+Awy7EU1qXUEWEbCAWJ/0nEC81T0BUjD
23JI8L1tnGAcr35CnnVNBrEvdZxHa1fwp7Y0320CQpq54hKhV5LplqK0Mcab7d9zSfV7w4YUoF6m
vnAQI5/vseN7ix8LN3fPJeH9s2k6v+J75cLb/jgalmIVUbkzsn+u6R5X+P/37t19sB3/ffCgiP9+
kFYy5rdX//rt1T/+T1//5gf8567kzyefmBaKtPumapr9zrADJ6LTfdSnE7r/tDl82m/jO7lGovYL
fJu2U+Dz3lPOM9IRuRTPeg3323Vy8HwwbD/5QwKb37SbjyGwb571+o8tLLnfPuo1WuYO/UcC7L1H
zNoDlF2n2d4l8L83PZXX9/rVjVgLoN8mh5OR/GgMG812dwi6xbec02/idA3sQQ0fPHz0bav7RqIu
C+ziTO+0WJ5xPIfvILAfL6R3LLDfO2q/Qw9f3fRUXt/rP8Dzfc9l6FibZ+EynAs6kYkdmprCFupR
qRJWFOxx33qCLrVPohpLZSyFte65IDRwtB9esF2gQ6IWsAzhqDEJYGCegtxK7FbGQy6z3Kha1w0w
39ozOaSM8HxGrls5RWSHvjg4mZLFbP5a+aH6w8EdfLwV5sTEc8Qp1/RLGhepoH76qqDTgYq6sInH
YE2Gs9kTZTOVVSbKdjm2ZqWl43zHmsoew1EZ7R4J4Big7cA434FciOtlFO7gETCVZQzkLSq866b3
2N19iwfhe89o4CHMPC0YWY2ZTGlsoQoqYzW6yMn4OoYiad2qMCJw2x8GsZ2ZjncIUmziFYIkRqKC
EAO5LKjRP25kxewWhGtZDKIZuWGWiIqPp8hC8GX+elGYv+pRRFr/mTAEsHckXo6pcBijlo1XHGCM
NTbghiSvbMNL5A6KJb6wGDfxCLNPv9fttGwQfJErvpQbZV1m9ZRNpQ2CJbjLXn3IQ8vc6iiOT7lD
+53OlDHafKfULQEtJM10YGcpKd8jA3nKXq/wuvyM7WrZWoGqpMdpvewE3lzVaDlLwsBSzs6iJF1n
ewUcDR5YUgwS4tyRXTCV6uhgZ65EUhUZ2uKSyRIXkynUHb4BaTGYCDxb8U2b9X+4MUf0NfF+2naF
//fl3fuC/7l77/79++QLkv8HSrjC//sA7bZpIt5Em7jn+4ROXGL5DDLs4inX06HiZCkLk1MktHAO
S7cl8pVMq/TWodmsuERWsjBa4iGmyyoEmkWZOTdLi8fBCoQlsw62KT5TFmoZkS3VOlkXL4Ua3aV6
pZ+c49EksuSBwIIlBwRg1CgklpKkdew6fOCjmOvwlIROZ2Cs45yxMwVNR+nHzbXZmh/ihbAp17eG
rmDc8qoL94u8h/NLsgOa5XjSG7bExPTB1bJkpuhII1+HGdOmwTwwje5zhLuroEr9Iiips1LHaK6S
6Aygop89wIiZ2Oz7YGMwJ/FikmYOuEMjxd+cMHFE7WrHHHpQ2OPG4NgyyPlmucV5U49m0VSRuHp/
+jsw7b+HWkk/VbQAxHgMFNf6j6dImajBMxICVqF95Y4NB52vlT0HndVAfFCCa1X3T5tTqSMmj4dC
3rhevJNCZX6PgFmh3Polq+qFMuS/IGcJ52w1nZLZxKx8yHQ8xIHMKLmIf0mCbOIkic/Fylf/ydtn
uQRJNOPj+cTV4Iph5bYE/E4DXPoZKsQq4dJlWlhlN6mseH/9QVASlv/655yA+PzhvS8OTa1G/+Sn
rZ9h/73x8X+7KfO6jT26JnWJugh9Xmud0JDOmDwrAWRz+yxAfpAWPl22MFnCedmnMTl+AjkNMqKD
BZWRK7uAMkWhJMtTENJ0O8KKE9qq+JTgAwPgI/eRyzdq/pL9CC4DDErH/c6TRv/538jt7NfzE4jN
7KOcuX3TyQEsBtfi8UfThWNJFVzUKFwuaUxoBSqMqOSDXDw2x54iLZ/6ekG+KZ/g4gZJKgnIIBQh
bq1l/r0WE8+Ed1IAKlJI5nLlmpR7pnEOa4BchpxxqRNpGs0jl6x9PLZ10YpWtKIVrWhFK1rRila0
ohWtaP8r7Xext/XSAHgAAA==
+48
View File
@@ -0,0 +1,48 @@
┌────────────────────────────────────────────────────────────────────────┐
│ !! DANGER - CRITICAL INFRASTRUCTURE !! │
│ │
│ YOU are connected to a CRITICAL SYSTEM │
│ CHECK YOUR WORK BEFORE RELOAD / RESTART OF THE SERVICE │
├────────────────────────────────────────────────────────────────────────┤
│ Data Center: @DATACENTER@ │
│ Hostname: @FQDN@ │
│ Node ID: @NODE_ID@ │
│ Role: @ROLE@ │
└────────────────────────────────────────────────────────────────────────┘
Zones are managed in GIT. Changes belong in a signed PR to the dns repo.
Anything you edit here is overwritten by the next deploy unless you also
land it in git. Use knsctl for anything that writes -- it audit-logs.
Will's Cheat Sheet \(^-^)/
Reload / check
knrl Reload config (use when a zone was ADDED or REMOVED)
knzr domain.tld Reload ONE zone's data (prefer this for record changes)
knsc Check server config is valid
knzc domain.tld Check a zone via the running server (no output = OK)
knst Zone status for everything
Navigate / edit
knpbz Go to public zones /var/lib/knot/zones/public
knsec Go to DNSSEC zones /var/lib/knot/zones/dnssec
knarpa Go to ARPA zones /var/lib/knot/zones/arpa
ozf domain.tld Open zone file here in $EDITOR
ozfa domain.tld Find + open it anywhere (warns: break-glass)
Inspect
knser domain.tld This node's serial for a zone
knsers Serial comparison across PEERS
knsid domain.tld Which node answered (EDNS NSID)
knstale Look for silently-ignored zone files
Domains
knsctl add domain.tld --class dnssec|public BREAK-GLASS
knsctl remove domain.tld BREAK-GLASS - read the
ordering warning first
knsctl help
Every knotc command here blocks (-b). Without it knotc reports OK when the
command was SENT, not when it succeeded -- a rejected zone looks green.
+51
View File
@@ -0,0 +1,51 @@
# knot.conf -- BOOTSTRAP SKELETON. Deployed once by deploy.sh; rarely changes.
#
# This file owns only what belongs to the BOX: identity, storage paths, listen
# addresses, logging, control socket. Everything that belongs to DNS POLICY --
# templates, dnssec policy, remotes/ACLs, modules, and the domain inventory --
# is delivered from the `dns` repo by its CD pipeline. Do not add zone: or
# template: sections here.
#
# Values in @UPPER@ are substituted by deploy.sh from .env.
#
# Reload: knotc -b conf-check && knotc -b reload
# The -b is not optional. Without it knotc returns OK when the command was
# *sent*, not when it succeeded -- a rejected zone file reports green.
server:
identity: "@NODE_ID@"
nsid: "@NODE_ID@"
rundir: "/run/knot"
user: knot:knot
automatic-acl: on
listen: [ @LISTEN@ ]
control:
listen: "/run/knot/knot.sock"
timeout: 0
log:
- target: syslog
server: warning
control: warning
zone: info # zone-load rejections and KSK-submission results log
# at info; at warning they are invisible
quic: warning
any: error
database:
storage: /var/lib/knot/database
journal-db: /var/lib/knot/journal
kasp-db: /var/lib/knot/keys
timer-db: /var/lib/knot/timer
catalog-db: /var/lib/knot/catalog
# --- delivered by the `dns` repo pipeline; see that repo's knot/ directory ---
include: /etc/knot/secrets.conf # rendered by deploy.sh, 0640 root:knot, NOT in git
include: /etc/knot/remotes.conf
include: /etc/knot/policy.conf
include: /etc/knot/modules.conf
include: /etc/knot/templates.conf
include: /etc/knot/arpa.conf
include: /etc/knot/dnssec.conf
include: /etc/knot/public.conf
+62
View File
@@ -0,0 +1,62 @@
# knotdns-aliases.sh -> /etc/profile.d/knotdns-aliases.sh
#
# The operator's existing muscle memory, preserved verbatim where it was
# already right and corrected only where it was silently wrong. Sourced by
# login shells on every Knot node.
# -----------------------------------------------------------------------------
# Reload / check
# -----------------------------------------------------------------------------
# -b is NOT cosmetic. Without it knotc returns OK as soon as the command has
# been *sent* to the server, not when it succeeded -- so a zone file the server
# rejected reports success. Every triggering command here blocks.
alias knrl='knotc -b reload' # config reload: picks up new/removed zones
alias knsc='knotc -b conf-check' # validate configuration
alias knst='knotc zone-status'
# Reload ONE zone's data. Prefer this over knrl for record changes: smaller
# blast radius, and a parse failure in one zone cannot disturb the others.
# knrl (full reload) is only needed when a zone was ADDED or REMOVED.
knzr() { knotc -b zone-reload "$1"; }
# Check a zone using the server's own load path with that zone's configured
# semantic-checks applied. knotc(8) marks zone-check "(*)" -- a LOCAL operation
# requiring only a configuration, so it works with knotd stopped, and CI can
# run it against a checked-out config without standing a server up.
# Stricter and more config-aware than standalone kzonecheck, which knows
# nothing about which template a zone uses.
knzc() { knotc -b zone-check "$1"; }
# -----------------------------------------------------------------------------
# Navigation
# -----------------------------------------------------------------------------
alias knpbz='cd /var/lib/knot/zones/public'
alias knsec='cd /var/lib/knot/zones/dnssec'
alias knarpa='cd /var/lib/knot/zones/arpa'
# Open a zone file in the editor, relative to the current directory -- composes
# with the cd aliases above (knsec; ozf srvno.de).
ozf() { "${EDITOR:-vi}" "$1".zone; }
# Same, but finds the file wherever it lives, and warns that a hand edit is a
# break-glass action that the next deploy will overwrite.
alias ozfa='knsctl edit'
# -----------------------------------------------------------------------------
# Inspection
# -----------------------------------------------------------------------------
# Serial for one zone from this node specifically. Always query the UNICAST
# address: asking the anycast service address reaches whichever node is
# nearest, which tells you nothing about which node is stale.
knser() { kdig +short @127.0.0.1 SOA "$1" | awk '{print $3}'; }
# Which node actually answered? NSID is set per node in knot.conf.
knsid() { kdig +nsid "@${2:-127.0.0.1}" SOA "$1" | grep -i nsid; }
# Serial comparison across the estate. Set PEERS in /etc/profile.d or the env.
alias knsers='knsctl serials'
# Catch a silently-ignored zone file: on a cold start Knot can log this and
# serve the journal copy instead, leaving git and the served zone diverged with
# a clean exit code everywhere.
alias knstale='logread 2>/dev/null | grep -i "ignoring zone file" || dmesg | grep -i "ignoring zone file"'
+256
View File
@@ -0,0 +1,256 @@
#!/usr/bin/env bash
#
# knsctl -- Knot DNS admin CLI. Replaces adddns.pl / adddnssec.pl.
#
# Fixes four real defects in the Perl scripts it replaces:
# 1. Their duplicate check searched for the domain wrapped in double quotes
# (/"$domain"/) -- BIND named.conf syntax. What they wrote is unquoted
# YAML (` - domain: example.com`), so the check could never match. Only
# the -f zone-file test ever caught a duplicate.
# 2. Neither script consulted the OTHER class's manifest, so a domain already
# in public.conf could be appended to dnssec.conf; the following
# `knotc reload` then failed on a duplicate zone -- after both files had
# already been written, leaving the server misconfigured.
# 3. Nothing validated before reloading.
# 4. The reload was non-blocking, so a rejected config reported success.
#
# NORMAL CHANGES GO THROUGH GIT. add/remove/edit here are BREAK-GLASS: they
# write directly to this server and the next pipeline deploy will overwrite
# them unless the change is also made in the `dns` repo. They warn and they
# audit-log.
set -euo pipefail
ZONES_ROOT=${ZONES_ROOT:-/var/lib/knot/zones}
CONF_DIR=${CONF_DIR:-/etc/knot}
SKELETON=${SKELETON:-$CONF_DIR/zone.tmpl}
AUDIT=${AUDIT:-/var/log/knsctl-audit.log}
KNOT_OWNER=${KNOT_OWNER:-knot:knot}
die() { printf 'knsctl: %s\n' "$*" >&2; exit 1; }
warn() { printf 'knsctl: %s\n' "$*" >&2; }
audit() {
printf '%s %s %s\n' "$(date -u +%FT%TZ)" "${SUDO_USER:-${USER:-root}}" "$*" \
>>"$AUDIT" 2>/dev/null || true
}
usage() {
cat <<'USAGE'
knsctl <command> [args]
Inspection (always safe):
list list configured zones
status [domain] zone status; all zones if omitted
check <domain> knotc zone-check -- the running server's own load path
serials compare this node's serials against PEERS
path <domain> print the zone file path
conf-check validate the configuration
Apply:
reload [domain] zone-reload <domain>; full config reload if omitted
Break-glass (writes to this server; must be reconciled into git):
add <domain> --class dnssec|public
remove <domain>
edit <domain> open the zone file in $EDITOR
Environment:
PEERS="10.1.24.68 10.1.24.69" peers for `serials`
USAGE
}
# Escape a domain for use as a literal in a POSIX ERE. valid_domain() already
# restricts input to [a-z0-9.-], so the dot is the only metacharacter that can
# appear. (Do NOT reach for a general bracket expression here: one starting
# "[." opens a POSIX collating symbol and silently breaks the pattern.)
ere_quote() { printf '%s' "$1" | sed 's/\./\\./g'; }
# Which manifest, if any, already lists this domain? Checks ALL of them --
# looking at only one was the Perl scripts' second bug.
find_in_manifests() {
local d rx f
d=$1; rx=$(ere_quote "$d")
for f in "$CONF_DIR"/dnssec.conf "$CONF_DIR"/public.conf "$CONF_DIR"/arpa.conf; do
[[ -f $f ]] || continue
# Anchored on the YAML key so `foo.net` cannot match `barfoo.net`.
if grep -qE "^[[:space:]]*-[[:space:]]*domain:[[:space:]]*${rx}[[:space:]]*\$" "$f"; then
printf '%s\n' "${f##*/}"
return 0
fi
done
return 1
}
find_zone_file() {
local d c
d=$1
for c in dnssec public arpa; do
if [[ -f "$ZONES_ROOT/$c/$d.zone" ]]; then
printf '%s\n' "$ZONES_ROOT/$c/$d.zone"
return 0
fi
done
return 1
}
valid_domain() {
[[ $1 =~ ^[a-z0-9]([a-z0-9-]*[a-z0-9])?(\.[a-z0-9]([a-z0-9-]*[a-z0-9])?)+$ ]]
}
zone_names() {
knotc zone-status 2>/dev/null | awk '/^\[/{gsub(/[][]/, "", $1); print $1}'
}
cmd_add() {
local domain class serial zf rx
domain=${1:-}; class=""
shift || true
while [[ $# -gt 0 ]]; do
case $1 in
--class) class=${2:-}; shift 2 ;;
*) die "unknown option: $1" ;;
esac
done
[[ -n $domain ]] || die "usage: knsctl add <domain> --class dnssec|public"
valid_domain "$domain" || die "invalid domain: $domain"
case $class in
dnssec|public) ;;
*) die "--class must be dnssec or public" ;;
esac
local existing
if existing=$(find_in_manifests "$domain"); then
die "$domain is already configured in $existing"
fi
zf="$ZONES_ROOT/$class/$domain.zone"
[[ -e $zf ]] && die "$zf already exists -- refusing to overwrite"
[[ -f $SKELETON ]] || die "skeleton not found at $SKELETON"
warn "BREAK-GLASS: adding $domain directly on this server."
warn " The next pipeline deploy will remove it unless you also add it to"
warn " the dns repo (zones/$class/$domain.zone + ci/gen-manifest.sh)."
# The serial is only ever a cold-start seed: Knot assigns the live serial
# under zonefile-load: difference-no-serial. It must still be large and
# RFC 1982-sane, so seed with today's dateserial rather than 1.
serial="$(date -u +%Y%m%d)01"
sed -e "s/@DOMAIN@/$domain/g" -e "s/@SERIAL@/$serial/g" "$SKELETON" >"$zf.tmp"
mv "$zf.tmp" "$zf"
chown "$KNOT_OWNER" "$zf"
chmod 0644 "$zf"
# Append to the manifest only after the zone file is in place, so a failure
# never leaves config referencing a file that does not exist.
printf ' - domain: %s\n template: %s-records\n' "$domain" "$class" \
>>"$CONF_DIR/$class.conf"
chown "$KNOT_OWNER" "$CONF_DIR/$class.conf"
# Validate BEFORE reloading -- the Perl scripts reloaded blind.
if ! knotc -b conf-check; then
warn "conf-check failed; rolling back"
rx=$(ere_quote "$domain")
sed -i "/^[[:space:]]*-[[:space:]]*domain:[[:space:]]*${rx}[[:space:]]*\$/,+1d" \
"$CONF_DIR/$class.conf"
rm -f "$zf"
die "configuration invalid; no changes applied"
fi
# A NEW zone needs a full config reload. zone-reload only reloads data for
# an already-configured zone and would not see this one.
knotc -b reload || die "reload failed"
knotc -b zone-check "$domain" || warn "zone-check reported problems for $domain"
audit "add $domain class=$class serial=$serial"
printf 'added %s (%s), cold-start seed %s\n' "$domain" "$class" "$serial"
if [[ $class == dnssec ]]; then
printf 'DS submission to the registrar is manual: keymgr %s ds\n' "$domain"
fi
}
cmd_remove() {
local domain manifest zf rx confirm
domain=${1:-}
[[ -n $domain ]] || die "usage: knsctl remove <domain>"
manifest=$(find_in_manifests "$domain") || die "$domain is not configured"
zf=$(find_zone_file "$domain") || warn "no zone file found for $domain"
cat >&2 <<EOF
knsctl: BREAK-GLASS removal of $domain
ORDER MATTERS. For a signed zone, removing it here before the parent's DS
record is withdrawn causes SERVFAIL for every validating resolver -- that is
an outage, not a graceful shutdown. The correct sequence is:
1. Publish the RFC 8078 delete signal, or remove the DS at the registrar.
2. Wait out the parent DS TTL (commonly 86400 at gTLDs).
3. Remove the NS delegation; wait out the NS TTL.
4. Only then remove the zone here.
This command performs step 4 only.
DNSSEC keys are deliberately NOT purged: 'zone-purge +keys' is irreversible
on Knot 3.5.x (the key trash bin arrived in 3.6.0). Orphaned keys are left
in place for a later, deliberate cleanup.
EOF
read -r -p "Type the domain to confirm: " confirm
[[ $confirm == "$domain" ]] || die "aborted"
rx=$(ere_quote "$domain")
sed -i "/^[[:space:]]*-[[:space:]]*domain:[[:space:]]*${rx}[[:space:]]*\$/,+1d" \
"$CONF_DIR/$manifest"
knotc -b conf-check || die "conf-check failed after edit -- inspect $CONF_DIR/$manifest"
knotc -b reload || die "reload failed"
if [[ -n ${zf:-} ]]; then
mv "$zf" "$zf.removed-$(date -u +%Y%m%d)"
warn "zone file retained as $zf.removed-$(date -u +%Y%m%d)"
fi
audit "remove $domain manifest=$manifest"
printf 'removed %s from %s. DNSSEC keys retained.\n' "$domain" "$manifest"
}
cmd_serials() {
local peers d local_serial p
peers=${PEERS:-}
[[ -n $peers ]] || warn "set PEERS='10.1.24.68 ...' to compare across nodes"
# Query each node's UNICAST address. Querying the anycast service address
# reaches whichever node is nearest, which tells you nothing about which
# node is stale.
while read -r d; do
[[ -n $d ]] || continue
local_serial=$(kdig +short @127.0.0.1 SOA "$d" 2>/dev/null | awk '{print $3}')
printf '%-34s local=%-12s' "$d" "${local_serial:-?}"
for p in $peers; do
printf ' %s=%-12s' "$p" \
"$(kdig +short "@$p" SOA "$d" 2>/dev/null | awk '{print $3}')"
done
printf '\n'
done < <(zone_names)
}
case ${1:-} in
add) shift; cmd_add "$@" ;;
remove) shift; cmd_remove "$@" ;;
list) zone_names ;;
status) shift; knotc zone-status "$@" ;;
check) shift; [[ -n ${1:-} ]] || die "usage: knsctl check <domain>"
knotc -b zone-check "$1" ;;
conf-check) knotc -b conf-check ;;
serials) cmd_serials ;;
path) shift; [[ -n ${1:-} ]] || die "usage: knsctl path <domain>"
find_zone_file "$1" || die "no zone file for $1" ;;
edit) shift; [[ -n ${1:-} ]] || die "usage: knsctl edit <domain>"
f=$(find_zone_file "$1") || die "no zone file for $1"
warn "BREAK-GLASS: edits here are overwritten by the next deploy"
warn " unless the same change is made in the dns repo."
audit "edit $1"
"${EDITOR:-vi}" "$f" ;;
reload) shift
if [[ -n ${1:-} ]]; then
# zone-reload for content; full reload only for config
# changes (a new or removed zone).
knotc -b zone-reload "$1"
else
knotc -b reload
fi ;;
''|-h|--help|help) usage ;;
*) die "unknown command: $1 (try: knsctl help)" ;;
esac
+15
View File
@@ -0,0 +1,15 @@
# secrets.conf.example -> /etc/knot/secrets.conf (0640 root:knot)
#
# NEVER COMMIT THE RENDERED FILE. It is in .gitignore.
# Generate a secret with: keymgr -t <keyname> hmac-sha256
#
# knot.conf includes this before every file that references these key ids.
key:
- id: authortive-tsig # DNS-1 <-> DNS-2 replication (AXFR/IXFR + NOTIFY)
algorithm: hmac-sha256
secret: REPLACE_ME
- id: admin-tsig # read-only AXFR for CI drift-checking and admins
algorithm: hmac-sha256
secret: REPLACE_ME
+26
View File
@@ -0,0 +1,26 @@
;
; BIND zone file for domain: @DOMAIN@
;
$ORIGIN @DOMAIN@.
$TTL 3600
@DOMAIN@. IN SOA dns-1.datacenter.gg. dns.alphacentri.com. (
@SERIAL@ ; Serial -- COLD-START SEED ONLY. Knot owns the live serial
; (zonefile-load: difference-no-serial). Never edit this
; value: CI fails on any change. It is only read when a
; node starts with no journal entry for this zone.
3600 ; Refresh -- fallback if a NOTIFY is missed
900 ; Retry
1209600 ; Expire -- 14d. MUST stay below rrsig-lifetime minus
; rrsig-refresh (23d), or a secondary serves records
; whose signatures have already expired.
300 ; Negative Cache TTL
)
@DOMAIN@. IN NS dns-1.datacenter.gg.
@DOMAIN@. IN NS dns-2.datacenter.gg.
;; Edit Below This Line.
+43
View File
@@ -0,0 +1,43 @@
# Copy to .env and fill in. docker compose picks .env up automatically.
# deploy.sh seeds .env from this on first run; an existing .env is never
# overwritten. Never commit the populated .env.
# ─── Reachability ───────────────────────────────────────────────────────────
# Address (IP or DNS name) the KANRISHA tape host uses to reach this vault on
# the LAN. Seeds the self-signed cert SAN below and is what you point
# [encryption.openbao].address at (https://${OPENBAO_ADDR}:8200).
OPENBAO_ADDR=10.0.0.10
# Interface the published API port binds to on the host. Leave blank/all and
# deploy.sh narrows it to OPENBAO_ADDR when that is an IP; set it explicitly to
# pin a specific LAN IP. (A published port bypasses the host INPUT firewall — see
# the README — so this bind is the main interface restriction.)
OPENBAO_BIND=0.0.0.0
# ─── TLS ────────────────────────────────────────────────────────────────────
# EXTRA SANs for the self-signed cert, beyond OPENBAO_ADDR + loopback (which
# deploy.sh always includes). This is read from the ENVIRONMENT at deploy time,
# so to add names export it before running deploy.sh, e.g.
# OPENBAO_TLS_SANS=DNS:vault.lan,IP:10.0.0.11 bash deploy.sh
# (deploy.sh records the final SAN list back into this .env for reference.) To
# use a CA-signed cert instead (e.g. Smallstep over ACME), drop tls.crt + tls.key
# into ./tls and this is ignored — see the README.
OPENBAO_TLS_SANS=
OPENBAO_TLS_DAYS=825
# ─── Web UI ─────────────────────────────────────────────────────────────────
# 1 = serve the built-in web UI on the same listener (default), 0 = API only.
# Not a new exposure: anything that can reach :8200 can already do everything
# via the API. With the UI on you can initialise and unseal from a browser,
# which keeps the unseal keys out of a server shell's history.
OPENBAO_UI=1
# ─── Image tag ──────────────────────────────────────────────────────────────
# Pin for reproducible deploys.
OPENBAO_TAG=2.6.2
# ─── Auto-unseal (optional; default is MANUAL unseal) ───────────────────────
# Only used when the seal "pkcs11" stanza is enabled in config.hcl (and the HSM
# module/device is mounted into the container). Otherwise leave blank and unseal
# manually after each restart.
OPENBAO_HSM_PIN=
+6
View File
@@ -0,0 +1,6 @@
# Never commit generated TLS material (the listener key), a populated .env, or
# raft snapshot backups (they contain vault data).
/tls/
.env
*.snap
*.snap.age
+293
View File
@@ -0,0 +1,293 @@
# openbao
Hardened [OpenBao](https://openbao.org) — the tape-encryption key store for
**Kanrisha** (the LTO tape-archive system; separate repo, separate host).
Deliberately **separate from the tape host**: a compromise of the tape node must
not reach the vault, and OpenBao's TLS/unseal lifecycle is cleaner on its
own box.
Unlike the other stacks here there is **no Caddy / Let's Encrypt** — a secrets
store terminates TLS itself and is reached over the **LAN**, not the public
internet.
- **Native TLS** on the listener — self-signed by default, or a CA-signed cert
from your Smallstep CA over ACME.
- **Integrated raft storage** — snapshot-based DR.
- **Swap disabled** — key material never hits disk. (Not mlock: OpenBao removed
mlock support, so `disable_mlock` must not appear in config.hcl at all.)
- **Manual unseal by default**, or PKCS#11 HSM auto-unseal.
> ⚠️ This vault is the **sole recovery path for encrypted tapes.** Losing the
> OpenBao data *and* the unseal keys/root token loses every encrypted tape. Take
> raft snapshots and store the unseal material out of band (below).
## Required `.env` values
| Variable | Notes |
|----------|-------|
| `OPENBAO_ADDR` | IP/DNS the Kanrisha tape host uses to reach this vault (goes in the cert SAN; you point the daemon at `https://$OPENBAO_ADDR:8200`). |
| `OPENBAO_BIND` | Host interface the API publishes on. Left at the default, deploy.sh narrows it to `OPENBAO_ADDR` when that is an IP (a published port bypasses the INPUT firewall, so this bind is the real restriction — see [Exposure](#exposure)). It must be an address this host actually holds — deploy.sh checks that before writing anything and refuses otherwise (`SKIP_BIND_CHECK=1` to override). |
| `OPENBAO_TLS_SANS` | **Extra** SANs beyond `OPENBAO_ADDR` + loopback (which deploy.sh always adds). Read from the environment at deploy time — `export` it before running deploy.sh. |
| `OPENBAO_TAG` | OpenBao image tag (pin it). |
| `OPENBAO_HSM_PIN` | Only for PKCS#11 auto-unseal. Leave blank for manual unseal. |
See [`.env.example`](.env.example) for the full list.
## Deploy
```bash
./automations.sh # Deploy on this host → deploy: openbao
# or, non-interactive:
OPENBAO_ADDR=10.0.0.10 SKIP_PROMPTS=1 bash deployments/openbao/deploy.sh
```
`deploy.sh` installs Docker, generates a self-signed TLS cert (if none present),
seeds `.env`, disables swap (keeping key material off disk), narrows the API
bind to the LAN IP,
opens `8200/tcp`, aligns file/volume ownership to the container's UID, and brings
the stack up. OpenBao starts **sealed** — initialise + unseal once (do this
immediately; an uninitialised vault reachable on the LAN can be init'd by anyone
who connects):
```bash
docker compose exec -e BAO_ADDR=https://127.0.0.1:8200 openbao \
bao operator init -tls-skip-verify # prints 5 unseal keys + root token
docker compose exec -e BAO_ADDR=https://127.0.0.1:8200 openbao \
bao operator unseal -tls-skip-verify # x3, prompts hidden each time
```
Pass **no key argument** — `unseal` prompts with echo off. A key given on the
command line lands in the docker process's argv, which `/proc/<pid>/cmdline`
exposes to every user on the host; three of them reconstruct the master key.
**Store the unseal keys + root token out of band** — ideally age-encrypted with
your backup recipient (`globals/age-pubkey.txt`), never on this host.
### Re-running after a wrong address
`deploy.sh` is idempotent, but two things it writes are deliberately **sticky**:
`.env` (never overwritten) and `tls/tls.{crt,key}` (never regenerated over an
existing pair). So passing a corrected `OPENBAO_ADDR` to a re-run does *not* move
the vault — the cert keeps the old SAN, and the old `.env` usually still decides
the bind. Which `OPENBAO_BIND` wins follows Compose's own precedence:
| How `OPENBAO_BIND` is set | What Compose uses |
|---------------------------|-------------------|
| Exported into deploy.sh's environment (`OPENBAO_BIND=… bash deploy.sh`, or answering the bind prompt — `automations.sh` passes answers via `env VAR=…`) | the **environment** value; `.env` is ignored for this run, so a later hand-run `docker compose up -d` can bind somewhere else |
| Derived by deploy.sh (bind prompt left blank → narrowed to `OPENBAO_ADDR`) | the **`.env`** value, since that assignment is never exported |
| Absent from both | `0.0.0.0` — **all interfaces** |
deploy.sh warns on each of those mismatches. To actually change the address:
```bash
cd /srv/openbao
docker compose down # keeps the raft volume
sed -i 's/OLD_IP/NEW_IP/g' .env # OPENBAO_ADDR + OPENBAO_BIND
rm -f tls/tls.crt tls/tls.key # force a new SAN
# then re-run deploy.sh with the corrected OPENBAO_ADDR
```
A bind address the host does not hold is caught up front; without that check
Docker fails the `up` with `cannot assign requested address` only *after* the
bad value is already in `.env` and the cert.
## TLS
- **Self-signed (default):** `deploy.sh` runs [`gen-tls.sh`](gen-tls.sh) to
create `./tls/tls.{crt,key}` with `OPENBAO_ADDR` in the SAN. Hand `tls.crt` to
the Kanrisha daemon as `[encryption.openbao].ca_cert`.
- **Smallstep CA over ACME (option):** issue a cert from your `step-ca` and drop
it in `./tls` instead — `gen-tls.sh` then no-ops. e.g. with the `step` client:
```bash
step ca certificate "$OPENBAO_ADDR" ./tls/tls.crt ./tls/tls.key \
--provisioner acme --acme https://ca.lan/acme/acme/directory
# renew on a timer: step ca renew --daemon ./tls/tls.crt ./tls/tls.key
```
Give the Kanrisha daemon your Smallstep **root** as `ca_cert` (then it
trusts the vault without `-tls-skip-verify`).
## Auto-unseal (optional)
Default is manual unseal after each restart. For hands-off restarts, enable the
`seal "pkcs11"` stanza in [`config.hcl`](config.hcl), mount the PKCS#11 module +
device into the `openbao` service, and set `OPENBAO_HSM_PIN` in `.env`.
## Bootstrap for Kanrisha
Once unsealed, enable KV v2 + AppRole and seed the domain key with the Kanrisha
bootstrap script (from the Kanrisha repo, `deploy/openbao/bootstrap.sh`). It calls
the `bao` CLI directly, so run it **from a host that has `bao`** (the Kanrisha
host or your workstation) pointed at this vault — this vault host only ships
Docker. Copy `tls.crt` to that host first and pass it as `BAO_CACERT`:
```bash
BAO_ADDR=https://$OPENBAO_ADDR:8200 BAO_CACERT=/path/to/openbao-ca.crt \
bash bootstrap.sh # export BAO_TOKEN beforehand rather than passing it inline
```
It enables the `kanrisha-tape` KV-v2 mount (with **effectively-unlimited**
`max_versions` so a key rotation never orphans old tapes), creates the AppRole +
policy, seeds the domain key, and prints the `[encryption.openbao]` block for the
Kanrisha config. Point the daemon at `address = "https://$OPENBAO_ADDR:8200"`.
## Exposure
This is a secrets store on the LAN, not a public service. Two things to know:
- **A Docker-published port is DNAT'd and bypasses the host `INPUT` firewall.** A
source-restricting rule on `INPUT` does **not** gate `:8200`. The real interface
restriction is the publish bind: `deploy.sh` defaults `OPENBAO_BIND` to
`OPENBAO_ADDR` when that is an IP, so the API listens only on that LAN IP. To
restrict by *source host*, use a `FORWARD`/`DOCKER-USER` rule or network
segmentation, not `INPUT`.
- **On a Proxmox guest, the per-VM firewall filters before the guest sees
anything.** If the VM's NIC has `firewall=1` and the datacenter firewall is
enabled, a Docker-published port needs its own `IN ACCEPT -p tcp -dport 8200`
rule in `/etc/pve/firewall/<vmid>.fw`. SSH working does **not** prove the path
— it only proves there is a rule for 22. The give-away is that everything
inside the guest looks perfect (`curl` to the bind address answers, the DNAT
and `FORWARD` jumps are present) while `tcpdump -ni eth0 'tcp port 8200'`
captures zero packets during a failed connection.
- **mTLS (optional):** to require the tape host to present a client cert, enable
the `tls_require_and_verify_client_cert` stanza in [`config.hcl`](config.hcl)
and issue the tape host a client cert from the same CA.
And **initialise the vault immediately** after `deploy.sh` — an uninitialised
vault reachable on the LAN can be `bao operator init`'d by anyone who connects,
handing them the root token and unseal keys.
### Reaching it without LAN access
The API publishes on `OPENBAO_BIND` only, so a browser on another subnet (or
behind a firewall you do not control) cannot reach it. Tunnel over SSH instead
of widening the publish — the generated cert already carries `DNS:localhost`
and `IP:127.0.0.1` in its SANs, so it validates as-is:
```bash
ssh -L 8200:<bind-addr>:8200 root@<host>
# then browse https://localhost:8200
```
This is also the better way to do the first `operator init`: the unseal keys and
root token are shown in your browser instead of a root shell's scrollback.
## Running bao commands
The host has no `bao` CLI, only Docker, so every authenticated command is a
`docker compose exec`. Getting a token in there naively leaks it: `-e
BAO_TOKEN=<value>` puts it in the docker process's argv, which
`/proc/<pid>/cmdline` exposes to every user on the box.
`baoctl` is installed alongside the stack for this. It is a **session** wrapper:
one hidden prompt, then your commands are typed verbatim.
```bash
cd /srv/openbao
bash baoctl login # prompts with echo off; nothing is echoed or stored on the host
bash baoctl secrets enable pki
bash baoctl write pki/root/generate/internal common_name="Example Root CA" key_bits=4096
bash baoctl policy write my-policy - < my-policy.hcl
bash baoctl logout
```
The token you type is **not** what the session holds. After login `baoctl` mints
a short-lived child token and swaps it in, so the session expires by itself and
`logout` revokes it without touching your root token. The session lives on the
container's `/dev/shm` (already a tmpfs), so it never reaches disk and dies with
the container.
Why a session rather than piping the token per command: `bao` needs stdin for
`policy write NAME -`, `write PATH -` and `key=-`. Spend stdin on the token and
a secret value like an OIDC client secret has nowhere left to go but argv —
reintroducing the leak. `update.sh` still pipes, because it runs one command
unattended and needs no stdin of its own.
`baoctl` verifies TLS rather than skipping it: `./tls` is mounted into the
container and the generated cert carries `IP:127.0.0.1` in its SANs, so
`BAO_CACERT` validates against the real listener cert.
## Updating
`update.sh` is installed alongside the stack. It is **not** run by `deploy.sh`,
because upgrading a live vault seals it — that is an operator's decision, not a
deploy step.
```bash
cd /srv/openbao
bash update.sh check # declared / running / latest + seal state
bash update.sh update # snapshot, then upgrade (prompts)
```
Unlike the copyparty and ergo updaters in this repo, this one **will not update
on a schedule by default**. Those services come back by themselves; OpenBao
comes back *sealed*, so an unattended 03:00 update would take the vault offline
until someone arrives with three unseal keys. `update.sh install` therefore
schedules a daily **check** (`UPDATE_POLICY=notify`). `UPDATE_POLICY=auto` opts
into unattended updates and is still refused unless a `seal` stanza is
configured — only auto-unseal makes the vault come back on its own.
Before it changes anything, `update` requires that the container is running, the
vault is **unsealed** (a sealed vault cannot produce a snapshot, so there would
be no rollback plan), and a token with `sys/storage/raft/snapshot` — the root
token works. It **prompts** for that token with echo off; never pass it as
`BAO_TOKEN=... bash update.sh`, which puts it in your shell history and in
`/proc/<pid>/cmdline`, where every user on the box can read it. For unattended
use set `BAO_TOKEN_FILE` to a `0600` file. It writes the snapshot to `/var/backups/openbao` at `0600`,
verifies it is a valid non-empty gzip archive, and refuses to continue if it is
not. **Copy it off the host**: it is the rollback plan, and OpenBao's upgrade
guide is explicit that reverting the image alone does not roll back the data
store. `SKIP_SNAPSHOT=1` exists and says loudly what you are giving up.
If the pull or the start fails, the `OPENBAO_TAG` pin is rolled back and the
previous version is started again — still sealed.
`update.sh` also refuses to cross into 2.7.x while an active built-in
`seal "pkcs11"` stanza is present: that stanza is *removed* in 2.7.0, not merely
deprecated, so the vault would come up with no way to unseal at all. Migrate to
the external `plugin "kms" "pkcs11"` first.
## Backup / DR
The vault is the sole recovery path for encrypted tapes — back it up:
Snapshot save/restore are **token-gated** (`sys/storage/raft/snapshot` is
sudo-capable) — they need a token that has that path (the root token works, or
mint a dedicated snapshot-policy token). The container has no ambient token.
**Do not put the token on a command line.** `docker ... -e BAO_TOKEN=<value>`
places it in the docker process's argv, and `/proc/<pid>/cmdline` is readable by
every user on the host — so it leaks further than your shell history does. Use
`update.sh snapshot`, which prompts with echo off and passes the token to the
container over stdin:
```bash
cd /srv/openbao
bash update.sh snapshot # prompts; writes to /var/backups/openbao, 0600
```
Then age-encrypt it and move it off-box:
```bash
age -r "$(cat /path/to/globals/age-pubkey.txt)" \
< /var/backups/openbao/openbao-pre-manual-*.snap \
> "openbao-$(date +%F).snap.age"
```
Restore (into a fresh, unsealed vault): decrypt, copy the raw `.snap` in, and
apply it (also token-gated):
```bash
age -d -i ~/.age/key.txt openbao-DATE.snap.age > openbao.snap
docker compose cp openbao.snap openbao:/tmp/openbao.snap
# Same argv problem as above, so hand the token over stdin rather than -e:
read -rs -p "token: " T; echo
printf '%s\n' "$T" | docker compose exec -T openbao sh -c \
'read -r t; BAO_TOKEN="$t" exec bao operator raft snapshot restore \
-address=https://127.0.0.1:8200 -tls-skip-verify /tmp/openbao.snap'
unset T
```
Snapshots do **not** contain the unseal keys or root token — you still need
those to unseal a restored vault, which is why they are stored separately.
+224
View File
@@ -0,0 +1,224 @@
#!/usr/bin/env bash
#
# baoctl -- run `bao` commands against the containerised vault without leaking
# credentials. Companion to deploy.sh, installed alongside the stack.
#
# The problem it solves: this host has no `bao` CLI, only Docker, so every
# authenticated command is a `docker compose exec`. The obvious ways to get a
# token in there all leak it:
# * `-e BAO_TOKEN=<value>` puts it in the docker process's argv, and
# /proc/<pid>/cmdline is world-readable -- every user on the box can read it.
# * `BAO_TOKEN=<value> some-command` additionally puts it in shell history.
# * piping the token to the container spends STDIN, and bao needs stdin for
# `policy write NAME -`, `write PATH -` (a JSON body) and `key=-` (a single
# secret value). Spend it on the token and an OIDC client secret has nowhere
# left to go but argv -- reintroducing the first leak.
#
# So baoctl is a SESSION wrapper, not a per-command token pump. `baoctl login`
# prompts once with echo off, stores a token inside the container, and every
# later command is typed verbatim with stdin free.
#
# The session token is NOT the one you type. After login, baoctl mints a
# short-lived child token and swaps it in, so what sits in the session expires
# on its own and `logout` can revoke it without touching your root token.
#
# The session lives at /dev/shm/.bao-session inside the container (BAO_TOKEN_PATH
# points bao's own token helper there). /dev/shm is already a per-container
# tmpfs, so the token never touches disk and dies with the container -- and no
# compose change was needed to arrange that, which matters because recreating
# this container means a seal cycle and three unseal keys typed by a human.
#
# Subcommands:
# login [--ttl 8h] [--policy NAME] prompt (hidden), then mint + store a child
# logout [--keep-token] revoke the session token AND remove it
# status seal state, and what the session can do
# <anything else> passed straight to `bao` in the container
#
# Usage:
# bash baoctl login
# bash baoctl write pki/root/generate/internal common_name="Example Root CA" ...
# bash baoctl policy write my-policy - < my-policy.hcl
# bash baoctl status
# bash baoctl logout
#
# Env:
# STACK_DIR=/srv/openbao SESSION_TTL=8h SESSION_POLICY=
set -euo pipefail
: "${STACK_DIR:=/srv/openbao}"
: "${SESSION_TTL:=8h}"
: "${SESSION_POLICY:=}"
# Inside the container. /dev/shm is tmpfs, so the session never hits disk.
TOKEN_PATH="/dev/shm/.bao-session"
ADDR_IN="https://127.0.0.1:8200"
CACERT_IN="/openbao/tls/tls.crt"
log() { printf '\033[1;32m[+]\033[0m %s\n' "$*"; }
warn() { printf '\033[1;33m[!]\033[0m %s\n' "$*" >&2; }
die() { printf '\033[1;31m[x]\033[0m %s\n' "$*" >&2; exit 1; }
# Before the stack check, so --help works anywhere and touches no docker.
case "${1:-}" in
-h|--help|help)
awk '/^set -euo/ { exit } NR > 1 { sub(/^# ?/, ""); print }' "$0"
exit 0 ;;
esac
[[ -d "$STACK_DIR" ]] || die "No stack at $STACK_DIR (set STACK_DIR)."
dc() { ( cd "$STACK_DIR" && docker compose "$@" ); }
# BAO_TOKEN in the environment beats the token file, always. baoctl never passes
# it through -- but if the operator has one set they will assume it is in play,
# and be debugging the wrong credential.
if [[ -n "${BAO_TOKEN:-}" ]]; then
warn "BAO_TOKEN is set in your shell. baoctl ignores it and uses its own session;"
warn "unset it to avoid confusion about which credential is being used."
fi
# Verify TLS properly rather than reaching for -tls-skip-verify: deploy.sh mounts
# ./tls read-only into the container and puts IP:127.0.0.1 in the cert's SANs, so
# pointing BAO_CACERT at it actually validates. Only fall back to skipping when
# the cert genuinely is not there.
bao_env() {
printf '%s\n' "-e" "BAO_ADDR=${ADDR_IN}" "-e" "BAO_TOKEN_PATH=${TOKEN_PATH}"
if [[ -f "$STACK_DIR/tls/tls.crt" ]]; then
printf '%s\n' "-e" "BAO_CACERT=${CACERT_IN}"
else
printf '%s\n' "-e" "BAO_SKIP_VERIFY=true"
fi
}
bao_run() { # bao_run <bao args...>
# Allocate a TTY only when our own stdin is one. An interactive command wants
# one; `baoctl write path - < body.json` must NOT have one, or stdin is not
# forwarded and the payload never arrives.
local -a envs=() tf=()
while IFS= read -r line; do envs+=("$line"); done < <(bao_env)
[[ -t 0 ]] || tf=(-T)
dc exec "${tf[@]}" "${envs[@]}" openbao bao "$@"
}
session_exists() {
local -a envs=()
while IFS= read -r line; do envs+=("$line"); done < <(bao_env)
dc exec -T "${envs[@]}" openbao sh -c '[ -s "$BAO_TOKEN_PATH" ]' >/dev/null 2>&1
}
do_login() {
local ttl="$SESSION_TTL" policy="$SESSION_POLICY"
while [[ $# -gt 0 ]]; do
case "$1" in
# `shift 2` with only one argument left FAILS, and under set -e that
# exits the script silently -- before any validation below runs.
# So check the count first rather than relying on ${2:-}.
--ttl) [[ $# -ge 2 ]] || die "login: --ttl needs a value (e.g. 8h)."
ttl="$2"; shift 2 ;;
--policy) [[ $# -ge 2 ]] || die "login: --policy needs a policy name."
policy="$2"; shift 2 ;;
*) die "login: unknown option '$1' (expected --ttl or --policy)." ;;
esac
done
[[ -n "$ttl" ]] || die "login: --ttl needs a value (e.g. 8h)."
# Checked after option parsing, so a typo'd flag is reported either way.
[[ -t 0 ]] || die "login needs a terminal: bao prompts for the token with echo off, which requires a TTY. Run it from an interactive shell."
local -a envs=()
while IFS= read -r line; do envs+=("$line"); done < <(bao_env)
# Step 1: the interactive login. -no-print is not optional -- without it bao
# prints the token in its success table, dumping whatever you just typed into
# this terminal's scrollback, which is worse than what we are replacing.
log "Logging in (the prompt is hidden; nothing is echoed or stored on the host)..."
dc exec "${envs[@]}" openbao bao login -no-print \
|| die "Login failed. The token was rejected, or the vault is sealed."
# Step 2: swap the token you typed for a short-lived child, so the session
# expires by itself and `logout` can revoke it without killing your root
# token. The child never reaches an argv or a stdout: it goes straight down a
# pipe into `bao login -`, which reads the token from stdin.
log "Minting a ${ttl} session token${policy:+ with policy '${policy}'}..."
if ! dc exec -T "${envs[@]}" openbao sh -c '
ttl="$1"; pol="$2"
if [ -n "$pol" ]; then
bao token create -ttl="$ttl" -policy="$pol" -field=token
else
bao token create -ttl="$ttl" -field=token
fi | bao login -no-print -
' sh "$ttl" "$policy"; then
# Do not leave the typed token sitting in the session: the whole point of
# the swap is that a session is disposable. Clear it and fail.
dc exec -T "${envs[@]}" openbao sh -c 'rm -f "$BAO_TOKEN_PATH"' >/dev/null 2>&1 || true
die "Could not mint a session token; the login has been discarded. If the token you used cannot create child tokens, pass --policy with one it can grant."
fi
log "Session ready. Run commands with: bash $(basename "$0") <bao args...>"
do_status || true
}
do_logout() {
local keep=0
while [[ $# -gt 0 ]]; do
case "$1" in
--keep-token) keep=1; shift ;;
*) die "logout: unknown option '$1'." ;;
esac
done
local -a envs=()
while IFS= read -r line; do envs+=("$line"); done < <(bao_env)
if ! session_exists; then
log "No session to end."
return 0
fi
# Both steps, deliberately. `bao token revoke -self` does not remove the token
# file and there is no `bao logout` in 2.6.x, so revoking alone leaves a stale
# file behind that fails with permission errors instead of "not logged in".
if [[ "$keep" == 0 ]]; then
dc exec -T "${envs[@]}" openbao bao token revoke -self >/dev/null 2>&1 \
|| warn "Could not revoke the session token (already expired?); removing it anyway."
fi
dc exec -T "${envs[@]}" openbao sh -c 'rm -f "$BAO_TOKEN_PATH"' >/dev/null 2>&1 || true
log "Session ended."
}
do_status() {
local -a envs=()
while IFS= read -r line; do envs+=("$line"); done < <(bao_env)
local rc=0
dc exec -T "${envs[@]}" openbao bao status >/dev/null 2>&1 || rc=$?
case "$rc" in
0) printf ' vault: unsealed\n' ;;
2) printf ' vault: SEALED (unseal before anything else works)\n' ;;
*) printf ' vault: unreachable (bao status rc=%s)\n' "$rc" ;;
esac
if ! session_exists; then
printf ' session: none -- run `%s login`\n' "$(basename "$0")"
return 0
fi
printf ' session: present at %s (in-container tmpfs)\n' "$TOKEN_PATH"
dc exec -T "${envs[@]}" openbao bao token lookup -format=json 2>/dev/null \
| awk -F'"' '
/"display_name"/ { printf " name: %s\n", $4 }
/"ttl"/ { gsub(/[^0-9]/, "", $0); if ($0 != "") printf " ttl: %ss\n", $0 }
' || warn "Session token present but lookup failed; it may have expired. Run logout, then login."
}
case "${1:-status}" in
login) shift; do_login "$@" ;;
logout) shift; do_logout "$@" ;;
status) shift; do_status "$@" ;;
*)
# Everything else is a bao command. Requiring a session first gives a
# clear error instead of bao's "missing client token".
session_exists || die "No session. Run: bash $(basename "$0") login"
bao_run "$@"
;;
esac
@@ -1,8 +1,8 @@
#!/usr/bin/env bash
#
# build.sh -- (re)embed docker-compose.yml, Caddyfile, .env.example into
# deploy.sh as a base64-encoded tar.gz payload after __ARCHIVE_BELOW__.
# Idempotent: strips any existing payload first.
# build.sh -- (re)embed docker-compose.yml, config.hcl, gen-tls.sh and
# .env.example into deploy.sh as a base64-encoded tar.gz payload after
# __ARCHIVE_BELOW__. Idempotent: strips any existing payload first.
set -euo pipefail
@@ -11,11 +11,11 @@ SCRIPT="$DIR/deploy.sh"
MARKER="__ARCHIVE_BELOW__"
[[ -f "$SCRIPT" ]] || { echo "deploy.sh not found at $SCRIPT" >&2; exit 1; }
for f in docker-compose.yml Caddyfile .env.example; do
for f in docker-compose.yml config.hcl gen-tls.sh update.sh baoctl .env.example; do
[[ -f "$DIR/$f" ]] || { echo "Missing $DIR/$f" >&2; exit 1; }
done
PAYLOAD=$(tar -czf - -C "$DIR" docker-compose.yml Caddyfile .env.example | base64)
PAYLOAD=$(tar -czf - -C "$DIR" docker-compose.yml config.hcl gen-tls.sh update.sh baoctl .env.example | base64)
TMP=$(mktemp)
trap 'rm -f "$TMP"' EXIT
+37
View File
@@ -0,0 +1,37 @@
#cloud-config
#
# OpenBao (Kanrisha tape-encryption key store) — harden SSH, then deploy, on a
# fresh Alpine/Debian/Alma host on the SAME LAN as the Kanrisha tape host.
#
# Fill in REPO_URL and the values in the runcmd block, then paste this as the
# instance user-data. No public DNS / 80 / 443 needed — this is a LAN vault with
# native TLS on 8200. After boot, initialise + unseal once (see the README).
packages:
- git
runcmd:
- hostnamectl set-hostname openbao || true
- |
set -e
REPO_URL=https://git.anomalous.dev/57_Wolve/automations.git
REPO_BRANCH=main
HARDEN_SSH=1 # harden SSH on this fresh VM (set 0 to skip)
SSH_PORT=22
ALLOWED_IP= # optional: whitelist your client IP in sshguard
git clone --depth 1 --branch "$REPO_BRANCH" "$REPO_URL" /opt/automations
cd /opt/automations
# Harden SSH: PQ KEX, key-only auth, sshguard. Seeds root from
# globals/authorized_keys (or SSH_KEYS_URL).
if [ "$HARDEN_SSH" = 1 ]; then
SSH_PORT="$SSH_PORT" ALLOWED_IP="$ALLOWED_IP" SKIP_PROMPTS=1 FORCE=1 \
bash scripts/harden-ssh.sh
fi
# Deploy OpenBao. OPENBAO_ADDR is the LAN address the Kanrisha tape host
# reaches this vault at — it goes in the TLS cert SAN. deploy.sh disables
# swap (for mlock) and firewalls 8200/tcp.
OPENBAO_ADDR=10.0.0.10 \
SKIP_PROMPTS=1 \
bash deployments/openbao/deploy.sh
+64
View File
@@ -0,0 +1,64 @@
# OpenBao — hardened, same-LAN tape-encryption key store for Kanrisha.
#
# Native TLS on the listener (no reverse proxy in front), integrated raft
# storage (clean snapshot-based DR). This vault is the SOLE
# recovery path for encrypted tapes — losing the OpenBao data AND the unseal
# material loses every encrypted tape. Back it up (raft snapshots) and store the
# unseal keys/root token out of band (see the README + DR runbook).
storage "raft" {
path = "/openbao/data"
node_id = "openbao-1"
}
# TLS-terminating listener. tls.crt/tls.key are self-signed by deploy.sh on first
# run, or a CA-signed pair you drop in ./tls (e.g. issued by your Smallstep CA
# over ACME — see the README). The Kanrisha daemon trusts this via
# [encryption.openbao].ca_cert.
listener "tcp" {
address = "0.0.0.0:8200"
tls_cert_file = "/openbao/tls/tls.crt"
tls_key_file = "/openbao/tls/tls.key"
# Optional mTLS — require + verify a client cert from the tape host:
# tls_require_and_verify_client_cert = true
# tls_client_ca_file = "/openbao/tls/client-ca.crt"
}
# Keeping key material off disk is done by DISABLING SWAP (deploy.sh does this),
# NOT by mlock. OpenBao removed mlock support: its own config parser carries the
# message "OpenBao has dropped support for mlock. Please remove the line
# \"disable_mlock\" = false from your config and disable or encrypt swap instead."
# So do NOT add a disable_mlock line of either value -- `disable_mlock = false`
# is a hard startup error. cap_add: [IPC_LOCK] and ulimits memlock in
# docker-compose.yml are inert leftovers.
# Single-node raft: advertising 127.0.0.1 is fine because a lone node is always
# ACTIVE and never redirects clients (they connect straight to the LAN address).
# For HA / multiple nodes, set BAO_API_ADDR + BAO_CLUSTER_ADDR to each node's
# reachable LAN address in the compose env instead.
api_addr = "https://127.0.0.1:8200"
cluster_addr = "https://127.0.0.1:8201"
# The built-in web UI. Substituted by deploy.sh from OPENBAO_UI. It adds NO new
# exposure: it is served on this same listener, and anything reachable on :8200
# can already do everything through the API. Turning it on is what lets an
# operator initialise/unseal and log in from a browser, so the unseal keys are
# shown in the browser instead of a server shell's scrollback.
ui = @UI@
# ── Auto-unseal (optional) — DEFAULT IS MANUAL UNSEAL ────────────────────────
# DEPRECATED: the BUILT-IN pkcs11 seal below is removed in OpenBao v2.7.0, and
# the HSM distribution is discontinued by then. PKCS#11 auto-unseal continues
# via the external KMS plugin (`plugin "kms" "pkcs11" {}`), so a stack that
# needs it must move before taking 2.7.x.
# With no seal stanza OpenBao starts SEALED and needs `bao operator unseal` (x3)
# after every restart. To auto-unseal against a PKCS#11 HSM/token, uncomment and
# fill this in, mount the PKCS#11 module + device into the container, and set
# OPENBAO_HSM_PIN in .env:
#
# seal "pkcs11" {
# lib = "/usr/lib/softhsm/libsofthsm2.so"
# slot = "0"
# pin = "env://BAO_HSM_PIN"
# key_label = "kanrisha-unseal"
# mechanism = "0x1087" # CKM_AES_GCM
# }
+793
View File
@@ -0,0 +1,793 @@
#!/usr/bin/env bash
#
# deploy.sh -- deploy the hardened OpenBao tape-encryption key store on a
# same-LAN host (Alpine / Debian / Alma). Single-node, runs as root.
#
# What this does:
# 1. Installs docker + compose if missing.
# 2. Lays down docker-compose.yml, config.hcl, gen-tls.sh, update.sh, baoctl
# in $STACK_DIR.
# 3. Seeds .env on first run (OPENBAO_ADDR into the cert SAN); never
# overwrites an existing .env.
# 4. Generates a self-signed TLS cert (if ./tls is empty) -- drop a CA-signed
# pair there instead to use your Smallstep CA.
# 5. Disables swap (so mlock is meaningful) and opens 8200/tcp.
# 6. Pulls images and brings the stack up. OpenBao starts SEALED -- initialise
# + unseal once afterwards (printed at the end).
#
# Idempotent: re-run to apply config changes / pull new images.
#
# Self-contained: docker-compose.yml, config.hcl, gen-tls.sh, update.sh, baoctl and .env.example are
# embedded as a base64 tar.gz at the bottom. Rebuild with build.sh after edits.
#
# Usage:
# OPENBAO_ADDR=10.0.0.10 bash deploy.sh # interactive prompt for the rest
# OPENBAO_ADDR=vault.lan SKIP_PROMPTS=1 bash deploy.sh
# STACK_DIR=/opt/openbao bash deploy.sh
set -euo pipefail
: "${STACK_DIR:=/srv/openbao}"
: "${SKIP_DOCKER_INSTALL:=0}"
: "${SKIP_BIND_CHECK:=0}" # 1 = publish on an address this host does not (yet) have
: "${FORCE:=0}"
: "${SKIP_PROMPTS:=0}" # non-interactive: require values via env, no prompts
[[ "$SKIP_PROMPTS" == "1" ]] && FORCE=1
: "${OPENBAO_ADDR:=}"
# Whether OPENBAO_BIND arrived in this script's ENVIRONMENT (automations.sh
# passes answers via `env VAR=... bash deploy.sh`, and a standalone run may
# export it too). If it did it is still exported when we reach compose, which
# prefers the shell environment over $STACK_DIR/.env; if this script derives it
# below instead, the assignment is NOT exported and the .env wins. Must be read
# before the := default, which would make an unset var look set.
BIND_FROM_ENV=0
[[ -n "${OPENBAO_BIND+x}" ]] && BIND_FROM_ENV=1
: "${OPENBAO_BIND:=0.0.0.0}"
# Same question for the UI switch: an explicit OPENBAO_UI=0 has to be told
# apart from "not mentioned", or a re-run would silently re-enable the UI.
UI_FROM_ENV=0
[[ -n "${OPENBAO_UI+x}" ]] && UI_FROM_ENV=1
: "${OPENBAO_UI:=}"
: "${DISABLE_SWAP:=1}" # set 0 to skip swapoff (mlock then only best-effort)
log() { printf '\033[1;32m[+]\033[0m %s\n' "$*"; }
warn() { printf '\033[1;33m[!]\033[0m %s\n' "$*" >&2; }
die() { printf '\033[1;31m[x]\033[0m %s\n' "$*" >&2; exit 1; }
[[ $EUID -eq 0 ]] || die "Run as root."
# ---------------------------------------------------------------------------
# OS detection + Docker install (Alpine / Debian / Alma). Inlined so this
# deploy.sh stays self-contained when scp'd standalone.
# ---------------------------------------------------------------------------
osfam() {
local id="" like=""
if [[ -r /etc/os-release ]]; then
id="$(. /etc/os-release 2>/dev/null && echo "${ID:-}")"
like="$(. /etc/os-release 2>/dev/null && echo "${ID_LIKE:-}")"
fi
case " $id $like " in
*" alpine "*) echo alpine ;;
*" debian "*|*" ubuntu "*) echo debian ;;
*" rhel "*|*" fedora "*|*" centos "*) echo rhel ;;
*) echo "${id:-unknown}" ;;
esac
}
install_docker() {
[[ "$SKIP_DOCKER_INSTALL" == "1" ]] && { log "Skipping Docker install."; return; }
if command -v docker >/dev/null 2>&1; then
log "Docker already installed: $(docker --version)"
else
log "Installing Docker (OS: $(osfam))..."
case "$(osfam)" in
alpine) apk add -q docker docker-cli-compose openrc ;;
debian|rhel) command -v curl >/dev/null 2>&1 || \
{ command -v apt-get >/dev/null 2>&1 && apt-get install -y -qq curl; } || \
{ command -v dnf >/dev/null 2>&1 && dnf install -y -q curl; }
curl -fsSL https://get.docker.com | sh ;;
*) die "Unsupported OS for auto Docker install. Set SKIP_DOCKER_INSTALL=1 and install Docker yourself." ;;
esac
fi
if command -v rc-update >/dev/null 2>&1; then
rc-update add docker default >/dev/null 2>&1 || true
rc-service docker status >/dev/null 2>&1 || rc-service docker start
elif command -v systemctl >/dev/null 2>&1; then
systemctl enable --now docker >/dev/null 2>&1 || systemctl start docker || true
fi
# dockerd is often started in the background (esp. openrc) and returns before
# the socket is listening -- poll so the first `docker compose` call doesn't
# race it and abort under set -e.
local i
for i in $(seq 1 30); do
docker info >/dev/null 2>&1 && return
sleep 1
done
warn "Docker daemon not ready after 30s; continuing (compose may fail -- check 'docker info')."
}
open_bao_port() {
# Register 8200/tcp (the vault API). Prefer the host firewall when present;
# else ufw/firewalld if active. Restrict the source to the tape host where
# you can -- this is a secrets store, not a public service.
if [[ -d /etc/firewall/ports.d && -x /usr/local/sbin/firewall-apply ]]; then
log "Registering 8200/tcp with host firewall..."
printf '8200/tcp\n' > /etc/firewall/ports.d/openbao.rule
/usr/local/sbin/firewall-apply
elif command -v ufw >/dev/null 2>&1 && ufw status 2>/dev/null | grep -q '^Status: active'; then
log "ufw active -- allowing 8200/tcp..."
ufw allow 8200/tcp >/dev/null
elif command -v firewall-cmd >/dev/null 2>&1 && firewall-cmd --state >/dev/null 2>&1; then
log "firewalld active -- allowing 8200/tcp..."
firewall-cmd -q --add-port=8200/tcp --permanent
firewall-cmd -q --reload
fi
}
# Addresses currently assigned to this host, one per line. Parses plain
# `ip addr show` output -- no -o/scope filters, since busybox ip (what Alpine
# ships by default) supports neither -- with an ifconfig fallback.
host_addrs() {
if command -v ip >/dev/null 2>&1; then
ip addr show 2>/dev/null | awk '$1=="inet"||$1=="inet6"{split($2,a,"/"); print a[1]}'
elif command -v ifconfig >/dev/null 2>&1; then
ifconfig 2>/dev/null | awk '$1=="inet"||$1=="inet6"{v=$2; if(v=="addr:") v=$3; sub(/^addr:/,"",v); split(v,a,"/"); if(a[1]!="") print a[1]}'
fi
}
# A published port can only bind an address this host actually owns. Docker does
# not find that out until `up`, where it fails with a bare "cannot assign
# requested address" -- by which point this script has seeded .env and burned the
# address into the cert SAN, neither of which a re-run rewrites. So check first.
check_bind_addr() {
local bind="$1" bare addrs
bare="${bind#[}"; bare="${bare%]}" # unwrap an [IPv6] publish literal
case "$bare" in ''|0.0.0.0|'::'|'*') return 0 ;; esac
if [[ "$SKIP_BIND_CHECK" == "1" ]]; then
warn "SKIP_BIND_CHECK=1 -- not checking whether ${bare} is local."
return 0
fi
# `|| true` is load-bearing: host_addrs ends in a pipeline, and under
# `set -o pipefail` a probe that fails AFTER printing usable addresses (or an
# absent awk) would make this plain assignment non-zero and kill the whole
# deploy at this line, silently -- before the fail-open below is ever reached.
addrs="$(host_addrs || true)"
# Empty means the probe found no tool to ask, not that the address is absent
# -- do not block a deploy on that.
if [[ -z "$addrs" ]]; then
warn "Could not list this host's addresses (no ip/ifconfig, or it failed); skipping the bind check."
return 0
fi
if printf '%s\n' "$addrs" | grep -qxF "$bare"; then
return 0
fi
warn "Addresses on this host: $(printf '%s\n' "$addrs" | tr '\n' ' ')"
die "Nothing here is assigned ${bare}, so Docker cannot publish 8200 on it. Fix OPENBAO_ADDR / OPENBAO_BIND (compose reads an exported OPENBAO_BIND first, then ${STACK_DIR}/.env), or set SKIP_BIND_CHECK=1 if the address only comes up later."
}
disable_swap() {
[[ "$DISABLE_SWAP" == "1" ]] || { warn "DISABLE_SWAP=0 -- mlock will be best-effort."; return; }
# Detect active swap via /proc/swaps (a header line + one line per device) so
# this works on musl/BusyBox too, where `swapon --show` does not exist.
if [[ -r /proc/swaps ]] && [[ "$(wc -l < /proc/swaps)" -gt 1 ]]; then
log "Disabling swap (mlock keeps key material off disk)..."
swapoff -a || warn "swapoff failed -- disable swap manually."
else
log "No active swap."
fi
# Persist: comment any swap lines in fstab so it stays off across reboots.
# [[:space:]] (not \s) so the match works under musl/BusyBox grep/sed.
if [[ -f /etc/fstab ]] && grep -qE '^[^#].*[[:space:]]swap[[:space:]]' /etc/fstab; then
sed -i.bak -E 's|^([^#].*[[:space:]]swap[[:space:]].*)$|# \1 # disabled for OpenBao mlock|' /etc/fstab
log "Commented swap entries in /etc/fstab (backup: /etc/fstab.bak)."
fi
}
# ----------------------------------------------------------------------------
# Extract embedded archive (docker-compose.yml, config.hcl, gen-tls.sh, .env.example)
# ----------------------------------------------------------------------------
SCRIPT_DIR=$(mktemp -d -t openbao-deploy.XXXXXX)
trap 'rm -rf "$SCRIPT_DIR"' EXIT
extract_archive() {
grep -a -A 9999999 '^__ARCHIVE_BELOW__$' "$0" \
| tail -n +2 \
| base64 -d \
| tar -xz -C "$SCRIPT_DIR" 2>/dev/null || true
}
extract_archive
# Fallback: run straight from the source dir (before build.sh embeds a payload).
if [[ ! -f "$SCRIPT_DIR/docker-compose.yml" ]]; then
SRC=$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)
for f in docker-compose.yml config.hcl gen-tls.sh update.sh baoctl .env.example; do
[[ -f "$SRC/$f" ]] || die "Missing $f (no embedded payload and not in $SRC -- run build.sh)."
cp "$SRC/$f" "$SCRIPT_DIR/$f"
done
fi
# ----------------------------------------------------------------------------
# Prompt for required values
# ----------------------------------------------------------------------------
prompt() {
local var="$1" msg="$2" cur="${!1}"
[[ -n "$cur" ]] && return
[[ "$SKIP_PROMPTS" == "1" ]] && die "$var required (set it in the environment; running with SKIP_PROMPTS=1)."
read -r -p "$msg: " "$var"
}
prompt OPENBAO_ADDR "LAN address the Kanrisha tape host reaches this vault at (IP or DNS)"
[[ -n "$OPENBAO_ADDR" ]] || die "OPENBAO_ADDR is required."
# Build the cert SAN list: loopback + whatever OPENBAO_ADDR is (IP vs DNS) +
# any extra SANs the operator exported in OPENBAO_TLS_SANS.
ADDR_KIND=dns
if [[ "$OPENBAO_ADDR" =~ ^[0-9.]+$ ]]; then
ADDR_SAN="IP:${OPENBAO_ADDR}"; ADDR_KIND=ipv4
elif [[ "$OPENBAO_ADDR" == *:* ]]; then
ADDR_SAN="IP:${OPENBAO_ADDR}"; ADDR_KIND=ipv6
else
ADDR_SAN="DNS:${OPENBAO_ADDR}"
fi
SANS="DNS:localhost,IP:127.0.0.1,${ADDR_SAN}"
[[ -n "${OPENBAO_TLS_SANS:-}" ]] && SANS="${SANS},${OPENBAO_TLS_SANS}"
# A Docker-published port bypasses the host INPUT firewall, so the interface bind
# is the real restriction. If the operator left OPENBAO_BIND at the all-interfaces
# default and OPENBAO_ADDR is an IP, narrow the publish to just that LAN IP. IPv6
# literals must be bracketed in the compose port mapping ([addr]:8200:8200).
if [[ "$OPENBAO_BIND" == "0.0.0.0" ]]; then
case "$ADDR_KIND" in
ipv4) OPENBAO_BIND="$OPENBAO_ADDR"; log "Binding the API to ${OPENBAO_BIND} only (set OPENBAO_BIND to override)." ;;
ipv6) OPENBAO_BIND="[${OPENBAO_ADDR}]"; log "Binding the API to ${OPENBAO_BIND} only (set OPENBAO_BIND to override)." ;;
*) warn "OPENBAO_BIND=0.0.0.0 and OPENBAO_ADDR is a DNS name -- API publishes on ALL interfaces. Set OPENBAO_BIND to a LAN IP to narrow it." ;;
esac
fi
# Which OPENBAO_BIND `docker compose` interpolates decides where the port lands,
# and compose reads the shell environment BEFORE $STACK_DIR/.env. So an exported
# value wins; one derived above does not, and the .env -- which deploy.sh never
# rewrites -- wins instead. Resolve which, say so when the two disagree, and
# confirm the address is really on this box before anything is written to disk.
EFFECTIVE_BIND="$OPENBAO_BIND"
if [[ -f "$STACK_DIR/.env" ]]; then
ENV_BIND=$(sed -n 's/^OPENBAO_BIND=//p' "$STACK_DIR/.env" | tail -n1)
ENV_ADDR=$(sed -n 's/^OPENBAO_ADDR=//p' "$STACK_DIR/.env" | tail -n1)
if [[ "$BIND_FROM_ENV" == "1" ]]; then
# Exported, so compose prefers it -- but a stale .env line still bites a
# later hand-run `docker compose up` that has no such environment.
if [[ -n "$ENV_BIND" && "$ENV_BIND" != "$OPENBAO_BIND" ]]; then
warn "OPENBAO_BIND=${OPENBAO_BIND} came from the environment, so compose prefers it: THIS run binds ${OPENBAO_BIND}."
warn "But ${STACK_DIR}/.env still says ${ENV_BIND} -- update that line, or a later plain 'docker compose up -d' will bind ${ENV_BIND}."
fi
elif [[ -n "$ENV_BIND" ]]; then
EFFECTIVE_BIND="$ENV_BIND"
if [[ "$ENV_BIND" != "$OPENBAO_BIND" ]]; then
warn "${STACK_DIR}/.env pins OPENBAO_BIND=${ENV_BIND}. Nothing was exported this run, so compose uses that, not the ${OPENBAO_BIND} derived here -- edit the .env to change the bind."
fi
else
# .env exists but has no OPENBAO_BIND line (hand-edited?): compose falls
# back to the compose-file default, which publishes on everything.
EFFECTIVE_BIND=0.0.0.0
warn "${STACK_DIR}/.env has no OPENBAO_BIND line and none was exported -- compose falls back to 0.0.0.0, publishing the API on ALL interfaces."
warn "Add 'OPENBAO_BIND=${OPENBAO_BIND}' to ${STACK_DIR}/.env to narrow it."
fi
if [[ -n "$ENV_ADDR" && "$ENV_ADDR" != "$OPENBAO_ADDR" ]]; then
warn "${STACK_DIR}/.env still says OPENBAO_ADDR=${ENV_ADDR}, and an existing cert in ${STACK_DIR}/tls is never regenerated over."
warn "To actually move the vault to ${OPENBAO_ADDR}: edit that .env, then 'rm -f ${STACK_DIR}/tls/tls.crt ${STACK_DIR}/tls/tls.key', then re-run."
fi
fi
check_bind_addr "$EFFECTIVE_BIND"
# config.hcl is rendered from OPENBAO_UI, so resolve which value actually
# applies: one passed to this run wins, else whatever .env already deploys,
# else the default. Without this, `OPENBAO_UI=0 bash deploy.sh` against an
# existing node would look like it worked and change nothing.
EFFECTIVE_UI="$OPENBAO_UI"
if [[ "$UI_FROM_ENV" != "1" && -f "$STACK_DIR/.env" ]]; then
_env_ui=$(sed -n 's/^OPENBAO_UI=//p' "$STACK_DIR/.env" | tail -n1)
[[ -n "$_env_ui" ]] && EFFECTIVE_UI="$_env_ui"
fi
[[ -n "$EFFECTIVE_UI" ]] || EFFECTIVE_UI=1
case "$EFFECTIVE_UI" in
1|true|yes|on) UI_HCL=true ;;
0|false|no|off) UI_HCL=false ;;
*) die "OPENBAO_UI must be 1 or 0 (got '${EFFECTIVE_UI}')." ;;
esac
# ----------------------------------------------------------------------------
# Lay down the stack
# ----------------------------------------------------------------------------
log "Setting up $STACK_DIR..."
install -d -m 0750 "$STACK_DIR"
install -m 0640 "$SCRIPT_DIR/docker-compose.yml" "$STACK_DIR/docker-compose.yml"
# config.hcl holds no secrets and is read by the in-container server process
# (which may be a non-root user) over a read-only mount -- keep it world-readable.
# Rendered rather than copied, so @UI@ reflects OPENBAO_UI. A bind-mounted
# file's CONTENTS are not part of the compose config hash, so `up -d` alone
# would leave a changed config.hcl unloaded -- track whether it actually
# changed and restart below only then, because restarting re-SEALS the vault.
CONFIG_CHANGED=0
_UI_HCL="$UI_HCL" awk '{ gsub(/@UI@/, ENVIRON["_UI_HCL"]); print }' \
"$SCRIPT_DIR/config.hcl" > "$SCRIPT_DIR/config.hcl.rendered"
# `cmp -s` exits 1 when they differ, so it stays inside an `if` condition:
# `cmp -s A B && CONFIG_CHANGED=1` would trip set -e whenever they matched.
if [[ -f "$STACK_DIR/config.hcl" ]] && ! cmp -s "$STACK_DIR/config.hcl" "$SCRIPT_DIR/config.hcl.rendered"; then
CONFIG_CHANGED=1
fi
install -m 0644 "$SCRIPT_DIR/config.hcl.rendered" "$STACK_DIR/config.hcl"
log "Web UI: ${UI_HCL} (OPENBAO_UI=${EFFECTIVE_UI})"
install -m 0750 "$SCRIPT_DIR/gen-tls.sh" "$STACK_DIR/gen-tls.sh"
# The updater. Deliberately NOT run by deploy.sh: upgrading a live vault seals
# it, so it is an operator-invoked step rather than part of a deploy.
install -m 0750 "$SCRIPT_DIR/update.sh" "$STACK_DIR/update.sh"
# baoctl: run bao commands without a token ever reaching an argv or history.
install -m 0750 "$SCRIPT_DIR/baoctl" "$STACK_DIR/baoctl"
ENV_FILE="$STACK_DIR/.env"
set_env() { # <KEY> <value>: update KEY in .env, or append if absent
# The value goes through the ENVIRONMENT, never interpolated into a sed
# script. Interpolating it corrupts any value containing & (sed expands it
# to the whole match) and aborts the run on one containing the s|||
# delimiter -- both reachable for an OIDC client secret or an issuer URL
# with a query string. Same helper as 947c899 gave the other stacks.
local key="$1" val="$2" tmp
if [[ ! -f "$ENV_FILE" ]]; then
printf '%s=%s\n' "$key" "$val" >> "$ENV_FILE"
return 0
fi
tmp="$(mktemp)"
_SE_KEY="$key" _SE_VAL="$val" awk '
BEGIN { k = ENVIRON["_SE_KEY"]; v = ENVIRON["_SE_VAL"]; seen = 0 }
!seen && index($0, k "=") == 1 { print k "=" v; seen = 1; next }
{ print }
END { if (!seen) print k "=" v }
' "$ENV_FILE" > "$tmp"
cat "$tmp" > "$ENV_FILE" # rewrite in place: keeps the 0600 mode/owner
rm -f "$tmp"
}
if [[ ! -f "$ENV_FILE" ]]; then
log "Seeding $ENV_FILE..."
install -m 0600 "$SCRIPT_DIR/.env.example" "$ENV_FILE"
set_env OPENBAO_ADDR "$OPENBAO_ADDR"
set_env OPENBAO_BIND "$OPENBAO_BIND"
set_env OPENBAO_TLS_SANS "$SANS"
set_env OPENBAO_UI "$EFFECTIVE_UI"
else
log ".env exists; leaving it alone."
# ...except a UI switch passed to THIS run: config.hcl is rendered from
# it above, so letting .env keep the old value means the next run
# silently reverts the UI. Only touch the key when it was actually passed.
if [[ "$UI_FROM_ENV" == "1" ]]; then
_cur_ui=$(sed -n 's/^OPENBAO_UI=//p' "$ENV_FILE" | tail -n1)
if [[ "$_cur_ui" != "$EFFECTIVE_UI" ]]; then
set_env OPENBAO_UI "$EFFECTIVE_UI"
log " OPENBAO_UI: ${_cur_ui:-<unset>} -> ${EFFECTIVE_UI}"
fi
fi
fi
# ----------------------------------------------------------------------------
# TLS: self-signed unless a cert is already present (drop in a CA-signed pair
# to use your Smallstep CA -- see the README).
# ----------------------------------------------------------------------------
install -d -m 0750 "$STACK_DIR/tls"
OPENBAO_TLS_SANS="$SANS" bash "$STACK_DIR/gen-tls.sh" "$STACK_DIR/tls"
disable_swap
open_bao_port
# ----------------------------------------------------------------------------
# Bring up the stack
# ----------------------------------------------------------------------------
if [[ "$FORCE" != "1" ]]; then
printf '\nAbout to start OpenBao from %s (TLS on https://%s:8200). Continue? [y/N] ' "$STACK_DIR" "$OPENBAO_ADDR"
read -r ans
[[ "${ans,,}" == "y" || "${ans,,}" == "yes" ]] || { warn "Aborted."; exit 0; }
fi
install_docker
cd "$STACK_DIR"
log "Pulling image..."
docker compose pull
# The OpenBao server process runs as the image's own user -- root in some image
# variants, a non-root service user in others -- and reads its config + TLS over
# read-only bind mounts and writes the raft data volume. So those must be owned by
# that UID. Detect it from the pulled image (-T: no TTY on the piped stdout) and
# align ownership; a no-op when the image runs as root.
# Ask for the account the ENTRYPOINT switches to, not the one a probe starts as.
# The image's entrypoint runs as root and then does `su-exec openbao "$@"` before
# exec'ing the server, so `--entrypoint id -u` bypasses that drop and reports 0.
# Chowning to 0 on the strength of that leaves the server running as the
# unprivileged account with a root-owned raft volume, and it crash-loops on
# "failed to open bolt file: /openbao/data/vault.db: permission denied".
BAO_UID=$(docker compose run --rm --no-deps -T --entrypoint sh openbao \
-c 'id -u openbao 2>/dev/null' 2>/dev/null | tr -dc '0-9')
if [[ -z "$BAO_UID" ]]; then
# No such account: this image runs the server as whatever the entrypoint
# started as, so the older probe is the right answer here.
BAO_UID=$(docker compose run --rm --no-deps -T --entrypoint id openbao -u 2>/dev/null | tr -dc '0-9')
fi
if [[ -z "$BAO_UID" ]]; then
# Both probes failed. Falling back to root is the DANGEROUS direction -- it is
# exactly what produced the crash loop -- so say how to recover.
warn "Could not detect the OpenBao service account; assuming root."
warn "If OpenBao crash-loops with a permission error on /openbao/data, run:"
warn " cd ${STACK_DIR} && docker compose down"
warn " docker run --rm -v openbao_openbao-data:/data -v ${STACK_DIR}/tls:/tls \\"
warn " --user 0:0 --entrypoint sh openbao/openbao:\${OPENBAO_TAG:-2.6.2} -c 'chown -R openbao:openbao /data /tls'"
BAO_UID=0
fi
log "OpenBao server runs as UID ${BAO_UID}; aligning file/volume ownership."
chown -R "${BAO_UID}:${BAO_UID}" "$STACK_DIR/config.hcl" "$STACK_DIR/tls"
# tls.key stays private to that UID; config.hcl + tls.crt are non-secret.
chmod 0600 "$STACK_DIR/tls/tls.key" 2>/dev/null || true
# Raft data volume. Gate on what the volume ACTUALLY is, not on whether this is
# the first run: a volume left root-owned by an earlier deploy (or by a botched
# UID detection) would otherwise never be repaired, since the first-run flag is
# false forever after. Reading the current owner costs one container start and
# still keeps the recursive chown off a healthy live raft dir.
_data_uid=$(docker compose run --rm --no-deps -T --user 0:0 --entrypoint stat openbao \
-c '%u' /openbao/data 2>/dev/null | tr -dc '0-9')
if [[ -z "$_data_uid" ]]; then
warn "Could not read the raft volume's ownership; skipping the data chown."
elif [[ "$_data_uid" != "$BAO_UID" ]]; then
log "Raft volume is owned by UID ${_data_uid}; chowning to ${BAO_UID}..."
docker compose run --rm --no-deps --user 0:0 --entrypoint chown openbao \
-R "${BAO_UID}:${BAO_UID}" /openbao/data 2>/dev/null || \
warn "Could not chown the raft data volume; OpenBao will fail to write storage. chown the openbao-data volume to UID ${BAO_UID}."
fi
log "Starting OpenBao..."
_bao_before="$(docker compose ps -q openbao 2>/dev/null || true)"
docker compose up -d --remove-orphans
_bao_after="$(docker compose ps -q openbao 2>/dev/null || true)"
if (( CONFIG_CHANGED )) && [[ -n "$_bao_before" && "$_bao_before" == "$_bao_after" ]]; then
warn "config.hcl changed and compose did not recreate the container; restarting to load it."
warn "NOTE: a restart re-SEALS the vault -- you will have to unseal again."
docker compose restart openbao || warn "Restart failed; run: cd ${STACK_DIR} && docker compose restart openbao"
fi
# OpenBao starts SEALED (and, first time, uninitialised), so it reports unhealthy
# until you init + unseal -- that is expected. Give it a moment to bind.
sleep 5
docker compose ps
# OpenBao starts SEALED and therefore reports UNHEALTHY until it is initialised
# and unsealed -- that is expected and the summary below explains it. A
# RESTARTING container is a different thing entirely: it crashed, compose is
# looping it, and every command in that summary will fail against it. Say so
# plainly rather than printing an unqualified DEPLOYED.
BAO_STATE="$(docker inspect -f '{{.State.Status}}' openbao 2>/dev/null || echo unknown)"
if [[ "$BAO_STATE" != "running" ]]; then
warn "Container state is '${BAO_STATE}', not 'running' -- OpenBao is crash-looping, not merely sealed."
warn "Nothing below will work until that is fixed. Start with:"
warn " cd ${STACK_DIR} && docker compose logs --tail=60 openbao"
fi
cat <<EOF
================================================================
DEPLOYED (OpenBao starts SEALED -- finish setup below)
Address: https://${OPENBAO_ADDR}:8200
Stack dir: ${STACK_DIR}
^ run every 'docker compose' command below from there:
cd ${STACK_DIR}
From anywhere else compose reports "no configuration file
provided: not found".
TLS: ${STACK_DIR}/tls/tls.crt (give this to the Kanrisha daemon as
[encryption.openbao].ca_cert)
>> Initialise NOW: an uninitialised vault reachable on the LAN can be init'd by
anyone who connects, capturing the root token + unseal keys. Do step 1 before
walking away.
1. Initialise + unseal (ONCE). Store the unseal keys + root token OUT OF BAND:
docker compose exec -e BAO_ADDR=https://127.0.0.1:8200 openbao \\
bao operator init -tls-skip-verify
docker compose exec -e BAO_ADDR=https://127.0.0.1:8200 openbao \\
bao operator unseal -tls-skip-verify # x3, three different keys
Pass NO key argument: it then prompts with echo off. A key on the command
line lands in the docker process's argv, and /proc/<pid>/cmdline is
world-readable -- three of those reconstruct the master key.
2. Bootstrap for Kanrisha. Run the Kanrisha repo's deploy/openbao/bootstrap.sh
from a host that HAS the 'bao' CLI (the Kanrisha host or your workstation --
this vault host only ships Docker), pointed at this vault. Copy tls.crt there
first as the BAO_CACERT:
BAO_ADDR=https://${OPENBAO_ADDR}:8200 BAO_CACERT=/path/to/openbao-ca.crt \\
bash bootstrap.sh # export BAO_TOKEN first; do not pass it inline
Then point the Kanrisha daemon at:
[encryption.openbao]
address = "https://${OPENBAO_ADDR}:8200"
ca_cert = "/etc/kanrisha/openbao-ca.crt" # = tls.crt above
3. Back it up (this vault is the sole recovery path for its secrets). Snapshot
save is token-gated; update.sh prompts for the token with echo off and hands
it to the container over stdin, so it never reaches a command line:
cd ${STACK_DIR} && bash update.sh snapshot
...then age-encrypt + copy it off-box. See the README for the full DR flow.
Do NOT use `-e BAO_TOKEN=<value>`: that puts the token in the docker
process's argv, and /proc/<pid>/cmdline is world-readable.
Manage (run these from ${STACK_DIR}):
bash baoctl login # one hidden prompt, then a session
bash baoctl <bao args...> # e.g. bash baoctl secrets enable pki
bash baoctl logout # revoke + remove the session
bash update.sh check # declared/running/latest + seal state
bash update.sh update # snapshot, then upgrade (comes back SEALED)
docker compose logs -f
docker compose pull && docker compose up -d # update
docker compose down # stop, keep the vault data
docker compose down -v # stop, WIPE the vault (DESTROYS keys)
Re-running this script is idempotent (it won't re-init or touch .env / tls).
================================================================
EOF
if [[ "$BAO_STATE" != "running" ]]; then
warn "Reminder: the container is '${BAO_STATE}'. Fix that before step 1 above."
fi
# IMPORTANT: nothing executable below this line. Everything after
# __ARCHIVE_BELOW__ is the embedded tar.gz payload (base64), added by build.sh.
exit 0
__ARCHIVE_BELOW__
H4sIAAAAAAAAA+xb63LbRpb2bz5Fh3JiMRFAUhc7oSNXaIm2WZIplSgllVK8FAg0SYxwGzQgieNo
ax5innCeZL9zunEhJSeT2kxmaytMyhKB7tOnz/073fJi91qmlhuHSaykvQyDJ7/7p4PP893dJ53n
ne6LvS7/7HS6/Byfve29nSdd/Lu7hzcvtp90ujt4/ER0fn9WHn5ylTmpEE/2Xkx+iIMb+Ues+X/o
syFOEhm9dmIBObjX4p9//4dYOKknI+mJzEmkJSM3XSaZH0fiWi4xLE6lmMWpOHKi1FcLx25sNDbE
eZzEQTxfik3lhNI67o+2xOhEpPJGpkqKJI3vlq0eBopyItMXi1hlwrIWWZYoy3olYvAzdeLe19ud
DgZvRk7m30hxfjwWYCFbSHHj5EEm/EzJYNbixS+iwL+W/DLGP6nejBI+TfAVuEhiegvO8S2KxYHj
eUvRFscye6bEQG+xJxzQUtJNZabMRjOZhj5YkIo50IsKJ/IEk3XcBeQUY4+8OO86ijP6AlJJPg18
F1yASiQzW5NQwhFExVL+nKTsyjQTm55Mgnhpq4WYQ/Ypr+hnLQFBE1cH/WJ44vipWMa58NI4oR3a
7SxQYlPac1vc+A69S8U4dIJAZTLBTM1f/+D9oGU3GhHU0yuk3GjcxEEeStVriOKZ5TmZ0xP82RCp
M8t4B3PiyWOxOHPJhlIqAzJMoGYfC0YZCT+TjYaS6Y3vrpCmX4XwQxAoWWgX755+PDkdjF73Tybn
/bc9a9t+bm/f8wQ3jjLHh1Qmq8zrd2EIdfQELYdtIpRFM3++X5Bt6+/mh71wA56WSvJ7qDyPAqmU
hW0lifQaetPnCwmbzW5j4cAGhqPB2bldOgrJHUNFGCB2CpUnSQwFYuN+KoPlllAxmcDCj+aQmyFI
kip3IWiiEqEM43TJtnTl+cqZBnKiae6LmRMoeUXarfhmyzH0yEUF7yBPhEzTOLWJ6SXzey2TDL4S
LMVUuk6u2DH0K8wLab+0qiGVgo0bYhaDQnEb54GHZ/ABKFGQ9y6Ftk0y1XILFADgR3N8sw2hI6wR
YlLqO8yrZmM2E9jctZguxeFw3H99PBy9FeMf+qdb4nbhuwtRGb4XS6WJuU4ygYdqcxHCEsPTg8nx
ycERP8gDP4QjFm8hR5JaT1hdfkLqKF9uiBHHD4tcr386tMUpOSVWM8EELgvqWlybhQG+Ho4OW6RI
P9MBIyvJyTtKlB5NzyMonJ2CPXzmwNhtBL3zASKJONSZNdHLkeOSmYDa4ah//swrCZL+p8vEUQoe
TxxxQByOTi/OxQwmdQtHZqNC1IBnu5JtN/VdDsk6qJHkqg2fnEMvUB6YJy+NHNjHrXBclxR/62cL
Ud+n+ErHTQ7x0FwISZW0NhFaJqn8aw5GJuB0AnvwZ8uJG/gw+AmFrhbZYfQ3Z9VYt3SMjPwM1uAr
WVKsBfAwlJ4PRknyDomzGu6ZMS6eTyXTeUZiAtVlHEmYTlxSpDEciSle2nV7kjMiola3m8Xl9/7h
4RlIyUjruaYSmIRdGl+zikxEoWd1bP7vnnMU/9PkwbVYqmfatbDT+3RE6qVxbQpEXo2lL7W3KyG6
HETfeIiMbvw0jkLopnKA06OD8Ua3K5w8i608UhJKPh2O2DbY7CE8T4uBlMPvm8m1q7rdZqnbSjgy
okjlrarbFu+KyLKZRwg5Xktn4lsos5AkCfDd+P0Ei/fqQi2eWffNcpV3fWTnENrzrSj2JDMKr/Jg
f5nPEY390pPPTBomptiZETfgIMouSbGmT4esbazLhUav3a7Wpxf3NTUWkw6OL8bng7N/YWK3aWJg
B7FbixgS2lzgZ7ZYtl6KbTw3T1Hi5JF5o92aNxhJOAOiMDsCDTUUsc0bX/nYnS0sWIOlrv3E0m5Y
j+8i4OyL0Bw6S0EPH1YZWiZ6bRQuiJlmvyg2kAovmwfvD5tbogmboh+Ux3NFv1lGqPuFDLrbL9gH
ulpsNGSNt+YHQ5tj440T9MROp7CizA9lnGPJveIJSq7Uh+uIHfOAk9sEVYUfI7l3MfU/XSr/v/xU
LvzvW+NX8N/zne3dCv/tYVyXAOCf+O+P+FT4r478UG8YEPebMODoAVQrg9ImisUVLEj5Y4ZclbW2
6vCC4AaBMIMyNt1AIiKqyEnUIs6sqUO1weFZi8pdCo26ktCF0/jkeNCggtaNuWxNHJQ6xKVh3yBa
xVsNYmWK3qqwRxoVfSoR8NAE4Y2qpMUMzNUV8SpFW7wm7IwiAtX4JiOmgmPV4kLIgElGhSYFQ5Cq
ncYEFuNr5F7EQ1TLYkrDN5XUQf1s0D98P0CJhkIlzaNpHF8ThivE06S1muIjYiZvFp990VypCygz
UYKZ+B69KwoIZKz7BsH247FVYFySR6EwWyCe226aUQViXxv4UE8o02Wt1IK6UasqUh3Y3GLU+q9i
Vl+pXNN7gFxBrsSurLVVubAVyKqf4DnAMjC8FE6tdIkAQAwil5UB20YCH2zXmeikWBppM3MTLU2T
73QqIrmZkq+sEqguptmTmY+6oy50vGkb2RUDIT497tGBeEsDyRNrJThv11Te0L9J+I7QlbduGsB/
QhZH2UjpMSG96i+X7WAFcpK18cVLx/C69llnXQ+3XEfvlK3pSMqErOi6DgVLCAh1eFS8P4CC9e4H
QRlWXWurobEMhjPAqxA4Q9Z1BN6j5oyIb4uyFEaHYJMCHqRUWRjnQ3mu2HMKWgtHlYi+wPIUM8yS
p4g/SpoVTUiLiNBPzRXM/lOzQO1aK2zKhhHyaDNYVNFIqFuHfAG253h2EyTHCEExbxnmB1WvNgVo
XQoQ0ucOFyqqXArL+lTvoLGhO02PdQoKeC0uC1z9gZk0wLpA1NS+2ABL6x1iDgXUBchEIGcZeSjK
bdoAdB9IXbBTZKoKdjKKsmYkxma0m6J8dRBb8ZXnEdPBrbNUoNc/OB9+P2DWIoq7UIMHi3bBorY+
hBBubUDQER5jn6njzxcUUktwb1wZYXNDvIH4K2SRBHpNhXQnsxWcAI9bRwBEk1GmRh2caB7BHbrn
SN0elhdBslLJDSfxSfBp4VCfKKcbboAIJlM99lPjumQzFP+muR9kFta9lVNxMbTFOJ+qzM/ybD1M
s2kW6IVGDjPiW1GnNpK3DdPeyFPZM60P7qp5Op3TV5QFZY7QGB+IXDe7KnFgNG+kocG5ExDCXJJx
c/bUw7NFGufzBQuLGzPneRpx04yaV7T27cIhC8uoW0WZIKG2KDRYtQnaJpcSH0E8NzVFCIOapvGt
IhZVXEvnnHTJfKnEWFCsMNoywwtNkZ85RUdRLWQQAGYqN42DYIpMbzdyfyUyfncx/I4c4J//+Dv+
F/0a0t4sGistjuiHgzf9i+NzMRyL9/3RRf9YXIzGA/zQU//3/4OLw8Hp2eCgfz447PHmXl8Mj88t
IH4N6jXCn8ogvtVNbB1MIYkiJt5s27AzVm9Dd2wA0SkcAaNNcy4DKZb7ivqBfmSyN8ZF9qPdhmIY
OQ11qImivKOmOF4evR+LJMjnWH/zyvzSvA5Vs2pCfLy/ahUtMD6lyGAZoEWYmbrkcGeFfyhCT+WM
Ky3nmkyJ9nFHjv8Ddb1QgjI7pqdRO/hIYWOkhcGhCTZE9wqZrjI6vZUrsXm30wJBRDeYhi4GTS8Z
FhyvbNuh7igYc0qhQIxtrva2QI8615TKtZSRcwPtYz7ehnEe8SlCOTWMvTygSsCT1FWnmjlebStr
d0QkA7W11gqXXIhEPS7TxVqP5yM/E/DraS3X5ypt40lbxbNsoUL63fy6bau4aeaoAAVsWSQVT1EE
VJRo3Xa7xk0xisqiwIEh0qhrU8QZ4RVjQomYEvkqZPp33c7XL5r1qmRDHBy9n/QH48nbg/c85/7P
RsFv+MxlRJ0bJId/3xoE8l/s7X0K/+90Xujz387znZ0XO4z/6Uj4T/z/B3w2PmNHn/pRmyoV4OsF
x4jKLjhtFWeSa01FwikMJr7ikp/q5jqgLgsFKkZngmo8fZpSlgQJYidioC1GXN9RKXmb+nT2Ke8w
mWI4gRGl83is+5orx6H69PTXgGTtCFT8hZIFA1ED08RXlOUMzC3jKsxijqoQBWdZFVB85RCNwjUG
kuDerRUnuvtxofhYk2VYl98lpDQ5HJ59KGLWpjmY0CiYsskAIZLelAegmDHuj8bCnG8KfGFpbgne
afNwNO5x78MOnGhreNrraoDa7fCXokwsoug67cP+j0QbGML3/Ix7MR5K7oqzr7f3yhPe0Lkr6iPF
p0hJ1qLT3UxYMo8R6xNM8oNGA3vcbz792O1ZvLH7ZoP2sF/r9Bf76lm0AWAMJyDUusIzphF769Po
Wc8CX3jfCK9JLVYimk+xKB7Aui4vhTXTD0r8Lb74YvUhYW3x4cNLLlaAfZFdYtG8NNr6wAZdAtc1
IyUpESHdPpLOjSlWHQIvNkF42GwmOo2Z32iYY2lh3fCZjVKBeNVG6m5HOfL89qsvuuLnn8XHBwwU
g8lJZigCvKZ49cX2S026+1IwzB4cHI774tTa3nv+sMfPkIDshYsiPwKW8IqD3AKjrPQMQHCTWD04
HrbYwbLYnKeVZ3Xc0ME3CJ0kxN0E33W4EiTdA/4VjKfyr8K62+t8IywgCvIo6UJT+AWFMH6fAJY7
IfY7cXOU170k9UOJjdx0BQNHJX6CJC0an2frmrNWn7GKLbZcPISBNPVklU//ggrmYLRvGhYYha2j
5BRNegeo2A+yEbDM/tOPZI9kU+4CJZZ43umsLVq82N1dW7nRWFfeWxMlvRWtFNGQI5WxoeaDuYJ0
pnrCMFTzTnpGm7v3Hs56S/3WOlcF8l3viTnqsSV/qS+GOKUPwOvttCKit2Dv/+nU9efnd/jkiQeD
/beWf79S/+11drvF/b9uZ7fznOq/7t7zP+u/P+LzqfqvtAtqMeovK6VdiTptcRCHCSAatYjiqte0
xVUTKjBJ6JZy5Fz5nibCIN5cG3w3HIuL08P++eCMmiKHg+Ph68EZvh7/KC5Gx8OjAcYMxAn+ORsL
QFmecTY4PbHBRLJEOsmWBkPLdE6chVJQl8b0JELE4huq1RixO+XdJru2F6rleIpuAvT4loxpfRQl
0XjhhMh/BE63mAyDftMryBapXG0yZctEN1Xo0opY5CH3sIr2RHkJhntMqmqicc/AFn2h6GZhThcG
jPQdlBY7PaQnvqKlmUvLbkedZjybcbM4jzI/QOEcSqq8szyNFB1Q8Q0gGk18ai2MddKo1uTTpNHg
eyiF8PfcXE1K84hbdNTxJYVzc5HFw61EumaZZkTTqHRyenI8PPhxn7oidMNINzeo2EA9wp03vnqD
beR0tqbv4q30UHzqGek+ep7S7Qmr7Brz3RUmYc4Miv1XFkCNKn0ooPdJ7c6AkAQERtX7ylFdqZw4
d7UujFKKm2qRab3HM3OYoO+aav2gNjamQF380riekcjnqQPLn+e+7m/LuyTwXcqs1ODkU1FCOsbe
+HakLir1xuhAgDqPektsknRQyUeKpr/p1HbhF0cdRbdSJIQRNOwiiAZYhKpC14krs6gxZc25fiEV
uU4UsVRYXGS9prdqLrKYynAKKZBItIx0l9hNfRR7fMWETY4PeiFtlhfYZWKoG0nKCyeY3TrLoh9s
DDKfmgpa6TvDTGsNP7VgfG7gkFncqMo2UaXQ7dlsq+z2ZVIb3e1CEgc1UKQ/fOfGyVDZeJW/6ZuQ
2L1yZiiU4WsIdcYXzP1OWzfTChnSORp54rpdcXcaeIFv5CEijEf90/G7k3MChUC+5q6oDOimFBE0
DPDH4xIsim97BEPg13S6Yb2qiON3jxrZ+EldN/pBCINuGSVMDXIpN8od9bqng50bqF29FPGUotaK
2/J0E8V1s6+YpY3Qh/9dgfyV5jqqhtaOy4o5rNbTGIa/FJsry5hb4ZApnXDW1asPfXQAeiz+vDQR
h0MBE+FAIyoZ8q02h842+PRIH+Zxa0FxKAGcOBkh0wA/Pog6tU8VgGxuJRMM8WFgvPsi7OlzvnjF
PDS2B76iqMP37NiH2jJzi8NUy6RaWuNl/QYfAnWkjHDG5/2DI7KX/bZKb4qpmrnVUKvFyJMYOp8c
DUb74sHHHA5zt+RKc3DVviqs6soWp7A2yIyCTxabrvIvf0zzmlMPZEtH9c1xniTBsrhowRGmqY8e
7FUWJ2+Gx4NVPllc5bQi+PDZ9CZKyE6r6MtoCdWcar9946CmQezLE1UIi0ed98/eDs4nyGvj4cma
XMh9ENwS6RK4LXNc7UhIxxUmdHj24+TsYrTfeUQQgLSkPfI1HUQWTpJQz99YsjEPzfbR8HRS8L5G
7bA/ejs4O7kY98occmvOMVZiOxN6+25CRVGBeKtdczR9oAVtVleViXxLV1Fe6eZVVfwZy8A+0mu1
xf5yeMJGwYGpZ7JukEumF1DAphDHDTg+OBOwCmSq5dphobkf5KTzG33YhpQLSSVpzLeUjQOSCfCJ
IqxUdzByOqUzk6fxHR8ntWlW+9vE91613dAjAzTXg0wi4jjCBJM8UzWbImAcLYu/IRCF6SLiLvRO
Sjxd3XyPC1oq8yjgWvqmM7krdgtWwkTfIOBHUBfDbroEgSDDd2D47wP4sgECKPWOmNwbPmSqspBC
0kpisqRVJ6FC0BHkAOwL9d4j01nTYD1xrn02yJK0IKh0kd5j8+vZ7cF8vVtVNn6Z2GNU6intN1BZ
Cw98XaptU9Dki1OPLvKYn/Pfknxi+GOdzPHg+M1+8+mm6wn86/kp/d0JXVl+3R+/m4xPLs4OBped
D/fNFrcYk1uv1X66SZfTPj2w0Vi59Gyi9sHJ6E1v/1P54L5ZdDZTzH1k6kozk2Q5WMsf1d8mEb2e
Lpppxxye1rrxeK6vCKOWg9MN34z3dRTG8hN2jp9/Zm4icMMP9PpecUFd0Nvy1f5/i/+6vOypxHFl
78OHLzdouvU3kgKPaLdrb9v3RIukWZwWlzQn1/vFlM8/3//yvvlSTG7KRxtf7vMTPeh6nWidM833
NS8EVh4sVAz5+Nnk+qu7+2Icx/MZtXL17Gefq2f0640mzneavv2EfrgZzIovU3hvJYeDQ369ksV7
Jo0XL0s/6O0/eMSuUT2v58Heo4mwGLnqIxUFk9p6+52SZj1J1Z6bnNNbTzrUTh2MvtdO23xa7rxN
Z9DNBglm+Hby7uB45WV1FRnTg3i+2RLiYyH8Zz91dnYuuy93tsPLrz7wl04oPlc/RaSKL5vUGb91
0giTHs7ZCS8/e2SO7qzfNzxfPr5WN7y8++S8lY68viRVXvpRbkzo3+QOfWcBUYTRWv1vQrbMRUjK
ICQb6ipQE8pueC7vZFNwCCqFxOFmbbHm0++aosWMwIEnIMNTEQy+PRr8+Ep8y/n5Va8IwnhY3AfQ
FzWpOPG4BJ7SOYeJJN/zBbO5htbVVR3odXh2Mno/GJ1vmQKdb/gncfmnUHxPg5rgOvtalqFIFeE3
uy/cr7/5RrRF9+u92W5nl6P+7WKpAw8fCVFjApbRbVJlgV+2myILE/2ngxwNP9OHOoWJrQXBsgSb
kafuF3r7n/aufquNK8n/r6foCHwk2agFTjyTActZYvCEHRv7GBxv1jjQSA10jD6ibowZW+fMQ8wT
zpNs/arqfrVaQBI82ezSZyYJUuv27Xvr1nf9igMI9UUakjbwkf9z+6tJCidJtMwf0MHFv+jJkASD
d0U6GLfk3v2dzX1axa4Oij+/Xwcx8+DJ+buoYcf8dvOvW9u0HcjM07V7U9cB6jTt9+XPaSR8Tos1
pK+Wo6kd6gv+jAggIw3hQ3NxeYlGrXfrrajbjVYM/cpn0Xs7BNHoEKEPN5C5032yub1Bn9LqNvkh
rXAovbERrDmWkBanrgV7hf4ZzS7tZCD7xXcTlRLh7R+nRUilZLBKSiPr+XwYmo2Go8pWzXDo8tZL
IM3seGPNbuOazhu0CL6ed35cXOl2OuN6MMInstozspiHK5LE2r65i+uioV9BnTxlnV9z9ozahfOD
A5hMctZGA100jg70rKcfEEVLPZNOTzVsbzEyxmyTmoQ2/Z3q042cNW3RwSvUZdG3Sck/7betzg0b
jr0PMj7Zzu84B3JGD+cQLisSHFoXrscZdp7+H0dbakhZk4Fz/CYZNN4ZNRv5dqxfS8w9FS+g/Q4e
Dxke083lxXUM34RmJxsNj/WTjK4DVsrjG97nWtKTZGveV1C2JVcoFHbbjI4TcBpV8YIb5zE2owvO
3ohjQJIsqpeU5rKqMDVan9no2DFAR17E8k5Y8xuuzD6tNatVhW/4S6ZCbLW4iC9nwrpCBbHDmRVR
BTXndADj0pfj1cyGOBUnWb9PPI1UFveCM5ICQoLk+jVfbNs8IoXsS/tXzD/8TfKe+A0fMSkNAZkb
TdwADpzG0U5aNv26HfjlO8Wok3RgAHZgALIMh5OPxXDSQwkOHOMj+C7FtILbav39KOt7GxzHsTm/
PsvhvF8W+nyeZu14reul76rMbq0DeEnT4fTJh8R2cnrUI9Sy2/iPM6bZQLaTggFCRm+S54GvwlSB
yEoJTzL54ABggP4kngQy5s+5INrjTHGNJrKPB5UOp9MRREHwdpp2uac8d9fU3EaIhfU8yW4to6Ly
/OwXjiaQVT0pKqiPnwEFTgUsPQOP4k8+gzDaKcRAvslRa8YJv68eM5brKuMjD9KBNVT14Pr3qnbI
9CIlpweyN3lU/16dcFHEJnw9jl7DE2VCEDYvlyjcDENHFIUtAcHlJiGexlrVh7ACn0f3xUcsgyy5
FC+teF5ALjnqHUCUxyJUSNcRIdiDtzvPCq1wrqAZ907RFTW0s8W9972EoT1LMZ9Ev+z8qEvTsdrc
ojFKpo1IF0FoyC9KXvJKkZecOw7BB/q8x5KeAUzoyMc13LfPARS7UaKhT3rd5c/xzu49K7Kl6KGL
35SOLg7upMeHxaGEKJGV5pwJ5kLWh3CjeY9pfj/bWfurraunYs5IgqxvRICy+BXzNcYTfYtoboxC
FFJSGx8/xnze+J9n+XTa0EFKj2L/4Nnw3ZCUmlYdurwJc7x9Kxu4PnTp4kT1B5K+TWy1Hn08MOns
gzQZ+lFQ2fFcIWtMDJMVPcZ7gCOIlqyfcRbXGZkUE72nMhyGhP/0KB1KKXoN0ZJ9ecS+i44Eig/0
dGfqV67d8SQd0yZsRo3AUYRRvb/v1RvhUFgV8cVX8BESsz/9LLo0CSwyvNdPx1BxIQvWpChM89+O
ueaSRVBuGEjvbEKmwFG+85SY5CD50EapenR/2dXgJOMsPqbfnB3GtCUdBH/yjvWFTDsaZco7MsP6
3HOsZkkj78R36zQbBrep++uw6v9Rf7/3cXlpZW+613zzY/3t3b1WPb7b2VvpjBvBsFZtC3jATXL8
BQ4OSSjynnUQ37RYKdU2+bUrRIsvN589/35zA+YOF6ugwLigMwAIHiShTNIe3BGtJc+AuKqIJY4e
T0ZG90DpEcjGZkvoMRMFCxnSGSdu6GnjH2jE3AQRHVaUC6VbcIjDCzYKtZwvrrGffl/eEqWYR0fW
PpZM5LaS+iOPpck34isZJD+NJvukO44mfIf3Nzyjcuud+O7UWOzsOvLuqhu4oojF7Z8/3Y+/pv//
5dOX8d1WtLZmv6O/jJJrPk7zpOdr6hVnn6ye6x53LUwpnfqS3s86tYeOhEJO6IG8P45ueCv2dMg9
C2iiVGGyvIkSzBJNETtKCod8MktrZBk7PivWsSQDmfAc5zOMdLPj6FnG1e3GyLWlTw0td9pDvRPN
zU2zIeXUalaQFUGHGLkF++a0zSMO+FEgUDQhxkZvYDjMIRya+mBMZyYXPZb/ZBGJKbfPont3frgz
uNPfvfPdnWd3dv5bpSLu75b8zVMbw6Dz17YLSv/Fg05jTKde8wilvhj4lln61Vcq7F74dqN6GC1d
kdw0REZxZnnxATnnx0bj6KXJr6nOrmFLeE52jWfVWffSXDuP/x24AcIXVbf6/FeU9Fy6bTXI7lj8
iKWe1gPrNFSAcMN8A1RmobkZbZINg2j5z5zm7O9dXanjOKrvSvFcOY2FqFfnwpqHmc5C9BoVG5X2
3ZJ4X4jlpolUqr9cfy0hfBK9cXRgfNi98YHTuBFePpNyaRhzbJKmXJicILAXnU/gMZsIz+Woq7gz
ZM94kNDkc1V84RvlyMlxsvMXa6udYjDugNQ1PM/0Xa3DGsfBjnk2Ao2Sz8Yiw1AxcoxSJcRCXYdk
7lzkHcV86OAVOuYVlEIrtHB4ZKvn98hQjJ3UY6v2eztlVwlK0sw05z5XnL1znjyr0ysGARK8kGW/
LPn3QtBKDutBlloSHf+ddkBoAPAbkxj6MXuRuEqBlGcW/UJm7A2wqrGOGHAIxwY4L0gy6pAhyrUd
OW1UAYPH5a4ReeeDtMh6sVV1c7ukno/IbjUNpQfH+rvmLSgxC69yBG9aXrQS0+Bb2oV9/v3LaC6c
iPibuMpAhkkmvRPIzjlzU18Wswg75PO/Rc3F5nkPrpGHZhYt0m5I920hN3h8gXV8/uSJmNXskXHM
o4qPfQ6/h0snvHHvx2ifVbfA9CnF7ELHoYQDsRf2FgC+FC7JqxUbbgxZbZMcjVkohoUkNnLamgpk
uY1Fd8klozJbfogbQuup5WTyjBFd9rbKpxij5MrxnMI8M9zi/Af6LfwMs2PN/BIZjQHRVZub0Cnx
TfeCTjCtMP/3cFQLSMutTNTksJLnkmqt2rDuR3PXavshHlQ8UkXZjWPWP2o6F0dLgVLdOHoXDcNS
Sf6aHUw3kQa7w4OY36sdiRDVR7mHRjodjd4hjZZlxqNQg+clC3T35Vb4KJcGaybrXQI/sLmBp3s6
/v1fMoaWzjclAfpLTixvlQbElrauO6BzEHtrWBrw7i+ZobpX/J2DO+mOTNMuYtmacaOX4AyUAFfD
vcdN5X0ec8prxZT8HwbJH4GCzOkyxpMwRyt24tuV7f01K747OzQprzkQNtaMGZqTdaKp0OzrEQeP
ZMVeKxADTi2HRbTZ+RNkYfFqzEnxcIVdoaXq7VLvYKIkq5F/OL+ZilIsj5zWg2nxBmBKZIPNWax1
W7nBmqPkZTtrTkk5LJ3gIE9FkYf3OjI44h4ImELpRapfj9ULVlRNYTBZTJtPn0xdRlyzInUwEd2v
5Zt/JGjkN76kWdx8tbURtdOfOTJmpQzmQaoLkuRi58S8AankiSUBnxqxcibTpA8uET7EmmeTjtoz
omgaOl2xn95EHan3U4mUId1ejXiTsdtU4g+4ZksCaqXEwA/xRfz3uGSS4kWUsOXxVUS9boqDs+FQ
Un3I3h1N1yz2s9SBXEHws14feah+WxbGdh12GbFXNX3LLE1gV26OEe6ZsBrLzoRV5BJ5eWjTilwi
Wq5239e8vVIPwz6ITvtnPYYCUC3QoOxMvKoJL0lZRxNz/HFSMFKvaOa+Wi0FWmF8kNGgDHrQ+dBP
EmKWvXaJrnGZfAwlnV1SCycobCD2bY9SOUx1Mcyl66Ar8EoYCO+JGshkKNGmxfWSfCuRvE2Id6Ku
LNDY01Kk0xb8HkdnuRKiSR8/PM2G/TiUcyXGi19oFB2MF2/pmG4x0uOJVcGiBz4pId1W/df7PUrU
iC1f42xVT2fr6jkr3Yuql5kPlZwvPYIL0WPBmDw5K7iSRsvAAM82DtxGS+CpDCYVVHQl+bs8dqu4
AyR58zvhqk3PWwhOf5gasjm3FXBcnteCZ8VY1piMTFHT+vxV8NkErfQX+IEsQYXMewFrcHQkG6j8
cI3J0swTtQ/iABvZTTdqsMfDDFcWDilO3erJ8R1O+evJVgQWP38BCjcC/3A04VQFnMTxJH2fjYio
bRlGLj5WWyS55vzszkAN1kMe2W5LNVJ7NBkTm8mrHYtm6D/IEm0dlVyVzCnhVWI2pL4wmifO5mrb
9+dMo7GyWQHRbOQeuGhpJZWgEzxUliZ2+X4PH24+f1KrdX/jVRPldyOq4ja1WiVbNjVzXJaiyhr0
LUsRq5h/SdzVojJ78LPqAPE3z/FnJO3ennEulnHBZn2Dei1EH75cUs2xT+KdZAPyKkl55FcbKrDn
6o1OzkTf58wpcpMz+RW82EukITFXYgTJWu2lyq0lnCg5GIMsP0xPSGXOV13h6jXLVmueY52kk4XE
jTw/oef5bpfAXkM6rbm6y1VL5mRxD1U8P5r+dsIEeYvi7UddbtrJ8xuVPN9TdKVKxAJZFaJuN7QZ
Kg5a1GQCN4mCrVVP52PbRpuMBBrR0MU4BaEo9GUhCK+xo5L2MEiGZ8lpKelixqu6PXJTOIc8xqPt
M3zXpJdaaryeRFS9ao+kbjStqpdK8HpeuazIqPOTEX0kNVpcEw5kqEky5DQLiU64vhKcv4pqf8bI
Xq1OsuBKOFmH04t2uI5atJ6Z/Pywyl8hBiKA3b43eFmmxCc1NUCmWN/Xj0Pfg68nSw1Ky/7N0kjd
ncGH5WgwLpji4U9ZLFc78koy+dKH4RKhXYExgALFEswjLYB7kMUDLcVuLyz8o0WJCCK6h6MArIy9
46Ise6K/vBzBF+pAMVdgQ4c3Xm5Gm+uXmdPeHhinEkxbOlRXm7tX7sb13ve4csO4KDSovI58sjCr
XKUblFbe+keuJM27JZJmBcufmkBnEV9ryBFoILDUwDQbUfOYNr9RqteaNlolq4rYCQPZMZFbllKO
wj54ENYXVlWQtUxvEZcmy7lYp9khg00RB4aoNdxOIqaMqHzcPs2A7OqBzOd0E8oQZMRMCg6GRxIr
4yqRwj8fgOIFrhenfiVaOusxPL3P8JOCY4uVZYqiLC5EFUWO4Mps2h5euOrMtZlUfr+akQGkkYbm
6uIDTS8ktPJm1YJa8VIKQ80UUruUrlpp7ZFoPRRQX1p8KU+Zv+S8yP761hauWmFZ1GhPC/P3Djp7
tjR/78B3FGr18GztsEW09gqHBfKXy7Ns+TCJv1PpO2NzrSUhbl7E2UQHGY7BviVNYn6tru3eIE8A
ITC9+BHfytpJkcl6aH6V29OdwcAnUZJ57DBm9IlqX9+OiH9ZJsG9ePzd5uO/RU3x7mu15qqtAr8K
uMLofkE4bN6jkNtvHlQm5VYYm7XFwlJQnLcN5EWQfxU1ftza6CacFtko31p6/wqOxT+QxjxZr8Nz
tLcLD6i4o1TdbPjBF4w8lZ/U2NphBzkp1ZtPiWPQa1tSceRy78M1RnezD8Lpk555+uUxdZYb9l5g
8feICfUtEtRlaQze77X/of5Y1ImrfupJTSaHLYOhdY23rqIpbz/oRBfpoK//Lv/WTFb25c2rYVa8
rW2kUsJJFNudAf4yuEe1dahMXbFezTi12psd+a+3td2LcdolRsVJopu0z+xQ6M7gjkW88bP7ft1X
QKbuZP4LbPBhMq+hmysKTe3NLn77tvZ8+JgU7mE/ofdhun5hu1p2eZ9eEimNBtnf0/5Geppc7KS9
7soJ/V736W3tdYIU7W8vujybPJbQRPBCMvtecarYjDh5o6Rf8b20WiOOOyS+Xfm2V9FTiYgKZ1GG
4+jKcmOuSVwOOw1/CwvWhKDrHdrwwLqVuPzAeiuqbTZ+w5J5870OvV3/bF1FAJfOLEi+MdaoQvfH
USWsxJS7c0D/GJ8mPa3NUqMPIL1M+1Nr8fGfHFO3mr7RaI0Ab/FXVhPQb+UtW/pDPVj6HR3mlr61
WNfmC6WolnxhpIwZcWi/9snPfG0iIwq4HABvLNeJRnnyn8w8P8mUPtHTP+lIn+yYU1bbWWX/vaEY
f5eLyJQI8fM+43L8z/tfrtj+3/gn8D+XH3x12//t33LNw/8UumDTTGs5Dwxnzm1jiyBW63VtPVeI
NJSMC9BVj6x52EzJaX4NwNBL4ELZXTwiFj+Ass1gGfmqc+RxtuxwpBN+/HRrSXyU0oyXvT9cugdI
0jPIkCKTZFEjdTjD9KDCFX8gPr/RoYSl0KEI00feAQDpra0n4d9EUkvecSgCgER3o4PK8n1BpLpO
3b7ir1VX71fU7lfjZWmzXABWxWZiM7PixJe2LskBNODMmrvefMP6fh1tnNngp6zJDIBWDriRPNrZ
3djalvIcOLgld0xKfMmy1bc9ELtH7Lhoe/3ZZtQ+WIoO5O8X67vf0d9RM4n+c+f5dgQvnfTbOwCE
iHyTc1Mq066Em7sL/AAyQgT5pDArpPXZDFQWPd/aeGzarenvhLq4xNj0TIGIBR2I75GrkNmfkQ0L
zpMwi8HhfqYJC7mqZ4xJbmdzBx48k9tu4DrH0nmLSVMmNz4bjGOmb/yW9IGM8RcNhtZo2EtD2LEl
CbnkxsKfk6ivoKx8NOBDnPhHQsBs6dvDhLQe9RDIVk3S1B3NnGhWzjU/SIJAUm2YcoAOA8URWwwy
9yWzCgMu+E2kMdOkUD+KoMK5bUFQXsmPDzOjWOVZYdtumRmkH8bZhBsOOQxYoQx6LPGmAz0JcOcb
pDawLAv+ypXvgYOj9I6nUv9VRKy15ScDdXbo15Xl7k3nJAHtYt9G/N6HjBSLOWrSf3qqUdoJ6NQ8
ws+HMcShY4MHDcZHufVw+wgj/FZpLm34pB1dmrvSNjdBRbMbjhhvQ2szxI2BYIvAxLELaTJR50ZS
mNrSARxQk9wC9CrSsoG2FdeuPkhS8iV9Jupd9BQMYR6OsgNRnoMPy8QUvWm3C6Klr0/e4j+Vd4Br
AOxe3WZNhYPQ1BvQHaoYORqpMIRmRBAEjfMuTcfiu7ItPNRUEPIpZigfTUQVAzUrFCyQQ7KXXbNo
tSb6pGNLNiEP9zCo3H5UNRx3lO8HLfFEKtoGdZZyKgAEff4y86lw3/G7TJx8pjlMhxEohpwhOyC7
hgtau/XND8kArfZe4jA9Xq9HcRzPDBnw+cGF2bt29ND9hTK/mR/KwlZNnHbPALBeCaKq7Hd/d/dp
9+uT4CN1IlfiBF4H2M0bepXGLn9sAN+AmbYAjJwZrhEe//CQG9KQY34CRoczHtcck+nWKxlUvYak
gv2t7e7cLoSP1x9vvtzlW6q7m/7RUNq+dejrEp9nQ5EXs90Gy9WEEgsjImJXeOdwpMpZ7BvRzn5u
n3ySQT7hHy7EpHANhnqA1yAoDdH2y+gR43blZ4fNzo8L0Tedpaheb60ZeK4GG7QuUY7by1h7tXZT
Ob2/DmsuCIoMZ8I1hzaaofC9DFUjrT1jc0qFbJlT4QgzWpNAzSk4c3ZkfFaSbQMFDLoEXqJA5092
t9PvzwapdqwUf8cFek6xbomo9+HZ8bFRxM5phseeQRLXPPwlD2mxPS3FQiXoHESCMI0A7da+WXY8
ZKUrE/hvhrc3aoiewLW6NywnkinebMJQPQhanPHhTg4Zp4BlrJt3xH2o8FZnOQMQIWmLM3oQyECb
HhJ4tG6IIHhZtlwiwH21aDnLGUKrXq9QbkGIXRGIQigdbbalZjsP8iuyYeA3crKCJp0UpNygvQ2O
mlF6MAfuscpcJtJ6u17BuQ9SloZQdhw9x0OPEptJRHRNM2YrA+mTGk7jljUkic5oQphkbqJUKAlF
OajiIvKal6B/2mldNpbzrBY/KmckAnBfeQx18aP7IyhA4LJ3B2jp88rLIArDOch60EMs+51WOPbn
/ZorpL/ffLn15Af2WPuuXKyCJrgg+qt/Rg8hBA1Ok8n8PkWCAvd72939QWxozlUFpTMNG9QqRCBR
gOmhT1nL4TwBBcmIdN+aNVtUgUCTCUh42G3xT/loeCCRd1gNSDLDjxjeysfI0vGIfM/Rrb5vS/nH
yQW7URVUTxJh/IzxdgIOlXdpAYoj+qccv1moW9jUa9ytlu6+123WFxnPtrVmoF4fNpWgHChhYd3g
GLq92zJZvIrw9LE4evMfb6dc1oVh5Q8fMsdgPil32Jd6nBKGjXuFm5q8V8JbOTVFvdIi1/AkEFU3
yl5rjRmw4liaPKnmAHx1+lBd1T7vU1MC5d4OsYaFqH1swN88sGGVwytB1pKQx0F+kpFlfl/g/oSA
8f5E52csn9hwf7K+9XRHtG0BwBEpzVZNaUBBi5JkMIY1zWlyw+L0QlqSsGIBLHFlXNwdhePtdMSU
Ct1gZPz3bJl3j3u9in8g5NSC2c7pmfdJHoWjsKnTUvrjFUqj+74CwFuwKvepeyVR4EXpl/j1Sasi
e0ku2av7pLHpQvoJNPJ42bvWlY9XVd7MwPxJtsHcp1uyuGQCJgqgzzEVftJuOWosrjToNT8AjYnL
7LEIkHhm2mGxha34wyEJEhDpd/VfsarKpbiYHzyKPR46N2AQcVsZ7mhMBu6o0Y+OTpNjacc85hxv
0+ed9KW4gsu4udhJGOTAVWYnlaDqZc+Q6BTaCyIXZh9HWsVm0jl9zi5KTlAHdrMsSddtB207V6S9
gP98fuMYnQLboiOrnDdNtnEYz22DDiyEDqiAcm4lMnEL5Wc9bjhQwG26FPXPBqpYkPCDGIHTSlBN
tZ9TYUZkh4ZZ86A9uMPagj3BcFO0jNIPAhwInhGE5LhFiCzkcVR/OhIdFX2wJSua3RVwbLPDwlWW
AUQAEC99EY2jiWnQLiBXQRGJET/VMkcIyK2mhyOlJPaU75B6iNhDroU3aJL+xKfLdId06cMKyVcP
tvP+KrvvvD0wHsG+Fj3OeP7Kdm7N8GJ27sEr5IEsXeHZe5dJLYd17Jl9ZPcev5p4G6uhJWWCRZ+G
chiZ1rnChUOJobVsrP25D7xFPnDHLekHlhHOGWs4HjU8Uw05Qa5jcToNHUyLH7Ws+J6caOWpDfP5
tDG1RGCqYa4j5cMMUxYBK8SBaVDmxTOppMIi6WtikXOyRg8tGqu0XYvaMi7z1bZl9DxG+yhLT/td
SRPzBwlU3+uNPG+ooyz6VEn7bXsbw3DqODwzmmIZfWoh2pA2Hq6AWCg5yLMseaWFn4Vp496IfCdO
iMkeT+yuC1IZmeAM2MsVaxNjW+JwOtXgmjttMLlLCt2MOlddjuRqEtl9moTUKbmbsr4w2A+BQw6g
NdbXbfGS4wEwYE19pe6l5/jPBRLWaRKq0fHpxlk/npCdUVm5pLNih7nINRvDxCCrmndUTiAILCLT
5GFfXbgerJ7oucQRSoouHMaKj/nrdVjf79ySIVeMMjRfFWL2VKELXabufF5hzvwnNGqqcizhrrJE
RINfBqlsEqJHHIZKx0Qhfj60hMeUvlQQtCEmDjzIL9ehzPGIBUmJVYNyot3KLRdn4cIgcH+KP7Bo
4sGZTXPdE/MCjmoUJP/8IQ9Tktx9OdXcMFn5NtQHDVhNJqQr+N2m6sxeSCtgvaMeBxmsoAdbw1NG
xruCBVQvzszJ3wtojMh+Fm1iTvijaWJUIqj737TWZMFNx+zhBdTaIHX1c/KtgBlwirQt+JFD/Vmt
bG/gKtjcuVuk/OZSIFzlHpNeucDdAZCwUgb8EVNONIv0UnGzKQDTUJwzcD28YHaWl3Fe7s55NutT
nJswD4MFLzFTmH419/ARYFTKchmgSV25k2uUXJ5S4vWXw51UjG2awdNBppGb2dCFYCU2o2/jEee1
d1xOj8JWtLlVZ9GFX+xSLOj2kwbRf6i7depQF06TCwGV7biwST1i05sBaeA8rC9Fi195nUDk19B9
Ou6Dj9Exxyje/Ljc/stbjlLQz5aJ3NEyZHE5+qKLwIX3DBpAMG/u5PqQ5eAhDcdQdgLeYdYX3v8A
wGMNnGOQXIhnUDmLCHXhzRrQFROxnOwo9OaiNXwXvCcsUHFsxWUlEQ4lQxm3VbpJetC7uxQIyr9L
qdu/yys6Wog2kWjhnSROAZEyT1ZPUB8Lm1wsAMNcxUWkuLPeaD3WBlmA+PJDcgrqLGAQ8ZBMFunB
6PTF8GwF0SP5hpd4nqbEa+bOkHEpW0R7XP+v8yv/t19ARItTidB/rmcgy/NPX301J/9z+cFXX97X
/E/65sFX6P/+4E8PbvM//x3XgkA0ksrN2HhsTyKkCi9bGWElQzdnvo2YMqohSTxl6IHAyX8ucpiz
Q5JvlALXE44TGRf3GRmIyVDwxsCY+EZTpqy9PrSaMY62pZKK2KK08iG7eXwmncaYdBHy/Nc//yH/
I67JSkaGBtfu4z/K/4D0Lwi0UXPrBTxOG9s7LK1b/Op/W99+ubXz3XpUJGNx9UlcuRhZ4DdaRu0B
b+KiT9e3gXyVqssJmn4bnSxgcSNkurO+rYEKzfVjPyWscvZQ0Chv0mFvcsGmZKzayttYoXKhBjVN
/oirQkAAdcpZJK245n/YXVmW4PCyZLsU6eQo6SkMxNnhaZaf0MTWX2xF3DDlMBtKR3W/n0v0lF0u
h6fJ8F0HYWFJjnX0N0xIFJ6b7q3+4yWGaeriiQi3Xgi+EBCGtSs7985BkDpsB0zrSHfHUXPdm6hM
8kISGOwMo63tF692Qe7pOeb3r3/8E2dCN0TQLOTDkYIGZJpmSV8Pkkxd7rwwqEOfZD1e/JZdym+3
tje6spLL4QlA3P93J+SbPROb/7X7cp3TB2wwo0zFS0TCF6iyCzb7HhTIMScNNNn1GlCJpKLQUvdO
z/oMVrsrfV+8Pt1F2PAQ1C4DcKEUUkywhSOuEcQ5zU2nH8QexGQylZ9enjvCRJyQZvGAnu7s0/vt
dOm8r0rxLxH3EpIp9LisiP5lB6GfN92rTNLeaKIH/AjhCD7VRKOFpExoxkZmmDKAsFOGoumlcSva
BbkjYTOJHq8HzMFokxLZ2hkQNcPfImXY64+fbbaWov5kNI40zYGWHP/1LkUWMT/V9bfUrjqaGdM3
p8I7E45V2AUJPtlY/2Gn+/X9ByHFv04Po1db/xeInl5rhZvzTNQzZZsYnMs7KhfMoX+fclkicoq1
LJW2Ao1+wDsR7QaFbXMO+TA9Z7LMzxgZwJjwpmOFyg4BE8LfDnBPEsJNs3K02sbj6Qlx9NrkDsu0
DNhlNsyQm5TlqQbVBepOoPIOwZXTCY6NRELgxFJENS/jF/YVA6HwQmhiVSN3hQa21m2ruxLSwhaD
EKGzy+++mb+FCl5IDQTChFJDALeJHPfcOyTrf5XG1uEarHtoGk0TF12zxct0Ap+tb79af6pr3rqZ
paIpcJ4Wu/JVyipQoWnn4QGySGEr3JqR18GjaTJ5vtt5RuMNRihyhN8Dhcn0K85IMy1mgwQ0Yt7P
4bE9B+GdOvXAI0IMqMA6Go1nslecF7eq9Oz9F1vb3Vtz9fa6vW6v2+v2ur1ur9vr9rq9bq/b6/a6
vW6v2+v2ur1urz/49T94yipqAMgAAA==
+59
View File
@@ -0,0 +1,59 @@
# OpenBao stack — hardened tape-encryption key store for Kanrisha.
#
# Topology (same-LAN, NO reverse proxy):
# Kanrisha tape host --https--> openbao:8200 (native TLS on the vault itself)
#
# Unlike the other stacks in this repo there is no Caddy / Let's Encrypt: a
# secrets store terminates TLS itself and is reached over the LAN, not the
# public internet. TLS is a self-signed cert (deploy.sh generates it) or a
# CA-signed pair you drop in ./tls (e.g. via your Smallstep CA over ACME).
name: openbao
volumes:
openbao-data: # raft integrated storage — the vault's persistent state
services:
openbao:
image: openbao/openbao:${OPENBAO_TAG:-2.6.2}
container_name: openbao
command: server -config=/openbao/config/config.hcl
restart: unless-stopped
# These two are INERT. OpenBao dropped mlock support entirely, so nothing in
# the container locks memory and `disable_mlock = false` in config.hcl is a
# hard startup error. They are kept only because they are harmless and
# removing them would recreate every deployed container for no gain.
# Key material is kept off disk by DISABLING SWAP, which deploy.sh does.
cap_add:
- IPC_LOCK
ulimits:
memlock: -1
ports:
# Native-TLS API. Publish on the LAN IP only (OPENBAO_BIND) so it is not
# exposed on unrelated interfaces. NOTE: a Docker-published port is DNAT'd
# and bypasses the host INPUT firewall, so a source restriction there does
# NOT gate it — narrow access with OPENBAO_BIND + the optional mTLS
# (tls_require_and_verify_client_cert) stanza in config.hcl, and initialise
# the vault immediately (an uninitialised vault can be init'd by anyone who
# can reach it). deploy.sh defaults OPENBAO_BIND to OPENBAO_ADDR when it is
# an IP.
- "${OPENBAO_BIND:-0.0.0.0}:8200:8200"
volumes:
- ./config.hcl:/openbao/config/config.hcl:ro
- ./tls:/openbao/tls:ro
- openbao-data:/openbao/data
environment:
# PKCS#11 auto-unseal PIN — only read when the seal "pkcs11" stanza is
# enabled in config.hcl. Harmless (unused) otherwise.
BAO_HSM_PIN: "${OPENBAO_HSM_PIN:-}"
# HA / multi-node only: advertise this node's reachable LAN address.
# BAO_API_ADDR: "https://${OPENBAO_ADDR}:8200"
# BAO_CLUSTER_ADDR: "https://${OPENBAO_ADDR}:8201"
# 0 = unsealed (healthy); 2 = sealed -> unhealthy so a node needing an unseal
# is visible. -tls-skip-verify because the listener may use a self-signed cert.
healthcheck:
test: ["CMD", "bao", "status", "-address=https://127.0.0.1:8200", "-tls-skip-verify"]
interval: 30s
timeout: 5s
retries: 3
start_period: 10s
+39
View File
@@ -0,0 +1,39 @@
#!/usr/bin/env bash
#
# gen-tls.sh — generate a self-signed TLS cert + key for the OpenBao listener,
# if one is not already present. Never overwrites existing files, so to use a
# CA-signed cert (e.g. issued by your Smallstep CA over ACME) just drop tls.crt +
# tls.key into the target dir instead and this becomes a no-op.
#
# Usage: bash gen-tls.sh [TLS_DIR] (default ./tls)
# Env: OPENBAO_TLS_SANS comma SAN list, e.g. "DNS:vault.lan,IP:10.0.0.10,IP:127.0.0.1"
# OPENBAO_TLS_DAYS validity in days (default 825 — the max browsers accept)
set -euo pipefail
DIR="${1:-./tls}"
SANS="${OPENBAO_TLS_SANS:-DNS:localhost,IP:127.0.0.1}"
DAYS="${OPENBAO_TLS_DAYS:-825}"
mkdir -p "$DIR"
if [[ -f "$DIR/tls.crt" && -f "$DIR/tls.key" ]]; then
echo "[gen-tls] TLS material already present in $DIR — leaving it alone."
exit 0
fi
command -v openssl >/dev/null 2>&1 || { echo "[gen-tls] openssl not found" >&2; exit 1; }
# ECDSA P-256 self-signed cert. The SAN must include every address the tape host
# (and CLI) use to reach the vault, or their TLS verification fails.
openssl req -x509 -newkey ec -pkeyopt ec_paramgen_curve:prime256v1 -nodes \
-keyout "$DIR/tls.key" -out "$DIR/tls.crt" -days "$DAYS" \
-subj "/CN=openbao" -addext "subjectAltName=${SANS}"
chmod 600 "$DIR/tls.key"
chmod 644 "$DIR/tls.crt"
echo "[gen-tls] Generated self-signed OpenBao cert in $DIR"
echo "[gen-tls] SANs: ${SANS} validity: ${DAYS}d"
echo "[gen-tls] Give $DIR/tls.crt to the Kanrisha daemon as"
echo "[gen-tls] [encryption.openbao].ca_cert (so it trusts this listener)."
+512
View File
@@ -0,0 +1,512 @@
#!/usr/bin/env bash
#
# update.sh -- update the OpenBao container. Companion to deploy.sh, installed
# alongside the stack.
#
# THIS UPDATER IS DELIBERATELY UNLIKE THE OTHERS IN THIS REPO. copyparty and
# ergo come back by themselves after a recreate. OpenBao comes back SEALED: with
# the default Shamir seal, a restart needs three unseal keys typed in by a human
# before the vault serves anything again. A scheduled update at 03:00 would
# therefore take the vault offline until someone turns up with the keys.
#
# So the scheduled path NEVER changes the running version by default. It reports.
# UPDATE_POLICY=auto opts in, and even then it refuses unless auto-unseal is
# configured -- because only then does the vault come back on its own.
#
# It also takes a raft snapshot before touching anything, which neither of the
# other updaters needs to do. OpenBao's upgrade guide is explicit that reverting
# the image alone does NOT roll back the data store, so that snapshot is the
# rollback plan, not a formality. The snapshot is token-gated and cannot be taken
# from a sealed vault, both of which this script checks up front rather than
# failing halfway through.
#
# Subcommands:
# check (default) declared vs running vs latest, seal state, and whether
# an unattended update would be safe here. Changes nothing.
# snapshot take a raft snapshot and leave it in SNAPSHOT_DIR. Nothing else.
# update do it now: preflight -> snapshot -> down -> pin -> pull -> up
# run what the schedule invokes; obeys UPDATE_POLICY
# install schedule the daily `run`
# uninstall remove the schedule
#
# Policy (UPDATE_POLICY):
# notify (default) never change the running version; report only
# auto update when a newer release exists -- but ONLY if auto-unseal is
# configured. Without it, `run` reports and does nothing.
#
# Env (also read from /etc/openbao-update.conf; environment wins):
# STACK_DIR=/srv/openbao UPDATE_POLICY=notify
# BAO_TOKEN= required by `update`/`snapshot`. Prefer NOT to set
# this inline -- see "Supplying the token" below.
# BAO_TOKEN_FILE= read the token from a file (0600) instead
# SNAPSHOT_DIR=/var/backups/openbao
# TARGET_VERSION= pin a specific version instead of latest
# DRY_RUN=0 print what would happen, change nothing
# SKIP_SNAPSHOT=0 DANGEROUS: upgrade with no rollback plan
# GH_REPO=openbao/openbao
#
# Supplying the token:
# `BAO_TOKEN=<root> bash update.sh update` works, but DO NOT do it: the value
# lands in your shell history, and anything on the argv of a child process is
# readable by every user on the box via /proc/<pid>/cmdline. This script never
# puts the token on any command line -- it hands it to the container over
# stdin -- and it will prompt for it with echo off if you do not supply one.
# For unattended use, point BAO_TOKEN_FILE at a 0600 file.
#
# Usage:
# bash update.sh check # no token needed
# bash update.sh snapshot # prompts for the token
# bash update.sh update # prompts for the token
# BAO_TOKEN_FILE=/root/.bao-root bash update.sh update
# TARGET_VERSION=2.6.2 bash update.sh update
set -euo pipefail
SELF="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)/$(basename "${BASH_SOURCE[0]}")"
: "${OPENBAO_UPDATE_CONF:=/etc/openbao-update.conf}"
if [[ -r "$OPENBAO_UPDATE_CONF" ]]; then
# Environment wins over the conf: only set what is not already set.
while IFS= read -r _line || [[ -n "$_line" ]]; do
[[ "$_line" =~ ^[[:space:]]*# || -z "${_line//[[:space:]]/}" ]] && continue
_k="${_line%%=*}"; _v="${_line#*=}"; _k="${_k//[[:space:]]/}"
[[ -n "$_k" ]] || continue
[[ -n "${!_k+x}" ]] || printf -v "$_k" '%s' "$_v"
done < "$OPENBAO_UPDATE_CONF"
fi
: "${STACK_DIR:=/srv/openbao}"
: "${UPDATE_POLICY:=notify}"
: "${BAO_TOKEN:=}"
: "${BAO_TOKEN_FILE:=}"
: "${SNAPSHOT_DIR:=/var/backups/openbao}"
: "${TARGET_VERSION:=}"
: "${DRY_RUN:=0}"
: "${SKIP_SNAPSHOT:=0}"
: "${GH_REPO:=openbao/openbao}"
ENV_FILE="$STACK_DIR/.env"
CONFIG_HCL="$STACK_DIR/config.hcl"
log() { printf '\033[1;32m[+]\033[0m %s\n' "$*"; }
warn() { printf '\033[1;33m[!]\033[0m %s\n' "$*" >&2; }
die() { printf '\033[1;31m[x]\033[0m %s\n' "$*" >&2; exit 1; }
# docker compose scoped to the stack dir, so ./config.hcl, ./tls and .env resolve.
dc() { ( cd "$STACK_DIR" && docker compose "$@" ); }
set_env() { # <KEY> <value>: update KEY in .env, or append if absent
# Value goes through the ENVIRONMENT, never interpolated into a sed script --
# see 947c899 / 185f404 for why.
local key="$1" val="$2" tmp
if [[ ! -f "$ENV_FILE" ]]; then
printf '%s=%s\n' "$key" "$val" >> "$ENV_FILE"
return 0
fi
tmp="$(mktemp)"
_SE_KEY="$key" _SE_VAL="$val" awk '
BEGIN { k = ENVIRON["_SE_KEY"]; v = ENVIRON["_SE_VAL"]; seen = 0 }
!seen && index($0, k "=") == 1 { print k "=" v; seen = 1; next }
{ print }
END { if (!seen) print k "=" v }
' "$ENV_FILE" > "$tmp"
cat "$tmp" > "$ENV_FILE"
rm -f "$tmp"
}
env_get() { # <KEY> -> value from .env ('' if absent)
[[ -f "$ENV_FILE" ]] || { printf ''; return 0; }
sed -n "s/^$1=//p" "$ENV_FILE" | tail -n1
}
# ---------------------------------------------------------------------------
# Token handling
#
# The token never appears on a command line. `docker exec -e BAO_TOKEN=<value>`
# would put it in the docker process's argv, and /proc/<pid>/cmdline is
# world-readable -- so it would leak to every user on the host, not just to the
# shell history. Instead the value is written to the container's stdin and the
# container's own shell reads it into the environment it then execs `bao` with.
# ---------------------------------------------------------------------------
acquire_token() {
[[ -n "$BAO_TOKEN" ]] && return 0
if [[ -n "$BAO_TOKEN_FILE" ]]; then
[[ -r "$BAO_TOKEN_FILE" ]] || die "BAO_TOKEN_FILE=${BAO_TOKEN_FILE} is not readable."
BAO_TOKEN="$(head -n1 "$BAO_TOKEN_FILE")"
[[ -n "$BAO_TOKEN" ]] || die "BAO_TOKEN_FILE=${BAO_TOKEN_FILE} is empty."
return 0
fi
if [[ -t 0 ]]; then
read -rs -p "OpenBao token (input hidden): " BAO_TOKEN
printf '\n' >&2
[[ -n "$BAO_TOKEN" ]] || die "No token entered."
return 0
fi
die "No token available and stdin is not a terminal. Set BAO_TOKEN_FILE=/path/to/a/0600/file, or run interactively to be prompted. Avoid BAO_TOKEN=... on the command line: it goes into your shell history and into /proc/<pid>/cmdline."
}
# Run `bao <args...>` inside the container with BAO_TOKEN set, passing the token
# over stdin so it never reaches an argv anywhere on the host.
bao_with_token() {
printf '%s\n' "$BAO_TOKEN" | dc exec -T openbao sh -c '
read -r _t
BAO_TOKEN="$_t"
export BAO_TOKEN
exec "$@"
' sh bao "$@"
}
# ---------------------------------------------------------------------------
# State
# ---------------------------------------------------------------------------
declared_version() { env_get OPENBAO_TAG; }
running_version() {
# `bao status` prints "Version 2.6.2". Works sealed or unsealed.
# `|| true` inside the substitution: status exits 2 when sealed, and this is
# a bare assignment at every call site.
dc exec -T openbao bao status -address=https://127.0.0.1:8200 -tls-skip-verify 2>/dev/null \
| awk '/^Version/ { print $2; exit }' || true
}
# 0 = unsealed, 2 = sealed, anything else = could not tell.
seal_state() {
local rc=0
dc exec -T openbao bao status -address=https://127.0.0.1:8200 -tls-skip-verify \
>/dev/null 2>&1 || rc=$?
printf '%s' "$rc"
}
container_running() {
local id
id="$(dc ps -q openbao 2>/dev/null || true)"
[[ -n "$id" ]] || return 1
[[ "$(docker inspect -f '{{.State.Status}}' "$id" 2>/dev/null || echo unknown)" == running ]]
}
# An uncommented `seal "..." {` stanza means the vault unseals itself, which is
# the only condition under which an unattended update is defensible.
auto_unseal_configured() {
[[ -f "$CONFIG_HCL" ]] || return 1
grep -qE '^[[:space:]]*seal[[:space:]]+"' "$CONFIG_HCL"
}
latest_version() {
# No jq on a stock Alpine host; parse the tag out with sed.
curl -fsSL --max-time 20 "https://api.github.com/repos/${GH_REPO}/releases/latest" 2>/dev/null \
| sed -n 's/.*"tag_name"[[:space:]]*:[[:space:]]*"v\{0,1\}\([^"]*\)".*/\1/p' \
| head -n1 || true
}
# ---------------------------------------------------------------------------
# Preflight + snapshot
# ---------------------------------------------------------------------------
# The built-in pkcs11 seal is REMOVED in 2.7.0 (not merely deprecated), and the
# HSM distribution is discontinued. Crossing that line with the stanza active
# gives a vault that cannot unseal itself and cannot be unsealed by hand either.
check_pkcs11_cliff() { # <target-version>
local target="$1" major_minor
major_minor="${target%.*}"
case "$major_minor" in
2.7|2.8|2.9|3.*) ;;
*) return 0 ;;
esac
if [[ -f "$CONFIG_HCL" ]] && grep -qE '^[[:space:]]*seal[[:space:]]+"pkcs11"' "$CONFIG_HCL"; then
die "config.hcl has an active built-in seal \"pkcs11\" stanza and the target is ${target}. That stanza is REMOVED in 2.7.0 -- the vault would start with no way to unseal. Migrate to the external 'plugin \"kms\" \"pkcs11\"' first."
fi
}
take_snapshot() { # <target-version> -> echoes the snapshot path
local target="$1" stamp dest
stamp="$(date -u +%Y%m%dT%H%M%SZ)"
dest="${SNAPSHOT_DIR}/openbao-pre-${target}-${stamp}.snap"
if [[ "$SKIP_SNAPSHOT" == "1" ]]; then
warn "SKIP_SNAPSHOT=1 -- upgrading with NO rollback plan. Reverting the image alone does not roll back the data store."
printf ''
return 0
fi
acquire_token
if [[ "$DRY_RUN" == "1" ]]; then
echo "DRY: snapshot -> ${dest}" >&2
printf '%s' "$dest"
return 0
fi
install -d -m 0700 "$SNAPSHOT_DIR"
log "Taking a raft snapshot to ${dest}..." >&2
# Write inside the container, then stream the RAW file out. `compose cp` is
# NOT usable here: it emits a TAR wrapper that will not restore.
bao_with_token operator raft snapshot save \
-address=https://127.0.0.1:8200 -tls-skip-verify /tmp/pre-upgrade.snap \
>/dev/null || die "Snapshot failed -- not upgrading. Check the token has sys/storage/raft/snapshot."
dc exec -T openbao cat /tmp/pre-upgrade.snap > "$dest" || die "Could not stream the snapshot out -- not upgrading."
dc exec -T openbao rm -f /tmp/pre-upgrade.snap >/dev/null 2>&1 || true
chmod 0600 "$dest"
# A snapshot is a gzip-wrapped tar. An empty or truncated file here means the
# rollback plan does not exist, so treat it as fatal rather than cosmetic.
[[ -s "$dest" ]] || die "Snapshot at ${dest} is empty -- not upgrading."
if command -v gzip >/dev/null 2>&1; then
gzip -t "$dest" 2>/dev/null || die "Snapshot at ${dest} is not a valid gzip archive -- not upgrading."
fi
log "Snapshot OK ($(wc -c < "$dest") bytes). Copy it OFF this host." >&2
printf '%s' "$dest"
}
# ---------------------------------------------------------------------------
# Subcommands
# ---------------------------------------------------------------------------
do_check() {
[[ -d "$STACK_DIR" ]] || die "No stack at $STACK_DIR (set STACK_DIR)."
local declared running latest state auto
declared="$(declared_version)"
latest="$(latest_version)"
if container_running; then
running="$(running_version)"
state="$(seal_state)"
else
running=""
state="down"
fi
auto_unseal_configured && auto=yes || auto=no
printf ' declared (.env OPENBAO_TAG): %s\n' "${declared:-<unset>}"
printf ' running (bao status): %s\n' "${running:-<not running>}"
printf ' latest (%s): %s\n' "$GH_REPO" "${latest:-<lookup failed>}"
case "$state" in
0) printf ' seal state: UNSEALED\n' ;;
2) printf ' seal state: SEALED (needs 3 keys)\n' ;;
down) printf ' seal state: container not running\n' ;;
*) printf ' seal state: unknown (bao status rc=%s)\n' "$state" ;;
esac
printf ' auto-unseal configured: %s\n' "$auto"
printf ' policy: %s\n' "$UPDATE_POLICY"
if [[ -z "$latest" ]]; then
warn "Could not reach the GitHub releases API; cannot say whether an update exists."
return 0
fi
if [[ "$declared" == "$latest" ]]; then
log "Up to date."
return 0
fi
log "Update available: ${declared:-?} -> ${latest}"
if [[ "$auto" == no ]]; then
warn "A restart will leave the vault SEALED until someone enters three unseal keys."
warn "Run it when you can do that: bash ${SELF} update (it will prompt for a token)"
fi
}
do_update() {
[[ $EUID -eq 0 ]] || die "Run as root."
[[ -d "$STACK_DIR" ]] || die "No stack at $STACK_DIR (set STACK_DIR)."
local from to snap
from="$(declared_version)"
to="${TARGET_VERSION:-$(latest_version)}"
[[ -n "$to" ]] || die "Could not determine a target version (GitHub lookup failed). Set TARGET_VERSION=x.y.z."
if [[ "$from" == "$to" ]]; then
log "Already pinned to ${to}; nothing to do."
return 0
fi
check_pkcs11_cliff "$to"
container_running || die "The openbao container is not running. Start it first: cd ${STACK_DIR} && docker compose up -d"
# A sealed vault cannot produce a snapshot, so there would be no rollback
# plan. Catch it here rather than after the container is already down.
local state; state="$(seal_state)"
case "$state" in
0) ;;
2) die "The vault is SEALED. A snapshot cannot be taken from a sealed vault, so there would be no rollback plan. Unseal first, then re-run." ;;
*) die "Could not read the seal state (bao status rc=${state}). Refusing to upgrade blind." ;;
esac
log "Updating OpenBao: ${from:-?} -> ${to}"
snap="$(take_snapshot "$to")"
if [[ "$DRY_RUN" == "1" ]]; then
echo "DRY: docker compose down; set OPENBAO_TAG=${to}; docker compose pull; docker compose up -d"
return 0
fi
# Clean shutdown before swapping the image, as the upgrade guide asks.
log "Stopping the stack (the vault will be sealed when it returns)..."
dc down
set_env OPENBAO_TAG "$to"
if ! dc pull; then
warn "Pull of ${to} failed; rolling the pin back to ${from:-<unset>}."
[[ -n "$from" ]] && set_env OPENBAO_TAG "$from"
dc up -d || true
die "Update aborted. The previous version is starting again; unseal it."
fi
if ! dc up -d --remove-orphans; then
warn "Starting ${to} failed; rolling the pin back to ${from:-<unset>}."
[[ -n "$from" ]] && set_env OPENBAO_TAG "$from"
dc up -d || true
die "Update aborted. If the data store is at fault, restore ${snap:-the snapshot} per the README's DR runbook."
fi
log "Started ${to}."
cat <<EOF
================================================================
UPDATED ${from:-?} -> ${to}
The vault is SEALED. Nothing works until you unseal it:
cd ${STACK_DIR}
docker compose exec -e BAO_ADDR=https://127.0.0.1:8200 openbao \\
bao operator unseal -tls-skip-verify # x3, three different keys
Then verify:
docker compose exec -e BAO_ADDR=https://127.0.0.1:8200 openbao \\
bao status -tls-skip-verify # Version ${to}, Sealed false
Rollback, if ${to} misbehaves: reverting the image alone does NOT roll back the
data store. Restore the pre-upgrade snapshot -- see the README's DR runbook.
snapshot: ${snap:-<none taken>}
================================================================
EOF
}
do_snapshot() {
[[ -d "$STACK_DIR" ]] || die "No stack at $STACK_DIR (set STACK_DIR)."
container_running || die "The openbao container is not running."
local state; state="$(seal_state)"
[[ "$state" == 0 ]] || die "The vault is SEALED (or unreadable): a snapshot can only be taken from an unsealed vault."
local dest; dest="$(take_snapshot manual)"
[[ -n "$dest" ]] || die "No snapshot was taken."
log "Snapshot written to ${dest} -- copy it OFF this host."
}
do_run() {
# What the schedule invokes. The whole point of this branch is that it is
# conservative: an unattended update of a manually-unsealed vault takes it
# offline until a human arrives, so that is never the default.
case "$UPDATE_POLICY" in
notify)
do_check
return 0 ;;
auto)
if ! auto_unseal_configured; then
do_check
warn "UPDATE_POLICY=auto but no seal stanza is configured, so the vault would stay SEALED after a restart with nobody present. Not updating."
return 0
fi
local from to
from="$(declared_version)"
to="${TARGET_VERSION:-$(latest_version)}"
if [[ -z "$to" || "$from" == "$to" ]]; then
do_check
return 0
fi
log "UPDATE_POLICY=auto and auto-unseal is configured; updating ${from:-?} -> ${to}."
do_update
return 0 ;;
*)
die "UPDATE_POLICY must be 'notify' or 'auto' (got '${UPDATE_POLICY}')." ;;
esac
}
write_conf() {
install -d -m 0755 "$(dirname "$OPENBAO_UPDATE_CONF")"
# BAO_TOKEN is deliberately NOT written here: a long-lived root token sitting
# in a conf file next to the vault it opens defeats the point of the vault.
cat > "$OPENBAO_UPDATE_CONF" <<EOF
# openbao-update.conf -- read by update.sh; the environment wins over this file.
STACK_DIR=${STACK_DIR}
UPDATE_POLICY=${UPDATE_POLICY}
SNAPSHOT_DIR=${SNAPSHOT_DIR}
GH_REPO=${GH_REPO}
# BAO_TOKEN is intentionally absent: a long-lived root token in a file next to
# the vault it opens defeats the vault. \`update\`/\`snapshot\` prompt for one.
# For unattended use set BAO_TOKEN_FILE below to a 0600 file holding a token
# with sys/storage/raft/snapshot -- not the root token.
#BAO_TOKEN_FILE=/root/.bao-snapshot-token
EOF
chmod 0600 "$OPENBAO_UPDATE_CONF"
}
do_install() {
[[ $EUID -eq 0 ]] || die "Run as root."
write_conf
if [[ "$UPDATE_POLICY" == notify ]]; then
log "Scheduling a daily CHECK (policy=notify: it will never change the running version)."
else
log "Scheduling a daily run (policy=${UPDATE_POLICY})."
fi
if [[ -r /etc/os-release ]] && grep -q '^ID=alpine' /etc/os-release; then
install -d -m 0755 /etc/periodic/daily
cat > /etc/periodic/daily/openbao-update <<EOF
#!/bin/sh
exec bash "$SELF" run
EOF
chmod +x /etc/periodic/daily/openbao-update
if command -v rc-update >/dev/null 2>&1; then
rc-update add crond default >/dev/null 2>&1 || true
rc-service crond start >/dev/null 2>&1 || true
fi
log "Installed /etc/periodic/daily/openbao-update."
else
cat > /etc/systemd/system/openbao-update.service <<EOF
[Unit]
Description=OpenBao container updater
After=docker.service
[Service]
Type=oneshot
ExecStart=/usr/bin/env bash $SELF run
EOF
cat > /etc/systemd/system/openbao-update.timer <<EOF
[Unit]
Description=Daily OpenBao update check
[Timer]
OnCalendar=daily
Persistent=true
RandomizedDelaySec=1h
[Install]
WantedBy=timers.target
EOF
systemctl daemon-reload
systemctl enable --now openbao-update.timer >/dev/null 2>&1 || true
log "Installed the openbao-update.timer systemd timer."
fi
}
do_uninstall() {
[[ $EUID -eq 0 ]] || die "Run as root."
rm -f /etc/periodic/daily/openbao-update
if command -v systemctl >/dev/null 2>&1; then
systemctl disable --now openbao-update.timer >/dev/null 2>&1 || true
rm -f /etc/systemd/system/openbao-update.timer /etc/systemd/system/openbao-update.service
systemctl daemon-reload >/dev/null 2>&1 || true
fi
log "Schedule removed. ${OPENBAO_UPDATE_CONF} left in place."
}
case "${1:-check}" in
check) do_check ;;
snapshot) do_snapshot ;;
update) do_update ;;
run) do_run ;;
install) do_install ;;
uninstall) do_uninstall ;;
*) die "Usage: $(basename "$0") {check|snapshot|update|run|install|uninstall}" ;;
esac
+45 -1
View File
@@ -40,6 +40,15 @@ set -euo pipefail
[[ "$SKIP_PROMPTS" == "1" ]] && FORCE=1
: "${POCKETID_DOMAIN:=}"
: "${ACME_EMAIL:=}"
# Whether the WebFinger keys actually arrived in this script's ENVIRONMENT has to
# be recorded BEFORE the ":=" defaults below, which make an unset variable look
# like an empty one. The .env is what decides whether the WebFinger block is
# served, so a value passed to a re-run has to be written through to it.
WF_KEYS_FROM_ENV=()
for _k in BASE_DOMAIN REDIRECT_URL; do
if [[ -n "${!_k+x}" ]]; then WF_KEYS_FROM_ENV+=("$_k"); fi
done
: "${BASE_DOMAIN:=}" # optional: enables the WebFinger block (set with REDIRECT_URL)
: "${REDIRECT_URL:=}" # optional: where the base domain 301s non-webfinger traffic
@@ -171,6 +180,28 @@ install -m 0640 "$SCRIPT_DIR/docker-compose.yml" "$STACK_DIR/docker-compose.yml"
# can be appended when BASE_DOMAIN is set.)
ENV_FILE="$STACK_DIR/.env"
set_env() { # <KEY> <value>: update KEY in .env, or append if absent
# The value goes through the ENVIRONMENT, never interpolated into a sed
# script. Interpolating it corrupts any value containing & (sed expands it to
# the whole match) and aborts the run on one containing the s||| delimiter --
# which is reachable for an OIDC secret, a password or a URL query string.
local key="$1" val="$2" tmp
if [[ ! -f "$ENV_FILE" ]]; then
printf '%s=%s
' "$key" "$val" >> "$ENV_FILE"
return 0
fi
tmp="$(mktemp)"
_SE_KEY="$key" _SE_VAL="$val" awk '
BEGIN { k = ENVIRON["_SE_KEY"]; v = ENVIRON["_SE_VAL"]; seen = 0 }
!seen && index($0, k "=") == 1 { print k "=" v; seen = 1; next }
{ print }
END { if (!seen) print k "=" v }
' "$ENV_FILE" > "$tmp"
cat "$tmp" > "$ENV_FILE" # rewrite in place: keeps the original mode/owner
rm -f "$tmp"
}
if [[ ! -f "$ENV_FILE" ]]; then
log "Seeding $ENV_FILE with generated secrets..."
install -m 0600 "$SCRIPT_DIR/.env.example" "$ENV_FILE"
@@ -183,7 +214,18 @@ if [[ ! -f "$ENV_FILE" ]]; then
-e "s|^REDIRECT_URL=.*|REDIRECT_URL=${REDIRECT_URL}|" \
"$ENV_FILE"
else
log ".env exists; leaving secrets alone."
log ".env exists; keeping it (secrets and values not passed are left alone)."
# The WebFinger decision below is read from .env, so a BASE_DOMAIN passed to a
# re-run has to land there -- otherwise enabling WebFinger is a silent no-op
# that still reports DEPLOYED. Only keys actually passed this run are touched:
# automations.sh drops a blank optional answer, so "blank" is indistinguishable
# from "not supplied" and must not be treated as "disable".
for _k in ${WF_KEYS_FROM_ENV[@]+"${WF_KEYS_FROM_ENV[@]}"}; do
_cur=$(sed -n "s/^${_k}=//p" "$ENV_FILE" | tail -n1)
if [[ "$_cur" == "${!_k}" ]]; then continue; fi
set_env "$_k" "${!_k}"
log " ${_k}: ${_cur:-<unset>} -> ${!_k:-<empty>}"
done
fi
# Validate required values are present.
@@ -202,6 +244,8 @@ ru="$(sed -n 's/^REDIRECT_URL=//p' "$ENV_FILE")"
if [[ -n "$bd" ]]; then
[[ -n "$ru" ]] || die "BASE_DOMAIN is set but REDIRECT_URL is empty in $ENV_FILE; set both or neither."
log "WebFinger enabled -- serving /.well-known/webfinger at ${bd}."
log " (to retire it, clear BASE_DOMAIN + REDIRECT_URL in $ENV_FILE and re-run --"
log " leaving the prompt blank cannot disable it, as a blank answer is dropped)"
cat "$SCRIPT_DIR/Caddyfile.webfinger" >> "$STACK_DIR/Caddyfile"
else
log "No BASE_DOMAIN -- pocket-id only (use the webfinger deployment for OIDC discovery)."
+1
View File
@@ -98,6 +98,7 @@ TARGETS=(
# Current environment and docker-compose setup
"$SIMPLEX_DIR/.env"
"$SIMPLEX_DIR/docker-compose.yml"
"$SIMPLEX_DIR/caddy_conf/Caddyfile"
"$SIMPLEX_DIR/print-addresses.sh"
# Tor configs
+56 -42
View File
@@ -221,7 +221,7 @@ done
# 5. Lay out /opt/simplex
# ----------------------------------------------------------------------------
log "Writing compose stack to ${INSTALL_DIR}..."
mkdir -p "$INSTALL_DIR"/tor_conf
mkdir -p "$INSTALL_DIR"/tor_conf "$INSTALL_DIR"/caddy_conf
cd "$INSTALL_DIR"
cat > .env <<EOF
@@ -235,53 +235,56 @@ CERT_PATH=${CERT_PATH}
EOF
chmod 600 .env
# Caddyfile. This used to be produced by a caddy-init container whose `command:`
# was a YAML FOLDED scalar (`>`): the closing `}` and the heredoc's `EOF` were
# both at the scalar's base indent, so YAML folded them into `} EOF`, leaving the
# heredoc unterminated and the script a parse error. caddy-init exited 2 on every
# run, `service_completed_successfully` never fired, and nothing downstream of
# caddy ever started. Generating the file here removes the whole class of problem
# -- and, unlike the old `if [ ! -f ]` guard, applies a changed DOMAIN/ACME_EMAIL/
# KEY_TYPE on a re-run instead of freezing them at the first deploy.
#
# The heredoc is deliberately UNQUOTED so ${DOMAIN}/${ACME_EMAIL}/${KEY_TYPE}
# expand here. Caddy's own {uri} has no `$` so it survives; any future Caddy
# {$ENV} placeholder must be escaped as \${...}.
CADDY_CHANGED=0
cat > caddy_conf/Caddyfile.new <<EOF
{
email ${ACME_EMAIL}
}
http://smp.${DOMAIN} {
redir https://smp.${DOMAIN}{uri} permanent
}
smp.${DOMAIN}:8443 {
tls { key_type ${KEY_TYPE} }
reverse_proxy smp-server:8000
}
http://xftp.${DOMAIN} {
redir https://xftp.${DOMAIN}{uri} permanent
}
xftp.${DOMAIN}:8443 {
tls { key_type ${KEY_TYPE} }
reverse_proxy xftp-server:8000
}
EOF
# `cmp -s` exits 1 when the files differ, so it stays inside an `if` condition --
# never `cmp -s A B && CADDY_CHANGED=1`, which trips set -e when they match.
if [[ -f caddy_conf/Caddyfile ]] && ! cmp -s caddy_conf/Caddyfile caddy_conf/Caddyfile.new; then
CADDY_CHANGED=1
fi
mv -f caddy_conf/Caddyfile.new caddy_conf/Caddyfile
chmod 0644 caddy_conf/Caddyfile
cat > docker-compose.yml <<'YAML'
name: simplex
services:
caddy-init:
image: alpine:latest
command: >
sh -c '
if [ ! -f /etc/caddy/Caddyfile ]; then
cat > /etc/caddy/Caddyfile <<EOF
{
email ${ACME_EMAIL}
}
http://smp.${DOMAIN} {
redir https://smp.${DOMAIN}{uri} permanent
}
smp.${DOMAIN}:8443 {
tls { key_type ${KEY_TYPE} }
reverse_proxy smp-server:8000
}
http://xftp.${DOMAIN} {
redir https://xftp.${DOMAIN}{uri} permanent
}
xftp.${DOMAIN}:8443 {
tls { key_type ${KEY_TYPE} }
reverse_proxy xftp-server:8000
}
EOF
fi
'
environment:
DOMAIN: ${DOMAIN:?}
ACME_EMAIL: ${ACME_EMAIL:?}
KEY_TYPE: ${KEY_TYPE:-rsa4096}
volumes:
- ./caddy_conf:/etc/caddy
restart: "no"
caddy:
image: caddy:2-alpine
depends_on:
caddy-init:
condition: service_completed_successfully
cap_add:
- NET_ADMIN
ports:
@@ -494,7 +497,18 @@ rc-update add simplex default
log "Pulling images and starting the stack..."
cd "$INSTALL_DIR"
docker compose pull
# A bind-mounted file's CONTENTS are not part of the compose config hash, so a
# changed Caddyfile does not by itself cause a recreate and `up -d` would leave
# it unloaded. Only restart when it actually changed AND compose did not already
# replace the container -- restarting Caddy seconds after a first start would
# interrupt the initial ACME issuance for nothing.
_caddy_before="$(docker compose ps -q caddy 2>/dev/null || true)"
docker compose up -d
_caddy_after="$(docker compose ps -q caddy 2>/dev/null || true)"
if (( CADDY_CHANGED )) && [[ -n "$_caddy_before" && "$_caddy_before" == "$_caddy_after" ]]; then
log "Caddyfile changed; restarting Caddy to load it..."
docker compose restart caddy || warn "Could not restart Caddy; run: cd $INSTALL_DIR && docker compose restart caddy"
fi
log "Waiting up to 90s for Tor to publish hidden services..."
for _ in $(seq 1 90); do
+1
View File
@@ -238,6 +238,7 @@ SIMPLEX_FILES=(
"docker-compose.yml"
"print-addresses.sh"
"tor_conf/"
"caddy_conf/"
)
for item in "${SIMPLEX_FILES[@]}"; do
+41
View File
@@ -0,0 +1,41 @@
# Copy to .env and fill in. docker compose picks .env up automatically.
#
# Squid SSL-bump caching forward proxy. Unlike the other stacks there is NO
# public hostname / Let's Encrypt cert -- this is a forward proxy. The TLS
# interception CA is generated on first deploy (deploy.sh) and never overwritten.
# ─── Who may use the proxy ──────────────────────────────────────────────────
# Space-separated CIDR(s) allowed to connect (Squid http_access). This is the
# REAL access gate -- keep it tight. Examples: 100.64.0.0/10 (Tailscale CGNAT),
# 10.0.0.0/8, 192.168.0.0/16.
TRUSTED_CIDR=100.64.0.0/10
# Host interface/IP to publish the proxy on, and the host-side port. Pin
# BIND_ADDR to a trusted interface (e.g. your Tailscale IP) -- a published
# Docker port BYPASSES the host INPUT firewall, so 0.0.0.0 exposes the proxy to
# every reachable network. Use 0.0.0.0 only if TRUSTED_CIDR + upstream
# firewalling already cover you.
BIND_ADDR=0.0.0.0
PROXY_PORT=3128
# ─── Cache sizing ───────────────────────────────────────────────────────────
CACHE_SIZE_MB=5000 # on-disk cache budget (MB)
MAX_OBJECT_SIZE_MB=256 # largest single object to cache (raise for ISOs/images)
CACHE_MEM_MB=256 # in-memory hot cache (MB)
# ─── Cache scope ────────────────────────────────────────────────────────────
# 0 = boost mode (default): cache everything per normal HTTP rules, and
# force-cache the domains in cache-domains.txt / .regex with long TTLs.
# 1 = strict allowlist: store ONLY the listed domains; pass the rest through.
CACHE_ONLY_LISTED=0
# ─── TLS interception CA (generated on first deploy) ────────────────────────
CA_CN=Squid Proxy CA
CA_O=automations
CA_DAYS=3650
DYNAMIC_CERT_MEM_MB=8 # in-RAM cache of generated per-host leaf certs (MB)
# ─── Misc ───────────────────────────────────────────────────────────────────
VISIBLE_HOSTNAME=squid-proxy
# Local build tag for the image (built from ./Dockerfile).
SQUID_IMAGE_TAG=automations/squid:latest
+22
View File
@@ -0,0 +1,22 @@
# Minimal Squid image with SSL-bump.
#
# Alpine ships its `squid` package built `--with-openssl`, so ssl-bump,
# https_port and security_file_certgen are all compiled in -- no
# compile-from-source needed (Debian/Ubuntu, by contrast, build squid against
# GnuTLS and need the separate `squid-openssl` package). openssl is included so
# deploy.sh can mint the CA via this image without a host openssl dependency.
FROM alpine:3.21
RUN apk add --no-cache squid ca-certificates openssl tini \
&& update-ca-certificates
COPY squid.conf.tmpl /etc/squid/squid.conf.tmpl
COPY entrypoint.sh /usr/local/bin/entrypoint.sh
RUN chmod +x /usr/local/bin/entrypoint.sh
# Explicit forward-proxy port. Caching/inspection happen here; clients set
# HTTP(S)_PROXY to this host:3128.
EXPOSE 3128
# tini reaps zombies and forwards signals so `docker stop` shuts squid cleanly.
ENTRYPOINT ["/sbin/tini", "--", "/usr/local/bin/entrypoint.sh"]
+143
View File
@@ -0,0 +1,143 @@
# squid — SSL-bump caching forward proxy
A [Squid](https://www.squid-cache.org/) **forward proxy** that caches static
content to cut bandwidth across a fleet and **intercepts TLS** (SSL-bump) with a
locally-generated CA so HTTPS can be cached and inspected.
This is the one deployment that **breaks the repo's Caddy/Let's-Encrypt
convention on purpose**: it is not an inbound web service, has no public
hostname, and uses no ACME cert. Clients point `HTTP(S)_PROXY` at it and trust
its CA.
> ⚠️ **TLS interception — read first.** SSL-bump decrypts your clients' HTTPS.
> The CA private key it generates can impersonate **any** site to **any** client
> that trusts it. Only run this on networks and devices **you own and are
> authorized to inspect**. Keep `ssl/squid-ca-key.pem` secret. Cert-pinned and
> HSTS sites (banking, app stores) will break unless you **splice** them
> (passthrough, see below).
## What it does
- **Explicit forward proxy** on port `3128`. Clients set `http_proxy` /
`https_proxy` (transparent/intercepting mode is future work — see end).
- **SSL-bump**: peek at the TLS SNI → **splice** (passthrough, no decryption)
the domains in `splice-domains.txt` → **bump** (decrypt) everything else,
minting per-host leaf certs on the fly from the local CA.
- **Hostname-targeted caching** with wildcards, and a storage gate that **never
caches HTML or dynamic content**.
## Deploy
**Via the launcher** (from a clone or the one-liner): pick `deploy: squid`, then
answer the prompts (trusted CIDR, etc.).
**Standalone** (self-contained `deploy.sh`, scp'd to a host):
```bash
TRUSTED_CIDR=100.64.0.0/10 BIND_ADDR=100.64.0.1 SKIP_PROMPTS=1 bash deploy.sh
```
**Fresh VM**: paste [`cloud-init.yml`](cloud-init.yml) as user-data (it hardens
SSH first, then deploys).
The deploy is idempotent. On first run it builds the local image, generates the
CA into `ssl/` (never overwritten), seeds `.env`, registers the port with the
host firewall if present, and brings the stack up.
On a **re-run**, an existing `.env` is kept — except for values you pass
explicitly that run (`TRUSTED_CIDR`, `BIND_ADDR`, `PROXY_PORT`, `CACHE_SIZE_MB`,
`CACHE_ONLY_LISTED`), which are written through and logged. That matters because
Compose reads the shell environment *before* `.env`: without the write-through, a
re-run that narrowed `BIND_ADDR` would apply only to that run, and the next plain
`docker compose up -d` — which has no such environment — would fall back to the
old `.env` and republish the intercepting proxy on `0.0.0.0`. Values you do not
pass are left untouched, so hand-edits to `.env` survive.
## Point clients at it
```bash
export http_proxy=http://<host>:3128
export https_proxy=http://<host>:3128
```
Per-tool: apt → `Acquire::http(s)::Proxy "http://<host>:3128";`; dnf →
`proxy=http://<host>:3128` in `/etc/dnf/dnf.conf`; apk → `http_proxy` env or
`--proxy`.
**Trust the CA** (so bumped HTTPS validates) — distribute `ssl/squid-ca-cert.pem`:
| Client | Install |
|---|---|
| Debian/Ubuntu | `cp squid-ca-cert.pem /usr/local/share/ca-certificates/squid-ca.crt && update-ca-certificates` |
| Alpine | `cp squid-ca-cert.pem /usr/local/share/ca-certificates/squid-ca.crt && update-ca-certificates` |
| Alma/RHEL | `cp squid-ca-cert.pem /etc/pki/ca-trust/source/anchors/squid-ca.pem && update-ca-trust` |
Browsers, Java, and some language runtimes keep their own trust stores — import
there too if needed.
## Caching model
Two knobs and three lists, all in the stack dir; they are bind-mounted, so edit
then `docker compose restart`.
- **`cache-domains.txt`** — hostnames to cache hard (long TTL, force-cache past
`Cache-Control: private/no-store`). A **leading dot** is a subdomain wildcard:
`.ubuntu.com` matches `ubuntu.com` and every subdomain.
- **`cache-domains.regex`** — optional `dstdom_regex` patterns for wildcards
*inside* a label (e.g. `^mirror[0-9]+\.example\.com$`). Comments-only = disabled.
- **`CACHE_ONLY_LISTED`** (`.env`):
- `0` (default, *boost*): cache everything per normal HTTP rules, and
force-cache the listed domains aggressively.
- `1` (*strict allowlist*): store **only** the listed domains; pass the rest
through uncached.
**Never cached** (storage gate, applies even to boosted domains): HTML (by
`.html` extension and `text/html` content-type) and dynamic content (script
endpoints + query strings). Query strings are **exempt on boosted domains**, so
versioned static assets like `app.js?v=123` still cache there.
### splice ⇄ cache are mutually exclusive
A spliced domain is passed through encrypted — there is nothing to cache or
inspect. **Do not list the same domain in both** `splice-domains.txt` and
`cache-domains.txt`; splice wins.
## Security posture
- **Access control is Squid's `http_access`** (`TRUSTED_CIDR`), deny-by-default.
This matters because a **published Docker port bypasses the host `INPUT`
firewall** — so also pin **`BIND_ADDR`** to a trusted interface (e.g. your
Tailscale IP). The `/etc/firewall/ports.d/squid.rule` entry is belt-and-braces.
- **CA key** lives at `ssl/squid-ca-key.pem`, mode `0600` root, mounted
read-only; the container stages a squid-readable copy into `/run` (tmpfs) at
start. The key never enters the embedded archive and is git-ignored.
- **Upstream certs are validated** (`sslproxy_cert_error deny all`) — the proxy
won't silently launder a broken origin certificate to clients.
## Caching caveats (be realistic)
Even with bump, much of the web is `Cache-Control: private/no-store`, dynamic,
or personalized. The real wins are **distro packages** (apk/apt/dnf),
**container layers**, **OS updates**, and large static assets across many hosts
— which is what the default `cache-domains.txt` targets. It is not a blanket
"cache the whole internet."
## Files
| File | Purpose |
|---|---|
| `Dockerfile` | Minimal Alpine image (`apk add squid` — ssl-bump compiled in). |
| `entrypoint.sh` | Renders `squid.conf`, generates the cache policy from the lists, stages the CA, inits the cert DB/cache, starts squid. |
| `squid.conf.tmpl` | Static config (ports, bump policy, access control, cache sizing). |
| `cache-domains.txt` / `.regex` | Wildcard hostnames to cache hard. |
| `splice-domains.txt` | Domains to pass through without decrypting. |
| `docker-compose.yml` / `.env.example` | Stack definition + tunables. |
| `deploy.sh` / `build.sh` | Self-contained installer + archive embedder. |
| `cloud-init.yml` | Fresh-VM bootstrap (harden SSH, then deploy). |
## Future work
- **Transparent / intercepting mode** — add `intercept`/`tproxy` ports and an
iptables `REDIRECT` recipe for when the box is the gateway (clients need no
proxy config). Out of scope for v1.
- Upstream `cache_peer` chaining; access-log shipping / dashboard.
+50
View File
@@ -0,0 +1,50 @@
#!/usr/bin/env bash
#
# build.sh -- (re)embed the loose deployment files into deploy.sh as a base64
# tar.gz payload after __ARCHIVE_BELOW__. Idempotent: strips any existing
# payload first.
#
# Run this after editing ANY embedded file below, then re-stage deploy.sh --
# the deployed stack uses the EMBEDDED copies, not the loose files.
set -euo pipefail
DIR=$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)
SCRIPT="$DIR/deploy.sh"
MARKER="__ARCHIVE_BELOW__"
# The CA (ssl/), README, cloud-init, and the build/deploy scripts are NOT
# embedded -- the CA is generated on the host at deploy time.
FILES=(
docker-compose.yml
Dockerfile
entrypoint.sh
squid.conf.tmpl
splice-domains.txt
cache-domains.txt
cache-domains.regex
.env.example
)
[[ -f "$SCRIPT" ]] || { echo "deploy.sh not found at $SCRIPT" >&2; exit 1; }
for f in "${FILES[@]}"; do
[[ -f "$DIR/$f" ]] || { echo "Missing $DIR/$f" >&2; exit 1; }
done
PAYLOAD=$(tar -czf - -C "$DIR" "${FILES[@]}" | base64)
TMP=$(mktemp)
trap 'rm -f "$TMP"' EXIT
sed "/^${MARKER}\$/,\$d" "$SCRIPT" > "$TMP"
{
echo "$MARKER"
echo "$PAYLOAD"
} >> "$TMP"
mv "$TMP" "$SCRIPT"
chmod +x "$SCRIPT"
trap - EXIT
size=$(wc -c < "$SCRIPT")
echo "Built $SCRIPT (${size} bytes)"
+14
View File
@@ -0,0 +1,14 @@
# cache-domains.regex
#
# OPTIONAL. dstdom_regex patterns for glob-style hostnames that the leading-dot
# matching in cache-domains.txt can't express (e.g. a wildcard INSIDE a label).
# One POSIX extended regex per line, matched case-insensitively against the
# request hostname. Full-line "#" comments only.
#
# Leave this file with comments only to disable it -- the entrypoint skips it
# when there are no active entries (an empty ACL would otherwise error).
#
# Examples (uncomment / adapt):
# ^mirror[0-9]+\.example\.com$
# ^cdn-.*\.fastly\.net$
# ^.*\.pkg\.dev$
+42
View File
@@ -0,0 +1,42 @@
# cache-domains.txt
#
# Hostnames to cache aggressively (long TTL + force-cache, overriding
# Cache-Control: private/no-store). One entry per line; a LEADING DOT matches
# the domain and all subdomains (.ubuntu.com matches archive.ubuntu.com,
# security.ubuntu.com, ...). For wildcards INSIDE a label, use
# cache-domains.regex instead.
#
# These are BUMPED (decrypted) so HTTPS bodies can be cached -- do not also
# list them in splice-domains.txt. HTML and dynamic content are still skipped
# even here (see the storage gate); only static objects are stored.
#
# Defaults target high-bandwidth distro / package / container mirrors for a
# fleet of VMs. Trim or extend for your environment.
# ── Debian / Ubuntu ──
.ubuntu.com
.debian.org
.launchpad.net
# ── Alpine ──
.alpinelinux.org
# ── RHEL family (Alma / Rocky / Fedora / EPEL) ──
.almalinux.org
.rockylinux.org
.fedoraproject.org
.centos.org
# ── Container registries (layer blobs) ──
.docker.io
.docker.com
.ghcr.io
.quay.io
.gcr.io
.k8s.io
# ── Language package registries ──
.pypi.org
.pythonhosted.org
.npmjs.org
.crates.io
+41
View File
@@ -0,0 +1,41 @@
#cloud-config
#
# Squid SSL-bump caching proxy — harden SSH, then deploy, on a fresh host.
#
# Fill in REPO_URL and the values in the runcmd block, then paste this as the
# instance user-data. Unlike the web stacks this is a FORWARD proxy: no public
# DNS record or Let's Encrypt cert is needed, but clients must be able to reach
# TRUSTED_CIDR and must trust the CA this generates on first boot.
#
# Only deploy this on networks/devices you own and are authorized to inspect.
packages:
- git
runcmd:
- hostnamectl set-hostname squid || true
- |
set -e
REPO_URL=https://git.anomalous.dev/57_Wolve/automations.git
REPO_BRANCH=main
HARDEN_SSH=1 # harden SSH on this fresh VM (set 0 to skip)
SSH_PORT=22
ALLOWED_IP= # optional: whitelist your client IP in sshguard
git clone --depth 1 --branch "$REPO_BRANCH" "$REPO_URL" /opt/automations
cd /opt/automations
# Harden SSH first (PQ KEX, key-only auth, sshguard + deny-by-default
# firewall). The firewall it installs is what deploy.sh registers the proxy
# port with.
if [ "$HARDEN_SSH" = 1 ]; then
SSH_PORT="$SSH_PORT" ALLOWED_IP="$ALLOWED_IP" SKIP_PROMPTS=1 FORCE=1 \
bash scripts/harden-ssh.sh
fi
# Deploy the proxy. Set TRUSTED_CIDR to the network allowed to use it, and
# BIND_ADDR to a trusted interface (a published Docker port bypasses the
# host firewall, so this is the real exposure control).
TRUSTED_CIDR=100.64.0.0/10 \
BIND_ADDR=0.0.0.0 \
SKIP_PROMPTS=1 \
bash deployments/squid/deploy.sh
+493
View File
@@ -0,0 +1,493 @@
#!/usr/bin/env bash
#
# deploy.sh -- deploy the Squid SSL-bump caching forward proxy.
#
# What this does:
# 1. Installs docker + compose if missing.
# 2. Lays down the stack files in $STACK_DIR and builds the local image.
# 3. Generates the TLS interception CA on first run (never overwritten).
# 4. Generates .env on first run. An existing .env is kept, except that values
# passed explicitly to a re-run are written through (they are access
# controls -- a stale .env is what a later `docker compose up -d` reads).
# 5. Prompts for required values not preset (TRUSTED_CIDR).
# 6. Registers the proxy port with the host firewall if present.
# 7. Brings the stack up and waits for health.
#
# Role: this is a FORWARD proxy, not an inbound web service -- there is no
# public hostname or Let's Encrypt cert. Clients set HTTP(S)_PROXY to this host
# and trust the generated CA. Only run it on networks/devices you own.
#
# Idempotent: re-run to apply config changes / rebuild the image.
#
# Self-contained: the compose file, Dockerfile, entrypoint, config template and
# domain lists are embedded as a base64 tar.gz at the bottom of this file.
# Rebuild with build.sh after editing the loose source files.
#
# Usage:
# bash deploy.sh # interactive prompts
# TRUSTED_CIDR=10.0.0.0/8 BIND_ADDR=100.64.0.1 \
# SKIP_PROMPTS=1 bash deploy.sh # non-interactive
# STACK_DIR=/opt/squid bash deploy.sh
# SKIP_DOCKER_INSTALL=1 bash deploy.sh
set -euo pipefail
: "${STACK_DIR:=/srv/squid}"
: "${SKIP_DOCKER_INSTALL:=0}"
: "${FORCE:=0}"
: "${SKIP_PROMPTS:=0}" # non-interactive: require values via env, no prompts
[[ "$SKIP_PROMPTS" == "1" ]] && FORCE=1
# Runtime keys the operator supplies and that compose interpolates. Which of them
# actually arrived in this script's ENVIRONMENT has to be recorded BEFORE the
# ":=" defaults below, which make an unset variable look supplied -- PROXY_PORT
# especially, since its default is a non-empty 3128.
RUNTIME_KEYS=(TRUSTED_CIDR BIND_ADDR PROXY_PORT CACHE_SIZE_MB CACHE_ONLY_LISTED)
KEYS_FROM_ENV=()
for _k in "${RUNTIME_KEYS[@]}"; do
[[ -n "${!_k+x}" ]] && KEYS_FROM_ENV+=("$_k")
done
: "${TRUSTED_CIDR:=}"
: "${BIND_ADDR:=}"
: "${PROXY_PORT:=3128}"
: "${CACHE_SIZE_MB:=}"
: "${CACHE_ONLY_LISTED:=}"
: "${CA_CN:=Squid Proxy CA}"
: "${CA_O:=automations}"
: "${CA_DAYS:=3650}"
: "${SQUID_IMAGE_TAG:=automations/squid:latest}"
log() { printf '\033[1;32m[+]\033[0m %s\n' "$*"; }
warn() { printf '\033[1;33m[!]\033[0m %s\n' "$*" >&2; }
die() { printf '\033[1;31m[x]\033[0m %s\n' "$*" >&2; exit 1; }
[[ $EUID -eq 0 ]] || die "Run as root."
# ---------------------------------------------------------------------------
# OS detection + Docker install (Alpine / Debian / Alma). This deploy.sh is
# self-contained (scp'd standalone), so the OS logic is inlined here instead
# of sourced from scripts/oslib.sh.
# ---------------------------------------------------------------------------
osfam() {
local id="" like=""
if [[ -r /etc/os-release ]]; then
id="$(. /etc/os-release 2>/dev/null && echo "${ID:-}")"
like="$(. /etc/os-release 2>/dev/null && echo "${ID_LIKE:-}")"
fi
case " $id $like " in
*" alpine "*) echo alpine ;;
*" debian "*|*" ubuntu "*) echo debian ;;
*" rhel "*|*" fedora "*|*" centos "*) echo rhel ;;
*) echo "${id:-unknown}" ;;
esac
}
install_docker() {
if command -v docker >/dev/null 2>&1; then
log "Docker already installed: $(docker --version)"
else
log "Installing Docker (OS: $(osfam))..."
case "$(osfam)" in
alpine) apk add -q docker docker-cli-compose openrc ;;
debian|rhel) command -v curl >/dev/null 2>&1 || \
{ command -v apt-get >/dev/null 2>&1 && apt-get install -y -qq curl; } || \
{ command -v dnf >/dev/null 2>&1 && dnf install -y -q curl; }
curl -fsSL https://get.docker.com | sh ;;
*) die "Unsupported OS for auto Docker install. Set SKIP_DOCKER_INSTALL=1 and install Docker yourself." ;;
esac
fi
if command -v rc-update >/dev/null 2>&1; then
rc-update add docker default >/dev/null 2>&1 || true
rc-service docker status >/dev/null 2>&1 || rc-service docker start
elif command -v systemctl >/dev/null 2>&1; then
systemctl enable --now docker >/dev/null 2>&1 || systemctl start docker || true
fi
}
open_proxy_port() {
# Register the proxy port. Prefer the host firewall (harden-firewall.sh)
# when present; else ufw/firewalld if active.
#
# NOTE: a Docker-published port reaches the host through nat/FORWARD and
# BYPASSES the INPUT firewall, so this is belt-and-braces -- the real access
# gate is Squid's http_access (TRUSTED_CIDR) plus pinning BIND_ADDR to a
# trusted interface. The ports.d format is port-only (no source CIDR), so
# restrict the source via BIND_ADDR / upstream firewalling, not here.
local port="${PROXY_PORT:-3128}"
if [[ -d /etc/firewall/ports.d && -x /usr/local/sbin/firewall-apply ]]; then
log "Registering ${port}/tcp with host firewall..."
printf '%s/tcp\n' "$port" > /etc/firewall/ports.d/squid.rule
/usr/local/sbin/firewall-apply
elif command -v ufw >/dev/null 2>&1 && ufw status 2>/dev/null | grep -q '^Status: active'; then
log "ufw active -- allowing ${port}/tcp..."
ufw allow "${port}/tcp" >/dev/null
elif command -v firewall-cmd >/dev/null 2>&1 && firewall-cmd --state >/dev/null 2>&1; then
log "firewalld active -- allowing ${port}/tcp..."
firewall-cmd -q --add-port="${port}/tcp" --permanent
firewall-cmd -q --reload
fi
}
# ----------------------------------------------------------------------------
# Extract embedded archive
# ----------------------------------------------------------------------------
SCRIPT_DIR=$(mktemp -d -t squid-deploy.XXXXXX)
trap 'rm -rf "$SCRIPT_DIR"' EXIT
extract_archive() {
grep -a -A 9999999 '^__ARCHIVE_BELOW__$' "$0" \
| tail -n +2 \
| base64 -d \
| tar -xz -C "$SCRIPT_DIR"
}
if grep -q -a '^__ARCHIVE_BELOW__$' "$0"; then
log "Extracting embedded deployment files..."
extract_archive
else
die "No embedded archive found. Run build.sh to embed deployment files."
fi
EMBEDDED=(docker-compose.yml Dockerfile entrypoint.sh squid.conf.tmpl
splice-domains.txt cache-domains.txt cache-domains.regex .env.example)
for f in "${EMBEDDED[@]}"; do
[[ -f "$SCRIPT_DIR/$f" ]] || die "Embedded archive missing $f"
done
# ----------------------------------------------------------------------------
# Prompt for required vars
# ----------------------------------------------------------------------------
prompt() {
local varname="$1" message="$2"
local -n ref="$varname"
if [[ -z "${ref:-}" ]]; then
[[ "$SKIP_PROMPTS" == "1" ]] && die "$varname required (set it in the environment; running with SKIP_PROMPTS=1)."
read -r -p "$message: " ref
[[ -n "$ref" ]] || die "$varname required."
fi
}
prompt TRUSTED_CIDR "Trusted client CIDR(s) allowed to use the proxy (e.g. 100.64.0.0/10)"
# ----------------------------------------------------------------------------
# Docker + firewall
# ----------------------------------------------------------------------------
if [[ "$SKIP_DOCKER_INSTALL" != "1" ]]; then
install_docker
fi
open_proxy_port
# ----------------------------------------------------------------------------
# Stack directory + files
# ----------------------------------------------------------------------------
log "Setting up $STACK_DIR..."
install -d -m 0750 "$STACK_DIR"
install -m 0644 "$SCRIPT_DIR/docker-compose.yml" "$STACK_DIR/docker-compose.yml"
install -m 0644 "$SCRIPT_DIR/Dockerfile" "$STACK_DIR/Dockerfile"
install -m 0755 "$SCRIPT_DIR/entrypoint.sh" "$STACK_DIR/entrypoint.sh"
install -m 0644 "$SCRIPT_DIR/squid.conf.tmpl" "$STACK_DIR/squid.conf.tmpl"
install -m 0644 "$SCRIPT_DIR/splice-domains.txt" "$STACK_DIR/splice-domains.txt"
install -m 0644 "$SCRIPT_DIR/cache-domains.txt" "$STACK_DIR/cache-domains.txt"
install -m 0644 "$SCRIPT_DIR/cache-domains.regex" "$STACK_DIR/cache-domains.regex"
ENV_FILE="$STACK_DIR/.env"
set_env() { # <KEY> <value>: update KEY in .env, or append if absent
# The value goes through the ENVIRONMENT, never interpolated into a sed
# script. Interpolating it corrupts any value containing & (sed expands it to
# the whole match) and aborts the run on one containing the s||| delimiter --
# which is reachable for an OIDC secret, a password or a URL query string.
local key="$1" val="$2" tmp
if [[ ! -f "$ENV_FILE" ]]; then
printf '%s=%s
' "$key" "$val" >> "$ENV_FILE"
return 0
fi
tmp="$(mktemp)"
_SE_KEY="$key" _SE_VAL="$val" awk '
BEGIN { k = ENVIRON["_SE_KEY"]; v = ENVIRON["_SE_VAL"]; seen = 0 }
!seen && index($0, k "=") == 1 { print k "=" v; seen = 1; next }
{ print }
END { if (!seen) print k "=" v }
' "$ENV_FILE" > "$tmp"
cat "$tmp" > "$ENV_FILE" # rewrite in place: keeps the original mode/owner
rm -f "$tmp"
}
if [[ ! -f "$ENV_FILE" ]]; then
log "Seeding $ENV_FILE..."
install -m 0600 "$SCRIPT_DIR/.env.example" "$ENV_FILE"
set_env TRUSTED_CIDR "$TRUSTED_CIDR"
if [[ -n "${BIND_ADDR:-}" ]]; then set_env BIND_ADDR "$BIND_ADDR"; fi
if [[ -n "${PROXY_PORT:-}" ]]; then set_env PROXY_PORT "$PROXY_PORT"; fi
if [[ -n "${CACHE_SIZE_MB:-}" ]]; then set_env CACHE_SIZE_MB "$CACHE_SIZE_MB"; fi
if [[ -n "${CACHE_ONLY_LISTED:-}" ]]; then set_env CACHE_ONLY_LISTED "$CACHE_ONLY_LISTED"; fi
set_env CA_CN "$CA_CN"
set_env CA_O "$CA_O"
set_env SQUID_IMAGE_TAG "$SQUID_IMAGE_TAG"
else
# Keep the file, but do NOT let it drift from what the operator passed this
# run. compose reads the shell environment before .env, so an exported value
# silently wins now while .env still holds the old one -- and a later plain
# `docker compose up -d` (the documented update path) has no such environment
# and falls back to .env. That is how a deliberately narrowed BIND_ADDR
# reverts to 0.0.0.0 on the next image update, republishing an intercepting
# proxy on every interface. These keys are access controls, so write them
# through and say so; anything not passed this run is left untouched.
log ".env exists; keeping it (only values passed this run are updated)."
for _k in ${KEYS_FROM_ENV[@]+"${KEYS_FROM_ENV[@]}"}; do
_cur=$(sed -n "s/^${_k}=//p" "$ENV_FILE" | tail -n1)
_new="${!_k}"
if [[ "$_cur" != "$_new" ]]; then
set_env "$_k" "$_new"
log " ${_k}: ${_cur:-<unset>} -> ${_new:-<empty>}"
fi
done
fi
# Validate required value landed.
grep -E "^TRUSTED_CIDR=.+$" "$ENV_FILE" >/dev/null || die "TRUSTED_CIDR missing in $ENV_FILE."
# Use the image tag actually recorded in .env (keeps build + CA-gen in sync).
IMAGE=$(grep -E '^SQUID_IMAGE_TAG=' "$ENV_FILE" | head -n1 | cut -d= -f2-)
IMAGE=${IMAGE:-$SQUID_IMAGE_TAG}
# ----------------------------------------------------------------------------
# Confirm (interception is a big deal)
# ----------------------------------------------------------------------------
if [[ "$FORCE" != "1" ]]; then
cat <<EOF
About to build the Squid image and start an SSL-bump forward proxy from
$STACK_DIR.
This generates a local CA and INTERCEPTS TLS for every client that trusts it
(except spliced domains). Only do this on networks/devices you own and are
authorized to inspect. Clients reach it at: http://<this-host>:${PROXY_PORT}
Continue? [y/N]
EOF
read -r ans
[[ "${ans,,}" == "y" || "${ans,,}" == "yes" ]] || { warn "Aborted."; exit 0; }
fi
cd "$STACK_DIR"
# ----------------------------------------------------------------------------
# Build image, then generate the CA (needs the built image's openssl)
# ----------------------------------------------------------------------------
log "Building image $IMAGE..."
docker compose build
SSL_DIR="$STACK_DIR/ssl"
install -d -m 0750 "$SSL_DIR"
if [[ ! -f "$SSL_DIR/squid-ca-key.pem" || ! -f "$SSL_DIR/squid-ca-cert.pem" ]]; then
log "Generating TLS interception CA (one-time; CN=${CA_CN})..."
docker run --rm -v "$SSL_DIR:/out" --entrypoint openssl "$IMAGE" \
req -x509 -newkey rsa:4096 -sha256 -days "$CA_DAYS" -nodes \
-keyout /out/squid-ca-key.pem -out /out/squid-ca-cert.pem \
-subj "/CN=${CA_CN}/O=${CA_O}" \
-addext "basicConstraints=critical,CA:TRUE" \
-addext "keyUsage=critical,keyCertSign,cRLSign"
chmod 0600 "$SSL_DIR/squid-ca-key.pem"
chmod 0644 "$SSL_DIR/squid-ca-cert.pem"
log "CA created. Distribute $SSL_DIR/squid-ca-cert.pem to client trust stores."
else
log "CA already present in $SSL_DIR; leaving it."
fi
# ----------------------------------------------------------------------------
# Bring up the stack
# ----------------------------------------------------------------------------
log "Starting stack..."
docker compose up -d --remove-orphans
# ----------------------------------------------------------------------------
# Wait for health
# ----------------------------------------------------------------------------
log "Waiting for squid to become healthy (up to 120s)..."
deadline=$(( $(date +%s) + 120 ))
while (( $(date +%s) < deadline )); do
status=$(docker compose ps --format '{{.Service}} {{.Health}}' 2>/dev/null || true)
unhealthy=$(echo "$status" | awk '$2 != "healthy" && $2 != "" {print $1}')
if [[ -z "$unhealthy" && -n "$status" ]]; then
log "Squid healthy."
break
fi
sleep 5
done
echo
log "Stack status:"
docker compose ps
echo
CERT="$SSL_DIR/squid-ca-cert.pem"
cat <<EOF
================================================================
DEPLOYED -- Squid SSL-bump caching forward proxy
Proxy: http://<this-host>:${PROXY_PORT}
Stack dir: ${STACK_DIR}
CA cert: ${CERT}
1. Point clients at the proxy:
export http_proxy=http://<this-host>:${PROXY_PORT}
export https_proxy=http://<this-host>:${PROXY_PORT}
(apk: /etc/apk/repositories via http_proxy; apt: Acquire::http(s)::Proxy;
dnf: proxy= in /etc/dnf/dnf.conf)
2. Trust the CA on each client (so bumped HTTPS validates):
Debian/Ubuntu: cp ${CERT##*/} /usr/local/share/ca-certificates/squid-ca.crt && update-ca-certificates
Alpine: cp ${CERT##*/} /usr/local/share/ca-certificates/squid-ca.crt && update-ca-certificates
Alma/RHEL: cp ${CERT##*/} /etc/pki/ca-trust/source/anchors/squid-ca.pem && update-ca-trust
Browsers/Java keep their own trust stores -- import there too if needed.
3. Tune behavior, then re-render (lists are bind-mounted):
${STACK_DIR}/cache-domains.txt domains to cache hard (wildcards)
${STACK_DIR}/cache-domains.regex glob/regex cache patterns (optional)
${STACK_DIR}/splice-domains.txt domains to NOT decrypt
edit, then: cd ${STACK_DIR} && docker compose restart
Logs / cache stats:
docker compose logs -f
docker exec squid squid -k parse # validate config
docker exec squid tail -f /var/log/squid/access.log # TCP_HIT / TCP_MISS
Manage:
docker compose pull >/dev/null; docker compose build && docker compose up -d # update
docker compose down # stop, keep cache + CA
docker compose down -v # stop, WIPE cache + leaf-cert DB (CA on host survives)
SECURITY: keep ${SSL_DIR}/squid-ca-key.pem secret (it can MITM every client
that trusts it). It is 0600 root and never leaves this host. Only intercept
traffic on networks/devices you own.
Or just re-run this script -- it's idempotent.
================================================================
EOF
# IMPORTANT: do not put any code below this exit. Everything after the
# __ARCHIVE_BELOW__ marker is the embedded tar.gz payload (base64), appended by
# build.sh.
exit 0
__ARCHIVE_BELOW__
H4sIAAAAAAAAA+xc63bbRpL2bzxFhVIiciKAF4mSlx7F0YWxtUeUNCKdsdcXGgSaJCIQQNCAJNry
OftrH2DPPuE8yVZVNwjwItmzI3tnzgmjSERfqrurq6u+qi7YDZ1LEZtOOIlCKazpxH/04J8afna2
tx/Vdmr13Wad/9ZqdS7Hz1aztvWo3mzUd7ZqW9s7zUe1emN3d+sR1B5+KsufVCZ2DPCoudv/a+hf
iW8x5j/RZw26v6eeC93uiTlIJxE4tjP2ghEMw/jajl2I4vBmahlrxhocpJ6fgB86tu9PYRiHE7Cq
RyxAQ88XUE7GAmIRhRsSwkCAFP4QiVInb2KPRAVME7BNLMCTSC8IIY1kEgt7AlL6anxuCdf4YwcJ
JCEkMW4ReAHYMMYJiUC4SDlJIwt6OF7vpIukvCARsSOixAsDONwH37sSNAkaDsahImBVcRQoj5BE
bCdIZjAFV0R+OLXkeBMmYYpUXKSGM3LNMPCnlSdMwBf20HREnMDRwSZzSGyCHbjIi5EEG9cT2BOk
dxX66URIxa19iNKB78kxVigmQRQiiYNX5/vdbrubT+349PxFD4ZeLK6Rs5sgQzg4Pj3q7x8dXYAc
h6lPk4o8tRokmdAKiDe24g6OwAwY2o7gefUuXnR77aP+4TFSKKsdHidJ1LcdR0hZIY5JUNtl+6BK
YYRMsQyD1tICSZ0MQy+pZYAqMXn1+SPtmzvIn4kjLcOQIr7ynEI/+gJAwqC/AjghzvkmaYHFBbzx
LVj/2P3Li+Oj/nFn/1m739t/1jLtNAknNm2trCpiPs5UJp+MjIztIUP6xYlTTSzobOMAaeDj+kyZ
hFEkVB1thcxmYkJp/eOM5S2zZvF/n1rrH88vzl6+6p+fXfRa5la98fhTi36XuKMIrrw4DCYiSDJK
Rca3iGqx4FNJtzrcP3ze7neP/6Pd7xxws/kSs4lqcda6s/+yf3bw7+3D3lyXFcVmo7mzMEin3Zkf
QxfMNT16dbrfOT7sH7YvesUeq8rNxwsjnJ2evOqfHNMqC8MUS818Lb8ed48PTtr952fdHtJuc4+l
QlPJEqse3bUgh2rH+DC3qiJxlEzwYxwW6yMfRdB0UXa8QFoJilqx+XLtXG+W87s6L1fe0zcWI3Fz
d29VXehfPGfVKztW7au5XOeN9OnjVr43yDnRdwcLTflgqobhqEBsjAogGeMAzmXG3DV4z/VgXgJX
vAfpjQLb1yojDQKyEBMbNU9MBoK1x5Vng5dI1FMusDkwzRm9WPiePfCVcrL9a3uK7fB44smBMpoB
1MIiQNXuTEln8+CO72FtNXAqliZDR74Fr0uHnSOz+7x9clLahNL8RMEcQnGP6beFCmIIt7cgbrwE
6qW3mhxrzCvbb8FWTWZDeBMRpjhKMyuJRRJ7KHewpQtYpfQjEXuh24IGdv3/NuH/0Cc3319vjM/g
v/rOdn2G/xq1BuK/+lZ95w/89y0+a9DxAjS9vsaBGn55yXgGCTWc8RGAIKYbe5Hkc650xHuIbOeS
+iig9940qbMZ4oFGPfSe0UwG7jaRDsEQ2WckRMpACieNvWTaJxHsE8ZCeMaQCqEQkFuC5QRvCDsG
IRLQZSbhT1OGKcI+CIRwsVX5SAw8O6i+GCCSSzcJ3xE6iFFRbSrsoXQL2CNSvQlSexakiCB5KkSE
FZwUkU0AUS9xtpZsqRULdBGhKC9w/JRGlzS7GZ5EmBjAxCMMiyQRkpJ+TMbUYcZiVDWEagkDZgRz
TWgZv1ycdZANxPfWltWoG8bFi1Owo0uwXdR5ZhAqK6EX5diMUb2h5xA8mpFMcIPhDauvH35AxO1i
rbnQ2DAOz85fQa4wrWQS+at1KVep9qih42kU4ippxQDVVJJ9QQehOvCC6lw1T94ZT0IXfry5vyXy
sX1D9hkVtvZDFBhg2GbBofJRqriHkXAY849thHb4B52LJ6BshyQ3AUk97/XOy91Kn6EcexS0C8R1
xnKW0X55ftZtAz3Q0MwvxMUo5x/CyQC1P4uHnoic2UIU7PeuAvaELNFIjlMaVG0G+gzoQCDx097F
q/Oz49Memq6qpMXSAGS7TJN+38cJNFZf+fzPDfeVxiAlv9ts3qH/69uN3P/f3m6y/1+v7f6h/7/F
Z+07ljo8dKTk588zatyZc6Vwj5lGdA5Ypd0RNyjECwDqFlyQOovvQGUcQ1hUOuWff92/+BnMn8jB
QvRHhBoWPNNuOw+uIHQUoo6YZhpYUVuC5lAFhbH5FGNnJrjKfUHdMBpRHGMQkkrePzyRm3S0Y9LX
5BxvAuI+TRt6vROpZ7dlQTehRlrVI//IN1fWg4IJjH1xcUOJHI2mUGadfymmxFbEwngiahCHYaIJ
bltwHHjKdLhT9Gs9J4s/wI9qiX3Xo02JUydJ0VqWPVdMojDBHdQ0mhb8avueO2MZMtgbbdL3ALGw
cLSm0mEF3FYxisM0cNXmXaQB6j3Js0KMjrYR5z4OrwPSugka58oTTYAJunGI+hKXrSYnhkNSzFei
n0oRWxQQSMAUqWH0Oucne/faldNfVtcbz9qn/aPji2It93OpZq+0ruurReHgniUDxbTfvTjcm/cX
uRjN0l4VfRrtE1HR0cHeapdKiczRWWf/+LS7d59DqJtetJ+1X97dkOWSbE4vDUhGJJ05dy4ki6BD
SsGel9QibuG5eIJIYWinPhqcSyEi3kIGFhS+QgeNJQ7tIiE13DQ+tSiRk/BSsCNlGUvhidZevbHL
kY96dauBfveKyMSejkzcFYHYU2GF5XhDsWJVWGHv8Xy3YvxgbzbiUqRgbz5SYKBvW67AR9RCeAaH
sPE612hv4Xv5JthAMn8qwU8/NJ7AJ4PDdBRQUx5raYwCTz3Q5pfAGyL3ECf4KPk/7dURfQbmwLeD
y03+ivtDkR/sHgjLwJ593ZMn8Jpc0dJ6vQRvCXWNYtwj8/c2bLx7/bolEUiK1tu3f3r9bi1/2uD2
elrmw32QWq6HGd8mduI5WiEokaLiBBUIBdYyBYwPjqDli1gpuoeck0TArDCpiXyXtz8XJfHn24XA
2e2oNN96Tiqp+VzBcvtlUaVOy6V3jaSkNB9IPS+3XiHa1GlF8XLfReGmjotlhV6ldVKlKMr4jZRm
6eHF5sutLhsqZRfpOEm0UsqYKluE5jG3oqSfJIW78diVr8eCbgTIcpClFWhVCjY2FsNYyHE/spNE
xIFUxCgyD0ka0NhYG1D0KRzC9dhOQKbxFZod5cipgPbDcmVySYbXjGBmb0rGRxUORvMIJdQoL3pn
Jta1L/YJVKATOg+rcJYLsEppfdQqCQjXS6zSHL0u+7itlchmhTlB1i+p0AJFw1AhuYMZxmFvm3kq
Jl5ClwnXZNGRgbjQAbq8fB1E+hHdJtJ8qCeS6SYIXwpNTAGBa7qpABHHaGvQ9KD/yy1pEBXDEzfO
XklNBZVrQWWSCBdNa4nvXQIdctOs2LAdHxQ0c2WiJaR0nxEubeQUaOh1/A3fMQk1i6F332TYeH9m
Kn3knZpNX3H/7glx/T1TQgKzWambDZSUP/95A9nX3TBoRJrk76mIp5DGPp6JsR7zzVOuRqTYFygW
sFBdrjojzySYf/vGKuP322gc3doyunl6+5uMKusVRT2Z+Kr/AnWLap6W3zy9LbZ0Eg6PRv2JNxH9
ZBoJMD14R1c6VWpA6qh47FsQiCvlKyPPnvc6Jwi0UeZwQJQXvgwKEpPoVKhAw96swjJUqMMVwXQ2
02KZXrxB3NJS+RdmFWJkFGAl5RnR9svD9nmPhJQ5jyKvNwnKOEXphXzJyGZS0yIQhmrN9xA/2VFk
/SafXiFe2tL3c9jYo3gVz4evOHXYWh3i4tyzPVz/iFv/qXDUN5gzyE5UZ/kynWTDyIT09Uw0C2e7
BHtQIpyxQlDnpRZ1yVLvvXpL7wgV5heMM45YpRmRLzqzei1qxegmhNeKxxtPstP2BScO7iWFcjlP
rdiOOYeNN+aO0t+vbpZYxwqzBX6I2hB9P1SziKkdoUNwhK8Ffsu4d+35rkP354qNBS4qKPhZJDg/
QbhFpUwqmLxJMGMiC3264ERHIJ+rktWSJW5sxHICnZ9JCa1Eae4ZHeus4LX1lovK+AUmdoIrwYOy
QM730PbaqF7QPFH8DV2LKVks+yr0KBxGF+loJaRvy7EppGNHZDBQAcaXcnZ5oz7u3no5g+bfS1qm
S8PfMpENWX33xqpW0a+s4l+cUnW0UZnrn/VdwAWkeL6XUN/erkG9Vvv+e9jeaqBHHeJhjtExNsVN
5FHWwShASUdvwbsiHKIfg9BUJ2DWHFeSTEKXvIU3c+PTp/SOQ9hPW9Vq+fW76tsfcaKVp+sf3U/V
fJNd1CCZAH5iiIaA4CsgtL838qByJmbRB68QelDYqqyzIOZyIDgGgN3p6HiclEEOM/mVKolCmX8Z
Ku8C2SpiRY1i4Cyx2CsLePt4jgRfIM4SIAhcPDROY3X5nfLCtPdfzW5WOZZiRWJC3tnt7T0NkUm6
XUFFoJMJpV/2e/snLeI5KWxaDqKRjIAKmMDSeKg0FklXMmzGt5MGyouOr4CJB30Ctd16DcxQ89gc
6S9qrhcvUKhm7bHx9vb2nY3v4MCMUrVYPE+1dvcUVrNrnmpW/PDy/48Hyh5yPpnMSb38o4Oqh173
DWPRZRGimXu2730g6aFbqGzmiPt0f8uyMkwYuDOqqMMnRDhKxlBHORG+QHXW+KnqiqtqkOKuoVDj
ipUKUhcMeShr5Y2b6RSnjQN07ojUfOoclEhKORYI5oVO7ymKBPVfTFyA+dSDlZOd8c9d6l5FCbyf
fyv2esY8LbjDzFMG8wPuex7xpBmsapNNjFDEQ0vuYng25vgMKo0fIU9Qm4vZjil+HQ6H7CtzVPSh
YzJZLkVkx1IUeGEwu7vkqhKrVbvylSc9tC59UuGUerW3IlxR4T0oRJvN0yKPS+s/l/618yf+1T8L
4fevMsZn8j92dpq7ef5Hc4vu/7bqzT/u/77FZ23x+k1d/mTKiKDeypu7L4hqtRgDmrAcTWZfXKYD
dJqTlPTcLIZYSKmEMt1zVJ5oKrP4I+j445fd0i1GECv69g+Rb65mP3dxyKiab1h0zHMg0BnV92Sh
L1rQfnl+cnx43JtPm17Io4HDPC9hdVYC5cfM8hKg3IvtAH3CmBLh8gxnSr0LXY7GDVPGNNdhfEke
mRQCLtr7R522VTGMRf0MyxFmMmt/+5//xB840WnFLUSiOvWC17CZX/EKvlVydY+v8UPZF75kBFel
L+SjkKxwWIJSJAhvp1IJJifX0FazD0E52gygZmnfQ0q/WGOHR/s6klwlV20nezAoxUrqMkmeOTFs
rOjusujDIDXlhVOTMDBnfhRHRxeuci2DM64504n2M09xJ58yW2bh6rGIv2dtkPpCEw2muUUmwjyT
uYSePWQCtdCgmHFefyImfYWTpPdB7K28regcGAbO1IkTt48CMIrtyRciSAkrL00JSt41UjaQg1x1
UczhsVIdabRXB8/1xV6dRHQGlF6cd3so4B2YS3SizDAVWsT5BIk/hc5xr4Nu8CAOL5FoGHsjVAXM
WxqS5VpxRAWrs5hR4TjMhF5pmxYKmrgkJdc9PYa//dd/g8ofVmEpvkGkQnVMcCrThJMhKESuKXLg
FI9YVEcqffoCXekfYIe6qmJ6fYrrUCZzizLodU77XEx5OW+5RGvqq8nSJHmQvExPNKefV/Gv+XXv
q1cBOG0u9FvMGnMwNfVt8+bshYMsx4ovDv7vB56WTj4D5+OrFyS2t7dUsT0UxfLHNaqCRh12a/ib
Qj2NprnTbG41oUGVjx9D89+wdneX+yPEPG0f9mAiknHoZo9G4T0ItfHf5QMtV2ZEvptNkph1yIZo
YgeoT+KWUk0qjY+4UWYNpXVLbhQLmdYVa24kFdvMqWjCy7PJKnh9eif6s2y32IH5a9T7BllRt0Bw
efh5UVFMkMr9+no24StaG6ULUSvC/DUvoGqa2DfeJJ30w8FvwklYYfa9gBqH8RS97tVNYNVNM7XN
3dN0KJe944ULbajvQKO5Yzjoorp4SLnj0rsAi1uRAaMZsGnBl2cyLe6gkWGdpWyb5fwamsmBLUWG
tiBDW1AeosxwdHjuXpenJfktrSz/KBsOnVykRkZfxaNddVNcseBcZd8ScEs4DL+p/GOXwCKt4cmS
6mVEhY1ImdqUjsQ5YLG40gZFWsaKUPIbq+yKwW0cTW7t6PI2uhzdolHC0tGH25sPtx9kcjv40Kjc
Xo/929/suLJOQefHKvSsIs9LVC2AGkAD67lBEXeFo9G/6gFaFBmlK/oULZCJ64UL75tUVb2FBfo8
UMuFNlzBTfB/BhhxmJDt3y0w7Tz2rrAKhZnyzhCo+95QOFPHz6ztP8OPsSoVTr+eFnkuoyc27zm+
s7DcoDRiN7wO+rQoeiMFmpQmHwbuV3jXZBlOPPgQn33/Y7eR+//bdWxX38IOf/j/3+KzBssSwH7t
kb4QV2m+E3oH9/SshzjAQVcpseCvRdeWXsRiMnQZneReK722akuZjOMwHY1bChAxutiQHPEehYLc
stj2RmP1ti95VozXN7UvhmocaSElGp6PlMtanB71CwDCteCFVJewBOvNyAvoasqOIrR7AzugFJZN
etQ3fVUkd9bV70KAM7axuS/VO712oC1IQvlEA3T9LiWknDxXfMtYef9ngfYeybpxKuAT9DtO0Ac/
Pn0GRzjF2a2q71PIY5ZnYOFs1I1s1oS8y+vrvHwTRrL4ZFkWWsFfUt83aSAorZVAJyEqGK6mdNyh
t1X3T3utLKfI9pGT7KPYWZIWv4S9f3pE1nc58EGhBN5OmhIW8lbQFR6nOihfS+lpSlwV6oUKlUaJ
OzAN07gYyLEKihuZXi3sw8IGQJnWwxi6sI0ZNDFyZhiW5/hh6qrvkw8qVUM94YRRe0pFWhXp7+gH
x6EMh4kqxc2OQ4pnKbhrjcJwlJGPfHuqC+zIk1yYr+JASRQuJbKniv9lfaPO3ii9nsNcQC2ezx7b
Rrav6FNSSiQWyHbQONojTVhJcPYe9YwIJblJSkAJRIJk+FUUK4xHhpUINIjorfPTF5//pc3/Cjrm
M/q/UW8W/v2HBul/xL5/xH+/yWdt+fizFnmug4Z5Tj/YoxFiWolIhnzc1Xkwm1kSBUpx5iebh1kw
QSdeVLOMC1RnX6xAKTSYJ5mq94gXNGrKr/0VVSpaDmeMEy5U0TuIWQCrWJypV1RgWe6OhOPT7vFR
Gyfk2wPhb1JA0FjkWBYiRtfZ1u9O9DKrBQcvOuftIyhrs0nKDFUpxX676AS5lFBB8cWByCxbnglK
WttYU2oVVz4hVb1sqy2VSUf8WMiYU0F2TkuTlx7/wwdr5CBp5V+mYLHKBs/dwsoTFcHQGeLKsdXh
etowvb6j7NUDPDrkjI3ReptoZt1rz0VU4OKU45C1o3o1tFqIgkw8CrlJ9UIC0hr6AimEQ/i1I9Gc
xN6Esv/+t51ra26bCsLv+hV66AwyjRWbJG4KowcTDHjGjT1JGGAwMIos28LukSrJbf3v2W/36GYr
k3aAlsLZl9gn0rl77/uFb2kJi0fFidR7Uv9S8VlyyepcLXfBzwiP3Po7FawTOiewz5rXS4pbSznD
X+ky7t7ye9WDN9+PJvaSthk/geH2pU9j38TBZk9/vw0XtJP0YTQbTWpCi56qunJTPF37vuS3kjTG
VktTQKuMs4ORr8odpAuHDcbdcUhMUcv9Nr7PqhGlgMSN4vIjb8RqHUjjqx0JN3xY6YbNZYa/NaPU
V6sdTq44wdqYpTjbJ5HMNyGFKVZrTqSUFpW8/CPTi0nZzqbuPzarM9RCLbzsbx/jMfuv96xXyf+z
M7b/Loz8/yDUKs3EtJndjafXw4nbzG8vXXvgzStiPCTK91tJ5NMKA+wmDZsEPYD6hiHIIlnnyR2b
HCQFP8sReoSKQdq0u3JJ7JZZtE1BzIlb0Bxm09vxT1pYcM4iz1ArEielBzHws7BLI4Uqi3LRXzQC
gK4HTcNXuzDLy0W8g5k1Cf3X2vLh6B+HehvPQXMiUcgZmFGusa8ascaNYClwWFEqLlJRG1Rs++wz
KqrRbKdeTaENsRgvvIlI1eAwVkcmNioMEYdkna5RI5tr4Sd5ByH530QC/9LrPv/16bxIFZ5DRjzB
v4OF6rqfz92ln+Xb/RySktvRlmxWcxKor58YZv4fIeRYFHfgnxrjEf5//ux8UPH/ix7wX4z994GI
bDQUpBOrwk0QoAuYDJFydSGyXRYiR8Emk8d2iV2AsjEYoHCed0IStH9QXETDjj4u/aMDQMcFMKB9
PYXbENh5QcmQiYNNQjgNR4oNKUmVZZYKRBUyUQ6HeRgbsJH6Q01SiSjALbDUNIBLR4PCwNkFo/ZN
iuI4VTM/oCn/uI5JzOxhGfKSJO/no8ce2gMSdESoMOkWADcLG2Fih0wHDgBLWg2ZaopMkTbYQOyq
bLeIzZvRcFIHD8SBcDU6MrLh0XVLafQlomPu4Bz15af9nu3c+dE2o8sT2lffXQ/vOrDL+xp47/Ty
xO4//8LtDy7l+YFr1QPbXqMvS7srKgzE0/EMK9H4i7VzidVJgQLBV6ubRYtQg8rMIoUwZgm8+AC8
oqgmbJRWSxjPGNvSryAf2YP+nqCPevXQgmJd8q/nncdiu6d7pCIFa9YpSDAj8Utc38W7rHhEyyYC
5NMSZxMGtx4SP01/i8wmoCThltOiXKvcAE93aVUQiF4BkVPd/087CcBqhN0Z3KCqe4pVl9S3jfZ+
3e8WcHU4L77uWMcxfsAb6GopuETofFESAggGdqFUTjQn9SMdqRjfTrNTxmzIOlY9/6DorJhIpLo6
7WCNOhzpB/NoO4kgTv5NQdD34k4929MZC5zg6Oiko46ufa6H96HkqzgFcBm8aXa6IyZzorMHuGCi
VpxXOQ6zduujzPRkLb5wbQLI1e7TnOD5CPIq1+uRksmvOOils/rZxuDwl2sdV1/2mmeIkNihuHIe
FFZHORt/5Xfw+9W1Jxx/xiznaojGqVcDX0XDN8Ofb72zwUXPaknp8y6rG3uDLD3e/XhZk7dt+ZrH
d/lFlAWf5CVu393DvNs6aglMSc5sFXy63F+VCEsCFOcIsN4R3DKZewdIud5DQLnGWjNkyJAhQ4YM
GTJkyJAhQ4YMGTJkyJAhQ4YMGTJk6H9BfwI4WgWBAHgAAA==
+51
View File
@@ -0,0 +1,51 @@
# Squid SSL-bump caching forward proxy.
#
# Built locally from ./Dockerfile (the repo's one self-built image) -- there is
# no upstream ssl-bump image we want to trust in a hardened setup. The TLS
# interception CA lives on the host in ./ssl (generated by deploy.sh, mounted
# read-only); the leaf-cert DB, cache, and logs are named volumes.
#
# A published Docker port BYPASSES the host INPUT firewall, so BIND_ADDR should
# pin the listener to a trusted interface and TRUSTED_CIDR (Squid http_access)
# is the real access gate.
name: squid
volumes:
squid-cache:
squid-ssl-db:
squid-logs:
services:
squid:
build:
context: .
image: ${SQUID_IMAGE_TAG:-automations/squid:latest}
container_name: squid
restart: unless-stopped
ports:
- "${BIND_ADDR:-0.0.0.0}:${PROXY_PORT:-3128}:3128"
environment:
TRUSTED_CIDR: "${TRUSTED_CIDR}"
CACHE_SIZE_MB: "${CACHE_SIZE_MB:-5000}"
MAX_OBJECT_SIZE_MB: "${MAX_OBJECT_SIZE_MB:-256}"
CACHE_MEM_MB: "${CACHE_MEM_MB:-256}"
DYNAMIC_CERT_MEM_MB: "${DYNAMIC_CERT_MEM_MB:-8}"
CACHE_ONLY_LISTED: "${CACHE_ONLY_LISTED:-0}"
VISIBLE_HOSTNAME: "${VISIBLE_HOSTNAME:-squid-proxy}"
volumes:
- ./ssl:/etc/squid/ssl:ro
- ./splice-domains.txt:/etc/squid/splice-domains.txt:ro
- ./cache-domains.txt:/etc/squid/cache-domains.txt:ro
- ./cache-domains.regex:/etc/squid/cache-domains.regex:ro
- squid-cache:/var/cache/squid
- squid-ssl-db:/var/lib/squid/ssl_db
- squid-logs:/var/log/squid
healthcheck:
# `squid -k check` signals the running master process via its pid file --
# reliable and always present (no dependency on squidclient/nc).
test: ["CMD-SHELL", "squid -k check -f /etc/squid/squid.conf || exit 1"]
interval: 30s
timeout: 5s
retries: 3
start_period: 20s
+145
View File
@@ -0,0 +1,145 @@
#!/bin/sh
#
# entrypoint.sh -- container start-up for the Squid SSL-bump caching proxy.
#
# 1. Render /etc/squid/squid.conf from squid.conf.tmpl (@VAR@ -> env).
# 2. Generate the cache-policy include from cache-domains.txt / .regex and the
# CACHE_ONLY_LISTED toggle (boost ACLs, storage gate, per-domain TTLs).
# 3. Stage the CA into a squid-readable tmpfs copy (host key stays 0600 root).
# 4. Init the dynamic-cert DB + cache_dir structure (idempotent).
# 5. Validate the config, then exec squid in the foreground.
#
# Runs as root (needs chown/install); squid then drops to cache_effective_user.
set -eu
TMPL=/etc/squid/squid.conf.tmpl
CONF=/etc/squid/squid.conf
GEN_DIR=/etc/squid/conf.d
GEN="$GEN_DIR/cache-policy.conf"
SSL_SRC=/etc/squid/ssl
SSL_RUN=/run/squid
SSL_DB=/var/lib/squid/ssl_db
CACHE_DOMAINS=/etc/squid/cache-domains.txt
CACHE_REGEX=/etc/squid/cache-domains.regex
# Tunables -- docker-compose passes these from .env; defaults keep the image
# runnable on its own for a smoke test.
: "${TRUSTED_CIDR:=127.0.0.1/32}"
: "${CACHE_SIZE_MB:=5000}"
: "${MAX_OBJECT_SIZE_MB:=256}"
: "${CACHE_MEM_MB:=256}"
: "${DYNAMIC_CERT_MEM_MB:=8}"
: "${CACHE_ONLY_LISTED:=0}"
: "${VISIBLE_HOSTNAME:=squid-proxy}"
log() { printf '[entrypoint] %s\n' "$*" >&2; }
# A list file "has entries" if it holds >=1 non-blank, non-comment line.
has_entries() { [ -f "$1" ] && grep -qE '^[[:space:]]*[^#[:space:]]' "$1"; }
# ---------------------------------------------------------------------------
# 1. Render the static config from the template (@VAR@ placeholders).
# ---------------------------------------------------------------------------
sed \
-e "s|@TRUSTED_CIDR@|${TRUSTED_CIDR}|g" \
-e "s|@CACHE_SIZE_MB@|${CACHE_SIZE_MB}|g" \
-e "s|@MAX_OBJECT_SIZE_MB@|${MAX_OBJECT_SIZE_MB}|g" \
-e "s|@CACHE_MEM_MB@|${CACHE_MEM_MB}|g" \
-e "s|@DYNAMIC_CERT_MEM_MB@|${DYNAMIC_CERT_MEM_MB}|g" \
-e "s|@VISIBLE_HOSTNAME@|${VISIBLE_HOSTNAME}|g" \
"$TMPL" > "$CONF"
# ---------------------------------------------------------------------------
# 2. Generate the cache-policy include from the domain lists + toggle.
# Storage gate runs first (whether to store); per-domain refresh_patterns
# only tune freshness of what survived the gate.
# ---------------------------------------------------------------------------
mkdir -p "$GEN_DIR"
{
echo "# AUTO-GENERATED by entrypoint.sh at container start -- do not edit."
echo "# Source: cache-domains.txt / cache-domains.regex + CACHE_ONLY_LISTED."
echo
# Boost ACLs are only emitted when their backing file is non-empty, else
# squid would error on an empty ACL.
exc=""
if has_entries "$CACHE_DOMAINS"; then
echo 'acl boost dstdomain "/etc/squid/cache-domains.txt"'
exc="$exc !boost"
fi
if has_entries "$CACHE_REGEX"; then
echo 'acl boost_re dstdom_regex "/etc/squid/cache-domains.regex"'
exc="$exc !boost_re"
fi
cat <<'ACLS'
acl has_query urlpath_regex \?
acl dyn_ext urlpath_regex (/cgi-bin/|\.(cgi|php|aspx?|jsp)$)
acl html_ext urlpath_regex \.html?(\?|$)
acl html_ct rep_mime_type -i ^text/html
# Storage gate: never store HTML (by ext or content-type) or dynamic content.
cache deny html_ext
cache deny dyn_ext
ACLS
# Query strings are dynamic EXCEPT on boosted domains (versioned static
# assets like app.js?v=123 should still cache there).
echo "cache deny has_query${exc}"
echo 'store_miss deny html_ct'
if [ "$CACHE_ONLY_LISTED" = "1" ]; then
echo
echo "# CACHE_ONLY_LISTED=1: store ONLY the listed domains."
if has_entries "$CACHE_DOMAINS"; then echo 'cache allow boost'; fi
if has_entries "$CACHE_REGEX"; then echo 'cache allow boost_re'; fi
echo 'cache deny all'
fi
if has_entries "$CACHE_DOMAINS"; then
echo
echo "# Boost: long TTL + force-cache for each listed wildcard domain."
grep -E '^[[:space:]]*[^#[:space:]]' "$CACHE_DOMAINS" | while read -r dom _rest; do
# ".example.com" / "example.com" -> "example[.]com" ([.] matches a
# literal dot portably -- avoids sed backslash-escaping quirks).
d=$(printf '%s' "$dom" | sed 's/^\.//; s/\./[.]/g')
printf 'refresh_pattern -i %s 1440 100%% 43200 override-expire ignore-private ignore-no-store override-lastmod\n' \
"^https?://([^/]+[.])?${d}/"
done
fi
} > "$GEN"
# ---------------------------------------------------------------------------
# 3. Stage the CA into a squid-readable tmpfs copy. The host key is 0600 root
# (mounted read-only); root copies it to /run owned by squid so the signer
# can read it without loosening the host file.
# ---------------------------------------------------------------------------
if [ ! -f "$SSL_SRC/squid-ca-cert.pem" ] || [ ! -f "$SSL_SRC/squid-ca-key.pem" ]; then
log "FATAL: CA missing in $SSL_SRC (need squid-ca-cert.pem + squid-ca-key.pem)."
exit 1
fi
install -d -m 0710 -o squid -g squid "$SSL_RUN"
install -m 0444 -o squid -g squid "$SSL_SRC/squid-ca-cert.pem" "$SSL_RUN/ca-cert.pem"
install -m 0400 -o squid -g squid "$SSL_SRC/squid-ca-key.pem" "$SSL_RUN/ca-key.pem"
# ---------------------------------------------------------------------------
# 4. Init the dynamic-cert DB + cache_dir structure (idempotent).
# ---------------------------------------------------------------------------
if [ ! -s "$SSL_DB/index.txt" ]; then
log "Initializing TLS cert DB at $SSL_DB..."
find "$SSL_DB" -mindepth 1 -delete 2>/dev/null || true
/usr/lib/squid/security_file_certgen -c -s "$SSL_DB" -M "${DYNAMIC_CERT_MEM_MB}MB"
fi
chown -R squid:squid "$SSL_DB" /var/cache/squid /var/log/squid 2>/dev/null || true
if [ ! -d /var/cache/squid/00 ]; then
log "Initializing cache_dir structure..."
squid -f "$CONF" -z --foreground || squid -f "$CONF" -z || true
fi
# ---------------------------------------------------------------------------
# 5. Validate the rendered + generated config, then hand off to squid.
# ---------------------------------------------------------------------------
squid -k parse -f "$CONF"
log "Starting squid (visible_hostname=${VISIBLE_HOSTNAME})..."
exec squid -N -f "$CONF" "$@"
+30
View File
@@ -0,0 +1,30 @@
# splice-domains.txt
#
# Domains Squid must NOT decrypt. With SSL-bump, "splice" = transparent
# passthrough: the client's TLS goes straight to the origin, so these are
# NOT cached and NOT inspected. Use for cert-pinned apps, banking, app-store /
# OS update channels, and anything that breaks under interception.
#
# One entry per line; a LEADING DOT matches all subdomains (.apple.com matches
# www.apple.com, gs.apple.com, ...). Full-line "#" comments only.
#
# IMPORTANT: do not also list a domain here AND in cache-domains.txt -- splice
# wins, so it would never cache. Tune this list for your environment.
# ── OS / app-store update channels (commonly cert-pinned) ──
.apple.com
.icloud.com
.mzstatic.com
.windowsupdate.com
.update.microsoft.com
.android.clients.google.com
.play.googleapis.com
# ── Banking / payments (examples -- add your own) ──
.paypal.com
.stripe.com
# ── Messaging / pinned services ──
.whatsapp.net
.signal.org
.telegram.org
+76
View File
@@ -0,0 +1,76 @@
# squid.conf.tmpl
#
# Rendered to /etc/squid/squid.conf by entrypoint.sh at container start:
# - @VAR@ placeholders are substituted from the environment (.env);
# - the cache policy (boost ACLs, storage gate, per-domain refresh_patterns)
# is generated from cache-domains.txt / .regex into the include below.
#
# Role: EXPLICIT forward proxy with SSL-bump. Clients set HTTP(S)_PROXY to this
# host:3128. (Transparent/intercepting mode is future work -- see README.)
visible_hostname @VISIBLE_HOSTNAME@
# ── Listener: explicit proxy, SSL-bump enabled ──────────────────────────────
# tls-cert/tls-key are the local CA used to mint per-host leaf certs on the fly.
# The key is staged into /run by the entrypoint so the squid user can read it
# while the on-host key file stays 0600 root.
http_port 3128 ssl-bump \
tls-cert=/run/squid/ca-cert.pem \
tls-key=/run/squid/ca-key.pem \
generate-host-certificates=on \
dynamic_cert_mem_cache_size=@DYNAMIC_CERT_MEM_MB@MB
sslcrtd_program /usr/lib/squid/security_file_certgen -s /var/lib/squid/ssl_db -M @DYNAMIC_CERT_MEM_MB@MB
sslcrtd_children 8 startup=1 idle=1
# Validate UPSTREAM certificates -- never silently MITM a broken origin cert.
sslproxy_cert_error deny all
# ── SSL-bump policy: peek at SNI → splice allowlist → bump everything else ──
acl step1 at_step SslBump1
acl splice_dom ssl::server_name "/etc/squid/splice-domains.txt"
ssl_bump peek step1
ssl_bump splice splice_dom
ssl_bump bump all
# ── Access control: deny-by-default, trusted clients only ───────────────────
acl SSL_ports port 443
acl Safe_ports port 80 443 21 70 210 1025-65535 280 488 591 777
acl CONNECT method CONNECT
http_access deny !Safe_ports
http_access deny CONNECT !SSL_ports
# Cache manager: localhost only (used by the container healthcheck).
http_access allow localhost manager
http_access deny manager
acl trusted_clients src @TRUSTED_CIDR@
http_access allow localhost
http_access allow trusted_clients
http_access deny all
# ── Cache sizing ────────────────────────────────────────────────────────────
cache_mem @CACHE_MEM_MB@ MB
maximum_object_size_in_memory 1 MB
maximum_object_size @MAX_OBJECT_SIZE_MB@ MB
cache_dir ufs /var/cache/squid @CACHE_SIZE_MB@ 16 256
coredump_dir /var/cache/squid
# ── Cache policy (generated: boost ACLs, storage gate, per-domain TTLs) ─────
include /etc/squid/conf.d/cache-policy.conf
# Base refresh patterns (fallbacks; per-domain boosts live in the include and
# are matched first). Packages get long, confident TTLs; everything else is
# conservative and revalidates.
refresh_pattern -i \.(deb|rpm|apk|pkg|tar\.(gz|xz|zst|bz2)|whl|jar)$ 10080 100% 43200
refresh_pattern . 0 20% 4320
# ── Logging ─────────────────────────────────────────────────────────────────
access_log stdio:/var/log/squid/access.log
cache_log /var/log/squid/cache.log
logfile_rotate 7
# ── Privilege drop / lifecycle ──────────────────────────────────────────────
cache_effective_user squid
pid_filename /run/squid.pid
shutdown_lifetime 5 seconds
+65 -2
View File
@@ -9,7 +9,9 @@
# What this does:
# 1. Installs docker + docker-cli-compose if missing.
# 2. Lays down docker-compose.yml, Caddyfile, .env.example in $STACK_DIR.
# 3. Generates .env on first run; existing .env is never overwritten.
# 3. Generates .env on first run. On a re-run an existing .env is kept,
# except for values passed explicitly that run, which are written
# through -- .env is what compose interpolates into the container.
# 4. Prompts for required values not preset (BASE_DOMAIN, ISSUER_URL,
# REDIRECT_URL, ACME_EMAIL).
# 5. Opens TCP 80/443 in UFW if active.
@@ -36,6 +38,18 @@ set -euo pipefail
: "${FORCE:=0}"
: "${SKIP_PROMPTS:=0}" # non-interactive: require values via env, no prompts
[[ "$SKIP_PROMPTS" == "1" ]] && FORCE=1
ENV_FILE="$STACK_DIR/.env"
# All four values reach Caddy only by compose interpolating .env, so .env has to absorb
# what this run passed or the run deploys the old values while reporting the
# new ones. Which keys arrived in the ENVIRONMENT must be recorded BEFORE the
# ":=" defaults below, which make an unset variable look like an empty one.
RUNTIME_KEYS=(BASE_DOMAIN ISSUER_URL REDIRECT_URL ACME_EMAIL)
KEYS_FROM_ENV=()
for _k in "${RUNTIME_KEYS[@]}"; do
if [[ -n "${!_k+x}" ]]; then KEYS_FROM_ENV+=("$_k"); fi
done
: "${BASE_DOMAIN:=}"
: "${ISSUER_URL:=}"
: "${REDIRECT_URL:=}"
@@ -148,6 +162,24 @@ prompt() {
fi
}
# On a re-run the deployed values live in .env. Load them into any key NOT
# passed this run, so the prompts below do not force the operator to retype
# everything (and so SKIP_PROMPTS=1 does not die demanding values .env has).
if [[ -f "$ENV_FILE" ]]; then
_preloaded=()
for _k in "${RUNTIME_KEYS[@]}"; do
if [[ " ${KEYS_FROM_ENV[*]-} " == *" ${_k} "* ]]; then continue; fi
_v=$(sed -n "s/^${_k}=//p" "$ENV_FILE" | tail -n1)
if [[ -z "$_v" ]]; then continue; fi
printf -v "$_k" '%s' "$_v"
_preloaded+=("$_k")
done
if (( ${#_preloaded[@]} > 0 )); then
log "Reusing ${ENV_FILE} values for: ${_preloaded[*]}"
log " (pass VAR=... to change one; it is then written back to .env)"
fi
fi
prompt BASE_DOMAIN "Apex domain to serve from (e.g. example.com)"
prompt ISSUER_URL "OIDC issuer URL (e.g. https://auth.example.com)"
prompt REDIRECT_URL "Where to redirect non-webfinger traffic (e.g. https://example.org)"
@@ -171,6 +203,28 @@ install -d -m 0750 "$STACK_DIR"
install -m 0640 "$SCRIPT_DIR/docker-compose.yml" "$STACK_DIR/docker-compose.yml"
install -m 0640 "$SCRIPT_DIR/Caddyfile" "$STACK_DIR/Caddyfile"
set_env() { # <KEY> <value>: update KEY in .env, or append if absent
# The value goes through the ENVIRONMENT, never interpolated into a sed
# script. Interpolating it corrupts any value containing & (sed expands it to
# the whole match) and aborts the run on one containing the s||| delimiter --
# which is reachable for an OIDC issuer URL or a redirect query string.
local key="$1" val="$2" tmp
if [[ ! -f "$ENV_FILE" ]]; then
printf '%s=%s
' "$key" "$val" >> "$ENV_FILE"
return 0
fi
tmp="$(mktemp)"
_SE_KEY="$key" _SE_VAL="$val" awk '
BEGIN { k = ENVIRON["_SE_KEY"]; v = ENVIRON["_SE_VAL"]; seen = 0 }
!seen && index($0, k "=") == 1 { print k "=" v; seen = 1; next }
{ print }
END { if (!seen) print k "=" v }
' "$ENV_FILE" > "$tmp"
cat "$tmp" > "$ENV_FILE" # rewrite in place: keeps the original mode/owner
rm -f "$tmp"
}
ENV_FILE="$STACK_DIR/.env"
if [[ ! -f "$ENV_FILE" ]]; then
log "Seeding $ENV_FILE..."
@@ -182,7 +236,16 @@ if [[ ! -f "$ENV_FILE" ]]; then
-e "s|^ACME_EMAIL=.*|ACME_EMAIL=${ACME_EMAIL}|" \
"$ENV_FILE"
else
log ".env exists; leaving it alone."
# Keep the file, but absorb what was passed this run: compose reads .env,
# so letting it drift means deploying the old values while the banner
# prints the new ones. Keys not passed are left untouched.
log ".env exists; keeping it (only values passed this run are updated)."
for _k in ${KEYS_FROM_ENV[@]+"${KEYS_FROM_ENV[@]}"}; do
_cur=$(sed -n "s/^${_k}=//p" "$ENV_FILE" | tail -n1)
if [[ "$_cur" == "${!_k}" ]]; then continue; fi
set_env "$_k" "${!_k}"
log " ${_k}: ${_cur:-<unset>} -> ${!_k:-<empty>}"
done
fi
# Validate required values are present.
-5
View File
@@ -15,11 +15,6 @@ REPO_BRANCH=main
ACME_EMAIL=admin@example.com
SSH_PORT=22
# ─── ScoutFS (kanrisha) ─────────────────────────────────────────────────────
# License key passed to setup-scoutfs.sh by the kanrisha deployment. Leave
# empty to be prompted per host; set it here to reuse one key unattended.
SCOUTFS_LICENSE=
# ─── SSH authorized keys ────────────────────────────────────────────────────
# resolve_ssh_keys() prefers this URL when set (fetched live, always current),
# e.g. a GitHub keys endpoint (https://github.com/<user>.keys) or a raw
+309 -14
View File
@@ -2,11 +2,22 @@
#
# harden-firewall.sh
#
# Deny-by-default host firewall (iptables) for Alpine, Debian, and Alma Linux.
# Hardens the INPUT chain to: loopback, established/related, ICMP, SSH (on the
# configured port), and any explicitly-registered ports -- everything else is
# dropped. OUTPUT stays open (egress allow); FORWARD is left untouched so Docker
# container networking is unaffected.
# Deny-by-default host firewall for Alpine, Debian, and Alma Linux.
#
# Backend per family (override with FW_BACKEND=iptables|firewalld):
# alpine/debian iptables -- hardens the INPUT chain to loopback,
# established/related, ICMP, SSH (configured port) + registered
# ports; everything else dropped. Persisted natively, no boot hook.
# rhel (Alma) firewalld -- the native RHEL firewall. The default zone is
# deny-by-default; we open SSH + registered ports and let
# sshguard block via the sshguard-firewalld backend.
# proxmox none -- SKIPPED. Proxmox VE/PMG are Debian, but pve-firewall
# already owns the host ruleset; see the Proxmox section below.
# On a host hardened BEFORE this skip existed, 'disable' removes
# our ruleset, config and boot restore, and hands it back.
# OUTPUT/egress stays open; FORWARD is left untouched so Docker networking is
# unaffected. The allow/deny/list/disable subcommands work on either managed
# backend; on Proxmox they explain themselves and change nothing.
#
# All distro-specific operations go through scripts/oslib.sh. The OS-specific
# surface here is exactly three things, all in oslib: which packages provide
@@ -38,6 +49,7 @@
# OPEN_PORTS="80/tcp 443/tcp" bash harden-firewall.sh # open extra ports at install
# FW_SSH_SOURCE=10.0.0.0/8 bash harden-firewall.sh # restrict SSH to a source CIDR
# FW_ALLOW_PING=0 bash harden-firewall.sh # drop ICMP echo (ping)
# FW_IGNORE_PVE=1 bash harden-firewall.sh # harden a Proxmox host anyway
#
# bash harden-firewall.sh allow 443/tcp 51820/udp # register + apply
# bash harden-firewall.sh allow web # preset: 80/tcp + 443/tcp
@@ -82,6 +94,22 @@ if [[ -z "${SSH_PORT:-}" ]]; then
[[ -n "$SSH_PORT" ]] || SSH_PORT=22
fi
# Firewall backend: firewalld on RHEL/Alma (native, and the sshguard-firewalld
# backend blocks there), iptables everywhere else. Override with
# FW_BACKEND=iptables|firewalld.
: "${FW_BACKEND:=}"
if [[ -z "$FW_BACKEND" ]]; then
[[ "$OS_FAMILY" == rhel ]] && FW_BACKEND=firewalld || FW_BACKEND=iptables
fi
# Proxmox VE / Proxmox Mail Gateway: pve-firewall owns the host ruleset, so we
# do not manage a firewall here at all. This overrides an explicit FW_BACKEND
# too -- FW_IGNORE_PVE=1 is the single, deliberate escape hatch.
: "${FW_IGNORE_PVE:=0}"
if is_proxmox && [[ "$FW_IGNORE_PVE" != "1" ]]; then
FW_BACKEND=pve
fi
# ============================================================================
# Spec parsing: PORT[/PROTO][@CIDR] and named presets -> "PORT/PROTO[ CIDR]" lines
# ============================================================================
@@ -144,6 +172,9 @@ seed_open_ports() {
# ============================================================================
write_apply() {
local save_cmd; save_cmd="$(fw_save_cmd)"
# Alpine's baselayout ships /usr/local/{bin,lib,share} but NOT sbin, so the
# redirect below fails with "No such file or directory" unless we make it.
install -d -m 0755 "${APPLY%/*}"
cat > "$APPLY" <<'ENGINE'
#!/bin/sh
# Managed by harden-firewall.sh -- do not edit by hand (regenerated on install).
@@ -259,21 +290,275 @@ ensure_installed() {
install -d -m 0755 "$PORTS_DIR"
}
# ============================================================================
# firewalld backend (RHEL/Alma) -- same CLI as the iptables engine, driven via
# firewall-cmd. firewalld zones are deny-by-default for unsolicited inbound and
# persist natively (--permanent), and sshguard blocks via the sshguard-firewalld
# backend -- so there is no INPUT->sshguard jump, no native save, no boot hook.
# The default zone is the managed surface.
# ============================================================================
_fw_zone() { firewall-cmd --get-default-zone 2>/dev/null || echo public; }
# Add/remove one "PORT/PROTO" spec on the zone (permanent). A range a:b becomes
# a-b; a source CIDR turns the rule into a rich rule.
_fw_spec() { # <add|remove> <zone> <port/proto> [cidr]
local op="$1" z="$2" spec="$3" cidr="${4:-}" port proto flag
port="${spec%%/*}"; proto="${spec##*/}"; [[ "$proto" == "$spec" ]] && proto=tcp
port="${port//:/-}"
[[ "$op" == add ]] && flag=--add || flag=--remove
if [[ -n "$cidr" ]]; then
firewall-cmd -q --permanent --zone="$z" \
"${flag}-rich-rule=rule family=\"ipv4\" source address=\"$cidr\" port port=\"$port\" protocol=\"$proto\" accept" 2>/dev/null || true
else
firewall-cmd -q --permanent --zone="$z" "${flag}-port=$port/$proto" 2>/dev/null || true
fi
}
# Apply each "PORT/PROTO [CIDR]" spec line from stdin to the zone.
_fw_apply_specs() { # <add|remove> <zone> (spec lines on stdin)
local op="$1" z="$2" line spec cidr
while IFS= read -r line; do
[[ -n "$line" ]] || continue
spec="$(awk '{print $1}' <<< "$line")"; cidr="$(awk '{print $2}' <<< "$line")"
_fw_spec "$op" "$z" "$spec" "$cidr"
done
}
apply_firewalld() {
log "Detected OS: ${OS_ID} (family ${OS_FAMILY}) -- firewall backend: firewalld"
install_firewalld
local z; z="$(_fw_zone)"
log "Hardening firewalld zone '$z' (deny-by-default; SSH/${SSH_PORT}${FW_SSH_SOURCE:+ from $FW_SSH_SOURCE}$([[ $FW_ALLOW_PING == 1 ]] || echo '; no ping'))..."
# We own SSH via an explicit port/rich-rule, so drop the stock services that
# would otherwise leave 22 (and cockpit's 9090) open on the zone.
local s
for s in ssh cockpit; do firewall-cmd -q --permanent --zone="$z" --remove-service="$s" 2>/dev/null || true; done
# SSH (source-restricted -> rich rule; otherwise a plain port).
_fw_spec add "$z" "$SSH_PORT/tcp" "$FW_SSH_SOURCE"
# Extra ports from OPEN_PORTS (same spec grammar as the iptables path).
if [[ -n "$OPEN_PORTS" ]]; then
# shellcheck disable=SC2086
_fw_apply_specs add "$z" <<< "$(_expand_args ${OPEN_PORTS//,/ })"
fi
# Ping: firewalld permits echo-request by default; block it only when asked.
if [[ "$FW_ALLOW_PING" == "1" ]]; then
firewall-cmd -q --permanent --zone="$z" --remove-icmp-block=echo-request 2>/dev/null || true
else
firewall-cmd -q --permanent --zone="$z" --add-icmp-block=echo-request 2>/dev/null || true
fi
firewall-cmd -q --reload
log "firewalld active on zone '$z'. Add ports: 'bash $0 allow <port/proto>' Recover: 'bash $0 disable'"
}
allow_firewalld() {
local z lines; z="$(_fw_zone)"; lines="$(_expand_args "$@")"
_fw_apply_specs add "$z" <<< "$lines"
firewall-cmd -q --reload
log "Allowed in firewalld zone '$z': $(printf '%s' "$lines" | tr '\n' ' ')"
}
deny_firewalld() {
local z lines; z="$(_fw_zone)"; lines="$(_expand_args "$@")"
_fw_apply_specs remove "$z" <<< "$lines"
firewall-cmd -q --reload
log "Removed from firewalld zone '$z': $(printf '%s' "$lines" | tr '\n' ' ')"
}
list_firewalld() {
local z; z="$(_fw_zone)"
echo "firewalld default zone: $z"
firewall-cmd --list-all --zone="$z" 2>/dev/null | sed 's/^/ /' || echo " (firewalld not running)"
}
# Recovery: re-open SSH so a misconfig can't lock you out. Leaves firewalld
# running (sshguard-firewalld keeps working). To turn the firewall off entirely:
# systemctl disable --now firewalld.
disable_firewalld() {
local z; z="$(_fw_zone)"
firewall-cmd -q --permanent --zone="$z" --add-service=ssh 2>/dev/null || true
firewall-cmd -q --permanent --zone="$z" --add-port="${SSH_PORT}/tcp" 2>/dev/null || true
firewall-cmd -q --reload || true
warn "Re-opened SSH (service ssh + ${SSH_PORT}/tcp) on firewalld zone '$z'; firewalld left running."
warn "Re-harden: 'bash $0 apply' | turn firewalld off: 'systemctl disable --now firewalld'"
}
# ============================================================================
# Proxmox backend (VE / Mail Gateway) -- deliberately does nothing.
#
# These are Debian hosts, so the iptables backend would happily install itself
# here. It must not -- though not for the reason you would guess. pve-firewall
# does NOT delete foreign rules: it applies every ruleset with
# `iptables-restore -n` (--noflush), only ever flushes chains matching its own
# patterns (PVEFW-*, tapNiM-*, vethNiM-*, fwbrN-*, GROUP-*), APPENDS
# `-A INPUT -j PVEFW-INPUT` only when that exact hook is missing, and never sets
# a built-in chain's policy at all. Our rules would survive it. The damage runs
# the other way -- we break Proxmox:
#
# * `-P INPUT DROP` is ours alone, and PVEFW-HOST-IN RETURNs on accept rather
# than ACCEPTing ("we use RETURN because we need to check also tap rules"),
# so traffic Proxmox explicitly allowed falls out of its chain and lands on
# our DROP. We silently override the platform's own accepts.
# * the persistence layer is worse. netfilter-persistent restores at boot with
# a full iptables-restore (no --noflush), wiping PVE's -j PVEFW-INPUT hook
# along with everything else, until the daemon re-appends it on its next
# ~10s cycle.
# * a deny-by-default chain has to enumerate the whole platform to stay usable:
# 8006/tcp web UI, 5405-5412/udp corosync (5405 + knet link number, up to 8
# links), 60000-60050/tcp migration, 5900-5999/tcp VNC, 3128/tcp SPICE,
# 22/tcp SSH, 111/udp rpcbind with NFS storage, and Ceph's 6789 / 3300 /
# 6800-7300 when hyperconverged. Miss one and you lose the GUI or the cluster.
# Miss `-i lo` and pveproxy can no longer reach pvedaemon on 127.0.0.1:85,
# which breaks the API locally, not just remotely.
# * under the opt-in nftables backend (PVE 8.2+) proxmox-firewall registers its
# own nft input hook, and an nft DROP beats an iptables ACCEPT -- so our rules
# would not even be authoritative.
#
# So: skip, and leave the host to the firewall the platform ships. Be clear-eyed
# that this is NOT the same as protected: that firewall is off until someone
# turns it on (see _pve_explain).
# ============================================================================
PVE_CLUSTER_FW=/etc/pve/firewall/cluster.fw
PVE_HOST_FW=/etc/pve/local/host.fw # == /etc/pve/nodes/<node>/host.fw
# The shared explanation. Printed by every subcommand so the reason is never a
# mystery, whichever entry point the operator came in through.
_pve_explain() {
warn "Proxmox detected -- pve-firewall owns the host ruleset; not touching it."
warn " Manage the firewall natively instead:"
warn " Datacenter -> Firewall -> ${PVE_CLUSTER_FW}"
warn " Node -> Firewall -> ${PVE_HOST_FW}"
warn " pve-firewall status | compile | localnet"
warn " Host rules go under a [RULES] section -- a line before one is discarded:"
warn " [RULES]"
warn " IN ACCEPT -p tcp -dport 443"
warn " HEADS UP: Proxmox's firewall is DISABLED by default (cluster-wide 'enable'"
warn " defaults to 0, and pve-firewall tears its chains down every ~10s while it is)."
warn " Until you enable it at Datacenter -> Firewall -> Options this host has NO"
warn " host firewall -- 8006, 22, 3128 and 111 are open on every interface. The node"
warn " panel reading 'Firewall: Yes' is ignored while the datacenter one reads 'No'."
}
# Did an EARLIER harden-firewall.sh run (from before this skip existed) leave our
# firewall on this host? Config, engine, or a DROP policy only we would have set.
_pve_leftovers() {
PVE_LEFTOVER_WHY=""
if [[ -e "$APPLY" || -d "$FW_DIR" ]]; then
PVE_LEFTOVER_WHY="our engine/config is still installed"
return 0
fi
# pve-firewall never sets the INPUT policy itself (its own default lives inside
# PVEFW-INPUT), so a DROP policy on a Proxmox host was set by something else --
# us, or an admin. Say which signal fired rather than assuming.
if iptables -S INPUT 2>/dev/null | grep -q '^-P INPUT DROP'; then
PVE_LEFTOVER_WHY="INPUT policy is DROP (not something pve-firewall sets)"
return 0
fi
return 1
}
apply_pve() {
log "Detected OS: ${OS_ID} (family ${OS_FAMILY}) on Proxmox -- firewall backend: none (skipped)"
_pve_explain
if _pve_leftovers; then
warn " NOTE: this host carries a deny-by-default ruleset (${PVE_LEFTOVER_WHY})."
warn " Remove them with: bash $0 disable"
fi
warn " Override (fights pve-firewall, can lock you out): FW_IGNORE_PVE=1 bash $0 apply"
}
# allow/deny must NOT report success when nothing was opened -- fail loudly and
# say where the rule really belongs.
allow_deny_pve() {
local verb="$1"; shift
_pve_explain
die "Refusing to $verb $* on a Proxmox host: add it to ${PVE_HOST_FW} (or the GUI) instead."
}
list_pve() {
echo "Proxmox host -- firewall managed by pve-firewall (harden-firewall.sh is a no-op here)."
echo " cluster: ${PVE_CLUSTER_FW}"
echo " host: ${PVE_HOST_FW}"
echo
echo "pve-firewall status:"
if command -v pve-firewall >/dev/null 2>&1; then
pve-firewall status 2>&1 | sed 's/^/ /' || true
else
echo " (pve-firewall not installed)"
fi
}
# On a host we skipped there is nothing to unlock -- but a host hardened BEFORE the
# skip existed still carries our rules, and `disable` is exactly what you reach for
# then, so do the real cleanup rather than claiming there is nothing to do.
disable_pve() {
if ! _pve_leftovers; then
_pve_explain
warn "Nothing to disable: this host was skipped, so no rules of ours are in place."
warn " If Proxmox's own firewall is locking you out: pve-firewall stop"
return
fi
warn "Proxmox host carrying a deny-by-default ruleset (${PVE_LEFTOVER_WHY}) -- removing it."
local ipt s f
# Policy BEFORE flush: flushing a chain whose policy is still DROP kills the SSH
# session this is running over.
for ipt in iptables ip6tables; do
command -v "$ipt" >/dev/null 2>&1 || continue
"$ipt" -P INPUT ACCEPT || true
"$ipt" -F INPUT || true
done
# Deliberately NOT fw_save_cmd here, unlike the normal disable path: on Proxmox
# that snapshot captures pve-firewall's own PVEFW-* chains and restores the stale
# copy at the next boot, before pve-firewall starts. Kill the restore instead.
for s in $(fw_restore_services); do svc_disable "$s"; done
for f in $(fw_saved_files); do
[[ -f "$f" ]] || continue
mv -f "$f" "$f.bak-harden-firewall"
log "Saved ruleset moved aside: $f -> $f.bak-harden-firewall"
done
rm -rf "$APPLY" "$FW_DIR"
log "Removed ${APPLY}, ${FW_DIR}, and the boot-time restore."
# The flush took the INPUT -> sshguard jump with it; put it back so brute-force
# protection keeps working until harden-ssh.sh installs its boot hook.
if command -v sshguard >/dev/null 2>&1; then
for ipt in iptables ip6tables; do
command -v "$ipt" >/dev/null 2>&1 || continue
"$ipt" -N sshguard 2>/dev/null || true
"$ipt" -C INPUT -p tcp --dport "$SSH_PORT" -j sshguard 2>/dev/null || "$ipt" -I INPUT -p tcp --dport "$SSH_PORT" -j sshguard || true
done
warn "Re-added the INPUT -> sshguard jump on ${SSH_PORT}; re-run harden-ssh.sh for its boot hook."
fi
# Hand the ruleset back to pve-firewall.
if command -v pve-firewall >/dev/null 2>&1; then
pve-firewall restart >/dev/null 2>&1 || warn "pve-firewall restart failed -- check 'pve-firewall status'."
fi
_pve_explain
}
# ============================================================================
# Subcommands
# ============================================================================
cmd_apply() {
log "Detected OS: ${OS_ID} (family ${OS_FAMILY}, init ${INIT_SYSTEM})"
if [[ "$FW_BACKEND" == pve ]]; then apply_pve; return; fi
if [[ "$FW_BACKEND" == firewalld ]]; then apply_firewalld; return; fi
# RHEL/Alma ships firewalld; it and iptables-services fight over nftables.
# Decoupled from FORCE on purpose: harden-ssh passes FORCE=1 to skip its SSH
# prompt, and we must NOT let that bulldoze an active firewalld. Bail clearly
# instead (harden-ssh treats this as a warning and leaves firewalld in place).
log "Detected OS: ${OS_ID} (family ${OS_FAMILY}, init ${INIT_SYSTEM}) -- firewall backend: iptables"
# FW_BACKEND=iptables was forced on a RHEL/Alma host where firewalld is the
# native firewall; the two fight over nftables. Bail clearly unless overridden.
if [[ "$OS_FAMILY" == rhel ]] && command -v firewall-cmd >/dev/null 2>&1 \
&& firewall-cmd --state >/dev/null 2>&1; then
warn "firewalld is active. It conflicts with this iptables firewall."
warn "Disable it first: systemctl disable --now firewalld"
[[ "${FW_IGNORE_FIREWALLD:-0}" == "1" ]] || die "Refusing to proceed while firewalld is active (set FW_IGNORE_FIREWALLD=1 to override)."
warn "firewalld is active but FW_BACKEND=iptables was forced; they conflict."
warn "Prefer the firewalld backend (unset FW_BACKEND), or: systemctl disable --now firewalld"
[[ "${FW_IGNORE_FIREWALLD:-0}" == "1" ]] || die "Refusing iptables while firewalld is active (set FW_IGNORE_FIREWALLD=1 to override)."
fi
log "Installing iptables + native persistence package..."
@@ -293,6 +578,8 @@ cmd_apply() {
cmd_allow() {
[[ "$#" -ge 1 ]] || die "usage: allow <port[/proto][@cidr]|web|http|https>..."
[[ "$FW_BACKEND" == pve ]] && allow_deny_pve allow "$@"
[[ "$FW_BACKEND" == firewalld ]] && { allow_firewalld "$@"; return; }
ensure_installed
local lines file="$PORTS_DIR/manual.rule"
lines="$(_expand_args "$@")"
@@ -309,6 +596,8 @@ cmd_allow() {
cmd_deny() {
[[ "$#" -ge 1 ]] || die "usage: deny <port[/proto][@cidr]|web|http|https>..."
[[ "$FW_BACKEND" == pve ]] && allow_deny_pve deny "$@"
[[ "$FW_BACKEND" == firewalld ]] && { deny_firewalld "$@"; return; }
ensure_installed
local lines l file tmp
lines="$(_expand_args "$@")"
@@ -332,6 +621,8 @@ cmd_deny() {
}
cmd_list() {
[[ "$FW_BACKEND" == pve ]] && { list_pve; return; }
[[ "$FW_BACKEND" == firewalld ]] && { list_firewalld; return; }
echo "Registered ports ($PORTS_DIR):"
if ls "$PORTS_DIR"/*.rule >/dev/null 2>&1; then
for f in "$PORTS_DIR"/*.rule; do
@@ -349,6 +640,8 @@ cmd_list() {
# Recovery escape hatch: open the policy and flush our rules, then persist that
# open state so a reboot does not re-DROP. The sshguard chain is left intact.
cmd_disable() {
[[ "$FW_BACKEND" == pve ]] && { disable_pve; return; }
[[ "$FW_BACKEND" == firewalld ]] && { disable_firewalld; return; }
local ipt
for ipt in iptables ip6tables; do
command -v "$ipt" >/dev/null 2>&1 || continue
@@ -361,7 +654,9 @@ cmd_disable() {
}
usage() {
sed -n '2,50p' "$0" | sed 's/^#\{0,1\} \{0,1\}//'
# The whole header comment block, however long it grows: line 2 until the
# first non-comment line.
awk 'NR == 1 { next } !/^#/ { exit } { print }' "$0" | sed 's/^#\{0,1\} \{0,1\}//'
}
# ============================================================================
+64 -16
View File
@@ -57,8 +57,11 @@ log "Detected OS: ${OS_ID} (family ${OS_FAMILY}, init ${INIT_SYSTEM})"
# 1. Packages
# ----------------------------------------------------------------------------
log "Installing OpenSSH + sshguard + iptables..."
install_openssh
install_bruteforce_protection
install_openssh || die "OpenSSH packages failed to install; cannot harden. Fix the package error above, then re-run."
# sshguard is best-effort (see harden-ssh.sh): never let a missing brute-force
# package abort the whole bastion hardening.
install_bruteforce_protection \
|| warn "sshguard not installed; brute-force protection is OFF. Add it later with: dnf install -y epel-release sshguard. Continuing with the rest of the hardening."
ensure_gum || warn "gum not installed; sshuser will use its CLI mode."
SFTP_PATH="$(sftp_server_path)"
@@ -66,6 +69,7 @@ NOLOGIN="$(nologin_path)" # /sbin/nologin (Alpine/Alma) or /usr/sbin/nologi
# Install the sshuser tool alongside this script if present.
if [[ -f "$SCRIPT_DIR/sshuser.sh" ]]; then
install -d -m 0755 /usr/local/bin
install -m 0755 "$SCRIPT_DIR/sshuser.sh" /usr/local/bin/sshuser
log "Installed /usr/local/bin/sshuser"
fi
@@ -73,19 +77,47 @@ fi
# ----------------------------------------------------------------------------
# 2. PQ KEX detection
# ----------------------------------------------------------------------------
log "Checking OpenSSH version supports PQ KEX..."
SSH_VER=$(ssh -V 2>&1 | grep -oE 'OpenSSH_[0-9]+\.[0-9]+' | head -1 | sed 's/OpenSSH_//')
SSH_MAJOR=${SSH_VER%%.*}
SSH_MINOR=${SSH_VER##*.}
HAS_MLKEM=0; HAS_SNTRUP=0
[[ $SSH_MAJOR -gt 9 || ( $SSH_MAJOR -eq 9 && $SSH_MINOR -ge 0 ) ]] && HAS_SNTRUP=1
[[ $SSH_MAJOR -gt 9 || ( $SSH_MAJOR -eq 9 && $SSH_MINOR -ge 9 ) ]] && HAS_MLKEM=1
[[ $HAS_SNTRUP -eq 1 || $HAS_MLKEM -eq 1 ]] || die "OpenSSH ${SSH_VER} has no PQ KEX. Need >= 9.0."
log "OpenSSH ${SSH_VER}: ML-KEM=${HAS_MLKEM} sntrup761=${HAS_SNTRUP}"
log "Checking which key exchange methods this OpenSSH supports..."
# Cosmetic only -- every decision below comes from `ssh -Q kex`, not this. A build
# whose banner does not match (OpenSSH_for_Windows_9.5p2, vendor forks) must not
# abort the run: without the guard, grep's non-match fails the pipeline under
# pipefail and set -e kills the script here with no message at all.
SSH_VER=$(ssh -V 2>&1 | grep -oE 'OpenSSH_[0-9]+[.][0-9]+' | head -1 | sed 's/OpenSSH_//' || true)
KEX_LIST=""
[[ $HAS_MLKEM -eq 1 ]] && KEX_LIST="mlkem768x25519-sha256"
[[ $HAS_SNTRUP -eq 1 ]] && KEX_LIST="${KEX_LIST:+$KEX_LIST,}sntrup761x25519-sha512"
# Ask the binary what it supports rather than deriving it from the version -- the
# same algorithm has two spellings and guessing wrong either breaks sshd_config or
# locks out clients that implement it under the other name. See oslib.
KEX_PQ="$(ssh_kex_pq_list)"
KEX_LIST="$(ssh_kex_list)"
if [[ -n "$KEX_PQ" ]]; then
if [[ "${SSH_ALLOW_CLASSIC_KEX:-0}" == "1" ]]; then
KEX_NOTE='# --- Key exchange: post-quantum hybrid + classical fallback ---
# curve25519-sha256 is offered for clients too old for any PQ method
# (SSH_ALLOW_CLASSIC_KEX=1). A session that negotiates it has NO store-now-
# decrypt-later protection -- drop the fallback once those clients are gone.'
warn "SSH_ALLOW_CLASSIC_KEX=1 -- offering curve25519-sha256 next to the PQ methods."
warn " Admits clients with no PQ KEX at all (Windows in-box ssh.exe), at the cost of"
warn " store-now-decrypt-later protection for any session that negotiates it."
else
KEX_NOTE='# --- Key exchange: post-quantum hybrid only ---
# Every classical-only method is rejected, which is what protects the session key
# against "store now, decrypt later". Clients with no PQ KEX cannot connect --
# re-run with SSH_ALLOW_CLASSIC_KEX=1 to also offer curve25519-sha256.'
fi
else
# No post-quantum method on this host at all. Decide on THAT, not on an empty
# list: without the opt-in the classical names are never collected, so an empty
# list here would otherwise be misreported as "no usable KEX".
[[ "${SSH_ALLOW_CLASSIC_KEX:-0}" == "1" ]] \
|| die "OpenSSH ${SSH_VER:-?} has no post-quantum KEX (needs >= 8.5 built with sntrup761). Re-run with SSH_ALLOW_CLASSIC_KEX=1 to accept classical-only."
[[ -n "$KEX_LIST" ]] \
|| die "OpenSSH ${SSH_VER:-?} reports no usable key exchange method at all ('ssh -Q kex' returned nothing)."
KEX_NOTE='# --- Key exchange: CLASSICAL ONLY ---
# This OpenSSH has no post-quantum method. No store-now-decrypt-later protection.'
warn "OpenSSH ${SSH_VER:-?} has no PQ KEX -- classical curve25519 only."
fi
log "OpenSSH ${SSH_VER:-?}: KexAlgorithms ${KEX_LIST}"
# ----------------------------------------------------------------------------
# 3. Host keys (Ed25519 only)
@@ -170,7 +202,7 @@ LogLevel VERBOSE
# --- Host key: Ed25519 only ---
HostKey /etc/ssh/ssh_host_ed25519_key
# --- Post-quantum hybrid KEX only ---
${KEX_NOTE}
KexAlgorithms ${KEX_LIST}
# --- Modern ciphers and MACs ---
@@ -303,6 +335,20 @@ EOF
# standalone boot hook. Otherwise fall back to the minimal init-agnostic boot
# hook that just (re)inserts the jump at every boot.
: "${ENABLE_FIREWALL:=1}"
# Proxmox (VE/PMG) is Debian underneath, but pve-firewall already owns the host
# ruleset -- harden-firewall.sh skips those hosts, so fall through to the boot
# hook and keep sshguard's jump. That jump is safe there: pve-firewall restores
# with --noflush and only ever flushes its own PVEFW-* chains, and the hook it
# adds to INPUT is APPENDED -- so a jump inserted with -I sits ahead of it and
# keeps getting first look at NEW connections, firewall enabled or not.
# FW_IGNORE_PVE=1 forces our firewall anyway.
if [[ "$ENABLE_FIREWALL" == "1" && "${FW_IGNORE_PVE:-0}" != "1" ]] && is_proxmox; then
warn "Proxmox detected -- pve-firewall owns the host firewall; installing only the sshguard jump hook."
# A host hardened before the skip existed still has a live DROP chain that only
# accepts the OLD port -- changing SSH_PORT here would lock you out of it.
[[ -x /usr/local/sbin/firewall-apply || -d /etc/firewall ]] && warn " This host still has an older harden-firewall.sh ruleset. Run 'harden-firewall.sh disable' FIRST, especially if you are changing SSH_PORT."
ENABLE_FIREWALL=0
fi
if [[ "$ENABLE_FIREWALL" == "1" && -f "$SCRIPT_DIR/harden-firewall.sh" ]]; then
log "Installing host firewall (deny-by-default INPUT; carries the sshguard jump)..."
SSH_PORT="$SSH_PORT" OPEN_PORTS="${OPEN_PORTS:-}" \
@@ -424,5 +470,7 @@ EOF
fi
log "Reloading ${SSHD_SVC}..."
svc_reload "$SSHD_SVC"
# Reload -- and on Alpine swap in the PAM sshd build if the running one
# predates this config (oslib explains why that can happen).
sshd_apply_config
log "Done."
+78 -25
View File
@@ -15,7 +15,9 @@
# 1. Generates fresh Ed25519 host keys; removes RSA/ECDSA/DSA host keys
# 2. Generates an Ed25519 root keypair, installs the public key into
# /root/.ssh/authorized_keys, and PRINTS the private key to stdout once.
# 3. Forces post-quantum hybrid KEX only (mlkem768x25519, sntrup761x25519).
# 3. Forces post-quantum hybrid KEX only (mlkem768x25519, sntrup761x25519 --
# every spelling this OpenSSH supports). SSH_ALLOW_CLASSIC_KEX=1 also
# offers curve25519-sha256 for clients with no PQ method at all.
# 4. Modern ciphers and MACs only.
# 5. Disables everything but an interactive terminal + SFTP (no forwarding,
# tunneling, X11, agent, password auth).
@@ -34,6 +36,16 @@
# SSH_PORT=2222 bash harden-ssh.sh # change port
# ALLOWED_IP=1.2.3.4 bash harden-ssh.sh # whitelist your client IP
# FORCE=1 bash harden-ssh.sh # skip the confirm prompt
# SSH_ALLOW_CLASSIC_KEX=1 bash harden-ssh.sh # also offer curve25519-sha256
#
# SSH_ALLOW_CLASSIC_KEX exists for one reason: clients too old for ANY
# post-quantum KEX. The ssh.exe bundled with Windows is the common case -- its
# `ssh -Q kex` lists no mlkem or sntrup at all, so it cannot connect to a
# PQ-only host however the server spells the algorithms. Turning this on is a
# real trade: a session that negotiates curve25519-sha256 is safe against a
# classical attacker but has no store-now-decrypt-later protection. Prefer
# upgrading the client (Git for Windows, WSL, or a current Win32-OpenSSH ship
# OpenSSH 9.x) and leave this off.
set -euo pipefail
@@ -64,33 +76,61 @@ log "Detected OS: ${OS_ID} (family ${OS_FAMILY}, init ${INIT_SYSTEM})"
# ----------------------------------------------------------------------------
if ! command -v ssh >/dev/null 2>&1; then
log "ssh not found; installing openssh..."
install_openssh
install_openssh || die "Could not install OpenSSH; cannot harden. Fix the package error above, then re-run."
fi
log "Checking OpenSSH version supports PQ KEX..."
SSH_VER=$(ssh -V 2>&1 | grep -oE 'OpenSSH_[0-9]+\.[0-9]+' | head -1 | sed 's/OpenSSH_//')
SSH_MAJOR=${SSH_VER%%.*}
SSH_MINOR=${SSH_VER##*.}
log "Checking which key exchange methods this OpenSSH supports..."
# Cosmetic only -- every decision below comes from `ssh -Q kex`, not this. A build
# whose banner does not match (OpenSSH_for_Windows_9.5p2, vendor forks) must not
# abort the run: without the guard, grep's non-match fails the pipeline under
# pipefail and set -e kills the script here with no message at all.
SSH_VER=$(ssh -V 2>&1 | grep -oE 'OpenSSH_[0-9]+[.][0-9]+' | head -1 | sed 's/OpenSSH_//' || true)
# OpenSSH 9.0+ has sntrup761x25519-sha512; 9.9+ adds mlkem768x25519-sha256.
HAS_MLKEM=0
HAS_SNTRUP=0
[[ $SSH_MAJOR -gt 9 || ( $SSH_MAJOR -eq 9 && $SSH_MINOR -ge 0 ) ]] && HAS_SNTRUP=1
[[ $SSH_MAJOR -gt 9 || ( $SSH_MAJOR -eq 9 && $SSH_MINOR -ge 9 ) ]] && HAS_MLKEM=1
[[ $HAS_SNTRUP -eq 1 || $HAS_MLKEM -eq 1 ]] \
|| die "OpenSSH ${SSH_VER} has no PQ KEX. Need >= 9.0. Upgrade the base OS first."
log "OpenSSH ${SSH_VER}: ML-KEM=${HAS_MLKEM} sntrup761=${HAS_SNTRUP}"
# Ask the binary what it supports rather than deriving it from the version -- the
# same algorithm has two spellings and guessing wrong either breaks sshd_config or
# locks out clients that implement it under the other name. See oslib.
KEX_PQ="$(ssh_kex_pq_list)"
KEX_LIST="$(ssh_kex_list)"
KEX_LIST=""
[[ $HAS_MLKEM -eq 1 ]] && KEX_LIST="mlkem768x25519-sha256"
[[ $HAS_SNTRUP -eq 1 ]] && KEX_LIST="${KEX_LIST:+$KEX_LIST,}sntrup761x25519-sha512"
if [[ -n "$KEX_PQ" ]]; then
if [[ "${SSH_ALLOW_CLASSIC_KEX:-0}" == "1" ]]; then
KEX_NOTE='# --- Key exchange: post-quantum hybrid + classical fallback ---
# curve25519-sha256 is offered for clients too old for any PQ method
# (SSH_ALLOW_CLASSIC_KEX=1). A session that negotiates it has NO store-now-
# decrypt-later protection -- drop the fallback once those clients are gone.'
warn "SSH_ALLOW_CLASSIC_KEX=1 -- offering curve25519-sha256 next to the PQ methods."
warn " Admits clients with no PQ KEX at all (Windows in-box ssh.exe), at the cost of"
warn " store-now-decrypt-later protection for any session that negotiates it."
else
KEX_NOTE='# --- Key exchange: post-quantum hybrid only ---
# Every classical-only method is rejected, which is what protects the session key
# against "store now, decrypt later". Clients with no PQ KEX cannot connect --
# re-run with SSH_ALLOW_CLASSIC_KEX=1 to also offer curve25519-sha256.'
fi
else
# No post-quantum method on this host at all. Decide on THAT, not on an empty
# list: without the opt-in the classical names are never collected, so an empty
# list here would otherwise be misreported as "no usable KEX".
[[ "${SSH_ALLOW_CLASSIC_KEX:-0}" == "1" ]] \
|| die "OpenSSH ${SSH_VER:-?} has no post-quantum KEX (needs >= 8.5 built with sntrup761). Re-run with SSH_ALLOW_CLASSIC_KEX=1 to accept classical-only."
[[ -n "$KEX_LIST" ]] \
|| die "OpenSSH ${SSH_VER:-?} reports no usable key exchange method at all ('ssh -Q kex' returned nothing)."
KEX_NOTE='# --- Key exchange: CLASSICAL ONLY ---
# This OpenSSH has no post-quantum method. No store-now-decrypt-later protection.'
warn "OpenSSH ${SSH_VER:-?} has no PQ KEX -- classical curve25519 only."
fi
log "OpenSSH ${SSH_VER:-?}: KexAlgorithms ${KEX_LIST}"
# ----------------------------------------------------------------------------
# 2. Install packages (OS-gated inside oslib)
# ----------------------------------------------------------------------------
log "Installing OpenSSH server + sshguard + iptables..."
install_openssh
install_bruteforce_protection
install_openssh || die "OpenSSH packages failed to install; cannot harden SSH. Fix the package error above, then re-run."
# sshguard is best-effort: a host where it can't install right now (e.g. EPEL
# momentarily unreachable) must still get the sshd_config hardening AND the login
# notifier -- not a silently half-configured box. Warn and press on.
install_bruteforce_protection \
|| warn "sshguard not installed; brute-force protection is OFF. Add it later with: dnf install -y epel-release sshguard. Continuing with the rest of the hardening."
# The external SFTP subsystem binary path differs per distro.
SFTP_PATH="$(sftp_server_path)"
@@ -171,9 +211,7 @@ PidFile /run/sshd.pid
# --- Host key: Ed25519 only ---
HostKey /etc/ssh/ssh_host_ed25519_key
# --- Post-quantum hybrid KEX only ---
# Anything not in this list (every classical-only KEX) is rejected, which is
# what protects the session key against "store now, decrypt later".
${KEX_NOTE}
KexAlgorithms ${KEX_LIST}
# --- Modern ciphers and MACs ---
@@ -281,6 +319,20 @@ EOF
# standalone boot hook. Otherwise fall back to the minimal init-agnostic boot
# hook that just (re)inserts the jump at every boot.
: "${ENABLE_FIREWALL:=1}"
# Proxmox (VE/PMG) is Debian underneath, but pve-firewall already owns the host
# ruleset -- harden-firewall.sh skips those hosts, so fall through to the boot
# hook and keep sshguard's jump. That jump is safe there: pve-firewall restores
# with --noflush and only ever flushes its own PVEFW-* chains, and the hook it
# adds to INPUT is APPENDED -- so a jump inserted with -I sits ahead of it and
# keeps getting first look at NEW connections, firewall enabled or not.
# FW_IGNORE_PVE=1 forces our firewall anyway.
if [[ "$ENABLE_FIREWALL" == "1" && "${FW_IGNORE_PVE:-0}" != "1" ]] && is_proxmox; then
warn "Proxmox detected -- pve-firewall owns the host firewall; installing only the sshguard jump hook."
# A host hardened before the skip existed still has a live DROP chain that only
# accepts the OLD port -- changing SSH_PORT here would lock you out of it.
[[ -x /usr/local/sbin/firewall-apply || -d /etc/firewall ]] && warn " This host still has an older harden-firewall.sh ruleset. Run 'harden-firewall.sh disable' FIRST, especially if you are changing SSH_PORT."
ENABLE_FIREWALL=0
fi
if [[ "$ENABLE_FIREWALL" == "1" && -f "$SCRIPT_DIR/harden-firewall.sh" ]]; then
log "Installing host firewall (deny-by-default INPUT; carries the sshguard jump)..."
SSH_PORT="$SSH_PORT" OPEN_PORTS="${OPEN_PORTS:-}" \
@@ -373,8 +425,7 @@ terminal -- before answering yes -- to verify the new keys, port, and PQ
KEX work. If something is wrong, this reload will end your current session.
Test in another terminal first:
ssh -i ~/.ssh/<your saved key> -p ${SSH_PORT} \\
-o KexAlgorithms=${KEX_LIST} root@<host>
ssh -i ~/.ssh/<your saved key> -p ${SSH_PORT} root@<host>
Reload sshd now? [y/N]
EOF
@@ -386,6 +437,8 @@ EOF
fi
log "Reloading ${SSHD_SVC}..."
svc_reload "$SSHD_SVC"
# Reload -- and on Alpine swap in the PAM sshd build if the running one
# predates this config (oslib explains why that can happen).
sshd_apply_config
log "Done. Your session, if any, should remain alive (reload preserves connections)."
log "Test from another machine before closing this session."
+16 -1
View File
@@ -123,5 +123,20 @@ tags="warning"
[ -n "${NTFY_REGION:-}" ] && tags="${tags},${NTFY_REGION}"
set -- "$@" -H "X-Tags: ${tags}"
curl "$@" -d "$body" "$NTFY_URL" >/dev/null 2>&1 || true
# Deliver. Failures are non-fatal -- a login must never be blocked by a notifier
# hiccup. Set NTFY_DEBUG=1 in the conf to log attempts + curl errors to
# /var/log/ssh-notify.log, so a silent failure (SELinux, egress, bad token, ...)
# leaves a trace instead of vanishing.
if [ "${NTFY_DEBUG:-0}" = "1" ]; then
log=/var/log/ssh-notify.log
printf '%s login user=%s rhost=%s -> %s\n' \
"$(date -u +%FT%TZ 2>/dev/null || echo)" "$user" "$rhost" "$NTFY_URL" >> "$log" 2>/dev/null || true
if curl "$@" -d "$body" "$NTFY_URL" >>"$log" 2>&1; then
echo " -> delivered" >> "$log" 2>/dev/null || true
else
echo " -> curl FAILED (exit $?)" >> "$log" 2>/dev/null || true
fi
else
curl "$@" -d "$body" "$NTFY_URL" >/dev/null 2>&1 || true
fi
exit 0
+253 -14
View File
@@ -65,6 +65,22 @@ os_detect() {
_require_detected() { [[ -n "${OS_FAMILY:-}" ]] || os_detect; }
# True on a Proxmox host that ships pve-firewall -- Proxmox VE and Proxmox Mail
# Gateway. Both are Debian underneath, so os_detect reports debian/systemd and
# every other helper here is correct for them; the ONE thing that is not is the
# firewall. pve-firewall owns the host ruleset, and while it leaves foreign rules
# alone (it restores with --noflush and touches only its own PVEFW-* chains), a
# deny-by-default INPUT chain underneath it overrides the accepts it RETURNs on
# and cuts the web UI, corosync and migration traffic -- and our boot-time
# iptables-restore would wipe its hook outright. harden-firewall.sh therefore
# skips these hosts -- see the Proxmox backend there for the full reasoning.
#
# Proxmox Backup Server ships no firewall of its own and is deliberately NOT
# matched: it is a plain Debian host as far as we are concerned.
is_proxmox() {
[[ -x /usr/sbin/pve-firewall || -x /usr/bin/pveversion || -d /etc/pve/nodes ]]
}
# ============================================================================
# Packages
# ============================================================================
@@ -136,6 +152,14 @@ svc_enable() { # enable at boot
esac
}
svc_disable() { # stop it being started at boot (the mirror of svc_enable)
_require_detected
case "$INIT_SYSTEM" in
openrc) rc-update del "$1" default >/dev/null 2>&1 || true ;;
systemd) systemctl disable "$1" >/dev/null 2>&1 || true ;;
esac
}
svc_start() { _require_detected; [[ "$INIT_SYSTEM" == openrc ]] && rc-service "$1" start || systemctl start "$1"; }
svc_restart() { _require_detected; [[ "$INIT_SYSTEM" == openrc ]] && { rc-service "$1" restart || rc-service "$1" start; } || systemctl restart "$1"; }
svc_reload() { _require_detected; [[ "$INIT_SYSTEM" == openrc ]] && { rc-service "$1" reload || rc-service "$1" restart; } || { systemctl reload "$1" || systemctl restart "$1"; }; }
@@ -148,6 +172,66 @@ sshd_service() {
[[ "$OS_FAMILY" == debian ]] && echo ssh || echo sshd
}
# ============================================================================
# SSH key exchange -- post-quantum hybrid, with an opt-in classical fallback.
# ============================================================================
# Build the KexAlgorithms list from what THIS OpenSSH build actually supports,
# asked via `ssh -Q kex`, instead of inferring it from a version number. The
# same algorithm has two spellings -- OpenSSH used
# sntrup761x25519-sha512@openssh.com before the method was standardised and
# sntrup761x25519-sha512 after -- and guessing wrong breaks in both directions:
# a name the local sshd does not know is a fatal sshd_config error, while a name
# the CLIENT does not know is an "Unable to negotiate ... no matching key
# exchange method found" lockout even though both ends implement the algorithm.
# Offering every spelling this host supports costs nothing and avoids both.
#
# SSH_ALLOW_CLASSIC_KEX=1 additionally offers curve25519-sha256 (and its older
# @libssh.org spelling): ordinary X25519 ECDH, secure against a classical
# attacker but with NO post-quantum protection. It exists for clients too old
# for any PQ method -- notably the ssh.exe bundled with Windows, which has none
# -- and the callers warn when it is on.
kex_supported() { # kex_supported <algorithm-name>
ssh -Q kex 2>/dev/null | grep -qxF "$1"
}
# The post-quantum hybrids this host supports, every spelling it has. Empty when
# this OpenSSH has none -- pre-8.5, or a build with sntrup761 compiled out (the
# Microsoft fork does exactly that: it needs C99 VLAs, which MSVC lacks).
ssh_kex_pq_list() {
local list="" k
for k in mlkem768x25519-sha256 \
sntrup761x25519-sha512 \
sntrup761x25519-sha512@openssh.com; do
if kex_supported "$k"; then list="${list:+$list,}$k"; fi
done
printf '%s\n' "$list"
}
# The classical fallback. Offered only when SSH_ALLOW_CLASSIC_KEX=1.
ssh_kex_classic_list() {
local list="" k
for k in curve25519-sha256 curve25519-sha256@libssh.org; do
if kex_supported "$k"; then list="${list:+$list,}$k"; fi
done
printf '%s\n' "$list"
}
# The KexAlgorithms value itself: PQ first, classical appended only on request.
ssh_kex_list() {
local pq classic=""
pq="$(ssh_kex_pq_list)"
if [[ "${SSH_ALLOW_CLASSIC_KEX:-0}" == "1" ]]; then classic="$(ssh_kex_classic_list)"; fi
if [[ -n "$pq" && -n "$classic" ]]; then
printf '%s,%s\n' "$pq" "$classic"
else
printf '%s\n' "${pq}${classic}"
fi
}
ssh_kex_has_pq() { # ssh_kex_has_pq <list>
case "$1" in *mlkem*|*sntrup*) return 0 ;; *) return 1 ;; esac
}
# ============================================================================
# SSH-specific paths
# ============================================================================
@@ -188,6 +272,77 @@ sshd_disable_keygen() {
fi
}
# ----------------------------------------------------------------------------
# Alpine keeps PAM support in a SEPARATE binary.
#
# openssh-server -> /usr/sbin/sshd (built WITHOUT PAM)
# openssh-server-pam -> /usr/sbin/sshd.pam (built WITH PAM)
#
# The OpenRC init picks between them in start_pre (checkconfig -> update_command:
# "sshd.pam if it is executable and the config says UsePAM yes"), and its
# reload/stop match the running process with `start-stop-daemon --exec "$command"`.
# So a host that is ALREADY running /usr/sbin/sshd never swaps over on its own:
# reload signals a process that doesn't match (nothing happens), stop matches
# nothing, and a follow-up start hits "address already in use". The daemon keeps
# serving without PAM -- which means /etc/pam.d/sshd, and therefore the pam_exec
# login notifier, is silently never consulted. Every other distro builds PAM into
# the one sshd binary, so this is Alpine-only.
# ----------------------------------------------------------------------------
sshd_wanted_binary() { # echo the sshd binary this host's config should be running
_require_detected
if [[ "$OS_FAMILY" == alpine ]] && [[ -x /usr/sbin/sshd.pam ]] && grep -qiE '^[[:space:]]*UsePAM[[:space:]]+yes' /etc/ssh/sshd_config 2>/dev/null; then
echo /usr/sbin/sshd.pam
else
echo /usr/sbin/sshd
fi
}
sshd_running_binary() { # echo the executable behind the running sshd master ('' if unknown)
local pid="" p
if [[ -r /run/sshd.pid ]]; then pid="$(cat /run/sshd.pid 2>/dev/null || true)"; fi
if [[ -z "$pid" ]]; then
for p in sshd.pam sshd; do
pid="$(pgrep -x "$p" 2>/dev/null | head -n1 || true)"
[[ -n "$pid" ]] && break
done
fi
[[ -n "$pid" ]] || return 0
readlink -f "/proc/$pid/exe" 2>/dev/null || true
}
# Apply a freshly written sshd_config. Normally a reload (keeps connections).
# On Alpine, when the running binary is not the one the config calls for, the
# service is stopped by pidfile and started again so the PAM build takes over --
# established sessions are separate processes and survive; only the listener
# blinks.
sshd_apply_config() {
_require_detected
local svc; svc="$(sshd_service)"
if [[ "$OS_FAMILY" == alpine ]]; then
local want run pid i
want="$(sshd_wanted_binary)"
run="$(sshd_running_binary)"
if [[ -n "$run" && "$run" != "$want" ]]; then
_warn "sshd is running $run, but this config needs $want -- restarting to swap it in."
_warn "(Existing SSH sessions survive; the listener is down for about a second.)"
pid=""
if [[ -r /run/sshd.pid ]]; then pid="$(cat /run/sshd.pid 2>/dev/null || true)"; fi
rc-service "$svc" stop >/dev/null 2>&1 || true
if [[ -n "$pid" ]]; then
kill "$pid" 2>/dev/null || true
i=0
while kill -0 "$pid" 2>/dev/null && [[ "$i" -lt 10 ]]; do sleep 1; i=$((i + 1)); done
kill -9 "$pid" 2>/dev/null || true
fi
rm -f /run/sshd.pid
svc_start "$svc" || _die "sshd failed to start as $want. Check: rc-service $svc start"
_log "sshd restarted as $(sshd_running_binary)."
return 0
fi
fi
svc_reload "$svc"
}
# ============================================================================
# Users & groups (busybox adduser/addgroup vs shadow useradd/groupadd)
# ============================================================================
@@ -286,6 +441,7 @@ install_boot_hook() { # install_boot_hook <name> <path-to-script>
rc-update add local default >/dev/null 2>&1 || true
"/etc/local.d/${name}.start" || true ;;
systemd)
install -d -m 0755 /usr/local/sbin
install -m 0755 "$src" "/usr/local/sbin/${name}"
cat > "/etc/systemd/system/${name}.service" <<UNIT
[Unit]
@@ -316,6 +472,8 @@ install_daily_job() { # install_daily_job <name> <script-src> [run-args...]
local args="$*"
# The job scripts use bash; ensure it's present (Alpine images often lack it).
command -v bash >/dev/null 2>&1 || pkg_install bash || true
# Alpine's baselayout ships /usr/local/{bin,lib,share} but NOT sbin.
install -d -m 0755 /usr/local/sbin
install -m 0755 "$src" "/usr/local/sbin/$name"
# Co-install oslib.sh so a script that sources it still works standalone.
local srcdir; srcdir="$(dirname "$src")"
@@ -397,20 +555,34 @@ install_openssh() {
fi
local sftp_pkg; sftp_pkg="$(pkg_name sftp-server)"
# shellcheck disable=SC2046
pkg_install $(pkg_name openssh-server) $(pkg_name openssh-client) ${sftp_pkg:+$sftp_pkg}
# Alpine needs linux-pam present for the PAM server build.
[[ "$OS_FAMILY" == alpine ]] && pkg_install linux-pam openrc
pkg_install $(pkg_name openssh-server) $(pkg_name openssh-client) ${sftp_pkg:+$sftp_pkg} || return 1
# Alpine needs linux-pam present for the PAM server build. Use an if-block,
# NOT `[[ ... ]] && ...`: as the LAST statement, that trailing test makes the
# whole function exit 1 on every non-Alpine OS (a false `[[ ]]` returns 1) --
# harmless to a bare call under `set -e`, but a caller guarding with `|| die`
# reads it as an OpenSSH install failure. The `|| return 1` above still
# surfaces a real package failure.
if [[ "$OS_FAMILY" == alpine ]]; then
pkg_install linux-pam openrc
fi
}
# Install sshguard + an iptables firewall backend. On RHEL/Alma sshguard lives
# in EPEL, so enable that first.
# in EPEL, so enable that first. The iptables backend is installed best-effort
# FIRST (it's usually already present as iptables-nft), then sshguard, and the
# function RETURNS sshguard's status -- so a caller can treat a sshguard miss
# (e.g. EPEL momentarily unreachable) as non-fatal and still apply the rest of
# the hardening instead of aborting the whole run.
install_bruteforce_protection() {
_require_detected
case "$OS_FAMILY" in
alpine) pkg_install sshguard iptables ip6tables ;;
debian) pkg_install sshguard iptables ;;
rhel) pkg_install epel-release || true
pkg_install sshguard iptables ;;
alpine) pkg_install iptables ip6tables || true
pkg_install sshguard ;;
debian) pkg_install iptables || true
pkg_install sshguard ;;
rhel) pkg_install iptables || true # el9+: provided by iptables-nft
pkg_install epel-release || true # sshguard lives in EPEL
pkg_install sshguard ;;
esac
}
@@ -455,16 +627,45 @@ fw_save_cmd() {
esac
}
# The service(s) that restore the saved ruleset at boot, and the file(s) they
# restore from -- the inverse of fw_save_cmd. Undoing the firewall on a host that
# should never have had it means disabling the former and clearing the latter
# (harden-firewall.sh's Proxmox `disable`), so name them once, here.
fw_restore_services() {
_require_detected
case "$OS_FAMILY" in
alpine) echo "iptables ip6tables" ;;
debian) echo "netfilter-persistent" ;;
rhel) echo "iptables ip6tables" ;;
esac
}
fw_saved_files() {
_require_detected
case "$OS_FAMILY" in
alpine) echo "/etc/iptables/rules-save /etc/iptables/rules6-save" ;;
debian) echo "/etc/iptables/rules.v4 /etc/iptables/rules.v6" ;;
rhel) echo "/etc/sysconfig/iptables /etc/sysconfig/ip6tables" ;;
esac
}
# Enable the family's native boot-time restore service(s). Rules are already
# live when this runs, so we only need them re-applied on the NEXT boot --
# enable, don't start.
fw_enable_restore() {
local s
for s in $(fw_restore_services); do svc_enable "$s"; done
}
# Install + enable firewalld (the native RHEL/Alma firewall). harden-firewall.sh
# uses this on the rhel family instead of raw iptables; sshguard then blocks via
# the sshguard-firewalld backend (no INPUT->sshguard jump, no boot hook).
install_firewalld() {
_require_detected
case "$OS_FAMILY" in
alpine) svc_enable iptables; svc_enable ip6tables ;;
debian) svc_enable netfilter-persistent ;;
rhel) svc_enable iptables; svc_enable ip6tables ;;
esac
command -v firewall-cmd >/dev/null 2>&1 || pkg_install firewalld
svc_enable firewalld
# Must be running before we push --permanent rules and --reload.
firewall-cmd --state >/dev/null 2>&1 || svc_start firewalld || true
}
# ============================================================================
@@ -538,6 +739,8 @@ NTFY_PRIORITY="${NTFY_PRIORITY:-min}"
NTFY_REGION="${NTFY_REGION:-}"
NOTIFY_GROUPS="${NOTIFY_GROUPS:-}"
NOTIFY_PRIORITY_MAP="${NOTIFY_PRIORITY_MAP:-}"
# Set to 1 to log every delivery attempt (and curl errors) to /var/log/ssh-notify.log.
NTFY_DEBUG="${NTFY_DEBUG:-0}"
CONF
)
chmod 600 /etc/ssh-notify.conf
@@ -556,11 +759,47 @@ CONF
_warn "$pam not found; add this line to your sshd PAM stack manually:"
_warn " $line"
fi
# Verify the hook actually landed and report loudly. A notifier that fails to
# install silently is worse than none -- you'd believe logins are watched
# when they aren't (exactly the trap that hid this on the first Alma run).
# "Wired into /etc/pam.d/sshd" is not the same as "will fire": on Alpine the
# PAM stack is only read when sshd IS the PAM build (see sshd_wanted_binary).
# Check that too, or we'd report ACTIVE for a hook that can never run.
if [[ "${OS_FAMILY:-}" == alpine ]]; then
if [[ ! -x /usr/sbin/sshd.pam ]]; then
_warn "Alpine: /usr/sbin/sshd.pam is missing, so /etc/pam.d/sshd is never read and this notifier cannot fire."
_warn " Fix: apk add openssh-server-pam && rc-service $(sshd_service) restart"
elif ! grep -qiE '^[[:space:]]*UsePAM[[:space:]]+yes' /etc/ssh/sshd_config 2>/dev/null; then
_warn "Alpine: sshd_config lacks 'UsePAM yes', so the PAM build is never selected and this notifier cannot fire."
fi
fi
if [[ -x /opt/scripts/ntfy-ssh-login.sh ]] \
&& grep -qF '/opt/scripts/ntfy-ssh-login.sh' "$pam" 2>/dev/null; then
_log "Login notifier ACTIVE -> ${NTFY_URL:-<NTFY_URL unset!>}"
return 0
fi
_warn "Login notifier did NOT fully install (script or pam hook missing)."
return 1
}
# Locate the sshguard iptables backend binary (path varies by packaging).
# Locate the sshguard firewall-backend binary (path varies by packaging).
# On RHEL/Alma we run firewalld, so prefer sshguard's firewalld backend
# (from the sshguard-firewalld package) -- blocks land in firewalld and there
# is no iptables INPUT->sshguard jump to maintain.
sshguard_backend() {
_require_detected
local c
if [[ "${OS_FAMILY:-}" == rhel ]]; then
for c in /usr/libexec/sshguard/sshg-fw-firewalld \
/usr/lib/sshguard/sshg-fw-firewalld; do
[[ -x "$c" ]] && { echo "$c"; return; }
done
# firewalld backend expected but not found yet; name it anyway so
# sshguard.conf points at the right binary once the package is in.
echo /usr/libexec/sshguard/sshg-fw-firewalld
return
fi
for c in /usr/libexec/sshguard/sshg-fw-iptables \
/usr/libexec/sshg-fw-iptables \
/usr/lib/sshguard/sshg-fw-iptables \