Files
automations/deployments/ergo/ergoctl
T
57_WolveandClaude Opus 5 85d945ccde feat(ergo): persistent message history, plus deployment-readiness audit fixes
Ergo keeps history in RAM by default, so it is lost on every restart --
including the ones this stack's updater performs. HISTORY now selects a backend
at first deploy: sqlite (default, a file beside ircd.db, no extra container),
postgres (a pinned container via docker-compose.postgres.yml, loopback-only,
POSTGRES_PASSWORD_FILE so no secret lands in .env), or off. Both SQL backends
need Ergo 2.18.0+, which deploy.sh enforces. HISTORY_EXPIRE (default 30d) sets
retention, because upstream's 1w expire-time DELETES from persistent storage --
persistence with the shipped default would buy only a week.

Ergo opens the history backend only at startup and, unlike MySQL, has no
"after launching the server" guard for sqlite/postgresql: a rehash that enables
one reports success and then silently discards every message. So `ergoctl
history` restarts, and `ergoctl edit` detects a backend change and refuses to
apply it by rehash. The config validator runs with no network by design, so it
neutralises postgresql in its copy and reachability is checked separately.

Audit fixes (six lenses over first-deploy, runtime, lifecycle, security, docs):

- SECURITY (blocker): ircd/ is writable by the container uid while the 15-minute
  cert sync and every ergoctl config edit run as root. cp/install/> follow a
  symlink, so code execution in Ergo could redirect a root write onto any host
  file. All such writes now stage under $STACK_DIR (0700 root) and land via mv
  (rename(2) replaces a symlink instead of following it); root reads use cp -P
  or refuse. Regression-tested with real symlinks.
- deploy.sh no longer swaps the image when ERGO_TAG changes on a deployed stack
  (that bypassed pre-flight, snapshot, user warning and rollback); it points at
  `TARGET_VERSION=... ergoctl update update`.
- restore was lossy and could lock you out: it now snapshots the current config
  as well as the database, restores message history, re-points the postgres
  password and re-hashes the local admin password into the restored config, and
  validates before starting.
- `install -d -m` re-modes existing directories: `ergoctl backup /var/backups`
  no longer chmods it 0700 root, and the updater no longer re-modes /var/log
  (0775 root:syslog on Debian, which rsyslog needs).
- The admin oper password is printed only on a first run at a TTY, so it stays
  out of cloud-init serial-console logs.
- A failed update is remembered, so the daily job stops repeating a disruptive
  warn/stop/swap/roll-back cycle every night; postgres readiness gates an update
  that would otherwise stop a healthy server it cannot restart.
- certsync no longer sends "TLS recovered" for outcomes that synced nothing.
- ergoctl history writes .env only after the restart is healthy, and returns 0.
- CR/LF is stripped before IRC framing, so a multi-line argument cannot inject a
  second command; the ntfy token moves out of curl's argv.
- ufw/firewalld are additive, so 6667 is now explicitly revoked when PLAINTEXT=0.
- build.sh refuses to build a deploy.sh whose archive is missing a file the
  script reads -- the failure mode that would have shipped a stack aborting on
  every host.

Docs corrected against the code throughout, including retention, the pre-connect
account-registration default, encrypted-restore (AGE_IDENTITY), what a re-run
really does to .env, and what the update log does and does not contain.

Verified locally: all six suites pass (config render for each backend against the
real 2.19.1 template, yaml/oper/version/env helpers, the IRC client against a
fake server, and the audit fixes including the symlink escalation). Still not
exercised on a Docker host: the containers, ACME issuance, cert sync and
PostgreSQL itself.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-02 14:32:37 -05:00

607 lines
33 KiB
Bash

#!/usr/bin/env bash
#
# ergoctl -- administer the Ergo IRC stack from the host.
#
# Installed by deploy.sh as $STACK_DIR/ergoctl with a wrapper at
# /usr/local/bin/ergoctl. Talks to Ergo over the loopback plaintext listener
# (127.0.0.1:6667) as the 'admin' oper, using the password in secrets/admin.pass.
# Config changes are validated in a throwaway container (ergo run --smoke on a
# copy) and applied with REHASH; on failure the previous file is restored.
#
# Run `ergoctl help` for the command list.
set -euo pipefail
SELF_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
log() { printf '\033[1;32m[+]\033[0m %s\n' "$*"; }
warn() { printf '\033[1;33m[!]\033[0m %s\n' "$*" >&2; }
die() { printf '\033[1;31m[x]\033[0m %s\n' "$*" >&2; exit 1; }
# shellcheck source=ergolib.sh
. "$SELF_DIR/ergolib.sh"
load_conf
: "${STACK_DIR:=$SELF_DIR}"
ergo_set_paths
CONF="$IRCD_DIR/ircd.yaml"
MOTD="$IRCD_DIR/ergo.motd"
usage() {
cat <<EOF
ergoctl -- administer the Ergo IRC server (stack: $STACK_DIR)
Status & logs
status containers, versions, TLS cert, user count, schedule
users LUSERS summary
logs [-f] [N] ergo log (health-probe noise filtered; default last 100)
version running and pinned version, latest release
cert [show|sync] cert on disk vs served on :6697; 'sync' pulls Caddy's now
Configuration (validated, then REHASH; reverted on failure)
edit \$EDITOR ircd/ircd.yaml
motd \$EDITOR ircd/ergo.motd
rehash reload config, MOTD and TLS certs (no disconnects)
restart | stop | start the ergo container (restart drops every user)
caddy-restart apply Caddyfile / conf.d changes
debug on|off switch the log level (info <-> debug)
history [sqlite|postgres|off]
show, or switch, the persistent-history backend.
Switching does NOT migrate existing messages and
needs a restart (Ergo fixes the backend at startup).
Operators
oper list
oper add <name> [class] class: chat-moderator (default) | server-admin; prints the password once
oper passwd <name> new random password for an oper
oper rm <name>
passwd [show|rotate] the 'admin' oper password (secrets/admin.pass)
Moderation (as the admin oper)
announce <text> NOTICE to all users
kill <nick> [reason]
ban add <ip|cidr|nick!user@host|account> [duration] [reason] e.g. ban add 203.0.113.9 1d spam
(a bare name is treated as an ACCOUNT to suspend, not a
connected nick -- use 'ban info <nick>' to get their IP first)
ban del <target> | ban list | ban info <target>
defcon [1-5] show or set the DEFCON level
cmd <raw IRC line> anything else, e.g. cmd NS SAREGISTER alice hunter2
(replies are printed; control characters stripped)
Data
backup [--live] [dir] tar.gz of ircd.yaml, ircd.db, ergo.motd and the message
history (SQLite file, or a pg_dump). Stops Ergo briefly
for a consistent copy; --live = crash-consistent.
Encrypted with age when $STACK_DIR/age-recipients.txt exists.
restore <file> restore a backup: stops Ergo, keeps the current config and
database as backups/*.pre-restore.<ts>, reconciles this
host's credentials, validates, then restarts.
Encrypted backup: AGE_IDENTITY=/path/to/age.key ergoctl restore <file>.age
Updates (update.sh)
update [check|update|run|certsync|caddy|install|uninstall]
shell sh inside the ergo container
EOF
}
require_root() { [[ $EUID -eq 0 ]] || die "Run as root."; }
require_stack() { [[ -f "$STACK_DIR/docker-compose.yml" && -f "$CONF" ]] || die "No deployed stack at $STACK_DIR (set STACK_DIR)."; }
# Validate the current ./ircd with the running image; on failure print the log.
validate_or_fail() { # <what>
if ergo_validate_config; then rm -f "$VALIDATE_LOG"; return 0; fi
warn "$1 did not load in $(ergo_image):"
tail -n 15 "$VALIDATE_LOG" >&2 || true
rm -f "$VALIDATE_LOG"
return 1
}
# Apply a config change: validate, then REHASH; revert the file(s) on failure.
# $1 = description; $2.. = files that were changed (each has a .last-good copy).
apply_change() {
local what="$1"; shift
local f rc=0
if ! validate_or_fail "$what"; then
for f in "$@"; do [[ -f "$f.last-good" ]] && cat "$f.last-good" > "$f"; done
die "$what rejected; previous version restored."
fi
if ! ergo_running; then
log "$what saved (Ergo is not running; it will load on start)."
for f in "$@"; do rm -f "$f.last-good"; done
return 0
fi
# Ergo opens the history backend ONLY at startup. Unlike MySQL, sqlite and
# postgresql have no "after launching the server" guard, so a rehash that
# turns one on reports success and then silently discards every message.
# Catch that here rather than letting the operator believe it worked.
for f in "$@"; do
[[ "$f" == "${CONF:-}" && -f "$f.last-good" ]] || continue
if [[ "$(history_backend "$f")" != "$(history_backend "$f.last-good")" ]]; then
warn "This change switches the message-history backend, which Ergo only reads at startup."
warn "A rehash would report success and silently do nothing, so a restart is required."
if [[ "${ERGOCTL_ASSUME_YES:-0}" != 1 ]]; then
read -r -p "Restart Ergo now (disconnects every user)? [y/N] " ans
[[ "${ans,,}" == y* ]] || { cat "$f.last-good" > "$f"; die "Reverted; nothing changed. Use 'ergoctl history <backend>' to switch it properly."; }
fi
if ! validate_or_fail "$what"; then cat "$f.last-good" > "$f"; die "$what rejected; previous version restored."; fi
irc_raw --oper --quiet 1 -- 'NOTICE $$* :Server restarting to change message-history storage.' >/dev/null 2>&1 || true
dc up -d --no-deps ergo >/dev/null 2>&1 || dc restart ergo >/dev/null 2>&1 || true
if wait_healthy ergo 120 && irc_probe; then
rm -f "$f.last-good"; log "$what applied (Ergo restarted)."; return 0
fi
dc logs --tail 30 ergo >&2 || true
cat "$f.last-good" > "$f"
die "Ergo is not healthy after the restart; config reverted. Run 'ergoctl restart'."
fi
done
# Take the status from the call itself: `$?` after an `if` whose condition
# failed and that has no else branch is 0, which would make the SIGHUP
# fallback below unreachable.
irc_rehash || rc=$?
if (( rc == 0 )); then
log "$what applied (rehash complete)."
for f in "$@"; do rm -f "$f.last-good"; done
return 0
fi
# Ergo refuses to rehash settings that are fixed at startup. It phrases them
# all as "... after launching the server", so match that, not one variant.
if [[ "$REHASH_MSG" == *"after launching the server"* || "$REHASH_MSG" == *"cannot be changed after launching"* ]]; then
warn "Rehash refused: $REHASH_MSG"
warn "The file is valid and saved; apply it with: ergoctl restart (disconnects everyone)"
for f in "$@"; do rm -f "$f.last-good"; done
return 0
fi
if (( rc == 2 )); then
warn "Could not REHASH over IRC ($REHASH_MSG); sending SIGHUP instead."
dc kill -s HUP ergo >/dev/null 2>&1 || true
sleep 2
if dc logs --since 15s ergo 2>/dev/null | grep -qiE 'rehash.*(complete|success)'; then
log "$what applied (rehash via SIGHUP)."; for f in "$@"; do rm -f "$f.last-good"; done; return 0
fi
fi
for f in "$@"; do [[ -f "$f.last-good" ]] && cat "$f.last-good" > "$f"; done
die "Rehash failed: ${REHASH_MSG}. Previous version restored (Ergo kept running on the old config)."
}
snapshot_file() { cp -p "$1" "$1.last-good"; }
edit_file() { # <file> <what>
local f="$1" what="$2" tmp
require_root; require_stack
tmp="$(mktemp)"; cat "$f" > "$tmp"
"${EDITOR:-vi}" "$tmp"
if cmp -s "$f" "$tmp"; then rm -f "$tmp"; log "No changes."; return 0; fi
snapshot_file "$f"
cat "$tmp" > "$f"; rm -f "$tmp"
apply_change "$what" "$f"
}
oper_name_ok() { [[ "$1" =~ ^[a-z][a-z0-9_-]{1,31}$ ]]; }
cmd_status() {
require_stack
echo "Stack: $STACK_DIR (domain $(ergo_domain), network $(env_get NETWORK_NAME))"
dc ps 2>/dev/null || true
echo
local st cur latest
st="$(svc_state ergo || true)"; cur="$(ergo_version_running || true)"
latest="$(normver "$(ergo_latest_tag || true)")"
printf 'Ergo: %s | running %s | pinned %s | latest %s\n' "${st:-not created}" "${cur:-?}" "$(ergo_tag)" "${latest:-?}"
if [[ "$(env_get PLAINTEXT)" == "1" ]]; then echo " PUBLIC PLAINTEXT LISTENER on :6667 (PLAINTEXT=1)"; fi
local hb; hb="$(history_backend)"
case "$hb" in
postgres) pg_check || true; printf 'History: persistent, PostgreSQL -- %s\n' "$PG_CHECK_MSG" ;;
sqlite) printf 'History: persistent, SQLite (%s)\n' "$IRCD_DIR/$(yaml_get_nested "$CONF" datastore sqlite database-path)" ;;
*) printf 'History: RAM only -- lost on every restart\n' ;;
esac
cmd_cert show
if ergo_running; then
echo; cmd_users || true
fi
echo
printf 'Policy: %s | schedule: ' "${UPDATE_POLICY:-$(env_get UPDATE_POLICY)}"
if [[ -x /etc/periodic/15min/ergo-certsync || -f /etc/systemd/system/ergo-certsync.timer ]]; then printf 'certsync 15min'; else printf 'certsync MISSING'; fi
if [[ -x /etc/periodic/daily/ergo-update || -f /etc/systemd/system/ergo-update.timer ]]; then printf ', update daily\n'; else printf ', update NOT scheduled\n'; fi
local s; s="$(state_get certsync)"; [[ -z "$s" ]] || echo "certsync: LAST FAILURE: $s"
}
cmd_users() {
local out
out="$(irc_raw --quiet 1 -- LUSERS)" || die "Could not talk to Ergo on 127.0.0.1:6667 (is it running?)."
# ":server 251 nick :There are N users..." -> "There are N users..."
printf '%s\n' "$out" | awk '$2=="251"||$2=="252"||$2=="254"||$2=="265"||$2=="266" {
sub(/^:[^ ]+ [0-9]+ [^ ]+ /, ""); sub(/^:/, ""); sub(/ :/, " "); print " " $0 }'
}
cmd_logs() {
require_stack
local follow="" n=100
while [[ $# -gt 0 ]]; do
case "$1" in -f|--follow) follow="-f" ;; *[0-9]*) n="$1" ;; esac; shift
done
# Each healthcheck probe logs a 'Client connecting: real IP 127.0.0.1' line; hide those.
dc logs --tail "$n" $follow ergo 2>&1 | grep -vE 'real IP 127\.0\.0\.1, proxied IP <nil>|Preexisting TLS cert' || true
}
cmd_cert() {
require_stack
local f="$IRCD_DIR/fullchain.pem" served
case "${1:-show}" in
sync) require_root; bash "$STACK_DIR/update.sh" certsync ;;
*)
if [[ ! -f "$f" ]]; then echo "TLS: no cert in ircd/ yet"; return 0; fi
if cert_is_selfsigned "$f"; then
echo "TLS: SELF-SIGNED (Caddy has not issued a cert yet -- DNS/80/443?) expires $(cert_enddate "$f")"
else
echo "TLS: $(cert_subject_cn "$f") expires $(cert_enddate "$f") (from Caddy)"
fi
if ergo_running; then
served="$(served_fingerprint || true)"
if [[ -z "$served" ]]; then echo " :6697 did not complete a TLS handshake"
elif [[ "$served" == "$(cert_fingerprint "$f")" ]]; then echo " :6697 serves this cert"
else echo " :6697 serves a DIFFERENT cert -- run: ergoctl rehash"; fi
fi ;;
esac
}
cmd_oper() {
require_root; require_stack
local sub="${1:-list}" name="${2:-}" class="${3:-chat-moderator}" pw hash
case "$sub" in
list) yaml_oper_list "$CONF" | awk '{printf " %-20s %s\n", $1, $2}' ;;
add)
[[ -n "$name" ]] || die "Usage: ergoctl oper add <name> [chat-moderator|server-admin]"
oper_name_ok "$name" || die "Oper names: lowercase letters, digits, _ - (2-32 chars)."
yaml_oper_has "$CONF" "$name" && die "Oper '$name' already exists."
grep -qE "^ \"?${class}\"?:" "$CONF" || warn "Class '$class' is not defined under oper-classes in ircd.yaml -- the rehash will fail unless you add it."
pw="$(ergo_random_password 24)"
hash="$(ergo_genpasswd "$(ergo_image)" "$pw")" || die "genpasswd failed."
snapshot_file "$CONF"
yaml_oper_add "$CONF" "$name" "$class" "$hash" || { cat "$CONF.last-good" > "$CONF"; die "Could not insert the oper block."; }
apply_change "oper '$name'" "$CONF"
echo
echo " Oper '$name' ($class) added. Log in with: /OPER $name $pw"
echo " (shown once; rotate with: ergoctl oper passwd $name)" ;;
passwd)
[[ -n "$name" ]] || die "Usage: ergoctl oper passwd <name>"
yaml_oper_has "$CONF" "$name" || die "No oper '$name' in ircd.yaml."
pw="$(ergo_random_password 24)"
hash="$(ergo_genpasswd "$(ergo_image)" "$pw")" || die "genpasswd failed."
snapshot_file "$CONF"
yaml_oper_set_password "$CONF" "$name" "$hash" || { cat "$CONF.last-good" > "$CONF"; die "Oper '$name' has no password line to replace."; }
apply_change "oper '$name' password" "$CONF"
if [[ "$name" == "admin" ]]; then ( umask 077; printf '%s\n' "$pw" > "$ADMIN_PASS_FILE" ); fi
echo
echo " New password for '$name': /OPER $name $pw" ;;
rm|remove|del)
[[ -n "$name" ]] || die "Usage: ergoctl oper rm <name>"
[[ "$name" != "admin" ]] || die "Refusing to remove 'admin' (ergoctl uses it). Rotate its password instead."
yaml_oper_has "$CONF" "$name" || die "No oper '$name' in ircd.yaml."
snapshot_file "$CONF"
yaml_oper_rm "$CONF" "$name" || { cat "$CONF.last-good" > "$CONF"; die "Could not remove the oper block."; }
apply_change "removal of oper '$name'" "$CONF" ;;
*) die "Usage: ergoctl oper list|add|passwd|rm" ;;
esac
}
cmd_passwd() {
require_root; require_stack
case "${1:-show}" in
show) [[ -f "$ADMIN_PASS_FILE" ]] || die "No $ADMIN_PASS_FILE -- run: ergoctl passwd rotate"
echo "/OPER admin $(head -n1 "$ADMIN_PASS_FILE")" ;;
rotate) cmd_oper passwd admin
if irc_raw --oper --quiet 1 -- >/dev/null 2>&1; then log "Verified: the new admin password works."; else warn "Could not verify the new password over IRC (is Ergo running?)."; fi ;;
*) die "Usage: ergoctl passwd [show|rotate]" ;;
esac
}
cmd_cmd() { # raw line as oper
require_root; require_stack
[[ $# -gt 0 ]] || die "Usage: ergoctl cmd <raw IRC line>"
local out rc=0
out="$(irc_raw --oper --quiet 2 -- "$*")" || rc=$?
case "$rc" in
0) ;;
3) die "OPER failed (wrong admin password? see secrets/admin.pass; the server logs details under type 'opers')." ;;
4) die "Cannot connect to 127.0.0.1:6667 -- is Ergo running?" ;;
*) die "Registration with the server failed (rc=$rc)." ;;
esac
# Drop the routine registration burst and show only what the command produced.
# The MODE filter is anchored to our own throwaway nick so that MODE replies
# caused by the command itself (e.g. SAMODE #chan +m) still print.
printf '%s\n' "$out" | grep -vE '^:[^ ]+ (00[1-5]|25[0-9]|26[56]|37[256]|422|381) |^:[^ ]+ MODE ergoctl[0-9]+ ' || true
}
cmd_backup() {
require_root; require_stack
local live=0 dest="$BACKUP_DIR" ts stage out was_running=0 users=""
while [[ $# -gt 0 ]]; do case "$1" in --live) live=1 ;; *) dest="$1" ;; esac; shift; done
# `install -d -m` also re-modes an EXISTING directory, so only create-and-mode
# our own default; never change the mode of a directory the operator named
# (`ergoctl backup /var/backups` must not make it 0700 root).
if [[ "$dest" == "$BACKUP_DIR" ]]; then install -d -m 0700 "$dest"
else [[ -d "$dest" ]] || install -d -m 0700 "$dest"; fi
# A recipients file means "encrypt these" -- refuse rather than silently
# writing account hashes and the oper hash out in the clear.
if [[ -f "$STACK_DIR/age-recipients.txt" ]] && ! command -v age >/dev/null 2>&1; then
die "$STACK_DIR/age-recipients.txt exists but 'age' is not installed, so the backup would be UNENCRYPTED. Install it (apk add age / apt install age / dnf install age) or move the recipients file away."
fi
ts="$(date -u +%Y%m%d-%H%M%S)"
stage="$(mktemp -d)"
if ergo_running; then
was_running=1
if [[ "$live" == 0 ]]; then
users="$(irc_raw --quiet 1 -- LUSERS 2>/dev/null | awk '$2=="251" {print}' | sed 's/.*:There are //' || true)"
log "Stopping Ergo for a consistent snapshot (${users:-users unknown})..."
irc_raw --oper --quiet 1 -- 'NOTICE $$* :Brief restart for a backup -- back in a few seconds.' >/dev/null 2>&1 || true
dc stop -t 15 ergo >/dev/null 2>&1 || true
else
warn "Live backup: ircd.db is copied while Ergo writes to it (crash-consistent; may miss the last second)."
fi
fi
# -P everywhere: ircd/ is writable by the container uid, so a symlink there
# must be copied as a link, never followed by this root process.
refuse_symlink "$CONF" "ircd/ircd.yaml" || die "Refusing to back up through a symlink."
cp -Pp "$CONF" "$stage/ircd.yaml"
[[ -f "$IRCD_DIR/ircd.db" ]] && cp -Pp "$IRCD_DIR/ircd.db" "$stage/ircd.db"
[[ -f "$MOTD" ]] && cp -Pp "$MOTD" "$stage/ergo.motd"
# Message history, whichever backend holds it.
local backend hf
backend="$(history_backend)"
case "$backend" in
sqlite)
hf="$IRCD_DIR/$(yaml_get_nested "$CONF" datastore sqlite database-path)"
# SQLite writes -wal/-shm siblings; with Ergo stopped they are already
# checkpointed, and on a --live backup we take them along.
for f in "$hf" "$hf-wal" "$hf-shm"; do
[[ -f "$f" ]] && cp -Pp "$f" "$stage/$(basename "$f")"
done ;;
postgres)
if pg_check; then
log "Dumping the PostgreSQL history database..."
# A file copy of a live PGDATA is not a valid backup; pg_dump is.
dc exec -T postgres pg_dump -U "$(env_get POSTGRES_USER)" -d "$(env_get POSTGRES_DB)" \
> "$stage/history.sql" 2>/dev/null || { rm -f "$stage/history.sql"; warn "pg_dump failed; the backup will NOT contain message history."; }
else
warn "PostgreSQL is not reachable ($PG_CHECK_MSG); the backup will NOT contain message history."
fi ;;
esac
cat > "$stage/meta" <<EOF
ERGO_TAG=$(ergo_tag)
ERGO_DOMAIN=$(ergo_domain)
HISTORY=$backend
DATE=$(date -u +%FT%TZ)
HOST=$(hostname -f 2>/dev/null || hostname)
CONSISTENT=$(( live == 0 ))
EOF
if [[ "$was_running" == 1 && "$live" == 0 ]]; then dc start ergo >/dev/null 2>&1 || true; fi
out="$dest/ergo-backup-${ts}.tar.gz"
( umask 077; tar -czf "$out" -C "$stage" . )
rm -rf "$stage"
if [[ -f "$STACK_DIR/age-recipients.txt" ]]; then
age -R "$STACK_DIR/age-recipients.txt" -o "$out.age" "$out" || { rm -f "$out.age"; die "age encryption failed; the plaintext tar is still at $out."; }
rm -f "$out"; out="$out.age"
fi
chmod 0600 "$out"
log "Backup written: $out ($(wc -c < "$out") bytes). Copy it off this host."
[[ "$out" == *.age ]] || warn "Unencrypted: it holds account hashes and the oper hash. Put an age public key in $STACK_DIR/age-recipients.txt (and install 'age') to encrypt future backups."
}
cmd_restore() {
require_root; require_stack
local file="${1:-}" stage tag
[[ -f "$file" ]] || die "Usage: ergoctl restore <backup.tar.gz[.age]>"
stage="$(mktemp -d)"
if [[ "$file" == *.age ]]; then
command -v age >/dev/null 2>&1 || die "age is required to decrypt $file."
age -d -i "${AGE_IDENTITY:?set AGE_IDENTITY=/path/to/age/key}" "$file" | tar -xzf - -C "$stage"
else
tar -xzf "$file" -C "$stage"
fi
[[ -f "$stage/ircd.yaml" && -f "$stage/ircd.db" ]] || { rm -rf "$stage"; die "Backup lacks ircd.yaml/ircd.db."; }
tag="$(grep '^ERGO_TAG=' "$stage/meta" 2>/dev/null | cut -d= -f2 || true)"
if [[ -n "$tag" ]] && ver_gt "$(normver "$tag")" "$(normver "$(ergo_tag)")"; then
warn "Backup was taken with Ergo $tag but $(ergo_tag) is pinned -- an older Ergo cannot read a newer database schema. Update first (ergoctl update update) or pin ERGO_TAG=$tag."
read -r -p "Continue anyway? [y/N] " ans; [[ "${ans,,}" == y* ]] || { rm -rf "$stage"; die "Aborted."; }
fi
log "Stopping Ergo and restoring from $file..."
dc stop -t 15 ergo >/dev/null 2>&1 || true
install -d -m 0700 "$BACKUP_DIR"
# Keep the CURRENT config and database, not just the database: if the restore
# turns out to be unloadable there has to be a way back.
local ts uid gid
ts="$(date -u +%Y%m%d%H%M%S)"; uid="$(ergo_uid)"; gid="$(ergo_gid)"
[[ -f "$CONF" ]] && cp -Pp "$CONF" "$BACKUP_DIR/ircd.yaml.pre-restore.$ts"
[[ -f "$IRCD_DIR/ircd.db" ]] && cp -Pp "$IRCD_DIR/ircd.db" "$BACKUP_DIR/ircd.db.pre-restore.$ts"
safe_replace "$stage/ircd.yaml" "$CONF" 0600 "$uid" "$gid"
safe_replace "$stage/ircd.db" "$IRCD_DIR/ircd.db" 0600 "$uid" "$gid"
[[ -f "$stage/ergo.motd" ]] && safe_replace "$stage/ergo.motd" "$MOTD" 0644 "$uid" "$gid"
# The backup carries the ORIGINAL host's credentials. Reconcile them with
# this host's, or the restored server is unreachable and unadministrable.
if [[ "$(history_backend)" == "postgres" && -f "$SECRETS_DIR/postgres.pass" ]]; then
log "Re-pointing datastore.postgresql at this host's database password..."
yaml_set_nested "$CONF" datastore postgresql password "\"$(cat "$SECRETS_DIR/postgres.pass")\"" \
|| warn "Could not update the postgresql password; Ergo may fail to reach its history database."
fi
if [[ -f "$ADMIN_PASS_FILE" ]]; then
local h
if h="$(ergo_genpasswd "$(ergo_image)" "$(head -n1 "$ADMIN_PASS_FILE")")"; then
yaml_oper_set_password "$CONF" admin "$h" \
&& log "Re-hashed this host's admin password into the restored config (ergoctl keeps working)." \
|| warn "Could not update the admin oper hash; 'ergoctl passwd rotate' will fix it."
else
warn "Could not hash the local admin password; the restored config keeps the backup's admin hash. If OPER fails, run 'ergoctl passwd rotate'."
fi
fi
# Message history
local f
for f in "$stage"/ergo_history.db*; do
[[ -f "$f" ]] || continue
install -m 0600 -o "$(ergo_uid)" -g "$(ergo_gid)" "$f" "$IRCD_DIR/$(basename "$f")"
done
if [[ -f "$stage/history.sql" ]]; then
log "Starting PostgreSQL and restoring the history dump..."
dc up -d --no-deps postgres >/dev/null 2>&1 || true
if wait_healthy postgres 120; then
dc exec -T postgres psql -U "$(env_get POSTGRES_USER)" -d "$(env_get POSTGRES_DB)" < "$stage/history.sql" >/dev/null 2>&1 \
|| warn "Restoring the history dump failed; Ergo will start with an empty history."
else
warn "PostgreSQL did not become healthy; message history was NOT restored."
fi
fi
rm -rf "$stage"
# Validate before starting, so a bad backup does not cost an outage.
if ! validate_or_fail "the restored config"; then
[[ -f "$BACKUP_DIR/ircd.yaml.pre-restore.$ts" ]] && cp -p "$BACKUP_DIR/ircd.yaml.pre-restore.$ts" "$CONF"
[[ -f "$BACKUP_DIR/ircd.db.pre-restore.$ts" ]] && cp -p "$BACKUP_DIR/ircd.db.pre-restore.$ts" "$IRCD_DIR/ircd.db"
dc up -d --no-deps ergo >/dev/null 2>&1 || true
die "The backup's config does not load; rolled back to what was here before (copies kept as $BACKUP_DIR/*.pre-restore.$ts)."
fi
dc up -d --no-deps ergo >/dev/null 2>&1 || true
if wait_healthy ergo 120 && irc_probe; then
log "Restored; Ergo is healthy."
log "Previous state kept at $BACKUP_DIR/{ircd.yaml,ircd.db}.pre-restore.$ts"
else
dc logs --tail 20 ergo >&2 || true
die "Ergo is not healthy after the restore. Roll back with: cp $BACKUP_DIR/ircd.yaml.pre-restore.$ts $CONF && cp $BACKUP_DIR/ircd.db.pre-restore.$ts $IRCD_DIR/ircd.db && ergoctl start"
fi
}
cmd_history() {
require_stack
local want="${1:-}" cur pw
cur="$(history_backend)"
if [[ -z "$want" ]]; then
echo "History backend: $cur"
case "$cur" in
sqlite) local f="$IRCD_DIR/$(yaml_get_nested "$CONF" datastore sqlite database-path)"
[[ -f "$f" ]] && echo " $f ($(wc -c < "$f") bytes)" || echo " (database not created yet)" ;;
postgres) pg_check && echo " $PG_CHECK_MSG" || echo " $PG_CHECK_MSG" ;;
off) echo " messages are kept in RAM only and lost on every restart" ;;
esac
echo " retention: history.restrictions.expire-time = $(yaml_get_nested "$CONF" history restrictions expire-time)"
return 0
fi
require_root
case "$want" in sqlite|postgres|off) ;; *) die "Usage: ergoctl history [sqlite|postgres|off]" ;; esac
[[ "$want" != "$cur" ]] || { log "History is already using '$cur'."; return 0; }
# .env and ircd.yaml must never disagree: deploy.sh reconciles them from .env
# on its next run, so a half-applied switch would let it tear down the wrong
# thing. Remember both and put them back on every failure path.
local prev_hist prev_compose
prev_hist="$(env_get HISTORY)"; prev_compose="$(env_get COMPOSE_FILE)"
_hist_revert() {
[[ -f "$CONF.last-good" ]] && cat "$CONF.last-good" > "$CONF"
[[ -n "$prev_hist" ]] && env_set HISTORY "$prev_hist"
[[ -n "$prev_compose" ]] && env_set COMPOSE_FILE "$prev_compose"
return 0
}
if [[ "$want" == "postgres" ]]; then
[[ -f "$STACK_DIR/docker-compose.postgres.yml" ]] || die "docker-compose.postgres.yml is missing; re-run deploy.sh to install it."
if [[ ! -f "$SECRETS_DIR/postgres.pass" ]]; then
log "Generating the PostgreSQL password..."
( umask 077; ergo_random_password 32 > "$SECRETS_DIR/postgres.pass" )
fi
chmod 0600 "$SECRETS_DIR/postgres.pass"; pw="$(cat "$SECRETS_DIR/postgres.pass")"
[[ -n "$(env_get POSTGRES_TAG)" ]] || env_set POSTGRES_TAG 17-alpine
[[ -n "$(env_get POSTGRES_USER)" ]] || env_set POSTGRES_USER ergo
[[ -n "$(env_get POSTGRES_DB)" ]] || env_set POSTGRES_DB ergo_history
[[ -n "$(env_get POSTGRES_PORT)" ]] || env_set POSTGRES_PORT 5432
env_set COMPOSE_FILE "docker-compose.yml:docker-compose.postgres.yml"
fi
warn "Switching the history backend does NOT migrate existing messages; the old store is left in place."
snapshot_file "$CONF"
# Turn every backend off first, then enable the requested one.
yaml_set_nested "$CONF" datastore sqlite enabled false || true
yaml_set_nested "$CONF" datastore postgresql enabled false || true
case "$want" in
sqlite) yaml_set_nested "$CONF" datastore sqlite enabled true || { cat "$CONF.last-good" > "$CONF"; die "Could not enable datastore.sqlite."; }
yaml_set_nested "$CONF" history persistent enabled true || true ;;
postgres) yaml_set_nested "$CONF" datastore postgresql enabled true || { cat "$CONF.last-good" > "$CONF"; die "Could not enable datastore.postgresql."; }
yaml_set_nested "$CONF" datastore postgresql host "\"127.0.0.1\"" || true
yaml_set_nested "$CONF" datastore postgresql port "$(env_get POSTGRES_PORT)" || true
yaml_set_nested "$CONF" datastore postgresql user "\"$(env_get POSTGRES_USER)\"" || true
yaml_set_nested "$CONF" datastore postgresql password "\"${pw}\"" || true
yaml_set_nested "$CONF" datastore postgresql history-database "\"$(env_get POSTGRES_DB)\"" || true
yaml_set_nested "$CONF" history persistent enabled true || true
log "Starting PostgreSQL..."
dc up -d --no-deps postgres >/dev/null 2>&1 || true
wait_healthy postgres 120 || { _hist_revert; die "PostgreSQL did not become healthy; config reverted."; } ;;
off) yaml_set_nested "$CONF" history persistent enabled false || true ;;
esac
if ! validate_or_fail "history backend $want"; then _hist_revert; die "Config rejected; reverted."; fi
log "Restarting Ergo (the history backend is fixed at startup; every user is disconnected)..."
irc_raw --oper --quiet 1 -- 'NOTICE $$* :Server restarting to change message-history storage.' >/dev/null 2>&1 || true
dc up -d --no-deps ergo >/dev/null 2>&1 || dc restart ergo >/dev/null 2>&1 || true
if wait_healthy ergo 120 && irc_probe; then
# Only now is the switch real, so only now does .env change.
env_set HISTORY "$want"
[[ "$want" == "postgres" ]] || env_set COMPOSE_FILE "docker-compose.yml"
rm -f "$CONF.last-good"
log "History backend is now '$want'."
if [[ "$want" == "off" ]]; then
warn "Existing persisted messages are still on disk; remove them yourself if that is the point."
fi
else
dc logs --tail 30 ergo >&2 || true
_hist_revert
dc up -d --no-deps ergo >/dev/null 2>&1 || true
die "Ergo is not healthy after the switch; config and .env reverted. Check 'ergoctl logs'."
fi
return 0 # never let a trailing false test become this command's exit status
}
cmd_debug() {
require_root; require_stack
local level
case "${1:-}" in on) level=debug ;; off) level=info ;; *) die "Usage: ergoctl debug on|off" ;; esac
snapshot_file "$CONF"
yaml_set_log_level "$CONF" "$level" || { cat "$CONF.last-good" > "$CONF"; die "No 'level:' line found in the logging block."; }
apply_change "log level $level" "$CONF"
}
case "${1:-help}" in
help|-h|--help) usage ;;
status) cmd_status ;;
users) require_stack; cmd_users ;;
logs) shift; cmd_logs "$@" ;;
version) require_stack; printf 'running %s | pinned %s | latest %s\n' "$(ergo_version_running)" "$(ergo_tag)" "$(ergo_latest_tag || echo '?')" ;;
cert) shift; cmd_cert "$@" ;;
edit) edit_file "$CONF" "ircd.yaml" ;;
motd) edit_file "$MOTD" "MOTD" ;;
rehash) require_root; require_stack
if irc_rehash; then log "Rehash complete."; else
if [[ "$REHASH_MSG" == *"could not reach"* ]]; then warn "$REHASH_MSG -- sending SIGHUP."; dc kill -s HUP ergo; else die "Rehash failed: $REHASH_MSG"; fi
fi ;;
restart) require_root; require_stack; irc_raw --oper --quiet 1 -- 'NOTICE $$* :Server restarting now.' >/dev/null 2>&1 || true
dc restart -t 15 ergo; wait_healthy ergo 120 && log "Ergo is healthy." || die "Ergo did not become healthy (ergoctl logs)." ;;
stop) require_root; require_stack; dc stop -t 15 ergo ;;
start) require_root; require_stack; dc start ergo; wait_healthy ergo 120 && log "Ergo is healthy." || die "Ergo did not become healthy (ergoctl logs)." ;;
caddy-restart) require_root; require_stack
dc run --rm --no-deps -T caddy caddy validate --config /etc/caddy/Caddyfile --adapter caddyfile >/dev/null || die "Caddyfile invalid (see caddy/etc/Caddyfile, conf.d/)."
dc restart caddy; wait_healthy caddy 60 && log "Caddy is healthy." || die "Caddy did not become healthy." ;;
debug) shift; cmd_debug "$@" ;;
history) shift; cmd_history "$@" ;;
oper) shift; cmd_oper "$@" ;;
passwd) shift; cmd_passwd "$@" ;;
cmd) shift; cmd_cmd "$@" ;;
announce) shift; [[ $# -gt 0 ]] || die "Usage: ergoctl announce <text>"; cmd_cmd "NOTICE \$\$* :$*" ;;
kill) shift; [[ $# -ge 1 ]] || die "Usage: ergoctl kill <nick> [reason]"; n="$1"; shift; cmd_cmd "KILL $n :${*:-Killed by an operator}" ;;
ban) shift; sub="${1:-list}"; shift || true
case "$sub" in
add) [[ $# -ge 1 ]] || die "Usage: ergoctl ban add <ip|cidr|nick!user@host|account> [duration e.g. 1d] [reason] -- a bare name means an ACCOUNT; use 'ban info <nick>' for a connected user's IP"
t="$1"; shift; d=""; if [[ "${1:-}" =~ ^[0-9]+[ymwdhs]$|^[0-9]+mo$ ]]; then d="DURATION $1"; shift; fi
cmd_cmd "UBAN ADD $t $d ${*:+:$*}" ;;
del) [[ $# -ge 1 ]] || die "Usage: ergoctl ban del <target>"; cmd_cmd "UBAN DEL $1" ;;
info) [[ $# -ge 1 ]] || die "Usage: ergoctl ban info <target>"; cmd_cmd "UBAN INFO $1" ;;
list|*) cmd_cmd "UBAN LIST" ;;
esac ;;
defcon) shift; cmd_cmd "DEFCON ${1:-}" ;;
backup) shift; cmd_backup "$@" ;;
restore) shift; cmd_restore "$@" ;;
update) shift; require_root; require_stack; exec bash "$STACK_DIR/update.sh" "${1:-check}" ;;
shell) require_root; require_stack; dc exec ergo sh ;;
*) usage; exit 1 ;;
esac