Files
automations/deployments/ergo/update.sh
T
57_WolveandClaude Opus 5 4cbd808eb1 feat(ergo): add Ergo IRC server deployment with ergoctl admin CLI
New deployments/ergo/: the Ergo IRC server behind Caddy for Let's Encrypt TLS
and the IRC-over-WebSocket endpoint. Docker rather than a native OpenRC/systemd
service, because Alpine's apk ergo trails upstream (3.24 ships 2.18.0 against a
2.19.1 security release) and Debian/Alma package it at all -- so native would
mean three install paths plus a per-distro ACME client.

Both containers run with network_mode: host. IRC bans, throttling and cloaking
key on the client's address, and Docker's userland proxy would hide every IPv6
client behind the bridge gateway; host mode also makes the repo's INPUT firewall
genuinely govern 80/443/6697. Caddy reaches Ergo over loopback, which is what
lets Ergo honour X-Forwarded-For (proxy-allowed-from defaults to localhost) and
mark web sessions secure.

- deploy.sh generates ircd.yaml ONCE from the pulled image's own default.yaml
  (version-matched), rewriting the listeners/websockets blocks wholesale rather
  than patching lines, then asserts hard post-conditions and validates with
  `ergo run --smoke` in a throwaway container before anything starts.
- update.sh: pinned vX.Y.Z tags, GHSA + "### Security" release-note policies,
  pre-flight against the new image, user NOTICE + grace, stop-consistent DB
  snapshot, health check (IRC-level, not a bare TCP connect) and rollback that
  restores the DB only when the schema actually moved. Compatibility-break
  releases are held for review. certsync copies Caddy's cert pairwise-atomically
  and verifies the fingerprint served on 6697 after the rehash.
- ergoctl: status/users/logs, validated edit+rehash, oper add/passwd/rm,
  moderation, backup/restore, cert and update passthrough. Talks IRC to the
  loopback listener over bash /dev/tcp and strips control characters from
  replies.
- Ergo runs as a non-root system user, read-only rootfs, all caps dropped;
  Caddy keeps only NET_BIND_SERVICE, with admin API and HTTP/3 off.

Reviewed adversarially across six lenses; 20 confirmed findings fixed, notably
a dead SIGHUP fallback (`rc=$?` after an `if` is always 0), several `set -e`
aborts from non-total pipelines, a release-list cache that only ever populated
in a subshell, and re-runs that used shell defaults instead of the deployed
.env. Verified locally: bash -n, LF endings, the ircd.yaml render against the
real 2.19.1 template in both PLAINTEXT modes, the yaml/oper/version/env helpers,
and the IRC client against a fake server (registration, oper, rehash success and
400-failure, control-character stripping, server-down paths). Not yet exercised
on a Docker host: the containers themselves, ACME issuance and cert sync.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-02 13:19:31 -05:00

571 lines
26 KiB
Bash

#!/usr/bin/env bash
#
# update.sh -- keep the Ergo stack current: container updates with a DB snapshot
# and health-checked rollback, plus the Caddy -> Ergo TLS certificate sync.
# Companion to deploy.sh; installed in $STACK_DIR next to ergolib.sh.
#
# Subcommands:
# check (default) current vs latest, advisories, cert state; changes nothing
# run what the daily schedule runs: apply UPDATE_POLICY, cert sync, expiry check
# update update now (to latest, or TARGET_VERSION=x.y.z); honours FORCE_UPDATE
# certsync copy Caddy's cert into ircd/ and rehash if it changed (15-min schedule)
# caddy pull a newer Caddy image and recreate it (health-checked)
# install write /etc/ergo-update.conf + schedule the jobs
# uninstall remove the schedule
#
# Policy (UPDATE_POLICY):
# latest update to the newest release whenever one exists (default)
# security update ONLY when a published GitHub security advisory covers the
# running version, or a release between current and latest has a
# "### Security" section in its notes
# off never change the running version (check/notify only)
#
# Safety rails (all policies):
# * releases whose notes announce "Compatibility breaks" are HELD (notify
# only) unless FORCE_UPDATE=1 -- review the notes, then `update update`
# * the new image must load the current ircd/ (ergo run --smoke on a copy,
# which also dry-runs any DB schema upgrade) before the live server is touched
# * users get a NOTICE and UPDATE_GRACE seconds before the restart
# * ircd.db is snapshotted with Ergo stopped (buntdb is append-only; a live
# copy can miss the last second of writes)
# * health = compose healthcheck (IRC-level) + a registration handshake; on
# failure the previous tag is restored and, if the new version bumped the
# DB schema, the snapshot is put back
#
# Env, in precedence order: environment > /etc/ergo-update.conf > the stack's
# .env > the built-in default.
# STACK_DIR=/srv/ergo UPDATE_POLICY=latest FORCE_UPDATE=0 UPDATE_GRACE=60
# CADDY_AUTOUPDATE=0 NOTIFY=1 SSH_NOTIFY_CONF=/etc/ssh-notify.conf DRY_RUN=0
# TARGET_VERSION= GH_REPO=ergochat/ergo LOG=/var/log/ergo-update.log
# SCHEDULE_UPDATES=1 (install: 0 schedules certsync only, no daily update)
# VERBOSE= (certsync: also log "nothing to do" outcomes)
set -euo pipefail
SELF="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)/$(basename "${BASH_SOURCE[0]}")"
SELF_DIR="$(dirname "$SELF")"
log() { printf '\033[1;32m[+]\033[0m %s\n' "$*"; }
warn() { printf '\033[1;33m[!]\033[0m %s\n' "$*" >&2; }
die() { printf '\033[1;31m[x]\033[0m %s\n' "$*" >&2; exit 1; }
# shellcheck source=ergolib.sh
. "$SELF_DIR/ergolib.sh"
load_conf
: "${STACK_DIR:=$SELF_DIR}"
ergo_set_paths
# .env is the stack's own settings file and documents these knobs, so honour it
# for anything the environment and /etc/ergo-update.conf did not already set.
env_defaults UPDATE_POLICY FORCE_UPDATE UPDATE_GRACE CADDY_AUTOUPDATE
: "${UPDATE_POLICY:=latest}"
: "${FORCE_UPDATE:=0}"
: "${UPDATE_GRACE:=60}"
: "${CADDY_AUTOUPDATE:=0}"
: "${TARGET_VERSION:=}"
: "${SCHEDULE_UPDATES:=1}"
: "${LOG:=/var/log/ergo-update.log}"
# ---------------------------------------------------------------------------
# Release metadata (GitHub). Unauthenticated: 60 requests/hour -- a run uses 2-3.
# ---------------------------------------------------------------------------
RELEASES_JSON=""
# Fetch the release list ONCE per process. This must be called as a plain
# statement, never inside a pipeline or $(): the assignment would land in a
# subshell and every use would re-download the list (and burn the 60/hour
# anonymous API budget).
load_releases() {
[[ -n "$RELEASES_JSON" ]] && return 0
RELEASES_JSON="$(fetch "https://api.github.com/repos/${GH_REPO}/releases?per_page=50" || true)"
printf '%s' "$RELEASES_JSON" | jq -e 'type=="array"' >/dev/null 2>&1 || RELEASES_JSON="[]"
return 0
}
# Bodies of stable releases with cur < version <= upto, newest first:
# "=== <tag>" header lines followed by the body.
release_notes_between() { # <cur> <upto>
local cur="$1" upto="$2" tag v
load_releases
while IFS= read -r tag; do
[[ -n "$tag" ]] || continue
v="$(normver "$tag")"; [[ -n "$v" ]] || continue
ver_gt "$v" "$cur" || continue
ver_le "$v" "$upto" || continue
printf '=== %s\n' "$tag"
printf '%s' "$RELEASES_JSON" | jq -r --arg t "$tag" '.[] | select(.tag_name==$t) | .body // ""'
printf '\n'
done < <(printf '%s' "$RELEASES_JSON" | jq -r '.[] | select(.prerelease==false and .draft==false) | .tag_name' 2>/dev/null)
}
# Published GitHub security advisories covering <cur>. Prints "GHSA-... <patched>"
# lines. Ranges look like "<= v2.19.0" or ">= v2.15.0, < v2.19.1".
ghsa_covering() { # <cur>
local cur="$1" json id rng patched part op ver ok
json="$(fetch "https://api.github.com/repos/${GH_REPO}/security-advisories" || true)"
printf '%s' "$json" | jq -e 'type=="array"' >/dev/null 2>&1 || return 0
while IFS=$'\t' read -r id rng patched; do
[[ -n "$rng" ]] || continue
ok=1
IFS=',' read -r -a parts <<< "$rng"
for part in "${parts[@]}"; do
part="$(printf '%s' "$part" | tr -d ' ')"
op="$(printf '%s' "$part" | grep -oE '^(<=|>=|==|<|>|=)' || echo '=')"
ver="$(normver "$part")"; [[ -n "$ver" ]] || { ok=0; break; }
case "$op" in
'<') ver_gt "$ver" "$cur" || ok=0 ;;
'<=') ver_ge "$ver" "$cur" || ok=0 ;;
'>') ver_gt "$cur" "$ver" || ok=0 ;;
'>=') ver_ge "$cur" "$ver" || ok=0 ;;
*) [[ "$(ver_cmp "$ver" "$cur")" == "0" ]] || ok=0 ;;
esac
(( ok )) || break
done
(( ok )) && printf '%s %s\n' "$id" "$(normver "$patched")"
done < <(printf '%s' "$json" | jq -r '.[] | select(.state=="published") | .ghsa_id as $id | (.vulnerabilities // [])[] | [$id, (.vulnerable_version_range // ""), ((.patched_versions // "") | tostring)] | @tsv' 2>/dev/null)
}
# ---------------------------------------------------------------------------
# Decide what to do. Sets globals (not via stdout -- a $() would lose them):
# TARGET (x.y.z or ""), REASON, HOLD (1 = compat break, notify only),
# NOTES (one-line summary for the notification), DBCHANGE (0/1)
# ---------------------------------------------------------------------------
TARGET=""; REASON=""; HOLD=0; NOTES=""; DBCHANGE=0
resolve_target() { # <current x.y.z>
local cur="$1" latest notes ghsa ids
TARGET=""; REASON=""; HOLD=0; NOTES=""; DBCHANGE=0
latest="$(normver "$(ergo_latest_tag || true)")"
if [[ -n "$TARGET_VERSION" ]]; then
TARGET="$(normver "$TARGET_VERSION")"; REASON="target override ${TARGET}"
else
case "$UPDATE_POLICY" in
off) REASON="policy=off (no changes)"; return 0 ;;
security)
ghsa="$(ghsa_covering "$cur")"
if [[ -n "$ghsa" ]]; then
ids="$(printf '%s\n' "$ghsa" | awk '{print $1}' | paste -sd, - || true)"
# NF>=2: an advisory with no patched version yields only an id.
TARGET="$(printf '%s\n' "$ghsa" | awk 'NF>=2 {print $2}' | tail -n1 || true)"
[[ -n "$TARGET" ]] || TARGET="$latest"
REASON="security advisory covers ${cur}: ${ids}"
elif [[ -n "$latest" ]] && ver_gt "$latest" "$cur" \
&& grep -qiE '^#+ *security|security vulnerabilit' <<< "$(release_notes_between "$cur" "$latest")"; then
TARGET="$latest"; REASON="a release since ${cur} carries a Security section"
else
REASON="no security advisory or security release covers ${cur}"; return 0
fi ;;
latest|*)
[[ -n "$latest" ]] || { REASON="could not resolve the latest release"; return 0; }
TARGET="$latest"; REASON="policy=latest" ;;
esac
fi
[[ -n "$TARGET" ]] || return 0
notes="$(release_notes_between "$cur" "$TARGET")"
# Feed grep from a here-string, not a pipe: `printf | grep -q` on notes
# larger than the pipe buffer makes printf die of SIGPIPE, and pipefail
# then reports "no match" for what was in fact a match.
if grep -qiE 'compatibility break|not backwards[- ]compatible|backwards[- ]incompatible' <<< "$notes"; then
HOLD=1
fi
if grep -qiE 'database change|database file format changes|schema' <<< "$notes" \
&& ! grep -qiE 'no changes to the config file format or database file format' <<< "$notes"; then
DBCHANGE=1
fi
NOTES="$(grep -iE 'config file format|database (file )?format|database change|compatibility break' <<< "$notes" \
| head -n 3 | sed 's/^[ *-]*//' | tr '\n' ' ' | cut -c1-300 || true)"
return 0
}
# ---------------------------------------------------------------------------
# Apply an update: pre-flight on a copy, warn users, stop, snapshot, up, verify;
# roll back (tag + DB snapshot when needed) on failure.
# ---------------------------------------------------------------------------
notify_users() { # <text> -- best effort, only when Ergo is up and we can oper
ergo_running || return 0
irc_raw --oper --quiet 1 -- "NOTICE \$\$* :$1" >/dev/null 2>&1 || true
}
apply_update() { # <from x.y.z> <to x.y.z> <reason>
local from="$1" to="$2" reason="$3" image_base new_image snap ts cause prev_tag
image_base="$(env_get ERGO_IMAGE)"; image_base="${image_base:-ghcr.io/ergochat/ergo}"
new_image="${image_base}:v${to}"
# Roll back to the tag that is actually pinned right now. Reconstructing
# "v${from}" would write a bare "v" when the running version is unknown
# (e.g. .env still holds the 'stable' fallback and the container is down).
prev_tag="$(ergo_tag)"
log "Updating Ergo ${from:-?} -> ${to} (${reason})..."
if [[ "$DRY_RUN" == "1" ]]; then
echo "DRY: pull ${new_image}; validate on a copy; NOTICE users; stop; snapshot ircd.db; ERGO_TAG=v${to}; up -d --no-deps ergo"
return 0
fi
cp -a "$ENV_FILE" "${ENV_FILE}.bak.$(date -u +%Y%m%d%H%M%S)" 2>/dev/null || true
log "Pulling ${new_image}..."
if ! docker pull -q "$new_image" >/dev/null 2>&1; then
send_notice "Ergo update FAILED" "high" "pull of ${new_image} failed; staying on ${from}"
return 1
fi
log "Pre-flight: loading the current ircd/ with ${new_image} (throwaway container)..."
if ! ergo_validate_config "$new_image"; then
cause="$(tail -n 5 "$VALIDATE_LOG" 2>/dev/null | tr '\n' ' ' | cut -c1-300)"
warn "Pre-flight failed: $cause"
send_notice "Ergo update HELD" "high" "${to} cannot load the current config/DB: ${cause}. Staying on ${from}; run 'ergoctl update check'."
return 1
fi
rm -f "$VALIDATE_LOG"
if (( UPDATE_GRACE > 0 )) && ergo_running; then
log "Warning connected users (${UPDATE_GRACE}s grace)..."
notify_users "Server restarting for an upgrade to Ergo ${to} in ${UPDATE_GRACE} seconds -- you will be disconnected briefly."
sleep "$UPDATE_GRACE"
fi
log "Stopping Ergo and snapshotting ircd.db..."
dc stop -t 15 ergo >/dev/null 2>&1 || true
ts="$(date -u +%Y%m%d%H%M%S)"
install -d -m 0700 "$BACKUP_DIR"
snap="$BACKUP_DIR/ircd.db.${from:-unknown}.${ts}"
if [[ -f "$IRCD_DIR/ircd.db" ]]; then
cp -p "$IRCD_DIR/ircd.db" "$snap" && chmod 0600 "$snap"
else
snap=""
fi
env_set ERGO_TAG "v${to}"
dc up -d --no-deps ergo >/dev/null 2>&1 || true
if wait_healthy ergo 120 && irc_probe; then
log "Ergo ${to} is healthy."
record_template "$new_image" "$from" "$to"
prune_snapshots
local b="updated ${from:-?} -> ${to} (${reason})"
[[ -n "$NOTES" ]] && b="${b}. Notes: ${NOTES}"
[[ -n "$TEMPLATE_NOTE" ]] && b="${b}. ${TEMPLATE_NOTE}"
send_notice "Ergo updated" "default" "$b"
return 0
fi
cause="$(dc logs --tail 60 ergo 2>/dev/null | grep -E 'Config file did not load|Database requires update|Could not load server|failed to|fatal|panic' | tail -n 2 | cut -c1-200 | tr '\n' ' ' || true)"
warn "Ergo ${to} did not become healthy (${cause:-no diagnostic in the log}); rolling back to ${prev_tag}."
dc stop -t 10 ergo >/dev/null 2>&1 || true
env_set ERGO_TAG "$prev_tag"
dc up -d --no-deps ergo >/dev/null 2>&1 || true
if ! wait_healthy ergo 90; then
# The old binary refuses a newer schema ("Database requires update") --
# only then does the snapshot go back (otherwise keep the live DB).
if [[ -n "$snap" ]] && { [[ "$DBCHANGE" == 1 ]] || dc logs --tail 40 ergo 2>/dev/null | grep -q 'Database requires update'; }; then
warn "Old version rejects the upgraded database; restoring the snapshot ${snap}."
dc stop -t 10 ergo >/dev/null 2>&1 || true
install -m 0600 -o "$(ergo_uid)" -g "$(ergo_gid)" "$snap" "$IRCD_DIR/ircd.db"
dc up -d --no-deps ergo >/dev/null 2>&1 || true
wait_healthy ergo 90 || true
fi
fi
if [[ "$(svc_state ergo)" == "running healthy" ]]; then
send_notice "Ergo update FAILED" "high" "${to} unhealthy (${cause:-see docker compose logs ergo}); rolled back to ${prev_tag}. Snapshot: ${snap:-none}"
else
send_notice "Ergo DOWN after failed update" "urgent" "${to} failed (${cause:-?}) and ${prev_tag} did not come back healthy. Snapshot: ${snap:-none}. Run 'ergoctl status' / 'docker compose logs ergo'."
fi
return 1
}
# Keep the image's default.yaml per version and diff it against the previous one
# so template drift (new options, renamed keys) is visible after an upgrade.
TEMPLATE_NOTE=""
record_template() { # <image> <from> <to>
local image="$1" from="$2" to="$3" new old diff added removed
TEMPLATE_NOTE=""
install -d -m 0700 "$TEMPLATES_DIR"
new="$TEMPLATES_DIR/default.yaml.v${to}"
ergo_extract_template "$image" "$new" 2>/dev/null || { rm -f "$new"; return 0; }
chmod 0600 "$new"
old="$TEMPLATES_DIR/default.yaml.v${from}"
[[ -f "$old" ]] || return 0
diff="$TEMPLATES_DIR/diff-v${from}-v${to}.txt"
if diff -u "$old" "$new" > "$diff" 2>/dev/null; then rm -f "$diff"; return 0; fi
chmod 0600 "$diff" 2>/dev/null || true
added="$(grep -cE '^\+[^+]' "$diff" || true)"; removed="$(grep -cE '^-[^-]' "$diff" || true)"
TEMPLATE_NOTE="default config changed upstream (+${added}/-${removed} lines): review ${diff} against ircd/ircd.yaml"
log "$TEMPLATE_NOTE"
}
prune_snapshots() { # keep the last 5 DB snapshots and .env backups
local f n=0
for f in $(ls -1t "$BACKUP_DIR"/ircd.db.* 2>/dev/null); do
n=$((n + 1)); (( n > 5 )) && rm -f "$f"
done
n=0
for f in $(ls -1t "$STACK_DIR"/.env.bak.* 2>/dev/null); do
n=$((n + 1)); (( n > 5 )) && rm -f "$f"
done
return 0
}
# The pinned tag must be an immutable vX.Y.Z: deploy.sh falls back to 'stable'
# when GitHub is unreachable, and a floating tag silently changes the running
# version on any pull, with no pre-flight, snapshot or rollback. Pin it to
# whatever is running as soon as we can see a version (no restart needed).
pin_floating_tag() { # <current x.y.z>
local cur="$1" tag
tag="$(ergo_tag)"
[[ -n "$cur" ]] || return 0
case "$tag" in v[0-9]*) return 0 ;; esac
env_set ERGO_TAG "v${cur}"
log "Pinned ERGO_TAG=v${cur} (was '${tag}', which floats)."
}
# ---------------------------------------------------------------------------
# Cert sync + expiry watch
# ---------------------------------------------------------------------------
do_certsync() {
local rc=0 prev
ergo_certsync || rc=$?
prev="$(state_get certsync)"
case "$rc" in
0) [[ "$CERTSYNC_MSG" == "up to date" ]] || log "certsync: $CERTSYNC_MSG"
if [[ -n "$prev" ]]; then state_rm certsync; send_notice "Ergo TLS recovered" "default" "cert sync ok: ${CERTSYNC_MSG}"; fi ;;
2) [[ -n "${VERBOSE:-}" ]] && log "certsync: $CERTSYNC_MSG"
return 0 ;;
*) warn "certsync: $CERTSYNC_MSG"
if [[ "$prev" != "$CERTSYNC_MSG" ]]; then state_set certsync "$CERTSYNC_MSG"; send_notice "Ergo TLS sync FAILED" "high" "$CERTSYNC_MSG"; fi
return 1 ;;
esac
return 0
}
cert_expiry_check() { # daily: warn once per day if the served cert is self-signed or expiring
local f="$IRCD_DIR/fullchain.pem" today msg=""
[[ -f "$f" ]] || return 0
today="$(date -u +%F)"
if cert_is_selfsigned "$f"; then
msg="Ergo is serving a SELF-SIGNED certificate on 6697 (Caddy has not issued one yet: check DNS for $(ergo_domain) and ports 80/443)."
elif cert_expires_within "$f" $((14 * 86400)); then
msg="Ergo's TLS cert expires soon ($(cert_enddate "$f")) and certsync has not replaced it -- is the 15-minute job running? ($SELF certsync)"
fi
[[ -n "$msg" ]] || { state_rm certexpiry; return 0; }
warn "$msg"
[[ "$(state_get certexpiry)" == "$today" ]] && return 0
state_set certexpiry "$today"
send_notice "Ergo TLS attention" "high" "$msg"
}
# ---------------------------------------------------------------------------
# Caddy: pull + recreate when the image moved (drops web-client websockets)
# ---------------------------------------------------------------------------
do_caddy() {
preflight
local tag before after
tag="$(env_get CADDY_TAG)"; tag="${tag:-2-alpine}"
before="$(docker inspect --format '{{.Image}}' ergo-caddy 2>/dev/null | head -n1 || true)"
log "Pulling caddy:${tag}..."
docker pull -q "caddy:${tag}" >/dev/null 2>&1 || { warn "pull failed"; return 1; }
after="$(docker image inspect --format '{{.Id}}' "caddy:${tag}" 2>/dev/null | head -n1 || true)"
if [[ -n "$before" && "$before" == "$after" ]]; then log "Caddy is current (${tag})."; return 0; fi
[[ "$DRY_RUN" == "1" ]] && { echo "DRY: docker compose up -d --no-deps caddy"; return 0; }
log "Recreating Caddy (web-client websockets will reconnect)..."
dc up -d --no-deps caddy >/dev/null 2>&1 || true
if wait_healthy caddy 60; then
send_notice "Caddy updated" "min" "caddy:${tag} recreated (image $(printf '%s' "$after" | cut -c8-19))"
return 0
fi
send_notice "Caddy update FAILED" "high" "caddy:${tag} is not healthy after recreate; check docker compose logs caddy"
return 1
}
# ---------------------------------------------------------------------------
# Subcommands
# ---------------------------------------------------------------------------
preflight() {
[[ $EUID -eq 0 ]] || die "Run as root."
[[ -f "$STACK_DIR/docker-compose.yml" ]] || die "No stack at $STACK_DIR (set STACK_DIR)."
command -v docker >/dev/null 2>&1 || die "docker not found."
command -v jq >/dev/null 2>&1 || die "jq is required (release metadata); install it."
}
do_check() {
preflight
local cur latest ghsa held="" prio="min" note
cur="$(ergo_version_running)"
pin_floating_tag "$cur"
latest="$(normver "$(ergo_latest_tag || true)")"
log "Running: ${cur:-unknown} (pinned $(ergo_tag)) | latest release: ${latest:-unknown} | policy: ${UPDATE_POLICY}"
ghsa="$(ghsa_covering "$cur" || true)"
if [[ -n "$ghsa" ]]; then
note="VULNERABLE: $(printf '%s\n' "$ghsa" | awk '{print $1}' | paste -sd, -) covers ${cur}"; prio="high"; warn "$note"
else
note="no published advisory covers ${cur}"; log "$note"
fi
if [[ -n "$cur" && -n "$latest" ]] && ver_gt "$latest" "$cur"; then
resolve_target "$cur"
log "A newer release is available: ${cur} -> ${latest}"
[[ "$prio" == "min" ]] && prio="default"
note="${note}; newer release ${latest} available"
if [[ "$HOLD" == 1 ]]; then held=" -- HELD: release notes announce compatibility breaks (review, then FORCE_UPDATE=1 or 'ergoctl update update')"; warn "${held# -- }"; fi
[[ -n "$NOTES" ]] && log "Release notes: $NOTES"
send_notice "Ergo check" "$prio" "${note}${held}"
fi
local f="$IRCD_DIR/fullchain.pem"
if [[ -f "$f" ]]; then
if cert_is_selfsigned "$f"; then warn "TLS: serving a SELF-SIGNED cert (no Let's Encrypt cert from Caddy yet)."
else log "TLS: $(cert_subject_cn "$f"), expires $(cert_enddate "$f") (served: $( [[ "$(served_fingerprint)" == "$(cert_fingerprint "$f")" ]] && echo matches || echo DIFFERS -- rehash needed ))"; fi
fi
}
do_run() {
preflight
if [[ "$DRY_RUN" != "1" ]]; then
install -d -m 0755 "$(dirname "$LOG")" 2>/dev/null || true
echo "=== ergo-update $(date -u +%FT%TZ) ===" >> "$LOG"
# Scheduled runs have nowhere to send stdout (busybox crond would try to
# mail it and these hosts have no MTA), so keep the whole run in the log.
[[ -t 1 ]] || exec >>"$LOG" 2>&1
fi
do_certsync || true
cert_expiry_check || true
local cur
cur="$(ergo_version_running)"
pin_floating_tag "$cur"
resolve_target "$cur"
log "current=${cur:-?} | ${REASON}"
if [[ -n "$TARGET" && -n "$cur" ]] && ver_gt "$TARGET" "$cur"; then
if [[ "$HOLD" == 1 && "$FORCE_UPDATE" != "1" ]]; then
warn "HELD: ${TARGET} announces compatibility breaks. Review the release notes, then 'ergoctl update update' (or FORCE_UPDATE=1)."
if [[ "$(state_get held)" != "$TARGET" ]]; then
state_set held "$TARGET"
send_notice "Ergo update HELD" "default" "${cur} -> ${TARGET} announces compatibility breaks; not applied automatically. Notes: ${NOTES:-see GitHub release}. Apply with 'ergoctl update update'."
fi
else
state_rm held
apply_update "$cur" "$TARGET" "$REASON" || true
fi
else
log "No update to apply."
fi
if [[ "$CADDY_AUTOUPDATE" == "1" ]]; then do_caddy || true; fi
}
do_update() {
preflight
local cur
cur="$(ergo_version_running)"
[[ -n "$cur" ]] || die "Cannot tell which version is running (container down and ERGO_TAG='$(ergo_tag)' is not a version). Start Ergo, or pin ERGO_TAG=vX.Y.Z in $ENV_FILE first."
pin_floating_tag "$cur"
if [[ -z "$TARGET_VERSION" && "$UPDATE_POLICY" == "off" ]]; then UPDATE_POLICY=latest; fi
resolve_target "$cur"
[[ -n "$TARGET" ]] || die "Could not determine a target version (${REASON})."
if [[ -n "$cur" ]] && ! ver_gt "$TARGET" "$cur" && [[ -z "$TARGET_VERSION" ]]; then
log "Already on ${cur} (latest ${TARGET}); nothing to do."
return 0
fi
if [[ "$HOLD" == 1 && "$FORCE_UPDATE" != "1" ]]; then
warn "Release notes between ${cur} and ${TARGET} announce compatibility breaks: ${NOTES}"
warn "Proceeding because you asked explicitly (this is 'update update'). Set FORCE_UPDATE=1 to also auto-apply such releases."
fi
apply_update "$cur" "$TARGET" "${REASON} (manual)"
}
# Only what the scheduled job needs to FIND the stack lives here. The update
# knobs (UPDATE_POLICY, FORCE_UPDATE, UPDATE_GRACE, CADDY_AUTOUPDATE) are read
# from the stack's .env so there is a single place to edit them; setting one
# here (or in the environment) still overrides .env for that run.
write_conf() {
cat > "$ERGO_UPDATE_CONF" <<CONF
# Where the scheduled Ergo jobs find the stack (deployments/ergo/update.sh).
# The update knobs live in \${STACK_DIR}/.env -- edit them there. Anything set
# here overrides .env, so leave it alone unless you mean to pin a value.
STACK_DIR="${STACK_DIR}"
NOTIFY="${NOTIFY}"
CONF
chmod 644 "$ERGO_UPDATE_CONF"
log "Wrote $ERGO_UPDATE_CONF (update knobs come from ${STACK_DIR}/.env)"
}
do_install() {
[[ $EUID -eq 0 ]] || die "Run as root."
write_conf
case "$(osfam)" in
alpine)
install -d -m 0755 /etc/periodic/15min /etc/periodic/daily
printf '#!/bin/sh\nexec bash "%s" certsync\n' "$SELF" > /etc/periodic/15min/ergo-certsync
chmod +x /etc/periodic/15min/ergo-certsync
if [[ "$SCHEDULE_UPDATES" == "1" ]]; then
printf '#!/bin/sh\nexec bash "%s" run\n' "$SELF" > /etc/periodic/daily/ergo-update
chmod +x /etc/periodic/daily/ergo-update
else
rm -f /etc/periodic/daily/ergo-update
fi
if command -v rc-update >/dev/null 2>&1; then
rc-update add crond default >/dev/null 2>&1 || true
rc-service crond status >/dev/null 2>&1 || rc-service crond start >/dev/null 2>&1 || true
fi
log "Scheduled: /etc/periodic/15min/ergo-certsync$( [[ "$SCHEDULE_UPDATES" == "1" ]] && printf ', /etc/periodic/daily/ergo-update (policy=%s)' "$UPDATE_POLICY" )." ;;
*)
cat > /etc/systemd/system/ergo-certsync.service <<EOF
[Unit]
Description=Ergo: sync Caddy's TLS cert into the IRC server
After=docker.service
[Service]
Type=oneshot
ExecStart=/usr/bin/env bash $SELF certsync
EOF
cat > /etc/systemd/system/ergo-certsync.timer <<EOF
[Unit]
Description=Ergo cert sync every 15 minutes
[Timer]
OnCalendar=*:0/15
Persistent=true
[Install]
WantedBy=timers.target
EOF
cat > /etc/systemd/system/ergo-update.service <<EOF
[Unit]
Description=Ergo container updater
After=docker.service
[Service]
Type=oneshot
ExecStart=/usr/bin/env bash $SELF run
EOF
cat > /etc/systemd/system/ergo-update.timer <<EOF
[Unit]
Description=Daily Ergo update check
[Timer]
OnCalendar=daily
Persistent=true
RandomizedDelaySec=1h
[Install]
WantedBy=timers.target
EOF
systemctl daemon-reload
systemctl enable --now ergo-certsync.timer >/dev/null 2>&1 || true
if [[ "$SCHEDULE_UPDATES" == "1" ]]; then systemctl enable --now ergo-update.timer >/dev/null 2>&1 || true
else systemctl disable --now ergo-update.timer >/dev/null 2>&1 || true; fi
log "Scheduled: ergo-certsync.timer (15 min)$( [[ "$SCHEDULE_UPDATES" == "1" ]] && printf ', ergo-update.timer (daily, policy=%s)' "$UPDATE_POLICY" )." ;;
esac
}
do_uninstall() {
[[ $EUID -eq 0 ]] || die "Run as root."
rm -f /etc/periodic/15min/ergo-certsync /etc/periodic/daily/ergo-update
if command -v systemctl >/dev/null 2>&1; then
systemctl disable --now ergo-certsync.timer ergo-update.timer >/dev/null 2>&1 || true
rm -f /etc/systemd/system/ergo-certsync.{timer,service} /etc/systemd/system/ergo-update.{timer,service}
systemctl daemon-reload >/dev/null 2>&1 || true
fi
log "Removed the scheduled Ergo jobs (config kept at $ERGO_UPDATE_CONF)."
}
case "${1:-check}" in
check) do_check ;;
run) do_run ;;
update) do_update ;;
certsync) preflight; do_certsync ;;
caddy) do_caddy ;;
install) do_install ;;
uninstall) do_uninstall ;;
*) die "Usage: update.sh [check|run|update|certsync|caddy|install|uninstall]" ;;
esac