Files
automations/deployments/openbao
57_WolveandClaude Opus 5 a02524ae6e fix(openbao): detect the account the entrypoint drops to, not the probe's
The image's entrypoint starts as root and then runs `su-exec openbao "$@"`
before exec'ing the server. deploy.sh probed with
`--entrypoint id -u`, which BYPASSES the entrypoint, so it reported UID 0 and
chowned config.hcl, tls/ and the raft volume to root -- while the server ran
as uid 100 and could write none of it:

  error initializing storage of type raft: failed to create fsm:
  failed to open bolt file: open /openbao/data/vault.db: permission denied

The container crash-looped on that, and tls.key (0600 root:root) would have
failed the listener straight afterwards. Confirmed on the affected host:
the probe reports 0, `id -u openbao` reports 100, and the entrypoint's line
92 is `set -- su-exec openbao "$@"`.

Ask for the account the entrypoint switches to, falling back to the old probe
when the image has no such account (then the server really does run as
whatever the entrypoint started as). Preferring the service account is also
the safe direction to be wrong in: root ignores file permissions, so chowning
to the unprivileged uid still works if the server turns out to run as root,
whereas the reverse is fatal.

Gate the raft chown on the volume's ACTUAL ownership rather than on a
first-run flag. The flag was false forever after the first deploy, so a volume
left root-owned by an earlier run -- exactly the state this bug created --
could never be repaired by re-running; the operator had to chown it by hand.
Reading the owner costs one container start and still keeps the recursive
chown off a healthy live raft dir. FIRST_RUN is now unused, so it is gone.

Also from the same run, two reporting failures:
- `docker compose ps` printed "Restarting (1)" and the script went on to print
  an unqualified DEPLOYED banner. It now inspects the container state and, when
  it is not `running`, says plainly that this is a crash loop rather than the
  expected sealed-and-unhealthy state -- before and after the banner, with the
  logs command to run.
- Every `docker compose` command in that banner assumed the project directory.
  deploy.sh runs them from $STACK_DIR itself, so the omission only bit the
  operator afterwards, with "no configuration file provided: not found". The
  banner now says to cd there first and quotes that error.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-16 14:49:30 -05:00
..

openbao

Hardened OpenBao — the tape-encryption key store for Kanrisha (the LTO tape-archive system; separate repo, separate host). Deliberately separate from the tape host: a compromise of the tape node must not reach the vault, and OpenBao's mlock/TLS/unseal lifecycle is cleaner on its own box.

Unlike the other stacks here there is no Caddy / Let's Encrypt — a secrets store terminates TLS itself and is reached over the LAN, not the public internet.

  • Native TLS on the listener — self-signed by default, or a CA-signed cert from your Smallstep CA over ACME.
  • Integrated raft storage — snapshot-based DR.
  • mlock on — key material never hits swap.
  • Manual unseal by default, or PKCS#11 HSM auto-unseal.

⚠️ This vault is the sole recovery path for encrypted tapes. Losing the OpenBao data and the unseal keys/root token loses every encrypted tape. Take raft snapshots and store the unseal material out of band (below).

Required .env values

Variable Notes
OPENBAO_ADDR IP/DNS the Kanrisha tape host uses to reach this vault (goes in the cert SAN; you point the daemon at https://$OPENBAO_ADDR:8200).
OPENBAO_BIND Host interface the API publishes on. Left at the default, deploy.sh narrows it to OPENBAO_ADDR when that is an IP (a published port bypasses the INPUT firewall, so this bind is the real restriction — see Exposure). It must be an address this host actually holds — deploy.sh checks that before writing anything and refuses otherwise (SKIP_BIND_CHECK=1 to override).
OPENBAO_TLS_SANS Extra SANs beyond OPENBAO_ADDR + loopback (which deploy.sh always adds). Read from the environment at deploy time — export it before running deploy.sh.
OPENBAO_TAG OpenBao image tag (pin it).
OPENBAO_HSM_PIN Only for PKCS#11 auto-unseal. Leave blank for manual unseal.

See .env.example for the full list.

Deploy

./automations.sh        # Deploy on this host → deploy: openbao
# or, non-interactive:
OPENBAO_ADDR=10.0.0.10 SKIP_PROMPTS=1 bash deployments/openbao/deploy.sh

deploy.sh installs Docker, generates a self-signed TLS cert (if none present), seeds .env, disables swap (for mlock), narrows the API bind to the LAN IP, opens 8200/tcp, aligns file/volume ownership to the container's UID, and brings the stack up. OpenBao starts sealed — initialise + unseal once (do this immediately; an uninitialised vault reachable on the LAN can be init'd by anyone who connects):

docker compose exec -e BAO_ADDR=https://127.0.0.1:8200 openbao \
  bao operator init -tls-skip-verify           # prints 5 unseal keys + root token
docker compose exec -e BAO_ADDR=https://127.0.0.1:8200 openbao \
  bao operator unseal -tls-skip-verify <key>   # x3, three different keys

Store the unseal keys + root token out of band — ideally age-encrypted with your backup recipient (globals/age-pubkey.txt), never on this host.

Re-running after a wrong address

deploy.sh is idempotent, but two things it writes are deliberately sticky: .env (never overwritten) and tls/tls.{crt,key} (never regenerated over an existing pair). So passing a corrected OPENBAO_ADDR to a re-run does not move the vault — the cert keeps the old SAN, and the old .env usually still decides the bind. Which OPENBAO_BIND wins follows Compose's own precedence:

How OPENBAO_BIND is set What Compose uses
Exported into deploy.sh's environment (OPENBAO_BIND=… bash deploy.sh, or answering the bind prompt — automations.sh passes answers via env VAR=…) the environment value; .env is ignored for this run, so a later hand-run docker compose up -d can bind somewhere else
Derived by deploy.sh (bind prompt left blank → narrowed to OPENBAO_ADDR) the .env value, since that assignment is never exported
Absent from both 0.0.0.0all interfaces

deploy.sh warns on each of those mismatches. To actually change the address:

cd /srv/openbao
docker compose down                                   # keeps the raft volume
sed -i 's/OLD_IP/NEW_IP/g' .env                       # OPENBAO_ADDR + OPENBAO_BIND
rm -f tls/tls.crt tls/tls.key                         # force a new SAN
# then re-run deploy.sh with the corrected OPENBAO_ADDR

A bind address the host does not hold is caught up front; without that check Docker fails the up with cannot assign requested address only after the bad value is already in .env and the cert.

TLS

  • Self-signed (default): deploy.sh runs gen-tls.sh to create ./tls/tls.{crt,key} with OPENBAO_ADDR in the SAN. Hand tls.crt to the Kanrisha daemon as [encryption.openbao].ca_cert.

  • Smallstep CA over ACME (option): issue a cert from your step-ca and drop it in ./tls instead — gen-tls.sh then no-ops. e.g. with the step client:

    step ca certificate "$OPENBAO_ADDR" ./tls/tls.crt ./tls/tls.key \
      --provisioner acme --acme https://ca.lan/acme/acme/directory
    # renew on a timer:  step ca renew --daemon ./tls/tls.crt ./tls/tls.key
    

    Give the Kanrisha daemon your Smallstep root as ca_cert (then it trusts the vault without -tls-skip-verify).

Auto-unseal (optional)

Default is manual unseal after each restart. For hands-off restarts, enable the seal "pkcs11" stanza in config.hcl, mount the PKCS#11 module + device into the openbao service, and set OPENBAO_HSM_PIN in .env.

Bootstrap for Kanrisha

Once unsealed, enable KV v2 + AppRole and seed the domain key with the Kanrisha bootstrap script (from the Kanrisha repo, deploy/openbao/bootstrap.sh). It calls the bao CLI directly, so run it from a host that has bao (the Kanrisha host or your workstation) pointed at this vault — this vault host only ships Docker. Copy tls.crt to that host first and pass it as BAO_CACERT:

BAO_ADDR=https://$OPENBAO_ADDR:8200 BAO_CACERT=/path/to/openbao-ca.crt \
  BAO_TOKEN=<root> bash bootstrap.sh

It enables the kanrisha-tape KV-v2 mount (with effectively-unlimited max_versions so a key rotation never orphans old tapes), creates the AppRole + policy, seeds the domain key, and prints the [encryption.openbao] block for the Kanrisha config. Point the daemon at address = "https://$OPENBAO_ADDR:8200".

Exposure

This is a secrets store on the LAN, not a public service. Two things to know:

  • A Docker-published port is DNAT'd and bypasses the host INPUT firewall. A source-restricting rule on INPUT does not gate :8200. The real interface restriction is the publish bind: deploy.sh defaults OPENBAO_BIND to OPENBAO_ADDR when that is an IP, so the API listens only on that LAN IP. To restrict by source host, use a FORWARD/DOCKER-USER rule or network segmentation, not INPUT.
  • mTLS (optional): to require the tape host to present a client cert, enable the tls_require_and_verify_client_cert stanza in config.hcl and issue the tape host a client cert from the same CA.

And initialise the vault immediately after deploy.sh — an uninitialised vault reachable on the LAN can be bao operator init'd by anyone who connects, handing them the root token and unseal keys.

Backup / DR

The vault is the sole recovery path for encrypted tapes — back it up:

Snapshot save/restore are token-gated (sys/storage/raft/snapshot is sudo-capable) — pass a token that has that path (the root token works, or mint a dedicated snapshot-policy token). The container has no ambient token, so supply it via -e BAO_TOKEN.

# Consistent raft snapshot (safe while running); write it to a scratch path, NOT
# into the live raft dir:
docker compose exec -T -e BAO_TOKEN=<token> openbao \
  bao operator raft snapshot save -address=https://127.0.0.1:8200 -tls-skip-verify /tmp/openbao.snap
# then stream the RAW file off-box, age-encrypted. Use `exec -T ... cat`, not
# `compose cp openbao:… -` (which emits a TAR wrapper that won't restore):
docker compose exec -T openbao cat /tmp/openbao.snap | \
  age -r "$(cat /path/to/globals/age-pubkey.txt)" > "openbao-$(date +%F).snap.age"

Restore (into a fresh, unsealed vault): decrypt, copy the raw .snap in, and apply it (also token-gated):

age -d -i ~/.age/key.txt openbao-DATE.snap.age > openbao.snap
docker compose cp openbao.snap openbao:/tmp/openbao.snap
docker compose exec -T -e BAO_TOKEN=<token> openbao \
  bao operator raft snapshot restore -address=https://127.0.0.1:8200 -tls-skip-verify /tmp/openbao.snap

Snapshots do not contain the unseal keys or root token — you still need those to unseal a restored vault, which is why they are stored separately.