Files
automations/deployments/openbao/README.md
T
57_WolveandClaude Opus 5 0eb7f26663 fix(openbao): never put the vault token on a command line
`BAO_TOKEN=<root> bash update.sh update` -- which the previous commit and the
pre-existing DR runbook both recommended -- leaks the token twice over. It goes
into the shell history, and `docker compose exec -e BAO_TOKEN=<value>` puts it
in the docker process's argv, where /proc/<pid>/cmdline makes it readable by
every user on the host. The second is the worse one and is not fixed by
anything the operator does at their prompt.

The token now reaches the container over STDIN and is never an argument to
anything: the container's own shell reads one line, exports it, and execs bao.
Verified locally that the child process sees the exact value -- including &, |
and a backslash -- while its argv contains zero occurrences of it.

How a token is supplied, in order: an already-exported BAO_TOKEN; a 0600 file
named by the new BAO_TOKEN_FILE; otherwise an echo-off prompt. With no token
and no terminal it dies telling the operator to use BAO_TOKEN_FILE rather than
an inline assignment, and says why.

Added a `snapshot` subcommand so the DR runbook no longer needs an inline
pipeline at all. That runbook and the deploy banner both carried the -e form
before this deployment had an updater, so both are corrected: taking a backup
is now `bash update.sh snapshot`, and restore keeps the stdin shape rather than
-e. The Kanrisha bootstrap lines say to export the token first instead of
passing it inline.

The conf file still refuses to hold a token, but now points at BAO_TOKEN_FILE
as the unattended answer -- and suggests a snapshot-policy token rather than
the root token for it.

Not verified without a live host: that `docker compose exec -T` forwards stdin
as expected. If it does not, the snapshot fails loudly with a permission error
and the upgrade aborts before touching anything, which is the safe direction.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-22 09:31:47 -05:00

12 KiB

openbao

Hardened OpenBao — the tape-encryption key store for Kanrisha (the LTO tape-archive system; separate repo, separate host). Deliberately separate from the tape host: a compromise of the tape node must not reach the vault, and OpenBao's TLS/unseal lifecycle is cleaner on its own box.

Unlike the other stacks here there is no Caddy / Let's Encrypt — a secrets store terminates TLS itself and is reached over the LAN, not the public internet.

  • Native TLS on the listener — self-signed by default, or a CA-signed cert from your Smallstep CA over ACME.
  • Integrated raft storage — snapshot-based DR.
  • Swap disabled — key material never hits disk. (Not mlock: OpenBao removed mlock support, so disable_mlock must not appear in config.hcl at all.)
  • Manual unseal by default, or PKCS#11 HSM auto-unseal.

⚠️ This vault is the sole recovery path for encrypted tapes. Losing the OpenBao data and the unseal keys/root token loses every encrypted tape. Take raft snapshots and store the unseal material out of band (below).

Required .env values

Variable Notes
OPENBAO_ADDR IP/DNS the Kanrisha tape host uses to reach this vault (goes in the cert SAN; you point the daemon at https://$OPENBAO_ADDR:8200).
OPENBAO_BIND Host interface the API publishes on. Left at the default, deploy.sh narrows it to OPENBAO_ADDR when that is an IP (a published port bypasses the INPUT firewall, so this bind is the real restriction — see Exposure). It must be an address this host actually holds — deploy.sh checks that before writing anything and refuses otherwise (SKIP_BIND_CHECK=1 to override).
OPENBAO_TLS_SANS Extra SANs beyond OPENBAO_ADDR + loopback (which deploy.sh always adds). Read from the environment at deploy time — export it before running deploy.sh.
OPENBAO_TAG OpenBao image tag (pin it).
OPENBAO_HSM_PIN Only for PKCS#11 auto-unseal. Leave blank for manual unseal.

See .env.example for the full list.

Deploy

./automations.sh        # Deploy on this host → deploy: openbao
# or, non-interactive:
OPENBAO_ADDR=10.0.0.10 SKIP_PROMPTS=1 bash deployments/openbao/deploy.sh

deploy.sh installs Docker, generates a self-signed TLS cert (if none present), seeds .env, disables swap (keeping key material off disk), narrows the API bind to the LAN IP, opens 8200/tcp, aligns file/volume ownership to the container's UID, and brings the stack up. OpenBao starts sealed — initialise + unseal once (do this immediately; an uninitialised vault reachable on the LAN can be init'd by anyone who connects):

docker compose exec -e BAO_ADDR=https://127.0.0.1:8200 openbao \
  bao operator init -tls-skip-verify           # prints 5 unseal keys + root token
docker compose exec -e BAO_ADDR=https://127.0.0.1:8200 openbao \
  bao operator unseal -tls-skip-verify <key>   # x3, three different keys

Store the unseal keys + root token out of band — ideally age-encrypted with your backup recipient (globals/age-pubkey.txt), never on this host.

Re-running after a wrong address

deploy.sh is idempotent, but two things it writes are deliberately sticky: .env (never overwritten) and tls/tls.{crt,key} (never regenerated over an existing pair). So passing a corrected OPENBAO_ADDR to a re-run does not move the vault — the cert keeps the old SAN, and the old .env usually still decides the bind. Which OPENBAO_BIND wins follows Compose's own precedence:

How OPENBAO_BIND is set What Compose uses
Exported into deploy.sh's environment (OPENBAO_BIND=… bash deploy.sh, or answering the bind prompt — automations.sh passes answers via env VAR=…) the environment value; .env is ignored for this run, so a later hand-run docker compose up -d can bind somewhere else
Derived by deploy.sh (bind prompt left blank → narrowed to OPENBAO_ADDR) the .env value, since that assignment is never exported
Absent from both 0.0.0.0 — all interfaces

deploy.sh warns on each of those mismatches. To actually change the address:

cd /srv/openbao
docker compose down                                   # keeps the raft volume
sed -i 's/OLD_IP/NEW_IP/g' .env                       # OPENBAO_ADDR + OPENBAO_BIND
rm -f tls/tls.crt tls/tls.key                         # force a new SAN
# then re-run deploy.sh with the corrected OPENBAO_ADDR

A bind address the host does not hold is caught up front; without that check Docker fails the up with cannot assign requested address only after the bad value is already in .env and the cert.

TLS

  • Self-signed (default): deploy.sh runs gen-tls.sh to create ./tls/tls.{crt,key} with OPENBAO_ADDR in the SAN. Hand tls.crt to the Kanrisha daemon as [encryption.openbao].ca_cert.

  • Smallstep CA over ACME (option): issue a cert from your step-ca and drop it in ./tls instead — gen-tls.sh then no-ops. e.g. with the step client:

    step ca certificate "$OPENBAO_ADDR" ./tls/tls.crt ./tls/tls.key \
      --provisioner acme --acme https://ca.lan/acme/acme/directory
    # renew on a timer:  step ca renew --daemon ./tls/tls.crt ./tls/tls.key
    

    Give the Kanrisha daemon your Smallstep root as ca_cert (then it trusts the vault without -tls-skip-verify).

Auto-unseal (optional)

Default is manual unseal after each restart. For hands-off restarts, enable the seal "pkcs11" stanza in config.hcl, mount the PKCS#11 module + device into the openbao service, and set OPENBAO_HSM_PIN in .env.

Bootstrap for Kanrisha

Once unsealed, enable KV v2 + AppRole and seed the domain key with the Kanrisha bootstrap script (from the Kanrisha repo, deploy/openbao/bootstrap.sh). It calls the bao CLI directly, so run it from a host that has bao (the Kanrisha host or your workstation) pointed at this vault — this vault host only ships Docker. Copy tls.crt to that host first and pass it as BAO_CACERT:

BAO_ADDR=https://$OPENBAO_ADDR:8200 BAO_CACERT=/path/to/openbao-ca.crt \
  bash bootstrap.sh   # export BAO_TOKEN beforehand rather than passing it inline

It enables the kanrisha-tape KV-v2 mount (with effectively-unlimited max_versions so a key rotation never orphans old tapes), creates the AppRole + policy, seeds the domain key, and prints the [encryption.openbao] block for the Kanrisha config. Point the daemon at address = "https://$OPENBAO_ADDR:8200".

Exposure

This is a secrets store on the LAN, not a public service. Two things to know:

  • A Docker-published port is DNAT'd and bypasses the host INPUT firewall. A source-restricting rule on INPUT does not gate :8200. The real interface restriction is the publish bind: deploy.sh defaults OPENBAO_BIND to OPENBAO_ADDR when that is an IP, so the API listens only on that LAN IP. To restrict by source host, use a FORWARD/DOCKER-USER rule or network segmentation, not INPUT.
  • On a Proxmox guest, the per-VM firewall filters before the guest sees anything. If the VM's NIC has firewall=1 and the datacenter firewall is enabled, a Docker-published port needs its own IN ACCEPT -p tcp -dport 8200 rule in /etc/pve/firewall/<vmid>.fw. SSH working does not prove the path — it only proves there is a rule for 22. The give-away is that everything inside the guest looks perfect (curl to the bind address answers, the DNAT and FORWARD jumps are present) while tcpdump -ni eth0 'tcp port 8200' captures zero packets during a failed connection.
  • mTLS (optional): to require the tape host to present a client cert, enable the tls_require_and_verify_client_cert stanza in config.hcl and issue the tape host a client cert from the same CA.

And initialise the vault immediately after deploy.sh — an uninitialised vault reachable on the LAN can be bao operator init'd by anyone who connects, handing them the root token and unseal keys.

Reaching it without LAN access

The API publishes on OPENBAO_BIND only, so a browser on another subnet (or behind a firewall you do not control) cannot reach it. Tunnel over SSH instead of widening the publish — the generated cert already carries DNS:localhost and IP:127.0.0.1 in its SANs, so it validates as-is:

ssh -L 8200:<bind-addr>:8200 root@<host>
# then browse https://localhost:8200

This is also the better way to do the first operator init: the unseal keys and root token are shown in your browser instead of a root shell's scrollback.

Updating

update.sh is installed alongside the stack. It is not run by deploy.sh, because upgrading a live vault seals it — that is an operator's decision, not a deploy step.

cd /srv/openbao
bash update.sh check                          # declared / running / latest + seal state
bash update.sh update                         # snapshot, then upgrade (prompts)

Unlike the copyparty and ergo updaters in this repo, this one will not update on a schedule by default. Those services come back by themselves; OpenBao comes back sealed, so an unattended 03:00 update would take the vault offline until someone arrives with three unseal keys. update.sh install therefore schedules a daily check (UPDATE_POLICY=notify). UPDATE_POLICY=auto opts into unattended updates and is still refused unless a seal stanza is configured — only auto-unseal makes the vault come back on its own.

Before it changes anything, update requires that the container is running, the vault is unsealed (a sealed vault cannot produce a snapshot, so there would be no rollback plan), and a token with sys/storage/raft/snapshot — the root token works. It prompts for that token with echo off; never pass it as BAO_TOKEN=... bash update.sh, which puts it in your shell history and in /proc/<pid>/cmdline, where every user on the box can read it. For unattended use set BAO_TOKEN_FILE to a 0600 file. It writes the snapshot to /var/backups/openbao at 0600, verifies it is a valid non-empty gzip archive, and refuses to continue if it is not. Copy it off the host: it is the rollback plan, and OpenBao's upgrade guide is explicit that reverting the image alone does not roll back the data store. SKIP_SNAPSHOT=1 exists and says loudly what you are giving up.

If the pull or the start fails, the OPENBAO_TAG pin is rolled back and the previous version is started again — still sealed.

update.sh also refuses to cross into 2.7.x while an active built-in seal "pkcs11" stanza is present: that stanza is removed in 2.7.0, not merely deprecated, so the vault would come up with no way to unseal at all. Migrate to the external plugin "kms" "pkcs11" first.

Backup / DR

The vault is the sole recovery path for encrypted tapes — back it up:

Snapshot save/restore are token-gated (sys/storage/raft/snapshot is sudo-capable) — they need a token that has that path (the root token works, or mint a dedicated snapshot-policy token). The container has no ambient token.

Do not put the token on a command line. docker ... -e BAO_TOKEN=<value> places it in the docker process's argv, and /proc/<pid>/cmdline is readable by every user on the host — so it leaks further than your shell history does. Use update.sh snapshot, which prompts with echo off and passes the token to the container over stdin:

cd /srv/openbao
bash update.sh snapshot          # prompts; writes to /var/backups/openbao, 0600

Then age-encrypt it and move it off-box:

age -r "$(cat /path/to/globals/age-pubkey.txt)" \
  < /var/backups/openbao/openbao-pre-manual-*.snap \
  > "openbao-$(date +%F).snap.age"

Restore (into a fresh, unsealed vault): decrypt, copy the raw .snap in, and apply it (also token-gated):

age -d -i ~/.age/key.txt openbao-DATE.snap.age > openbao.snap
docker compose cp openbao.snap openbao:/tmp/openbao.snap
# Same argv problem as above, so hand the token over stdin rather than -e:
read -rs -p "token: " T; echo
printf '%s\n' "$T" | docker compose exec -T openbao sh -c \
  'read -r t; BAO_TOKEN="$t" exec bao operator raft snapshot restore \
     -address=https://127.0.0.1:8200 -tls-skip-verify /tmp/openbao.snap'
unset T

Snapshots do not contain the unseal keys or root token — you still need those to unseal a restored vault, which is why they are stored separately.