Files
automations/deployments/openbao/README.md
T
57_WolveandClaude Opus 5 0eb7f26663 fix(openbao): never put the vault token on a command line
`BAO_TOKEN=<root> bash update.sh update` -- which the previous commit and the
pre-existing DR runbook both recommended -- leaks the token twice over. It goes
into the shell history, and `docker compose exec -e BAO_TOKEN=<value>` puts it
in the docker process's argv, where /proc/<pid>/cmdline makes it readable by
every user on the host. The second is the worse one and is not fixed by
anything the operator does at their prompt.

The token now reaches the container over STDIN and is never an argument to
anything: the container's own shell reads one line, exports it, and execs bao.
Verified locally that the child process sees the exact value -- including &, |
and a backslash -- while its argv contains zero occurrences of it.

How a token is supplied, in order: an already-exported BAO_TOKEN; a 0600 file
named by the new BAO_TOKEN_FILE; otherwise an echo-off prompt. With no token
and no terminal it dies telling the operator to use BAO_TOKEN_FILE rather than
an inline assignment, and says why.

Added a `snapshot` subcommand so the DR runbook no longer needs an inline
pipeline at all. That runbook and the deploy banner both carried the -e form
before this deployment had an updater, so both are corrected: taking a backup
is now `bash update.sh snapshot`, and restore keeps the stdin shape rather than
-e. The Kanrisha bootstrap lines say to export the token first instead of
passing it inline.

The conf file still refuses to hold a token, but now points at BAO_TOKEN_FILE
as the unattended answer -- and suggests a snapshot-policy token rather than
the root token for it.

Not verified without a live host: that `docker compose exec -T` forwards stdin
as expected. If it does not, the snapshot fails loudly with a permission error
and the upgrade aborts before touching anything, which is the safe direction.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-22 09:31:47 -05:00

255 lines
12 KiB
Markdown

# openbao
Hardened [OpenBao](https://openbao.org) — the tape-encryption key store for
**Kanrisha** (the LTO tape-archive system; separate repo, separate host).
Deliberately **separate from the tape host**: a compromise of the tape node must
not reach the vault, and OpenBao's TLS/unseal lifecycle is cleaner on its
own box.
Unlike the other stacks here there is **no Caddy / Let's Encrypt** — a secrets
store terminates TLS itself and is reached over the **LAN**, not the public
internet.
- **Native TLS** on the listener — self-signed by default, or a CA-signed cert
from your Smallstep CA over ACME.
- **Integrated raft storage** — snapshot-based DR.
- **Swap disabled** — key material never hits disk. (Not mlock: OpenBao removed
mlock support, so `disable_mlock` must not appear in config.hcl at all.)
- **Manual unseal by default**, or PKCS#11 HSM auto-unseal.
> ⚠️ This vault is the **sole recovery path for encrypted tapes.** Losing the
> OpenBao data *and* the unseal keys/root token loses every encrypted tape. Take
> raft snapshots and store the unseal material out of band (below).
## Required `.env` values
| Variable | Notes |
|----------|-------|
| `OPENBAO_ADDR` | IP/DNS the Kanrisha tape host uses to reach this vault (goes in the cert SAN; you point the daemon at `https://$OPENBAO_ADDR:8200`). |
| `OPENBAO_BIND` | Host interface the API publishes on. Left at the default, deploy.sh narrows it to `OPENBAO_ADDR` when that is an IP (a published port bypasses the INPUT firewall, so this bind is the real restriction — see [Exposure](#exposure)). It must be an address this host actually holds — deploy.sh checks that before writing anything and refuses otherwise (`SKIP_BIND_CHECK=1` to override). |
| `OPENBAO_TLS_SANS` | **Extra** SANs beyond `OPENBAO_ADDR` + loopback (which deploy.sh always adds). Read from the environment at deploy time — `export` it before running deploy.sh. |
| `OPENBAO_TAG` | OpenBao image tag (pin it). |
| `OPENBAO_HSM_PIN` | Only for PKCS#11 auto-unseal. Leave blank for manual unseal. |
See [`.env.example`](.env.example) for the full list.
## Deploy
```bash
./automations.sh # Deploy on this host → deploy: openbao
# or, non-interactive:
OPENBAO_ADDR=10.0.0.10 SKIP_PROMPTS=1 bash deployments/openbao/deploy.sh
```
`deploy.sh` installs Docker, generates a self-signed TLS cert (if none present),
seeds `.env`, disables swap (keeping key material off disk), narrows the API
bind to the LAN IP,
opens `8200/tcp`, aligns file/volume ownership to the container's UID, and brings
the stack up. OpenBao starts **sealed** — initialise + unseal once (do this
immediately; an uninitialised vault reachable on the LAN can be init'd by anyone
who connects):
```bash
docker compose exec -e BAO_ADDR=https://127.0.0.1:8200 openbao \
bao operator init -tls-skip-verify # prints 5 unseal keys + root token
docker compose exec -e BAO_ADDR=https://127.0.0.1:8200 openbao \
bao operator unseal -tls-skip-verify <key> # x3, three different keys
```
**Store the unseal keys + root token out of band** — ideally age-encrypted with
your backup recipient (`globals/age-pubkey.txt`), never on this host.
### Re-running after a wrong address
`deploy.sh` is idempotent, but two things it writes are deliberately **sticky**:
`.env` (never overwritten) and `tls/tls.{crt,key}` (never regenerated over an
existing pair). So passing a corrected `OPENBAO_ADDR` to a re-run does *not* move
the vault — the cert keeps the old SAN, and the old `.env` usually still decides
the bind. Which `OPENBAO_BIND` wins follows Compose's own precedence:
| How `OPENBAO_BIND` is set | What Compose uses |
|---------------------------|-------------------|
| Exported into deploy.sh's environment (`OPENBAO_BIND=… bash deploy.sh`, or answering the bind prompt — `automations.sh` passes answers via `env VAR=…`) | the **environment** value; `.env` is ignored for this run, so a later hand-run `docker compose up -d` can bind somewhere else |
| Derived by deploy.sh (bind prompt left blank → narrowed to `OPENBAO_ADDR`) | the **`.env`** value, since that assignment is never exported |
| Absent from both | `0.0.0.0` — **all interfaces** |
deploy.sh warns on each of those mismatches. To actually change the address:
```bash
cd /srv/openbao
docker compose down # keeps the raft volume
sed -i 's/OLD_IP/NEW_IP/g' .env # OPENBAO_ADDR + OPENBAO_BIND
rm -f tls/tls.crt tls/tls.key # force a new SAN
# then re-run deploy.sh with the corrected OPENBAO_ADDR
```
A bind address the host does not hold is caught up front; without that check
Docker fails the `up` with `cannot assign requested address` only *after* the
bad value is already in `.env` and the cert.
## TLS
- **Self-signed (default):** `deploy.sh` runs [`gen-tls.sh`](gen-tls.sh) to
create `./tls/tls.{crt,key}` with `OPENBAO_ADDR` in the SAN. Hand `tls.crt` to
the Kanrisha daemon as `[encryption.openbao].ca_cert`.
- **Smallstep CA over ACME (option):** issue a cert from your `step-ca` and drop
it in `./tls` instead — `gen-tls.sh` then no-ops. e.g. with the `step` client:
```bash
step ca certificate "$OPENBAO_ADDR" ./tls/tls.crt ./tls/tls.key \
--provisioner acme --acme https://ca.lan/acme/acme/directory
# renew on a timer: step ca renew --daemon ./tls/tls.crt ./tls/tls.key
```
Give the Kanrisha daemon your Smallstep **root** as `ca_cert` (then it
trusts the vault without `-tls-skip-verify`).
## Auto-unseal (optional)
Default is manual unseal after each restart. For hands-off restarts, enable the
`seal "pkcs11"` stanza in [`config.hcl`](config.hcl), mount the PKCS#11 module +
device into the `openbao` service, and set `OPENBAO_HSM_PIN` in `.env`.
## Bootstrap for Kanrisha
Once unsealed, enable KV v2 + AppRole and seed the domain key with the Kanrisha
bootstrap script (from the Kanrisha repo, `deploy/openbao/bootstrap.sh`). It calls
the `bao` CLI directly, so run it **from a host that has `bao`** (the Kanrisha
host or your workstation) pointed at this vault — this vault host only ships
Docker. Copy `tls.crt` to that host first and pass it as `BAO_CACERT`:
```bash
BAO_ADDR=https://$OPENBAO_ADDR:8200 BAO_CACERT=/path/to/openbao-ca.crt \
bash bootstrap.sh # export BAO_TOKEN beforehand rather than passing it inline
```
It enables the `kanrisha-tape` KV-v2 mount (with **effectively-unlimited**
`max_versions` so a key rotation never orphans old tapes), creates the AppRole +
policy, seeds the domain key, and prints the `[encryption.openbao]` block for the
Kanrisha config. Point the daemon at `address = "https://$OPENBAO_ADDR:8200"`.
## Exposure
This is a secrets store on the LAN, not a public service. Two things to know:
- **A Docker-published port is DNAT'd and bypasses the host `INPUT` firewall.** A
source-restricting rule on `INPUT` does **not** gate `:8200`. The real interface
restriction is the publish bind: `deploy.sh` defaults `OPENBAO_BIND` to
`OPENBAO_ADDR` when that is an IP, so the API listens only on that LAN IP. To
restrict by *source host*, use a `FORWARD`/`DOCKER-USER` rule or network
segmentation, not `INPUT`.
- **On a Proxmox guest, the per-VM firewall filters before the guest sees
anything.** If the VM's NIC has `firewall=1` and the datacenter firewall is
enabled, a Docker-published port needs its own `IN ACCEPT -p tcp -dport 8200`
rule in `/etc/pve/firewall/<vmid>.fw`. SSH working does **not** prove the path
— it only proves there is a rule for 22. The give-away is that everything
inside the guest looks perfect (`curl` to the bind address answers, the DNAT
and `FORWARD` jumps are present) while `tcpdump -ni eth0 'tcp port 8200'`
captures zero packets during a failed connection.
- **mTLS (optional):** to require the tape host to present a client cert, enable
the `tls_require_and_verify_client_cert` stanza in [`config.hcl`](config.hcl)
and issue the tape host a client cert from the same CA.
And **initialise the vault immediately** after `deploy.sh` — an uninitialised
vault reachable on the LAN can be `bao operator init`'d by anyone who connects,
handing them the root token and unseal keys.
### Reaching it without LAN access
The API publishes on `OPENBAO_BIND` only, so a browser on another subnet (or
behind a firewall you do not control) cannot reach it. Tunnel over SSH instead
of widening the publish — the generated cert already carries `DNS:localhost`
and `IP:127.0.0.1` in its SANs, so it validates as-is:
```bash
ssh -L 8200:<bind-addr>:8200 root@<host>
# then browse https://localhost:8200
```
This is also the better way to do the first `operator init`: the unseal keys and
root token are shown in your browser instead of a root shell's scrollback.
## Updating
`update.sh` is installed alongside the stack. It is **not** run by `deploy.sh`,
because upgrading a live vault seals it — that is an operator's decision, not a
deploy step.
```bash
cd /srv/openbao
bash update.sh check # declared / running / latest + seal state
bash update.sh update # snapshot, then upgrade (prompts)
```
Unlike the copyparty and ergo updaters in this repo, this one **will not update
on a schedule by default**. Those services come back by themselves; OpenBao
comes back *sealed*, so an unattended 03:00 update would take the vault offline
until someone arrives with three unseal keys. `update.sh install` therefore
schedules a daily **check** (`UPDATE_POLICY=notify`). `UPDATE_POLICY=auto` opts
into unattended updates and is still refused unless a `seal` stanza is
configured — only auto-unseal makes the vault come back on its own.
Before it changes anything, `update` requires that the container is running, the
vault is **unsealed** (a sealed vault cannot produce a snapshot, so there would
be no rollback plan), and a token with `sys/storage/raft/snapshot` — the root
token works. It **prompts** for that token with echo off; never pass it as
`BAO_TOKEN=... bash update.sh`, which puts it in your shell history and in
`/proc/<pid>/cmdline`, where every user on the box can read it. For unattended
use set `BAO_TOKEN_FILE` to a `0600` file. It writes the snapshot to `/var/backups/openbao` at `0600`,
verifies it is a valid non-empty gzip archive, and refuses to continue if it is
not. **Copy it off the host**: it is the rollback plan, and OpenBao's upgrade
guide is explicit that reverting the image alone does not roll back the data
store. `SKIP_SNAPSHOT=1` exists and says loudly what you are giving up.
If the pull or the start fails, the `OPENBAO_TAG` pin is rolled back and the
previous version is started again — still sealed.
`update.sh` also refuses to cross into 2.7.x while an active built-in
`seal "pkcs11"` stanza is present: that stanza is *removed* in 2.7.0, not merely
deprecated, so the vault would come up with no way to unseal at all. Migrate to
the external `plugin "kms" "pkcs11"` first.
## Backup / DR
The vault is the sole recovery path for encrypted tapes — back it up:
Snapshot save/restore are **token-gated** (`sys/storage/raft/snapshot` is
sudo-capable) — they need a token that has that path (the root token works, or
mint a dedicated snapshot-policy token). The container has no ambient token.
**Do not put the token on a command line.** `docker ... -e BAO_TOKEN=<value>`
places it in the docker process's argv, and `/proc/<pid>/cmdline` is readable by
every user on the host — so it leaks further than your shell history does. Use
`update.sh snapshot`, which prompts with echo off and passes the token to the
container over stdin:
```bash
cd /srv/openbao
bash update.sh snapshot # prompts; writes to /var/backups/openbao, 0600
```
Then age-encrypt it and move it off-box:
```bash
age -r "$(cat /path/to/globals/age-pubkey.txt)" \
< /var/backups/openbao/openbao-pre-manual-*.snap \
> "openbao-$(date +%F).snap.age"
```
Restore (into a fresh, unsealed vault): decrypt, copy the raw `.snap` in, and
apply it (also token-gated):
```bash
age -d -i ~/.age/key.txt openbao-DATE.snap.age > openbao.snap
docker compose cp openbao.snap openbao:/tmp/openbao.snap
# Same argv problem as above, so hand the token over stdin rather than -e:
read -rs -p "token: " T; echo
printf '%s\n' "$T" | docker compose exec -T openbao sh -c \
'read -r t; BAO_TOKEN="$t" exec bao operator raft snapshot restore \
-address=https://127.0.0.1:8200 -tls-skip-verify /tmp/openbao.snap'
unset T
```
Snapshots do **not** contain the unseal keys or root token — you still need
those to unseal a restored vault, which is why they are stored separately.