`BAO_TOKEN=<root> bash update.sh update` -- which the previous commit and the pre-existing DR runbook both recommended -- leaks the token twice over. It goes into the shell history, and `docker compose exec -e BAO_TOKEN=<value>` puts it in the docker process's argv, where /proc/<pid>/cmdline makes it readable by every user on the host. The second is the worse one and is not fixed by anything the operator does at their prompt. The token now reaches the container over STDIN and is never an argument to anything: the container's own shell reads one line, exports it, and execs bao. Verified locally that the child process sees the exact value -- including &, | and a backslash -- while its argv contains zero occurrences of it. How a token is supplied, in order: an already-exported BAO_TOKEN; a 0600 file named by the new BAO_TOKEN_FILE; otherwise an echo-off prompt. With no token and no terminal it dies telling the operator to use BAO_TOKEN_FILE rather than an inline assignment, and says why. Added a `snapshot` subcommand so the DR runbook no longer needs an inline pipeline at all. That runbook and the deploy banner both carried the -e form before this deployment had an updater, so both are corrected: taking a backup is now `bash update.sh snapshot`, and restore keeps the stdin shape rather than -e. The Kanrisha bootstrap lines say to export the token first instead of passing it inline. The conf file still refuses to hold a token, but now points at BAO_TOKEN_FILE as the unattended answer -- and suggests a snapshot-policy token rather than the root token for it. Not verified without a live host: that `docker compose exec -T` forwards stdin as expected. If it does not, the snapshot fails loudly with a permission error and the upgrade aborts before touching anything, which is the safe direction. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
255 lines
12 KiB
Markdown
255 lines
12 KiB
Markdown
# openbao
|
|
|
|
Hardened [OpenBao](https://openbao.org) — the tape-encryption key store for
|
|
**Kanrisha** (the LTO tape-archive system; separate repo, separate host).
|
|
Deliberately **separate from the tape host**: a compromise of the tape node must
|
|
not reach the vault, and OpenBao's TLS/unseal lifecycle is cleaner on its
|
|
own box.
|
|
|
|
Unlike the other stacks here there is **no Caddy / Let's Encrypt** — a secrets
|
|
store terminates TLS itself and is reached over the **LAN**, not the public
|
|
internet.
|
|
|
|
- **Native TLS** on the listener — self-signed by default, or a CA-signed cert
|
|
from your Smallstep CA over ACME.
|
|
- **Integrated raft storage** — snapshot-based DR.
|
|
- **Swap disabled** — key material never hits disk. (Not mlock: OpenBao removed
|
|
mlock support, so `disable_mlock` must not appear in config.hcl at all.)
|
|
- **Manual unseal by default**, or PKCS#11 HSM auto-unseal.
|
|
|
|
> ⚠️ This vault is the **sole recovery path for encrypted tapes.** Losing the
|
|
> OpenBao data *and* the unseal keys/root token loses every encrypted tape. Take
|
|
> raft snapshots and store the unseal material out of band (below).
|
|
|
|
## Required `.env` values
|
|
|
|
| Variable | Notes |
|
|
|----------|-------|
|
|
| `OPENBAO_ADDR` | IP/DNS the Kanrisha tape host uses to reach this vault (goes in the cert SAN; you point the daemon at `https://$OPENBAO_ADDR:8200`). |
|
|
| `OPENBAO_BIND` | Host interface the API publishes on. Left at the default, deploy.sh narrows it to `OPENBAO_ADDR` when that is an IP (a published port bypasses the INPUT firewall, so this bind is the real restriction — see [Exposure](#exposure)). It must be an address this host actually holds — deploy.sh checks that before writing anything and refuses otherwise (`SKIP_BIND_CHECK=1` to override). |
|
|
| `OPENBAO_TLS_SANS` | **Extra** SANs beyond `OPENBAO_ADDR` + loopback (which deploy.sh always adds). Read from the environment at deploy time — `export` it before running deploy.sh. |
|
|
| `OPENBAO_TAG` | OpenBao image tag (pin it). |
|
|
| `OPENBAO_HSM_PIN` | Only for PKCS#11 auto-unseal. Leave blank for manual unseal. |
|
|
|
|
See [`.env.example`](.env.example) for the full list.
|
|
|
|
## Deploy
|
|
|
|
```bash
|
|
./automations.sh # Deploy on this host → deploy: openbao
|
|
# or, non-interactive:
|
|
OPENBAO_ADDR=10.0.0.10 SKIP_PROMPTS=1 bash deployments/openbao/deploy.sh
|
|
```
|
|
|
|
`deploy.sh` installs Docker, generates a self-signed TLS cert (if none present),
|
|
seeds `.env`, disables swap (keeping key material off disk), narrows the API
|
|
bind to the LAN IP,
|
|
opens `8200/tcp`, aligns file/volume ownership to the container's UID, and brings
|
|
the stack up. OpenBao starts **sealed** — initialise + unseal once (do this
|
|
immediately; an uninitialised vault reachable on the LAN can be init'd by anyone
|
|
who connects):
|
|
|
|
```bash
|
|
docker compose exec -e BAO_ADDR=https://127.0.0.1:8200 openbao \
|
|
bao operator init -tls-skip-verify # prints 5 unseal keys + root token
|
|
docker compose exec -e BAO_ADDR=https://127.0.0.1:8200 openbao \
|
|
bao operator unseal -tls-skip-verify <key> # x3, three different keys
|
|
```
|
|
|
|
**Store the unseal keys + root token out of band** — ideally age-encrypted with
|
|
your backup recipient (`globals/age-pubkey.txt`), never on this host.
|
|
|
|
### Re-running after a wrong address
|
|
|
|
`deploy.sh` is idempotent, but two things it writes are deliberately **sticky**:
|
|
`.env` (never overwritten) and `tls/tls.{crt,key}` (never regenerated over an
|
|
existing pair). So passing a corrected `OPENBAO_ADDR` to a re-run does *not* move
|
|
the vault — the cert keeps the old SAN, and the old `.env` usually still decides
|
|
the bind. Which `OPENBAO_BIND` wins follows Compose's own precedence:
|
|
|
|
| How `OPENBAO_BIND` is set | What Compose uses |
|
|
|---------------------------|-------------------|
|
|
| Exported into deploy.sh's environment (`OPENBAO_BIND=… bash deploy.sh`, or answering the bind prompt — `automations.sh` passes answers via `env VAR=…`) | the **environment** value; `.env` is ignored for this run, so a later hand-run `docker compose up -d` can bind somewhere else |
|
|
| Derived by deploy.sh (bind prompt left blank → narrowed to `OPENBAO_ADDR`) | the **`.env`** value, since that assignment is never exported |
|
|
| Absent from both | `0.0.0.0` — **all interfaces** |
|
|
|
|
deploy.sh warns on each of those mismatches. To actually change the address:
|
|
|
|
```bash
|
|
cd /srv/openbao
|
|
docker compose down # keeps the raft volume
|
|
sed -i 's/OLD_IP/NEW_IP/g' .env # OPENBAO_ADDR + OPENBAO_BIND
|
|
rm -f tls/tls.crt tls/tls.key # force a new SAN
|
|
# then re-run deploy.sh with the corrected OPENBAO_ADDR
|
|
```
|
|
|
|
A bind address the host does not hold is caught up front; without that check
|
|
Docker fails the `up` with `cannot assign requested address` only *after* the
|
|
bad value is already in `.env` and the cert.
|
|
|
|
## TLS
|
|
|
|
- **Self-signed (default):** `deploy.sh` runs [`gen-tls.sh`](gen-tls.sh) to
|
|
create `./tls/tls.{crt,key}` with `OPENBAO_ADDR` in the SAN. Hand `tls.crt` to
|
|
the Kanrisha daemon as `[encryption.openbao].ca_cert`.
|
|
- **Smallstep CA over ACME (option):** issue a cert from your `step-ca` and drop
|
|
it in `./tls` instead — `gen-tls.sh` then no-ops. e.g. with the `step` client:
|
|
|
|
```bash
|
|
step ca certificate "$OPENBAO_ADDR" ./tls/tls.crt ./tls/tls.key \
|
|
--provisioner acme --acme https://ca.lan/acme/acme/directory
|
|
# renew on a timer: step ca renew --daemon ./tls/tls.crt ./tls/tls.key
|
|
```
|
|
|
|
Give the Kanrisha daemon your Smallstep **root** as `ca_cert` (then it
|
|
trusts the vault without `-tls-skip-verify`).
|
|
|
|
## Auto-unseal (optional)
|
|
|
|
Default is manual unseal after each restart. For hands-off restarts, enable the
|
|
`seal "pkcs11"` stanza in [`config.hcl`](config.hcl), mount the PKCS#11 module +
|
|
device into the `openbao` service, and set `OPENBAO_HSM_PIN` in `.env`.
|
|
|
|
## Bootstrap for Kanrisha
|
|
|
|
Once unsealed, enable KV v2 + AppRole and seed the domain key with the Kanrisha
|
|
bootstrap script (from the Kanrisha repo, `deploy/openbao/bootstrap.sh`). It calls
|
|
the `bao` CLI directly, so run it **from a host that has `bao`** (the Kanrisha
|
|
host or your workstation) pointed at this vault — this vault host only ships
|
|
Docker. Copy `tls.crt` to that host first and pass it as `BAO_CACERT`:
|
|
|
|
```bash
|
|
BAO_ADDR=https://$OPENBAO_ADDR:8200 BAO_CACERT=/path/to/openbao-ca.crt \
|
|
bash bootstrap.sh # export BAO_TOKEN beforehand rather than passing it inline
|
|
```
|
|
|
|
It enables the `kanrisha-tape` KV-v2 mount (with **effectively-unlimited**
|
|
`max_versions` so a key rotation never orphans old tapes), creates the AppRole +
|
|
policy, seeds the domain key, and prints the `[encryption.openbao]` block for the
|
|
Kanrisha config. Point the daemon at `address = "https://$OPENBAO_ADDR:8200"`.
|
|
|
|
## Exposure
|
|
|
|
This is a secrets store on the LAN, not a public service. Two things to know:
|
|
|
|
- **A Docker-published port is DNAT'd and bypasses the host `INPUT` firewall.** A
|
|
source-restricting rule on `INPUT` does **not** gate `:8200`. The real interface
|
|
restriction is the publish bind: `deploy.sh` defaults `OPENBAO_BIND` to
|
|
`OPENBAO_ADDR` when that is an IP, so the API listens only on that LAN IP. To
|
|
restrict by *source host*, use a `FORWARD`/`DOCKER-USER` rule or network
|
|
segmentation, not `INPUT`.
|
|
- **On a Proxmox guest, the per-VM firewall filters before the guest sees
|
|
anything.** If the VM's NIC has `firewall=1` and the datacenter firewall is
|
|
enabled, a Docker-published port needs its own `IN ACCEPT -p tcp -dport 8200`
|
|
rule in `/etc/pve/firewall/<vmid>.fw`. SSH working does **not** prove the path
|
|
— it only proves there is a rule for 22. The give-away is that everything
|
|
inside the guest looks perfect (`curl` to the bind address answers, the DNAT
|
|
and `FORWARD` jumps are present) while `tcpdump -ni eth0 'tcp port 8200'`
|
|
captures zero packets during a failed connection.
|
|
- **mTLS (optional):** to require the tape host to present a client cert, enable
|
|
the `tls_require_and_verify_client_cert` stanza in [`config.hcl`](config.hcl)
|
|
and issue the tape host a client cert from the same CA.
|
|
|
|
And **initialise the vault immediately** after `deploy.sh` — an uninitialised
|
|
vault reachable on the LAN can be `bao operator init`'d by anyone who connects,
|
|
handing them the root token and unseal keys.
|
|
|
|
### Reaching it without LAN access
|
|
|
|
The API publishes on `OPENBAO_BIND` only, so a browser on another subnet (or
|
|
behind a firewall you do not control) cannot reach it. Tunnel over SSH instead
|
|
of widening the publish — the generated cert already carries `DNS:localhost`
|
|
and `IP:127.0.0.1` in its SANs, so it validates as-is:
|
|
|
|
```bash
|
|
ssh -L 8200:<bind-addr>:8200 root@<host>
|
|
# then browse https://localhost:8200
|
|
```
|
|
|
|
This is also the better way to do the first `operator init`: the unseal keys and
|
|
root token are shown in your browser instead of a root shell's scrollback.
|
|
|
|
## Updating
|
|
|
|
`update.sh` is installed alongside the stack. It is **not** run by `deploy.sh`,
|
|
because upgrading a live vault seals it — that is an operator's decision, not a
|
|
deploy step.
|
|
|
|
```bash
|
|
cd /srv/openbao
|
|
bash update.sh check # declared / running / latest + seal state
|
|
bash update.sh update # snapshot, then upgrade (prompts)
|
|
```
|
|
|
|
Unlike the copyparty and ergo updaters in this repo, this one **will not update
|
|
on a schedule by default**. Those services come back by themselves; OpenBao
|
|
comes back *sealed*, so an unattended 03:00 update would take the vault offline
|
|
until someone arrives with three unseal keys. `update.sh install` therefore
|
|
schedules a daily **check** (`UPDATE_POLICY=notify`). `UPDATE_POLICY=auto` opts
|
|
into unattended updates and is still refused unless a `seal` stanza is
|
|
configured — only auto-unseal makes the vault come back on its own.
|
|
|
|
Before it changes anything, `update` requires that the container is running, the
|
|
vault is **unsealed** (a sealed vault cannot produce a snapshot, so there would
|
|
be no rollback plan), and a token with `sys/storage/raft/snapshot` — the root
|
|
token works. It **prompts** for that token with echo off; never pass it as
|
|
`BAO_TOKEN=... bash update.sh`, which puts it in your shell history and in
|
|
`/proc/<pid>/cmdline`, where every user on the box can read it. For unattended
|
|
use set `BAO_TOKEN_FILE` to a `0600` file. It writes the snapshot to `/var/backups/openbao` at `0600`,
|
|
verifies it is a valid non-empty gzip archive, and refuses to continue if it is
|
|
not. **Copy it off the host**: it is the rollback plan, and OpenBao's upgrade
|
|
guide is explicit that reverting the image alone does not roll back the data
|
|
store. `SKIP_SNAPSHOT=1` exists and says loudly what you are giving up.
|
|
|
|
If the pull or the start fails, the `OPENBAO_TAG` pin is rolled back and the
|
|
previous version is started again — still sealed.
|
|
|
|
`update.sh` also refuses to cross into 2.7.x while an active built-in
|
|
`seal "pkcs11"` stanza is present: that stanza is *removed* in 2.7.0, not merely
|
|
deprecated, so the vault would come up with no way to unseal at all. Migrate to
|
|
the external `plugin "kms" "pkcs11"` first.
|
|
|
|
## Backup / DR
|
|
|
|
The vault is the sole recovery path for encrypted tapes — back it up:
|
|
|
|
Snapshot save/restore are **token-gated** (`sys/storage/raft/snapshot` is
|
|
sudo-capable) — they need a token that has that path (the root token works, or
|
|
mint a dedicated snapshot-policy token). The container has no ambient token.
|
|
|
|
**Do not put the token on a command line.** `docker ... -e BAO_TOKEN=<value>`
|
|
places it in the docker process's argv, and `/proc/<pid>/cmdline` is readable by
|
|
every user on the host — so it leaks further than your shell history does. Use
|
|
`update.sh snapshot`, which prompts with echo off and passes the token to the
|
|
container over stdin:
|
|
|
|
```bash
|
|
cd /srv/openbao
|
|
bash update.sh snapshot # prompts; writes to /var/backups/openbao, 0600
|
|
```
|
|
|
|
Then age-encrypt it and move it off-box:
|
|
|
|
```bash
|
|
age -r "$(cat /path/to/globals/age-pubkey.txt)" \
|
|
< /var/backups/openbao/openbao-pre-manual-*.snap \
|
|
> "openbao-$(date +%F).snap.age"
|
|
```
|
|
|
|
Restore (into a fresh, unsealed vault): decrypt, copy the raw `.snap` in, and
|
|
apply it (also token-gated):
|
|
|
|
```bash
|
|
age -d -i ~/.age/key.txt openbao-DATE.snap.age > openbao.snap
|
|
docker compose cp openbao.snap openbao:/tmp/openbao.snap
|
|
# Same argv problem as above, so hand the token over stdin rather than -e:
|
|
read -rs -p "token: " T; echo
|
|
printf '%s\n' "$T" | docker compose exec -T openbao sh -c \
|
|
'read -r t; BAO_TOKEN="$t" exec bao operator raft snapshot restore \
|
|
-address=https://127.0.0.1:8200 -tls-skip-verify /tmp/openbao.snap'
|
|
unset T
|
|
```
|
|
|
|
Snapshots do **not** contain the unseal keys or root token — you still need
|
|
those to unseal a restored vault, which is why they are stored separately.
|