scanner: reclaim scan directories a killed process left behind

buildOCILayout already removes its scan dir on every error path, and syft.go
defers the stereoscope generator's Cleanup. What neither can do is clean up
after a process that dies mid-scan: the deferred call never runs, and nothing
afterwards ever looks at what was left. Every restart therefore leaks the
in-flight layout and extraction permanently, and a restart is routine — a
deploy is one.

On seamark-hold that reached 8.8 GB of orphaned scan-*, syft-scan-* and
syft-cataloger-* directories under a 20 GB disk, at which point the disk was
97% full and scans began failing on it:

  failed to load OCI image: unable to populate layer cache
    dir="/var/lib/seamark/scanner/tmp/syft-scan-1546187834/..."
    : no space left on device

  failed to download layer 5: failed to write blob:
    write /var/lib/seamark/scanner/tmp/scan-4160414849/blobs/sha256/...
    : no space left on device

The leaked directories cluster at the scanner's restart timestamps, which is
what identifies the killed process rather than the error paths as the source.
Manual removal reclaimed 8.8 GB and took the disk from 97% to 50%.

Startup is where this belongs: it is the one moment the previous process is
known to be gone, and it is immediately after the event that caused the leak.
The sweep runs in WorkerPool.Start after TMPDIR is set and before any worker
can dequeue, so nothing it removes can be work in progress here.

Three constraints shape what it will touch:

  - Only the three per-job prefixes, only as direct children, only
    directories. The Grype database lives beside the tmp dir at
    <parent>/vulndb and go-getter unpacks into grype-dl underneath it; both
    are state the scanner needs and neither matches a prefix. The prefixes now
    have one definition each, used by both the creator and the sweeper, so
    renaming a directory cannot silently take it out of the sweep's scope.

  - An age threshold, vuln.sweep_max_age, default 1h. A second scanner sharing
    the directory has an in-flight scan-* dir that is minutes old, and
    scanner.job_timeout is 8m, so an hour clears both with room to spare. 0
    disables the sweep rather than removing a peer's live work.

  - Nothing is fatal. A stat or removal failure is a WARN and the sweep moves
    on, so a permission problem in the tmp dir cannot keep the scanner from
    starting.

The sweep only runs at startup, so a scanner that is killed twice between
deploys carries the first leak until its next restart. That is the tradeoff
for never racing a live peer; a periodic sweep would be the follow-up if
processes ever live long enough for it to matter.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TA9D4DjaLZTvzQ7dJbu4eg
This commit is contained in:
Evan Jarrett
2026-09-08 21:59:58 -05:00
co-authored by Claude Opus 5
parent bd208661c6
commit 265e533ad3
8 changed files with 306 additions and 2 deletions
+6
View File
@@ -110,6 +110,7 @@ a YAML file or pure env vars with the `SCANNER_` prefix. Run with
| `vuln.db_path` | `SCANNER_VULN_DB_PATH` | `/var/lib/atcr-scanner/vulndb` | Directory for the Grype vulnerability database. |
| `vuln.tmp_dir` | `SCANNER_VULN_TMP_DIR` | `/var/lib/atcr-scanner/tmp` | Directory for layer extraction and DB download. Also exported as `TMPDIR`; point it at a large partition, **not** tmpfs. |
| `vuln.max_image_size`| `SCANNER_VULN_MAX_IMAGE_SIZE`| `2147483648` (2 GiB) | Max total compressed image size. Larger images are skipped with an error. `0` = no limit. |
| `vuln.sweep_max_age`| `SCANNER_VULN_SWEEP_MAX_AGE` | `1h` | Age threshold for the startup sweep of `vuln.tmp_dir`. Leftover `scan-*`, `syft-scan-*` and `syft-cataloger-*` directories older than this are removed before workers start; a scan killed by a restart never cleans up after itself. Keep it above the longest scan so a second scanner sharing the directory keeps its live work. `0` disables the sweep. |
| `server.addr` | `SCANNER_SERVER_ADDR` | `:9090` | Listen address for the scanner's health endpoint. |
Both `hold.url` and `hold.secret` are required; `LoadConfig` errors out if either is
@@ -420,6 +421,11 @@ admin endpoint `POST /admin/api/scan-backfill` instead.
- **Layer extraction or Grype DB download fails mid-process.** `vuln.tmp_dir` is too
small or on tmpfs. Point it at a large persistent partition; the scanner sets
`TMPDIR` to this directory.
- **`vuln.tmp_dir` fills up with `scan-*` / `syft-scan-*` directories.** These are
scans the process was killed in the middle of, which never ran their own cleanup.
The startup sweep reclaims them on the next restart; if it is not doing so, check
that `vuln.sweep_max_age` is not `0` and look for `Swept scanner tmp dir` in the
startup logs.
- **SBOM present but no vulnerability counts.** `vuln.enabled` is false on the scanner,
or the Grype DB failed to initialize (check startup logs).
- **Helm/attestation artifacts show "scanning isn't applied".** Expected — these are