mirror of
https://tangled.org/evan.jarrett.net/at-container-registry
synced 2026-09-29 05:25:35 +00:00
scanner: reclaim scan directories a killed process left behind
buildOCILayout already removes its scan dir on every error path, and syft.go
defers the stereoscope generator's Cleanup. What neither can do is clean up
after a process that dies mid-scan: the deferred call never runs, and nothing
afterwards ever looks at what was left. Every restart therefore leaks the
in-flight layout and extraction permanently, and a restart is routine — a
deploy is one.
On seamark-hold that reached 8.8 GB of orphaned scan-*, syft-scan-* and
syft-cataloger-* directories under a 20 GB disk, at which point the disk was
97% full and scans began failing on it:
failed to load OCI image: unable to populate layer cache
dir="/var/lib/seamark/scanner/tmp/syft-scan-1546187834/..."
: no space left on device
failed to download layer 5: failed to write blob:
write /var/lib/seamark/scanner/tmp/scan-4160414849/blobs/sha256/...
: no space left on device
The leaked directories cluster at the scanner's restart timestamps, which is
what identifies the killed process rather than the error paths as the source.
Manual removal reclaimed 8.8 GB and took the disk from 97% to 50%.
Startup is where this belongs: it is the one moment the previous process is
known to be gone, and it is immediately after the event that caused the leak.
The sweep runs in WorkerPool.Start after TMPDIR is set and before any worker
can dequeue, so nothing it removes can be work in progress here.
Three constraints shape what it will touch:
- Only the three per-job prefixes, only as direct children, only
directories. The Grype database lives beside the tmp dir at
<parent>/vulndb and go-getter unpacks into grype-dl underneath it; both
are state the scanner needs and neither matches a prefix. The prefixes now
have one definition each, used by both the creator and the sweeper, so
renaming a directory cannot silently take it out of the sweep's scope.
- An age threshold, vuln.sweep_max_age, default 1h. A second scanner sharing
the directory has an in-flight scan-* dir that is minutes old, and
scanner.job_timeout is 8m, so an hour clears both with room to spare. 0
disables the sweep rather than removing a peer's live work.
- Nothing is fatal. A stat or removal failure is a WARN and the sweep moves
on, so a permission problem in the tmp dir cannot keep the scanner from
starting.
The sweep only runs at startup, so a scanner that is killed twice between
deploys carries the first leak until its next restart. That is the tradeoff
for never racing a live peer; a periodic sweep would be the follow-up if
processes ever live long enough for it to matter.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TA9D4DjaLZTvzQ7dJbu4eg
This commit is contained in:
co-authored by
Claude Opus 5
parent
bd208661c6
commit
265e533ad3
@@ -110,6 +110,7 @@ a YAML file or pure env vars with the `SCANNER_` prefix. Run with
|
||||
| `vuln.db_path` | `SCANNER_VULN_DB_PATH` | `/var/lib/atcr-scanner/vulndb` | Directory for the Grype vulnerability database. |
|
||||
| `vuln.tmp_dir` | `SCANNER_VULN_TMP_DIR` | `/var/lib/atcr-scanner/tmp` | Directory for layer extraction and DB download. Also exported as `TMPDIR`; point it at a large partition, **not** tmpfs. |
|
||||
| `vuln.max_image_size`| `SCANNER_VULN_MAX_IMAGE_SIZE`| `2147483648` (2 GiB) | Max total compressed image size. Larger images are skipped with an error. `0` = no limit. |
|
||||
| `vuln.sweep_max_age`| `SCANNER_VULN_SWEEP_MAX_AGE` | `1h` | Age threshold for the startup sweep of `vuln.tmp_dir`. Leftover `scan-*`, `syft-scan-*` and `syft-cataloger-*` directories older than this are removed before workers start; a scan killed by a restart never cleans up after itself. Keep it above the longest scan so a second scanner sharing the directory keeps its live work. `0` disables the sweep. |
|
||||
| `server.addr` | `SCANNER_SERVER_ADDR` | `:9090` | Listen address for the scanner's health endpoint. |
|
||||
|
||||
Both `hold.url` and `hold.secret` are required; `LoadConfig` errors out if either is
|
||||
@@ -420,6 +421,11 @@ admin endpoint `POST /admin/api/scan-backfill` instead.
|
||||
- **Layer extraction or Grype DB download fails mid-process.** `vuln.tmp_dir` is too
|
||||
small or on tmpfs. Point it at a large persistent partition; the scanner sets
|
||||
`TMPDIR` to this directory.
|
||||
- **`vuln.tmp_dir` fills up with `scan-*` / `syft-scan-*` directories.** These are
|
||||
scans the process was killed in the middle of, which never ran their own cleanup.
|
||||
The startup sweep reclaims them on the next restart; if it is not doing so, check
|
||||
that `vuln.sweep_max_age` is not `0` and look for `Swept scanner tmp dir` in the
|
||||
startup logs.
|
||||
- **SBOM present but no vulnerability counts.** `vuln.enabled` is false on the scanner,
|
||||
or the Grype DB failed to initialize (check startup logs).
|
||||
- **Helm/attestation artifacts show "scanning isn't applied".** Expected — these are
|
||||
|
||||
Reference in New Issue
Block a user