mirror of
https://tangled.org/evan.jarrett.net/at-container-registry
synced 2026-09-28 05:04:15 +00:00
buildOCILayout already removes its scan dir on every error path, and syft.go
defers the stereoscope generator's Cleanup. What neither can do is clean up
after a process that dies mid-scan: the deferred call never runs, and nothing
afterwards ever looks at what was left. Every restart therefore leaks the
in-flight layout and extraction permanently, and a restart is routine — a
deploy is one.
On seamark-hold that reached 8.8 GB of orphaned scan-*, syft-scan-* and
syft-cataloger-* directories under a 20 GB disk, at which point the disk was
97% full and scans began failing on it:
failed to load OCI image: unable to populate layer cache
dir="/var/lib/seamark/scanner/tmp/syft-scan-1546187834/..."
: no space left on device
failed to download layer 5: failed to write blob:
write /var/lib/seamark/scanner/tmp/scan-4160414849/blobs/sha256/...
: no space left on device
The leaked directories cluster at the scanner's restart timestamps, which is
what identifies the killed process rather than the error paths as the source.
Manual removal reclaimed 8.8 GB and took the disk from 97% to 50%.
Startup is where this belongs: it is the one moment the previous process is
known to be gone, and it is immediately after the event that caused the leak.
The sweep runs in WorkerPool.Start after TMPDIR is set and before any worker
can dequeue, so nothing it removes can be work in progress here.
Three constraints shape what it will touch:
- Only the three per-job prefixes, only as direct children, only
directories. The Grype database lives beside the tmp dir at
<parent>/vulndb and go-getter unpacks into grype-dl underneath it; both
are state the scanner needs and neither matches a prefix. The prefixes now
have one definition each, used by both the creator and the sweeper, so
renaming a directory cannot silently take it out of the sweep's scope.
- An age threshold, vuln.sweep_max_age, default 1h. A second scanner sharing
the directory has an in-flight scan-* dir that is minutes old, and
scanner.job_timeout is 8m, so an hour clears both with room to spare. 0
disables the sweep rather than removing a peer's live work.
- Nothing is fatal. A stat or removal failure is a WARN and the sweep moves
on, so a permission problem in the tmp dir cannot keep the scanner from
starting.
The sweep only runs at startup, so a scanner that is killed twice between
deploys carries the first leak until its next restart. That is the tradeoff
for never racing a live peer; a periodic sweep would be the follow-up if
processes ever live long enough for it to matter.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TA9D4DjaLZTvzQ7dJbu4eg
56 lines
2.5 KiB
Cheetah
56 lines
2.5 KiB
Cheetah
version: "0.1"
|
|
log_level: info
|
|
log_shipper:
|
|
backend: ""
|
|
url: ""
|
|
batch_size: 100
|
|
flush_interval: 5s
|
|
username: ""
|
|
password: ""
|
|
server:
|
|
addr: :9090
|
|
hold:
|
|
url: "ws://localhost:8080"
|
|
secret: "{{.ScannerSecret}}"
|
|
scanner:
|
|
# One worker, deliberately. Two reasons, and the first one may go away:
|
|
# - The hold's proactive dispatch loop gates on waitForCapacity() and
|
|
# hands out one job at a time hold-wide, so a second worker only ever
|
|
# receives work when two pushes coincide.
|
|
# - Peak RSS is per concurrent scan. Two concurrent scans of a
|
|
# node:22-class image measured 687 MiB with the 512 MiB GOMEMLIMIT in
|
|
# force and 1357 MiB without it. On a host this size that is the
|
|
# difference between working and OOM-killing the hold.
|
|
# Raise this only together with MemoryMax in the unit file and the memory
|
|
# available on the host.
|
|
workers: 1
|
|
queue_size: 100
|
|
# Must stay below the hold's 10m scanning timeout, which it measures from
|
|
# the "started" message this scanner sends when a worker picks the job up.
|
|
# Whichever fires first decides: this one stops the scan and tells the hold
|
|
# why, the hold's marks the row failed and re-dispatches work that is still
|
|
# running. Two minutes of margin for the terminal message to land.
|
|
#
|
|
# Note what the deadline does not cover: stereoscope's layer extraction
|
|
# takes no context and cannot be interrupted, and it is ~81% of a scan. The
|
|
# bound on that stage is max_image_size below, not this.
|
|
job_timeout: 8m
|
|
vuln:
|
|
enabled: true
|
|
db_path: "{{.BasePath}}/scanner/vulndb"
|
|
# Must be on disk, not tmpfs. Layers are extracted here and amplify
|
|
# roughly 3.8x over their compressed size (measured: node:22, 389 MiB ->
|
|
# 1493 MiB), so a tmpfs would spend host memory to hold them.
|
|
tmp_dir: "{{.BasePath}}/scanner/tmp"
|
|
# 512 MiB compressed. The shipped default is 2 GiB, which no small host
|
|
# can survive: peak RSS tracks image size, and 389 MiB compressed already
|
|
# reached 561 MiB RSS with the memory limit applied. Images above this are
|
|
# rejected before any blob is downloaded.
|
|
max_image_size: 536870912
|
|
# Reclaim scan directories left behind by a scanner that was killed
|
|
# mid-scan (a restart or a deploy). Those never run their own cleanup, and
|
|
# on this host they reached 8.8 GB on a 20 GB disk before scans started
|
|
# failing with "no space left on device". Above the 8m job timeout, so a
|
|
# scan in flight elsewhere in the same directory is never touched.
|
|
sweep_max_age: 1h
|