Files
at-container-registry/docs/SBOM_SCANNING.md
T
Evan Jarrett 6758996300 add SBOM package diffing, verify hold-service captain records
- diff view gains a Packages tab with added/removed/changed/unchanged
  package tables and purl-derived type/license/upstream links
- captain records verified against the DID's atcr_hold service before
  caching (processor + batch backfill), preventing forged holds
- fix empty-handle updates clobbering cached handles and colliding on
  the UNIQUE constraint
- move fillPrevCIDs into repo.go; DirectRepoOperator is now canonical,
  repomgr kept as a test oracle
- surface read-only crew status in hold selector
- reconcile docs
2026-06-13 12:49:03 -05:00

397 lines
19 KiB
Markdown

# SBOM Scanning and Vulnerability Analysis
ATCR generates Software Bills of Materials (SBOMs) and scans container images for
vulnerabilities. Scanning runs in a separate `atcr-scanner` service that connects to
a hold over a WebSocket, so the hold itself never runs Syft or Grype. Results are
stored as `io.atcr.hold.scan` records in the hold's embedded PDS.
## Overview
- **Separate scanner binary**: Scanning is performed by `atcr-scanner` (the `scanner/`
Go module), not by the hold. The scanner connects out to the hold and pulls jobs.
- **Syft for SBOMs, Grype for vulnerabilities**: Each job runs Syft to produce an
SPDX-JSON SBOM, then Grype to scan that SBOM for CVEs. Grype is enabled by default.
- **WebSocket dispatch**: The hold pushes jobs to connected scanners over
`/xrpc/io.atcr.hold.subscribeScanJobs`. A shared secret authenticates the scanner.
- **ATProto result storage**: Results land as `io.atcr.hold.scan` records in the
hold's embedded PDS, with the SBOM and full Grype report uploaded as PDS blobs.
- **Tier-gated scan-on-push plus proactive rescans**: Pushes from eligible tiers
trigger an immediate scan; the hold also discovers never-scanned manifests and
re-scans stale ones on an interval.
### Tools
- [Anchore Syft](https://github.com/anchore/syft) generates the SBOM. Output format is
SPDX JSON, hardcoded in `scanner/internal/scan/syft.go` (not configurable).
- [Anchore Grype](https://github.com/anchore/grype) scans the SBOM for known
vulnerabilities and produces critical/high/medium/low/total counts plus a full
JSON report with CVE detail.
## Architecture
Three pieces cooperate:
```
io.atcr.hold.subscribeScanJobs (WebSocket, ?secret=...)
┌───────────┐ ◄──────────────────────────────────────────── ┌──────────────┐
│ Hold │ job: {seq, manifestDigest, repo, tier, │ atcr-scanner │
│ (Scan │ config, layers, holdEndpoint, ...} │ (Syft + │
│ Broadcaster)│ ────────────────────────────────────────────► │ Grype) │
│ │ │ │
│ │ result/error/skipped: {seq, sbom, │ │
│ │ ◄──── vulnReport, summary{critical,high,...}} └──────────────┘
└─────┬─────┘
│ stores io.atcr.hold.scan record + SBOM/vuln blobs
Hold embedded PDS (CAR store)
```
1. **Hold (`pkg/hold/pds/scan_broadcaster.go`)** owns the `ScanBroadcaster`. It
persists pending jobs in SQLite (`scan_jobs` table), accepts scanner WebSocket
connections, and dispatches jobs **round-robin** across all connected scanners
using a competing-consumer pattern. It re-dispatches timed-out jobs, and (when a
rescan interval is set) runs background discovery and stale-scan loops. On receiving
a result, the hold uploads the SBOM and vuln report as blobs and writes the
`io.atcr.hold.scan` record.
2. **Scanner (`scanner/` module)** dials the hold's WebSocket, acks jobs, runs the
Syft → Grype pipeline, and sends back a result, error, or skipped message. It keeps
a local **priority queue** so paid tiers jump ahead of free ones (see Scheduling).
3. **AppView** reads the scan records and blobs from the hold's PDS to render
vulnerability badges, SBOM details, and download links in the web UI.
### Why the hold's PDS?
Scan results are stored in the **hold's embedded PDS** rather than the user's PDS:
- No OAuth/service-token plumbing: the hold owns and signs its own records.
- Hold-scoped metadata (scanner version, scan time) stays with the operator.
- Different holds can independently scan the same image for cross-verification.
- The user's PDS stays lean: SBOM and Grype JSON live in hold blob storage.
The trust model is the same as Docker Hub: you trust the hold operator's scanner
version and scan integrity. The hold's DID signs the records, and anyone can re-scan a
digest to verify the result.
## Configuration
### Hold side
The hold's scanner integration is configured under `scanner:` in the hold config
(`pkg/hold/config.go`). Env-var prefix is `HOLD_`.
| YAML key | Env var | Default | Meaning |
|--------------------------|-------------------------------|---------|---------|
| `scanner.secret` | `HOLD_SCANNER_SECRET` | `""` | Shared secret a scanner must present (as `?secret=`) on the WebSocket. **Empty disables scanning entirely** — no scanner can connect and no jobs are dispatched. |
| `scanner.rescan_interval`| `HOLD_SCANNER_RESCAN_INTERVAL`| `168h` | Minimum interval between re-scans of the same manifest. When > 0 the hold runs proactive discovery + stale-scan loops. Set to `0` to disable proactive scanning (push-triggered scans still work). |
```yaml
# config-hold.yaml
scanner:
secret: "a-long-random-shared-secret"
rescan_interval: 168h
```
Whether a push triggers an immediate scan is decided by the quota tier (see
[Scan-on-push tier gate](#scan-on-push-tier-gate)).
### Scanner side
The scanner is configured via Viper (`scanner/internal/config/config.go`); it accepts
a YAML file or pure env vars with the `SCANNER_` prefix. Run with
`SCANNER_HOLD_URL=... SCANNER_HOLD_SECRET=... atcr-scanner serve`.
| YAML key | Env var | Default | Meaning |
|---------------------|------------------------------|----------------------------------|---------|
| `hold.url` | `SCANNER_HOLD_URL` | — (**required**) | WebSocket URL of the hold, e.g. `ws://localhost:8080` or `wss://hold01.atcr.io`. `http(s)` is auto-converted to `ws(s)`. |
| `hold.secret` | `SCANNER_HOLD_SECRET` | — (**required**) | Must match the hold's `scanner.secret`. Sent as `?secret=`. |
| `scanner.workers` | `SCANNER_SCANNER_WORKERS` | `1` | Number of concurrent scan workers. |
| `scanner.queue_size`| `SCANNER_SCANNER_QUEUE_SIZE` | `100` | Max depth of the local priority queue. |
| `vuln.enabled` | `SCANNER_VULN_ENABLED` | `true` | Run Grype after Syft. When false, only the SBOM is produced (no counts). |
| `vuln.db_path` | `SCANNER_VULN_DB_PATH` | `/var/lib/atcr-scanner/vulndb` | Directory for the Grype vulnerability database. |
| `vuln.tmp_dir` | `SCANNER_VULN_TMP_DIR` | `/var/lib/atcr-scanner/tmp` | Directory for layer extraction and DB download. Also exported as `TMPDIR`; point it at a large partition, **not** tmpfs. |
| `vuln.max_image_size`| `SCANNER_VULN_MAX_IMAGE_SIZE`| `2147483648` (2 GiB) | Max total compressed image size. Larger images are skipped with an error. `0` = no limit. |
| `server.addr` | `SCANNER_SERVER_ADDR` | `:9090` | Listen address for the scanner's health endpoint. |
Both `hold.url` and `hold.secret` are required; `LoadConfig` errors out if either is
empty.
```bash
# Minimal scanner invocation (env-only)
SCANNER_HOLD_URL=wss://hold01.atcr.io \
SCANNER_HOLD_SECRET=a-long-random-shared-secret \
./bin/atcr-scanner serve
```
## Scanning Workflow
### 1. Push → scan-on-push tier gate
When an image is pushed and the manifest is recorded, the hold's OCI XRPC handler
(`pkg/hold/oci/xrpc.go`) decides whether to enqueue a scan. Multi-arch manifest lists
and artifacts with a `subject` (attestations) are skipped — they have no scannable
content. For everything else, the tier of the pusher decides:
- **Captain / owner**: always scanned.
- **Quotas disabled** (`quotaMgr == nil` or quotas not enabled): all pushes scanned
(backwards compatible).
- **Quotas enabled**: scanned only if the pusher's tier has `scan_on_push: true`.
In the default config (`pkg/hold/config.go`), `bosun` and `quartermaster` have
`scan_on_push: true`; `deckhand` does not. So a free-tier (deckhand) push is **not**
scanned on push — it gets picked up later by the proactive discovery loop.
### 2. Dispatch
The `ScanBroadcaster.Enqueue` inserts the job into the `scan_jobs` SQLite table
(status `pending`) and immediately tries to dispatch it round-robin to one of the
connected scanners. Jobs survive hold restarts. If no scanner is connected, the job
waits; newly connected scanners drain pending jobs. Assigned-but-unacked jobs time out
after 5 minutes and are re-dispatched; jobs stuck in `processing` for 10 minutes are
marked failed (scanner likely crashed).
### 3. Scan pipeline (scanner)
For each job (`scanner/internal/scan/worker.go`):
1. **Artifact-type check** — if `config.mediaType` is in `unscannableConfigTypes` the
job returns a `SkipError` and the scanner sends a `skipped` message (see below).
2. **Size check** — if total compressed size exceeds `vuln.max_image_size`, the job
fails.
3. **Build OCI layout** — layers are fetched from the hold via presigned URLs and
assembled into an OCI image layout in `vuln.tmp_dir`.
4. **Syft** — generates the SBOM and encodes it to SPDX JSON.
5. **Grype** (if `vuln.enabled`) — scans the SBOM, producing the full JSON report and
a severity summary (critical/high/medium/low/total).
The scanner then sends one of three messages back over the WebSocket: `result`
(SBOM + optional vuln report + summary), `error`, or `skipped` (with a reason).
### 4. Result storage (hold)
On `result` (`scan_broadcaster.go` `handleResult`):
1. Upload the SBOM bytes as a PDS blob (`application/spdx+json`).
2. Upload the Grype report as a PDS blob (`application/vnd.atcr.vulnerabilities+json`).
3. Create an `io.atcr.hold.scan` record (`CreateScanRecord`) keyed by the manifest
digest, referencing both blobs and carrying the severity counts.
4. Mark the `scan_jobs` row `completed`.
On `error`, a failed scan record is written (`NewFailedScanRecord`) and the job is
marked `failed`. On `skipped`, a skipped record is written (`NewSkippedScanRecord`)
and the job is marked `completed`.
## Scan Record Schema
Results are `io.atcr.hold.scan` records in the hold's embedded PDS
(`pkg/atproto/lexicon.go`, `ScanRecord`). The record key is the manifest digest hex
(without the `sha256:` prefix), so there is exactly one scan record per manifest and
re-scans upsert it.
```json
{
"$type": "io.atcr.hold.scan",
"manifest": "at://did:plc:alice123/io.atcr.manifest/abc123...",
"repository": "myapp",
"userDid": "did:plc:alice123",
"sbomBlob": {
"$type": "blob",
"ref": { "$link": "bafkrei..." },
"mimeType": "application/spdx+json",
"size": 51234
},
"vulnReportBlob": {
"$type": "blob",
"ref": { "$link": "bafkrei..." },
"mimeType": "application/vnd.atcr.vulnerabilities+json",
"size": 18567
},
"critical": 2,
"high": 15,
"medium": 42,
"low": 8,
"total": 67,
"scannerVersion": "atcr-scanner-v1.0.0",
"scannedAt": "2026-06-11T12:34:56Z",
"status": "ok",
"reason": ""
}
```
| Field | Notes |
|------------------|-------|
| `manifest` | AT-URI of the scanned manifest in the user's PDS. |
| `userDid` | DID of the image owner. |
| `sbomBlob` | Reference to the SPDX-JSON SBOM in hold blob storage. Absent for failed/skipped scans. |
| `vulnReportBlob` | Reference to the full Grype JSON report. Absent if Grype disabled or scan failed/skipped. |
| `critical`/`high`/`medium`/`low`/`total` | Vulnerability counts from Grype. Zero on failed/skipped scans. |
| `scannerVersion` | Scanner identifier for reproducibility (currently `atcr-scanner-v1.0.0`). |
| `scannedAt` | RFC3339 scan completion timestamp. |
| `status` | `ok`, `failed`, or `skipped`. |
| `reason` | Populated for `failed` (error text) and `skipped` (why it was bypassed). |
### Status field
| Status | Meaning | Rescan behavior |
|-------------|-------------------------------------------------------------------------|-----------------|
| `ok` (or empty) | Scanner produced an SBOM; counts and SBOM blob populated. | Re-scanned on the rescan interval (default 7d). |
| `failed` | Scanner ran but errored (network, OOM, parse failure). No SBOM/counts. | Re-scanned on the rescan interval — failures may be transient. |
| `skipped` | Scanner intentionally bypassed the artifact (helm chart, in-toto, DSSE). `reason` explains why. | **Never re-queued.** Won't change without a code change in the scanner. |
Records written before the `status` field existed have an empty status. The appview
treats empty + nil-blob + zero-count as failed (legacy fallback).
### Unscannable artifact types
The scanner skips artifacts whose config media type appears in
`unscannableConfigTypes` (`scanner/internal/scan/worker.go`). Currently:
- `application/vnd.cncf.helm.config.v1+json` — Helm charts. Rendered with a
helm-aware digest page (`pkg/appview/handlers/digest.go`) that shows Chart.yaml
metadata instead of layers / vulns / SBOM.
- `application/vnd.in-toto+json` — in-toto attestations.
- `application/vnd.dsse.envelope.v1+json` — DSSE envelopes (SLSA provenance).
For these types the appview's vuln/SBOM tabs render *"Vulnerability scanning isn't
applied to this artifact type."* — no retry hint.
To add a new unscannable type: append the media type to `unscannableConfigTypes`.
Existing records won't auto-rewrite — run the backfill tool (below) once to convert
any pre-existing failure records into skipped records.
## Scheduling and Priority
### Scanner-side priority queue
Each scanner keeps a local priority heap (`scanner/internal/queue/priority_queue.go`).
Jobs are ordered by tier priority, FIFO within a tier (lower number = higher priority):
| Tier | Priority |
|-----------------|----------|
| `owner` | 0 |
| `quartermaster` | 1 |
| `bosun` | 2 |
| anything else (`deckhand`) | 3 |
So when a scanner has a backlog, owner and paid-tier jobs are processed before
free-tier ones.
### Hold-side dispatch
The hold dispatches jobs **round-robin** across connected scanners (no priority at the
hold level — that is the scanner's job). Each scanner pulls its assigned jobs into its
own priority queue. With multiple scanners, the competing-consumer pattern spreads load.
### Proactive scanning
When `scanner.rescan_interval > 0`, the hold runs three background loops:
- **Discovery loop**: every 4 hours (and on scanner reconnect), queries relays for
DIDs with `io.atcr.manifest` records, walks each user's PDS, and queues manifests
that belong to this hold but have no scan record yet. These are dispatched at the
`deckhand` tier.
- **Stale-scan loop**: walks the local scan records and re-queues any `ok`/`failed`
record older than `rescan_interval`. Skipped records are left alone.
- **Dispatch loop**: drains the unscanned queue (higher priority) before the stale
queue, throttled to one proactive job at a time so push-triggered scans aren't
starved.
## Accessing Results
There is **no** `io.atcr.hold.getSBOM` XRPC endpoint. Results are read directly from
the hold's PDS using standard ATProto XRPC, and the appview UI wraps these calls.
### From the AppView web UI
The appview exposes HTMX endpoints that render scan data on repository/digest pages
(`pkg/appview/routes/routes.go`, handlers in `pkg/appview/handlers/`):
- `GET /api/scan-result` — vulnerability badge for a digest (`scan_result.go`).
- `GET /api/scan-results` — batch badges for a tag list (`scan_result.go`).
- `GET /api/vuln-details` — full vulnerability detail modal (`vuln_details.go`).
- `GET /api/sbom-details` — SBOM summary modal (`sbom_details.go`).
- `GET /api/scan-download?digest=...&holdEndpoint=...&type=sbom|vuln` — downloads the
raw SBOM or Grype JSON as a file (`scan_download.go`).
These handlers resolve the hold, fetch the `io.atcr.hold.scan` record, and pull the
SBOM/vuln blobs.
### Directly from the hold's PDS
The appview handlers do exactly this under the hood:
```bash
# 1. Fetch the scan record (rkey = manifest digest hex, no "sha256:" prefix)
curl "https://hold01.atcr.io/xrpc/com.atproto.repo.getRecord?\
repo=did:web:hold01.atcr.io&\
collection=io.atcr.hold.scan&\
rkey=abc123..."
# Response value contains sbomBlob.ref.$link, vulnReportBlob.ref.$link, and counts.
# 2. Download the SBOM blob by its CID
curl "https://hold01.atcr.io/xrpc/com.atproto.sync.getBlob?\
did=did:web:hold01.atcr.io&\
cid=bafkrei..." > sbom.spdx.json
# 3. Scan locally with another tool if desired
grype sbom:./sbom.spdx.json
osv-scanner --sbom sbom.spdx.json
```
You can also list all scan records on a hold via
`com.atproto.repo.listRecords?repo=<holdDid>&collection=io.atcr.hold.scan`.
## Backfill and Rescan
### Rescans
Re-scanning is automatic when `scanner.rescan_interval > 0` — the stale-scan loop
re-queues records older than the interval (default 7 days). Failed scans are retried;
skipped scans are not.
### Backfill tool
`atcr-hold scan-backfill --config <path>` walks every `io.atcr.hold.scan` record and
rewrites legacy ones (empty status + nil SBOM blob + zero counts) by assigning a status
from the manifest's layer media types:
- Layer media type contains `helm.chart.content`, `in-toto`, or `dsse.envelope`
`status="skipped"`.
- Otherwise → `status="failed"`.
The tool is idempotent and preserves each record's original `scannedAt`. It opens the
hold's CAR store directly, so the hold service must be **stopped** first (the embedded
PDS holds an exclusive lock). For zero-downtime backfill on a running hold, use the
admin endpoint `POST /admin/api/scan-backfill` instead.
## Troubleshooting
- **No scans happening at all.** Check that `scanner.secret` is set on the hold (empty
disables scanning) and that a scanner is connected. Scanner connection failures log
`dial failed` / `WebSocket read error`.
- **Scanner connects then immediately disconnects.** Usually a secret mismatch —
`SCANNER_HOLD_SECRET` must equal the hold's `scanner.secret`.
- **Free-tier pushes never get scanned on push.** Expected: `deckhand` has
`scan_on_push: false` by default. They are picked up by the discovery loop instead
(requires `rescan_interval > 0`).
- **Large images skipped.** Total compressed size exceeds `vuln.max_image_size`
(2 GiB default). Raise it or set `0` for no limit.
- **Layer extraction or Grype DB download fails mid-process.** `vuln.tmp_dir` is too
small or on tmpfs. Point it at a large persistent partition; the scanner sets
`TMPDIR` to this directory.
- **SBOM present but no vulnerability counts.** `vuln.enabled` is false on the scanner,
or the Grype DB failed to initialize (check startup logs).
- **Helm/attestation artifacts show "scanning isn't applied".** Expected — these are
in `unscannableConfigTypes` and recorded as `skipped`.
## References
- [Syft](https://github.com/anchore/syft)
- [Grype](https://github.com/anchore/grype)
- [SPDX Specification](https://spdx.dev/)
- [Hold XRPC Endpoints](./HOLD_XRPC_ENDPOINTS.md)
- [Quotas](./QUOTAS.md)
- [ATProto Specification](https://atproto.com/)
</content>
</invoke>