# SBOM Scanning and Vulnerability Analysis ATCR generates Software Bills of Materials (SBOMs) and scans container images for vulnerabilities. Scanning runs in a separate `atcr-scanner` service that connects to a hold over a WebSocket, so the hold itself never runs Syft or Grype. Results are stored as `io.atcr.hold.scan` records in the hold's embedded PDS. ## Overview - **Separate scanner binary**: Scanning is performed by `atcr-scanner` (the `scanner/` Go module), not by the hold. The scanner connects out to the hold and pulls jobs. - **Syft for SBOMs, Grype for vulnerabilities**: Each job runs Syft to produce an SPDX-JSON SBOM, then Grype to scan that SBOM for CVEs. Grype is enabled by default. - **WebSocket dispatch**: The hold pushes jobs to connected scanners over `/xrpc/io.atcr.hold.subscribeScanJobs`. A shared secret authenticates the scanner. - **ATProto result storage**: Results land as `io.atcr.hold.scan` records in the hold's embedded PDS, with the SBOM and full Grype report uploaded as PDS blobs. - **Tier-gated scan-on-push plus proactive rescans**: Pushes from eligible tiers trigger an immediate scan; the hold also discovers never-scanned manifests and re-scans stale ones on an interval. ### Tools - [Anchore Syft](https://github.com/anchore/syft) generates the SBOM. Output format is SPDX JSON, hardcoded in `scanner/internal/scan/syft.go` (not configurable). - [Anchore Grype](https://github.com/anchore/grype) scans the SBOM for known vulnerabilities and produces critical/high/medium/low/total counts plus a full JSON report with CVE detail. ## Architecture Three pieces cooperate: ``` io.atcr.hold.subscribeScanJobs (WebSocket, ?secret=...) ┌───────────┐ ◄──────────────────────────────────────────── ┌──────────────┐ │ Hold │ job: {seq, manifestDigest, repo, tier, │ atcr-scanner │ │ (Scan │ config, layers, holdEndpoint, ...} │ (Syft + │ │ Broadcaster)│ ────────────────────────────────────────────► │ Grype) │ │ │ │ │ │ │ result/error/skipped: {seq, sbom, │ │ │ │ ◄──── vulnReport, summary{critical,high,...}} └──────────────┘ └─────┬─────┘ │ stores io.atcr.hold.scan record + SBOM/vuln blobs ▼ Hold embedded PDS (CAR store) ``` 1. **Hold (`pkg/hold/pds/scan_broadcaster.go`)** owns the `ScanBroadcaster`. It persists pending jobs in SQLite (`scan_jobs` table), accepts scanner WebSocket connections, and dispatches jobs **round-robin** across all connected scanners using a competing-consumer pattern. It re-dispatches timed-out jobs, and (when a rescan interval is set) runs background discovery and stale-scan loops. On receiving a result, the hold uploads the SBOM and vuln report as blobs and writes the `io.atcr.hold.scan` record. 2. **Scanner (`scanner/` module)** dials the hold's WebSocket, acks jobs, runs the Syft → Grype pipeline, and sends back a result, error, or skipped message. It keeps a local **priority queue** so paid tiers jump ahead of free ones (see Scheduling). 3. **AppView** reads the scan records and blobs from the hold's PDS to render vulnerability badges, SBOM details, and download links in the web UI. ### Why the hold's PDS? Scan results are stored in the **hold's embedded PDS** rather than the user's PDS: - No OAuth/service-token plumbing: the hold owns and signs its own records. - Hold-scoped metadata (scanner version, scan time) stays with the operator. - Different holds can independently scan the same image for cross-verification. - The user's PDS stays lean: SBOM and Grype JSON live in hold blob storage. The trust model is the same as Docker Hub: you trust the hold operator's scanner version and scan integrity. The hold's DID signs the records, and anyone can re-scan a digest to verify the result. ## Configuration ### Hold side The hold's scanner integration is configured under `scanner:` in the hold config (`pkg/hold/config.go`). Env-var prefix is `HOLD_`. | YAML key | Env var | Default | Meaning | |--------------------------|-------------------------------|---------|---------| | `scanner.secret` | `HOLD_SCANNER_SECRET` | `""` | Shared secret a scanner must present (as `?secret=`) on the WebSocket. **Empty disables scanning entirely** — no scanner can connect and no jobs are dispatched. | | `scanner.rescan_interval`| `HOLD_SCANNER_RESCAN_INTERVAL`| `168h` | Minimum interval between re-scans of the same manifest. When > 0 the hold runs proactive discovery + stale-scan loops. Set to `0` to disable proactive scanning (push-triggered scans still work). | ```yaml # config-hold.yaml scanner: secret: "a-long-random-shared-secret" rescan_interval: 168h ``` Whether a push triggers an immediate scan is decided by the quota tier (see [Scan-on-push tier gate](#scan-on-push-tier-gate)). ### Scanner side The scanner is configured via Viper (`scanner/internal/config/config.go`); it accepts a YAML file or pure env vars with the `SCANNER_` prefix. Run with `SCANNER_HOLD_URL=... SCANNER_HOLD_SECRET=... atcr-scanner serve`. | YAML key | Env var | Default | Meaning | |---------------------|------------------------------|----------------------------------|---------| | `hold.url` | `SCANNER_HOLD_URL` | — (**required**) | WebSocket URL of the hold, e.g. `ws://localhost:8080` or `wss://hold01.atcr.io`. `http(s)` is auto-converted to `ws(s)`. | | `hold.secret` | `SCANNER_HOLD_SECRET` | — (**required**) | Must match the hold's `scanner.secret`. Sent as `?secret=`. | | `scanner.workers` | `SCANNER_SCANNER_WORKERS` | `1` | Number of concurrent scan workers. | | `scanner.queue_size`| `SCANNER_SCANNER_QUEUE_SIZE` | `100` | Max depth of the local priority queue. | | `vuln.enabled` | `SCANNER_VULN_ENABLED` | `true` | Run Grype after Syft. When false, only the SBOM is produced (no counts). | | `vuln.db_path` | `SCANNER_VULN_DB_PATH` | `/var/lib/atcr-scanner/vulndb` | Directory for the Grype vulnerability database. | | `vuln.tmp_dir` | `SCANNER_VULN_TMP_DIR` | `/var/lib/atcr-scanner/tmp` | Directory for layer extraction and DB download. Also exported as `TMPDIR`; point it at a large partition, **not** tmpfs. | | `vuln.max_image_size`| `SCANNER_VULN_MAX_IMAGE_SIZE`| `2147483648` (2 GiB) | Max total compressed image size. Larger images are skipped with an error. `0` = no limit. | | `server.addr` | `SCANNER_SERVER_ADDR` | `:9090` | Listen address for the scanner's health endpoint. | Both `hold.url` and `hold.secret` are required; `LoadConfig` errors out if either is empty. ```bash # Minimal scanner invocation (env-only) SCANNER_HOLD_URL=wss://hold01.atcr.io \ SCANNER_HOLD_SECRET=a-long-random-shared-secret \ ./bin/atcr-scanner serve ``` ## Scanning Workflow ### 1. Push → scan-on-push tier gate When an image is pushed and the manifest is recorded, the hold's OCI XRPC handler (`pkg/hold/oci/xrpc.go`) decides whether to enqueue a scan. Multi-arch manifest lists and artifacts with a `subject` (attestations) are skipped — they have no scannable content. For everything else, the tier of the pusher decides: - **Captain / owner**: always scanned. - **Quotas disabled** (`quotaMgr == nil` or quotas not enabled): all pushes scanned (backwards compatible). - **Quotas enabled**: scanned only if the pusher's tier has `scan_on_push: true`. In the default config (`pkg/hold/config.go`), `bosun` and `quartermaster` have `scan_on_push: true`; `deckhand` does not. So a free-tier (deckhand) push is **not** scanned on push — it gets picked up later by the proactive discovery loop. ### 2. Dispatch The `ScanBroadcaster.Enqueue` inserts the job into the `scan_jobs` SQLite table (status `pending`) and immediately tries to dispatch it round-robin to one of the connected scanners. Jobs survive hold restarts. If no scanner is connected, the job waits; newly connected scanners drain pending jobs. Assigned-but-unacked jobs time out after 5 minutes and are re-dispatched; jobs stuck in `processing` for 10 minutes are marked failed (scanner likely crashed). ### 3. Scan pipeline (scanner) For each job (`scanner/internal/scan/worker.go`): 1. **Artifact-type check** — if `config.mediaType` is in `unscannableConfigTypes` the job returns a `SkipError` and the scanner sends a `skipped` message (see below). 2. **Size check** — if total compressed size exceeds `vuln.max_image_size`, the job fails. 3. **Build OCI layout** — layers are fetched from the hold via presigned URLs and assembled into an OCI image layout in `vuln.tmp_dir`. 4. **Syft** — generates the SBOM and encodes it to SPDX JSON. 5. **Grype** (if `vuln.enabled`) — scans the SBOM, producing the full JSON report and a severity summary (critical/high/medium/low/total). The scanner then sends one of three messages back over the WebSocket: `result` (SBOM + optional vuln report + summary), `error`, or `skipped` (with a reason). ### 4. Result storage (hold) On `result` (`scan_broadcaster.go` `handleResult`): 1. Upload the SBOM bytes as a PDS blob (`application/spdx+json`). 2. Upload the Grype report as a PDS blob (`application/vnd.atcr.vulnerabilities+json`). 3. Create an `io.atcr.hold.scan` record (`CreateScanRecord`) keyed by the manifest digest, referencing both blobs and carrying the severity counts. 4. Mark the `scan_jobs` row `completed`. On `error`, a failed scan record is written (`NewFailedScanRecord`) and the job is marked `failed`. On `skipped`, a skipped record is written (`NewSkippedScanRecord`) and the job is marked `completed`. ## Scan Record Schema Results are `io.atcr.hold.scan` records in the hold's embedded PDS (`pkg/atproto/lexicon.go`, `ScanRecord`). The record key is the manifest digest hex (without the `sha256:` prefix), so there is exactly one scan record per manifest and re-scans upsert it. ```json { "$type": "io.atcr.hold.scan", "manifest": "at://did:plc:alice123/io.atcr.manifest/abc123...", "repository": "myapp", "userDid": "did:plc:alice123", "sbomBlob": { "$type": "blob", "ref": { "$link": "bafkrei..." }, "mimeType": "application/spdx+json", "size": 51234 }, "vulnReportBlob": { "$type": "blob", "ref": { "$link": "bafkrei..." }, "mimeType": "application/vnd.atcr.vulnerabilities+json", "size": 18567 }, "critical": 2, "high": 15, "medium": 42, "low": 8, "total": 67, "scannerVersion": "atcr-scanner-v1.0.0", "scannedAt": "2026-06-11T12:34:56Z", "status": "ok", "reason": "" } ``` | Field | Notes | |------------------|-------| | `manifest` | AT-URI of the scanned manifest in the user's PDS. | | `userDid` | DID of the image owner. | | `sbomBlob` | Reference to the SPDX-JSON SBOM in hold blob storage. Absent for failed/skipped scans. | | `vulnReportBlob` | Reference to the full Grype JSON report. Absent if Grype disabled or scan failed/skipped. | | `critical`/`high`/`medium`/`low`/`total` | Vulnerability counts from Grype. Zero on failed/skipped scans. | | `scannerVersion` | Scanner identifier for reproducibility (currently `atcr-scanner-v1.0.0`). | | `scannedAt` | RFC3339 scan completion timestamp. | | `status` | `ok`, `failed`, or `skipped`. | | `reason` | Populated for `failed` (error text) and `skipped` (why it was bypassed). | ### Status field | Status | Meaning | Rescan behavior | |-------------|-------------------------------------------------------------------------|-----------------| | `ok` (or empty) | Scanner produced an SBOM; counts and SBOM blob populated. | Re-scanned on the rescan interval (default 7d). | | `failed` | Scanner ran but errored (network, OOM, parse failure). No SBOM/counts. | Re-scanned on the rescan interval — failures may be transient. | | `skipped` | Scanner intentionally bypassed the artifact (helm chart, in-toto, DSSE). `reason` explains why. | **Never re-queued.** Won't change without a code change in the scanner. | Records written before the `status` field existed have an empty status. The appview treats empty + nil-blob + zero-count as failed (legacy fallback). ### Unscannable artifact types The scanner skips artifacts whose config media type appears in `unscannableConfigTypes` (`scanner/internal/scan/worker.go`). Currently: - `application/vnd.cncf.helm.config.v1+json` — Helm charts. Rendered with a helm-aware digest page (`pkg/appview/handlers/digest.go`) that shows Chart.yaml metadata instead of layers / vulns / SBOM. - `application/vnd.in-toto+json` — in-toto attestations. - `application/vnd.dsse.envelope.v1+json` — DSSE envelopes (SLSA provenance). For these types the appview's vuln/SBOM tabs render *"Vulnerability scanning isn't applied to this artifact type."* — no retry hint. To add a new unscannable type: append the media type to `unscannableConfigTypes`. Existing records won't auto-rewrite — run the backfill tool (below) once to convert any pre-existing failure records into skipped records. ## Scheduling and Priority ### Scanner-side priority queue Each scanner keeps a local priority heap (`scanner/internal/queue/priority_queue.go`). Jobs are ordered by tier priority, FIFO within a tier (lower number = higher priority): | Tier | Priority | |-----------------|----------| | `owner` | 0 | | `quartermaster` | 1 | | `bosun` | 2 | | anything else (`deckhand`) | 3 | So when a scanner has a backlog, owner and paid-tier jobs are processed before free-tier ones. ### Hold-side dispatch The hold dispatches jobs **round-robin** across connected scanners (no priority at the hold level — that is the scanner's job). Each scanner pulls its assigned jobs into its own priority queue. With multiple scanners, the competing-consumer pattern spreads load. ### Proactive scanning When `scanner.rescan_interval > 0`, the hold runs three background loops: - **Discovery loop**: every 4 hours (and on scanner reconnect), queries relays for DIDs with `io.atcr.manifest` records, walks each user's PDS, and queues manifests that belong to this hold but have no scan record yet. These are dispatched at the `deckhand` tier. - **Stale-scan loop**: walks the local scan records and re-queues any `ok`/`failed` record older than `rescan_interval`. Skipped records are left alone. - **Dispatch loop**: drains the unscanned queue (higher priority) before the stale queue, throttled to one proactive job at a time so push-triggered scans aren't starved. ## Accessing Results There is **no** `io.atcr.hold.getSBOM` XRPC endpoint. Results are read directly from the hold's PDS using standard ATProto XRPC, and the appview UI wraps these calls. ### From the AppView web UI The appview exposes HTMX endpoints that render scan data on repository/digest pages (`pkg/appview/routes/routes.go`, handlers in `pkg/appview/handlers/`): - `GET /api/scan-result` — vulnerability badge for a digest (`scan_result.go`). - `GET /api/scan-results` — batch badges for a tag list (`scan_result.go`). - `GET /api/vuln-details` — full vulnerability detail modal (`vuln_details.go`). - `GET /api/sbom-details` — SBOM summary modal (`sbom_details.go`). - `GET /api/scan-download?digest=...&holdEndpoint=...&type=sbom|vuln` — downloads the raw SBOM or Grype JSON as a file (`scan_download.go`). These handlers resolve the hold, fetch the `io.atcr.hold.scan` record, and pull the SBOM/vuln blobs. ### Directly from the hold's PDS The appview handlers do exactly this under the hood: ```bash # 1. Fetch the scan record (rkey = manifest digest hex, no "sha256:" prefix) curl "https://hold01.atcr.io/xrpc/com.atproto.repo.getRecord?\ repo=did:web:hold01.atcr.io&\ collection=io.atcr.hold.scan&\ rkey=abc123..." # Response value contains sbomBlob.ref.$link, vulnReportBlob.ref.$link, and counts. # 2. Download the SBOM blob by its CID curl "https://hold01.atcr.io/xrpc/com.atproto.sync.getBlob?\ did=did:web:hold01.atcr.io&\ cid=bafkrei..." > sbom.spdx.json # 3. Scan locally with another tool if desired grype sbom:./sbom.spdx.json osv-scanner --sbom sbom.spdx.json ``` You can also list all scan records on a hold via `com.atproto.repo.listRecords?repo=&collection=io.atcr.hold.scan`. ## Backfill and Rescan ### Rescans Re-scanning is automatic when `scanner.rescan_interval > 0` — the stale-scan loop re-queues records older than the interval (default 7 days). Failed scans are retried; skipped scans are not. ### Backfill tool `atcr-hold scan-backfill --config ` walks every `io.atcr.hold.scan` record and rewrites legacy ones (empty status + nil SBOM blob + zero counts) by assigning a status from the manifest's layer media types: - Layer media type contains `helm.chart.content`, `in-toto`, or `dsse.envelope` → `status="skipped"`. - Otherwise → `status="failed"`. The tool is idempotent and preserves each record's original `scannedAt`. It opens the hold's CAR store directly, so the hold service must be **stopped** first (the embedded PDS holds an exclusive lock). For zero-downtime backfill on a running hold, use the admin endpoint `POST /admin/api/scan-backfill` instead. ## Troubleshooting - **No scans happening at all.** Check that `scanner.secret` is set on the hold (empty disables scanning) and that a scanner is connected. Scanner connection failures log `dial failed` / `WebSocket read error`. - **Scanner connects then immediately disconnects.** Usually a secret mismatch — `SCANNER_HOLD_SECRET` must equal the hold's `scanner.secret`. - **Free-tier pushes never get scanned on push.** Expected: `deckhand` has `scan_on_push: false` by default. They are picked up by the discovery loop instead (requires `rescan_interval > 0`). - **Large images skipped.** Total compressed size exceeds `vuln.max_image_size` (2 GiB default). Raise it or set `0` for no limit. - **Layer extraction or Grype DB download fails mid-process.** `vuln.tmp_dir` is too small or on tmpfs. Point it at a large persistent partition; the scanner sets `TMPDIR` to this directory. - **SBOM present but no vulnerability counts.** `vuln.enabled` is false on the scanner, or the Grype DB failed to initialize (check startup logs). - **Helm/attestation artifacts show "scanning isn't applied".** Expected — these are in `unscannableConfigTypes` and recorded as `skipped`. ## References - [Syft](https://github.com/anchore/syft) - [Grype](https://github.com/anchore/grype) - [SPDX Specification](https://spdx.dev/) - [Hold XRPC Endpoints](./HOLD_XRPC_ENDPOINTS.md) - [Quotas](./QUOTAS.md) - [ATProto Specification](https://atproto.com/)