diff --git a/docs/HOLD_XRPC_ENDPOINTS.md b/docs/HOLD_XRPC_ENDPOINTS.md
index b9a9e9c..6cde6fb 100644
--- a/docs/HOLD_XRPC_ENDPOINTS.md
+++ b/docs/HOLD_XRPC_ENDPOINTS.md
@@ -72,7 +72,7 @@ The `requireAuth` middleware validates Bearer service tokens only. `requestCrew`
| Endpoint | Method | Description |
|----------|--------|-------------|
-| `/xrpc/io.atcr.hold.subscribeScanJobs` | GET (WebSocket) | Scanner job subscription. Auth via `?secret=` query param or `X-Scanner-Secret` header (shared secret). Supports `?cursor=` for backfill. |
+| `/xrpc/io.atcr.hold.subscribeScanJobs` | GET (WebSocket) | Scanner job subscription. Auth via `?secret=` query param or `X-Scanner-Secret` header (shared secret). Supports `?cursor=` for backfill, `?workers=` to declare how many scans the process runs at once (default 1), and `?instance=` to declare a stable process identity so a reconnecting scanner resumes its own in-flight jobs. |
---
@@ -108,7 +108,7 @@ All require `blob:write` permission via service token:
| `/xrpc/io.atcr.hold.purgeManifest` | POST | inline (service token or DPoP; captain, crew:admin, or manifest owner) | Purge layer/scan/image-config records for a single manifest URI. Called by appview on UI delete; called internally on takedown receipt. Does not delete S3 blobs (GC handles those). |
| `/xrpc/io.atcr.hold.listTiers` | GET | none | List hold's available tiers with quotas and features (scanOnPush) |
| `/xrpc/io.atcr.hold.updateCrewTier` | POST | appview token (ES256 JWT; 503 if appview DID not configured) | Update crew member's tier |
-| `/xrpc/io.atcr.hold.subscribeScanJobs` | GET (WebSocket) | shared secret (`?secret=` or `X-Scanner-Secret`) | Scanner job subscription; supports `?cursor=` for backfill |
+| `/xrpc/io.atcr.hold.subscribeScanJobs` | GET (WebSocket) | shared secret (`?secret=` or `X-Scanner-Secret`) | Scanner job subscription; supports `?cursor=` for backfill, `?workers=` for concurrency, `?instance=` for reconnect resumption |
---
diff --git a/docs/SBOM_SCANNING.md b/docs/SBOM_SCANNING.md
index 0cf06c1..3011312 100644
--- a/docs/SBOM_SCANNING.md
+++ b/docs/SBOM_SCANNING.md
@@ -104,7 +104,7 @@ a YAML file or pure env vars with the `SCANNER_` prefix. Run with
|---------------------|------------------------------|----------------------------------|---------|
| `hold.url` | `SCANNER_HOLD_URL` | — (**required**) | WebSocket URL of the hold, e.g. `ws://localhost:8080` or `wss://hold01.atcr.io`. `http(s)` is auto-converted to `ws(s)`. |
| `hold.secret` | `SCANNER_HOLD_SECRET` | — (**required**) | Must match the hold's `scanner.secret`. Sent as `?secret=`. |
-| `scanner.workers` | `SCANNER_SCANNER_WORKERS` | `1` | Number of concurrent scan workers. |
+| `scanner.workers` | `SCANNER_SCANNER_WORKERS` | `1` | Number of concurrent scan workers. Declared to the hold on connect, which sizes the hold's dispatch budget for this process; raise it only alongside `vuln.max_image_size` and a cgroup memory cap. |
| `scanner.queue_size`| `SCANNER_SCANNER_QUEUE_SIZE` | `100` | Max depth of the local priority queue. |
| `vuln.enabled` | `SCANNER_VULN_ENABLED` | `true` | Run Grype after Syft. When false, only the SBOM is produced (no counts). |
| `vuln.db_path` | `SCANNER_VULN_DB_PATH` | `/var/lib/atcr-scanner/vulndb` | Directory for the Grype vulnerability database. |
@@ -143,11 +143,34 @@ scanned on push — it gets picked up later by the proactive discovery loop.
### 2. Dispatch
The `ScanBroadcaster.Enqueue` inserts the job into the `scan_jobs` SQLite table
-(status `pending`) and immediately tries to dispatch it round-robin to one of the
-connected scanners. Jobs survive hold restarts. If no scanner is connected, the job
-waits; newly connected scanners drain pending jobs. Assigned-but-unacked jobs time out
-after 5 minutes and are re-dispatched; jobs stuck in `processing` for 10 minutes are
-marked failed (scanner likely crashed).
+(status `pending`) and immediately tries to dispatch it to a connected scanner. Jobs
+survive hold restarts. If no scanner is connected, the job waits.
+
+**Which scanner gets it.** Selection is by spare capacity, not position: each
+connection declares how many scans it runs at once (`?workers=`, default 1) and the
+hold prefers the scanner with the smallest fraction of its capacity committed, with
+ties resolved round-robin. A job that no connected scanner has room for stays
+`pending` rather than being pushed into a scanner's own queue, and is offered again
+the moment any scanner finishes something. Keeping the queue on the hold is what
+makes the job re-routable to whichever process frees up first, and what makes the
+deadlines below mean anything.
+
+**Deadlines.** Assigned-but-unacked jobs time out after 5 minutes and are
+re-dispatched. Once a job is acked the scanner has it, but it may be queued behind
+that scanner's workers, so there are two further budgets: 10 minutes from the
+`started` message a worker sends when it actually begins the scan, and 60 minutes
+from dispatch for a job that was acked but never reported as started (which is also
+what a scanner too old to send `started` gets). Both write a failed scan record and
+release the manifest for re-scanning.
+
+**Disconnects.** A dropped WebSocket does not return a scanner's in-flight jobs to
+the pool: its worker pool never learns the socket went away and keeps scanning, so
+handing that work to another process would have two scanners scanning the same image.
+The rows are marked instead. A scanner sends a stable per-process identity
+(`?instance=`) on every connect and resumes its own jobs on reconnect; a scanner that
+does not come back within 2 minutes has them reclaimed and re-offered. A scanner that
+actually restarted comes back with a new identity, so its old work is reclaimed
+rather than resumed — which is right, since a restart really did lose it.
### 3. Scan pipeline (scanner)
@@ -163,8 +186,16 @@ For each job (`scanner/internal/scan/worker.go`):
5. **Grype** (if `vuln.enabled`) — scans the SBOM, producing the full JSON report and
a severity summary (critical/high/medium/low/total).
-The scanner then sends one of three messages back over the WebSocket: `result`
-(SBOM + optional vuln report + summary), `error`, or `skipped` (with a reason).
+A worker sends `started` when it dequeues a job, before step 1. This is distinct
+from the `ack`, which the WebSocket reader sends the instant a job frame arrives:
+the gap between them is however long the job waits in this scanner's own queue, and
+the hold measures its scanning deadline from `started` so that queueing does not
+count against it.
+
+The scanner then sends one of three terminal messages back over the WebSocket:
+`result` (SBOM + optional vuln report + summary), `error`, or `skipped` (with a
+reason). All four messages are ignored by the hold unless the job is currently
+assigned to the scanner sending them.
### 4. Result storage (hold)
@@ -292,8 +323,18 @@ When `scanner.rescan_interval > 0`, the hold runs three background loops:
- **Stale-scan loop**: walks the local scan records and re-queues any `ok`/`failed`
record older than `rescan_interval`. Skipped records are left alone.
- **Dispatch loop**: drains the unscanned queue (higher priority) before the stale
- queue, throttled to one proactive job at a time so push-triggered scans aren't
- starved.
+ queue, throttled to one proactive job per connected scanner worker — the sum of
+ every connected scanner's declared `workers`. The throttle counts only proactive
+ jobs: push-triggered scans bypass it entirely, so counting them meant a hold with
+ steady pushes never dispatched a proactive scan at all. With no scanner connected
+ the budget is zero and nothing is dispatched.
+
+Scaling this out is therefore a matter of running more scanner processes against the
+same hold, raising `scanner.workers`, or both: the dispatch budget, the drain on
+connect and the choice of scanner all follow the declared capacity. Do this only
+after bounding scanner memory — concurrency is what holds peak RSS down on a small
+host, and two concurrent scans of a `node:22`-class image measured 687 MiB with a
+512 MiB `GOMEMLIMIT` in force and 1357 MiB without.
## Accessing Results
diff --git a/lexicons/io/atcr/hold/subscribeScanJobs.json b/lexicons/io/atcr/hold/subscribeScanJobs.json
index fbdff3a..8bc0534 100644
--- a/lexicons/io/atcr/hold/subscribeScanJobs.json
+++ b/lexicons/io/atcr/hold/subscribeScanJobs.json
@@ -11,6 +11,17 @@
"cursor": {
"type": "integer",
"description": "Sequence number to resume from. If omitted, starts from latest. Use -1 to receive only new jobs."
+ },
+ "workers": {
+ "type": "integer",
+ "minimum": 1,
+ "maximum": 32,
+ "description": "How many scans this scanner runs concurrently. The hold keeps this many jobs in flight for the connection. Omitted or unusable means one."
+ },
+ "instance": {
+ "type": "string",
+ "maxLength": 64,
+ "description": "Stable identity of the scanner process, sent on every connect. A scanner that reconnects with the same value resumes the jobs it was holding when the connection dropped, instead of having them offered to another scanner. Omitted means the hold assigns a per-connection identity and the scanner's in-flight work is reclaimed rather than resumed."
}
}
},
diff --git a/pkg/appview/handlers/image_advisor.go b/pkg/appview/handlers/image_advisor.go
index 02e4b0a..b3c6e73 100644
--- a/pkg/appview/handlers/image_advisor.go
+++ b/pkg/appview/handlers/image_advisor.go
@@ -469,8 +469,15 @@ func generateAdvisorPrompt(w io.Writer, r *advisorReportData) {
// Vulnerability summary
if r.ScanRecord != nil {
sr := r.ScanRecord
- fmt.Fprintf(w, "vulns: {critical: %d, high: %d, medium: %d, low: %d, total: %d}\n",
- sr.Critical, sr.High, sr.Medium, sr.Low, sr.Total)
+ if vulnScanDidNotRun(sr) {
+ // Counts of zero here would be read as "no vulnerabilities", but
+ // this record was written by a scan that never ran a vulnerability
+ // database against the image.
+ fmt.Fprintf(w, "vulns: not scanned (SBOM only, no vulnerability data)\n")
+ } else {
+ fmt.Fprintf(w, "vulns: {critical: %d, high: %d, medium: %d, low: %d, total: %d}\n",
+ sr.Critical, sr.High, sr.Medium, sr.Low, sr.Total)
+ }
}
// Fixable critical/high vulns
diff --git a/pkg/appview/handlers/scan_result.go b/pkg/appview/handlers/scan_result.go
index c0d19d2..c5db8d0 100644
--- a/pkg/appview/handlers/scan_result.go
+++ b/pkg/appview/handlers/scan_result.go
@@ -25,50 +25,74 @@ type ScanResultHandler struct {
}
// vulnBadgeData is the template data for the vuln-badge partial.
-// The badge renders one of five states, in priority order:
-// 1. Error — we couldn't reach the hold at all (network/5xx)
-// 2. NotScanned — hold reachable, no scan record for this digest (404)
-// 3. Skipped — scan record explicitly marks this artifact as not-scannable
-// 4. ScanFailed — scan record exists but the scanner errored
-// 5. Found — scan succeeded; render tier counts (or "Clean" when zero)
+// The badge renders one of six states, in priority order:
+// 1. Error — we couldn't reach the hold at all (network/5xx)
+// 2. NotScanned — hold reachable, no scan record for this digest (404)
+// 3. Skipped — scan record explicitly marks this artifact as not-scannable
+// 4. ScanFailed — scan record exists but the scanner errored
+// 5. VulnsNotScanned — scan succeeded but no vulnerability data was produced
+// 6. Found — scan succeeded; render tier counts (or "Clean" when zero)
//
// These states must stay distinct so users can tell "hold is down" from
// "this hasn't been scanned yet" from "scanner errored on this image" from
-// "this artifact type is intentionally not scanned".
+// "this artifact type is intentionally not scanned" from "we catalogued the
+// image but never matched it against a vulnerability database".
type vulnBadgeData struct {
- Critical int64
- High int64
- Medium int64
- Low int64
- Total int64
- ScannedAt string
- Found bool // true if scan record exists and succeeded
- Error bool // true if hold unreachable (network/5xx)
- NotScanned bool // true if hold is up but no scan record (404)
- ScanFailed bool // true if scan record exists but scan failed
- Skipped bool // true if scan record marks the artifact as intentionally not scanned (helm, in-toto, etc.)
- Digest string // for the detail modal link
- HoldEndpoint string // for the detail modal link
+ Critical int64
+ High int64
+ Medium int64
+ Low int64
+ Total int64
+ ScannedAt string
+ Found bool // true if scan record exists and succeeded
+ Error bool // true if hold unreachable (network/5xx)
+ NotScanned bool // true if hold is up but no scan record (404)
+ ScanFailed bool // true if scan record exists but scan failed
+ Skipped bool // true if scan record marks the artifact as intentionally not scanned (helm, in-toto, etc.)
+ // VulnsNotScanned means the scan produced an SBOM but no vulnerability
+ // data, which is what a scanner running with vulnerability scanning off
+ // reports. Its zero counts are the absence of a measurement, not a finding,
+ // so the badge must not render them as "Clean".
+ VulnsNotScanned bool
+ Digest string // for the detail modal link
+ HoldEndpoint string // for the detail modal link
+}
+
+// vulnScanDidNotRun reports whether a successful scan record carries no
+// vulnerability data at all: no report blob and no counts. That is what the
+// hold writes when the scanner ran with vulnerability scanning disabled, and
+// its zeros must never be presented as "no vulnerabilities found".
+//
+// It deliberately requires an explicit "ok" status. Records written before the
+// vulnReportBlob field existed carry counts and an SBOM but no status, and
+// those really were scanned, so the legacy shape stays out of this branch.
+func vulnScanDidNotRun(scanRecord *atproto.ScanRecord) bool {
+ return scanRecord.Status == atproto.ScanStatusOK &&
+ scanRecord.VulnReportBlob == nil &&
+ scanRecord.Total == 0
}
// classifyScanRecord maps a scan record's Status field to badge data flags.
// An empty Status is treated as a legacy record from before the status field
// existed: nil-blob + zero-counts = treat as failed (preserves the prior badge
// for un-backfilled holds); otherwise treat as success.
-func classifyScanRecord(scanRecord *atproto.ScanRecord) (found, skipped, failed bool) {
+func classifyScanRecord(scanRecord *atproto.ScanRecord) (found, skipped, failed, vulnsNotScanned bool) {
switch scanRecord.Status {
case atproto.ScanStatusSkipped:
- return false, true, false
+ return false, true, false, false
case atproto.ScanStatusFailed:
- return false, false, true
+ return false, false, true, false
case atproto.ScanStatusOK:
- return true, false, false
+ if vulnScanDidNotRun(scanRecord) {
+ return false, false, false, true
+ }
+ return true, false, false, false
default:
// Legacy record (status field didn't exist when this was written).
if scanRecord.SbomBlob == nil && scanRecord.Total == 0 {
- return false, false, true
+ return false, false, true, false
}
- return true, false, false
+ return true, false, false, false
}
}
@@ -146,19 +170,21 @@ func (h *ScanResultHandler) ServeHTTP(w http.ResponseWriter, r *http.Request) {
return
}
- found, skipped, failed := classifyScanRecord(&scanRecord)
+ found, skipped, failed, vulnsNotScanned := classifyScanRecord(&scanRecord)
h.renderBadge(w, vulnBadgeData{
- Critical: scanRecord.Critical,
- High: scanRecord.High,
- Medium: scanRecord.Medium,
- Low: scanRecord.Low,
- Total: scanRecord.Total,
- ScannedAt: scanRecord.ScannedAt,
- Found: found,
- Skipped: skipped,
- ScanFailed: failed,
- Digest: digest,
- HoldEndpoint: holdDID,
+ Critical: scanRecord.Critical,
+ High: scanRecord.High,
+ Medium: scanRecord.Medium,
+ Low: scanRecord.Low,
+ Total: scanRecord.Total,
+ ScannedAt: scanRecord.ScannedAt,
+ Found: found,
+ Skipped: skipped,
+ ScanFailed: failed,
+
+ VulnsNotScanned: vulnsNotScanned,
+ Digest: digest,
+ HoldEndpoint: holdDID,
})
}
@@ -211,19 +237,21 @@ func fetchScanRecord(ctx context.Context, holdEndpoint, holdDID, hexDigest strin
return vulnBadgeData{Error: true}
}
- found, skipped, failed := classifyScanRecord(&scanRecord)
+ found, skipped, failed, vulnsNotScanned := classifyScanRecord(&scanRecord)
return vulnBadgeData{
- Critical: scanRecord.Critical,
- High: scanRecord.High,
- Medium: scanRecord.Medium,
- Low: scanRecord.Low,
- Total: scanRecord.Total,
- ScannedAt: scanRecord.ScannedAt,
- Found: found,
- Skipped: skipped,
- ScanFailed: failed,
- Digest: fullDigest,
- HoldEndpoint: holdDID,
+ Critical: scanRecord.Critical,
+ High: scanRecord.High,
+ Medium: scanRecord.Medium,
+ Low: scanRecord.Low,
+ Total: scanRecord.Total,
+ ScannedAt: scanRecord.ScannedAt,
+ Found: found,
+ Skipped: skipped,
+ ScanFailed: failed,
+
+ VulnsNotScanned: vulnsNotScanned,
+ Digest: fullDigest,
+ HoldEndpoint: holdDID,
}
}
diff --git a/pkg/appview/handlers/scan_result_test.go b/pkg/appview/handlers/scan_result_test.go
index f728d0c..73c188a 100644
--- a/pkg/appview/handlers/scan_result_test.go
+++ b/pkg/appview/handlers/scan_result_test.go
@@ -503,3 +503,95 @@ func TestBatchScanResult_SingleDigest(t *testing.T) {
t.Error("Expected critical count of 1")
}
}
+
+// mockSBOMOnlyScanRecord is what the hold writes when the scanner finished a
+// scan with vulnerability scanning turned off: status "ok", an SBOM blob, no
+// vulnerability report, and zero counts because Grype never ran.
+//
+// The zero counts are not a finding. Rendering them as "Clean" would tell every
+// user of a vuln-disabled scanner that their images have no vulnerabilities,
+// which is the misreading this record shape has to avoid.
+func mockSBOMOnlyScanRecord() string {
+ record := map[string]any{
+ "$type": "io.atcr.hold.scan",
+ "manifest": "at://did:plc:test/io.atcr.manifest/abc123",
+ "repository": "myapp",
+ "userDid": "did:plc:test",
+ "status": "ok",
+ "critical": 0,
+ "high": 0,
+ "medium": 0,
+ "low": 0,
+ "total": 0,
+ "sbomBlob": map[string]any{
+ "$type": "blob",
+ "ref": map[string]any{"$link": "bafkreigv3xw47pk7cbeahkmttetf4smxyluwlu3jmteo2nzke2oa7dbhhm"},
+ "mimeType": "application/spdx+json",
+ "size": 1234,
+ },
+ "scannerVersion": "atcr-scanner-v1.0.0",
+ "scannedAt": "2025-01-15T10:30:00Z",
+ }
+ envelope := map[string]any{
+ "uri": "at://did:web:hold.example.com/io.atcr.hold.scan/abc123",
+ "cid": "bafyreiabc123",
+ "value": record,
+ }
+ b, _ := json.Marshal(envelope)
+ return string(b)
+}
+
+func TestScanResult_SBOMWithoutVulnScanIsNotClean(t *testing.T) {
+ hold := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
+ if handleMockDID(w, r) {
+ return
+ }
+ w.Header().Set("Content-Type", "application/json")
+ w.Write([]byte(mockSBOMOnlyScanRecord()))
+ }))
+ defer hold.Close()
+
+ handler := setupScanResultHandler(t, hold.URL)
+
+ req := httptest.NewRequest("GET", "/api/scan-result?digest=sha256:abc123&holdEndpoint="+hold.URL, nil)
+ rr := httptest.NewRecorder()
+ handler.ServeHTTP(rr, req)
+
+ body := rr.Body.String()
+
+ if strings.Contains(body, "Clean") || strings.Contains(body, "badge-success") {
+ t.Errorf("a record with no vulnerability data was rendered as clean: %s", body)
+ }
+ if strings.Contains(body, "vuln-strip") {
+ t.Errorf("a record with no vulnerability data rendered severity counts: %s", body)
+ }
+ if !strings.Contains(body, "SBOM only") {
+ t.Errorf("expected the SBOM-only badge, got: %s", body)
+ }
+}
+
+// TestScanResult_LegacyCleanRecordStillReadsClean guards the discriminator from
+// the other side. Records written before the vulnReportBlob field existed carry
+// an SBOM, zero counts and no status, and they really were scanned clean, so
+// the "no vulnerability data" rule must key on the explicit status "ok" and not
+// swallow them.
+func TestScanResult_LegacyCleanRecordStillReadsClean(t *testing.T) {
+ hold := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
+ if handleMockDID(w, r) {
+ return
+ }
+ w.Header().Set("Content-Type", "application/json")
+ w.Write([]byte(mockScanRecord(0, 0, 0, 0, 0)))
+ }))
+ defer hold.Close()
+
+ handler := setupScanResultHandler(t, hold.URL)
+
+ req := httptest.NewRequest("GET", "/api/scan-result?digest=sha256:abc123&holdEndpoint="+hold.URL, nil)
+ rr := httptest.NewRecorder()
+ handler.ServeHTTP(rr, req)
+
+ if body := rr.Body.String(); !strings.Contains(body, "Clean") {
+ t.Errorf("legacy zero-count record no longer reads clean: %s", body)
+ }
+}
diff --git a/pkg/appview/handlers/vuln_details.go b/pkg/appview/handlers/vuln_details.go
index 20b51d2..1edb0cd 100644
--- a/pkg/appview/handlers/vuln_details.go
+++ b/pkg/appview/handlers/vuln_details.go
@@ -60,12 +60,17 @@ type vulnDetailsData struct {
// NotScanned means the hold answered but holds no scan record for this
// manifest. Distinct from Error: nothing failed, the image was simply
// never scanned, and the UI must not present it as a failure.
- NotScanned bool
- Status string // scan record's status field (ok | failed | skipped); empty for legacy records
- Reason string // scan record's reason field (only meaningful when Status != ok)
- ScannedAt string
- Digest string // image digest (for download URLs)
- HoldEndpoint string // hold DID (for download URLs)
+ NotScanned bool
+ // VulnsNotScanned means the scan succeeded and produced an SBOM, but no
+ // vulnerability data: the scanner ran with vulnerability scanning off. The
+ // zero counts are an absence of measurement, so the panel must say that
+ // rather than report zero findings or blame a failed fetch.
+ VulnsNotScanned bool
+ Status string // scan record's status field (ok | failed | skipped); empty for legacy records
+ Reason string // scan record's reason field (only meaningful when Status != ok)
+ ScannedAt string
+ Digest string // image digest (for download URLs)
+ HoldEndpoint string // hold DID (for download URLs)
}
type vulnMatch struct {
@@ -171,6 +176,18 @@ func (h *VulnDetailsHandler) ServeHTTP(w http.ResponseWriter, r *http.Request) {
Total: scanRecord.Total,
}
+ // A successful record with no report blob and no counts is a scan that
+ // never ran Grype. Distinct from a missing report: there is nothing to
+ // fetch and nothing went wrong.
+ if vulnScanDidNotRun(&scanRecord) {
+ h.renderDetails(w, vulnDetailsData{
+ ScannedAt: scanRecord.ScannedAt,
+ Status: scanRecord.Status,
+ VulnsNotScanned: true,
+ })
+ return
+ }
+
// Step 2: Fetch the vulnerability report blob
if scanRecord.VulnReportBlob == nil || scanRecord.VulnReportBlob.Ref.String() == "" {
h.renderDetails(w, vulnDetailsData{
@@ -343,6 +360,17 @@ func FetchVulnDetails(ctx context.Context, holdEndpoint, digest string) vulnDeta
}
}
+ // A successful record with no report blob and no counts is a scan that
+ // never ran Grype. Distinct from a missing report: there is nothing to
+ // fetch and nothing went wrong.
+ if vulnScanDidNotRun(&scanRecord) {
+ return vulnDetailsData{
+ ScannedAt: scanRecord.ScannedAt,
+ Status: scanRecord.Status,
+ VulnsNotScanned: true,
+ }
+ }
+
// Fetch the vulnerability report blob
if scanRecord.VulnReportBlob == nil || scanRecord.VulnReportBlob.Ref.String() == "" {
return vulnDetailsData{
diff --git a/pkg/appview/handlers/vuln_details_test.go b/pkg/appview/handlers/vuln_details_test.go
index e2d6086..631f7d4 100644
--- a/pkg/appview/handlers/vuln_details_test.go
+++ b/pkg/appview/handlers/vuln_details_test.go
@@ -356,3 +356,63 @@ func TestVulnDetails_MissingParams(t *testing.T) {
t.Error("Expected error message for missing parameters")
}
}
+
+// mockSBOMOnlyRecordEnvelope is the detail-modal counterpart of the badge
+// fixture in scan_result_test.go: status "ok", an SBOM, no vulnerability report
+// and zero counts, which is what a scanner running with vuln.enabled=false
+// produces.
+func mockSBOMOnlyRecordEnvelope() string {
+ record := map[string]any{
+ "$type": "io.atcr.hold.scan",
+ "manifest": "at://did:plc:test/io.atcr.manifest/abc123",
+ "repository": "myapp",
+ "userDid": "did:plc:test",
+ "status": "ok",
+ "critical": 0,
+ "high": 0,
+ "medium": 0,
+ "low": 0,
+ "total": 0,
+ "scannerVersion": "atcr-scanner-v1.0.0",
+ "scannedAt": "2025-01-15T10:30:00Z",
+ }
+ envelope := map[string]any{
+ "uri": "at://did:web:hold.example.com/io.atcr.hold.scan/abc123",
+ "cid": "bafyreiabc123",
+ "value": record,
+ }
+ b, _ := json.Marshal(envelope)
+ return string(b)
+}
+
+// TestVulnDetails_VulnScanDidNotRun is the modal's half of the same rule: a
+// record with no vulnerability data must say so, not report zero findings and
+// not blame the hold for a fetch that never happened.
+func TestVulnDetails_VulnScanDidNotRun(t *testing.T) {
+ hold := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
+ if handleMockDID(w, r) {
+ return
+ }
+ w.Header().Set("Content-Type", "application/json")
+ w.Write([]byte(mockSBOMOnlyRecordEnvelope()))
+ }))
+ defer hold.Close()
+
+ handler := setupVulnDetailsHandler(t)
+
+ req := httptest.NewRequest("GET", "/api/vuln-details?digest=sha256:abc123&holdEndpoint="+hold.URL, nil)
+ rr := httptest.NewRecorder()
+ handler.ServeHTTP(rr, req)
+
+ body := rr.Body.String()
+
+ if !strings.Contains(body, "Vulnerability scanning did not run") {
+ t.Errorf("expected copy saying vulnerability scanning did not run, got: %s", body)
+ }
+ if strings.Contains(body, "0 vulnerabilities") {
+ t.Errorf("a record with no vulnerability data claimed zero vulnerabilities: %s", body)
+ }
+ if strings.Contains(body, "No detailed vulnerability report") {
+ t.Errorf("a scan that never ran Grype was reported as a missing report: %s", body)
+ }
+}
diff --git a/pkg/appview/templates/partials/vuln-badge.html b/pkg/appview/templates/partials/vuln-badge.html
index fa979fb..725d86e 100644
--- a/pkg/appview/templates/partials/vuln-badge.html
+++ b/pkg/appview/templates/partials/vuln-badge.html
@@ -14,6 +14,11 @@
{{ else if .ScanFailed }}
{{ icon "alert-triangle" "size-3" }} Scan failed
+{{ else if .VulnsNotScanned }}
+{{/* The image was catalogued but never matched against a vulnerability
+ database, so its zero counts mean "unmeasured", not "clean". Ghost, like
+ "Not scanned": this is an absence of data, not a good result. */}}
+{{ icon "file-text" "size-3" }} SBOM only
{{ else if eq .Total 0 }}
{{ icon "shield-check" "size-3" }} Clean
{{ else }}
diff --git a/pkg/appview/templates/partials/vuln-details.html b/pkg/appview/templates/partials/vuln-details.html
index da8174d..ed2ea13 100644
--- a/pkg/appview/templates/partials/vuln-details.html
+++ b/pkg/appview/templates/partials/vuln-details.html
@@ -4,6 +4,12 @@
No vulnerability scan available yet
Scans run automatically shortly after a push. Check back in a few minutes, or push a new tag to trigger a scan.
+{{ else if .VulnsNotScanned }}
+
+
Vulnerability scanning did not run for this image
+
The scanner catalogued the image contents, so an SBOM is available, but it was never matched against a vulnerability database. No result here means no data, not a clean bill of health.
+ {{ if .ScannedAt }}
Scanned: {{ .ScannedAt }}
{{ end }}
+
{{ else if .Error }}
{{ if gt .Summary.Total 0 }}
diff --git a/pkg/hold/pds/scan_broadcaster.go b/pkg/hold/pds/scan_broadcaster.go
index 25d7012..a5bc4e7 100644
--- a/pkg/hold/pds/scan_broadcaster.go
+++ b/pkg/hold/pds/scan_broadcaster.go
@@ -13,6 +13,7 @@ import (
"net/url"
"strings"
"sync"
+ "sync/atomic"
"time"
"atcr.io/pkg/atproto"
@@ -38,6 +39,76 @@ const (
// capacity before it says so. A silent stall is what made this class of
// failure invisible until users noticed missing scans.
capacityStallWarnAfter = 10 * time.Minute
+
+ // scanningTimeout is how long a scan may run before the hold gives up on
+ // it. Measured from started_at — the moment a worker told us it picked the
+ // job up — so it budgets scanning and nothing else.
+ //
+ // It used to be measured from assigned_at, which the ack does not refresh.
+ // The scanner acks on receipt, off its WebSocket reader, and the job then
+ // waits in its own queue behind its workers, so the budget covered
+ // queueing: a 100-deep queue of no-op jobs drains in 16m40s and crosses
+ // ten minutes at position 59, and every job past that point was failed
+ // underneath a perfectly healthy scanner.
+ scanningTimeout = 10 * time.Minute
+
+ // queuedTimeout is the fallback budget for a job a scanner acked but never
+ // reported starting. A scanner built before the 'started' message exists
+ // never sends one, so for those the hold can only observe dispatch and this
+ // is the honest bound: long enough for a full default queue (100 jobs) of
+ // real scans to drain ahead of it, short enough that a connected-but-wedged
+ // scanner does not hold its share of dispatch capacity for the life of the
+ // process.
+ queuedTimeout = 60 * time.Minute
+
+ // reconnectGrace is how long a disconnected scanner's in-flight jobs stay
+ // its own.
+ //
+ // A dropped WebSocket is not evidence that a scanner stopped scanning: its
+ // worker pool never learns the socket went away and keeps going. Handing
+ // that work to another process immediately means two processes scan the
+ // same image and both report a verdict. A scanner keeps one identity for
+ // the life of its process and resumes its own rows on reconnect, so this
+ // only has to outlast a reconnect — the client retries every five seconds.
+ // A scanner that actually restarted comes back with a new identity and its
+ // old rows are reclaimed here instead, which is right: that work is gone.
+ reconnectGrace = 2 * time.Minute
+
+ // drainSendTimeout is how long the drain waits for room in a scanner's
+ // send buffer before giving the row back.
+ drainSendTimeout = 5 * time.Second
+
+ // activeJobsErrorBudget is how many consecutive database failures
+ // activeProactiveJobs answers with "assume busy" before it starts
+ // answering "assume idle" instead. See activeProactiveJobs.
+ activeJobsErrorBudget = 3
+
+ // defaultScannerCapacity is how many concurrent scans a scanner that
+ // declares nothing is assumed to run. Every scanner built before the
+ // workers parameter existed lands here, and one is exactly the depth the
+ // hold used to allow hold-wide, so an old scanner against a new hold
+ // behaves as it always did.
+ defaultScannerCapacity = 1
+
+ // maxScannerCapacity caps what a single connection may declare. The value
+ // arrives over the wire behind nothing but a shared secret and is used as
+ // a dispatch budget, so it is clamped rather than trusted.
+ maxScannerCapacity = 32
+
+ // maxScannerInstanceID bounds the scanner-supplied identity that becomes
+ // assigned_to. Same reasoning: it is client input that ends up in a
+ // database column.
+ maxScannerInstanceID = 64
+)
+
+// Job origins. The proactive dispatch gate counts only proactive rows: push
+// scans bypass the gate entirely (oci/xrpc.go calls Enqueue directly), so
+// counting them meant a hold with steady pushes never dispatched a proactive
+// scan at all. Rows that predate the column read as push, which errs toward
+// dispatching rather than toward the stall this whole gate exists to avoid.
+const (
+ originPush = "push"
+ originProactive = "proactive"
)
// ScanBroadcaster manages scanner WebSocket connections and dispatches scan jobs
@@ -71,15 +142,36 @@ type ScanBroadcaster struct {
inflight map[string]struct{} // Manifest digests currently queued or being scanned
inflightMu sync.Mutex
completionSignal chan struct{} // Signaled when a scan job completes (wakes dispatchLoop)
+ capacityFreed chan struct{} // Signaled when a scanner frees a slot (wakes reDispatchLoop)
discoverNow chan struct{} // Signaled to trigger an early discovery pass
+
+ // activeJobsErrs counts consecutive hasActiveJobs query failures, so a
+ // persistent database fault can fail open instead of freezing dispatch.
+ activeJobsErrs atomic.Int64
}
// ScanSubscriber represents a connected scanner WebSocket client
type ScanSubscriber struct {
conn *websocket.Conn
send chan *ScanJobEvent
- id string // Unique subscriber ID
+ id string // Scanner instance identity; also the assigned_to value
done chan struct{}
+
+ // capacity is how many scans this scanner runs at once, as declared on
+ // connect. It is the unit both the proactive dispatch depth and
+ // per-scanner admission control are counted in.
+ capacity int
+}
+
+// effectiveCapacity is capacity with the pre-declaration default applied.
+func (s *ScanSubscriber) effectiveCapacity() int {
+ if s.capacity <= 0 {
+ return defaultScannerCapacity
+ }
+ if s.capacity > maxScannerCapacity {
+ return maxScannerCapacity
+ }
+ return s.capacity
}
// ScanJobEvent is the message sent from hold to scanner over WebSocket
@@ -100,7 +192,7 @@ type ScanJobEvent struct {
// ScannerMessage is a message received from scanner over WebSocket
type ScannerMessage struct {
- Type string `json:"type"` // "ack", "result", "error", "skipped"
+ Type string `json:"type"` // "ack", "started", "result", "error", "skipped"
Seq int64 `json:"seq"` // Job sequence number
SBOM string `json:"sbom,omitempty"`
VulnReport string `json:"vulnReport,omitempty"`
@@ -166,6 +258,7 @@ func NewScanBroadcaster(holdDID, holdEndpoint, secret string, relayEndpoints []s
staleQueue: make(chan *scanCandidate, 200),
inflight: make(map[string]struct{}),
completionSignal: make(chan struct{}, 1),
+ capacityFreed: make(chan struct{}, 1),
discoverNow: make(chan struct{}, 1),
}
@@ -173,6 +266,8 @@ func NewScanBroadcaster(holdDID, holdEndpoint, secret string, relayEndpoints []s
db.Close()
return nil, fmt.Errorf("failed to initialize scan_jobs schema: %w", err)
}
+ sb.reconcileOnBoot()
+
// Start re-dispatch loop for timed-out jobs
sb.wg.Add(1)
go sb.reDispatchLoop()
@@ -212,12 +307,15 @@ func NewScanBroadcasterWithDB(holdDID, holdEndpoint, secret string, relayEndpoin
staleQueue: make(chan *scanCandidate, 200),
inflight: make(map[string]struct{}),
completionSignal: make(chan struct{}, 1),
+ capacityFreed: make(chan struct{}, 1),
discoverNow: make(chan struct{}, 1),
}
if err := sb.initSchema(); err != nil {
return nil, fmt.Errorf("failed to initialize scan_jobs schema: %w", err)
}
+ sb.reconcileOnBoot()
+
sb.wg.Add(1)
go sb.reDispatchLoop()
@@ -252,9 +350,8 @@ func normalizeRelayEndpoints(endpoints []string) []string {
// initSchema creates the scan_jobs table if it doesn't exist
func (sb *ScanBroadcaster) initSchema() error {
- // Execute statements individually for go-libsql compatibility
- stmts := []string{
- `CREATE TABLE IF NOT EXISTS scan_jobs (
+ // Executed individually for go-libsql compatibility
+ if _, err := sb.db.Exec(`CREATE TABLE IF NOT EXISTS scan_jobs (
seq INTEGER PRIMARY KEY AUTOINCREMENT,
manifest_digest TEXT NOT NULL,
repository TEXT NOT NULL,
@@ -271,11 +368,36 @@ func (sb *ScanBroadcaster) initSchema() error {
assigned_at TIMESTAMP,
created_at TIMESTAMP DEFAULT CURRENT_TIMESTAMP,
completed_at TIMESTAMP
- )`,
+ )`); err != nil {
+ return err
+ }
+
+ // Columns added after the table shipped. CREATE TABLE IF NOT EXISTS does
+ // nothing for a database that already has the old shape, so every one of
+ // these has to be added on its own.
+ added := []struct{ name, def string }{
+ // Which dispatcher created the row. The proactive gate counts only its
+ // own work; existing rows default to push so they never throttle it.
+ {"origin", "TEXT NOT NULL DEFAULT 'push'"},
+ // When a worker reported it actually started scanning, as opposed to
+ // when the row was handed out. NULL means the scanner never said.
+ {"started_at", "TIMESTAMP"},
+ // When the scanner holding this row dropped its connection. NULL means
+ // it is connected, or the row was never assigned.
+ {"disconnected_at", "TIMESTAMP"},
+ }
+ for _, col := range added {
+ if err := sb.ensureColumn("scan_jobs", col.name, col.def); err != nil {
+ return err
+ }
+ }
+
+ indexes := []string{
`CREATE INDEX IF NOT EXISTS idx_scan_jobs_status ON scan_jobs(status)`,
`CREATE INDEX IF NOT EXISTS idx_scan_jobs_assigned ON scan_jobs(assigned_to, status)`,
+ `CREATE INDEX IF NOT EXISTS idx_scan_jobs_origin_status ON scan_jobs(origin, status)`,
}
- for _, stmt := range stmts {
+ for _, stmt := range indexes {
if _, err := sb.db.Exec(stmt); err != nil {
return err
}
@@ -283,8 +405,84 @@ func (sb *ScanBroadcaster) initSchema() error {
return nil
}
-// Enqueue inserts a scan job into SQLite and dispatches to the next available scanner
+// reconcileOnBoot treats every job still assigned or processing as belonging to
+// a disconnected scanner, because it does: whatever connections held them died
+// with the previous process.
+//
+// Nothing used to reconcile at boot at all. Rows left mid-flight by a restart
+// sat holding dispatch capacity until their own deadlines fired, which is up
+// to an hour for a job that was queued inside a scanner. Marking them puts
+// them under the same two-minute grace a live disconnect gets: a scanner whose
+// workers are still running them redials with the same identity and resumes
+// them, and one that is not comes back to find them re-offered.
+func (sb *ScanBroadcaster) reconcileOnBoot() {
+ res, err := sb.db.Exec(`
+ UPDATE scan_jobs SET disconnected_at = ?
+ WHERE status IN ('assigned', 'processing') AND disconnected_at IS NULL
+ `, time.Now())
+ if err != nil {
+ slog.Error("Failed to reconcile in-flight scan jobs at boot", "error", err)
+ return
+ }
+ if n, err := res.RowsAffected(); err == nil && n > 0 {
+ slog.Info("Marked in-flight scan jobs from the previous process",
+ "jobs", n, "reclaimAfter", reconnectGrace)
+ }
+}
+
+// ensureColumn adds a column if the table does not already have it. SQLite has
+// no ADD COLUMN IF NOT EXISTS, and the error text for a duplicate is not
+// something worth matching on.
+func (sb *ScanBroadcaster) ensureColumn(table, column, definition string) error {
+ rows, err := sb.db.Query(fmt.Sprintf("PRAGMA table_info(%s)", table))
+ if err != nil {
+ return fmt.Errorf("inspect %s: %w", table, err)
+ }
+ present := false
+ for rows.Next() {
+ var (
+ cid int
+ name string
+ ctype sql.NullString
+ notNull sql.NullInt64
+ defaultVal sql.NullString
+ pk sql.NullInt64
+ )
+ if err := rows.Scan(&cid, &name, &ctype, ¬Null, &defaultVal, &pk); err != nil {
+ rows.Close()
+ return fmt.Errorf("inspect %s: %w", table, err)
+ }
+ if name == column {
+ present = true
+ }
+ }
+ rows.Close()
+ if err := rows.Err(); err != nil {
+ return fmt.Errorf("inspect %s: %w", table, err)
+ }
+ if present {
+ return nil
+ }
+
+ if _, err := sb.db.Exec(fmt.Sprintf("ALTER TABLE %s ADD COLUMN %s %s", table, column, definition)); err != nil {
+ return fmt.Errorf("add %s.%s: %w", table, column, err)
+ }
+ slog.Info("Added scan job column", "table", table, "column", column)
+ return nil
+}
+
+// Enqueue inserts a push-triggered scan job and dispatches it.
+//
+// Push scans are not subject to the proactive dispatch gate — a user who just
+// pushed is waiting for the answer — but they are still subject to
+// per-scanner admission control, so a job that nobody has room for now waits
+// on the hold's queue rather than in a scanner's.
func (sb *ScanBroadcaster) Enqueue(job *ScanJobEvent) error {
+ return sb.enqueue(job, originPush)
+}
+
+// enqueue inserts a scan job into SQLite and dispatches to an available scanner.
+func (sb *ScanBroadcaster) enqueue(job *ScanJobEvent, origin string) error {
job.Type = "job"
job.HoldDID = sb.holdDID
job.HoldEndpoint = sb.holdEndpoint
@@ -294,9 +492,9 @@ func (sb *ScanBroadcaster) Enqueue(job *ScanJobEvent) error {
// Insert into database
result, err := sb.db.Exec(`
- INSERT INTO scan_jobs (manifest_digest, repository, tag, user_did, user_handle, hold_did, hold_endpoint, tier, config_json, layers_json, status)
- VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, ?, 'pending')
- `, job.ManifestDigest, job.Repository, job.Tag, job.UserDID, job.UserHandle, job.HoldDID, job.HoldEndpoint, job.Tier, string(job.Config), string(job.Layers))
+ INSERT INTO scan_jobs (manifest_digest, repository, tag, user_did, user_handle, hold_did, hold_endpoint, tier, config_json, layers_json, status, origin)
+ VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, ?, 'pending', ?)
+ `, job.ManifestDigest, job.Repository, job.Tag, job.UserDID, job.UserHandle, job.HoldDID, job.HoldEndpoint, job.Tier, string(job.Config), string(job.Layers), origin)
if err != nil {
sb.removeInflight(job.ManifestDigest)
return fmt.Errorf("failed to insert scan job: %w", err)
@@ -313,7 +511,8 @@ func (sb *ScanBroadcaster) Enqueue(job *ScanJobEvent) error {
"seq", seq,
"repository", job.Repository,
"tag", job.Tag,
- "tier", job.Tier)
+ "tier", job.Tier,
+ "origin", origin)
// Try to dispatch immediately
sb.dispatchJob(job)
@@ -321,25 +520,42 @@ func (sb *ScanBroadcaster) Enqueue(job *ScanJobEvent) error {
return nil
}
-// Subscribe adds a new scanner WebSocket subscriber and drains pending jobs to it
-func (sb *ScanBroadcaster) Subscribe(conn *websocket.Conn, cursor int64) *ScanSubscriber {
- id := generateSubscriberID()
+// Subscribe adds a new scanner WebSocket subscriber and drains pending jobs to it.
+//
+// instanceID is the scanner's own identity, stable for the life of its process
+// and sent on every connect. It becomes the subscriber id and therefore the
+// assigned_to value, which is what lets a scanner that briefly lost its socket
+// resume the jobs its workers never stopped running. A scanner that declares
+// nothing gets a per-connection id, which is the old behaviour: its in-flight
+// work is not resumable, only reclaimable.
+//
+// capacity is how many scans the scanner runs at once (its worker count).
+func (sb *ScanBroadcaster) Subscribe(conn *websocket.Conn, cursor int64, instanceID string, capacity int) *ScanSubscriber {
+ id := sb.subscriberID(instanceID)
sub := &ScanSubscriber{
conn: conn,
send: make(chan *ScanJobEvent, 20),
id: id,
done: make(chan struct{}),
}
+ sub.capacity = capacity
+
+ // Before anything is dispatched: reclaim whatever this instance was
+ // holding when it dropped, so the drain below counts it against the
+ // scanner's capacity rather than treating it as idle.
+ sb.resumeInstance(id)
sb.mu.Lock()
sb.subscribers = append(sb.subscribers, sub)
+ total := len(sb.subscribers)
sb.mu.Unlock()
slog.Info("Scanner subscribed",
"id", id,
"remote", conn.RemoteAddr(),
"cursor", cursor,
- "totalSubscribers", len(sb.subscribers))
+ "capacity", sub.effectiveCapacity(),
+ "totalSubscribers", total)
// Start writer goroutine (sends jobs to scanner)
go sb.handleWriter(sub)
@@ -356,6 +572,72 @@ func (sb *ScanBroadcaster) Subscribe(conn *websocket.Conn, cursor int64) *ScanSu
return sub
}
+// subscriberID turns a scanner-declared instance identity into the id used for
+// assigned_to, falling back to a random per-connection id.
+//
+// The value arrives over the wire, so it is bounded and restricted to
+// characters that read cleanly in a log line and a database column. A
+// duplicate is refused rather than shared: two processes answering to one id
+// would each accept the other's jobs, which is precisely the confusion the
+// ownership guards exist to prevent.
+func (sb *ScanBroadcaster) subscriberID(instanceID string) string {
+ if instanceID == "" {
+ return generateSubscriberID()
+ }
+ if len(instanceID) > maxScannerInstanceID || !isSafeInstanceID(instanceID) {
+ slog.Warn("Scanner declared an unusable instance id, assigning one",
+ "declared", instanceID)
+ return generateSubscriberID()
+ }
+
+ sb.mu.RLock()
+ defer sb.mu.RUnlock()
+ for _, existing := range sb.subscribers {
+ if existing.id == instanceID {
+ slog.Warn("Two scanners declared the same instance id; assigning one",
+ "instanceId", instanceID)
+ return generateSubscriberID()
+ }
+ }
+ return instanceID
+}
+
+func isSafeInstanceID(id string) bool {
+ for _, r := range id {
+ switch {
+ case r >= 'a' && r <= 'z', r >= 'A' && r <= 'Z', r >= '0' && r <= '9':
+ case r == '-', r == '_', r == '.':
+ default:
+ return false
+ }
+ }
+ return true
+}
+
+// resumeInstance hands a reconnecting scanner back the jobs it was holding
+// when its connection dropped.
+//
+// Clearing disconnected_at is all it takes: the rows never left the scanner,
+// because Unsubscribe marks a disconnect rather than acting on it. The
+// scanner's workers carried on through the outage, so returning the rows to
+// the pool would have had another process scan the same images.
+func (sb *ScanBroadcaster) resumeInstance(id string) {
+ res, err := sb.db.Exec(`
+ UPDATE scan_jobs SET disconnected_at = NULL
+ WHERE assigned_to = ? AND disconnected_at IS NOT NULL
+ AND status IN ('assigned', 'processing')
+ `, id)
+ if err != nil {
+ slog.Error("Failed to resume jobs for reconnecting scanner",
+ "subscriberId", id, "error", err)
+ return
+ }
+ if n, err := res.RowsAffected(); err == nil && n > 0 {
+ slog.Info("Scanner reconnected and resumed its in-flight jobs",
+ "subscriberId", id, "jobs", n)
+ }
+}
+
// Unsubscribe removes a scanner subscriber and makes its jobs re-dispatchable
func (sb *ScanBroadcaster) Unsubscribe(sub *ScanSubscriber) {
sb.mu.Lock()
@@ -378,14 +660,37 @@ func (sb *ScanBroadcaster) Unsubscribe(sub *ScanSubscriber) {
return
}
- // Mark assigned/processing jobs as pending again so they can be re-dispatched.
- // Including 'processing' handles scanner crashes mid-scan.
- _, err := sb.db.Exec(`
- UPDATE scan_jobs SET status = 'pending', assigned_to = NULL, assigned_at = NULL
- WHERE assigned_to = ? AND status IN ('pending', 'assigned', 'processing')
- `, sub.id)
- if err != nil {
- slog.Error("Failed to unassign jobs from disconnected scanner",
+ // Mark this scanner's jobs as belonging to a disconnected scanner, and stop
+ // there.
+ //
+ // This used to flip them straight back to 'pending'. A dropped WebSocket
+ // tells the hold nothing about what the scanner is doing: its worker pool
+ // never learns the socket went away, so it keeps downloading layers and
+ // running Syft on jobs the hold has just put back in the pool. With one
+ // scanner that was a duplicate against itself. With several it is a
+ // duplicate nothing can dedupe — the next process to connect drains the
+ // rows out from under a scanner that is still mid-scan, and both report a
+ // verdict for the same image.
+ //
+ // So the disconnect is recorded and the work is left alone. The same
+ // instance reconnecting resumes it (resumeInstance); a scanner that does
+ // not come back inside reconnectGrace has it reclaimed by
+ // reDispatchTimedOut. Rows that were never handed over — still 'pending'
+ // but stamped with this subscriber — are released outright, since nothing
+ // is running them.
+ if _, err := sb.db.Exec(`
+ UPDATE scan_jobs SET status = 'pending', assigned_to = NULL, assigned_at = NULL, disconnected_at = NULL
+ WHERE assigned_to = ? AND status = 'pending'
+ `, sub.id); err != nil {
+ slog.Error("Failed to release undispatched jobs from disconnected scanner",
+ "subscriberId", sub.id,
+ "error", err)
+ }
+ if _, err := sb.db.Exec(`
+ UPDATE scan_jobs SET disconnected_at = ?
+ WHERE assigned_to = ? AND status IN ('assigned', 'processing') AND disconnected_at IS NULL
+ `, time.Now(), sub.id); err != nil {
+ slog.Error("Failed to mark jobs from disconnected scanner",
"subscriberId", sub.id,
"error", err)
}
@@ -400,7 +705,17 @@ func (sb *ScanBroadcaster) Unsubscribe(sub *ScanSubscriber) {
"totalSubscribers", len(sb.subscribers))
}
-// dispatchJob sends a job to the next available scanner via round-robin
+// dispatchJob hands a job to the connected scanner with the most room for it.
+//
+// Two things changed here when more than one scanner became a supported
+// deployment. Selection is by spare capacity rather than by position, because
+// plain round-robin hands work to a saturated scanner as readily as an idle
+// one, and with heterogeneous processes that is most of the fleet's work going
+// to the wrong place. And a job nobody has room for stays 'pending' instead of
+// being pushed into a scanner's own queue: the hold cannot see into that
+// queue, so anything sitting in it is work it cannot schedule, cannot
+// re-route to a process that freed up first, and cannot put a meaningful
+// deadline on.
func (sb *ScanBroadcaster) dispatchJob(job *ScanJobEvent) {
sb.mu.Lock()
defer sb.mu.Unlock()
@@ -410,9 +725,11 @@ func (sb *ScanBroadcaster) dispatchJob(job *ScanJobEvent) {
return
}
- // Round-robin dispatch
- sub := sb.subscribers[sb.nextIdx%len(sb.subscribers)]
- sb.nextIdx++
+ sub := sb.selectSubscriberLocked()
+ if sub == nil {
+ slog.Debug("Every scanner is at capacity, job stays pending", "seq", job.Seq)
+ return
+ }
// Mark as assigned in database
res, err := sb.db.Exec(`
@@ -442,10 +759,92 @@ func (sb *ScanBroadcaster) dispatchJob(job *ScanJobEvent) {
slog.Warn("Scanner buffer full, re-marking job as pending",
"seq", job.Seq,
"subscriberId", sub.id)
- if _, err := sb.db.Exec(`UPDATE scan_jobs SET status = 'pending', assigned_to = NULL, assigned_at = NULL WHERE seq = ?`, job.Seq); err != nil {
- slog.Error("Failed to re-mark scan job as pending", "seq", job.Seq, "error", err)
+ // Guarded, for the same reason the assign above is. Between our UPDATE
+ // and here another dispatcher can have taken the row; an unguarded
+ // reset would return a job that scanner is already holding to the
+ // pool, and a third scanner would be handed it.
+ sb.releaseUndelivered(sub, job.Seq)
+ }
+}
+
+// selectSubscriberLocked picks the connected scanner with the most spare
+// capacity, or nil when every one of them is full. Caller holds sb.mu.
+//
+// Ties resolve in round-robin order from nextIdx, so a fleet of equal, idle
+// scanners still gets work spread evenly across it.
+func (sb *ScanBroadcaster) selectSubscriberLocked() *ScanSubscriber {
+ n := len(sb.subscribers)
+ if n == 0 {
+ return nil
+ }
+
+ load, ok := sb.subscriberLoads()
+ if !ok {
+ // The load query is the only thing that can say a scanner is full, so
+ // without it admission control has nothing to stand on. Fall back to
+ // plain round-robin: over-dispatching is worse than it was, but it is
+ // far better than dispatching nothing while the database misbehaves.
+ sub := sb.subscribers[sb.nextIdx%n]
+ sb.nextIdx++
+ return sub
+ }
+
+ best := -1
+ bestScore := 0.0
+ for i := 0; i < n; i++ {
+ idx := (sb.nextIdx + i) % n
+ sub := sb.subscribers[idx]
+ capacity := sub.effectiveCapacity()
+ outstanding := load[sub.id]
+ if outstanding >= capacity {
+ continue
+ }
+ // Fraction of the scanner's capacity already committed, so a
+ // four-worker process holding one job outranks a one-worker process
+ // holding none only when it is genuinely emptier.
+ score := float64(outstanding) / float64(capacity)
+ if best < 0 || score < bestScore {
+ best, bestScore = idx, score
}
}
+ if best < 0 {
+ return nil
+ }
+
+ sb.nextIdx = best + 1
+ return sb.subscribers[best]
+}
+
+// subscriberLoads counts the jobs each scanner is currently holding. The
+// second return reports whether the count is usable; callers must not treat a
+// failed query as "everyone is idle".
+func (sb *ScanBroadcaster) subscriberLoads() (map[string]int, bool) {
+ rows, err := sb.db.Query(`
+ SELECT assigned_to, COUNT(*) FROM scan_jobs
+ WHERE status IN ('assigned', 'processing') AND assigned_to IS NOT NULL
+ GROUP BY assigned_to
+ `)
+ if err != nil {
+ slog.Error("Failed to count per-scanner load", "error", err)
+ return nil, false
+ }
+ defer rows.Close()
+
+ load := make(map[string]int)
+ for rows.Next() {
+ var id string
+ var n int
+ if err := rows.Scan(&id, &n); err != nil {
+ slog.Error("Failed to scan per-scanner load row", "error", err)
+ return nil, false
+ }
+ load[id] = n
+ }
+ if err := rows.Err(); err != nil {
+ slog.Error("Failed to read per-scanner load", "error", err)
+ return nil, false
+ }
+ return load, true
}
// handleWriter sends jobs to a scanner over its WebSocket connection
@@ -506,6 +905,8 @@ func (sb *ScanBroadcaster) handleReader(sub *ScanSubscriber) {
switch msg.Type {
case "ack":
sb.handleAck(sub, msg.Seq)
+ case "started":
+ sb.handleStarted(sub, msg.Seq)
case "result":
sb.handleResult(sub, msg)
case "error":
@@ -520,7 +921,14 @@ func (sb *ScanBroadcaster) handleReader(sub *ScanSubscriber) {
}
}
-// handleAck marks a job as processing (scanner received and queued it)
+// handleAck marks a job as processing (scanner received and queued it).
+//
+// The ack means "I have it", nothing more: the scanner sends it off its
+// WebSocket reader the moment a frame arrives, before the job is even queued.
+// It deliberately does not touch assigned_at or started_at — a job can sit
+// acked behind a scanner's workers for a long time without anything being
+// wrong, and the deadlines have to be able to tell that apart from a wedge.
+// The signal for "a worker is on it" is 'started'.
func (sb *ScanBroadcaster) handleAck(sub *ScanSubscriber, seq int64) {
_, err := sb.db.Exec(`
UPDATE scan_jobs SET status = 'processing'
@@ -539,10 +947,83 @@ func (sb *ScanBroadcaster) handleAck(sub *ScanSubscriber, seq int64) {
"subscriberId", sub.id)
}
+// handleStarted records that a worker has actually begun this scan, which is
+// the moment the scanning deadline is measured from.
+//
+// Without it the hold could only observe dispatch, and the ten-minute budget
+// covered however long the job spent queued inside the scanner — so a healthy
+// scanner working through a backlog had its work cancelled underneath it, each
+// cancellation now writing a failed scan record the user can see. A scanner
+// too old to send this message is not broken by its absence: started_at stays
+// NULL and the job falls under queuedTimeout instead, a budget sized for what
+// the hold can actually see.
+//
+// The assigned_to guard is what makes the message safe with several scanners
+// connected, and started_at is stamped once: a job has one beginning, and a
+// repeat must not roll the deadline forward.
+func (sb *ScanBroadcaster) handleStarted(sub *ScanSubscriber, seq int64) {
+ res, err := sb.db.Exec(`
+ UPDATE scan_jobs SET status = 'processing', started_at = ?
+ WHERE seq = ? AND assigned_to = ?
+ AND status IN ('assigned', 'processing') AND started_at IS NULL
+ `, time.Now(), seq, sub.id)
+ if err != nil {
+ slog.Error("Failed to record scan start",
+ "seq", seq, "subscriberId", sub.id, "error", err)
+ return
+ }
+ if n, err := res.RowsAffected(); err == nil && n == 0 {
+ slog.Debug("Ignoring 'started' for a job this scanner does not hold",
+ "seq", seq, "subscriberId", sub.id)
+ return
+ }
+
+ slog.Info("Scan job started",
+ "seq", seq,
+ "subscriberId", sub.id)
+}
+
+// claimJobForTerminal reads the row a terminal message refers to and refuses
+// it unless this scanner is the one holding it.
+//
+// Only handleAck used to check. With one scanner that was harmless; with
+// several it means any scanner can complete, fail or skip another's job —
+// writing a scan record for an image it never looked at, releasing a digest a
+// different process is still scanning, and freeing dispatch capacity that is
+// still in use. The verdict of whichever message lands first wins.
+//
+// A row whose assigned_to no longer matches has moved on: reclaimed after a
+// disconnect, or handed to another process. Its old holder's answer is stale
+// by definition and is dropped rather than applied.
+func (sb *ScanBroadcaster) claimJobForTerminal(sub *ScanSubscriber, seq int64, kind string) bool {
+ var owner sql.NullString
+ err := sb.db.QueryRow(`SELECT assigned_to FROM scan_jobs WHERE seq = ?`, seq).Scan(&owner)
+ if err != nil {
+ slog.Error("Failed to check scan job ownership",
+ "seq", seq, "type", kind, "subscriberId", sub.id, "error", err)
+ return false
+ }
+ if !owner.Valid || owner.String != sub.id {
+ slog.Warn("Ignoring scan message for a job assigned to another scanner",
+ "seq", seq,
+ "type", kind,
+ "subscriberId", sub.id,
+ "assignedTo", owner.String)
+ return false
+ }
+ return true
+}
+
// handleResult processes a completed scan result: uploads SBOM blob + stores scan record in PDS
func (sb *ScanBroadcaster) handleResult(sub *ScanSubscriber, msg ScannerMessage) {
ctx := context.Background()
+ // Before the S3 uploads, not after: a scanner that does not hold this job
+ // should not get its payload stored either.
+ if !sb.claimJobForTerminal(sub, msg.Seq, "result") {
+ return
+ }
+
slog.Info("Scan result received",
"seq", msg.Seq,
"subscriberId", sub.id,
@@ -595,29 +1076,49 @@ func (sb *ScanBroadcaster) handleResult(sub *ScanSubscriber, msg ScannerMessage)
}
}
- // Store scan result as a record in the hold's embedded PDS
+ // Store scan result as a record in the hold's embedded PDS.
+ //
+ // A result with no summary is a completed scan from a scanner running with
+ // vulnerability scanning off: it produced an SBOM, Grype never ran. The
+ // record is written either way, because the SBOM blob is already in S3 by
+ // this point and nothing but the record would reference it. The counts stay
+ // zero and vulnReportBlob stays nil, which is how the appview tells "not
+ // scanned for vulnerabilities" apart from "scanned, found none" — see
+ // classifyScanRecord in pkg/appview/handlers/scan_result.go. Do not fill in
+ // a zeroed summary here: that would report every image as clean.
+ var critical, high, medium, low, total int
if msg.Summary != nil {
- scanRecord := atproto.NewScanRecord(
- manifestDigest, repository, userDID,
- sbomBlob, vulnReportBlob,
- msg.Summary.Critical, msg.Summary.High, msg.Summary.Medium, msg.Summary.Low, msg.Summary.Total,
- "atcr-scanner-v1.0.0",
- )
+ critical = msg.Summary.Critical
+ high = msg.Summary.High
+ medium = msg.Summary.Medium
+ low = msg.Summary.Low
+ total = msg.Summary.Total
+ }
- rpath, _, err := sb.pds.CreateScanRecord(ctx, scanRecord)
- if err != nil {
- slog.Error("Failed to store scan record in PDS",
- "seq", msg.Seq,
- "error", err)
- } else {
- slog.Info("Scan record stored in PDS",
- "rpath", rpath,
- "manifest", scanRecord.Manifest,
- "critical", msg.Summary.Critical,
- "high", msg.Summary.High,
- "total", msg.Summary.Total)
- }
+ scanRecord := atproto.NewScanRecord(
+ manifestDigest, repository, userDID,
+ sbomBlob, vulnReportBlob,
+ critical, high, medium, low, total,
+ "atcr-scanner-v1.0.0",
+ )
+ rpath, _, err := sb.pds.CreateScanRecord(ctx, scanRecord)
+ if err != nil {
+ slog.Error("Failed to store scan record in PDS",
+ "seq", msg.Seq,
+ "error", err)
+ } else if msg.Summary != nil {
+ slog.Info("Scan record stored in PDS",
+ "rpath", rpath,
+ "manifest", scanRecord.Manifest,
+ "critical", critical,
+ "high", high,
+ "total", total)
+ } else {
+ slog.Info("Scan record stored in PDS",
+ "rpath", rpath,
+ "manifest", scanRecord.Manifest,
+ "vulnerabilities", "not scanned")
}
// Mark job as completed
@@ -635,13 +1136,21 @@ func (sb *ScanBroadcaster) handleResult(sub *ScanSubscriber, msg ScannerMessage)
sb.removeInflight(manifestDigest)
sb.signalCompletion()
- slog.Info("Scan job completed",
- "seq", msg.Seq,
- "repository", repository,
- "tag", tag,
- "critical", msg.Summary.Critical,
- "high", msg.Summary.High,
- "total", msg.Summary.Total)
+ if msg.Summary != nil {
+ slog.Info("Scan job completed",
+ "seq", msg.Seq,
+ "repository", repository,
+ "tag", tag,
+ "critical", msg.Summary.Critical,
+ "high", msg.Summary.High,
+ "total", msg.Summary.Total)
+ } else {
+ slog.Info("Scan job completed",
+ "seq", msg.Seq,
+ "repository", repository,
+ "tag", tag,
+ "vulnerabilities", "not scanned")
+ }
}
// handleError marks a job as failed and creates a scan record so the stale
@@ -650,6 +1159,10 @@ func (sb *ScanBroadcaster) handleResult(sub *ScanSubscriber, msg ScannerMessage)
func (sb *ScanBroadcaster) handleError(sub *ScanSubscriber, msg ScannerMessage) {
ctx := context.Background()
+ if !sb.claimJobForTerminal(sub, msg.Seq, "error") {
+ return
+ }
+
var manifestDigest, repository, userDID string
err := sb.db.QueryRow(`
SELECT manifest_digest, repository, user_did
@@ -695,6 +1208,10 @@ func (sb *ScanBroadcaster) handleError(sub *ScanSubscriber, msg ScannerMessage)
func (sb *ScanBroadcaster) handleSkipped(sub *ScanSubscriber, msg ScannerMessage) {
ctx := context.Background()
+ if !sb.claimJobForTerminal(sub, msg.Seq, "skipped") {
+ return
+ }
+
var manifestDigest, repository, userDID string
err := sb.db.QueryRow(`
SELECT manifest_digest, repository, user_did
@@ -737,6 +1254,12 @@ func (sb *ScanBroadcaster) handleSkipped(sub *ScanSubscriber, msg ScannerMessage
// drainPendingJobs sends pending/timed-out jobs to a newly connected scanner.
// Collects all pending rows first, closes cursor, then assigns and dispatches
// to avoid holding a SELECT cursor open during UPDATEs (prevents SQLite BUSY).
+//
+// It stops at the scanner's capacity. It used to walk every pending row, so
+// with a backlog and several scanner processes the first one to connect took
+// all of it and the rest stayed idle — horizontal scaling defeated at the
+// connect path rather than at the dispatch gate. What it leaves behind is
+// picked up by offerPendingJobs as soon as anything frees up.
func (sb *ScanBroadcaster) drainPendingJobs(sub *ScanSubscriber, cursor int64) {
rows, err := sb.db.Query(`
SELECT seq, manifest_digest, repository, tag, user_did, user_handle, hold_did, hold_endpoint, tier, config_json, layers_json
@@ -770,23 +1293,45 @@ func (sb *ScanBroadcaster) drainPendingJobs(sub *ScanSubscriber, cursor int64) {
}
rows.Close()
+ // Anything this scanner is already holding — jobs it resumed after a
+ // reconnect, most often — counts against what it can take now.
+ budget := sub.effectiveCapacity()
+ if load, ok := sb.subscriberLoads(); ok {
+ budget -= load[sub.id]
+ }
+
count := 0
for _, job := range jobs {
- _, err = sb.db.Exec(`
+ if count >= budget {
+ slog.Debug("Drain reached the scanner's capacity, leaving the rest pending",
+ "subscriberId", sub.id, "capacity", sub.effectiveCapacity())
+ break
+ }
+ res, err := sb.db.Exec(`
UPDATE scan_jobs SET status = 'assigned', assigned_to = ?, assigned_at = ?
WHERE seq = ? AND status = 'pending'
`, sub.id, time.Now(), job.Seq)
if err != nil {
continue
}
+ // The rows were selected up front, so dispatchJob or the re-dispatch
+ // loop can have claimed one in between. Sending it anyway would scan
+ // it twice — and the losing scanner's ack is dropped by handleAck's
+ // own guard, so nothing downstream would notice the duplicate.
+ if n, err := res.RowsAffected(); err == nil && n == 0 {
+ slog.Debug("Scan job no longer pending, skipping drain", "seq", job.Seq)
+ continue
+ }
select {
case sub.send <- job:
count++
case <-sub.done:
+ sb.releaseUndelivered(sub, job.Seq)
return
- case <-time.After(5 * time.Second):
- slog.Warn("Drain timeout for scanner", "subscriberId", sub.id)
+ case <-time.After(drainSendTimeout):
+ slog.Warn("Drain timeout for scanner", "subscriberId", sub.id, "seq", job.Seq)
+ sb.releaseUndelivered(sub, job.Seq)
return
}
}
@@ -798,7 +1343,28 @@ func (sb *ScanBroadcaster) drainPendingJobs(sub *ScanSubscriber, cursor int64) {
}
}
-// reDispatchLoop periodically checks for timed-out jobs and re-dispatches them
+// releaseUndelivered puts back a row this drain claimed but could not hand to
+// the scanner. Without it the row stays 'assigned' to a subscriber that was
+// never sent it, counting as active dispatch capacity until the five-minute
+// ack timeout reclaims it.
+//
+// The assigned_to guard is what makes this safe to run after the subscriber is
+// already gone: if Unsubscribe's bulk reset or another dispatcher has since
+// taken the row, this matches nothing.
+func (sb *ScanBroadcaster) releaseUndelivered(sub *ScanSubscriber, seq int64) {
+ _, err := sb.db.Exec(`
+ UPDATE scan_jobs SET status = 'pending', assigned_to = NULL, assigned_at = NULL
+ WHERE seq = ? AND assigned_to = ? AND status = 'assigned'
+ `, seq, sub.id)
+ if err != nil {
+ slog.Error("Failed to release undelivered scan job",
+ "seq", seq, "subscriberId", sub.id, "error", err)
+ }
+}
+
+// reDispatchLoop periodically checks for timed-out jobs and re-dispatches them,
+// and fills freed scanner capacity as soon as it is freed rather than on the
+// next tick.
func (sb *ScanBroadcaster) reDispatchLoop() {
defer sb.wg.Done()
@@ -811,13 +1377,79 @@ func (sb *ScanBroadcaster) reDispatchLoop() {
return
case <-ticker.C:
sb.reDispatchTimedOut()
+ case <-sb.capacityFreed:
+ sb.offerPendingJobs()
}
}
}
+// offerPendingJobs hands waiting rows to whichever scanners have room, oldest
+// first, and stops as soon as nobody does.
+//
+// This is the other half of admission control. dispatchJob now leaves a row
+// pending rather than pushing it into a saturated scanner's own queue, which
+// is only affordable if the row is offered again the moment something frees
+// up. Waiting for the thirty-second re-dispatch tick would have cost more
+// throughput than the scanner-side queue ever bought.
+//
+// No age guard is needed here, unlike the reclaim in reDispatchTimedOut:
+// dispatchJob's assign is conditional on the row still being pending and reads
+// RowsAffected, so racing another dispatcher costs a skipped row, not a
+// double send.
+func (sb *ScanBroadcaster) offerPendingJobs() {
+ if !sb.hasConnectedScanners() {
+ return
+ }
+
+ rows, err := sb.db.Query(`
+ SELECT seq, manifest_digest, repository, tag, user_did, user_handle, hold_did, hold_endpoint, tier, config_json, layers_json
+ FROM scan_jobs
+ WHERE status = 'pending'
+ ORDER BY seq ASC
+ `)
+ if err != nil {
+ slog.Error("Failed to query pending scan jobs", "error", err)
+ return
+ }
+
+ var jobs []*ScanJobEvent
+ for rows.Next() {
+ job := &ScanJobEvent{Type: "job"}
+ var configJSON, layersJSON string
+ if err := rows.Scan(
+ &job.Seq, &job.ManifestDigest, &job.Repository, &job.Tag,
+ &job.UserDID, &job.UserHandle, &job.HoldDID, &job.HoldEndpoint,
+ &job.Tier, &configJSON, &layersJSON,
+ ); err != nil {
+ slog.Error("Failed to scan pending job row", "error", err)
+ continue
+ }
+ job.Config = json.RawMessage(configJSON)
+ job.Layers = json.RawMessage(layersJSON)
+ jobs = append(jobs, job)
+ }
+ rows.Close()
+
+ for _, job := range jobs {
+ if !sb.hasFreeScannerCapacity() {
+ return
+ }
+ sb.dispatchJob(job)
+ }
+}
+
+// hasFreeScannerCapacity reports whether any connected scanner has room for
+// another job right now.
+func (sb *ScanBroadcaster) hasFreeScannerCapacity() bool {
+ sb.mu.Lock()
+ defer sb.mu.Unlock()
+ return sb.selectSubscriberLocked() != nil
+}
+
// reDispatchTimedOut finds jobs that were assigned but not acked/completed within timeout,
// re-offers jobs that have been sitting in 'pending' with nobody to hand them to,
-// and also marks stuck processing jobs as failed.
+// reclaims work from a scanner that disconnected and did not come back inside
+// reconnectGrace, and also marks stuck processing jobs as failed.
// Collects timed-out rows first, closes cursor, then resets and re-dispatches
// to avoid holding a SELECT cursor open during UPDATEs (prevents SQLite BUSY).
//
@@ -829,25 +1461,17 @@ func (sb *ScanBroadcaster) reDispatchLoop() {
func (sb *ScanBroadcaster) reDispatchTimedOut() {
timeout := time.Now().Add(-sb.ackTimeout)
- // Fail processing jobs stuck for >10 minutes (scanner likely crashed mid-scan)
- processingTimeout := time.Now().Add(-10 * time.Minute)
- res, err := sb.db.Exec(`
- UPDATE scan_jobs SET status = 'failed', completed_at = ?
- WHERE status = 'processing' AND assigned_at < ?
- `, time.Now(), processingTimeout)
- if err != nil {
- slog.Error("Failed to clean up stuck processing jobs", "error", err)
- } else if n, _ := res.RowsAffected(); n > 0 {
- slog.Warn("Cleaned up stuck processing jobs", "count", n)
- }
+ // Fail processing jobs stuck past the deadline (scanner likely crashed mid-scan)
+ sb.failStuckProcessingJobs()
rows, err := sb.db.Query(`
SELECT seq, manifest_digest, repository, tag, user_did, user_handle, hold_did, hold_endpoint, tier, config_json, layers_json, status
FROM scan_jobs
- WHERE (status = 'assigned' AND assigned_at < ?)
+ WHERE (status = 'assigned' AND disconnected_at IS NULL AND assigned_at < ?)
OR (status = 'pending' AND datetime(created_at) < datetime('now', ?))
+ OR (status IN ('assigned', 'processing') AND disconnected_at IS NOT NULL AND disconnected_at < ?)
ORDER BY seq ASC
- `, timeout, sqliteAgoModifier(pendingReclaimAfter))
+ `, timeout, sqliteAgoModifier(pendingReclaimAfter), time.Now().Add(-reconnectGrace))
if err != nil {
slog.Error("Failed to query timed-out scan jobs", "error", err)
return
@@ -886,7 +1510,8 @@ func (sb *ScanBroadcaster) reDispatchTimedOut() {
// so dispatchJob's status guard can drop it instead of double-sending.
if r.status != "pending" {
_, err = sb.db.Exec(`
- UPDATE scan_jobs SET status = 'pending', assigned_to = NULL, assigned_at = NULL
+ UPDATE scan_jobs SET status = 'pending', assigned_to = NULL,
+ assigned_at = NULL, started_at = NULL, disconnected_at = NULL
WHERE seq = ?
`, job.Seq)
if err != nil {
@@ -903,6 +1528,109 @@ func (sb *ScanBroadcaster) reDispatchTimedOut() {
}
}
+// failStuckProcessingJobs gives up on jobs a scanner acknowledged but never
+// answered for, and retires them the same way every other terminal transition
+// does: a scan record, the in-flight digest released, dispatch capacity
+// signalled.
+//
+// It used to be a bare UPDATE. That made it the only terminal transition that
+// wrote no record and never called removeInflight, and both halves of that
+// were user-visible. Without the record the appview cannot tell a job the hold
+// gave up on from one enqueued thirty seconds ago, so the image shows a grey
+// "Not scanned" forever. Without removeInflight the digest stays in sb.inflight
+// for the life of the process, and discoverUnscannedForUser and runStalePass
+// both skip any manifest already in that set — so every timeout permanently
+// retired one image from scanning.
+//
+// The record is a failure rather than a skip because a hung or crashed scanner
+// is a transient condition: the stale loop should retry it on the rescan
+// interval, which is exactly what it does for failed records and not for
+// skipped ones.
+func (sb *ScanBroadcaster) failStuckProcessingJobs() {
+ type stuckJob struct {
+ seq int64
+ manifestDigest string
+ repository string
+ userDID string
+ started bool
+ }
+
+ // Select first, close the cursor, then write — the same shape as the
+ // sibling loops below, and required here because the digests are needed
+ // after the UPDATE has already erased which rows were affected.
+ // Two deadlines, because there are two things a 'processing' row can be.
+ //
+ // A row with started_at is being scanned right now and gets scanningTimeout
+ // from that moment. A row without it was acked and is queued inside the
+ // scanner — either because it genuinely has not reached a worker yet, or
+ // because the scanner is too old to say — and gets queuedTimeout from
+ // dispatch, which is the only clock the hold has for it.
+ scanDeadline := time.Now().Add(-scanningTimeout)
+ queueDeadline := time.Now().Add(-queuedTimeout)
+
+ rows, err := sb.db.Query(`
+ SELECT seq, manifest_digest, repository, user_did, started_at IS NOT NULL
+ FROM scan_jobs
+ WHERE status = 'processing'
+ AND ((started_at IS NOT NULL AND started_at < ?)
+ OR (started_at IS NULL AND assigned_at < ?))
+ `, scanDeadline, queueDeadline)
+ if err != nil {
+ slog.Error("Failed to query stuck processing jobs", "error", err)
+ return
+ }
+
+ var jobs []stuckJob
+ for rows.Next() {
+ var j stuckJob
+ if err := rows.Scan(&j.seq, &j.manifestDigest, &j.repository, &j.userDID, &j.started); err != nil {
+ slog.Error("Failed to scan stuck processing job row", "error", err)
+ continue
+ }
+ jobs = append(jobs, j)
+ }
+ rows.Close()
+
+ for _, j := range jobs {
+ res, err := sb.db.Exec(`
+ UPDATE scan_jobs SET status = 'failed', completed_at = ?
+ WHERE seq = ? AND status = 'processing'
+ AND ((started_at IS NOT NULL AND started_at < ?)
+ OR (started_at IS NULL AND assigned_at < ?))
+ `, time.Now(), j.seq, scanDeadline, queueDeadline)
+ if err != nil {
+ slog.Error("Failed to fail stuck processing job", "seq", j.seq, "error", err)
+ continue
+ }
+ // The scanner answered — or started, which moves the row onto the
+ // other deadline — between the SELECT and here.
+ if n, err := res.RowsAffected(); err == nil && n == 0 {
+ continue
+ }
+
+ reason := fmt.Sprintf("scanner did not report a result within %s of starting the scan", scanningTimeout)
+ if !j.started {
+ reason = fmt.Sprintf("scanner acknowledged the job but never started it within %s", queuedTimeout)
+ }
+ record := atproto.NewFailedScanRecord(
+ j.manifestDigest, j.repository, j.userDID,
+ reason,
+ "atcr-scanner-v1.0.0",
+ )
+ if _, _, err := sb.pds.CreateScanRecord(context.Background(), record); err != nil {
+ slog.Error("Failed to store timeout scan record", "seq", j.seq, "error", err)
+ }
+
+ sb.removeInflight(j.manifestDigest)
+ sb.signalCompletion()
+
+ slog.Warn("Scan job timed out in processing",
+ "seq", j.seq,
+ "repository", j.repository,
+ "manifest", j.manifestDigest)
+ }
+}
+
// sqliteAgoModifier renders a duration as a SQLite datetime() modifier that
// walks backwards from 'now', e.g. 15m becomes "-900 seconds".
func sqliteAgoModifier(d time.Duration) string {
@@ -1259,8 +1987,8 @@ func (sb *ScanBroadcaster) runStalePass() {
}
// dispatchLoop pops candidates from the work queues with strict priority
-// (unscanned before stale) and enqueues them as scan jobs. Throttled to one
-// proactive job at a time via hasActiveJobs().
+// (unscanned before stale) and enqueues them as scan jobs, throttled by
+// waitForProactiveCapacity to one proactive job per connected scanner worker.
func (sb *ScanBroadcaster) dispatchLoop() {
defer sb.wg.Done()
@@ -1281,12 +2009,10 @@ func (sb *ScanBroadcaster) dispatchLoop() {
default:
}
- // Wait until there's capacity (no active proactive jobs)
- if !sb.waitForCapacity() {
- return // stopCh closed
- }
-
- // Wait until at least one scanner is connected
+ // Wait until at least one scanner is connected. This comes first
+ // because the capacity gate is derived from what is connected: with
+ // nothing there the budget is zero and the gate has nothing to wait
+ // for.
if !sb.hasConnectedScanners() {
select {
case <-sb.stopCh:
@@ -1296,6 +2022,13 @@ func (sb *ScanBroadcaster) dispatchLoop() {
continue
}
+ // Wait until the connected scanners have room for another proactive
+ // job. Returns false on shutdown and on the fleet emptying out, both
+ // of which are handled by looping back to the top.
+ if !sb.waitForProactiveCapacity() {
+ continue
+ }
+
// Pop from highest-priority non-empty queue
var candidate *scanCandidate
@@ -1323,14 +2056,29 @@ func (sb *ScanBroadcaster) dispatchLoop() {
}
}
-// waitForCapacity blocks until there are no active proactive scan jobs.
-// Returns false if stopCh is closed.
-func (sb *ScanBroadcaster) waitForCapacity() bool {
+// waitForProactiveCapacity blocks until the connected scanners have room for
+// another proactive job. Returns false when the caller should re-evaluate from
+// the top instead — shutdown, or every scanner having gone away.
+//
+// It must not block on zero capacity: the budget is derived from the
+// subscriber list, and a gate parked inside its own loop cannot notice a
+// scanner arriving or the loop being asked to stop for up to five seconds.
+func (sb *ScanBroadcaster) waitForProactiveCapacity() bool {
blockedSince := time.Now()
var lastWarn time.Time
for {
- if !sb.hasActiveJobs() {
+ select {
+ case <-sb.stopCh:
+ return false
+ default:
+ }
+
+ limit := sb.proactiveDispatchLimit()
+ if limit == 0 {
+ return false
+ }
+ if n, ok := sb.activeProactiveJobs(); ok && n < limit {
return true
}
@@ -1351,6 +2099,42 @@ func (sb *ScanBroadcaster) waitForCapacity() bool {
}
}
+// proactiveDispatchLimit is how many proactive scan jobs may be in flight at
+// once: one per worker, summed over every connected scanner.
+//
+// The depth used to be one, hold-wide, which defeated both ways of scaling at
+// the same time — a second worker in a scanner and a second scanner process
+// were equally unable to receive proactive work. Deriving it from declared
+// capacity is what makes `scanner.workers: 2` and a second scanner process
+// mean something.
+//
+// One per worker and not more. The scanner acks on receipt and queues
+// internally, so anything beyond one per worker is backlog the hold cannot
+// see, cannot re-route to a process that frees up first, and cannot put an
+// honest deadline on. Zero when nothing is connected, which is what stops the
+// loop from manufacturing work for a fleet that is not there.
+func (sb *ScanBroadcaster) proactiveDispatchLimit() int {
+ sb.mu.RLock()
+ defer sb.mu.RUnlock()
+
+ total := 0
+ for _, sub := range sb.subscribers {
+ total += sub.effectiveCapacity()
+ }
+ return total
+}
+
+// hasProactiveCapacity is the non-blocking form of the gate, for callers that
+// want an answer rather than a wait.
+func (sb *ScanBroadcaster) hasProactiveCapacity() bool {
+ limit := sb.proactiveDispatchLimit()
+ if limit == 0 {
+ return false
+ }
+ n, ok := sb.activeProactiveJobs()
+ return ok && n < limit
+}
+
// dispatchCandidate resolves manifest details if needed and enqueues a scan job.
func (sb *ScanBroadcaster) dispatchCandidate(candidate *scanCandidate) {
ctx, cancel := context.WithTimeout(context.Background(), 30*time.Second)
@@ -1392,7 +2176,7 @@ func (sb *ScanBroadcaster) dispatchCandidate(candidate *scanCandidate) {
"userDID", candidate.userDID,
"reason", reason)
- if err := sb.Enqueue(&ScanJobEvent{
+ if err := sb.enqueue(&ScanJobEvent{
ManifestDigest: candidate.manifestDigest,
Repository: candidate.manifest.Repository,
UserDID: candidate.userDID,
@@ -1400,7 +2184,7 @@ func (sb *ScanBroadcaster) dispatchCandidate(candidate *scanCandidate) {
Tier: "deckhand",
Config: configJSON,
Layers: layersJSON,
- }); err != nil {
+ }, originProactive); err != nil {
slog.Error("Dispatch: failed to enqueue",
"manifest", candidate.manifestDigest, "error", err)
// removeInflight not needed — Enqueue already cleans up on error
@@ -1551,8 +2335,15 @@ func (sb *ScanBroadcaster) hasConnectedScanners() bool {
return len(sb.subscribers) > 0
}
-// hasActiveJobs returns true if there are any assigned or processing scan jobs,
-// or any recently-created pending ones.
+// activeProactiveJobs counts the proactive scan jobs currently holding
+// dispatch capacity. The second return reports whether the count can be acted
+// on; a caller must not read (0, false) as "idle".
+//
+// Only proactive rows are counted. Push-triggered scans do not go through this
+// gate at all — oci/xrpc.go calls Enqueue directly, because a user who just
+// pushed is waiting — so counting them meant a hold with steady pushes never
+// dispatched a proactive scan, and "one proactive job at a time" was really
+// "none while anyone is pushing".
//
// Pending rows older than pendingStaleAfter are deliberately not counted. A job
// that has been pending that long is one no scanner can be given (dispatch
@@ -1561,18 +2352,38 @@ func (sb *ScanBroadcaster) hasConnectedScanners() bool {
// undispatchable job stopped scanning deployment-wide. The re-dispatch loop
// keeps re-offering such rows, so ignoring them here costs nothing when a
// scanner is available.
-func (sb *ScanBroadcaster) hasActiveJobs() bool {
+//
+// A query failure is not evidence of activity. The first few are answered
+// "busy" anyway, because guessing "idle" during a blip piles another job onto
+// a scanner that may already have one. A persistent failure is different: this
+// is the only gate on proactive dispatch and waitForProactiveCapacity spins on
+// it, so answering "busy" forever halts scanning for the life of the process
+// with no recovery and nothing but one log line every five seconds to show for
+// it — logStalledCapacity runs the same database and takes its own error
+// branch. After activeJobsErrorBudget consecutive failures it therefore fails
+// open, reporting an idle fleet so dispatch resumes.
+func (sb *ScanBroadcaster) activeProactiveJobs() (int, bool) {
var count int
err := sb.db.QueryRow(`
SELECT COUNT(*) FROM scan_jobs
- WHERE status IN ('assigned', 'processing')
- OR (status = 'pending' AND datetime(created_at) > datetime('now', ?))
- `, sqliteAgoModifier(pendingStaleAfter)).Scan(&count)
+ WHERE origin = ?
+ AND (status IN ('assigned', 'processing')
+ OR (status = 'pending' AND datetime(created_at) > datetime('now', ?)))
+ `, originProactive, sqliteAgoModifier(pendingStaleAfter)).Scan(&count)
if err != nil {
- slog.Error("Failed to check active scan jobs", "error", err)
- return true // Assume busy on error
+ failures := sb.activeJobsErrs.Add(1)
+ if failures <= activeJobsErrorBudget {
+ slog.Error("Failed to check active scan jobs, assuming busy",
+ "error", err, "consecutiveFailures", failures)
+ return 0, false
+ }
+ slog.Error("Failed to check active scan jobs; proceeding as if idle so a "+
+ "database fault does not halt proactive scanning outright",
+ "error", err, "consecutiveFailures", failures)
+ return 0, true
}
- return count > 0
+ sb.activeJobsErrs.Store(0)
+ return count, true
}
// logStalledCapacity reports what is holding the dispatch loop back, so a stall
@@ -1586,8 +2397,8 @@ func (sb *ScanBroadcaster) logStalledCapacity(blockedFor time.Duration) {
err := sb.db.QueryRow(`
SELECT COUNT(*), MIN(seq), GROUP_CONCAT(DISTINCT status)
FROM scan_jobs
- WHERE status IN ('pending', 'assigned', 'processing')
- `).Scan(&count, &oldestSeq, &statuses)
+ WHERE origin = ? AND status IN ('pending', 'assigned', 'processing')
+ `, originProactive).Scan(&count, &oldestSeq, &statuses)
if err != nil {
slog.Warn("Proactive scan dispatch stalled; could not inspect active jobs",
"blockedFor", blockedFor.Truncate(time.Minute), "error", err)
@@ -1596,7 +2407,8 @@ func (sb *ScanBroadcaster) logStalledCapacity(blockedFor time.Duration) {
slog.Warn("Proactive scan dispatch stalled waiting on active jobs",
"blockedFor", blockedFor.Truncate(time.Minute),
- "activeJobs", count,
+ "activeProactiveJobs", count,
+ "dispatchLimit", sb.proactiveDispatchLimit(),
"oldestSeq", oldestSeq.Int64,
"statuses", statuses.String)
}
@@ -1625,12 +2437,19 @@ func (sb *ScanBroadcaster) removeInflight(digest string) {
delete(sb.inflight, digest)
}
-// signalCompletion non-blocking signal to wake the dispatch loop.
+// signalCompletion non-blocking signal to wake the dispatch loop, and the
+// re-dispatch loop with it: a finished job frees a slot on some scanner, and
+// what is waiting for that slot may be a push-triggered row that the proactive
+// dispatch loop will never look at.
func (sb *ScanBroadcaster) signalCompletion() {
select {
case sb.completionSignal <- struct{}{}:
default:
}
+ select {
+ case sb.capacityFreed <- struct{}{}:
+ default:
+ }
}
// triggerDiscovery non-blocking signal to trigger an early discovery pass.
diff --git a/pkg/hold/pds/scan_broadcaster_concurrency_test.go b/pkg/hold/pds/scan_broadcaster_concurrency_test.go
new file mode 100644
index 0000000..c23e434
--- /dev/null
+++ b/pkg/hold/pds/scan_broadcaster_concurrency_test.go
@@ -0,0 +1,962 @@
+package pds
+
+import (
+ "context"
+ "database/sql"
+ "testing"
+ "time"
+
+ "atcr.io/pkg/atproto"
+)
+
+// These tests describe a hold that can keep more than one scan running at a
+// time — across the workers of a single scanner process (vertical) and across
+// several scanner processes (horizontal).
+//
+// Both were defeated by the same thing: dispatchLoop called waitForCapacity()
+// before every dispatch, that gate blocked until no row anywhere was
+// 'assigned' or 'processing', and then exactly one candidate was dispatched.
+// Proactive scanning was therefore depth one hold-wide, so a second worker
+// could never be given proactive work and neither could a second scanner
+// process, even though dispatchJob already round-robins across subscribers.
+//
+// Two separate defects live in that one gate. It counted every row rather than
+// only the proactive ones, so a push-triggered scan starved proactive dispatch
+// and vice versa; and the depth was hardcoded at one rather than derived from
+// how much scanning capacity is actually connected.
+
+// newConcurrencyBroadcaster is newRecordingScanBroadcaster plus the channels
+// the dispatch gate selects on. The bare helpers leave stopCh nil, which makes
+// every select in waitForProactiveCapacity block forever.
+func newConcurrencyBroadcaster(t *testing.T) *ScanBroadcaster {
+ t.Helper()
+
+ sb := newRecordingScanBroadcaster(t)
+ sb.stopCh = make(chan struct{})
+ sb.completionSignal = make(chan struct{}, 1)
+ t.Cleanup(func() {
+ select {
+ case <-sb.stopCh:
+ default:
+ close(sb.stopCh)
+ }
+ })
+ return sb
+}
+
+// seedJob inserts one pending job with an explicit origin, which is what the
+// proactive capacity gate keys on.
+func seedJob(t *testing.T, sb *ScanBroadcaster, digest, origin string) int64 {
+ t.Helper()
+
+ res, err := sb.db.Exec(`
+ INSERT INTO scan_jobs
+ (manifest_digest, repository, tag, user_did, user_handle,
+ hold_did, hold_endpoint, tier, config_json, layers_json, status, origin)
+ VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, ?, 'pending', ?)
+ `, digest, "repo", "latest", "did:plc:user", "user.example.com",
+ sb.holdDID, sb.holdEndpoint, "deckhand", "{}", "[]", origin)
+ if err != nil {
+ t.Fatalf("seed %s job: %v", origin, err)
+ }
+ seq, err := res.LastInsertId()
+ if err != nil {
+ t.Fatalf("seq: %v", err)
+ }
+ return seq
+}
+
+func setStatus(t *testing.T, sb *ScanBroadcaster, seq int64, status string) {
+ t.Helper()
+
+ if _, err := sb.db.Exec(`UPDATE scan_jobs SET status = ? WHERE seq = ?`, status, seq); err != nil {
+ t.Fatalf("set status %q on %d: %v", status, seq, err)
+ }
+}
+
+func assignedTo(t *testing.T, sb *ScanBroadcaster, seq int64) string {
+ t.Helper()
+
+ var to sql.NullString
+ if err := sb.db.QueryRow(`SELECT assigned_to FROM scan_jobs WHERE seq = ?`, seq).Scan(&to); err != nil {
+ t.Fatalf("query assigned_to for %d: %v", seq, err)
+ }
+ return to.String
+}
+
+// markStarted sets started_at the given number of minutes in the past, as a
+// 'started' message from a worker would have done at that time.
+func markStarted(t *testing.T, sb *ScanBroadcaster, seq int64, agoMinutes int) {
+ t.Helper()
+
+ at := time.Now().Add(-time.Duration(agoMinutes) * time.Minute)
+ if _, err := sb.db.Exec(`UPDATE scan_jobs SET started_at = ? WHERE seq = ?`, at, seq); err != nil {
+ t.Fatalf("mark started %d: %v", seq, err)
+ }
+}
+
+// jobIsDisconnected reports whether the row is marked as belonging to a
+// scanner that has dropped its connection but may still be running it.
+func jobIsDisconnected(t *testing.T, sb *ScanBroadcaster, seq int64) bool {
+ t.Helper()
+
+ var at sql.NullTime
+ if err := sb.db.QueryRow(`SELECT disconnected_at FROM scan_jobs WHERE seq = ?`, seq).Scan(&at); err != nil {
+ t.Fatalf("query disconnected_at for %d: %v", seq, err)
+ }
+ return at.Valid
+}
+
+func startedAt(t *testing.T, sb *ScanBroadcaster, seq int64) sql.NullTime {
+ t.Helper()
+
+ var at sql.NullTime
+ if err := sb.db.QueryRow(`SELECT started_at FROM scan_jobs WHERE seq = ?`, seq).Scan(&at); err != nil {
+ t.Fatalf("query started_at for %d: %v", seq, err)
+ }
+ return at
+}
+
+// ---------------------------------------------------------------------------
+// Dispatch depth
+// ---------------------------------------------------------------------------
+
+// TestScanProactiveDispatchLimit_TracksConnectedScannerWorkers pins what the
+// proactive dispatch depth is derived from: the scanning capacity actually
+// connected, summed over subscribers.
+//
+// A scanner declares its worker count when it subscribes. One that declares
+// nothing — every scanner built before the parameter existed — counts as one
+// worker, which is exactly the depth the hold had before, so an old scanner
+// against a new hold behaves as it always did.
+func TestScanProactiveDispatchLimit_TracksConnectedScannerWorkers(t *testing.T) {
+ sb := newConcurrencyBroadcaster(t)
+
+ if got := sb.proactiveDispatchLimit(); got != 0 {
+ t.Errorf("limit with no scanner connected = %d, want 0: there is "+
+ "nothing to dispatch to", got)
+ }
+
+ old := newTestScanSubscriber(t, sb, 4)
+ old.id = "legacy"
+ if got := sb.proactiveDispatchLimit(); got != 1 {
+ t.Errorf("limit for a scanner that declares no worker count = %d, want 1", got)
+ }
+
+ old.capacity = 2 // the same process reconnecting with workers: 2
+ if got := sb.proactiveDispatchLimit(); got != 2 {
+ t.Errorf("limit for one two-worker scanner = %d, want 2: workers above "+
+ "one must buy proactive throughput", got)
+ }
+
+ second := newTestScanSubscriber(t, sb, 4)
+ second.id = "second-process"
+ second.capacity = 3
+ if got := sb.proactiveDispatchLimit(); got != 5 {
+ t.Errorf("limit across two scanner processes = %d, want 5", got)
+ }
+
+ sb.Unsubscribe(second)
+ if got := sb.proactiveDispatchLimit(); got != 2 {
+ t.Errorf("limit after a scanner disconnected = %d, want 2", got)
+ }
+}
+
+// TestScanProactiveCapacity_AllowsOneJobPerWorker is the vertical scaling case:
+// `scanner.workers: 2` must genuinely produce two concurrent proactive scans.
+//
+// Depth is one proactive job per connected worker. Not more: the scanner acks
+// on receipt and queues internally, so anything beyond one per worker is a
+// backlog the hold cannot see into, which is what made the ten-minute deadline
+// fire under a healthy scanner. Not fewer: one worker would sit idle.
+func TestScanProactiveCapacity_AllowsOneJobPerWorker(t *testing.T) {
+ sb := newConcurrencyBroadcaster(t)
+ sub := newTestScanSubscriber(t, sb, 4)
+ sub.capacity = 2
+
+ first := seedJob(t, sb, "sha256:first", originProactive)
+ sb.dispatchJob(&ScanJobEvent{Type: "job", Seq: first, Repository: "repo"})
+
+ if n, ok := sb.activeProactiveJobs(); !ok || n != 1 {
+ t.Fatalf("active proactive jobs = %d (ok=%v), want 1", n, ok)
+ }
+ if !sb.hasProactiveCapacity() {
+ t.Fatal("a two-worker scanner with one job running has capacity for a " +
+ "second; depth-one dispatch is what defeats scanner.workers")
+ }
+
+ second := seedJob(t, sb, "sha256:second", originProactive)
+ sb.dispatchJob(&ScanJobEvent{Type: "job", Seq: second, Repository: "repo"})
+
+ if n, _ := sb.activeProactiveJobs(); n != 2 {
+ t.Fatalf("active proactive jobs = %d, want 2", n)
+ }
+ if sb.hasProactiveCapacity() {
+ t.Error("both workers are busy; the hold must stop dispatching rather " +
+ "than pile a backlog into the scanner's own queue")
+ }
+
+ // Both jobs really are out with the scanner, not merely counted.
+ if len(sub.send) != 2 {
+ t.Errorf("scanner received %d jobs, want 2", len(sub.send))
+ }
+ if got := assignedTo(t, sb, first); got != sub.id {
+ t.Errorf("job %d assigned_to = %q, want %q", first, got, sub.id)
+ }
+ if got := assignedTo(t, sb, second); got != sub.id {
+ t.Errorf("job %d assigned_to = %q, want %q", second, got, sub.id)
+ }
+}
+
+// TestScanProactiveCapacity_ScalesAcrossScannerProcesses is the horizontal
+// case. Two single-worker scanners are two units of capacity, and dispatchJob
+// already round-robins, so the two jobs must land on different scanners.
+func TestScanProactiveCapacity_ScalesAcrossScannerProcesses(t *testing.T) {
+ sb := newConcurrencyBroadcaster(t)
+ a := newTestScanSubscriber(t, sb, 4)
+ a.id = "scanner-a"
+ b := newTestScanSubscriber(t, sb, 4)
+ b.id = "scanner-b"
+
+ first := seedJob(t, sb, "sha256:first", originProactive)
+ sb.dispatchJob(&ScanJobEvent{Type: "job", Seq: first, Repository: "repo"})
+
+ if !sb.hasProactiveCapacity() {
+ t.Fatal("a second scanner process is a second unit of capacity and " +
+ "could never receive proactive work under a depth-one gate")
+ }
+
+ second := seedJob(t, sb, "sha256:second", originProactive)
+ sb.dispatchJob(&ScanJobEvent{Type: "job", Seq: second, Repository: "repo"})
+
+ if len(a.send) != 1 || len(b.send) != 1 {
+ t.Errorf("jobs per scanner = a:%d b:%d, want 1 each", len(a.send), len(b.send))
+ }
+ if sb.hasProactiveCapacity() {
+ t.Error("both scanners are busy; dispatch must stop")
+ }
+}
+
+// TestScanProactiveCapacity_IgnoresPushTriggeredJobs covers the second defect
+// in the gate. Push-triggered scans bypass it entirely — oci/xrpc.go calls
+// Enqueue directly — but they were counted by it, so on a hold with steady
+// pushes "one proactive job at a time" was really "none while anyone pushes".
+func TestScanProactiveCapacity_IgnoresPushTriggeredJobs(t *testing.T) {
+ sb := newConcurrencyBroadcaster(t)
+ newTestScanSubscriber(t, sb, 4) // one worker: depth one
+
+ push := seedJob(t, sb, "sha256:pushed", originPush)
+ setStatus(t, sb, push, "processing")
+
+ if n, ok := sb.activeProactiveJobs(); !ok || n != 0 {
+ t.Fatalf("active proactive jobs = %d (ok=%v), want 0: a push-triggered "+
+ "scan is not proactive work", n, ok)
+ }
+ if !sb.hasProactiveCapacity() {
+ t.Error("a push-triggered scan is throttling proactive dispatch; the " +
+ "two paths must not consume each other's budget")
+ }
+
+ // And the mirror: a proactive job in flight fills the proactive budget.
+ proactive := seedJob(t, sb, "sha256:proactive", originProactive)
+ setStatus(t, sb, proactive, "assigned")
+
+ if n, _ := sb.activeProactiveJobs(); n != 1 {
+ t.Errorf("active proactive jobs = %d, want 1", n)
+ }
+ if sb.hasProactiveCapacity() {
+ t.Error("the one worker is busy with a proactive job already")
+ }
+}
+
+// TestScanWaitForProactiveCapacity_BlocksAtTheLimitAndReleasesOnCompletion
+// checks the gate the dispatch loop actually calls: it must park while the
+// budget is full, and wake on the completion signal rather than on a timer.
+func TestScanWaitForProactiveCapacity_BlocksAtTheLimitAndReleasesOnCompletion(t *testing.T) {
+ sb := newConcurrencyBroadcaster(t)
+ newTestScanSubscriber(t, sb, 4) // one worker
+
+ seq := seedJob(t, sb, "sha256:running", originProactive)
+ setStatus(t, sb, seq, "processing")
+
+ got := make(chan bool, 1)
+ go func() { got <- sb.waitForProactiveCapacity() }()
+
+ select {
+ case <-got:
+ t.Fatal("the gate returned while the only worker was busy")
+ case <-time.After(250 * time.Millisecond):
+ }
+
+ setStatus(t, sb, seq, "completed")
+ sb.signalCompletion()
+
+ select {
+ case ok := <-got:
+ if !ok {
+ t.Error("the gate reported no capacity after the job completed")
+ }
+ case <-time.After(5 * time.Second):
+ t.Fatal("the gate did not wake on the completion signal")
+ }
+}
+
+// TestScanWaitForProactiveCapacity_YieldsWhenNoScannerIsConnected keeps the
+// gate safe at zero capacity. Blocking inside it forever would be correct-ish
+// today but leaves the dispatch loop unable to notice a scanner arriving, and
+// unable to re-check anything else; it returns instead so the loop can wait on
+// the connection.
+func TestScanWaitForProactiveCapacity_YieldsWhenNoScannerIsConnected(t *testing.T) {
+ sb := newConcurrencyBroadcaster(t)
+
+ done := make(chan bool, 1)
+ go func() { done <- sb.waitForProactiveCapacity() }()
+
+ select {
+ case ok := <-done:
+ if ok {
+ t.Error("the gate granted capacity with no scanner connected")
+ }
+ case <-time.After(5 * time.Second):
+ t.Fatal("the gate blocked with no scanner connected; the dispatch loop " +
+ "cannot re-check the subscriber list from in there")
+ }
+}
+
+// TestScanWaitForProactiveCapacity_YieldsWhenTheLastScannerDisconnects is the
+// mid-flight version: capacity that vanishes while the gate is parked must
+// release it, not strand the loop.
+func TestScanWaitForProactiveCapacity_YieldsWhenTheLastScannerDisconnects(t *testing.T) {
+ sb := newConcurrencyBroadcaster(t)
+ sub := newTestScanSubscriber(t, sb, 4)
+
+ seq := seedJob(t, sb, "sha256:running", originProactive)
+ setStatus(t, sb, seq, "processing")
+
+ got := make(chan bool, 1)
+ go func() { got <- sb.waitForProactiveCapacity() }()
+
+ select {
+ case <-got:
+ t.Fatal("the gate returned while the only worker was busy")
+ case <-time.After(250 * time.Millisecond):
+ }
+
+ sb.Unsubscribe(sub)
+
+ select {
+ case ok := <-got:
+ if ok {
+ t.Error("the gate granted capacity after the last scanner left")
+ }
+ case <-time.After(10 * time.Second):
+ t.Fatal("the gate stayed parked after the last scanner disconnected")
+ }
+}
+
+// ---------------------------------------------------------------------------
+// F4: the deadline must measure scanning, not queueing
+// ---------------------------------------------------------------------------
+
+// TestScanStarted_StartsTheScanningClock covers the new signal.
+//
+// The ack means "I have it": the scanner sends it off the WebSocket reader the
+// moment a frame arrives, before the job is even queued. The job then waits in
+// the scanner's own 100-deep queue behind its workers. Measuring the scanning
+// deadline from dispatch therefore budgets queueing, and a 100-deep queue of
+// no-op jobs crosses ten minutes at position 59.
+//
+// 'started' is sent by the worker that dequeues the job, so it marks the one
+// moment the hold could not otherwise observe.
+func TestScanStarted_StartsTheScanningClock(t *testing.T) {
+ sb := newConcurrencyBroadcaster(t)
+ sub := newTestScanSubscriber(t, sb, 4)
+
+ seq := seedJob(t, sb, "sha256:queued", originProactive)
+ assignJob(t, sb, seq, sub, 0)
+ sb.handleAck(sub, seq)
+
+ if at := startedAt(t, sb, seq); at.Valid {
+ t.Fatal("the ack started the scanning clock; it means \"I have it\", " +
+ "not \"a worker is on it\"")
+ }
+
+ sb.handleStarted(sub, seq)
+
+ at := startedAt(t, sb, seq)
+ if !at.Valid {
+ t.Fatal("'started' did not stamp started_at, so the deadline still " +
+ "measures queueing rather than scanning")
+ }
+ if d := time.Since(at.Time); d > time.Minute || d < -time.Minute {
+ t.Errorf("started_at is %s away from now, want ~0", d)
+ }
+ if got := jobStatus(t, sb, seq); got != "processing" {
+ t.Errorf("status = %q, want processing", got)
+ }
+}
+
+// TestScanStarted_KeepsAQueuedJobInsideItsScanningDeadline is the behaviour
+// the whole change is for: a job that sat in a scanner's queue for longer than
+// the scanning deadline, and has only just started, must not be cancelled out
+// from under the worker now running it.
+func TestScanStarted_KeepsAQueuedJobInsideItsScanningDeadline(t *testing.T) {
+ sb := newConcurrencyBroadcaster(t)
+ sub := newTestScanSubscriber(t, sb, 4)
+
+ const digest = "sha256:longqueue"
+ seq := seedJob(t, sb, digest, originProactive)
+ assignJob(t, sb, seq, sub, 20) // dispatched twenty minutes ago
+ sb.handleAck(sub, seq)
+ sb.addInflight(digest)
+
+ sb.handleStarted(sub, seq) // a worker picked it up just now
+
+ sb.reDispatchTimedOut()
+
+ if got := jobStatus(t, sb, seq); got != "processing" {
+ t.Fatalf("status = %q, want processing: a scan that started seconds "+
+ "ago was cancelled because it had queued for twenty minutes", got)
+ }
+ if _, _, err := sb.pds.GetScanRecord(context.Background(), digest); err == nil {
+ t.Error("a failure record was written for a scan that is actively running")
+ }
+}
+
+// TestScanProcessingTimeout_FailsAJobStuckSinceItStarted keeps the deadline
+// real. Once a worker has said it started, the scanning budget applies from
+// that moment and a wedged scan is still retired.
+func TestScanProcessingTimeout_FailsAJobStuckSinceItStarted(t *testing.T) {
+ sb := newConcurrencyBroadcaster(t)
+ sub := newTestScanSubscriber(t, sb, 4)
+
+ const digest = "sha256:wedgedscan"
+ seq := seedJob(t, sb, digest, originProactive)
+ assignJob(t, sb, seq, sub, 30)
+ sb.handleAck(sub, seq)
+ sb.addInflight(digest)
+ markStarted(t, sb, seq, 11) // started eleven minutes ago, still nothing
+
+ sb.reDispatchTimedOut()
+
+ if got := jobStatus(t, sb, seq); got != "failed" {
+ t.Fatalf("status = %q, want failed", got)
+ }
+ _, record, err := sb.pds.GetScanRecord(context.Background(), digest)
+ if err != nil {
+ t.Fatalf("no scan record for a job the hold gave up on: %v", err)
+ }
+ if record.Status != atproto.ScanStatusFailed {
+ t.Errorf("record status = %q, want %q", record.Status, atproto.ScanStatusFailed)
+ }
+ if !sb.addInflight(digest) {
+ t.Error("the timed-out digest is still in flight")
+ }
+}
+
+// TestScanProcessingTimeout_ToleratesAScannerThatNeverReportsStarts is the
+// compatibility half. A scanner built before 'started' existed never sends it,
+// so started_at stays NULL and the hold can only observe dispatch. The budget
+// for that case is the one the hold can actually justify: long enough for a
+// full scanner queue to drain, so a healthy backlogged scanner is not killed,
+// and still bounded so a wedged-but-connected scanner does not hold capacity
+// for the life of the process.
+func TestScanProcessingTimeout_ToleratesAScannerThatNeverReportsStarts(t *testing.T) {
+ sb := newConcurrencyBroadcaster(t)
+ sub := newTestScanSubscriber(t, sb, 4)
+
+ const digest = "sha256:oldscanner"
+ seq := seedJob(t, sb, digest, originProactive)
+ assignJob(t, sb, seq, sub, 15) // well past the ten-minute scanning deadline
+ sb.handleAck(sub, seq)
+ sb.addInflight(digest)
+
+ sb.reDispatchTimedOut()
+
+ if got := jobStatus(t, sb, seq); got != "processing" {
+ t.Fatalf("status = %q, want processing: an acked job queued fifteen "+
+ "minutes behind a busy scanner is not evidence of anything wrong", got)
+ }
+
+ // Past the queue budget it is retired like any other stuck job.
+ if _, err := sb.db.Exec(
+ `UPDATE scan_jobs SET assigned_at = ? WHERE seq = ?`,
+ time.Now().Add(-queuedTimeout-time.Minute), seq,
+ ); err != nil {
+ t.Fatalf("age the row: %v", err)
+ }
+
+ sb.reDispatchTimedOut()
+
+ if got := jobStatus(t, sb, seq); got != "failed" {
+ t.Errorf("status = %q, want failed: an unstarted job must still have a "+
+ "bound, or a connected-but-wedged scanner holds capacity forever", got)
+ }
+}
+
+// ---------------------------------------------------------------------------
+// Ownership of terminal transitions
+// ---------------------------------------------------------------------------
+
+// TestScanTerminalHandlers_IgnoreAnotherScannersJob closes the hole that opens
+// the moment horizontal scaling works. handleAck guards on assigned_to; the
+// handlers that write records and retire the row did not, so any scanner could
+// complete, fail or skip a job belonging to another one — writing a scan
+// record for an image it never looked at and releasing a digest a different
+// scanner is still working on.
+func TestScanTerminalHandlers_IgnoreAnotherScannersJob(t *testing.T) {
+ cases := []struct {
+ name string
+ send func(sb *ScanBroadcaster, sub *ScanSubscriber, seq int64)
+ }{
+ {"result", func(sb *ScanBroadcaster, sub *ScanSubscriber, seq int64) {
+ sb.handleResult(sub, ScannerMessage{Type: "result", Seq: seq, SBOM: testSBOM})
+ }},
+ {"error", func(sb *ScanBroadcaster, sub *ScanSubscriber, seq int64) {
+ sb.handleError(sub, ScannerMessage{Type: "error", Seq: seq, Error: "boom"})
+ }},
+ {"skipped", func(sb *ScanBroadcaster, sub *ScanSubscriber, seq int64) {
+ sb.handleSkipped(sub, ScannerMessage{Type: "skipped", Seq: seq, Reason: "nope"})
+ }},
+ {"started", func(sb *ScanBroadcaster, sub *ScanSubscriber, seq int64) {
+ sb.handleStarted(sub, seq)
+ }},
+ }
+
+ for _, tc := range cases {
+ t.Run(tc.name, func(t *testing.T) {
+ sb := newConcurrencyBroadcaster(t)
+ owner := newTestScanSubscriber(t, sb, 4)
+ owner.id = "owner"
+ intruder := newTestScanSubscriber(t, sb, 4)
+ intruder.id = "intruder"
+
+ digest := "sha256:owned" + tc.name
+ seq := seedJob(t, sb, digest, originProactive)
+ assignJob(t, sb, seq, owner, 0)
+ sb.handleAck(owner, seq)
+ sb.addInflight(digest)
+
+ tc.send(sb, intruder, seq)
+
+ if got := jobStatus(t, sb, seq); got != "processing" {
+ t.Errorf("status = %q, want processing: a scanner that does not "+
+ "own the job retired it", got)
+ }
+ if sb.addInflight(digest) {
+ t.Error("another scanner's message released the in-flight digest")
+ }
+ if _, _, err := sb.pds.GetScanRecord(context.Background(), digest); err == nil {
+ t.Error("a scan record was written by a scanner that never had the job")
+ }
+ if tc.name == "started" {
+ if startedAt(t, sb, seq).Valid {
+ t.Error("another scanner restarted the scanning clock")
+ }
+ }
+ })
+ }
+}
+
+// TestScanTerminalHandlers_AcceptTheOwningScanner is the other half: the guard
+// must not reject the scanner that legitimately holds the job.
+func TestScanTerminalHandlers_AcceptTheOwningScanner(t *testing.T) {
+ sb := newConcurrencyBroadcaster(t)
+ owner := newTestScanSubscriber(t, sb, 4)
+ owner.id = "owner"
+
+ const digest = "sha256:legitimate"
+ seq := seedJob(t, sb, digest, originProactive)
+ assignJob(t, sb, seq, owner, 0)
+ sb.handleAck(owner, seq)
+ sb.handleStarted(owner, seq)
+ sb.addInflight(digest)
+
+ sb.handleResult(owner, ScannerMessage{Type: "result", Seq: seq, SBOM: testSBOM})
+
+ if got := jobStatus(t, sb, seq); got != "completed" {
+ t.Fatalf("status = %q, want completed", got)
+ }
+ if !sb.addInflight(digest) {
+ t.Error("the owning scanner's result did not release the digest")
+ }
+}
+
+// ---------------------------------------------------------------------------
+// The buffer-full reset
+// ---------------------------------------------------------------------------
+
+// TestScanDispatchJob_BufferFullResetLeavesAnotherScannersClaimAlone is the
+// mirror of the double-dispatch window closed in drainPendingJobs.
+//
+// dispatchJob assigns the row, finds the scanner's send buffer full, and puts
+// the row back to 'pending' — with no assigned_to or status guard. If another
+// dispatcher claimed the row in between, that reset hands a job a second
+// scanner is already holding back to the pool, and a third scanner gets it.
+//
+// The interleaving is built with a SQLite trigger rather than a race: the
+// trigger fires on dispatchJob's own assign UPDATE and reassigns the row to
+// another scanner, so by the time the buffer-full branch runs, the row is
+// provably not ours. SQLite does not recurse triggers by default, so the reset
+// UPDATE does not re-fire it.
+func TestScanDispatchJob_BufferFullResetLeavesAnotherScannersClaimAlone(t *testing.T) {
+ sb := newConcurrencyBroadcaster(t)
+ sub := newTestScanSubscriber(t, sb, 0) // unbuffered, nobody receiving: always full
+ sub.id = "loser"
+
+ seq := seedJob(t, sb, "sha256:contended", originProactive)
+
+ if _, err := sb.db.Exec(`
+ CREATE TRIGGER steal_assignment AFTER UPDATE OF status ON scan_jobs
+ WHEN NEW.status = 'assigned' AND NEW.assigned_to = 'loser'
+ BEGIN
+ UPDATE scan_jobs SET assigned_to = 'winner' WHERE seq = NEW.seq;
+ END
+ `); err != nil {
+ t.Fatalf("create trigger: %v", err)
+ }
+
+ sb.dispatchJob(&ScanJobEvent{Type: "job", Seq: seq, Repository: "repo"})
+
+ if got := jobStatus(t, sb, seq); got != "assigned" {
+ t.Errorf("status = %q, want assigned: the buffer-full reset returned a "+
+ "row another dispatcher holds to the pool, so a third scanner will "+
+ "be handed a job that is already out", got)
+ }
+ if got := assignedTo(t, sb, seq); got != "winner" {
+ t.Errorf("assigned_to = %q, want \"winner\": the claim was stolen back", got)
+ }
+}
+
+// TestScanDispatchJob_BufferFullReturnsOurOwnRowToPending is the guard's other
+// side: when the row really is ours, a full buffer must still release it so
+// the re-dispatch loop can offer it again rather than leaving it assigned to a
+// scanner that was never sent it.
+func TestScanDispatchJob_BufferFullReturnsOurOwnRowToPending(t *testing.T) {
+ sb := newConcurrencyBroadcaster(t)
+ sub := newTestScanSubscriber(t, sb, 0) // unbuffered, nobody receiving
+ sub.id = "solo"
+
+ seq := seedJob(t, sb, "sha256:nobuffer", originProactive)
+ sb.dispatchJob(&ScanJobEvent{Type: "job", Seq: seq, Repository: "repo"})
+
+ if got := jobStatus(t, sb, seq); got != "pending" {
+ t.Errorf("status = %q, want pending", got)
+ }
+ if got := assignedTo(t, sb, seq); got != "" {
+ t.Errorf("assigned_to = %q, want empty", got)
+ }
+}
+
+// ---------------------------------------------------------------------------
+// Fairness and admission control across N scanner processes
+// ---------------------------------------------------------------------------
+
+// TestScanDispatch_PrefersTheLeastLoadedScanner covers plain round-robin's
+// blind spot. nextIdx hands the next job to whichever subscriber is next in
+// the slice, saturated or idle, which with heterogeneous processes (different
+// worker counts, different hosts, one mid-scan and one just connected) piles
+// work onto a scanner that cannot take it while another sits idle.
+func TestScanDispatch_PrefersTheLeastLoadedScanner(t *testing.T) {
+ sb := newConcurrencyBroadcaster(t)
+ busy := newTestScanSubscriber(t, sb, 4)
+ busy.id = "busy"
+ busy.capacity = 2
+ idle := newTestScanSubscriber(t, sb, 4)
+ idle.id = "idle"
+ idle.capacity = 2
+
+ // "busy" is already running one scan.
+ running := seedJob(t, sb, "sha256:running", originPush)
+ if _, err := sb.db.Exec(
+ `UPDATE scan_jobs SET status='processing', assigned_to=?, assigned_at=? WHERE seq=?`,
+ busy.id, time.Now(), running,
+ ); err != nil {
+ t.Fatalf("prime load: %v", err)
+ }
+
+ seq := seedJob(t, sb, "sha256:next", originProactive)
+ sb.dispatchJob(&ScanJobEvent{Type: "job", Seq: seq, Repository: "repo"})
+
+ if got := assignedTo(t, sb, seq); got != idle.id {
+ t.Errorf("job went to %q, want %q: dispatch must weigh what each "+
+ "scanner is already holding", got, idle.id)
+ }
+ if len(idle.send) != 1 || len(busy.send) != 0 {
+ t.Errorf("sends: idle=%d busy=%d, want 1/0", len(idle.send), len(busy.send))
+ }
+}
+
+// TestScanDispatch_LeavesTheRowPendingWhenEveryScannerIsSaturated is the
+// admission control the deadline depends on.
+//
+// The hold cannot see into a scanner's own queue: the scanner acks on receipt
+// and the job then waits behind its workers, which is exactly why a deadline
+// measured from dispatch cancels healthy work. Keeping the queue on the hold
+// instead of pushing it into the scanner keeps the row observable, keeps it
+// available to whichever process frees up first, and bounds what any one
+// process is holding.
+func TestScanDispatch_LeavesTheRowPendingWhenEveryScannerIsSaturated(t *testing.T) {
+ sb := newConcurrencyBroadcaster(t)
+ sub := newTestScanSubscriber(t, sb, 20)
+ sub.capacity = 1
+
+ first := seedJob(t, sb, "sha256:one", originPush)
+ sb.dispatchJob(&ScanJobEvent{Type: "job", Seq: first, Repository: "repo"})
+ if got := jobStatus(t, sb, first); got != "assigned" {
+ t.Fatalf("first job status = %q, want assigned", got)
+ }
+
+ second := seedJob(t, sb, "sha256:two", originPush)
+ sb.dispatchJob(&ScanJobEvent{Type: "job", Seq: second, Repository: "repo"})
+
+ if got := jobStatus(t, sb, second); got != "pending" {
+ t.Errorf("second job status = %q, want pending: the only worker is "+
+ "busy, so the row belongs on the hold's queue, not the scanner's", got)
+ }
+ if len(sub.send) != 1 {
+ t.Errorf("scanner received %d jobs, want 1", len(sub.send))
+ }
+}
+
+// TestScanDispatch_OffersPendingWorkAsSoonAsCapacityFrees is the other side of
+// admission control: a row held back must not wait for the thirty-second
+// re-dispatch tick, or holding it back would cost more throughput than the
+// scanner-side queue ever did.
+func TestScanDispatch_OffersPendingWorkAsSoonAsCapacityFrees(t *testing.T) {
+ sb := newConcurrencyBroadcaster(t)
+ sub := newTestScanSubscriber(t, sb, 20)
+ sub.capacity = 1
+
+ first := seedJob(t, sb, "sha256:one", originPush)
+ sb.dispatchJob(&ScanJobEvent{Type: "job", Seq: first, Repository: "repo"})
+ second := seedJob(t, sb, "sha256:two", originPush)
+ sb.dispatchJob(&ScanJobEvent{Type: "job", Seq: second, Repository: "repo"})
+ <-sub.send // drop the first job so the assertion below is unambiguous
+
+ setStatus(t, sb, first, "completed")
+ sb.offerPendingJobs()
+
+ if got := jobStatus(t, sb, second); got != "assigned" {
+ t.Errorf("held job status = %q, want assigned once the worker freed up", got)
+ }
+ select {
+ case job := <-sub.send:
+ if job.Seq != second {
+ t.Errorf("scanner received seq %d, want %d", job.Seq, second)
+ }
+ default:
+ t.Error("the freed worker was given nothing")
+ }
+}
+
+// TestScanDrain_StopsAtTheSubscriberCapacity stops the first scanner to
+// connect from swallowing a whole backlog. drainPendingJobs walked every
+// pending row and pushed it at the new subscriber, so with a backlog and two
+// scanner processes the first one took all of it and the second stayed idle —
+// horizontal scaling defeated at the connect path rather than at the gate.
+func TestScanDrain_StopsAtTheSubscriberCapacity(t *testing.T) {
+ sb := newConcurrencyBroadcaster(t)
+ sub := newTestScanSubscriber(t, sb, 20)
+ sub.capacity = 2
+
+ var seqs []int64
+ for i := 0; i < 5; i++ {
+ seqs = append(seqs, seedJob(t, sb, "sha256:backlog", originProactive))
+ }
+
+ sb.drainPendingJobs(sub, 0)
+
+ var assigned, pending int
+ for _, seq := range seqs {
+ switch got := jobStatus(t, sb, seq); got {
+ case "assigned":
+ assigned++
+ case "pending":
+ pending++
+ default:
+ t.Errorf("job %d in unexpected status %q", seq, got)
+ }
+ }
+ if assigned != 2 || pending != 3 {
+ t.Errorf("assigned=%d pending=%d, want 2/3: a two-worker scanner takes "+
+ "two jobs, and the rest stay available to other processes",
+ assigned, pending)
+ }
+ if len(sub.send) != 2 {
+ t.Errorf("scanner received %d jobs, want 2", len(sub.send))
+ }
+}
+
+// ---------------------------------------------------------------------------
+// Disconnect, reconnect, and who owns the work in between
+// ---------------------------------------------------------------------------
+
+// TestScanUnsubscribe_DoesNotImmediatelyHandOffWorkStillRunning is the N-process
+// version of the duplicate-scan problem.
+//
+// Unsubscribe flipped every assigned and processing row belonging to the
+// dropped subscriber straight back to 'pending'. The scanner's worker pool
+// never learns the socket dropped, so it keeps scanning. With one scanner that
+// produced a duplicate against itself. With several it is worse and cannot be
+// deduped anywhere: scanner A blips, B claims A's rows out of drainPendingJobs,
+// and both processes scan the same images and both report a verdict.
+//
+// A disconnect is not evidence that a scanner is gone. It is marked as
+// disconnected and its work is left alone for a grace period instead.
+func TestScanUnsubscribe_DoesNotImmediatelyHandOffWorkStillRunning(t *testing.T) {
+ sb := newConcurrencyBroadcaster(t)
+ first := newTestScanSubscriber(t, sb, 4)
+ first.id = "scanner-a"
+
+ seq := seedJob(t, sb, "sha256:midscan", originProactive)
+ assignJob(t, sb, seq, first, 0)
+ sb.handleAck(first, seq)
+ sb.handleStarted(first, seq)
+
+ sb.Unsubscribe(first)
+
+ if got := jobStatus(t, sb, seq); got != "processing" {
+ t.Fatalf("status = %q, want processing: a five-second blip must not "+
+ "put a running scan back in the pool", got)
+ }
+ if got := assignedTo(t, sb, seq); got != first.id {
+ t.Errorf("assigned_to = %q, want %q", got, first.id)
+ }
+
+ // A second scanner process connecting must not be handed it.
+ second := newTestScanSubscriber(t, sb, 4)
+ second.id = "scanner-b"
+ sb.drainPendingJobs(second, 0)
+
+ select {
+ case job := <-second.send:
+ t.Fatalf("seq %d was handed to a second scanner process while the first "+
+ "is still scanning it", job.Seq)
+ default:
+ }
+}
+
+// TestScanReconnect_LetsAScannerResumeItsOwnWork is why the disconnect is only
+// marked rather than acted on. A scanner keeps one identity for the life of the
+// process and sends it on every connect, so a reconnection inside the grace
+// window reclaims the jobs its workers never stopped running.
+//
+// A scanner that actually restarted comes back with a new identity, so its old
+// rows are not resumed and are reclaimed by the grace timeout instead — which
+// is right, because a restarted process really did lose that work.
+func TestScanReconnect_LetsAScannerResumeItsOwnWork(t *testing.T) {
+ sb := newConcurrencyBroadcaster(t)
+ sub := newTestScanSubscriber(t, sb, 4)
+ sub.id = "instance-1"
+
+ seq := seedJob(t, sb, "sha256:resumed", originProactive)
+ assignJob(t, sb, seq, sub, 0)
+ sb.handleAck(sub, seq)
+ sb.handleStarted(sub, seq)
+ sb.Unsubscribe(sub)
+
+ if !jobIsDisconnected(t, sb, seq) {
+ t.Fatal("the row was not marked as belonging to a disconnected scanner")
+ }
+
+ sb.resumeInstance("instance-1")
+
+ if jobIsDisconnected(t, sb, seq) {
+ t.Error("reconnecting did not clear the disconnect mark, so the job " +
+ "will be reclaimed from a scanner that never stopped running it")
+ }
+ if got := jobStatus(t, sb, seq); got != "processing" {
+ t.Errorf("status = %q, want processing", got)
+ }
+
+ // A different process must not adopt it.
+ sb.resumeInstance("instance-2")
+ if got := assignedTo(t, sb, seq); got != "instance-1" {
+ t.Errorf("assigned_to = %q, want instance-1", got)
+ }
+}
+
+// TestScanDisconnect_ReclaimsWorkOnceTheGraceExpires bounds the wait. A scanner
+// that does not come back has its work returned to the pool, so a genuinely
+// dead process costs one grace period rather than a permanent hole.
+func TestScanDisconnect_ReclaimsWorkOnceTheGraceExpires(t *testing.T) {
+ sb := newConcurrencyBroadcaster(t)
+ gone := newTestScanSubscriber(t, sb, 4)
+ gone.id = "departed"
+
+ seq := seedJob(t, sb, "sha256:abandoned", originProactive)
+ assignJob(t, sb, seq, gone, 0)
+ sb.handleAck(gone, seq)
+ sb.handleStarted(gone, seq)
+ sb.Unsubscribe(gone)
+
+ // Inside the grace: still theirs.
+ sb.reDispatchTimedOut()
+ if got := jobStatus(t, sb, seq); got != "processing" {
+ t.Fatalf("status = %q, want processing inside the grace window", got)
+ }
+
+ if _, err := sb.db.Exec(
+ `UPDATE scan_jobs SET disconnected_at = ? WHERE seq = ?`,
+ time.Now().Add(-reconnectGrace-time.Minute), seq,
+ ); err != nil {
+ t.Fatalf("age the disconnect: %v", err)
+ }
+
+ replacement := newTestScanSubscriber(t, sb, 4)
+ replacement.id = "replacement"
+
+ sb.reDispatchTimedOut()
+
+ if got := assignedTo(t, sb, seq); got != replacement.id {
+ t.Errorf("assigned_to = %q, want %q: a scanner that never came back "+
+ "must not hold its work forever", got, replacement.id)
+ }
+ if len(replacement.send) != 1 {
+ t.Errorf("replacement received %d jobs, want 1", len(replacement.send))
+ }
+}
+
+// TestScanReconcileOnBoot_TreatsInFlightJobsAsDisconnected covers the restart
+// case, which nothing reconciled before.
+//
+// Every connection the previous process held died with it, so a row still
+// 'assigned' or 'processing' at boot belongs to a scanner the hold has no
+// connection to — whether or not that scanner is still running it. Without the
+// mark those rows sat holding dispatch capacity until their own deadlines
+// fired, which is up to an hour for a job that was queued inside a scanner.
+// With it they fall under the same grace as a live disconnect: resumed by the
+// scanner that redials with the same identity, reclaimed otherwise.
+func TestScanReconcileOnBoot_TreatsInFlightJobsAsDisconnected(t *testing.T) {
+ sb := newConcurrencyBroadcaster(t)
+
+ assigned := seedJob(t, sb, "sha256:assigned", originProactive)
+ setStatus(t, sb, assigned, "assigned")
+ processing := seedJob(t, sb, "sha256:processing", originProactive)
+ setStatus(t, sb, processing, "processing")
+ done := seedJob(t, sb, "sha256:done", originProactive)
+ setStatus(t, sb, done, "completed")
+
+ sb.reconcileOnBoot()
+
+ for _, seq := range []int64{assigned, processing} {
+ if !jobIsDisconnected(t, sb, seq) {
+ t.Errorf("job %d was left holding dispatch capacity across a restart", seq)
+ }
+ }
+ if jobIsDisconnected(t, sb, done) {
+ t.Error("a completed job was marked as in flight")
+ }
+
+ // And the mark is what the grace acts on, so a scanner that comes back
+ // with the same identity still keeps its work.
+ if _, err := sb.db.Exec(`UPDATE scan_jobs SET assigned_to = ? WHERE seq = ?`,
+ "instance-1", processing); err != nil {
+ t.Fatalf("set owner: %v", err)
+ }
+ sb.resumeInstance("instance-1")
+ if jobIsDisconnected(t, sb, processing) {
+ t.Error("a scanner that redialed did not get its own job back")
+ }
+}
diff --git a/pkg/hold/pds/scan_broadcaster_stall_test.go b/pkg/hold/pds/scan_broadcaster_stall_test.go
index d590b41..1028126 100644
--- a/pkg/hold/pds/scan_broadcaster_stall_test.go
+++ b/pkg/hold/pds/scan_broadcaster_stall_test.go
@@ -1,6 +1,7 @@
package pds
import (
+ "database/sql"
"testing"
"time"
)
@@ -28,62 +29,85 @@ func jobStatus(t *testing.T, sb *ScanBroadcaster, seq int64) string {
return status
}
-// TestScanHasActiveJobs_IgnoresLongPendingJob is the regression test for the
-// nine-day deployment-wide scanning outage. A single job sat in 'pending' with
-// nothing left to dispatch it, hasActiveJobs() counted it forever, and
-// waitForCapacity() therefore never let the proactive dispatch loop enqueue
-// another job — so discovery kept finding unscanned images and creating none.
-func TestScanHasActiveJobs_IgnoresLongPendingJob(t *testing.T) {
- sb := newTestScanBroadcaster(t)
- seedPendingJobs(t, sb, 1)
+// activeProactive is the count half of activeProactiveJobs, failing the test if
+// the answer is not usable.
+func activeProactive(t *testing.T, sb *ScanBroadcaster) int {
+ t.Helper()
- if !sb.hasActiveJobs() {
+ n, ok := sb.activeProactiveJobs()
+ if !ok {
+ t.Fatal("activeProactiveJobs reported an unusable answer")
+ }
+ return n
+}
+
+// TestScanActiveProactiveJobs_IgnoresLongPendingJob is the regression test for
+// the nine-day deployment-wide scanning outage. A single job sat in 'pending'
+// with nothing left to dispatch it, the capacity check counted it forever, and
+// the dispatch gate therefore never let the proactive loop enqueue another job
+// — so discovery kept finding unscanned images and creating none.
+func TestScanActiveProactiveJobs_IgnoresLongPendingJob(t *testing.T) {
+ sb := newTestScanBroadcaster(t)
+ seq := seedJob(t, sb, "sha256:pendingforever", originProactive)
+
+ if activeProactive(t, sb) != 1 {
t.Fatal("a freshly enqueued pending job must count as active")
}
- backdateJob(t, sb, 1, 60)
+ backdateJob(t, sb, seq, 60)
- if sb.hasActiveJobs() {
- t.Error("a job pending for an hour must not block dispatch capacity")
+ if n := activeProactive(t, sb); n != 0 {
+ t.Errorf("active = %d, want 0: a job pending for an hour must not block "+
+ "dispatch capacity", n)
}
}
-// TestScanHasActiveJobs_CountsAssignedAndProcessing guards the other half:
-// assigned and processing jobs have their own reclaim timeouts, so they must
-// still hold capacity no matter how old the row is.
-func TestScanHasActiveJobs_CountsAssignedAndProcessing(t *testing.T) {
+// TestScanActiveProactiveJobs_CountsAssignedAndProcessing guards the other
+// half: assigned and processing jobs have their own reclaim timeouts, so they
+// must still hold capacity no matter how old the row is.
+func TestScanActiveProactiveJobs_CountsAssignedAndProcessing(t *testing.T) {
for _, status := range []string{"assigned", "processing"} {
t.Run(status, func(t *testing.T) {
sb := newTestScanBroadcaster(t)
- seedPendingJobs(t, sb, 1)
- backdateJob(t, sb, 1, 60)
+ seq := seedJob(t, sb, "sha256:inflight", originProactive)
+ backdateJob(t, sb, seq, 60)
+ setStatus(t, sb, seq, status)
- if _, err := sb.db.Exec(`UPDATE scan_jobs SET status = ? WHERE seq = 1`, status); err != nil {
- t.Fatalf("set status: %v", err)
- }
-
- if !sb.hasActiveJobs() {
+ if activeProactive(t, sb) != 1 {
t.Errorf("%s job must hold dispatch capacity", status)
}
})
}
}
-// TestScanHasActiveJobs_IgnoresTerminalJobs keeps completed and failed rows out
-// of the capacity check — the table holds tens of thousands of them.
-func TestScanHasActiveJobs_IgnoresTerminalJobs(t *testing.T) {
+// TestScanActiveProactiveJobs_IgnoresTerminalJobs keeps completed and failed
+// rows out of the capacity check — the table holds tens of thousands of them.
+func TestScanActiveProactiveJobs_IgnoresTerminalJobs(t *testing.T) {
sb := newTestScanBroadcaster(t)
- seedPendingJobs(t, sb, 2)
+ setStatus(t, sb, seedJob(t, sb, "sha256:done", originProactive), "completed")
+ setStatus(t, sb, seedJob(t, sb, "sha256:dead", originProactive), "failed")
- if _, err := sb.db.Exec(`UPDATE scan_jobs SET status = 'completed' WHERE seq = 1`); err != nil {
- t.Fatalf("complete: %v", err)
- }
- if _, err := sb.db.Exec(`UPDATE scan_jobs SET status = 'failed' WHERE seq = 2`); err != nil {
- t.Fatalf("fail: %v", err)
+ if n := activeProactive(t, sb); n != 0 {
+ t.Errorf("active = %d, want 0: terminal jobs must not hold capacity", n)
}
+}
- if sb.hasActiveJobs() {
- t.Error("terminal jobs must not hold dispatch capacity")
+// TestScanActiveProactiveJobs_IgnoresPushTriggeredWork is the second half of
+// the gate's original defect, alongside its hardcoded depth of one. Push scans
+// never pass through this gate — a user who just pushed is waiting for the
+// answer, so oci/xrpc.go enqueues directly — but they were counted by it, so
+// on a hold with steady pushes the proactive loop had capacity approximately
+// never.
+func TestScanActiveProactiveJobs_IgnoresPushTriggeredWork(t *testing.T) {
+ sb := newTestScanBroadcaster(t)
+ setStatus(t, sb, seedJob(t, sb, "sha256:pushed", originPush), "processing")
+ // Rows written before the column existed read as push for the same reason:
+ // unknown provenance must not throttle proactive dispatch.
+ seedPendingJobs(t, sb, 1)
+
+ if n := activeProactive(t, sb); n != 0 {
+ t.Errorf("active = %d, want 0: push-triggered work must not spend the "+
+ "proactive budget", n)
}
}
@@ -165,3 +189,140 @@ func TestScanDispatchJob_SkipsClaimedJob(t *testing.T) {
t.Errorf("assignment stolen from the first dispatcher, assigned_to=%q", assignedTo)
}
}
+
+// TestScanDrainPendingJobs_SkipsClaimedJob is the sibling of
+// TestScanDispatchJob_SkipsClaimedJob for the other dispatcher.
+//
+// drainPendingJobs collects every pending row up front, then assigns and sends
+// them one at a time, so a row can be claimed by dispatchJob or the re-dispatch
+// loop in between. Its UPDATE carries the same `AND status = 'pending'` guard,
+// but the result was never read: the job was pushed onto this scanner's queue
+// whether or not this scanner had won it, and both scanners then scanned the
+// same image. The losing scanner's ack is silently dropped by handleAck (which
+// does guard on assigned_to), so nothing downstream notices either.
+//
+// The interleaving is built deterministically: an unbuffered send channel
+// parks the drain inside the send for job 1, which is after job 1's UPDATE and
+// before job 2's, and job 2 is claimed from under it there.
+func TestScanDrainPendingJobs_SkipsClaimedJob(t *testing.T) {
+ sb := newTestScanBroadcaster(t)
+ sub := newTestScanSubscriber(t, sb, 0) // unbuffered: the drain parks in the send
+ sub.capacity = 2 // enough room that the drain would reach job 2
+ seedPendingJobs(t, sb, 2)
+
+ done := make(chan struct{})
+ go func() {
+ defer close(done)
+ sb.drainPendingJobs(sub, 0)
+ }()
+
+ // Wait until job 1 is assigned, which means the drain is now parked in the
+ // send for it and has not yet looked at job 2.
+ deadline := time.Now().Add(10 * time.Second)
+ for jobStatus(t, sb, 1) != "assigned" {
+ if time.Now().After(deadline) {
+ t.Fatal("drain never assigned the first job")
+ }
+ time.Sleep(time.Millisecond)
+ }
+
+ // Another dispatcher claims job 2 while the drain is stuck on job 1.
+ if _, err := sb.db.Exec(
+ `UPDATE scan_jobs SET status='assigned', assigned_to='other' WHERE seq = 2 AND status = 'pending'`,
+ ); err != nil {
+ t.Fatalf("claim: %v", err)
+ }
+
+ // Release the drain. Job 1 is legitimately ours.
+ select {
+ case job := <-sub.send:
+ if job.Seq != 1 {
+ t.Fatalf("first send was seq %d, want 1", job.Seq)
+ }
+ case <-time.After(10 * time.Second):
+ t.Fatal("drain never sent the first job")
+ }
+
+ // Job 2 must not follow it. The receive is what makes this decisive: a
+ // drain that sends anyway is parked in that send right now, and would
+ // otherwise give up after its own timeout and look identical to a drain
+ // that correctly skipped it.
+ select {
+ case job := <-sub.send:
+ t.Fatalf("job %d was sent to a scanner that did not claim it; both "+
+ "scanners now scan the same image", job.Seq)
+ case <-time.After(time.Second):
+ }
+
+ <-done
+
+ var assignedTo string
+ if err := sb.db.QueryRow(`SELECT assigned_to FROM scan_jobs WHERE seq = 2`).Scan(&assignedTo); err != nil {
+ t.Fatalf("query assigned_to: %v", err)
+ }
+ if assignedTo != "other" {
+ t.Errorf("assignment stolen from the first dispatcher, assigned_to=%q", assignedTo)
+ }
+}
+
+// TestScanHasActiveJobs_FailsOpenAfterPersistentDBErrors covers the last way
+// proactive scanning halts for the life of the process with no recovery.
+//
+// hasActiveJobs returned true on any query error ("assume busy") and
+// waitForCapacity spins on it, so a persistently failing query — a locked
+// database, a handle closed under a shared connection — stopped proactive
+// dispatch entirely. logStalledCapacity could not report it either: it runs
+// the same database and takes its own error branch.
+//
+// A transient error should still be treated as busy, since guessing "idle"
+// piles work onto a scanner that may already have some. A persistent one must
+// not: after a small budget of consecutive failures the check fails open, and
+// says so distinctly in the log.
+func TestScanHasActiveJobs_FailsOpenAfterPersistentDBErrors(t *testing.T) {
+ dbPath := "file:" + t.TempDir() + "/scan.db"
+ sb := newTestScanBroadcaster(t)
+ seedPendingJobs(t, sb, 1)
+
+ // Closing the handle is the cheapest persistent query failure there is.
+ if err := sb.db.Close(); err != nil {
+ t.Fatalf("close db: %v", err)
+ }
+
+ if _, ok := sb.activeProactiveJobs(); ok {
+ t.Error("the first query failure must be treated as busy; a blip is " +
+ "no reason to pile another job on a scanner")
+ }
+ for i := 1; i < activeJobsErrorBudget; i++ {
+ if _, ok := sb.activeProactiveJobs(); ok {
+ t.Errorf("failure %d is still within the budget and must read as busy", i+1)
+ }
+ }
+ if _, ok := sb.activeProactiveJobs(); !ok {
+ t.Fatalf("the capacity check still reports busy after %d consecutive "+
+ "database errors: proactive dispatch is halted for the life of the "+
+ "process with no recovery", activeJobsErrorBudget+1)
+ }
+
+ // A working database restores the budget, so a later blip is absorbed
+ // rather than landing on an already-exhausted counter.
+ db, err := sql.Open("libsql", dbPath)
+ if err != nil {
+ t.Fatalf("reopen db: %v", err)
+ }
+ t.Cleanup(func() { _ = db.Close() })
+ sb.db = db
+ if err := sb.initSchema(); err != nil {
+ t.Fatalf("initSchema: %v", err)
+ }
+ seedJob(t, sb, "sha256:fresh", originProactive)
+
+ if activeProactive(t, sb) != 1 {
+ t.Fatal("a fresh pending job must count as active")
+ }
+ if err := sb.db.Close(); err != nil {
+ t.Fatalf("close db again: %v", err)
+ }
+ if _, ok := sb.activeProactiveJobs(); ok {
+ t.Error("the error budget was not reset by a successful query")
+ }
+}
diff --git a/pkg/hold/pds/scan_broadcaster_stuck_test.go b/pkg/hold/pds/scan_broadcaster_stuck_test.go
new file mode 100644
index 0000000..1b62848
--- /dev/null
+++ b/pkg/hold/pds/scan_broadcaster_stuck_test.go
@@ -0,0 +1,616 @@
+package pds
+
+import (
+ "context"
+ "testing"
+ "time"
+
+ "atcr.io/pkg/atproto"
+)
+
+// These tests cover the ways a scan job wedges on the hold side: the clock the
+// processing deadline is measured from, what happens to a job that blows it,
+// and what the hold does with a job whose scanner vanished mid-scan.
+//
+// They are characterization tests. Where the behaviour they pin is wrong, the
+// comment says so and the assertion still describes what the code does today,
+// so the suite stays green until a fix lands. A test that can only pass after a
+// fix is marked with t.Skip and names the bug.
+//
+// The scanner-side half of the disconnect story lives in
+// scanner/internal/e2e/stuck_test.go; where a scenario here encodes an
+// assumption about what the scanner does with a re-offered job, that file pins
+// it against the real scanner.
+
+// newStuckBroadcaster is the broadcaster these scenarios need: the in-flight
+// digest set and the ack timeout the constructor fills in but the bare helper
+// leaves zeroed, plus the PDS and S3 stand-in behind every terminal
+// transition. The processing timeout writes a scan record now, so a
+// broadcaster with a nil pds is no longer a usable stand-in for one.
+func newStuckBroadcaster(t *testing.T) *ScanBroadcaster {
+ t.Helper()
+
+ sb := newRecordingScanBroadcaster(t)
+ // A buffered signal so a test can assert dispatch capacity was actually
+ // released; production wires the same channel to the dispatch loop.
+ sb.completionSignal = make(chan struct{}, 1)
+ return sb
+}
+
+// seedJobWithDigest inserts one pending job carrying a specific manifest
+// digest. seedPendingJobs gives every row the same digest, which is fine for
+// status bookkeeping but useless for anything that keys on the digest.
+func seedJobWithDigest(t *testing.T, sb *ScanBroadcaster, digest string) int64 {
+ t.Helper()
+
+ res, err := sb.db.Exec(`
+ INSERT INTO scan_jobs
+ (manifest_digest, repository, tag, user_did, user_handle,
+ hold_did, hold_endpoint, tier, config_json, layers_json, status)
+ VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, ?, 'pending')
+ `, digest, "repo", "latest", "did:plc:user", "user.example.com",
+ sb.holdDID, sb.holdEndpoint, "deckhand", "{}", "[]")
+ if err != nil {
+ t.Fatalf("seed job: %v", err)
+ }
+ seq, err := res.LastInsertId()
+ if err != nil {
+ t.Fatalf("seq: %v", err)
+ }
+ return seq
+}
+
+// assignJob puts a row in 'assigned' as dispatchJob would, with assigned_at set
+// the given number of minutes in the past.
+func assignJob(t *testing.T, sb *ScanBroadcaster, seq int64, sub *ScanSubscriber, agoMinutes int) {
+ t.Helper()
+
+ at := time.Now().Add(-time.Duration(agoMinutes) * time.Minute)
+ _, err := sb.db.Exec(`
+ UPDATE scan_jobs SET status = 'assigned', assigned_to = ?, assigned_at = ?
+ WHERE seq = ?
+ `, sub.id, at, seq)
+ if err != nil {
+ t.Fatalf("assign job %d: %v", seq, err)
+ }
+}
+
+func assignedAt(t *testing.T, sb *ScanBroadcaster, seq int64) time.Time {
+ t.Helper()
+
+ var at time.Time
+ if err := sb.db.QueryRow(`SELECT assigned_at FROM scan_jobs WHERE seq = ?`, seq).Scan(&at); err != nil {
+ t.Fatalf("query assigned_at for %d: %v", seq, err)
+ }
+ return at
+}
+
+// TestScanAck_DoesNotStartTheScanningClock is the hold-side half of the ack
+// timing mismatch.
+//
+// The scanner acks the moment a job comes off the WebSocket, before it is even
+// queued (client/hold.go handleFrame sends the ack, then Enqueue). handleAck
+// moves the row 'assigned' → 'processing' and touches nothing else, which is
+// correct — the ack means "I have it", not "a worker is on it".
+//
+// What was wrong was the deadline. The ten-minute processing budget was
+// measured from assigned_at, so it covered however long the job spent queued
+// inside the scanner: a 100-deep queue of no-op jobs drains in 16m40s and
+// crosses the deadline at position 59, and with 16-second scans at position
+// 23. Every job past that point was failed underneath a scanner that was
+// working perfectly, and since the timeout now writes a scan record, each one
+// is a "scan failed" the user can see.
+//
+// The scanning deadline is measured from 'started' instead, and a job that has
+// only been acked falls under the much larger queue budget.
+func TestScanAck_DoesNotStartTheScanningClock(t *testing.T) {
+ sb := newStuckBroadcaster(t)
+ sub := newTestScanSubscriber(t, sb, 4)
+
+ seq := seedJobWithDigest(t, sb, "sha256:aaaa")
+ assignJob(t, sb, seq, sub, 11)
+ before := assignedAt(t, sb, seq)
+
+ sb.handleAck(sub, seq)
+
+ if got := jobStatus(t, sb, seq); got != "processing" {
+ t.Fatalf("after ack: status = %q, want processing", got)
+ }
+ if after := assignedAt(t, sb, seq); !after.Equal(before) {
+ t.Errorf("ack moved assigned_at from %s to %s; it is the dispatch "+
+ "timestamp and the queue budget is measured from it", before, after)
+ }
+ if startedAt(t, sb, seq).Valid {
+ t.Error("the ack stamped started_at; only a worker picking the job up does that")
+ }
+
+ // The scanner acked eleven minutes after dispatch and is still holding the
+ // job. Nothing here is evidence of a problem.
+ sb.reDispatchTimedOut()
+
+ if got := jobStatus(t, sb, seq); got != "processing" {
+ t.Fatalf("status = %q, want processing: a job acked and queued for "+
+ "eleven minutes was cancelled underneath a working scanner", got)
+ }
+}
+
+// TestScanProcessingTimeout_RetiresTheJobProperly is the compounding half of
+// the stuck-scan story, and the reason a single hung scan used to remove an
+// image from scanning until the hold process restarted.
+//
+// Every proactive enqueue path adds the manifest digest to sb.inflight and
+// relies on a terminal transition to take it out again. The ten-minute
+// processing sweep in reDispatchTimedOut was the one terminal transition that
+// did neither of the two things all the others do: it wrote no scan record
+// and never called removeInflight. So the digest stayed in the set —
+// discoverUnscannedForUser and runStalePass both skip any manifest whose
+// addInflight returns false — and nothing in the system recorded that the
+// manifest had ever been attempted, which the appview renders as a grey "Not
+// scanned" indefinitely.
+//
+// The sweep is now a terminal transition like handleError: a failed scan
+// record with a reason, the digest released, and dispatch capacity signalled.
+// "Failed" rather than "skipped" is deliberate — a hung scanner is a
+// transient condition, so the stale loop should retry it on the rescan
+// interval.
+func TestScanProcessingTimeout_RetiresTheJobProperly(t *testing.T) {
+ sb := newStuckBroadcaster(t)
+ sub := newTestScanSubscriber(t, sb, 4)
+
+ const digest = "sha256:hungscan"
+ seq := seedJobWithDigest(t, sb, digest)
+ assignJob(t, sb, seq, sub, 11)
+ sb.handleAck(sub, seq)
+ markStarted(t, sb, seq, 11) // a worker took it and then went quiet
+
+ if !sb.addInflight(digest) {
+ // The enqueue paths do this; do it here so the state matches.
+ t.Fatal("digest was already in flight before the test started")
+ }
+
+ sb.reDispatchTimedOut()
+
+ if got := jobStatus(t, sb, seq); got != "failed" {
+ t.Fatalf("status = %q, want failed", got)
+ }
+
+ if !sb.addInflight(digest) {
+ t.Error("the timed-out digest is still in flight: discovery and the " +
+ "stale loop will skip this manifest for the life of the process")
+ }
+ sb.removeInflight(digest)
+
+ _, record, err := sb.pds.GetScanRecord(context.Background(), digest)
+ if err != nil {
+ t.Fatalf("no scan record was written for a job the hold gave up on, so "+
+ "the appview cannot tell it apart from one enqueued a minute ago: %v", err)
+ }
+ if record.Status != atproto.ScanStatusFailed {
+ t.Errorf("record status = %q, want %q", record.Status, atproto.ScanStatusFailed)
+ }
+ if record.Reason == "" {
+ t.Error("failed record carries no reason; it is the only thing a user sees")
+ }
+
+ select {
+ case <-sb.completionSignal:
+ default:
+ t.Error("no completion signal: the dispatch loop sleeps up to 5s longer than it needs to")
+ }
+
+ // And nothing re-offers the row to the scanner that is still connected.
+ select {
+ case job := <-sub.send:
+ t.Fatalf("timed-out processing job %d was re-dispatched", job.Seq)
+ default:
+ }
+}
+
+// TestScanProcessingTimeout_LeavesAJobInsideTheDeadlineAlone is the guard on
+// the other side of the sweep: it must only touch rows that have actually
+// blown the ten minutes, and it must not write records for jobs a scanner is
+// still legitimately working on.
+func TestScanProcessingTimeout_LeavesAJobInsideTheDeadlineAlone(t *testing.T) {
+ sb := newStuckBroadcaster(t)
+ sub := newTestScanSubscriber(t, sb, 4)
+
+ const digest = "sha256:stillworking"
+ seq := seedJobWithDigest(t, sb, digest)
+ assignJob(t, sb, seq, sub, 2)
+ sb.handleAck(sub, seq)
+ markStarted(t, sb, seq, 2)
+ sb.addInflight(digest)
+
+ sb.reDispatchTimedOut()
+
+ if got := jobStatus(t, sb, seq); got != "processing" {
+ t.Errorf("status = %q, want processing", got)
+ }
+ if sb.addInflight(digest) {
+ t.Error("digest was released while the scan is still inside its deadline")
+ }
+ if _, _, err := sb.pds.GetScanRecord(context.Background(), digest); err == nil {
+ t.Error("a failure record was written for a scan that is still running")
+ }
+}
+
+// TestScanProcessingTimeout_ReleasesCapacityWhileTheScannerIsStillWedged is
+// what turns one hung scan into a slow leak rather than a single lost job.
+//
+// hasActiveJobs counts 'processing' rows with no age bound, so a wedged scan
+// holds the proactive dispatch loop still — until the ten-minute timeout marks
+// it 'failed', at which point capacity is free again and dispatchLoop picks the
+// next candidate and hands it to the same scanner, whose only worker is still
+// stuck on the first one. Repeat every ten minutes: each new job is acked,
+// queued behind the wedge, failed by the timeout, and leaks its digest out of
+// the in-flight set for good.
+//
+// Before dfd604b the same wedge froze dispatch outright, which is the outage
+// that commit was written for. It bounded 'pending' but left 'processing'
+// alone, so the freeze became this drip instead.
+func TestScanProcessingTimeout_ReleasesCapacityWhileTheScannerIsStillWedged(t *testing.T) {
+ sb := newStuckBroadcaster(t)
+ sub := newTestScanSubscriber(t, sb, 4)
+
+ seq := seedJob(t, sb, "sha256:wedged", originProactive)
+ assignJob(t, sb, seq, sub, 11)
+ sb.handleAck(sub, seq)
+ markStarted(t, sb, seq, 11)
+
+ if n := activeProactive(t, sb); n != 1 {
+ t.Fatalf("active proactive jobs = %d, want 1: a processing job must "+
+ "hold dispatch capacity", n)
+ }
+
+ sb.reDispatchTimedOut()
+
+ if activeProactive(t, sb) != 0 {
+ t.Fatal("capacity is still held after the processing timeout; the " +
+ "drip this test describes cannot happen, update it")
+ }
+ t.Logf("job %d is failed and capacity is free, but the scanner that never "+
+ "answered for it is unchanged: the next candidate goes to the same "+
+ "wedged worker. The digest is released and a failed record written, so "+
+ "the manifest is at least retried on the rescan interval", seq)
+}
+
+// TestScanResult_ArrivingAfterTheTimeoutHealsTheRow bounds the previous test:
+// the leak is permanent only when the scanner never speaks for that seq again.
+// handleResult has no status guard, so a result that arrives after the deadline
+// flips 'failed' back to 'completed' and releases the digest.
+//
+// That is what makes the leak hard to see in production. It bites exactly in
+// the cases where the scanner is wedged (an unbounded Syft extraction) or where
+// its terminal message was written to a dead socket and dropped
+// (client/hold.go sendJSON) — the same cases where scanning was already stuck.
+func TestScanResult_ArrivingAfterTheTimeoutHealsTheRow(t *testing.T) {
+ sb := newRecordingScanBroadcaster(t)
+ sub := newTestScanSubscriber(t, sb, 4)
+
+ const digest = "sha256:lateresult"
+ seq := seedJobWithDigest(t, sb, digest)
+ assignJob(t, sb, seq, sub, 11)
+ sb.handleAck(sub, seq)
+ markStarted(t, sb, seq, 11)
+ sb.addInflight(digest)
+
+ sb.reDispatchTimedOut()
+ if got := jobStatus(t, sb, seq); got != "failed" {
+ t.Fatalf("status = %q, want failed", got)
+ }
+
+ // handleResult writes a scan record, so this needs the broadcaster that
+ // has a PDS behind it. The row bookkeeping under test happens after that
+ // write, which is exactly why the write must not be able to panic.
+ sb.handleResult(sub, ScannerMessage{Type: "result", Seq: seq, SBOM: testSBOM})
+
+ if got := jobStatus(t, sb, seq); got != "completed" {
+ t.Errorf("status = %q, want completed: a late result overwrites the "+
+ "failure with no status guard", got)
+ }
+ if !sb.addInflight(digest) {
+ t.Error("late result did not release the in-flight digest")
+ }
+}
+
+// TestScanHandleResult_SurvivesAResultWithoutSummary covers what used to be a
+// hold crash reachable from any scanner running with vuln.enabled=false.
+//
+// The scanner only fills ScanResult.Summary when Grype ran (worker.go
+// processJob, step 3), and SendResult copies it straight through, so with
+// vulnerability scanning disabled every successful scan sends a result with no
+// summary. handleResult guarded the record-writing branch with
+// `if msg.Summary != nil` and then dereferenced msg.Summary unguarded in its
+// final log line. That panic was not in an HTTP handler; it was in the
+// subscriber's reader goroutine, so it took the whole hold process down, and
+// the job was re-dispatched on restart into the same crash.
+//
+// Two things have to hold now. The obvious one is that nothing panics. The
+// less obvious one is that the SBOM still lands in a scan record: the blob was
+// uploaded to S3 before the record write, so skipping the record (as the old
+// `if msg.Summary != nil` guard did) leaves that blob orphaned with nothing
+// referencing it.
+//
+// A nil summary means "not scanned for vulnerabilities", which is NOT the same
+// as "scanned, found zero". The record therefore carries no vulnerability
+// report blob, and the appview keys off that to avoid claiming the image is
+// clean (see classifyScanRecord in pkg/appview/handlers/scan_result.go).
+func TestScanHandleResult_SurvivesAResultWithoutSummary(t *testing.T) {
+ sb := newRecordingScanBroadcaster(t)
+ sub := newTestScanSubscriber(t, sb, 4)
+
+ const digest = "sha256:nosummary"
+ seq := seedJobWithDigest(t, sb, digest)
+ assignJob(t, sb, seq, sub, 0)
+ sb.handleAck(sub, seq)
+
+ func() {
+ defer func() {
+ if r := recover(); r != nil {
+ t.Fatalf("handleResult panicked on a summary-less result: %v", r)
+ }
+ }()
+ // Exactly the message a vuln-disabled scanner sends: an SBOM, no
+ // summary, no vulnerability report.
+ sb.handleResult(sub, ScannerMessage{Type: "result", Seq: seq, SBOM: testSBOM})
+ }()
+
+ if got := jobStatus(t, sb, seq); got != "completed" {
+ t.Errorf("status = %q, want completed", got)
+ }
+
+ _, record, err := sb.pds.GetScanRecord(context.Background(), digest)
+ if err != nil {
+ t.Fatalf("no scan record was written for a summary-less result, so the "+
+ "SBOM blob already in S3 is orphaned: %v", err)
+ }
+ if record.SbomBlob == nil {
+ t.Error("scan record carries no SBOM blob; the uploaded blob is orphaned")
+ }
+ if record.VulnReportBlob != nil {
+ t.Error("scan record carries a vulnerability report that was never produced")
+ }
+ if record.Status != atproto.ScanStatusOK {
+ t.Errorf("status = %q, want %q: the scan itself succeeded", record.Status, atproto.ScanStatusOK)
+ }
+ if record.Total != 0 || record.Critical != 0 || record.High != 0 || record.Medium != 0 || record.Low != 0 {
+ t.Errorf("counts = %d/%d/%d/%d total %d, want all zero: Grype never ran",
+ record.Critical, record.High, record.Medium, record.Low, record.Total)
+ }
+}
+
+// TestScanHandleResult_RecordsSummaryCountsWhenGrypeRan is the other half of
+// the pair: with a summary present the counts must reach the record, so the
+// nil-tolerant path above cannot be satisfied by dropping them everywhere.
+func TestScanHandleResult_RecordsSummaryCountsWhenGrypeRan(t *testing.T) {
+ sb := newRecordingScanBroadcaster(t)
+ sub := newTestScanSubscriber(t, sb, 4)
+
+ const digest = "sha256:withsummary"
+ seq := seedJobWithDigest(t, sb, digest)
+ assignJob(t, sb, seq, sub, 0)
+ sb.handleAck(sub, seq)
+
+ sb.handleResult(sub, ScannerMessage{
+ Type: "result",
+ Seq: seq,
+ SBOM: testSBOM,
+ VulnReport: `{"matches":[]}`,
+ Summary: &VulnerabilitySummary{Critical: 1, High: 2, Medium: 3, Low: 4, Total: 10},
+ })
+
+ _, record, err := sb.pds.GetScanRecord(context.Background(), digest)
+ if err != nil {
+ t.Fatalf("get scan record: %v", err)
+ }
+ if record.Critical != 1 || record.High != 2 || record.Medium != 3 || record.Low != 4 || record.Total != 10 {
+ t.Errorf("counts = %d/%d/%d/%d total %d, want 1/2/3/4 total 10",
+ record.Critical, record.High, record.Medium, record.Low, record.Total)
+ }
+ if record.SbomBlob == nil || record.VulnReportBlob == nil {
+ t.Errorf("blobs = sbom:%v vuln:%v, want both", record.SbomBlob != nil, record.VulnReportBlob != nil)
+ }
+}
+
+// testSBOM is a stand-in for the SPDX document the scanner sends. Only its
+// bytes matter here: the hold hashes them into a blob CID and never parses it.
+const testSBOM = `{"spdxVersion":"SPDX-2.3","packages":[]}`
+
+// TestScanUnsubscribe_HoldsAJobForAReconnectingScanner is the hold-side
+// assumption behind the duplicate-scan scenario in
+// scanner/internal/e2e/stuck_test.go, updated for a hold that expects several
+// scanner processes.
+//
+// Unsubscribe used to flip every assigned and processing row belonging to the
+// dropped subscriber back to 'pending'. Nothing tells the scanner: its worker
+// pool still has the job queued, or a worker is halfway through downloading
+// blobs for it. drainPendingJobs then handed the same seq straight to whoever
+// connected next. Against a single scanner that was a duplicate against
+// itself, which a local dedupe could in principle catch. Against N processes
+// nothing can catch it, because the two copies are in different processes.
+//
+// A disconnect is now recorded rather than acted on, and a scanner keeps one
+// identity for the life of its process, so the reconnecting scanner gets its
+// own work back and a different process is never offered it.
+func TestScanUnsubscribe_HoldsAJobForAReconnectingScanner(t *testing.T) {
+ sb := newStuckBroadcaster(t)
+ first := newTestScanSubscriber(t, sb, 4)
+ first.id = "instance-1"
+
+ seq := seedJobWithDigest(t, sb, "sha256:duplicated")
+ assignJob(t, sb, seq, first, 0)
+ sb.handleAck(first, seq)
+ sb.handleStarted(first, seq)
+ if got := jobStatus(t, sb, seq); got != "processing" {
+ t.Fatalf("status = %q, want processing", got)
+ }
+
+ sb.Unsubscribe(first)
+
+ if got := jobStatus(t, sb, seq); got != "processing" {
+ t.Fatalf("after disconnect: status = %q, want processing", got)
+ }
+
+ // A different process connecting must not be offered it.
+ other := newTestScanSubscriber(t, sb, 4)
+ other.id = "instance-2"
+ sb.drainPendingJobs(other, 0)
+
+ select {
+ case job := <-other.send:
+ t.Fatalf("seq %d was handed to a second scanner process while the first "+
+ "is still scanning it", job.Seq)
+ default:
+ }
+
+ // The original process redialing resumes it. Nothing is re-sent: the
+ // scanner never lost the job, only the socket.
+ sb.resumeInstance("instance-1")
+ if jobIsDisconnected(t, sb, seq) {
+ t.Error("the reconnecting scanner did not get its own job back")
+ }
+ if got := jobStatus(t, sb, seq); got != "processing" {
+ t.Errorf("status = %q, want processing", got)
+ }
+}
+
+// TestScanUnsubscribe_DoesNotReleaseTheInFlightDigest is a smaller sibling of
+// the timeout leak: a disconnect leaves the digest in flight. That is harmless
+// — the job is still the disconnected scanner's, and will either be resumed or
+// reclaimed — but it means the set cannot be read as "jobs a scanner is
+// working on".
+func TestScanUnsubscribe_DoesNotReleaseTheInFlightDigest(t *testing.T) {
+ sb := newStuckBroadcaster(t)
+ sub := newTestScanSubscriber(t, sb, 4)
+
+ const digest = "sha256:disconnected"
+ seq := seedJobWithDigest(t, sb, digest)
+ assignJob(t, sb, seq, sub, 0)
+ sb.addInflight(digest)
+
+ sb.Unsubscribe(sub)
+
+ if got := jobStatus(t, sb, seq); got != "assigned" {
+ t.Fatalf("status = %q, want assigned: the row stays the disconnected "+
+ "scanner's until it resumes or the grace expires", got)
+ }
+ if !jobIsDisconnected(t, sb, seq) {
+ t.Error("the row was not marked as belonging to a disconnected scanner")
+ }
+ if sb.addInflight(digest) {
+ t.Error("Unsubscribe released the in-flight digest: behaviour changed")
+ }
+}
+
+// TestScanDispatchQueue_ReturnsUndeliverableJobsToPending covers what the
+// drain does when it cannot hand a scanner everything it claimed.
+//
+// sub.send is 20 deep in production. dispatchJob's default branch resets the
+// row to 'pending' when the buffer is full; drainPendingJobs used to block for
+// up to five seconds and then simply return, leaving every row it had already
+// marked 'assigned' owned by a subscriber that was never sent them. Those rows
+// sat in 'assigned' — counted as active dispatch capacity — until the
+// five-minute ack timeout reclaimed them.
+//
+// The drain now puts the row it could not deliver back to 'pending', so the
+// re-dispatch loop can offer it again on its next tick.
+func TestScanDispatchQueue_ReturnsUndeliverableJobsToPending(t *testing.T) {
+ sb := newStuckBroadcaster(t)
+ sub := newTestScanSubscriber(t, sb, 2) // tiny buffer, nobody draining it
+ sub.capacity = 4 // the send buffer, not capacity, is the constraint here
+
+ var seqs []int64
+ for i := 0; i < 4; i++ {
+ seqs = append(seqs, seedJobWithDigest(t, sb, "sha256:burst"))
+ }
+
+ done := make(chan struct{})
+ go func() {
+ defer close(done)
+ sb.drainPendingJobs(sub, 0)
+ }()
+
+ select {
+ case <-done:
+ case <-time.After(30 * time.Second):
+ t.Fatal("drainPendingJobs never returned")
+ }
+
+ var assigned, pending int
+ for _, seq := range seqs {
+ switch got := jobStatus(t, sb, seq); got {
+ case "assigned":
+ assigned++
+ case "pending":
+ pending++
+ default:
+ t.Errorf("job %d in unexpected status %q", seq, got)
+ }
+ }
+
+ // Two made it into the buffer and are genuinely assigned. The third blocked
+ // out the five-second window and must have been handed back, and the fourth
+ // was never claimed at all.
+ if assigned != 2 || pending != 2 {
+ t.Errorf("assigned=%d pending=%d, want 2 assigned / 2 pending: a row the "+
+ "scanner was never sent must not stay assigned to it", assigned, pending)
+ }
+}
+
+// TestScanSkipped_RetiresAnUndecodableFrame is the hold-side half of the
+// scanner's answer to a frame it cannot parse
+// (TestUnparseableFramesAreAnsweredWithSkipped in
+// scanner/internal/e2e/protocol_test.go).
+//
+// The scanner now replies "skipped" for any job frame carrying a usable seq,
+// which is only worth doing if it actually retires the row. It does: the job
+// goes to 'completed', a skipped scan record lands in the PDS so the stale
+// loop leaves it alone, and the in-flight digest is released. Before that
+// reply existed, the row stayed 'assigned', timed out after five minutes, was
+// re-offered to the same scanner within thirty seconds, and was dropped again
+// — while hasActiveJobs counted it and no proactive scan was dispatched
+// anywhere in the deployment.
+func TestScanSkipped_RetiresAnUndecodableFrame(t *testing.T) {
+ sb := newStuckBroadcaster(t)
+ sub := newTestScanSubscriber(t, sb, 4)
+
+ const digest = "sha256:undecodable"
+ seq := seedJobWithDigest(t, sb, digest)
+ assignJob(t, sb, seq, sub, 0)
+ sb.addInflight(digest)
+
+ // Exactly what connectOnce sends when the config sub-document does not
+ // decode: no ack ever arrives, so the row is still 'assigned'.
+ sb.handleSkipped(sub, ScannerMessage{
+ Type: "skipped",
+ Seq: seq,
+ Reason: "malformed job config: json: cannot unmarshal string into Go value of type scanner.BlobDescriptor",
+ })
+
+ if got := jobStatus(t, sb, seq); got != "completed" {
+ t.Errorf("status = %q, want completed: a skip must be terminal", got)
+ }
+ if !sb.addInflight(digest) {
+ t.Error("skip did not release the in-flight digest")
+ }
+ sb.removeInflight(digest)
+
+ _, record, err := sb.pds.GetScanRecord(context.Background(), digest)
+ if err != nil {
+ t.Fatalf("no scan record for a skipped job: %v", err)
+ }
+ if record.Status != atproto.ScanStatusSkipped {
+ t.Errorf("record status = %q, want %q", record.Status, atproto.ScanStatusSkipped)
+ }
+
+ // Terminal means terminal: the re-dispatch loop must not pick it back up.
+ sb.reDispatchTimedOut()
+ if got := jobStatus(t, sb, seq); got != "completed" {
+ t.Errorf("after a re-dispatch tick: status = %q, want completed", got)
+ }
+ select {
+ case job := <-sub.send:
+ t.Fatalf("a skipped job (%d) was re-offered to the scanner", job.Seq)
+ default:
+ }
+}
diff --git a/pkg/hold/pds/scan_broadcaster_test.go b/pkg/hold/pds/scan_broadcaster_test.go
index a6aeae2..099e7b6 100644
--- a/pkg/hold/pds/scan_broadcaster_test.go
+++ b/pkg/hold/pds/scan_broadcaster_test.go
@@ -4,6 +4,8 @@ import (
"database/sql"
"testing"
"time"
+
+ "atcr.io/pkg/s3"
)
// newTestScanBroadcaster builds a ScanBroadcaster with just a database, no
@@ -30,6 +32,26 @@ func newTestScanBroadcaster(t *testing.T) *ScanBroadcaster {
return sb
}
+// newRecordingScanBroadcaster is newTestScanBroadcaster plus the two
+// dependencies handleResult needs to finish its work: an embedded PDS to write
+// the scan record into, and an S3 stand-in to take the SBOM blob. The bare
+// helper leaves both nil, which is fine for row bookkeeping and fatal for
+// anything that asserts on what was stored.
+func newRecordingScanBroadcaster(t *testing.T) *ScanBroadcaster {
+ t.Helper()
+
+ sb := newTestScanBroadcaster(t)
+ sb.inflight = make(map[string]struct{})
+ sb.ackTimeout = 5 * time.Minute
+
+ pds, _ := setupTestPDS(t)
+ sb.pds = pds
+ sb.holdDID = pds.did
+ sb.s3 = &s3.S3Service{Client: s3.NewMockS3Client(""), Bucket: "test-bucket"}
+
+ return sb
+}
+
// newTestScanSubscriber mirrors what Subscribe builds, registered with the
// broadcaster so Unsubscribe finds it.
func newTestScanSubscriber(t *testing.T, sb *ScanBroadcaster, bufSize int) *ScanSubscriber {
@@ -40,6 +62,9 @@ func newTestScanSubscriber(t *testing.T, sb *ScanBroadcaster, bufSize int) *Scan
send: make(chan *ScanJobEvent, bufSize),
id: "test-subscriber",
done: make(chan struct{}),
+ // One worker unless a test says otherwise, which is what a scanner
+ // that declares nothing is treated as.
+ capacity: 1,
}
sb.mu.Lock()
@@ -93,41 +118,66 @@ func TestScanUnsubscribe_IsIdempotent(t *testing.T) {
}
}
-// TestScanUnsubscribe_UnassignsJobsOnce verifies the idempotency guard protects
-// the job-reassignment UPDATE too. Re-running it would unassign jobs that a
-// replacement scanner had already been given.
-func TestScanUnsubscribe_UnassignsJobsOnce(t *testing.T) {
+// TestScanUnsubscribe_MarksItsOwnJobsOnce verifies the idempotency guard
+// protects the disconnect bookkeeping too, and that the bookkeeping is scoped
+// to this subscriber's rows.
+//
+// Unsubscribe used to flip every assigned and processing row straight back to
+// 'pending', which handed a running scan to whichever process connected next
+// (see TestScanUnsubscribe_DoesNotImmediatelyHandOffWorkStillRunning). It now
+// marks the disconnect and leaves the work where it is. Either way the guard
+// matters for the same reason: a dropped scanner unwinds both handleWriter and
+// handleReader, and each calls Unsubscribe.
+func TestScanUnsubscribe_MarksItsOwnJobsOnce(t *testing.T) {
sb := newTestScanBroadcaster(t)
sub := newTestScanSubscriber(t, sb, 4)
seedPendingJobs(t, sb, 1)
- if _, err := sb.db.Exec(`UPDATE scan_jobs SET status='assigned', assigned_to=?`, sub.id); err != nil {
+ if _, err := sb.db.Exec(`UPDATE scan_jobs SET status='processing', assigned_to=?`, sub.id); err != nil {
t.Fatalf("assign: %v", err)
}
sb.Unsubscribe(sub)
- var status string
- if err := sb.db.QueryRow(`SELECT status FROM scan_jobs LIMIT 1`).Scan(&status); err != nil {
- t.Fatalf("query: %v", err)
- }
- if status != "pending" {
- t.Errorf("expected job returned to pending, got %q", status)
+ var (
+ status string
+ assignedTo sql.NullString
+ disconnectedAt sql.NullTime
+ )
+ row := func() {
+ t.Helper()
+ if err := sb.db.QueryRow(
+ `SELECT status, assigned_to, disconnected_at FROM scan_jobs LIMIT 1`,
+ ).Scan(&status, &assignedTo, &disconnectedAt); err != nil {
+ t.Fatalf("query: %v", err)
+ }
}
+ row()
+ if status != "processing" || assignedTo.String != sub.id {
+ t.Errorf("job = %q/%q, want processing/%q: a disconnect is not evidence "+
+ "the scanner stopped scanning", status, assignedTo.String, sub.id)
+ }
+ if !disconnectedAt.Valid {
+ t.Error("the row was not marked as belonging to a disconnected scanner")
+ }
+ firstMark := disconnectedAt.Time
+
// Hand the job to a "replacement" scanner, then unsubscribe the dead one
- // again. The guard must stop it from stealing the job back.
- if _, err := sb.db.Exec(`UPDATE scan_jobs SET status='assigned', assigned_to=?`, "replacement"); err != nil {
+ // again. The guard must stop it from touching a row that has moved on.
+ if _, err := sb.db.Exec(
+ `UPDATE scan_jobs SET status='assigned', assigned_to=?, disconnected_at=NULL`, "replacement",
+ ); err != nil {
t.Fatalf("reassign: %v", err)
}
sb.Unsubscribe(sub)
- if err := sb.db.QueryRow(`SELECT status FROM scan_jobs LIMIT 1`).Scan(&status); err != nil {
- t.Fatalf("query: %v", err)
- }
- if status != "assigned" {
- t.Errorf("second Unsubscribe stole the replacement's job, status=%q", status)
+ row()
+ if status != "assigned" || assignedTo.String != "replacement" || disconnectedAt.Valid {
+ t.Errorf("second Unsubscribe touched the replacement's job: %q/%q "+
+ "disconnected=%v", status, assignedTo.String, disconnectedAt.Valid)
}
+ _ = firstMark
}
// TestScanDrainPendingJobs_ConcurrentUnsubscribe is the regression test for
@@ -143,6 +193,7 @@ func TestScanDrainPendingJobs_ConcurrentUnsubscribe(t *testing.T) {
func() {
sb := newTestScanBroadcaster(t)
sub := newTestScanSubscriber(t, sb, 1)
+ sub.capacity = 50 // the drain must walk rows, not stop at capacity
seedPendingJobs(t, sb, 50)
done := make(chan struct{})
diff --git a/pkg/hold/pds/scan_broadcaster_ws_test.go b/pkg/hold/pds/scan_broadcaster_ws_test.go
index 262e670..dc547e7 100644
--- a/pkg/hold/pds/scan_broadcaster_ws_test.go
+++ b/pkg/hold/pds/scan_broadcaster_ws_test.go
@@ -67,7 +67,7 @@ func TestScanBroadcaster_UnsubscribeIsIdempotent(t *testing.T) {
sb := newTestScanBroadcaster(t)
conn, _ := wsPair(t)
- sub := sb.Subscribe(conn, 0)
+ sub := sb.Subscribe(conn, 0, "", 0)
if got := subscriberCount(sb); got != 1 {
t.Fatalf("subscriber count after Subscribe = %d, want 1", got)
}
@@ -88,7 +88,7 @@ func TestScanBroadcaster_WriterExitsOnUnsubscribe(t *testing.T) {
sb := newTestScanBroadcaster(t)
conn, _ := wsPair(t)
- sub := sb.Subscribe(conn, 0)
+ sub := sb.Subscribe(conn, 0, "", 0)
sb.Unsubscribe(sub)
deadline := time.Now().Add(5 * time.Second)
@@ -108,7 +108,7 @@ func TestScanBroadcaster_DroppedScannerUnwindsOnce(t *testing.T) {
sb := newTestScanBroadcaster(t)
conn, closeClient := wsPair(t)
- sub := sb.Subscribe(conn, 0)
+ sub := sb.Subscribe(conn, 0, "", 0)
closeClient()
deadline := time.Now().Add(5 * time.Second)
diff --git a/pkg/hold/pds/xrpc.go b/pkg/hold/pds/xrpc.go
index 11cdabb..af56205 100644
--- a/pkg/hold/pds/xrpc.go
+++ b/pkg/hold/pds/xrpc.go
@@ -1101,6 +1101,24 @@ func (h *XRPCHandler) HandleSubscribeScanJobs(w http.ResponseWriter, r *http.Req
}
}
+ // How much scanning this process can run at once, and who it is.
+ //
+ // Both are optional: a scanner built before they existed sends neither,
+ // and the broadcaster then treats it as one worker with a
+ // per-connection identity — exactly the behaviour it had before. A
+ // malformed workers value is ignored rather than rejected, since the
+ // connection is still perfectly usable at the default.
+ workers := 0
+ if raw := r.URL.Query().Get("workers"); raw != "" {
+ if n, err := strconv.Atoi(raw); err == nil && n > 0 {
+ workers = n
+ } else {
+ slog.Warn("Scanner declared an unusable worker count, using the default",
+ "workers", raw)
+ }
+ }
+ instanceID := r.URL.Query().Get("instance")
+
// Upgrade to WebSocket
conn, err := upgrader.Upgrade(w, r, nil)
if err != nil {
@@ -1108,7 +1126,7 @@ func (h *XRPCHandler) HandleSubscribeScanJobs(w http.ResponseWriter, r *http.Req
return
}
- h.scanBroadcaster.Subscribe(conn, cursor)
+ h.scanBroadcaster.Subscribe(conn, cursor, instanceID, workers)
}
// ScanBroadcasterRef returns the scan broadcaster (used by OCI handler to enqueue jobs)
diff --git a/scanner/cmd/scanner/main.go b/scanner/cmd/scanner/main.go
index b2d8c01..884704f 100644
--- a/scanner/cmd/scanner/main.go
+++ b/scanner/cmd/scanner/main.go
@@ -77,8 +77,11 @@ Environment variables always override file values (SCANNER_ prefix).`,
// Create priority queue
q := queue.NewJobQueue(cfg.Scanner.QueueSize)
- // Create hold WebSocket client
+ // Create hold WebSocket client. Declaring the worker count is what
+ // lets the hold keep that many scans in flight for this process
+ // instead of one.
holdClient := client.NewHoldClient(cfg.Hold.URL, cfg.Hold.Secret, q)
+ holdClient.SetWorkers(cfg.Scanner.Workers)
// Start WebSocket connection (feeds queue)
go holdClient.Connect()
diff --git a/scanner/digest.go b/scanner/digest.go
new file mode 100644
index 0000000..f73a189
--- /dev/null
+++ b/scanner/digest.go
@@ -0,0 +1,63 @@
+package scanner
+
+import (
+ "fmt"
+ "strings"
+)
+
+// Digest is a content digest that has been validated, and is therefore safe to
+// use both as a blob name on the wire and as a single path element on disk.
+//
+// The validation is deliberately narrow. A digest reaches the scanner from an
+// io.atcr.manifest record in a user's own PDS, which the user can write
+// directly, and the hold's dispatch guards check the hold DID, the layer count,
+// the subject and the config but never the digest format. Downstream the string
+// is joined onto the blobs directory and handed to os.Create, so anything that
+// is not exactly an algorithm and a hex string is a filesystem primitive
+// wearing a digest's clothes.
+//
+// Only sha256 is accepted. It is the only algorithm the OCI layout the scanner
+// builds uses (blobs/sha256/) and the only one stereoscope can read, so
+// anything else is unscannable however well formed it is; refusing it here
+// turns a late, retried failure into an early, permanent one.
+type Digest struct {
+ Algorithm string // always "sha256" today
+ Hex string // lowercase hex, exactly HexLen characters
+}
+
+// SHA256 is the only digest algorithm the scanner accepts.
+const SHA256 = "sha256"
+
+// HexLen is the number of hex characters in a sha256 digest.
+const HexLen = 64
+
+// String renders the digest back into its "algorithm:hex" form.
+func (d Digest) String() string { return d.Algorithm + ":" + d.Hex }
+
+// ParseDigest validates a digest string and returns its parts.
+//
+// It accepts exactly "sha256:" followed by 64 lowercase hex characters, and
+// nothing else: no other algorithm, no uppercase, no other length, no leading
+// or trailing anything. Because the result is constrained to [0-9a-f], the Hex
+// field cannot contain a separator, a dot, or a NUL, and so cannot escape the
+// directory it is joined onto.
+func ParseDigest(digest string) (Digest, error) {
+ algorithm, hex, ok := strings.Cut(digest, ":")
+ if !ok {
+ return Digest{}, fmt.Errorf("digest %q has no algorithm prefix", digest)
+ }
+ if algorithm != SHA256 {
+ return Digest{}, fmt.Errorf("digest %q uses unsupported algorithm %q, want %s", digest, algorithm, SHA256)
+ }
+ if len(hex) != HexLen {
+ return Digest{}, fmt.Errorf("digest %q has %d hex characters, want %d", digest, len(hex), HexLen)
+ }
+ for i := 0; i < len(hex); i++ {
+ c := hex[i]
+ if (c >= '0' && c <= '9') || (c >= 'a' && c <= 'f') {
+ continue
+ }
+ return Digest{}, fmt.Errorf("digest %q is not lowercase hex", digest)
+ }
+ return Digest{Algorithm: algorithm, Hex: hex}, nil
+}
diff --git a/scanner/digest_test.go b/scanner/digest_test.go
new file mode 100644
index 0000000..f7e9155
--- /dev/null
+++ b/scanner/digest_test.go
@@ -0,0 +1,61 @@
+package scanner
+
+import (
+ "strings"
+ "testing"
+)
+
+func TestParseDigestAcceptsOnlyAlgorithmAndHex(t *testing.T) {
+ valid := "sha256:" + strings.Repeat("ab", 32)
+
+ d, err := ParseDigest(valid)
+ if err != nil {
+ t.Fatalf("ParseDigest(%q) = %v, want it accepted", valid, err)
+ }
+ if d.Algorithm != "sha256" {
+ t.Errorf("Algorithm = %q, want sha256", d.Algorithm)
+ }
+ if d.Hex != strings.Repeat("ab", 32) {
+ t.Errorf("Hex = %q, want the 64 hex characters", d.Hex)
+ }
+ if got := d.String(); got != valid {
+ t.Errorf("String() = %q, want %q", got, valid)
+ }
+}
+
+// TestParseDigestRejectsEverythingElse is the security boundary: anything that
+// is not exactly "sha256:<64 lowercase hex>" must be refused before it can be
+// joined onto a filesystem path or sent to the hold as a blob name.
+func TestParseDigestRejectsEverythingElse(t *testing.T) {
+ hex64 := strings.Repeat("ab", 32)
+
+ cases := []struct {
+ name string
+ digest string
+ }{
+ {"empty", ""},
+ {"no algorithm prefix", hex64},
+ {"path traversal in the hex", "sha256:../../../escaped-marker"},
+ {"traversal with no algorithm", "../../../escaped-marker"},
+ {"absolute path", "sha256:/etc/passwd"},
+ {"separator in the hex", "sha256:ab/cd"},
+ {"null byte", "sha256:" + hex64 + "\x00"},
+ {"unsupported algorithm", "sha512:" + strings.Repeat("cd", 64)},
+ {"uppercase hex", "sha256:" + strings.ToUpper(hex64)},
+ {"short hex", "sha256:abcd"},
+ {"long hex", "sha256:" + hex64 + "ab"},
+ {"non-hex characters", "sha256:" + strings.Repeat("zz", 32)},
+ {"empty hex", "sha256:"},
+ {"empty algorithm", ":" + hex64},
+ {"double colon", "sha256:sha256:" + hex64},
+ {"leading space", " sha256:" + hex64},
+ }
+
+ for _, tc := range cases {
+ t.Run(tc.name, func(t *testing.T) {
+ if d, err := ParseDigest(tc.digest); err == nil {
+ t.Errorf("ParseDigest(%q) accepted it as %+v, want an error", tc.digest, d)
+ }
+ })
+ }
+}
diff --git a/scanner/internal/client/hold.go b/scanner/internal/client/hold.go
index 0a3f8e3..d03290e 100644
--- a/scanner/internal/client/hold.go
+++ b/scanner/internal/client/hold.go
@@ -3,7 +3,11 @@
package client
import (
+ "crypto/rand"
+ "crypto/sha256"
+ "encoding/hex"
"encoding/json"
+ "errors"
"fmt"
"io"
"log/slog"
@@ -30,18 +34,54 @@ type HoldClient struct {
conn *websocket.Conn
mu sync.Mutex // protects conn writes
done chan struct{}
+
+ // instanceID identifies this scanner process to the hold, and is sent on
+ // every connect. It is deliberately per-process and not persisted: a hold
+ // hands a reconnecting instance back the jobs its workers never stopped
+ // running, and a process that actually restarted has lost that work and
+ // must not claim it.
+ instanceID string
+
+ // workers is how many scans this process runs at once. The hold uses it as
+ // the dispatch budget for this connection; a hold that does not know the
+ // parameter ignores it and assumes one.
+ workers int
}
// NewHoldClient creates a new hold client
func NewHoldClient(holdURL, secret string, q *queue.JobQueue) *HoldClient {
return &HoldClient{
- holdURL: holdURL,
- secret: secret,
- queue: q,
- done: make(chan struct{}),
+ holdURL: holdURL,
+ secret: secret,
+ queue: q,
+ done: make(chan struct{}),
+ instanceID: newInstanceID(),
+ workers: 1,
}
}
+// SetWorkers tells the hold how many scans this process runs concurrently, so
+// it can keep that many jobs in flight instead of one.
+//
+// This is a setter rather than a constructor argument on purpose: the
+// connection is entirely usable without it, and a caller that never calls it
+// gets the single-worker default the hold assumes anyway.
+func (c *HoldClient) SetWorkers(n int) {
+ if n > 0 {
+ c.workers = n
+ }
+}
+
+func newInstanceID() string {
+ b := make([]byte, 8)
+ if _, err := rand.Read(b); err != nil {
+ // The hold falls back to a per-connection identity when none is sent,
+ // which costs resumption on reconnect and nothing else.
+ return ""
+ }
+ return hex.EncodeToString(b)
+}
+
// Connect establishes the WebSocket connection with auto-reconnect
func (c *HoldClient) Connect() {
var cursor int64 = -1
@@ -91,6 +131,12 @@ func (c *HoldClient) connectOnce(cursor int64) error {
if cursor >= 0 {
q.Set("cursor", fmt.Sprintf("%d", cursor))
}
+ if c.instanceID != "" {
+ q.Set("instance", c.instanceID)
+ }
+ if c.workers > 0 {
+ q.Set("workers", fmt.Sprintf("%d", c.workers))
+ }
u.RawQuery = q.Encode()
slog.Info("Connecting to hold service", "url", u.Host)
@@ -117,62 +163,118 @@ func (c *HoldClient) connectOnce(cursor int64) error {
return err
}
- var raw scanner.ScanJobRaw
- if err := json.Unmarshal(data, &raw); err != nil {
- slog.Error("Failed to unmarshal message", "error", err)
- continue
- }
-
- if raw.Type != "job" {
- slog.Warn("Unknown message type from hold", "type", raw.Type)
- continue
- }
-
- // Parse config and layers from raw JSON
- var config scanner.BlobDescriptor
- if err := json.Unmarshal(raw.Config, &config); err != nil {
- slog.Error("Failed to unmarshal config", "seq", raw.Seq, "error", err)
- continue
- }
-
- var layers []scanner.BlobDescriptor
- if err := json.Unmarshal(raw.Layers, &layers); err != nil {
- slog.Error("Failed to unmarshal layers", "seq", raw.Seq, "error", err)
- continue
- }
-
- job := &scanner.ScanJob{
- Seq: raw.Seq,
- ManifestDigest: raw.ManifestDigest,
- Repository: raw.Repository,
- Tag: raw.Tag,
- UserDID: raw.UserDID,
- UserHandle: raw.UserHandle,
- HoldDID: raw.HoldDID,
- HoldEndpoint: raw.HoldEndpoint,
- Tier: raw.Tier,
- Config: config,
- Layers: layers,
- }
-
- // Send ack immediately
- c.SendAck(job.Seq)
-
- // Enqueue into priority queue
- if !c.queue.Enqueue(job) {
- slog.Warn("Queue full, sending error",
- "seq", job.Seq,
- "repository", job.Repository)
- c.SendError(job.Seq, "scanner queue full")
- }
+ c.handleFrame(data)
}
}
+// handleFrame decodes one frame from the hold and either enqueues the job or
+// answers for it.
+//
+// The invariant, which every branch below keeps: a frame carrying a usable seq
+// is never dropped in silence. The hold writes status='assigned' for that seq
+// before it writes the frame, and its only escape from that row is an
+// acknowledgement or a verdict from us. A frame we drop instead sits assigned
+// until the five-minute ack timeout, is re-offered within thirty seconds — to
+// this same scanner, which cannot decode it this time either — and holds the
+// hold's single proactive dispatch slot the entire time. One such frame stops
+// proactive scanning for every user of that hold, permanently.
+func (c *HoldClient) handleFrame(data []byte) {
+ var raw scanner.ScanJobRaw
+ if err := json.Unmarshal(data, &raw); err != nil {
+ // encoding/json records the first type error and keeps decoding the
+ // rest of the object, so the seq that addresses the row is often still
+ // usable even when the frame as a whole was rejected.
+ c.rejectJob(raw.Seq, "malformed job frame", err)
+ return
+ }
+
+ if raw.Type != "job" {
+ // Not a job, so no row of ours is waiting on an answer. The hold sends
+ // nothing but "job" today; a future message type is better ignored by
+ // an old scanner than answered with a verdict about a job.
+ slog.Warn("Unknown message type from hold", "type", raw.Type, "seq", raw.Seq)
+ return
+ }
+
+ // Parse config and layers from raw JSON
+ var config scanner.BlobDescriptor
+ if err := json.Unmarshal(raw.Config, &config); err != nil {
+ c.rejectJob(raw.Seq, "malformed job config", err)
+ return
+ }
+
+ var layers []scanner.BlobDescriptor
+ if err := json.Unmarshal(raw.Layers, &layers); err != nil {
+ c.rejectJob(raw.Seq, "malformed job layers", err)
+ return
+ }
+
+ job := &scanner.ScanJob{
+ Seq: raw.Seq,
+ ManifestDigest: raw.ManifestDigest,
+ Repository: raw.Repository,
+ Tag: raw.Tag,
+ UserDID: raw.UserDID,
+ UserHandle: raw.UserHandle,
+ HoldDID: raw.HoldDID,
+ HoldEndpoint: raw.HoldEndpoint,
+ Tier: raw.Tier,
+ Config: config,
+ Layers: layers,
+ }
+
+ // Send ack immediately
+ c.SendAck(job.Seq)
+
+ // Enqueue into priority queue
+ if !c.queue.Enqueue(job) {
+ slog.Warn("Queue full, sending error",
+ "seq", job.Seq,
+ "repository", job.Repository)
+ c.SendError(job.Seq, "scanner queue full")
+ }
+}
+
+// rejectJob answers a job frame this scanner cannot decode.
+//
+// "skipped" and not "error": the hold treats a failure as transient and
+// re-queues it on the rescan interval, but a frame that does not decode will
+// not decode on the next attempt either. A skip is terminal — the hold marks
+// the row completed, writes a scan record carrying the reason, and releases
+// the manifest from its in-flight set — so the job leaves the rotation and the
+// reason reaches the user instead of vanishing into this process's log.
+//
+// A seq of zero is the one frame that cannot be answered: the hold's rows
+// start at 1, so there is no job to address. Nothing is stranded by staying
+// quiet there either, because the hold assigns the row by seq before it sends,
+// and a seq that never survived the wire is not the seq it assigned.
+func (c *HoldClient) rejectJob(seq int64, reason string, err error) {
+ if seq <= 0 {
+ slog.Error("Dropping undecodable frame with no usable seq",
+ "reason", reason, "error", err)
+ return
+ }
+
+ slog.Error("Rejecting undecodable job frame",
+ "seq", seq, "reason", reason, "error", err)
+ c.SendSkipped(seq, fmt.Sprintf("%s: %v", reason, err))
+}
+
// SendAck sends an acknowledgement for a received job
func (c *HoldClient) SendAck(seq int64) {
c.sendJSON(scanner.AckMessage{Type: "ack", Seq: seq})
}
+// SendStarted tells the hold a worker has begun this scan.
+//
+// The ack this job already got was sent from the reader goroutine on receipt,
+// before the job was queued. This is the signal the hold measures its scanning
+// deadline from; a hold too old to know the message logs and ignores it, and
+// falls back to its own budget measured from dispatch.
+func (c *HoldClient) SendStarted(seq int64) {
+ c.sendJSON(scanner.StartedMessage{Type: "started", Seq: seq})
+}
+
// SendResult sends scan results back to the hold
func (c *HoldClient) SendResult(seq int64, result *scanner.ScanResult) {
msg := scanner.ResultMessage{
@@ -227,11 +329,15 @@ func (c *HoldClient) Close() {
// GetBlobPresignedURL gets a presigned download URL from the hold service.
// If secret is non-empty, it is sent as a Bearer token for private hold access.
-func GetBlobPresignedURL(holdEndpoint, holdDID, digest, secret string) (string, error) {
+//
+// The digest is a scanner.Digest rather than a string so that only a validated
+// "sha256:" can ever be asked for. Callers parse once, at the boundary,
+// and the same value then names the blob on the wire and the file on disk.
+func GetBlobPresignedURL(holdEndpoint, holdDID string, digest scanner.Digest, secret string) (string, error) {
reqURL := fmt.Sprintf("%s/xrpc/com.atproto.sync.getBlob?did=%s&cid=%s&method=GET",
holdEndpoint,
url.QueryEscape(holdDID),
- url.QueryEscape(digest))
+ url.QueryEscape(digest.String()))
req, err := http.NewRequest("GET", reqURL, nil)
if err != nil {
@@ -262,27 +368,93 @@ func GetBlobPresignedURL(holdEndpoint, holdDID, digest, secret string) (string,
return result.URL, nil
}
-// DownloadBlob downloads a blob from a presigned URL to a local file
-func DownloadBlob(presignedURL, destPath string) error {
+// Errors a blob download can fail with that no retry can fix. Both are
+// decided by bytes the scanner has already seen, so a caller mapping them to a
+// permanent skip is not guessing.
+var (
+ // ErrBlobCorrupt means the bytes that arrived are not the bytes the
+ // descriptor described. It deliberately does not distinguish "the digest
+ // is wrong" from "the size is wrong": both fields come from the same
+ // user-writable manifest record, so a disagreement condemns the whole
+ // descriptor rather than proving anything about one field.
+ ErrBlobCorrupt = errors.New("digest mismatch")
+
+ // ErrBlobTooLarge means the transfer would take the job past the byte
+ // ceiling the configuration set.
+ ErrBlobTooLarge = errors.New("image too large")
+)
+
+// BlobExpectation is what a download has to turn out to be.
+type BlobExpectation struct {
+ // Digest is the validated digest the bytes must hash to. This is the
+ // authority: it is checked always, and it is what makes serving the wrong
+ // object detectable.
+ Digest scanner.Digest
+
+ // DeclaredSize is the descriptor's size field, corroborating evidence
+ // rather than authority. Zero or negative means the record claimed no
+ // size, in which case the digest alone decides.
+ DeclaredSize int64
+
+ // MaxBytes caps this transfer. Negative means unbounded. Zero is a real
+ // ceiling of zero bytes: it is what remains when an earlier blob in the
+ // same job has spent the whole budget.
+ MaxBytes int64
+}
+
+// DownloadBlob downloads a blob from a presigned URL to a local file and
+// verifies it against want, returning the number of bytes written.
+//
+// Verification happens while streaming, not by re-reading the file: the hash
+// and the byte count are both accumulated by the same io.Copy that writes the
+// blob, so a scan can never catalog bytes that were not checked, and a blob
+// over the ceiling stops costing bandwidth one byte past it.
+func DownloadBlob(presignedURL, destPath string, want BlobExpectation) (int64, error) {
resp, err := httpClient.Get(presignedURL)
if err != nil {
- return fmt.Errorf("failed to download blob: %w", err)
+ return 0, fmt.Errorf("failed to download blob: %w", err)
}
defer resp.Body.Close()
if resp.StatusCode != http.StatusOK {
- return fmt.Errorf("download returned status %d", resp.StatusCode)
+ return 0, fmt.Errorf("download returned status %d", resp.StatusCode)
}
out, err := os.Create(destPath)
if err != nil {
- return fmt.Errorf("failed to create file: %w", err)
+ return 0, fmt.Errorf("failed to create file: %w", err)
}
defer out.Close()
- if _, err := io.Copy(out, resp.Body); err != nil {
- return fmt.Errorf("failed to write blob: %w", err)
+ // One byte past the ceiling is all the proof needed that the ceiling was
+ // broken, and reading no further is the point: the claimed sizes that let
+ // the job through the pre-check are the attacker's to choose.
+ body := io.Reader(resp.Body)
+ if want.MaxBytes >= 0 {
+ body = io.LimitReader(resp.Body, want.MaxBytes+1)
}
- return nil
+ hasher := sha256.New()
+ n, err := io.Copy(io.MultiWriter(out, hasher), body)
+ if err != nil {
+ return n, fmt.Errorf("failed to write blob: %w", err)
+ }
+
+ if want.MaxBytes >= 0 && n > want.MaxBytes {
+ return n, fmt.Errorf("%w: blob %s exceeds the %d bytes left in the job's budget",
+ ErrBlobTooLarge, want.Digest, want.MaxBytes)
+ }
+
+ gotHex := hex.EncodeToString(hasher.Sum(nil))
+ sizeDisagrees := want.DeclaredSize > 0 && n != want.DeclaredSize
+ if gotHex != want.Digest.Hex || sizeDisagrees {
+ claimedSize := "no declared size"
+ if want.DeclaredSize > 0 {
+ claimedSize = fmt.Sprintf("%d bytes", want.DeclaredSize)
+ }
+ return n, fmt.Errorf("%w: descriptor claims %s and %s, received %d bytes hashing to %s:%s",
+ ErrBlobCorrupt, want.Digest, claimedSize, n, want.Digest.Algorithm, gotHex)
+ }
+
+ return n, nil
}
diff --git a/scanner/internal/config/config_edge_test.go b/scanner/internal/config/config_edge_test.go
new file mode 100644
index 0000000..0c287c9
--- /dev/null
+++ b/scanner/internal/config/config_edge_test.go
@@ -0,0 +1,270 @@
+package config
+
+import (
+ "os"
+ "path/filepath"
+ "strings"
+ "testing"
+)
+
+// writeYAML writes a scanner config file containing the two required keys plus
+// whatever extra YAML the case supplies, and returns its path.
+func writeYAML(t *testing.T, extra string) string {
+ t.Helper()
+ path := filepath.Join(t.TempDir(), "scanner.yaml")
+ body := "hold:\n url: \"ws://hold.example\"\n secret: \"s3cret\"\n" + extra
+ if err := os.WriteFile(path, []byte(body), 0o600); err != nil {
+ t.Fatalf("write config: %v", err)
+ }
+ return path
+}
+
+// TestLoadConfigRejectsOnlyURLAndSecret is the positive control: these are the
+// only two fields LoadConfig validates, and they do fail loudly.
+func TestLoadConfigRejectsOnlyURLAndSecret(t *testing.T) {
+ clearScannerEnv(t)
+
+ if _, err := LoadConfig(""); err == nil {
+ t.Error("empty config was accepted; hold.url should be required")
+ }
+
+ path := filepath.Join(t.TempDir(), "scanner.yaml")
+ if err := os.WriteFile(path, []byte("hold:\n url: \"ws://hold.example\"\n"), 0o600); err != nil {
+ t.Fatalf("write config: %v", err)
+ }
+ if _, err := LoadConfig(path); err == nil {
+ t.Error("config without hold.secret was accepted; it should be required")
+ }
+}
+
+// TestLoadConfigAcceptsSilentlyBrokenValues is the finding: every one of these
+// loads clean, logs nothing at boot, and produces a scanner that connects,
+// reports healthy on /healthz, and then either fails or strands every job it
+// is given. LoadConfig validates two strings for emptiness and nothing else.
+//
+// Each case names the runtime consequence; the behavioural proof for the ones
+// that reach the pipeline is in scanner/internal/e2e/protocol_test.go.
+func TestLoadConfigAcceptsSilentlyBrokenValues(t *testing.T) {
+ cases := []struct {
+ name string
+ yaml string
+ env map[string]string
+ check func(*testing.T, *Config)
+ breakage string
+ }{
+ {
+ name: "workers: 0",
+ yaml: "scanner:\n workers: 0\n",
+ breakage: "WorkerPool.Start launches no goroutines; jobs are acked and then sit in the queue until the hold's 10 minute processing timeout fails them",
+ check: func(t *testing.T, c *Config) {
+ if c.Scanner.Workers != 0 {
+ t.Errorf("workers = %d, want 0", c.Scanner.Workers)
+ }
+ },
+ },
+ {
+ name: "workers: negative",
+ yaml: "scanner:\n workers: -4\n",
+ breakage: "same as workers: 0 — the for loop simply does not run",
+ check: func(t *testing.T, c *Config) {
+ if c.Scanner.Workers != -4 {
+ t.Errorf("workers = %d, want -4", c.Scanner.Workers)
+ }
+ },
+ },
+ {
+ name: "queue_size: 0",
+ yaml: "scanner:\n queue_size: 0\n",
+ breakage: "every Enqueue returns false, so every job is acked and then failed with \"scanner queue full\", which the hold retries forever",
+ check: func(t *testing.T, c *Config) {
+ if c.Scanner.QueueSize != 0 {
+ t.Errorf("queue_size = %d, want 0", c.Scanner.QueueSize)
+ }
+ },
+ },
+ {
+ name: "queue_size: negative",
+ yaml: "scanner:\n queue_size: -1\n",
+ breakage: "same as queue_size: 0; Len() >= maxSize is true for an empty queue",
+ check: func(t *testing.T, c *Config) {
+ if c.Scanner.QueueSize != -1 {
+ t.Errorf("queue_size = %d, want -1", c.Scanner.QueueSize)
+ }
+ },
+ },
+ {
+ name: "tmp_dir: empty",
+ yaml: "vuln:\n tmp_dir: \"\"\n",
+ breakage: "WorkerPool.Start skips the TMPDIR export, and processJob's ensureDir(\"\") fails, so every job errors retryably",
+ check: func(t *testing.T, c *Config) {
+ if c.Vuln.TmpDir != "" {
+ t.Errorf("tmp_dir = %q, want empty", c.Vuln.TmpDir)
+ }
+ },
+ },
+ {
+ name: "db_path: empty with vuln enabled",
+ yaml: "vuln:\n enabled: true\n db_path: \"\"\n",
+ breakage: "initializeVulnDatabase failure is only logged (\"scanning will be disabled\"), but nothing actually disables it: each scan then fails in scanVulnerabilities",
+ check: func(t *testing.T, c *Config) {
+ if c.Vuln.DBPath != "" {
+ t.Errorf("db_path = %q, want empty", c.Vuln.DBPath)
+ }
+ },
+ },
+ {
+ name: "hold.url is not a URL",
+ env: map[string]string{"SCANNER_HOLD_URL": "hold.example:8080"},
+ breakage: "url.Parse accepts it, the ws/wss switch does not match, and Connect redials a bad address every 5 seconds forever",
+ check: func(t *testing.T, c *Config) {
+ if c.Hold.URL != "hold.example:8080" {
+ t.Errorf("hold.url = %q", c.Hold.URL)
+ }
+ },
+ },
+ {
+ name: "max_image_size: negative",
+ yaml: "vuln:\n max_image_size: -1\n",
+ breakage: "the guard is `if MaxImageSize > 0`, so a negative ceiling silently means no ceiling at all",
+ check: func(t *testing.T, c *Config) {
+ if c.Vuln.MaxImageSize != -1 {
+ t.Errorf("max_image_size = %d, want -1", c.Vuln.MaxImageSize)
+ }
+ },
+ },
+ {
+ name: "log_level: nonsense",
+ yaml: "log_level: \"verbose\"\n",
+ breakage: "InitLoggerWithShipper's default branch quietly falls back to info, so a typo in the level is invisible",
+ check: func(t *testing.T, c *Config) {
+ if c.LogLevel != "verbose" {
+ t.Errorf("log_level = %q", c.LogLevel)
+ }
+ },
+ },
+ {
+ name: "server.addr is unbindable",
+ yaml: "server:\n addr: \"nope:not-a-port\"\n",
+ breakage: "ListenAndServe fails in a goroutine and only logs; the scanner keeps running with no health endpoint",
+ check: func(t *testing.T, c *Config) {
+ if c.Server.Addr != "nope:not-a-port" {
+ t.Errorf("server.addr = %q", c.Server.Addr)
+ }
+ },
+ },
+ }
+
+ for _, tc := range cases {
+ t.Run(tc.name, func(t *testing.T) {
+ clearScannerEnv(t)
+ for k, v := range tc.env {
+ t.Setenv(k, v)
+ }
+ cfg, err := LoadConfig(writeYAML(t, tc.yaml))
+ if err != nil {
+ t.Fatalf("LoadConfig rejected %s — this finding is fixed: %v", tc.name, err)
+ }
+ tc.check(t, cfg)
+ t.Logf("accepted; runtime consequence: %s", tc.breakage)
+ })
+ }
+}
+
+// TestEmptyEnvVarDoesNotOverrideDefaults pins the one thing that saves
+// SCANNER_VULN_TMP_DIR="" from being the same hazard as the YAML key: Viper is
+// constructed without AllowEmptyEnv, so an env var set to the empty string is
+// treated as unset. The YAML path is the reachable one.
+//
+// It is worth pinning because it is accidental rather than intended: enabling
+// AllowEmptyEnv anywhere in pkg/config would silently make every empty-string
+// env var in every deployment start overriding defaults.
+func TestEmptyEnvVarDoesNotOverrideDefaults(t *testing.T) {
+ clearScannerEnv(t)
+ t.Setenv("SCANNER_HOLD_URL", "ws://hold.example")
+ t.Setenv("SCANNER_HOLD_SECRET", "s3cret")
+ t.Setenv("SCANNER_VULN_TMP_DIR", "")
+
+ cfg, err := LoadConfig("")
+ if err != nil {
+ t.Fatalf("LoadConfig: %v", err)
+ }
+ if cfg.Vuln.TmpDir != "/var/lib/atcr-scanner/tmp" {
+ t.Errorf("tmp_dir = %q; an empty env var now overrides the default, which makes "+
+ "SCANNER_VULN_TMP_DIR= as dangerous as the YAML key", cfg.Vuln.TmpDir)
+ }
+}
+
+// TestDefaultsAreSane guards the values the deployment relies on when nothing
+// is set, so a regression in setScannerDefaults shows up here rather than in
+// production.
+func TestDefaultsAreSane(t *testing.T) {
+ clearScannerEnv(t)
+ cfg := DefaultConfig()
+
+ if cfg.Scanner.Workers < 1 {
+ t.Errorf("default workers = %d, want at least 1", cfg.Scanner.Workers)
+ }
+ if cfg.Scanner.QueueSize < 1 {
+ t.Errorf("default queue_size = %d, want at least 1", cfg.Scanner.QueueSize)
+ }
+ if cfg.Vuln.TmpDir == "" {
+ t.Error("default tmp_dir is empty; extraction would land on /tmp")
+ }
+ if cfg.Vuln.TmpDir == "/tmp" {
+ t.Error("default tmp_dir is /tmp, which the code comments call out as a small tmpfs")
+ }
+}
+
+// clearScannerEnv removes SCANNER_* and LOG_SHIPPER_* from the environment for
+// the duration of a test, so a developer's shell cannot change what the
+// configuration loader sees.
+func clearScannerEnv(t *testing.T) {
+ t.Helper()
+ for _, kv := range os.Environ() {
+ key, _, ok := strings.Cut(kv, "=")
+ if !ok {
+ continue
+ }
+ if !strings.HasPrefix(key, "SCANNER_") && !strings.HasPrefix(key, "LOG_SHIPPER_") {
+ continue
+ }
+ // Setenv registers the restore; Unsetenv is what the test actually
+ // needs, since Viper treats an empty value as unset anyway.
+ t.Setenv(key, "")
+ os.Unsetenv(key)
+ }
+}
+
+// TestMalformedYAMLIsIgnoredEntirely is the loudest silent failure in the
+// configuration path, and it is not scanner-specific: NewViper in
+// pkg/config/viper.go discards the result of ReadInConfig, so a config file
+// that does not parse is skipped whole. Nothing logs, nothing fails; the
+// process boots on defaults plus environment.
+//
+// In the shipped deployment hold.url and hold.secret come from the
+// environment, so even LoadConfig's two checks pass. A typo in scanner.yaml
+// therefore silently reverts workers, queue_size, tmp_dir, db_path,
+// max_image_size, log level and log shipping to their defaults on a scanner
+// that reports itself healthy.
+func TestMalformedYAMLIsIgnoredEntirely(t *testing.T) {
+ clearScannerEnv(t)
+ t.Setenv("SCANNER_HOLD_URL", "ws://hold.example")
+ t.Setenv("SCANNER_HOLD_SECRET", "s3cret")
+
+ path := filepath.Join(t.TempDir(), "scanner.yaml")
+ body := "scanner:\n workers: 8\n workers: 9\nvuln:\n tmp_dir: \"/data/scanner/tmp\"\n"
+ if err := os.WriteFile(path, []byte(body), 0o600); err != nil {
+ t.Fatalf("write config: %v", err)
+ }
+
+ cfg, err := LoadConfig(path)
+ if err != nil {
+ t.Fatalf("LoadConfig rejected an unparseable file — this finding is fixed: %v", err)
+ }
+ if cfg.Vuln.TmpDir != "/data/scanner/tmp" {
+ t.Logf("unparseable YAML silently ignored: tmp_dir fell back to %q, workers to %d",
+ cfg.Vuln.TmpDir, cfg.Scanner.Workers)
+ } else {
+ t.Error("the file parsed after all; pick a different malformed body for this test")
+ }
+}
diff --git a/scanner/internal/e2e/bench_test.go b/scanner/internal/e2e/bench_test.go
new file mode 100644
index 0000000..e856b90
--- /dev/null
+++ b/scanner/internal/e2e/bench_test.go
@@ -0,0 +1,1595 @@
+package e2e
+
+// Performance and resource scenarios for the scan pipeline.
+//
+// Nothing in here runs during an ordinary `go test ./...`: every Test is gated
+// on ATCR_SCANNER_PERF=1 and every Benchmark needs -bench, and both skip when
+// the image fixture they need is absent. Fixtures are the gitignored OCI
+// layouts under ../mockhold/testdata/blobs; see fixtureLayout for how to pull
+// the ones these scenarios use.
+//
+// # Why several scenarios drive a replica of the pipeline rather than the
+// # WorkerPool
+//
+// worker.go used to dereference result.Summary.Total on every successful scan
+// while Summary is only populated when cfg.Vuln.Enabled. The harness disables
+// Grype (enabling it would download a multi-hundred-MB database), so the first
+// successful scan through the real WorkerPool panicked and took the whole test
+// binary with it. That is fixed; the scenarios below were written under it and
+// have not been re-cut, so they still avoid successful scans through the pool.
+//
+// The consequence for measurement was concrete: no scenario that drives the
+// real worker could observe more than one successful scan, so throughput,
+// sustained-memory and concurrency numbers cannot come from that path. Those
+// scenarios instead run pipelineOnce, which performs the same four steps
+// against the same libraries and the same mock hold: presign + download, OCI
+// layout assembly, stereoscope load/extract, Syft catalog, SPDX encode. It is
+// a replica, not the production function — buildOCILayout and generateSBOM are
+// unexported in package scan and this file may not add a seam to them — so
+// treat its absolute numbers as "what the work costs", and the WorkerPool
+// scenarios below as "what the worker adds on top".
+//
+// The scenarios that do drive the real WorkerPool (queue burst, cooldown
+// cadence, reconnect leaks) are all built from jobs that terminate in an error
+// or a skip, which is the only way to run many jobs through the real worker
+// today.
+
+import (
+ "archive/tar"
+ "bufio"
+ "bytes"
+ "compress/gzip"
+ "context"
+ "crypto/sha256"
+ "encoding/json"
+ "fmt"
+ "io/fs"
+ "math"
+ "net/http"
+ "net/http/httptest"
+ "os"
+ "os/exec"
+ "path/filepath"
+ "runtime"
+ "runtime/debug"
+ "strconv"
+ "strings"
+ "sync"
+ "testing"
+ "time"
+
+ scanner "atcr.io/scanner"
+ "atcr.io/scanner/internal/client"
+ "atcr.io/scanner/internal/mockhold"
+ "atcr.io/scanner/internal/scan"
+
+ "github.com/anchore/stereoscope/pkg/file"
+ "github.com/anchore/stereoscope/pkg/image/oci"
+ "github.com/anchore/syft/syft"
+ "github.com/anchore/syft/syft/format"
+ "github.com/anchore/syft/syft/format/spdxjson"
+ "github.com/anchore/syft/syft/source/stereoscopesource"
+)
+
+// perfEnv gates every Test in this file. Benchmarks are gated by -bench.
+const perfEnv = "ATCR_SCANNER_PERF"
+
+func requirePerf(t *testing.T) {
+ t.Helper()
+ if os.Getenv(perfEnv) != "1" {
+ t.Skipf("set %s=1 to run performance scenarios", perfEnv)
+ }
+}
+
+// ---------------------------------------------------------------------------
+// Fixtures
+// ---------------------------------------------------------------------------
+
+// perfFixtures are the image layouts the scaling scenarios sweep over, in
+// increasing cost. Each is an OCI layout under ../mockhold/testdata/blobs,
+// which is what `skopeo copy docker://[ oci::img` writes:
+//
+// cd scanner/internal/mockhold/testdata/blobs
+// skopeo copy docker://docker.io/library/alpine:3.20 oci:perf-alpine:img
+// skopeo copy docker://docker.io/library/python:3.12-slim oci:perf-python:img
+// skopeo copy docker://docker.io/library/node:22 oci:perf-node:img
+//
+// hsm-secrets-operator comes from testdata/fetch-blobs.sh and is the one
+// fixture whose descriptors are also in corpus.json. The others are pulled
+// straight from Docker Hub because the corpus's larger images live on a
+// private hold: the scan job for them is reconstructed from the layout itself
+// by jobFromLayout, so no corpus entry is needed.
+//
+// perf-loom19 is optional: a 19-layer, ~37 MB corpus image that would separate
+// layer count from total bytes on real data. Pulling it with the authfile
+// dance fetch-blobs.sh performs (corpus digest
+// sha256:30d7f33c7f15ff3c6a1e4302575dcebfe31f3e2403486a2902a6bf30d44c7cdd)
+// failed here with "blob unknown to registry" — the hold no longer has one of
+// its layers — so the layer-count sweep in TestPerfSyntheticScaling covers
+// that axis synthetically instead. The name is kept so the fixture drops in if
+// a pullable equivalent turns up.
+var perfFixtures = []string{
+ "hsm-secrets-operator",
+ "perf-alpine",
+ "perf-python",
+ "perf-loom19",
+ "perf-node",
+}
+
+func fixtureDir(name string) string {
+ return filepath.Join("..", "mockhold", "testdata", "blobs", name)
+}
+
+func hasFixture(name string) bool {
+ _, err := os.Stat(filepath.Join(fixtureDir(name), "oci-layout"))
+ return err == nil
+}
+
+// availableFixtures returns the subset of perfFixtures present on disk.
+func availableFixtures(tb testing.TB) []string {
+ tb.Helper()
+ var out []string
+ for _, f := range perfFixtures {
+ if hasFixture(f) {
+ out = append(out, f)
+ }
+ }
+ if len(out) == 0 {
+ tb.Skip("no image fixtures present; see the perfFixtures comment for the skopeo commands")
+ }
+ return out
+}
+
+// jobFromLayout reconstructs the scan job for the single image in an OCI
+// layout: the descriptors the hold would have sent, read back out of the
+// bytes skopeo wrote. This is what lets any pulled image act as a fixture
+// without a matching corpus record.
+func jobFromLayout(tb testing.TB, dir string) *scanner.ScanJob {
+ tb.Helper()
+
+ var index struct {
+ Manifests []struct {
+ MediaType string `json:"mediaType"`
+ Digest string `json:"digest"`
+ } `json:"manifests"`
+ }
+ readJSON(tb, filepath.Join(dir, "index.json"), &index)
+ if len(index.Manifests) == 0 {
+ tb.Fatalf("%s: index.json declares no manifests", dir)
+ }
+
+ digest := index.Manifests[0].Digest
+ var manifest struct {
+ MediaType string `json:"mediaType"`
+ Config scanner.BlobDescriptor `json:"config"`
+ Layers []scanner.BlobDescriptor `json:"layers"`
+ Manifests []struct {
+ Digest string `json:"digest"`
+ } `json:"manifests"`
+ }
+ readJSON(tb, blobPath(dir, digest), &manifest)
+
+ // A layout pulled without --all still occasionally carries an index at the
+ // top: follow one hop into the first child manifest.
+ if manifest.Config.Digest == "" && len(manifest.Manifests) > 0 {
+ readJSON(tb, blobPath(dir, manifest.Manifests[0].Digest), &manifest)
+ }
+ if manifest.Config.Digest == "" {
+ tb.Fatalf("%s: no image manifest found in layout", dir)
+ }
+
+ return &scanner.ScanJob{
+ ManifestDigest: digest,
+ Repository: filepath.Base(dir),
+ Tag: "img",
+ Tier: "deckhand",
+ Config: manifest.Config,
+ Layers: manifest.Layers,
+ }
+}
+
+func blobPath(dir, digest string) string {
+ return filepath.Join(dir, "blobs", "sha256", mockhold.DigestHex(digest))
+}
+
+func readJSON(tb testing.TB, path string, v any) {
+ tb.Helper()
+ data, err := os.ReadFile(path)
+ if err != nil {
+ tb.Fatalf("read %s: %v", path, err)
+ }
+ if err := json.Unmarshal(data, v); err != nil {
+ tb.Fatalf("parse %s: %v", path, err)
+ }
+}
+
+// compressedBytes is the size the scanner's own MaxImageSize check would see.
+func compressedBytes(job *scanner.ScanJob) int64 {
+ total := job.Config.Size
+ for _, l := range job.Layers {
+ total += l.Size
+ }
+ return total
+}
+
+// ---------------------------------------------------------------------------
+// Pipeline replica
+// ---------------------------------------------------------------------------
+
+// stageTimes is one pass of the pipeline, broken into the four phases that
+// production runs back to back inside processJob.
+type stageTimes struct {
+ Download time.Duration // presign + fetch every blob, write the layout
+ Load time.Duration // stereoscope: read the layout, extract layers
+ Catalog time.Duration // syft.CreateSBOM
+ Encode time.Duration // SPDX JSON encode
+ Total time.Duration
+
+ Packages int
+ SBOMSize int
+ TmpPeak int64 // peak bytes under TMPDIR while this pass ran
+}
+
+func (s stageTimes) String() string {
+ return fmt.Sprintf("total=%s download=%s load=%s catalog=%s encode=%s packages=%d sbom=%dKiB tmpPeak=%.1fMiB",
+ round(s.Total), round(s.Download), round(s.Load), round(s.Catalog), round(s.Encode),
+ s.Packages, s.SBOMSize/1024, float64(s.TmpPeak)/(1<<20))
+}
+
+func round(d time.Duration) time.Duration { return d.Round(time.Millisecond) }
+
+// pipelineOnce runs one scan the way processJob does and reports what each
+// phase cost. See the file comment for why this is a replica.
+//
+// tmpDir stands in for cfg.Vuln.TmpDir: the layout is assembled there and
+// TMPDIR points at it, so stereoscope's extraction lands there too and a
+// single directory walk measures the whole disk footprint of a scan.
+//
+// It returns an error rather than calling Fatalf so the concurrency scenario
+// can run it from several goroutines at once.
+func pipelineOnce(job *scanner.ScanJob, tmpDir string) (st stageTimes, err error) {
+ watcher := watchDir(tmpDir, 20*time.Millisecond)
+ defer func() { st.TmpPeak = watcher.stop() }()
+
+ start := time.Now()
+
+ layoutDir, cleanup, aerr := assembleLayout(job, tmpDir)
+ if aerr != nil {
+ return st, fmt.Errorf("assemble layout: %w", aerr)
+ }
+ defer cleanup()
+ st.Download = time.Since(start)
+
+ ctx := context.Background()
+ t0 := time.Now()
+ tmpGen := file.NewTempDirGenerator("syft-scan")
+ defer tmpGen.Cleanup()
+
+ img, err := oci.NewDirectoryProvider(tmpGen, layoutDir).Provide(ctx)
+ if err != nil {
+ return st, fmt.Errorf("provide image: %w", err)
+ }
+ if err := img.Read(); err != nil {
+ img.Cleanup()
+ return st, fmt.Errorf("read image: %w", err)
+ }
+ src := stereoscopesource.New(img, stereoscopesource.ImageConfig{Reference: layoutDir})
+ defer src.Close()
+ st.Load = time.Since(t0)
+
+ t0 = time.Now()
+ sbomResult, err := syft.CreateSBOM(ctx, src, nil)
+ if err != nil {
+ return st, fmt.Errorf("create sbom: %w", err)
+ }
+ st.Catalog = time.Since(t0)
+ st.Packages = sbomResult.Artifacts.Packages.PackageCount()
+
+ t0 = time.Now()
+ encoder, err := spdxjson.NewFormatEncoderWithConfig(spdxjson.DefaultEncoderConfig())
+ if err != nil {
+ return st, fmt.Errorf("encoder: %w", err)
+ }
+ sbomJSON, err := format.Encode(*sbomResult, encoder)
+ if err != nil {
+ return st, fmt.Errorf("encode sbom: %w", err)
+ }
+ _ = sha256.Sum256(sbomJSON)
+ st.Encode = time.Since(t0)
+ st.SBOMSize = len(sbomJSON)
+
+ st.Total = time.Since(start)
+ return st, nil
+}
+
+// mustPipeline runs pipelineOnce and fails the test on error.
+func mustPipeline(tb testing.TB, job *scanner.ScanJob, tmpDir string) stageTimes {
+ tb.Helper()
+ st, err := pipelineOnce(job, tmpDir)
+ if err != nil {
+ tb.Fatalf("pipeline: %v", err)
+ }
+ return st
+}
+
+// assembleLayout mirrors buildOCILayout: download the config and every tar
+// layer through the hold's presign indirection, then write the manifest,
+// index.json and oci-layout beside them.
+func assembleLayout(job *scanner.ScanJob, tmpDir string) (string, func(), error) {
+ scanDir, err := os.MkdirTemp(tmpDir, "bench-scan-*")
+ if err != nil {
+ return "", nil, err
+ }
+ cleanup := func() { os.RemoveAll(scanDir) }
+
+ blobsDir := filepath.Join(scanDir, "blobs", "sha256")
+ if err := os.MkdirAll(blobsDir, 0o755); err != nil {
+ cleanup()
+ return "", nil, err
+ }
+
+ // Mirrors buildOCILayout's boundary too: parse the digest once, then let
+ // the same value name the blob on the wire and the file on disk.
+ fetch := func(digest string) error {
+ parsed, err := scanner.ParseDigest(digest)
+ if err != nil {
+ return err
+ }
+ url, err := client.GetBlobPresignedURL(job.HoldEndpoint, job.HoldDID, parsed, "")
+ if err != nil {
+ return err
+ }
+ _, err = client.DownloadBlob(url, filepath.Join(blobsDir, parsed.Hex), client.BlobExpectation{
+ Digest: parsed,
+ MaxBytes: -1,
+ })
+ return err
+ }
+
+ if err := fetch(job.Config.Digest); err != nil {
+ cleanup()
+ return "", nil, fmt.Errorf("config blob: %w", err)
+ }
+
+ type desc struct {
+ MediaType string `json:"mediaType"`
+ Digest string `json:"digest"`
+ Size int64 `json:"size"`
+ }
+ layers := make([]desc, 0, len(job.Layers))
+ for i, l := range job.Layers {
+ if l.Digest == "" || (l.MediaType != "" && !strings.Contains(l.MediaType, "tar")) {
+ continue
+ }
+ if err := fetch(l.Digest); err != nil {
+ cleanup()
+ return "", nil, fmt.Errorf("layer %d: %w", i, err)
+ }
+ mt := l.MediaType
+ if mt == "" {
+ mt = "application/vnd.oci.image.layer.v1.tar+gzip"
+ }
+ layers = append(layers, desc{MediaType: mt, Digest: l.Digest, Size: l.Size})
+ }
+
+ cfgType := job.Config.MediaType
+ if cfgType == "" {
+ cfgType = "application/vnd.oci.image.config.v1+json"
+ }
+ manifestJSON, _ := json.Marshal(map[string]any{
+ "schemaVersion": 2,
+ "mediaType": "application/vnd.oci.image.manifest.v1+json",
+ "config": desc{MediaType: cfgType, Digest: job.Config.Digest, Size: job.Config.Size},
+ "layers": layers,
+ })
+ sum := sha256.Sum256(manifestJSON)
+ if err := os.WriteFile(filepath.Join(blobsDir, fmt.Sprintf("%x", sum)), manifestJSON, 0o644); err != nil {
+ cleanup()
+ return "", nil, err
+ }
+ indexJSON, _ := json.Marshal(map[string]any{
+ "schemaVersion": 2,
+ "manifests": []desc{{
+ MediaType: "application/vnd.oci.image.manifest.v1+json",
+ Digest: fmt.Sprintf("sha256:%x", sum),
+ Size: int64(len(manifestJSON)),
+ }},
+ })
+ if err := os.WriteFile(filepath.Join(scanDir, "index.json"), indexJSON, 0o644); err != nil {
+ cleanup()
+ return "", nil, err
+ }
+ if err := os.WriteFile(filepath.Join(scanDir, "oci-layout"),
+ []byte(`{"imageLayoutVersion":"1.0.0"}`), 0o644); err != nil {
+ cleanup()
+ return "", nil, err
+ }
+ return scanDir, cleanup, nil
+}
+
+// perfHold starts a mock hold serving a fixture layout and points the job at
+// it, and redirects TMPDIR at a scratch directory the way WorkerPool.Start
+// redirects it at cfg.Vuln.TmpDir. It returns that directory.
+func perfHold(tb testing.TB, source mockhold.BlobSource, job *scanner.ScanJob) string {
+ tb.Helper()
+ hold := mockhold.New(source)
+ tb.Cleanup(hold.Close)
+ job.HoldEndpoint = hold.URL()
+
+ tmp := tb.TempDir()
+ setEnv(tb, "TMPDIR", tmp)
+ return tmp
+}
+
+// setEnv is t.Setenv, spelled out because testing.TB does not carry it and
+// several of these scenarios are benchmarks.
+func setEnv(tb testing.TB, key, value string) {
+ prev, had := os.LookupEnv(key)
+ os.Setenv(key, value)
+ tb.Cleanup(func() {
+ if had {
+ os.Setenv(key, prev)
+ return
+ }
+ os.Unsetenv(key)
+ })
+}
+
+// ---------------------------------------------------------------------------
+// Process-level sampling
+// ---------------------------------------------------------------------------
+
+// procSample is one observation of the whole process, which is what matters
+// here: Go's heap accounting misses the mmap'd regions stereoscope and
+// SQLite bring in, and RSS is what the container's memory limit counts.
+type procSample struct {
+ At time.Time
+ RSS int64 // bytes, /proc/self/statm
+ HeapAlloc int64
+ HeapSys int64
+ Goroutines int
+ FDs int
+}
+
+type monitor struct {
+ stop chan struct{}
+ done chan struct{}
+ mu sync.Mutex
+ samples []procSample
+ interval time.Duration
+}
+
+func startMonitor(interval time.Duration) *monitor {
+ m := &monitor{stop: make(chan struct{}), done: make(chan struct{}), interval: interval}
+ go func() {
+ defer close(m.done)
+ tick := time.NewTicker(interval)
+ defer tick.Stop()
+ m.record()
+ for {
+ select {
+ case <-m.stop:
+ m.record()
+ return
+ case <-tick.C:
+ m.record()
+ }
+ }
+ }()
+ return m
+}
+
+func (m *monitor) record() {
+ var ms runtime.MemStats
+ runtime.ReadMemStats(&ms)
+ s := procSample{
+ At: time.Now(),
+ RSS: rssBytes(),
+ HeapAlloc: int64(ms.HeapAlloc),
+ HeapSys: int64(ms.HeapSys),
+ Goroutines: runtime.NumGoroutine(),
+ FDs: fdCount(),
+ }
+ m.mu.Lock()
+ m.samples = append(m.samples, s)
+ m.mu.Unlock()
+}
+
+func (m *monitor) finish() []procSample {
+ close(m.stop)
+ <-m.done
+ m.mu.Lock()
+ defer m.mu.Unlock()
+ return append([]procSample(nil), m.samples...)
+}
+
+func peakRSS(samples []procSample) int64 {
+ var max int64
+ for _, s := range samples {
+ if s.RSS > max {
+ max = s.RSS
+ }
+ }
+ return max
+}
+
+func peakHeap(samples []procSample) int64 {
+ var max int64
+ for _, s := range samples {
+ if s.HeapAlloc > max {
+ max = s.HeapAlloc
+ }
+ }
+ return max
+}
+
+// settle forces the flattest heap this process can be talked into, so a
+// measurement is not reading the previous scenario's garbage. It is what makes
+// the peak numbers below comparable across iterations of a sweep.
+func settle() {
+ runtime.GC()
+ runtime.GC()
+ time.Sleep(100 * time.Millisecond)
+}
+
+// growth reports the peak RSS and heap of a run measured from its own first
+// sample, which is the part attributable to the work rather than to whatever
+// the process was already holding.
+func growth(samples []procSample) (rss, heap int64) {
+ if len(samples) == 0 {
+ return 0, 0
+ }
+ return peakRSS(samples) - samples[0].RSS, peakHeap(samples) - samples[0].HeapAlloc
+}
+
+// rssBytes reads resident set size from /proc/self/statm. Returns 0 where
+// procfs is unavailable, which is the honest answer rather than a guess.
+func rssBytes() int64 {
+ data, err := os.ReadFile("/proc/self/statm")
+ if err != nil {
+ return 0
+ }
+ fields := strings.Fields(string(data))
+ if len(fields) < 2 {
+ return 0
+ }
+ pages, err := strconv.ParseInt(fields[1], 10, 64)
+ if err != nil {
+ return 0
+ }
+ return pages * int64(os.Getpagesize())
+}
+
+// vmHWM is the kernel's own peak-RSS watermark for the process, which no
+// sampling interval can miss. It never decreases, so it is only meaningful as
+// "the highest this process ever reached", not per-scenario.
+func vmHWM() int64 {
+ f, err := os.Open("/proc/self/status")
+ if err != nil {
+ return 0
+ }
+ defer f.Close()
+ sc := bufio.NewScanner(f)
+ for sc.Scan() {
+ if !strings.HasPrefix(sc.Text(), "VmHWM:") {
+ continue
+ }
+ fields := strings.Fields(sc.Text())
+ if len(fields) < 2 {
+ return 0
+ }
+ kb, err := strconv.ParseInt(fields[1], 10, 64)
+ if err != nil {
+ return 0
+ }
+ return kb * 1024
+ }
+ return 0
+}
+
+func fdCount() int {
+ entries, err := os.ReadDir("/proc/self/fd")
+ if err != nil {
+ return -1
+ }
+ return len(entries)
+}
+
+// dirBytes sums the apparent size of every regular file under root.
+func dirBytes(root string) int64 {
+ var total int64
+ _ = filepath.WalkDir(root, func(_ string, d fs.DirEntry, err error) error {
+ if err != nil || d.IsDir() {
+ return nil //nolint:nilerr // a file vanishing mid-walk is expected
+ }
+ if info, err := d.Info(); err == nil {
+ total += info.Size()
+ }
+ return nil
+ })
+ return total
+}
+
+type dirWatcher struct {
+ stopCh chan struct{}
+ done chan struct{}
+ peak int64
+}
+
+// watchDir samples the on-disk footprint of a tree until stopped, and reports
+// the high-water mark. Sampling can undershoot a short spike; the ratios it is
+// used for are large enough that this does not change the conclusion.
+func watchDir(root string, interval time.Duration) *dirWatcher {
+ w := &dirWatcher{stopCh: make(chan struct{}), done: make(chan struct{})}
+ go func() {
+ defer close(w.done)
+ tick := time.NewTicker(interval)
+ defer tick.Stop()
+ for {
+ select {
+ case <-w.stopCh:
+ return
+ case <-tick.C:
+ if n := dirBytes(root); n > w.peak {
+ w.peak = n
+ }
+ }
+ }
+ }()
+ return w
+}
+
+func (w *dirWatcher) stop() int64 {
+ close(w.stopCh)
+ <-w.done
+ return w.peak
+}
+
+// ---------------------------------------------------------------------------
+// 1. Per-stage cost on real images
+// ---------------------------------------------------------------------------
+
+// TestPerfStageBreakdown times each phase of the pipeline for every fixture
+// present, which is the only way to see which one actually dominates. It runs
+// each image three times and reports every pass rather than an average: on a
+// shared machine the spread between passes is the honest error bar, and the
+// first pass also carries the page-cache cost of reading the fixture off disk.
+func TestPerfStageBreakdown(t *testing.T) {
+ requirePerf(t)
+
+ for _, name := range availableFixtures(t) {
+ t.Run(name, func(t *testing.T) {
+ dir := fixtureDir(name)
+ job := jobFromLayout(t, dir)
+ tmp := perfHold(t, mockhold.NewOCILayout(dir), job)
+
+ t.Logf("%s: %d layers, %.1f MiB compressed", name,
+ len(job.Layers), float64(compressedBytes(job))/(1<<20))
+
+ for i := 0; i < 3; i++ {
+ settle()
+ m := startMonitor(25 * time.Millisecond)
+ st := mustPipeline(t, job, tmp)
+ samples := m.finish()
+ dRSS, dHeap := growth(samples)
+ t.Logf("pass %d: %s peakRSS=%.0fMiB(+%.0f) peakHeap=%.0fMiB(+%.0f)",
+ i, st, float64(peakRSS(samples))/(1<<20), float64(dRSS)/(1<<20),
+ float64(peakHeap(samples))/(1<<20), float64(dHeap)/(1<<20))
+ }
+ })
+ }
+}
+
+// BenchmarkPipeline is the same work under the benchmark harness, for when a
+// stable per-op number matters more than the stage split.
+func BenchmarkPipeline(b *testing.B) {
+ for _, name := range availableFixtures(b) {
+ b.Run(name, func(b *testing.B) {
+ dir := fixtureDir(name)
+ job := jobFromLayout(b, dir)
+ tmp := perfHold(b, mockhold.NewOCILayout(dir), job)
+
+ b.ReportAllocs()
+ b.ResetTimer()
+ for i := 0; i < b.N; i++ {
+ if _, err := pipelineOnce(job, tmp); err != nil {
+ b.Fatal(err)
+ }
+ }
+ b.StopTimer()
+ // RSS at the end of the run, not a peak: a benchmark loop offers no
+ // quiet moment to sample one, and the peak is what
+ // TestPerfStageBreakdown reports.
+ b.ReportMetric(float64(rssBytes())/(1<<20), "endRSS_MiB")
+ })
+ }
+}
+
+// ---------------------------------------------------------------------------
+// 2. Sustained load, and whether the GC-plus-cooldown pause earns its keep
+// ---------------------------------------------------------------------------
+
+// cooldownMode selects what TestPerfSustainedLoad does between jobs.
+//
+// The two modes must be compared across processes, not within one: whichever
+// runs second inherits the other's already-grown heap and page-cache state, so
+// a single process cannot answer "does memory return to baseline" for both.
+// Run it twice, once per mode:
+//
+// ATCR_SCANNER_PERF=1 ATCR_SCANNER_PERF_COOLDOWN=prod go test ./internal/e2e/ -run TestPerfSustainedLoad -v -timeout 30m
+// ATCR_SCANNER_PERF=1 ATCR_SCANNER_PERF_COOLDOWN=none go test ./internal/e2e/ -run TestPerfSustainedLoad -v -timeout 30m
+const cooldownEnv = "ATCR_SCANNER_PERF_COOLDOWN"
+
+// TestPerfSustainedLoad runs the same image repeatedly and records RSS and
+// heap after every job, with and without the production `runtime.GC()` plus
+// ten second pause. It answers two separate questions that are easy to
+// conflate: whether memory ratchets across jobs, and whether the pause is what
+// stops it.
+func TestPerfSustainedLoad(t *testing.T) {
+ requirePerf(t)
+
+ mode := os.Getenv(cooldownEnv)
+ if mode == "" {
+ mode = "prod"
+ }
+ if mode != "prod" && mode != "none" {
+ t.Fatalf("%s must be prod or none, got %q", cooldownEnv, mode)
+ }
+
+ name := heaviestFixture(t)
+ dir := fixtureDir(name)
+ job := jobFromLayout(t, dir)
+ tmp := perfHold(t, mockhold.NewOCILayout(dir), job)
+
+ const jobs = 6
+ m := startMonitor(25 * time.Millisecond)
+ baselineRSS := rssBytes()
+ start := time.Now()
+ var work time.Duration
+
+ t.Logf("fixture=%s mode=%s jobs=%d baselineRSS=%.0fMiB", name, mode, jobs,
+ float64(baselineRSS)/(1<<20))
+
+ for i := 0; i < jobs; i++ {
+ t0 := time.Now()
+ st := mustPipeline(t, job, tmp)
+ work += time.Since(t0)
+
+ var before runtime.MemStats
+ runtime.ReadMemStats(&before)
+ rssBefore := rssBytes()
+
+ if mode == "prod" {
+ // Exactly what worker.go does between jobs.
+ runtime.GC()
+ time.Sleep(10 * time.Second)
+ }
+
+ var after runtime.MemStats
+ runtime.ReadMemStats(&after)
+ t.Logf("job %d: %s | RSS %.0f→%.0fMiB heap %.0f→%.0fMiB goroutines=%d fds=%d",
+ i, round(st.Total),
+ float64(rssBefore)/(1<<20), float64(rssBytes())/(1<<20),
+ float64(before.HeapAlloc)/(1<<20), float64(after.HeapAlloc)/(1<<20),
+ runtime.NumGoroutine(), fdCount())
+ }
+
+ wall := time.Since(start)
+ samples := m.finish()
+ t.Logf("RESULT mode=%s jobs=%d wall=%s work=%s throughput=%.2f jobs/min peakRSS=%.0fMiB peakHeap=%.0fMiB endRSS=%.0fMiB vmHWM=%.0fMiB",
+ mode, jobs, round(wall), round(work), float64(jobs)/wall.Minutes(),
+ float64(peakRSS(samples))/(1<<20), float64(peakHeap(samples))/(1<<20),
+ float64(rssBytes())/(1<<20), float64(vmHWM())/(1<<20))
+}
+
+// heaviestFixture picks the largest fixture present, since memory behaviour
+// only shows up on an image big enough to allocate.
+func heaviestFixture(tb testing.TB) string {
+ tb.Helper()
+ available := availableFixtures(tb)
+ best, bestSize := "", int64(-1)
+ for _, name := range available {
+ if n := dirBytes(fixtureDir(name)); n > bestSize {
+ best, bestSize = name, n
+ }
+ }
+ return best
+}
+
+// ---------------------------------------------------------------------------
+// 3. Scaling with layer count and with bytes
+// ---------------------------------------------------------------------------
+
+// syntheticImage builds an image out of nothing: layers layers, each holding
+// files of fileSize bytes, with content chosen by fill.
+//
+// Deterministic pseudo-random content ("random") is close to incompressible,
+// so compressed and extracted sizes stay within a few percent of each other
+// and the sweep measures bytes rather than the gzip ratio. Zero content
+// ("zeros") is the opposite extreme, and is what the disk-amplification
+// scenario uses.
+func syntheticImage(tb testing.TB, layers, filesPerLayer int, fileSize int64, fill string) (*scanner.ScanJob, *mockhold.Memory) {
+ tb.Helper()
+
+ mem := mockhold.NewMemory()
+ job := &scanner.ScanJob{
+ ManifestDigest: "sha256:" + strings.Repeat("0", 64),
+ Repository: "synthetic",
+ Tag: "img",
+ Tier: "deckhand",
+ }
+
+ var diffIDs []string
+ for i := 0; i < layers; i++ {
+ raw := buildTar(tb, i, filesPerLayer, fileSize, fill)
+ diffIDs = append(diffIDs, fmt.Sprintf("sha256:%x", sha256.Sum256(raw)))
+
+ var gz bytes.Buffer
+ zw, err := gzip.NewWriterLevel(&gz, gzip.BestSpeed)
+ if err != nil {
+ tb.Fatalf("gzip writer: %v", err)
+ }
+ if _, err := zw.Write(raw); err != nil {
+ tb.Fatalf("gzip write: %v", err)
+ }
+ if err := zw.Close(); err != nil {
+ tb.Fatalf("gzip close: %v", err)
+ }
+
+ data := gz.Bytes()
+ digest := fmt.Sprintf("sha256:%x", sha256.Sum256(data))
+ mem.Add(digest, append([]byte(nil), data...))
+ job.Layers = append(job.Layers, scanner.BlobDescriptor{
+ Digest: digest,
+ Size: int64(len(data)),
+ MediaType: "application/vnd.oci.image.layer.v1.tar+gzip",
+ })
+ }
+
+ cfg, err := json.Marshal(map[string]any{
+ "architecture": "amd64",
+ "os": "linux",
+ "config": map[string]any{},
+ "rootfs": map[string]any{"type": "layers", "diff_ids": diffIDs},
+ })
+ if err != nil {
+ tb.Fatalf("marshal config: %v", err)
+ }
+ cfgDigest := fmt.Sprintf("sha256:%x", sha256.Sum256(cfg))
+ mem.Add(cfgDigest, cfg)
+ job.Config = scanner.BlobDescriptor{
+ Digest: cfgDigest,
+ Size: int64(len(cfg)),
+ MediaType: "application/vnd.oci.image.config.v1+json",
+ }
+
+ return job, mem
+}
+
+func buildTar(tb testing.TB, layer, files int, size int64, fill string) []byte {
+ tb.Helper()
+ var buf bytes.Buffer
+ tw := tar.NewWriter(&buf)
+ for f := 0; f < files; f++ {
+ content := fillBytes(size, fill, int64(layer*1000+f))
+ hdr := &tar.Header{
+ Name: fmt.Sprintf("layer%02d/file%04d.bin", layer, f),
+ Mode: 0o644,
+ Size: int64(len(content)),
+ Typeflag: tar.TypeReg,
+ ModTime: time.Unix(0, 0),
+ }
+ if err := tw.WriteHeader(hdr); err != nil {
+ tb.Fatalf("tar header: %v", err)
+ }
+ if _, err := tw.Write(content); err != nil {
+ tb.Fatalf("tar write: %v", err)
+ }
+ }
+ if err := tw.Close(); err != nil {
+ tb.Fatalf("tar close: %v", err)
+ }
+ return buf.Bytes()
+}
+
+// fillBytes generates content without importing math/rand: a xorshift over a
+// seed is enough to defeat gzip while staying reproducible.
+func fillBytes(n int64, fill string, seed int64) []byte {
+ out := make([]byte, n)
+ if fill == "zeros" {
+ return out
+ }
+ x := uint64(seed*2654435761 + 12345)
+ for i := range out {
+ x ^= x << 13
+ x ^= x >> 7
+ x ^= x << 17
+ out[i] = byte(x)
+ }
+ return out
+}
+
+// TestPerfSyntheticScaling sweeps layer count at fixed total bytes, then total
+// bytes at fixed layer count. Splitting them is the point: the two are
+// confounded in any real corpus, where more layers usually also means more
+// bytes.
+func TestPerfSyntheticScaling(t *testing.T) {
+ requirePerf(t)
+
+ const totalMiB = 64
+
+ t.Run("layers_at_fixed_bytes", func(t *testing.T) {
+ for _, layers := range []int{1, 2, 4, 8, 16, 32} {
+ perLayer := int64(totalMiB<<20) / int64(layers)
+ job, mem := syntheticImage(t, layers, 4, perLayer/4, "random")
+ tmp := perfHold(t, mem, job)
+
+ settle()
+ m := startMonitor(25 * time.Millisecond)
+ st := mustPipeline(t, job, tmp)
+ samples := m.finish()
+ dRSS, dHeap := growth(samples)
+ t.Logf("layers=%2d compressed=%.1fMiB %s ΔRSS=%.0fMiB ΔHeap=%.0fMiB peakRSS=%.0fMiB",
+ layers, float64(compressedBytes(job))/(1<<20), st,
+ float64(dRSS)/(1<<20), float64(dHeap)/(1<<20), float64(peakRSS(samples))/(1<<20))
+ }
+ })
+
+ t.Run("bytes_at_fixed_layers", func(t *testing.T) {
+ for _, mib := range []int64{8, 32, 128, 512} {
+ job, mem := syntheticImage(t, 4, 4, (mib<<20)/16, "random")
+ tmp := perfHold(t, mem, job)
+
+ settle()
+ m := startMonitor(25 * time.Millisecond)
+ st := mustPipeline(t, job, tmp)
+ samples := m.finish()
+ dRSS, dHeap := growth(samples)
+ t.Logf("bytes=%4dMiB compressed=%.1fMiB %s ΔRSS=%.0fMiB ΔHeap=%.0fMiB peakRSS=%.0fMiB",
+ mib, float64(compressedBytes(job))/(1<<20), st,
+ float64(dRSS)/(1<<20), float64(dHeap)/(1<<20), float64(peakRSS(samples))/(1<<20))
+ }
+ })
+
+ t.Run("files_at_fixed_bytes", func(t *testing.T) {
+ // Entry count, not byte count, is what a filesystem catalog walks.
+ for _, files := range []int{16, 256, 4096} {
+ job, mem := syntheticImage(t, 4, files, (16<<20)/int64(files), "random")
+ tmp := perfHold(t, mem, job)
+
+ settle()
+ m := startMonitor(25 * time.Millisecond)
+ st := mustPipeline(t, job, tmp)
+ samples := m.finish()
+ dRSS, dHeap := growth(samples)
+ t.Logf("filesPerLayer=%5d (%d total) %s ΔRSS=%.0fMiB ΔHeap=%.0fMiB peakRSS=%.0fMiB",
+ files, files*4, st,
+ float64(dRSS)/(1<<20), float64(dHeap)/(1<<20), float64(peakRSS(samples))/(1<<20))
+ }
+ })
+}
+
+// ---------------------------------------------------------------------------
+// 4. Disk: what MaxImageSize does not bound
+// ---------------------------------------------------------------------------
+
+// TestPerfDiskAmplification measures the gap between the compressed bytes the
+// MaxImageSize guard checks and the bytes a scan actually puts on disk.
+// A layer of zeros is the extreme case, but the direction is the same for any
+// real layer, and the guard sees only the compressed figure.
+func TestPerfDiskAmplification(t *testing.T) {
+ requirePerf(t)
+
+ for _, mib := range []int64{64, 512} {
+ job, mem := syntheticImage(t, 1, 8, (mib<<20)/8, "zeros")
+ tmp := perfHold(t, mem, job)
+
+ compressed := compressedBytes(job)
+ st := mustPipeline(t, job, tmp)
+ leftover := dirBytes(tmp)
+
+ t.Logf("uncompressed=%dMiB compressed=%.2fMiB ratio=%.0fx peakTmp=%.1fMiB leftoverAfterScan=%dB %s",
+ mib, float64(compressed)/(1<<20), float64(mib<<20)/float64(compressed),
+ float64(st.TmpPeak)/(1<<20), leftover, st)
+ }
+}
+
+// TestPerfTempCleanup checks that nothing accumulates under the scan tmp dir
+// across a run of jobs, which is the failure mode a long-lived scanner would
+// hit long before it hit a memory limit.
+func TestPerfTempCleanup(t *testing.T) {
+ requirePerf(t)
+
+ name := availableFixtures(t)[0]
+ dir := fixtureDir(name)
+ job := jobFromLayout(t, dir)
+ tmp := perfHold(t, mockhold.NewOCILayout(dir), job)
+
+ for i := 0; i < 5; i++ {
+ mustPipeline(t, job, tmp)
+ entries, err := os.ReadDir(tmp)
+ if err != nil {
+ t.Fatalf("read tmp: %v", err)
+ }
+ t.Logf("after job %d: %d entries, %d bytes under %s", i, len(entries), dirBytes(tmp), tmp)
+ if len(entries) != 0 {
+ names := make([]string, 0, len(entries))
+ for _, e := range entries {
+ names = append(names, e.Name())
+ }
+ t.Errorf("job %d left %d entries behind: %v", i, len(entries), names)
+ }
+ }
+}
+
+// ---------------------------------------------------------------------------
+// 5. Concurrency
+// ---------------------------------------------------------------------------
+
+// TestPerfConcurrency runs the same image through 1, 2 and 4 concurrent
+// pipelines and reports throughput and peak RSS for each.
+//
+// This measures the work, not the worker pool: the pool cannot be driven to
+// completion on a successful scan (see the file comment), so what it shows is
+// the ceiling raising scanner.workers could reach if the pool itself adds no
+// contention of its own. The vulnerability-database lock, which the pool does
+// add once Grype is enabled, is modelled separately below.
+func TestPerfConcurrency(t *testing.T) {
+ requirePerf(t)
+
+ name := heaviestFixture(t)
+ dir := fixtureDir(name)
+ tmpl := jobFromLayout(t, dir)
+ tmp := perfHold(t, mockhold.NewOCILayout(dir), tmpl)
+
+ const perWorker = 2
+ for _, workers := range []int{1, 2, 4} {
+ settle()
+ m := startMonitor(25 * time.Millisecond)
+ start := time.Now()
+
+ var wg sync.WaitGroup
+ errs := make(chan error, workers*perWorker)
+ for w := 0; w < workers; w++ {
+ wg.Add(1)
+ go func() {
+ defer wg.Done()
+ job := *tmpl // each pipeline gets its own job value
+ for i := 0; i < perWorker; i++ {
+ if _, err := pipelineOnce(&job, tmp); err != nil {
+ errs <- err
+ return
+ }
+ }
+ }()
+ }
+ wg.Wait()
+ close(errs)
+ for err := range errs {
+ t.Fatalf("workers=%d: %v", workers, err)
+ }
+
+ wall := time.Since(start)
+ samples := m.finish()
+ total := workers * perWorker
+ dRSS, dHeap := growth(samples)
+ t.Logf("workers=%d scans=%d wall=%s perScan=%s throughput=%.2f scans/min peakRSS=%.0fMiB ΔRSS=%.0fMiB ΔHeap=%.0fMiB",
+ workers, total, round(wall), round(wall/time.Duration(total)),
+ float64(total)/wall.Minutes(), float64(peakRSS(samples))/(1<<20),
+ float64(dRSS)/(1<<20), float64(dHeap)/(1<<20))
+
+ runtime.GC()
+ }
+}
+
+// TestPerfVulnDBLockModel measures what the RWMutex discipline in grype.go
+// costs when a reload lands during steady-state scanning.
+//
+// It is a model, not the real code: loadVulnDB is unexported in package scan
+// and cannot be stubbed from here, and running the real thing would download
+// the database this file is forbidden to fetch. What the model reproduces is
+// exactly the structure at grype.go:105 and grype.go:187 — every scan holds
+// the read lock for the whole of FindMatches, a reload holds the write lock
+// for the whole of the download — and what it measures is the stall that
+// structure imposes on workers that are not reloading anything.
+func TestPerfVulnDBLockModel(t *testing.T) {
+ requirePerf(t)
+
+ const (
+ workers = 4
+ scanWork = 50 * time.Millisecond
+ reloadWork = 2 * time.Second // stands in for a database download
+ )
+
+ var lock sync.RWMutex
+ var stalls sync.Map // worker id -> longest wait for the read lock
+
+ stop := make(chan struct{})
+ var wg sync.WaitGroup
+
+ for w := 0; w < workers; w++ {
+ wg.Add(1)
+ go func(id int) {
+ defer wg.Done()
+ var worst time.Duration
+ for {
+ select {
+ case <-stop:
+ stalls.Store(id, worst)
+ return
+ default:
+ }
+ t0 := time.Now()
+ lock.RLock()
+ if wait := time.Since(t0); wait > worst {
+ worst = wait
+ }
+ time.Sleep(scanWork)
+ lock.RUnlock()
+ }
+ }(w)
+ }
+
+ time.Sleep(500 * time.Millisecond) // steady state
+ before := time.Now()
+ lock.Lock()
+ acquired := time.Since(before)
+ time.Sleep(reloadWork)
+ lock.Unlock()
+ time.Sleep(500 * time.Millisecond)
+
+ close(stop)
+ wg.Wait()
+
+ var worst time.Duration
+ stalls.Range(func(_, v any) bool {
+ if d := v.(time.Duration); d > worst {
+ worst = d
+ }
+ return true
+ })
+
+ t.Logf("model: %d workers, %s per scan, %s reload — writer waited %s to acquire, "+
+ "worst reader stall %s (%.1fx the reload)",
+ workers, scanWork, reloadWork, round(acquired), round(worst),
+ float64(worst)/float64(reloadWork))
+ t.Logf("in production the reload is a database download of several hundred MB, " +
+ "so the reader stall scales with that download, not with the model's 2s")
+}
+
+// ---------------------------------------------------------------------------
+// 6. The real worker pool: cadence, queue burst, reconnect
+// ---------------------------------------------------------------------------
+
+// failFastJob is a job the pipeline rejects before it opens a socket: no
+// config digest, no layers, so buildOCILayout fails on the first check. It is
+// the cheapest way to push many jobs through the real worker without paying
+// for a real image per job (see the file comment).
+func failFastJob(repo string) *scanner.ScanJob {
+ return &scanner.ScanJob{
+ ManifestDigest: "sha256:" + strings.Repeat("1", 64),
+ Repository: repo,
+ Tag: "latest",
+ Tier: "deckhand",
+ }
+}
+
+// TestPerfWorkerCadence measures the real inter-job gap at the production
+// JobCooldown. The jobs themselves do no work, so what is left is exactly the
+// per-job overhead the worker adds: the GC call plus the sleep.
+func TestPerfWorkerCadence(t *testing.T) {
+ requirePerf(t)
+
+ h := Start(t, mockhold.NewMemory())
+ // Restore the production value the harness shortens. Start already
+ // registered a cleanup that puts back whatever it found, and cleanups run
+ // last-in-first-out, so this needs no undo of its own.
+ scan.JobCooldown = 10 * time.Second
+
+ const jobs = 4
+ start := time.Now()
+ var seqs []int64
+ for i := 0; i < jobs; i++ {
+ seq, err := h.Hold.SendJob(failFastJob("cadence"))
+ if err != nil {
+ t.Fatalf("send job %d: %v", i, err)
+ }
+ seqs = append(seqs, seq)
+ }
+
+ var prev time.Time
+ for i, seq := range seqs {
+ msg := h.AwaitTerminal(t, seq, 2*time.Minute)
+ gap := time.Duration(0)
+ if i > 0 {
+ gap = msg.At.Sub(prev)
+ }
+ prev = msg.At
+ t.Logf("job %d terminal at +%s (gap %s) type=%s", i,
+ round(msg.At.Sub(start)), round(gap), msg.Type)
+ }
+
+ wall := time.Since(start)
+ t.Logf("RESULT %d no-op jobs took %s at the production cooldown: %s per job, "+
+ "%.1f jobs/min ceiling for a single worker independent of scan cost",
+ jobs, round(wall), round(wall/jobs), float64(jobs)/wall.Minutes())
+}
+
+// TestPerfQueueBurst fills the queue to its configured depth and measures how
+// long the last job in it waits, which is the number that has to be compared
+// against the hold's ten minute processing timeout: the scanner acks a job the
+// moment it arrives (client/hold.go:159), which moves the hold's row to
+// 'processing' without moving assigned_at, and reDispatchTimedOut fails any
+// processing row older than ten minutes
+// (pkg/hold/pds/scan_broadcaster.go:833). Queue time is spent against that
+// deadline.
+//
+// The first job is pointed at a server that never answers, which parks the
+// single worker while the burst arrives. Every burst job then fails instantly,
+// so what the last one's latency measures is queue position plus cooldown and
+// nothing else — a floor for the real thing, where each job ahead of it also
+// costs a scan.
+func TestPerfQueueBurst(t *testing.T) {
+ requirePerf(t)
+
+ release := make(chan struct{})
+ var releaseOnce sync.Once
+ stall := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
+ <-release
+ http.Error(w, "gone", http.StatusNotFound)
+ }))
+ defer stall.Close()
+ defer releaseOnce.Do(func() { close(release) })
+
+ h := Start(t, mockhold.NewMemory())
+ depth := h.Cfg.Scanner.QueueSize
+
+ // Park the worker on a presign request that never returns.
+ parked := failFastJob("parked")
+ parked.Config = scanner.BlobDescriptor{
+ Digest: "sha256:" + strings.Repeat("2", 64),
+ Size: 10,
+ MediaType: "application/vnd.oci.image.config.v1+json",
+ }
+ parked.HoldEndpoint = stall.URL
+ if _, err := h.Hold.SendJob(parked); err != nil {
+ t.Fatalf("send parking job: %v", err)
+ }
+ deadline := time.Now().Add(10 * time.Second)
+ for len(h.Hold.BlobRequests()) == 0 && h.Queue.Len() == 0 && time.Now().Before(deadline) {
+ time.Sleep(10 * time.Millisecond)
+ }
+ time.Sleep(200 * time.Millisecond)
+
+ // Exactly enough to fill the queue, then a few more that cannot fit.
+ const overflow = 10
+ sent := make([]int64, 0, depth+overflow)
+ sentAt := time.Now()
+ for i := 0; i < depth+overflow; i++ {
+ seq, err := h.Hold.SendJob(failFastJob(fmt.Sprintf("burst-%03d", i)))
+ if err != nil {
+ t.Fatalf("send burst job %d: %v", i, err)
+ }
+ sent = append(sent, seq)
+ }
+ dispatchTook := time.Since(sentAt)
+
+ // Wait for the client's read loop to drain the socket, then look at how
+ // deep the queue actually got.
+ maxDepth := 0
+ deadline = time.Now().Add(10 * time.Second)
+ for time.Now().Before(deadline) {
+ if n := h.Queue.Len(); n > maxDepth {
+ maxDepth = n
+ }
+ if maxDepth >= depth {
+ break
+ }
+ time.Sleep(5 * time.Millisecond)
+ }
+ t.Logf("dispatched %d jobs in %s; queue reached %d of %d",
+ len(sent), round(dispatchTook), maxDepth, depth)
+
+ // The read loop enqueues in arrival order with the worker parked, so the
+ // first `depth` jobs are the accepted ones and the rest are rejected.
+ lastAccepted := sent[depth-1]
+ releaseOnce.Do(func() { close(release) })
+
+ last := h.AwaitTerminal(t, lastAccepted, 5*time.Minute)
+ first := h.AwaitTerminal(t, sent[0], time.Minute)
+
+ var full int
+ for _, m := range h.Hold.Transcript() {
+ if m.Type == "error" && strings.Contains(m.Error, "queue full") {
+ full++
+ }
+ }
+
+ drain := last.At.Sub(first.At)
+ perJob := drain / time.Duration(depth-1)
+ t.Logf("RESULT queueDepth=%d rejectedAsQueueFull=%d firstTerminal=+%s lastTerminal=+%s drain=%s perJob=%s (cooldown=%s)",
+ depth, full, round(first.At.Sub(sentAt)), round(last.At.Sub(sentAt)),
+ round(drain), round(perJob), scan.JobCooldown)
+
+ overhead := perJob - scan.JobCooldown
+ t.Logf("EXTRAPOLATION per-job overhead outside the cooldown is %s; at the production "+
+ "10s cooldown a full %d-deep queue of no-op jobs drains in %s, and a queue of "+
+ "16s node:22 scans in %s — against the hold's 10m processing timeout, which is "+
+ "exceeded by queue position %d and %d respectively",
+ round(overhead), depth,
+ round(time.Duration(depth)*(overhead+10*time.Second)),
+ round(time.Duration(depth)*(overhead+10*time.Second+16*time.Second)),
+ int(10*time.Minute/(overhead+10*time.Second)),
+ int(10*time.Minute/(overhead+26*time.Second)))
+}
+
+// TestPerfConnectionChurn watches goroutines and file descriptors across
+// reconnects. The client redials on a fixed five second backoff, so each cycle
+// costs that much; the count is kept low deliberately.
+func TestPerfConnectionChurn(t *testing.T) {
+ requirePerf(t)
+
+ h := Start(t, mockhold.NewMemory())
+
+ settle := func() {
+ runtime.GC()
+ time.Sleep(200 * time.Millisecond)
+ }
+ settle()
+ baseGoroutines, baseFDs := runtime.NumGoroutine(), fdCount()
+ t.Logf("baseline: goroutines=%d fds=%d", baseGoroutines, baseFDs)
+
+ const cycles = 4
+ for i := 0; i < cycles; i++ {
+ mode := mockhold.DropAbrupt
+ if i%2 == 1 {
+ mode = mockhold.DropClean
+ }
+ h.Hold.DropConnections(mode)
+
+ deadline := time.Now().Add(30 * time.Second)
+ want := len(h.Hold.Dials()) + 1
+ for len(h.Hold.Dials()) < want && time.Now().Before(deadline) {
+ time.Sleep(100 * time.Millisecond)
+ }
+ if len(h.Hold.Dials()) < want {
+ t.Fatalf("cycle %d: scanner never redialled", i)
+ }
+ if _, err := h.Hold.SendJob(failFastJob("churn")); err != nil {
+ t.Fatalf("cycle %d: send job: %v", i, err)
+ }
+ settle()
+ t.Logf("cycle %d (%v): goroutines=%d fds=%d", i, mode, runtime.NumGoroutine(), fdCount())
+ }
+
+ settle()
+ t.Logf("RESULT after %d reconnects: goroutines %d→%d fds %d→%d",
+ cycles, baseGoroutines, runtime.NumGoroutine(), baseFDs, fdCount())
+}
+
+// TestPerfJobChurn pushes many no-op jobs through the real worker and watches
+// for goroutine or descriptor growth. It cannot cover the scanning path, which
+// is where a leak would most plausibly live; that limitation is the BLOCKER's,
+// not the scenario's.
+func TestPerfJobChurn(t *testing.T) {
+ requirePerf(t)
+
+ h := Start(t, mockhold.NewMemory())
+ runtime.GC()
+ time.Sleep(200 * time.Millisecond)
+ baseGoroutines, baseFDs := runtime.NumGoroutine(), fdCount()
+ var ms runtime.MemStats
+ runtime.ReadMemStats(&ms)
+ baseHeap := ms.HeapAlloc
+
+ const jobs = 200
+ var lastSeq int64
+ for i := 0; i < jobs; i++ {
+ seq, err := h.Hold.SendJob(failFastJob("churn"))
+ if err != nil {
+ t.Fatalf("send job %d: %v", i, err)
+ }
+ lastSeq = seq
+ if i%20 == 0 {
+ time.Sleep(50 * time.Millisecond) // stay inside the 100-deep queue
+ }
+ }
+ h.AwaitTerminal(t, lastSeq, 3*time.Minute)
+
+ runtime.GC()
+ time.Sleep(200 * time.Millisecond)
+ runtime.ReadMemStats(&ms)
+ t.Logf("RESULT after %d jobs: goroutines %d→%d fds %d→%d heap %.1f→%.1fMiB",
+ jobs, baseGoroutines, runtime.NumGoroutine(), baseFDs, fdCount(),
+ float64(baseHeap)/(1<<20), float64(ms.HeapAlloc)/(1<<20))
+}
+
+// ---------------------------------------------------------------------------
+// 7. The real worker on a real image, in a subprocess
+// ---------------------------------------------------------------------------
+
+// TestPerfRealWorkerScan measures one successful scan through the actual
+// WorkerPool in a child process. It was written when the worker panicked on
+// result.Summary.Total the instant a scan succeeded, so the child died with a
+// SIGSEGV that would have taken the whole test binary with it; the child is
+// kept because it also isolates the measurement from this process.
+//
+// The measurement is the child's own log line "Scan pipeline completed
+// duration=...", produced by production code, so it can be compared directly
+// against the replica's number to see what the worker adds.
+func TestPerfRealWorkerScan(t *testing.T) {
+ requirePerf(t)
+
+ name := heaviestFixture(t)
+ cmd := exec.Command(os.Args[0],
+ "-test.run", "^TestPerfRealWorkerChild$",
+ "-test.v",
+ "-test.timeout", "10m")
+ cmd.Env = append(os.Environ(),
+ perfEnv+"=1",
+ "ATCR_SCANNER_PERF_CHILD="+name)
+
+ start := time.Now()
+ out, err := cmd.CombinedOutput()
+ wall := time.Since(start)
+
+ var duration, panicked string
+ for _, line := range strings.Split(string(out), "\n") {
+ if strings.Contains(line, "Scan pipeline completed") {
+ duration = line
+ }
+ if strings.HasPrefix(line, "panic:") {
+ panicked = line
+ }
+ }
+
+ t.Logf("child fixture=%s wall=%s exit=%v", name, round(wall), err)
+ if duration != "" {
+ t.Logf("production log line: %s", strings.TrimSpace(duration))
+ }
+ if panicked != "" {
+ t.Errorf("child panicked: %s\noutput tail:\n%s",
+ strings.TrimSpace(panicked), tail(string(out), 20))
+ }
+}
+
+// TestPerfRealWorkerChild is the child half of TestPerfRealWorkerScan. It is
+// inert unless ATCR_SCANNER_PERF_CHILD names a fixture.
+func TestPerfRealWorkerChild(t *testing.T) {
+ name := os.Getenv("ATCR_SCANNER_PERF_CHILD")
+ if name == "" {
+ t.Skip("child process scenario; driven by TestPerfRealWorkerScan")
+ }
+ dir := fixtureDir(name)
+ if !hasFixture(name) {
+ t.Skipf("fixture %q absent", name)
+ }
+
+ job := jobFromLayout(t, dir)
+ h := Start(t, mockhold.NewOCILayout(dir))
+ seq, err := h.Hold.SendJob(job)
+ if err != nil {
+ t.Fatalf("send job: %v", err)
+ }
+ msg := h.AwaitTerminal(t, seq, 5*time.Minute)
+ t.Logf("terminal=%s", msg.Type)
+}
+
+func tail(s string, n int) string {
+ lines := strings.Split(strings.TrimRight(s, "\n"), "\n")
+ if len(lines) > n {
+ lines = lines[len(lines)-n:]
+ }
+ return strings.Join(lines, "\n")
+}
+
+// ---------------------------------------------------------------------------
+// 8. The production soft memory limit
+// ---------------------------------------------------------------------------
+
+// TestPerfMemoryLimit runs the same scan under the soft memory limit the
+// scanner sets on itself at cmd/scanner/main.go:46 and again with the limit
+// effectively removed, and reports what the limit costs.
+//
+// GOMEMLIMIT is a soft limit: the runtime does not fail an allocation that
+// crosses it, it runs the collector harder. When a scan's live heap is at or
+// above the limit there is nothing left to collect, so the collector runs
+// continuously against a heap it cannot shrink. GCCPUFraction is the number
+// that shows this happening.
+//
+// The limit is process-wide, so scanner.workers > 1 divides it between
+// concurrent scans rather than multiplying it.
+func TestPerfMemoryLimit(t *testing.T) {
+ requirePerf(t)
+
+ name := heaviestFixture(t)
+ dir := fixtureDir(name)
+ job := jobFromLayout(t, dir)
+ tmp := perfHold(t, mockhold.NewOCILayout(dir), job)
+
+ const prodLimit = 512 * 1024 * 1024 // cmd/scanner/main.go:46
+
+ run := func(label string, limit int64) {
+ prev := debug.SetMemoryLimit(limit)
+ defer debug.SetMemoryLimit(prev)
+
+ settle()
+ var before runtime.MemStats
+ runtime.ReadMemStats(&before)
+
+ m := startMonitor(25 * time.Millisecond)
+ st := mustPipeline(t, job, tmp)
+ samples := m.finish()
+
+ var after runtime.MemStats
+ runtime.ReadMemStats(&after)
+ dRSS, dHeap := growth(samples)
+ t.Logf("%s (limit=%s): %s peakRSS=%.0fMiB ΔRSS=%.0fMiB ΔHeap=%.0fMiB gcCycles=%d gcPause=%s gcCPU=%.1f%%",
+ label, limitLabel(limit), st,
+ float64(peakRSS(samples))/(1<<20), float64(dRSS)/(1<<20), float64(dHeap)/(1<<20),
+ after.NumGC-before.NumGC,
+ round(time.Duration(after.PauseTotalNs-before.PauseTotalNs)),
+ after.GCCPUFraction*100)
+ }
+
+ // The same again with two scans in flight, because the limit is
+ // process-wide: scanner.workers > 1 divides it rather than multiplying it,
+ // and the production template sets workers: 2.
+ runConcurrent := func(label string, limit int64, n int) {
+ prev := debug.SetMemoryLimit(limit)
+ defer debug.SetMemoryLimit(prev)
+
+ settle()
+ var before runtime.MemStats
+ runtime.ReadMemStats(&before)
+
+ m := startMonitor(25 * time.Millisecond)
+ start := time.Now()
+ var wg sync.WaitGroup
+ errs := make(chan error, n)
+ for i := 0; i < n; i++ {
+ wg.Add(1)
+ go func() {
+ defer wg.Done()
+ j := *job
+ if _, err := pipelineOnce(&j, tmp); err != nil {
+ errs <- err
+ }
+ }()
+ }
+ wg.Wait()
+ close(errs)
+ for err := range errs {
+ t.Fatalf("%s: %v", label, err)
+ }
+ wall := time.Since(start)
+ samples := m.finish()
+
+ var after runtime.MemStats
+ runtime.ReadMemStats(&after)
+ dRSS, dHeap := growth(samples)
+ t.Logf("%s x%d (limit=%s): wall=%s perScan=%s peakRSS=%.0fMiB ΔRSS=%.0fMiB ΔHeap=%.0fMiB gcCycles=%d gcPause=%s gcCPU=%.1f%%",
+ label, n, limitLabel(limit), round(wall), round(wall/time.Duration(n)),
+ float64(peakRSS(samples))/(1<<20), float64(dRSS)/(1<<20), float64(dHeap)/(1<<20),
+ after.NumGC-before.NumGC,
+ round(time.Duration(after.PauseTotalNs-before.PauseTotalNs)),
+ after.GCCPUFraction*100)
+ }
+
+ t.Logf("fixture=%s %d layers %.0fMiB compressed", name, len(job.Layers),
+ float64(compressedBytes(job))/(1<<20))
+ run("unlimited", math.MaxInt64)
+ run("production", prodLimit)
+ run("unlimited-again", math.MaxInt64) // ordering control
+ runConcurrent("unlimited", math.MaxInt64, 2)
+ runConcurrent("production", prodLimit, 2)
+ runConcurrent("unlimited", math.MaxInt64, 4)
+ runConcurrent("production", prodLimit, 4)
+}
+
+func limitLabel(limit int64) string {
+ if limit == math.MaxInt64 {
+ return "off"
+ }
+ return fmt.Sprintf("%dMiB", limit/(1<<20))
+}
diff --git a/scanner/internal/e2e/blob_edge_test.go b/scanner/internal/e2e/blob_edge_test.go
new file mode 100644
index 0000000..af03abe
--- /dev/null
+++ b/scanner/internal/e2e/blob_edge_test.go
@@ -0,0 +1,898 @@
+package e2e
+
+import (
+ "archive/tar"
+ "bytes"
+ "compress/gzip"
+ "context"
+ "crypto/sha256"
+ "encoding/json"
+ "fmt"
+ "math/rand"
+ "net/http"
+ "net/http/httptest"
+ "os"
+ "strings"
+ "sync/atomic"
+ "testing"
+ "time"
+
+ scanner "atcr.io/scanner"
+ "atcr.io/scanner/internal/client"
+ "atcr.io/scanner/internal/config"
+ "atcr.io/scanner/internal/mockhold"
+ "atcr.io/scanner/internal/queue"
+ "atcr.io/scanner/internal/scan"
+)
+
+// This file covers the blob download and OCI layout construction stages:
+// what the scanner does when a blob is gone, truncated, corrupt, slow, or
+// named by a digest a user made up.
+//
+// Historical note on the shape of these scenarios. worker.go used to
+// dereference result.Summary unconditionally, and Summary is only populated
+// when cfg.Vuln.Enabled; with Grype off in the harness, a *successful* scan
+// panicked and took the test binary with it, so no e2e test could observe
+// anything but a terminal "error" or "skipped". That is fixed, and
+// TestZeroSizeLayerProducesACleanScan now runs a success through to the end.
+// Scenarios below that would otherwise succeed (a duplicated layer, a lying
+// size) still break one blob deliberately, but now only to keep the assertion
+// on the download stage rather than on Syft's verdict.
+//
+// A second move since: the scenarios that pinned the *absence* of digest
+// validation and content verification (the traversal digest, unverified bytes,
+// a short layer, malformed digests, the claimed-bytes ceiling) now live in
+// blob_integrity_test.go, where they assert the checks that replaced them.
+// What is left here is the transport: what happens when a blob is gone, slow,
+// redirected, duplicated, or served with a body the transport itself rejects.
+
+// --- local harness ---------------------------------------------------------
+
+// startScanner is Start with the mock hold supplied by the caller. Start builds
+// its own Hold and so cannot install a presign or blob-response hook, and every
+// transport fault here needs one. Everything else matches Start, including the
+// TMPDIR restore and the shortened cooldown.
+func startScanner(t *testing.T, hold *mockhold.Hold, opts ...Option) *Harness {
+ t.Helper()
+ t.Cleanup(hold.Close)
+
+ cfg := config.DefaultConfig()
+ cfg.Hold.URL = hold.URL()
+ cfg.Hold.Secret = testSecret
+ cfg.Scanner.Workers = 1
+ cfg.Vuln.Enabled = false
+ cfg.Vuln.TmpDir = t.TempDir()
+
+ origTmpDir, hadTmpDir := os.LookupEnv("TMPDIR")
+ t.Cleanup(func() {
+ if hadTmpDir {
+ os.Setenv("TMPDIR", origTmpDir)
+ return
+ }
+ os.Unsetenv("TMPDIR")
+ })
+ for _, opt := range opts {
+ opt(cfg)
+ }
+
+ restoreCooldown := scan.JobCooldown
+ scan.JobCooldown = testJobCooldown
+
+ q := queue.NewJobQueue(cfg.Scanner.QueueSize)
+ c := client.NewHoldClient(cfg.Hold.URL, cfg.Hold.Secret, q)
+ pool := scan.NewWorkerPool(cfg, q, c)
+
+ ctx, cancel := context.WithCancel(context.Background())
+ pool.Start(ctx)
+ go c.Connect()
+
+ // Teardown waits for the worker to actually exit before restoring
+ // JobCooldown. Start does not, and the race detector sees it: a worker from
+ // the finished test is still reading the package variable while the next
+ // test's cleanup writes it. Waiting also keeps one test's in-flight scan
+ // from writing into the next test's temp directory.
+ t.Cleanup(func() {
+ cancel()
+ c.Close()
+ q.Close()
+ pool.Wait()
+ scan.JobCooldown = restoreCooldown
+ })
+
+ if err := hold.WaitForScanner(10 * time.Second); err != nil {
+ t.Fatalf("scanner never connected: %v", err)
+ }
+ return &Harness{Hold: hold, Queue: q, Client: c, Cfg: cfg}
+}
+
+// newHold builds a mock hold with the harness secret plus any hooks.
+func newHold(blobs mockhold.BlobSource, opts ...mockhold.Option) *mockhold.Hold {
+ return mockhold.New(blobs, append([]mockhold.Option{mockhold.WithSecret(testSecret)}, opts...)...)
+}
+
+// --- synthetic blobs -------------------------------------------------------
+
+const (
+ layerType = "application/vnd.oci.image.layer.v1.tar+gzip"
+ configType = "application/vnd.oci.image.config.v1+json"
+)
+
+func digestOf(b []byte) string {
+ return fmt.Sprintf("sha256:%x", sha256.Sum256(b))
+}
+
+// gzTar builds a one-file gzipped tar, the shape of a real image layer.
+func gzTar(t *testing.T, name string, content []byte) []byte {
+ t.Helper()
+ var buf bytes.Buffer
+ zw := gzip.NewWriter(&buf)
+ tw := tar.NewWriter(zw)
+ if err := tw.WriteHeader(&tar.Header{Name: name, Mode: 0644, Size: int64(len(content))}); err != nil {
+ t.Fatalf("tar header: %v", err)
+ }
+ if _, err := tw.Write(content); err != nil {
+ t.Fatalf("tar write: %v", err)
+ }
+ if err := tw.Close(); err != nil {
+ t.Fatalf("tar close: %v", err)
+ }
+ if err := zw.Close(); err != nil {
+ t.Fatalf("gzip close: %v", err)
+ }
+ return buf.Bytes()
+}
+
+// unparseableConfig is a config blob that is not JSON at all. Scenarios that
+// would otherwise scan clean through use it to stop the pipeline at Syft, so
+// the assertion stays on the download stage (see the file comment).
+func unparseableConfig() []byte { return []byte("{ this is not a container config") }
+
+// validConfig is a well-formed OCI image config, for scenarios where the
+// *layer* is the thing under test and the config must not be what fails.
+func validConfig(t *testing.T, diffIDs ...string) []byte {
+ t.Helper()
+ cfg := map[string]any{
+ "architecture": "amd64",
+ "os": "linux",
+ "config": map[string]any{},
+ "rootfs": map[string]any{"type": "layers", "diff_ids": diffIDs},
+ }
+ b, err := json.Marshal(cfg)
+ if err != nil {
+ t.Fatalf("marshal config: %v", err)
+ }
+ return b
+}
+
+func desc(digest string, size int64, mediaType string) scanner.BlobDescriptor {
+ return scanner.BlobDescriptor{Digest: digest, Size: size, MediaType: mediaType}
+}
+
+// jobFor assembles a scan job from raw descriptors. HoldEndpoint is left empty
+// so SendJob points it at the mock.
+func jobFor(cfg scanner.BlobDescriptor, layers ...scanner.BlobDescriptor) *scanner.ScanJob {
+ return &scanner.ScanJob{
+ ManifestDigest: "sha256:" + strings.Repeat("ab", 32),
+ Repository: "edge-case",
+ Tag: "latest",
+ UserDID: "did:plc:testuser",
+ UserHandle: "test.example",
+ HoldDID: "did:web:hold.test",
+ Tier: "deckhand",
+ Config: cfg,
+ Layers: layers,
+ }
+}
+
+// assertNoLeakedScanDirs checks the scanner's tmp dir is empty. The OCI layout
+// lives in tmpDir/scan-*, and stereoscope's extraction lands in the same
+// directory because WorkerPool.Start points TMPDIR at it, so a leak on any
+// error path shows up here.
+func assertNoLeakedScanDirs(t *testing.T, h *Harness) {
+ t.Helper()
+ entries, err := os.ReadDir(h.Cfg.Vuln.TmpDir)
+ if err != nil {
+ t.Fatalf("read tmp dir: %v", err)
+ }
+ for _, e := range entries {
+ t.Errorf("leftover in scanner tmp dir after failure: %s", e.Name())
+ }
+}
+
+func blobDigests(h *Harness) []string {
+ var out []string
+ for _, r := range h.Hold.BlobRequests() {
+ out = append(out, r.Digest)
+ }
+ return out
+}
+
+func countOf(list []string, want string) int {
+ n := 0
+ for _, s := range list {
+ if s == want {
+ n++
+ }
+ }
+ return n
+}
+
+// --- blob transport --------------------------------------------------------
+
+// TestMissingBlobIsRetryableError covers a layer garbage collected out from
+// under a queued job: getBlob answers 404 and there is nothing to download,
+// ever. The scanner reports "error", the hold records the scan as failed, and
+// failed records are re-offered by the stale-scan loop on every pass. Nothing
+// about the outcome can change, so this is an unbounded retry loop for a
+// permanently unscannable manifest.
+func TestMissingBlobIsRetryableError(t *testing.T) {
+ layer := gzTar(t, "usr/bin/app", []byte("hello"))
+ layerDigest := digestOf(layer)
+ cfgBytes := unparseableConfig()
+ cfgDigest := digestOf(cfgBytes)
+
+ blobs := mockhold.NewMemory().Add(cfgDigest, cfgBytes)
+
+ var hold *mockhold.Hold
+ hold = newHold(blobs, mockhold.WithPresignHook(func(digest string) (string, bool) {
+ if digest == layerDigest {
+ return "", false // garbage collected
+ }
+ return hold.URL() + "/blobs/" + mockhold.DigestHex(digest), true
+ }))
+ h := startScanner(t, hold)
+
+ seq, err := h.Hold.SendJob(jobFor(
+ desc(cfgDigest, int64(len(cfgBytes)), configType),
+ desc(layerDigest, int64(len(layer)), layerType),
+ ))
+ if err != nil {
+ t.Fatalf("send job: %v", err)
+ }
+
+ msg := h.AwaitTerminal(t, seq, 30*time.Second)
+ if msg.Type != "error" {
+ t.Fatalf("want the current retryable-error behaviour, got %s: %s%s", msg.Type, msg.Error, msg.Reason)
+ }
+ if !strings.Contains(msg.Error, "404") {
+ t.Errorf("error does not mention the 404: %q", msg.Error)
+ }
+ t.Logf("missing blob -> retryable error: %q", msg.Error)
+ assertNoLeakedScanDirs(t, h)
+}
+
+// TestGetBlobNon200 points the job at a hold endpoint whose getBlob is broken
+// in two ways a real hold can be: a 5xx, and a 200 carrying something that is
+// not the expected JSON. Both surface as retryable errors, which is right for
+// the 5xx and wrong-but-harmless for the malformed body.
+func TestGetBlobNon200(t *testing.T) {
+ cases := []struct {
+ name string
+ handler http.HandlerFunc
+ want string
+ }{
+ {
+ name: "503 from getBlob",
+ handler: func(w http.ResponseWriter, r *http.Request) {
+ http.Error(w, "hold is restarting", http.StatusServiceUnavailable)
+ },
+ want: "status 503",
+ },
+ {
+ name: "200 with a non-JSON body",
+ handler: func(w http.ResponseWriter, r *http.Request) {
+ w.Write([]byte("proxy error"))
+ },
+ want: "failed to decode response",
+ },
+ }
+
+ for _, tc := range cases {
+ t.Run(tc.name, func(t *testing.T) {
+ broken := httptest.NewServer(http.HandlerFunc(tc.handler))
+ defer broken.Close()
+
+ h := startScanner(t, newHold(mockhold.NewMemory()))
+
+ cfgBytes := unparseableConfig()
+ job := jobFor(desc(digestOf(cfgBytes), int64(len(cfgBytes)), configType))
+ job.HoldEndpoint = broken.URL
+
+ seq, err := h.Hold.SendJob(job)
+ if err != nil {
+ t.Fatalf("send job: %v", err)
+ }
+ msg := h.AwaitTerminal(t, seq, 30*time.Second)
+ if msg.Type != "error" {
+ t.Fatalf("want error, got %s: %s%s", msg.Type, msg.Error, msg.Reason)
+ }
+ if !strings.Contains(msg.Error, tc.want) {
+ t.Errorf("error %q does not mention %q", msg.Error, tc.want)
+ }
+ assertNoLeakedScanDirs(t, h)
+ })
+ }
+}
+
+// TestExpiredPresignedURL stands in for an S3 URL whose 15 minute window closed
+// while the job sat in the queue: getBlob succeeds, the download 403s.
+//
+// The assertion worth keeping is the second one. DownloadBlob reports only the
+// status code and drops the body, so the S3 ]AccessDenied
+// explanation never reaches the hold's scan record: an operator sees
+// "download returned status 403" with no indication of which blob, from where,
+// or why. GetBlobPresignedURL includes the body in its error; DownloadBlob does
+// not.
+func TestExpiredPresignedURL(t *testing.T) {
+ s3 := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
+ w.WriteHeader(http.StatusForbidden)
+ w.Write([]byte(`AccessDeniedRequest has expired`))
+ }))
+ defer s3.Close()
+
+ hold := newHold(mockhold.NewMemory(), mockhold.WithPresignHook(func(digest string) (string, bool) {
+ return s3.URL + "/bucket/blob", true
+ }))
+ h := startScanner(t, hold)
+
+ cfgBytes := unparseableConfig()
+ seq, err := h.Hold.SendJob(jobFor(desc(digestOf(cfgBytes), int64(len(cfgBytes)), configType)))
+ if err != nil {
+ t.Fatalf("send job: %v", err)
+ }
+
+ msg := h.AwaitTerminal(t, seq, 30*time.Second)
+ if msg.Type != "error" {
+ t.Fatalf("want error, got %s: %s%s", msg.Type, msg.Error, msg.Reason)
+ }
+ if !strings.Contains(msg.Error, "403") {
+ t.Errorf("error does not mention the 403: %q", msg.Error)
+ }
+ if strings.Contains(msg.Error, "AccessDenied") || strings.Contains(msg.Error, "expired") {
+ t.Errorf("DownloadBlob now includes the response body; update this test and the finding: %q", msg.Error)
+ }
+ t.Logf("expired presigned URL -> %q (body discarded)", msg.Error)
+ assertNoLeakedScanDirs(t, h)
+}
+
+// TestTruncatedBodyIsDetected sends fewer bytes than the declared
+// Content-Length. Go's client turns that into an unexpected EOF on the body, so
+// io.Copy does notice and the blob never reaches Syft. This is the transport
+// fault the code handles correctly, and it is here so a regression is visible.
+func TestTruncatedBodyIsDetected(t *testing.T) {
+ layer := gzTar(t, "usr/bin/app", bytes.Repeat([]byte("payload"), 4096))
+ layerDigest := digestOf(layer)
+ cfgBytes := unparseableConfig()
+ cfgDigest := digestOf(cfgBytes)
+
+ blobs := mockhold.NewMemory().Add(cfgDigest, cfgBytes)
+ hold := newHold(blobs, mockhold.WithBlobResponseHook(func(w http.ResponseWriter, r *http.Request, digest string) bool {
+ if digest != mockhold.DigestHex(layerDigest) {
+ return false
+ }
+ w.Header().Set("Content-Length", fmt.Sprintf("%d", len(layer)))
+ w.WriteHeader(http.StatusOK)
+ w.Write(layer[:len(layer)/4])
+ if f, ok := w.(http.Flusher); ok {
+ f.Flush()
+ }
+ return true // hang up short of Content-Length
+ }))
+ h := startScanner(t, hold)
+
+ seq, err := h.Hold.SendJob(jobFor(
+ desc(cfgDigest, int64(len(cfgBytes)), configType),
+ desc(layerDigest, int64(len(layer)), layerType),
+ ))
+ if err != nil {
+ t.Fatalf("send job: %v", err)
+ }
+
+ msg := h.AwaitTerminal(t, seq, 30*time.Second)
+ if msg.Type != "error" {
+ t.Fatalf("want error, got %s: %s%s", msg.Type, msg.Error, msg.Reason)
+ }
+ if !strings.Contains(msg.Error, "failed to write blob") {
+ t.Errorf("truncation was not caught by io.Copy; error was %q", msg.Error)
+ }
+ t.Logf("short body vs Content-Length -> %q", msg.Error)
+ assertNoLeakedScanDirs(t, h)
+}
+
+// TestStalledDownloadHasNoShortDeadline holds a response body open and shows
+// the scanner simply waits. client.httpClient's only bound is a 5 minute
+// per-request Timeout, and DownloadBlob builds no request context, so
+// cancelling the worker context (shutdown, SIGTERM) does not abort a download
+// in flight either.
+//
+// The stall here is deliberately short. The point is not to sit out the real
+// timeout but to show there is no shorter one: a 2 second stall costs the
+// worker 2 seconds, and a hold that accepts the connection and then says
+// nothing costs it five minutes per blob, times the number of layers.
+func TestStalledDownloadHasNoShortDeadline(t *testing.T) {
+ const stall = 2 * time.Second
+
+ cfgBytes := unparseableConfig()
+ cfgDigest := digestOf(cfgBytes)
+
+ hold := newHold(mockhold.NewMemory(), mockhold.WithBlobResponseHook(func(w http.ResponseWriter, r *http.Request, digest string) bool {
+ w.Header().Set("Content-Length", "4096")
+ w.WriteHeader(http.StatusOK)
+ w.Write([]byte("partial"))
+ if f, ok := w.(http.Flusher); ok {
+ f.Flush()
+ }
+ time.Sleep(stall)
+ return true
+ }))
+ h := startScanner(t, hold)
+
+ started := time.Now()
+ seq, err := h.Hold.SendJob(jobFor(desc(cfgDigest, int64(len(cfgBytes)), configType)))
+ if err != nil {
+ t.Fatalf("send job: %v", err)
+ }
+
+ msg := h.AwaitTerminal(t, seq, 60*time.Second)
+ elapsed := time.Since(started)
+ if msg.Type != "error" {
+ t.Fatalf("want error, got %s: %s%s", msg.Type, msg.Error, msg.Reason)
+ }
+ if elapsed < stall {
+ t.Errorf("scanner gave up after %s, before the stall ended; a shorter deadline now exists", elapsed)
+ }
+ t.Logf("stalled body held the worker for %s, then: %q", elapsed.Round(100*time.Millisecond), msg.Error)
+ assertNoLeakedScanDirs(t, h)
+}
+
+// TestDownloadFollowsRedirectToAnotherHost shows the scanner will fetch blob
+// bytes from wherever the getBlob response, or any redirect from it, points.
+// DownloadBlob sends no credentials, so nothing leaks; what it means is that
+// the hold (or anything able to answer as the hold, or an open redirect on the
+// presigned URL's host) chooses which server the scanner talks to. That is
+// acceptable now for the reason B-10 gave: the bytes are verified against the
+// digest on arrival, so where they came from does not decide what they are.
+func TestDownloadFollowsRedirectToAnotherHost(t *testing.T) {
+ cfgBytes := unparseableConfig()
+ cfgDigest := digestOf(cfgBytes)
+
+ var elsewhereHits atomic.Int32
+ var sawAuth atomic.Bool
+ elsewhere := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
+ elsewhereHits.Add(1)
+ if r.Header.Get("Authorization") != "" {
+ sawAuth.Store(true)
+ }
+ w.Write(cfgBytes)
+ }))
+ defer elsewhere.Close()
+
+ hold := newHold(mockhold.NewMemory(), mockhold.WithBlobResponseHook(func(w http.ResponseWriter, r *http.Request, digest string) bool {
+ http.Redirect(w, r, elsewhere.URL+"/somewhere-else", http.StatusFound)
+ return true
+ }))
+ h := startScanner(t, hold)
+
+ seq, err := h.Hold.SendJob(jobFor(desc(cfgDigest, int64(len(cfgBytes)), configType)))
+ if err != nil {
+ t.Fatalf("send job: %v", err)
+ }
+
+ msg := h.AwaitTerminal(t, seq, 30*time.Second)
+ if msg.Type != "error" {
+ t.Fatalf("want error (the redirected bytes are an unparseable config), got %s: %s%s",
+ msg.Type, msg.Error, msg.Reason)
+ }
+ if elsewhereHits.Load() == 0 {
+ t.Error("redirect to a different host was not followed")
+ }
+ if sawAuth.Load() {
+ t.Error("Authorization header was sent to the redirect target")
+ }
+ t.Logf("cross-host redirect followed %d time(s), no credentials attached", elsewhereHits.Load())
+ assertNoLeakedScanDirs(t, h)
+}
+
+// TestRedirectChainStopsAtTen pins the only bound on a redirect chain: Go's
+// default client policy. The scanner sets no policy of its own, so a hold
+// pointing a blob at a redirect loop costs eleven round trips per blob before
+// erroring, and the error names the URL rather than the blob.
+func TestRedirectChainStopsAtTen(t *testing.T) {
+ cfgBytes := unparseableConfig()
+ cfgDigest := digestOf(cfgBytes)
+
+ var hold *mockhold.Hold
+ hold = newHold(mockhold.NewMemory(), mockhold.WithBlobResponseHook(func(w http.ResponseWriter, r *http.Request, digest string) bool {
+ http.Redirect(w, r, hold.URL()+"/blobs/"+digest, http.StatusFound)
+ return true
+ }))
+ h := startScanner(t, hold)
+
+ seq, err := h.Hold.SendJob(jobFor(desc(cfgDigest, int64(len(cfgBytes)), configType)))
+ if err != nil {
+ t.Fatalf("send job: %v", err)
+ }
+
+ msg := h.AwaitTerminal(t, seq, 30*time.Second)
+ if msg.Type != "error" {
+ t.Fatalf("want error, got %s: %s%s", msg.Type, msg.Error, msg.Reason)
+ }
+ if !strings.Contains(msg.Error, "redirect") {
+ t.Errorf("error does not mention redirects: %q", msg.Error)
+ }
+ if n := len(h.Hold.BlobRequests()); n < 10 {
+ t.Errorf("redirect loop cost %d requests, expected the client's 10-hop limit", n)
+ }
+ t.Logf("redirect loop: %d requests, then %q", len(h.Hold.BlobRequests()), msg.Error)
+ assertNoLeakedScanDirs(t, h)
+}
+
+// --- layout construction ---------------------------------------------------
+
+// TestDuplicateLayerIsDownloadedTwice feeds the same layer digest twice, which
+// a hand-written manifest record can do. Both copies are fetched over the
+// network and written to the same path, and the layout's manifest lists the
+// layer twice.
+//
+// The config is deliberately unparseable so the run ends at Syft rather than
+// in a scan verdict; the download accounting is what this asserts.
+func TestDuplicateLayerIsDownloadedTwice(t *testing.T) {
+ layer := gzTar(t, "usr/bin/app", []byte("hello"))
+ layerDigest := digestOf(layer)
+ cfgBytes := unparseableConfig()
+ cfgDigest := digestOf(cfgBytes)
+
+ blobs := mockhold.NewMemory().
+ Add(cfgDigest, cfgBytes).
+ Add(layerDigest, layer)
+ h := startScanner(t, newHold(blobs))
+
+ seq, err := h.Hold.SendJob(jobFor(
+ desc(cfgDigest, int64(len(cfgBytes)), configType),
+ desc(layerDigest, int64(len(layer)), layerType),
+ desc(layerDigest, int64(len(layer)), layerType),
+ ))
+ if err != nil {
+ t.Fatalf("send job: %v", err)
+ }
+
+ msg := h.AwaitTerminal(t, seq, 30*time.Second)
+ if msg.Type != "error" {
+ t.Fatalf("want error, got %s: %s%s", msg.Type, msg.Error, msg.Reason)
+ }
+ if n := countOf(blobDigests(h), mockhold.DigestHex(layerDigest)); n != 2 {
+ t.Errorf("duplicate layer fetched %d times, want the current behaviour of 2", n)
+ }
+ t.Logf("duplicate layer digest fetched %d times", countOf(blobDigests(h), mockhold.DigestHex(layerDigest)))
+ assertNoLeakedScanDirs(t, h)
+}
+
+// TestUnparseableConfigBlobIsRetryableError states outright what several
+// scenarios above lean on: a config blob whose bytes are not JSON fails inside
+// generateSBOM and comes back as "error". That is a permanent property of the
+// image, so the hold's stale-scan loop will offer it again on every pass and
+// get the same answer forever. The same is true of a layer whose bytes are not
+// the tar its media type claims (TestDigestIsNeverVerified).
+func TestUnparseableConfigBlobIsRetryableError(t *testing.T) {
+ cfgBytes := unparseableConfig()
+ cfgDigest := digestOf(cfgBytes)
+ layer := gzTar(t, "usr/bin/app", []byte("hello"))
+ layerDigest := digestOf(layer)
+
+ blobs := mockhold.NewMemory().
+ Add(cfgDigest, cfgBytes).
+ Add(layerDigest, layer)
+ h := startScanner(t, newHold(blobs))
+
+ seq, err := h.Hold.SendJob(jobFor(
+ desc(cfgDigest, int64(len(cfgBytes)), configType),
+ desc(layerDigest, int64(len(layer)), layerType),
+ ))
+ if err != nil {
+ t.Fatalf("send job: %v", err)
+ }
+
+ msg := h.AwaitTerminal(t, seq, 60*time.Second)
+ if msg.Type != "error" {
+ t.Fatalf("want the current retryable-error behaviour, got %s: %s%s", msg.Type, msg.Error, msg.Reason)
+ }
+ t.Logf("unparseable config -> retryable error: %q", msg.Error)
+ assertNoLeakedScanDirs(t, h)
+}
+
+// TestEmptyLayerDigestIsSkippedNotFetched pins the one malformed-digest case
+// the layout builder does handle: an empty digest among valid ones is dropped
+// from both the download loop and the layout manifest.
+//
+// The surviving layer is served as garbage so the run ends in an error, which
+// keeps the assertion on which blobs were fetched.
+func TestEmptyLayerDigestIsSkipped(t *testing.T) {
+ layer := []byte("not actually a gzip stream")
+ layerDigest := digestOf(layer)
+ cfgBytes := unparseableConfig()
+ cfgDigest := digestOf(cfgBytes)
+
+ blobs := mockhold.NewMemory().
+ Add(cfgDigest, cfgBytes).
+ Add(layerDigest, layer)
+ h := startScanner(t, newHold(blobs))
+
+ seq, err := h.Hold.SendJob(jobFor(
+ desc(cfgDigest, int64(len(cfgBytes)), configType),
+ desc("", 0, layerType), // empty digest
+ desc(layerDigest, int64(len(layer)), layerType),
+ ))
+ if err != nil {
+ t.Fatalf("send job: %v", err)
+ }
+
+ msg := h.AwaitTerminal(t, seq, 30*time.Second)
+ if msg.Type != "error" {
+ t.Fatalf("want error, got %s: %s%s", msg.Type, msg.Error, msg.Reason)
+ }
+ if n := len(h.Hold.BlobRequests()); n != 2 {
+ t.Errorf("fetched %d blobs, want config + one real layer: %v", n, blobDigests(h))
+ }
+ assertNoLeakedScanDirs(t, h)
+}
+
+// TestZeroSizeLayerIsFetched covers a descriptor claiming size 0 whose blob is
+// genuinely empty. The size field is not consulted at download time, so the
+// blob is fetched and an empty file is written into the layout.
+func TestZeroSizeLayerIsFetched(t *testing.T) {
+ empty := []byte{}
+ emptyDigest := digestOf(empty)
+ cfgBytes := unparseableConfig()
+ cfgDigest := digestOf(cfgBytes)
+
+ blobs := mockhold.NewMemory().
+ Add(cfgDigest, cfgBytes).
+ Add(emptyDigest, empty)
+ h := startScanner(t, newHold(blobs))
+
+ seq, err := h.Hold.SendJob(jobFor(
+ desc(cfgDigest, int64(len(cfgBytes)), configType),
+ desc(emptyDigest, 0, layerType),
+ ))
+ if err != nil {
+ t.Fatalf("send job: %v", err)
+ }
+
+ msg := h.AwaitTerminal(t, seq, 30*time.Second)
+ if msg.Type != "error" {
+ t.Fatalf("want error, got %s: %s%s", msg.Type, msg.Error, msg.Reason)
+ }
+ if n := countOf(blobDigests(h), mockhold.DigestHex(emptyDigest)); n != 1 {
+ t.Errorf("zero-size layer fetched %d times, want 1", n)
+ }
+ t.Logf("zero-size layer -> %q", msg.Error)
+ assertNoLeakedScanDirs(t, h)
+}
+
+// TestZeroSizeLayerProducesACleanScan is the same image with a config the
+// scanner can actually read, and it is the cheapest end-to-end proof that a
+// successful scan no longer kills the process.
+//
+// stereoscope accepts a zero-byte layer as a valid, empty layer, so the scan
+// succeeds with an SBOM containing no packages and, with Grype off, no
+// vulnerability summary. That combination used to reach the unconditional
+// result.Summary dereference in worker.go and panic the test binary. It is
+// worth keeping for a second reason too: an image whose layer descriptors are
+// all zero-size blobs comes back as a clean scan rather than a failure, and
+// nothing in the scanner distinguishes "no vulnerabilities" from "nothing was
+// there".
+func TestZeroSizeLayerProducesACleanScan(t *testing.T) {
+ empty := []byte{}
+ emptyDigest := digestOf(empty)
+ cfgBytes := validConfig(t, emptyDigest)
+ cfgDigest := digestOf(cfgBytes)
+
+ blobs := mockhold.NewMemory().
+ Add(cfgDigest, cfgBytes).
+ Add(emptyDigest, empty)
+ h := startScanner(t, newHold(blobs))
+
+ seq, err := h.Hold.SendJob(jobFor(
+ desc(cfgDigest, int64(len(cfgBytes)), configType),
+ desc(emptyDigest, 0, layerType),
+ ))
+ if err != nil {
+ t.Fatalf("send job: %v", err)
+ }
+
+ msg := h.AwaitTerminal(t, seq, 60*time.Second)
+ if msg.Type != "result" {
+ t.Fatalf("want result, got %s: %s%s", msg.Type, msg.Error, msg.Reason)
+ }
+ if msg.Summary != nil {
+ t.Errorf("vuln scanning is disabled but the result carried a summary: %+v", msg.Summary)
+ }
+ t.Logf("zero-byte layer scanned clean, SBOM %d bytes", len(msg.SBOM))
+}
+
+// TestManyLayersAreFetchedSequentially runs an image as wide as the widest
+// manifest in the corpus (19 layers) through the download stage. Each layer is
+// one sequential HTTP round trip with its own 5 minute ceiling, so the
+// worst-case time for this one job is not five minutes but ninety-five, and
+// nothing caps it.
+//
+// The width comes from the corpus but the blobs are synthesised, because every
+// blob is now verified against its digest and no fixture can serve bytes that
+// hash to a real registry's digests. The layer count is what this measures.
+func TestManyLayersAreFetchedSequentially(t *testing.T) {
+ width := widestCorpusWidth(t)
+ if width < 10 {
+ t.Skipf("corpus has no wide manifest; widest is %d layers", width)
+ }
+
+ blobs := mockhold.NewMemory()
+ layers := make([]scanner.BlobDescriptor, 0, width)
+ diffIDs := make([]string, 0, width)
+ for i := 0; i < width; i++ {
+ b := gzTar(t, fmt.Sprintf("usr/bin/app%d", i), []byte(fmt.Sprintf("layer %d", i)))
+ d := digestOf(b)
+ blobs.Add(d, b)
+ layers = append(layers, desc(d, int64(len(b)), layerType))
+ diffIDs = append(diffIDs, d)
+ }
+ cfgBytes := validConfig(t, diffIDs...)
+ cfgDigest := digestOf(cfgBytes)
+ blobs.Add(cfgDigest, cfgBytes)
+
+ h := startScanner(t, newHold(blobs))
+
+ send := func() int64 {
+ t.Helper()
+ seq, err := h.Hold.SendJob(jobFor(desc(cfgDigest, int64(len(cfgBytes)), configType), layers...))
+ if err != nil {
+ t.Fatalf("send job: %v", err)
+ }
+ return seq
+ }
+
+ msg := h.AwaitTerminal(t, send(), 60*time.Second)
+ if msg.Type != "result" {
+ t.Fatalf("want result, got %s: %s%s", msg.Type, msg.Error, msg.Reason)
+ }
+ reqs := h.Hold.BlobRequests()
+ if len(reqs) != width+1 {
+ t.Errorf("fetched %d blobs, want %d (config + %d layers)", len(reqs), width+1, width)
+ }
+ // Sequential, not concurrent: no two fetches overlap.
+ for i := 1; i < len(reqs); i++ {
+ if reqs[i].At.Before(reqs[i-1].At) {
+ t.Errorf("blob requests are out of order at %d", i)
+ }
+ }
+ t.Logf("%d-layer manifest: %d sequential fetches", width, len(reqs))
+
+ // Per-layer resource accounting: downloadBlob opens a file and a response
+ // body per layer. Run the same job twice more and require the process's
+ // file descriptor count to hold steady, which is what rules out a per-layer
+ // handle leak (a leak would grow by ~20 per run).
+ if fds, ok := openFDs(); ok {
+ for i := 0; i < 2; i++ {
+ h.AwaitTerminal(t, send(), 60*time.Second)
+ }
+ time.Sleep(200 * time.Millisecond)
+ after, _ := openFDs()
+ if after > fds+8 {
+ t.Errorf("file descriptors grew from %d to %d across two more %d-layer scans",
+ fds, after, width)
+ }
+ t.Logf("file descriptors: %d after one scan, %d after three", fds, after)
+ }
+
+ assertNoLeakedScanDirs(t, h)
+}
+
+// widestCorpusWidth reports the layer count of the widest manifest in the
+// corpus, so the shape under test stays tied to what a hold really holds.
+func widestCorpusWidth(t *testing.T) int {
+ t.Helper()
+ all, err := mockhold.Corpus()
+ if err != nil {
+ t.Fatalf("load corpus: %v", err)
+ }
+ widest := 0
+ for _, m := range all {
+ if len(m.Layers) > widest {
+ widest = len(m.Layers)
+ }
+ }
+ return widest
+}
+
+// openFDs returns the number of open file descriptors, and whether the count
+// was available (it is not on platforms without /proc).
+func openFDs() (int, bool) {
+ entries, err := os.ReadDir("/proc/self/fd")
+ if err != nil {
+ return 0, false
+ }
+ return len(entries), true
+}
+
+// --- size guard ------------------------------------------------------------
+
+// TestMaxImageSizeZeroDisablesTheGuard confirms the documented "0 = no limit"
+// reading of the config. The guard now spends a budget down as bytes arrive,
+// so the case worth pinning is that a zero limit spends nothing: an image far
+// larger than any of the other scenarios here transfers in full.
+//
+// The descriptors are honest, because a manifest that lies about its size is
+// now rejected by the descriptor check rather than by the ceiling, which would
+// prove nothing about the ceiling.
+func TestMaxImageSizeZeroDisablesTheGuard(t *testing.T) {
+ noise := make([]byte, 64*1024)
+ rand.New(rand.NewSource(2)).Read(noise)
+ big := gzTar(t, "usr/bin/app", noise)
+ bigDigest := digestOf(big)
+
+ // Unparseable on purpose: the run ends at Syft, so the assertion stays on
+ // the download stage rather than on a scan verdict.
+ cfgBytes := unparseableConfig()
+ cfgDigest := digestOf(cfgBytes)
+
+ blobs := mockhold.NewMemory().
+ Add(cfgDigest, cfgBytes).
+ Add(bigDigest, big)
+ h := startScanner(t, newHold(blobs), WithMaxImageSize(0))
+
+ seq, err := h.Hold.SendJob(jobFor(
+ desc(cfgDigest, int64(len(cfgBytes)), configType),
+ desc(bigDigest, int64(len(big)), layerType),
+ ))
+ if err != nil {
+ t.Fatalf("send job: %v", err)
+ }
+
+ msg := h.AwaitTerminal(t, seq, 30*time.Second)
+ if msg.Type != "error" {
+ t.Fatalf("want error from Syft, got %s: %s%s", msg.Type, msg.Error, msg.Reason)
+ }
+ if strings.Contains(msg.Error, "too large") || strings.Contains(msg.Reason, "too large") {
+ t.Fatalf("MaxImageSize=0 rejected a job; the 0-means-unlimited contract changed: %q%q",
+ msg.Error, msg.Reason)
+ }
+ if n := len(h.Hold.BlobRequests()); n != 2 {
+ t.Errorf("fetched %d blobs, want config + layer with no ceiling in force", n)
+ }
+ assertNoLeakedScanDirs(t, h)
+}
+
+// --- cleanup ---------------------------------------------------------------
+
+// TestTmpDirIsCleanAfterAFailureInsideSyft covers the one cleanup path that is
+// not a direct return from buildOCILayout: the layout is built, downloads
+// succeed, and generateSBOM fails. processJob's deferred cleanup is responsible
+// for the directory, and stereoscope is responsible for its own extraction
+// scratch space under the same TMPDIR.
+func TestTmpDirIsCleanAfterAFailureInsideSyft(t *testing.T) {
+ layer := gzTar(t, "usr/bin/app", []byte("hello"))
+ layerDigest := digestOf(layer)
+ cfgBytes := unparseableConfig()
+ cfgDigest := digestOf(cfgBytes)
+
+ blobs := mockhold.NewMemory().
+ Add(cfgDigest, cfgBytes).
+ Add(layerDigest, layer)
+ h := startScanner(t, newHold(blobs))
+
+ seq, err := h.Hold.SendJob(jobFor(
+ desc(cfgDigest, int64(len(cfgBytes)), configType),
+ desc(layerDigest, int64(len(layer)), layerType),
+ ))
+ if err != nil {
+ t.Fatalf("send job: %v", err)
+ }
+
+ msg := h.AwaitTerminal(t, seq, 60*time.Second)
+ if msg.Type != "error" {
+ t.Fatalf("want error, got %s: %s%s", msg.Type, msg.Error, msg.Reason)
+ }
+ if !strings.Contains(msg.Error, "SBOM") && !strings.Contains(msg.Error, "OCI image") {
+ t.Logf("failure came from somewhere other than Syft: %q", msg.Error)
+ }
+ assertNoLeakedScanDirs(t, h)
+}
diff --git a/scanner/internal/e2e/blob_integrity_test.go b/scanner/internal/e2e/blob_integrity_test.go
new file mode 100644
index 0000000..c553adb
--- /dev/null
+++ b/scanner/internal/e2e/blob_integrity_test.go
@@ -0,0 +1,422 @@
+package e2e
+
+import (
+ "bytes"
+ "fmt"
+ "math/rand"
+ "net/http"
+ "os"
+ "path/filepath"
+ "strings"
+ "sync/atomic"
+ "testing"
+ "time"
+
+ "atcr.io/scanner/internal/mockhold"
+)
+
+// This file states what the scanner must do with blob bytes and blob names it
+// did not choose. Both come from an io.atcr.manifest record in a user's own
+// PDS, so both are attacker-controlled text, and the hold's dispatch guards
+// check neither.
+//
+// Two invariants, and everything here is one of them:
+//
+// 1. A digest names a blob and a file. It is only ever "sha256:<64 hex>";
+// anything else is rejected before it reaches a path or the network.
+// 2. The bytes that land on disk are the bytes the digest names, in the
+// amount the descriptor claims, within the budget the config allows.
+//
+// A violation of either is a permanent property of the record or of the stored
+// blob, so it must come back as "skipped" (which the hold records once) rather
+// than "error" (which the stale-scan loop re-offers forever).
+
+// --- invariant 1: a digest is a name, never a path -------------------------
+
+// TestTraversalDigestNeverBecomesAPath is the security case. A layer digest of
+// "sha256:../../../escaped-marker" used to be split on ":" and filepath.Joined
+// onto blobs/sha256, which is exactly three levels below the scanner's tmp
+// dir, so the downloaded bytes landed in the tmp dir itself: outside the scan
+// directory, beyond the reach of cleanup, and over whatever file was already
+// there (os.Create truncates).
+//
+// The escape is aimed at the harness's own t.TempDir so the test proves the
+// boundary holds without writing anywhere real.
+func TestTraversalDigestNeverBecomesAPath(t *testing.T) {
+ payload := []byte("bytes the scanner must never place here")
+ traversal := "sha256:../../../escaped-marker"
+
+ cfgBytes := validConfig(t)
+ cfgDigest := digestOf(cfgBytes)
+
+ blobs := mockhold.NewMemory().
+ Add(cfgDigest, cfgBytes).
+ Add("payload", payload)
+
+ // A hold willing to serve a 200 for the traversal digest, which is the
+ // precondition the old write needed.
+ var hold *mockhold.Hold
+ hold = newHold(blobs, mockhold.WithPresignHook(func(digest string) (string, bool) {
+ if digest == traversal {
+ return hold.URL() + "/blobs/payload", true
+ }
+ return hold.URL() + "/blobs/" + mockhold.DigestHex(digest), true
+ }))
+ h := startScanner(t, hold)
+
+ seq, err := h.Hold.SendJob(jobFor(
+ desc(cfgDigest, int64(len(cfgBytes)), configType),
+ desc(traversal, int64(len(payload)), layerType),
+ ))
+ if err != nil {
+ t.Fatalf("send job: %v", err)
+ }
+
+ msg := h.AwaitTerminal(t, seq, 30*time.Second)
+ if msg.Type != "skipped" {
+ t.Fatalf("want skipped (a malformed digest can never succeed), got %s: %s%s",
+ msg.Type, msg.Error, msg.Reason)
+ }
+ if !strings.Contains(msg.Reason, "digest") {
+ t.Errorf("skip reason %q does not name the digest as the problem", msg.Reason)
+ }
+
+ // Nothing may have been written outside the scan directory. The tmp dir is
+ // the escape's target, and after a scan it must hold nothing at all.
+ escaped := filepath.Join(h.Cfg.Vuln.TmpDir, "escaped-marker")
+ if got, err := os.ReadFile(escaped); err == nil {
+ os.Remove(escaped)
+ t.Fatalf("the traversal still wrote %d bytes to %s", len(got), escaped)
+ }
+ assertNoLeakedScanDirs(t, h)
+
+ // And the rejection happens before any network call: a digest that cannot
+ // name a file must not be sent to the hold as a blob name either.
+ if n := len(h.Hold.BlobRequests()); n != 0 {
+ t.Errorf("scanner made %d blob requests for a job it must reject up front: %v",
+ n, blobDigests(h))
+ }
+}
+
+// TestMalformedDigestsAreSkippedBeforeAnyDownload walks the digest shapes a
+// hand-written manifest record can carry. Every one is permanently
+// unscannable, so every one is a skip, and none of them costs a round trip.
+func TestMalformedDigestsAreSkippedBeforeAnyDownload(t *testing.T) {
+ layer := gzTar(t, "usr/bin/app", []byte("hello"))
+
+ cases := []struct {
+ name string
+ digest string
+ }{
+ {"no algorithm prefix", strings.Repeat("ef", 32)},
+ {"unsupported algorithm", "sha512:" + strings.Repeat("cd", 64)},
+ {"uppercase hex", "sha256:" + strings.ToUpper(strings.Repeat("ab", 32))},
+ {"truncated hex", "sha256:abcd"},
+ {"path traversal", "sha256:../../../escaped-marker"},
+ }
+
+ for _, tc := range cases {
+ t.Run(tc.name, func(t *testing.T) {
+ cfgBytes := validConfig(t)
+ cfgDigest := digestOf(cfgBytes)
+
+ blobs := mockhold.NewMemory().
+ Add(cfgDigest, cfgBytes).
+ Add(tc.digest, layer)
+ h := startScanner(t, newHold(blobs))
+
+ seq, err := h.Hold.SendJob(jobFor(
+ desc(cfgDigest, int64(len(cfgBytes)), configType),
+ desc(tc.digest, int64(len(layer)), layerType),
+ ))
+ if err != nil {
+ t.Fatalf("send job: %v", err)
+ }
+ msg := h.AwaitTerminal(t, seq, 30*time.Second)
+ if msg.Type != "skipped" {
+ t.Fatalf("want skipped, got %s: %s%s", msg.Type, msg.Error, msg.Reason)
+ }
+ if n := len(h.Hold.BlobRequests()); n != 0 {
+ t.Errorf("made %d blob requests before rejecting a malformed digest: %v",
+ n, blobDigests(h))
+ }
+ t.Logf("%s -> skipped: %s", tc.name, msg.Reason)
+ assertNoLeakedScanDirs(t, h)
+ })
+ }
+}
+
+// TestMalformedConfigDigestIsSkipped covers the same validation on the config
+// descriptor, which reaches downloadBlob by a separate call site.
+func TestMalformedConfigDigestIsSkipped(t *testing.T) {
+ h := startScanner(t, newHold(mockhold.NewMemory()))
+
+ seq, err := h.Hold.SendJob(jobFor(desc("sha256:../../../escaped-config", 10, configType)))
+ if err != nil {
+ t.Fatalf("send job: %v", err)
+ }
+ msg := h.AwaitTerminal(t, seq, 30*time.Second)
+ if msg.Type != "skipped" {
+ t.Fatalf("want skipped, got %s: %s%s", msg.Type, msg.Error, msg.Reason)
+ }
+ if got, err := os.ReadFile(filepath.Join(h.Cfg.Vuln.TmpDir, "escaped-config")); err == nil {
+ os.Remove(filepath.Join(h.Cfg.Vuln.TmpDir, "escaped-config"))
+ t.Fatalf("config traversal wrote %d bytes outside the scan directory", len(got))
+ }
+ assertNoLeakedScanDirs(t, h)
+}
+
+// --- invariant 2: the bytes are the bytes ----------------------------------
+
+// TestDigestMismatchIsDetectedAndSkipped serves bytes that do not hash to the
+// digest naming them, which is what a mixed-up S3 key, a poisoned cache or a
+// compromised BYOS hold produces. A compliant OCI client verifies on pull and
+// refuses these bytes; the scanner is the one consumer in the system that used
+// to trust them, and then vouched for them in a scan record.
+//
+// The failure must name the digest, not surface as whatever stereoscope makes
+// of a stream it cannot parse.
+func TestDigestMismatchIsDetectedAndSkipped(t *testing.T) {
+ layer := gzTar(t, "usr/bin/app", []byte("real content"))
+ layerDigest := digestOf(layer)
+ imposter := []byte("this is not a gzip stream at all")
+
+ cfgBytes := validConfig(t, digestOf(layer))
+ cfgDigest := digestOf(cfgBytes)
+
+ blobs := mockhold.NewMemory().
+ Add(cfgDigest, cfgBytes).
+ Add(layerDigest, imposter) // wrong bytes under the right name
+ h := startScanner(t, newHold(blobs))
+
+ seq, err := h.Hold.SendJob(jobFor(
+ desc(cfgDigest, int64(len(cfgBytes)), configType),
+ desc(layerDigest, int64(len(imposter)), layerType),
+ ))
+ if err != nil {
+ t.Fatalf("send job: %v", err)
+ }
+
+ msg := h.AwaitTerminal(t, seq, 30*time.Second)
+ if msg.Type != "skipped" {
+ t.Fatalf("want skipped, got %s: %s%s", msg.Type, msg.Error, msg.Reason)
+ }
+ if !strings.Contains(msg.Reason, "digest mismatch") {
+ t.Errorf("skip reason %q does not say the bytes failed their digest", msg.Reason)
+ }
+ if !strings.Contains(msg.Reason, mockhold.DigestHex(layerDigest)[:16]) {
+ t.Errorf("skip reason %q does not name the offending blob", msg.Reason)
+ }
+ assertNoLeakedScanDirs(t, h)
+}
+
+// TestConfigDigestMismatchIsDetectedAndSkipped is the same check on the config
+// blob, whose download is a separate call site from the layer loop.
+func TestConfigDigestMismatchIsDetectedAndSkipped(t *testing.T) {
+ cfgBytes := validConfig(t)
+ cfgDigest := digestOf(cfgBytes)
+ imposter := []byte(`{"architecture":"amd64","os":"linux","rootfs":{"type":"layers"}}`)
+
+ blobs := mockhold.NewMemory().Add(cfgDigest, imposter)
+ h := startScanner(t, newHold(blobs))
+
+ seq, err := h.Hold.SendJob(jobFor(desc(cfgDigest, int64(len(imposter)), configType)))
+ if err != nil {
+ t.Fatalf("send job: %v", err)
+ }
+ msg := h.AwaitTerminal(t, seq, 30*time.Second)
+ if msg.Type != "skipped" {
+ t.Fatalf("want skipped, got %s: %s%s", msg.Type, msg.Error, msg.Reason)
+ }
+ if !strings.Contains(msg.Reason, "digest mismatch") {
+ t.Errorf("skip reason %q does not say the bytes failed their digest", msg.Reason)
+ }
+ assertNoLeakedScanDirs(t, h)
+}
+
+// TestShortLayerIsRejectedByTheDescriptorCheck is the truncation the transport
+// cannot see: a body that is internally consistent (Content-Length matches
+// what is sent) but shorter than the size the descriptor declares.
+//
+// Both the digest and the size disagree with the bytes here, which is the
+// point: descriptor.Size is attacker-supplied too, so the right verdict is
+// that the descriptor as a whole does not describe the blob, and the message
+// says so rather than picking one field to blame.
+func TestShortLayerIsRejectedByTheDescriptorCheck(t *testing.T) {
+ layer := gzTar(t, "usr/bin/app", bytes.Repeat([]byte("payload"), 4096))
+ layerDigest := digestOf(layer)
+ short := layer[:len(layer)/4]
+
+ cfgBytes := validConfig(t, digestOf(layer))
+ cfgDigest := digestOf(cfgBytes)
+
+ blobs := mockhold.NewMemory().
+ Add(cfgDigest, cfgBytes).
+ Add(layerDigest, short)
+ h := startScanner(t, newHold(blobs))
+
+ seq, err := h.Hold.SendJob(jobFor(
+ desc(cfgDigest, int64(len(cfgBytes)), configType),
+ desc(layerDigest, int64(len(layer)), layerType), // declares the full size
+ ))
+ if err != nil {
+ t.Fatalf("send job: %v", err)
+ }
+
+ msg := h.AwaitTerminal(t, seq, 30*time.Second)
+ if msg.Type != "skipped" {
+ t.Fatalf("want skipped, got %s: %s%s", msg.Type, msg.Error, msg.Reason)
+ }
+ if !strings.Contains(msg.Reason, fmt.Sprintf("%d", len(short))) {
+ t.Errorf("skip reason %q does not report the %d bytes actually received", msg.Reason, len(short))
+ }
+ if !strings.Contains(msg.Reason, fmt.Sprintf("%d", len(layer))) {
+ t.Errorf("skip reason %q does not report the %d bytes claimed", msg.Reason, len(layer))
+ }
+ assertNoLeakedScanDirs(t, h)
+}
+
+// TestHonestBlobsStillScan is the control. Verification must not reject an
+// image whose bytes are what they say they are: config, layer, digests and
+// sizes all agree, and the scan runs through to a result.
+func TestHonestBlobsStillScan(t *testing.T) {
+ layer := gzTar(t, "usr/bin/app", []byte("hello from a real layer"))
+ layerDigest := digestOf(layer)
+ cfgBytes := validConfig(t, digestOf(layer))
+ cfgDigest := digestOf(cfgBytes)
+
+ blobs := mockhold.NewMemory().
+ Add(cfgDigest, cfgBytes).
+ Add(layerDigest, layer)
+ h := startScanner(t, newHold(blobs))
+
+ seq, err := h.Hold.SendJob(jobFor(
+ desc(cfgDigest, int64(len(cfgBytes)), configType),
+ desc(layerDigest, int64(len(layer)), layerType),
+ ))
+ if err != nil {
+ t.Fatalf("send job: %v", err)
+ }
+
+ msg := h.AwaitTerminal(t, seq, 60*time.Second)
+ if msg.Type != "result" {
+ t.Fatalf("verification rejected an honest image: %s: %s%s", msg.Type, msg.Error, msg.Reason)
+ }
+ assertNoLeakedScanDirs(t, h)
+}
+
+// TestUndeclaredSizeIsAcceptedWhenTheDigestHolds pins the one relaxation in
+// the size check. A descriptor whose Size is zero has not claimed a length, so
+// the digest alone decides. Real records do carry sizes; this keeps a record
+// that omits one scannable rather than making the weaker field authoritative.
+func TestUndeclaredSizeIsAcceptedWhenTheDigestHolds(t *testing.T) {
+ layer := gzTar(t, "usr/bin/app", []byte("hello"))
+ layerDigest := digestOf(layer)
+ cfgBytes := validConfig(t, digestOf(layer))
+ cfgDigest := digestOf(cfgBytes)
+
+ blobs := mockhold.NewMemory().
+ Add(cfgDigest, cfgBytes).
+ Add(layerDigest, layer)
+ h := startScanner(t, newHold(blobs))
+
+ seq, err := h.Hold.SendJob(jobFor(
+ desc(cfgDigest, 0, configType),
+ desc(layerDigest, 0, layerType),
+ ))
+ if err != nil {
+ t.Fatalf("send job: %v", err)
+ }
+ msg := h.AwaitTerminal(t, seq, 60*time.Second)
+ if msg.Type != "result" {
+ t.Fatalf("a size-less descriptor with correct bytes was rejected: %s: %s%s",
+ msg.Type, msg.Error, msg.Reason)
+ }
+ assertNoLeakedScanDirs(t, h)
+}
+
+// --- the size budget is measured, not claimed ------------------------------
+
+// TestMaxImageSizeCountsTransferredBytes closes the ceiling bypass. The guard
+// summed BlobDescriptor.Size, which comes from the same user-writable record
+// as the digests, so a manifest claiming one byte per blob passed any limit
+// and the scanner then wrote the real bytes to its tmp volume: 65,696 bytes
+// measured against a 1,024 byte limit. The layer below declares one byte and
+// the hold offers 65,696.
+//
+// The claimed-size pre-check stays (it is still worth refusing an honestly
+// large image before a byte moves), but the budget is now enforced against
+// what actually arrives.
+func TestMaxImageSizeCountsTransferredBytes(t *testing.T) {
+ const limit = 1024
+
+ noise := make([]byte, 64*1024)
+ rand.New(rand.NewSource(1)).Read(noise)
+ big := gzTar(t, "usr/bin/app", noise)
+ bigDigest := digestOf(big)
+ cfgBytes := validConfig(t, digestOf(big))
+ cfgDigest := digestOf(cfgBytes)
+
+ var served atomic.Int64
+ blobs := mockhold.NewMemory().
+ Add(cfgDigest, cfgBytes).
+ Add(bigDigest, big)
+
+ hold := newHold(blobs, mockhold.WithBlobResponseHook(func(w http.ResponseWriter, r *http.Request, digest string) bool {
+ if digest != mockhold.DigestHex(bigDigest) {
+ return false
+ }
+ w.Header().Set("Content-Length", fmt.Sprintf("%d", len(big)))
+ n, _ := w.Write(big)
+ served.Add(int64(n))
+ return true
+ }))
+ h := startScanner(t, hold, WithMaxImageSize(limit))
+
+ // The layer lies about its size, the way a hand-written record can. The
+ // config is honest, so the job reaches the download stage and the budget
+ // is what has to stop it rather than the descriptor check.
+ seq, err := h.Hold.SendJob(jobFor(
+ desc(cfgDigest, int64(len(cfgBytes)), configType),
+ desc(bigDigest, 1, layerType),
+ ))
+ if err != nil {
+ t.Fatalf("send job: %v", err)
+ }
+
+ msg := h.AwaitTerminal(t, seq, 30*time.Second)
+ if msg.Type != "skipped" {
+ t.Fatalf("want skipped (an image over the ceiling can never be scanned), got %s: %s%s",
+ msg.Type, msg.Error, msg.Reason)
+ }
+ if !strings.Contains(msg.Reason, "too large") {
+ t.Errorf("skip reason %q does not name the size ceiling", msg.Reason)
+ }
+ t.Logf("layer declared 1 byte against a %d byte limit, hold offered %d: %s",
+ limit, served.Load(), msg.Reason)
+ assertNoLeakedScanDirs(t, h)
+}
+
+// TestHonestlyOversizedImageIsSkipped keeps the cheap pre-check honest: a
+// record that declares its real, over-limit size is refused before a byte is
+// downloaded, and refused permanently.
+func TestHonestlyOversizedImageIsSkipped(t *testing.T) {
+ cfgBytes := validConfig(t)
+ cfgDigest := digestOf(cfgBytes)
+
+ blobs := mockhold.NewMemory().Add(cfgDigest, cfgBytes)
+ h := startScanner(t, newHold(blobs), WithMaxImageSize(1024))
+
+ seq, err := h.Hold.SendJob(jobFor(desc(cfgDigest, 1<<20, configType)))
+ if err != nil {
+ t.Fatalf("send job: %v", err)
+ }
+ msg := h.AwaitTerminal(t, seq, 30*time.Second)
+ if msg.Type != "skipped" {
+ t.Fatalf("want skipped, got %s: %s%s", msg.Type, msg.Error, msg.Reason)
+ }
+ if n := len(h.Hold.BlobRequests()); n != 0 {
+ t.Errorf("downloaded %d blobs for an image refused by the pre-check", n)
+ }
+ assertNoLeakedScanDirs(t, h)
+}
diff --git a/scanner/internal/e2e/harness.go b/scanner/internal/e2e/harness.go
new file mode 100644
index 0000000..5c94186
--- /dev/null
+++ b/scanner/internal/e2e/harness.go
@@ -0,0 +1,218 @@
+// Package e2e wires the real scanner components (WebSocket client, priority
+// queue, worker pool) to a mock hold, so scan scenarios run through the same
+// code path production uses.
+//
+// What this can and cannot cover is worth stating plainly. It drives the
+// scanner through hold behaviour: job dispatch, disconnects, blob faults,
+// artifact shapes. It cannot cover the hold's own state machine (the
+// pending/assigned/processing rows, the ack timeout, the stale-scan loop),
+// because a mock hold has no rows. That half belongs to the ScanBroadcaster
+// tests in pkg/hold/pds, and the two must be kept honest against each other:
+// whenever a scenario here encodes an assumption about what the real hold
+// does on reconnect, a hold-side test should pin that assumption.
+package e2e
+
+import (
+ "context"
+ "os"
+ "path/filepath"
+ "testing"
+ "time"
+
+ "atcr.io/scanner/internal/client"
+ "atcr.io/scanner/internal/config"
+ "atcr.io/scanner/internal/mockhold"
+ "atcr.io/scanner/internal/queue"
+ "atcr.io/scanner/internal/scan"
+)
+
+const testSecret = "test-scanner-secret"
+
+// Harness holds a running scanner attached to a mock hold.
+type Harness struct {
+ Hold *mockhold.Hold
+ Queue *queue.JobQueue
+ Client *client.HoldClient
+ Cfg *config.Config
+
+ pool *scan.WorkerPool
+}
+
+// Option customizes the harness before the scanner starts.
+type Option func(*config.Config)
+
+// WithWorkers sets the worker pool size. The default of 1 matches the shipped
+// config default and keeps job ordering deterministic.
+func WithWorkers(n int) Option {
+ return func(c *config.Config) { c.Scanner.Workers = n }
+}
+
+// WithMaxImageSize sets the compressed-size ceiling checked before download.
+func WithMaxImageSize(n int64) Option {
+ return func(c *config.Config) { c.Vuln.MaxImageSize = n }
+}
+
+// VulnDBEnv gates every scenario that runs Grype against the real
+// vulnerability database. It follows the ATCR_SCANNER_PERF precedent: a clean
+// checkout runs the ordinary suite offline and in the same time it always did,
+// and the expensive coverage is opt-in.
+const VulnDBEnv = "ATCR_SCANNER_VULNDB"
+
+// VulnDBDir is where internal/mockhold/testdata/fetch-vulndb.sh caches the
+// Grype database. It is gitignored: ~2 GB of SQLite.
+func VulnDBDir() string {
+ return filepath.Join("..", "mockhold", "testdata", "vulndb")
+}
+
+// WithVulnDB turns Grype on and points it at the cached database, skipping the
+// test when either the env gate is unset or the database has not been fetched
+// — the same shape layoutFor uses for blob fixtures, for the same reason:
+// this fixture is too large to commit and too slow to build implicitly.
+//
+// The database is a package-level global in scan guarded by an RWMutex, so the
+// first scenario in a test binary pays the load (tens of seconds, and the
+// mmap'd file is ~2 GB) and every one after it reuses the provider. That is
+// per test binary, not per package: the stubbed-loader tests in
+// internal/scan/vulndb_refresh_test.go link their own binary and their own
+// copy of these globals, so nothing here reaches them and they stay offline.
+//
+// Note that db_path is deliberately the shared cache rather than a t.TempDir:
+// a per-test copy would re-download 2 GB every run. The consequence is that a
+// scenario running against a database more than 14 days old will refresh it
+// in-process (Grype's MaxAllowedBuiltAge), which is slow but correct. Run
+// fetch-vulndb.sh to refresh it out of band.
+func WithVulnDB(t *testing.T) Option {
+ t.Helper()
+
+ if os.Getenv(VulnDBEnv) != "1" {
+ t.Skipf("set %s=1 to run scenarios against the real Grype database", VulnDBEnv)
+ }
+ dir := VulnDBDir()
+ if _, err := os.Stat(filepath.Join(dir, "6", "import.json")); err != nil {
+ t.Skipf("no Grype database at %s; run scanner/internal/mockhold/testdata/fetch-vulndb.sh", dir)
+ }
+ abs, err := filepath.Abs(dir)
+ if err != nil {
+ t.Fatalf("resolve vulndb path: %v", err)
+ }
+
+ return func(c *config.Config) {
+ c.Vuln.Enabled = true
+ c.Vuln.DBPath = abs
+ }
+}
+
+// testJobCooldown replaces the production 10s inter-job pause. Scenarios here
+// run several jobs through one worker, and the real cooldown would dominate
+// their runtime entirely. A scenario that genuinely depends on the production
+// pause should restore it explicitly and say why.
+const testJobCooldown = 10 * time.Millisecond
+
+// Start brings up a mock hold serving blobs, then a real scanner connected to
+// it. Everything is torn down via t.Cleanup.
+//
+// Grype is disabled by default. Enabling it would pull a multi-hundred-MB
+// vulnerability database on first run and make results move as the upstream
+// feed changes; the pipeline still runs blob download plus a real Syft
+// catalog, which is what these scenarios exercise. Database refresh behaviour
+// is covered separately by stubbing loadVulnDB in the scan package, and the
+// matcher itself by the opt-in scenarios in vulnreport_test.go, which pass
+// WithVulnDB to turn Grype on against a locally cached database.
+func Start(t *testing.T, blobs mockhold.BlobSource, opts ...Option) *Harness {
+ t.Helper()
+
+ hold := mockhold.New(blobs, mockhold.WithSecret(testSecret))
+ t.Cleanup(hold.Close)
+
+ cfg := config.DefaultConfig()
+ cfg.Hold.URL = hold.URL()
+ cfg.Hold.Secret = testSecret
+ cfg.Scanner.Workers = 1
+ cfg.Vuln.Enabled = false
+ cfg.Vuln.TmpDir = t.TempDir()
+
+ // WorkerPool.Start exports TMPDIR process-wide and deliberately never
+ // restores it, which is right in production (Grype's database download and
+ // stereoscope's extraction must not land on a small tmpfs) but leaks
+ // between tests: the next test's t.TempDir() would resolve against this
+ // test's directory, which cleanup has already removed. Restore it here.
+ // These tests must not run in parallel for the same reason.
+ origTmpDir, hadTmpDir := os.LookupEnv("TMPDIR")
+ t.Cleanup(func() {
+ if hadTmpDir {
+ os.Setenv("TMPDIR", origTmpDir)
+ return
+ }
+ os.Unsetenv("TMPDIR")
+ })
+ for _, opt := range opts {
+ opt(cfg)
+ }
+
+ restoreCooldown := scan.JobCooldown
+ scan.JobCooldown = testJobCooldown
+
+ q := queue.NewJobQueue(cfg.Scanner.QueueSize)
+ c := client.NewHoldClient(cfg.Hold.URL, cfg.Hold.Secret, q)
+ pool := scan.NewWorkerPool(cfg, q, c)
+
+ ctx, cancel := context.WithCancel(context.Background())
+ pool.Start(ctx)
+ go c.Connect()
+
+ t.Cleanup(func() {
+ // Order matters. Workers read scan.JobCooldown on every loop
+ // iteration, so restoring it before they have exited is a data race
+ // that -race reports against worker.go. Cancelling the context and
+ // closing the queue is what releases them (Dequeue returns nil once
+ // closed, and the cooldown select wakes on ctx.Done), so join them
+ // before touching the global back.
+ cancel()
+ c.Close()
+ q.Close()
+ pool.Wait()
+ scan.JobCooldown = restoreCooldown
+ })
+
+ if err := hold.WaitForScanner(10 * time.Second); err != nil {
+ t.Fatalf("scanner never connected: %v", err)
+ }
+
+ return &Harness{Hold: hold, Queue: q, Client: c, Cfg: cfg, pool: pool}
+}
+
+// AwaitTerminal waits for the scanner's final word on a job: a result, an
+// error, or a skip. Returning whichever arrived (rather than asserting a type)
+// lets a test report what actually happened, which matters most for the
+// scenarios where the current behaviour is the thing under examination.
+func (h *Harness) AwaitTerminal(t *testing.T, seq int64, timeout time.Duration) mockhold.Message {
+ t.Helper()
+ msg, err := h.Hold.WaitForMessage(func(m mockhold.Message) bool {
+ return m.Seq == seq && (m.Type == "result" || m.Type == "error" || m.Type == "skipped")
+ }, timeout)
+ if err != nil {
+ t.Fatalf("no terminal message for seq %d within %s; transcript: %s",
+ seq, timeout, describe(h.Hold.Transcript()))
+ }
+ return msg
+}
+
+func describe(msgs []mockhold.Message) string {
+ if len(msgs) == 0 {
+ return "(empty)"
+ }
+ out := ""
+ for _, m := range msgs {
+ if out != "" {
+ out += ", "
+ }
+ out += m.Type
+ if m.Error != "" {
+ out += "(" + m.Error + ")"
+ }
+ if m.Reason != "" {
+ out += "(" + m.Reason + ")"
+ }
+ }
+ return out
+}
diff --git a/scanner/internal/e2e/pipeline_test.go b/scanner/internal/e2e/pipeline_test.go
new file mode 100644
index 0000000..f4f9ae4
--- /dev/null
+++ b/scanner/internal/e2e/pipeline_test.go
@@ -0,0 +1,223 @@
+package e2e
+
+import (
+ "os"
+ "path/filepath"
+ "slices"
+ "testing"
+ "time"
+
+ "atcr.io/scanner/internal/mockhold"
+)
+
+// layoutFor returns an OCI layout source for a fixture pulled by
+// testdata/fetch-blobs.sh, skipping the test when it is absent. Blob fixtures
+// are gitignored because they are megabytes of real container layers, so a
+// clean checkout runs every descriptor-only scenario and skips only these.
+func layoutFor(t *testing.T, name string) mockhold.BlobSource {
+ t.Helper()
+ dir := filepath.Join("..", "mockhold", "testdata", "blobs", name)
+ if _, err := os.Stat(filepath.Join(dir, "oci-layout")); err != nil {
+ t.Skipf("fixture %q not present; run scanner/internal/mockhold/testdata/fetch-blobs.sh", name)
+ }
+ return mockhold.NewOCILayout(dir)
+}
+
+// corpusOne returns the single corpus manifest with the given digest.
+func corpusOne(t *testing.T, digest string) mockhold.Manifest {
+ t.Helper()
+ all, err := mockhold.Corpus()
+ if err != nil {
+ t.Fatalf("load corpus: %v", err)
+ }
+ for _, m := range all {
+ if m.Digest == digest {
+ return m
+ }
+ }
+ t.Fatalf("digest %s not in corpus", digest)
+ return mockhold.Manifest{}
+}
+
+// hsmOperator is the fixture image fetch-blobs.sh pulls: a real single-layer
+// image whose descriptors are in the corpus and whose bytes are in the layout.
+const hsmOperator = "sha256:1cfa4e2b09e127b9c4ed43578d3f3c18e7d44ea47b9ea98475c0cbe9086525f8"
+
+// TestScanRealImage runs the whole pipeline against a real image: the scanner
+// downloads the config and layer through the mock's getBlob indirection,
+// reassembles an OCI layout, and hands it to Syft.
+//
+// The assertion that matters most is not "an SBOM came back" but that the
+// blobs the scanner asked for are exactly the ones the manifest declares. That
+// is the join between two independently sourced halves of the fixture: the
+// descriptors come from the PDS, the bytes from skopeo, and nothing keeps them
+// in agreement except the digests being content addresses.
+func TestScanRealImage(t *testing.T) {
+ // The harness disables Grype to avoid a multi-hundred-MB database
+ // download, so this scan completes with no vulnerability summary. That is
+ // the case worker.go used to dereference unconditionally, taking the whole
+ // test binary down with it; a successful summary-less scan must now log
+ // and return like any other.
+ blobs := layoutFor(t, "hsm-secrets-operator")
+ h := Start(t, blobs)
+
+ manifest := corpusOne(t, hsmOperator)
+ seq, err := h.Hold.SendJob(manifest.Job())
+ if err != nil {
+ t.Fatalf("send job: %v", err)
+ }
+
+ msg := h.AwaitTerminal(t, seq, 3*time.Minute)
+ if msg.Type != "result" {
+ t.Fatalf("want result, got %s: %s%s", msg.Type, msg.Error, msg.Reason)
+ }
+ if msg.SBOM == "" {
+ t.Error("result carried no SBOM")
+ }
+ // Grype is off, so there is no summary. A nil summary means "not scanned
+ // for vulnerabilities", which is not the same as "scanned, found none" —
+ // the scanner must not invent a zeroed one to fill the hole.
+ if msg.Summary != nil {
+ t.Errorf("vuln scanning is disabled but the result carried a summary: %+v", msg.Summary)
+ }
+
+ want := manifest.Digests()
+ var got []string
+ for _, r := range h.Hold.BlobRequests() {
+ got = append(got, r.Digest)
+ }
+ if len(got) != len(want) {
+ t.Errorf("fetched %d blobs, manifest declares %d: got %v", len(got), len(want), got)
+ }
+ for _, w := range want {
+ if !slices.Contains(got, mockhold.DigestHex(w)) {
+ t.Errorf("declared blob %s was never fetched", w)
+ }
+ }
+}
+
+// TestSkipsAttestation feeds the scanner the shape that reached production: an
+// in-toto SLSA provenance manifest carrying an ordinary image config, so the
+// config media type alone waves it through. It must come back as "skipped"
+// rather than "error", because the hold retries failures on the stale-scan
+// loop and never retries skips.
+//
+// No blob bytes are involved: the scanner refuses before it downloads
+// anything, which the blob-request count asserts directly.
+func TestSkipsAttestation(t *testing.T) {
+ h := Start(t, mockhold.NewMemory())
+
+ attestations, err := mockhold.CorpusByShape(mockhold.ShapeAttestation)
+ if err != nil {
+ t.Fatalf("load corpus: %v", err)
+ }
+ if len(attestations) == 0 {
+ t.Fatal("corpus contains no attestation manifests")
+ }
+
+ for _, m := range mockhold.Representatives(attestations) {
+ seq, err := h.Hold.SendJob(m.Job())
+ if err != nil {
+ t.Fatalf("send job: %v", err)
+ }
+ msg := h.AwaitTerminal(t, seq, 30*time.Second)
+ if msg.Type != "skipped" {
+ t.Errorf("%s (%s): want skipped, got %s: %s%s",
+ m.Repository, m.Digest[:19], msg.Type, msg.Error, msg.Reason)
+ }
+ }
+
+ if n := len(h.Hold.BlobRequests()); n != 0 {
+ t.Errorf("scanner downloaded %d blobs for artifacts it refused", n)
+ }
+}
+
+// TestSkipsHelm covers the other refusal path, where the config media type is
+// itself unscannable.
+func TestSkipsHelm(t *testing.T) {
+ h := Start(t, mockhold.NewMemory())
+
+ charts, err := mockhold.CorpusByShape(mockhold.ShapeHelm)
+ if err != nil {
+ t.Fatalf("load corpus: %v", err)
+ }
+ if len(charts) == 0 {
+ t.Fatal("corpus contains no helm manifests")
+ }
+
+ for _, m := range mockhold.Representatives(charts) {
+ seq, err := h.Hold.SendJob(m.Job())
+ if err != nil {
+ t.Fatalf("send job: %v", err)
+ }
+ msg := h.AwaitTerminal(t, seq, 30*time.Second)
+ if msg.Type != "skipped" {
+ t.Errorf("%s: want skipped, got %s: %s%s",
+ m.Repository, msg.Type, msg.Error, msg.Reason)
+ }
+ }
+}
+
+// TestIndexManifestIsRetriedForever documents current behaviour rather than
+// desired behaviour, and is deliberately written to pass today.
+//
+// A multi-arch index has no config and no layers. skipReason returns empty for
+// it: there is no config media type to match, and the layer check is skipped
+// when the layer list is empty (worker_skip_test.go asserts exactly that, as
+// "no layers at all is left to the pipeline"). The pipeline then fails in
+// buildOCILayout with "config blob has empty digest", and an error is
+// retryable, so the hold's stale-scan loop will re-offer this job forever.
+//
+// Nothing in the scanner prevents this. The only thing that does is on the far
+// side of the WebSocket: both of the hold's enqueue paths gate on
+// HasScannableContent (!IsMultiArch() && !IsReferrer()), so an index never
+// gets dispatched in production. That guard is load-bearing and nothing on the
+// scanner side would notice if it regressed, which is the point of pinning the
+// behaviour here.
+func TestIndexManifestIsRetriedForever(t *testing.T) {
+ h := Start(t, mockhold.NewMemory())
+
+ indexes, err := mockhold.CorpusByShape(mockhold.ShapeIndex)
+ if err != nil {
+ t.Fatalf("load corpus: %v", err)
+ }
+ if len(indexes) == 0 {
+ t.Fatal("corpus contains no index manifests")
+ }
+
+ m := indexes[0]
+ seq, err := h.Hold.SendJob(m.Job())
+ if err != nil {
+ t.Fatalf("send job: %v", err)
+ }
+
+ msg := h.AwaitTerminal(t, seq, 30*time.Second)
+ if msg.Type != "error" {
+ t.Fatalf("index manifest: expected the current retryable-error behaviour, got %s (%s%s)",
+ msg.Type, msg.Error, msg.Reason)
+ }
+ t.Logf("index manifest returns a retryable error, not a skip: %q", msg.Error)
+}
+
+// TestOversizeImageIsRejectedBeforeDownload pins the cheap half of the size
+// guard: it sums the sizes the manifest claims, so no blob source is needed
+// and nothing is fetched. An image does not shrink, so the verdict is a
+// permanent skip rather than a failure the stale loop re-offers forever. The
+// transferred-bytes half of the ceiling lives in blob_integrity_test.go.
+func TestOversizeImageIsRejectedBeforeDownload(t *testing.T) {
+ h := Start(t, mockhold.NewMemory(), WithMaxImageSize(1024))
+
+ manifest := corpusOne(t, hsmOperator)
+ seq, err := h.Hold.SendJob(manifest.Job())
+ if err != nil {
+ t.Fatalf("send job: %v", err)
+ }
+
+ msg := h.AwaitTerminal(t, seq, 30*time.Second)
+ if msg.Type != "skipped" {
+ t.Fatalf("want skipped for oversize image, got %s: %s%s", msg.Type, msg.Error, msg.Reason)
+ }
+ if n := len(h.Hold.BlobRequests()); n != 0 {
+ t.Errorf("oversize image still fetched %d blobs", n)
+ }
+}
diff --git a/scanner/internal/e2e/protocol_test.go b/scanner/internal/e2e/protocol_test.go
new file mode 100644
index 0000000..e698351
--- /dev/null
+++ b/scanner/internal/e2e/protocol_test.go
@@ -0,0 +1,722 @@
+package e2e
+
+import (
+ "context"
+ "encoding/json"
+ "net/http"
+ "net/http/httptest"
+ "os"
+ "strings"
+ "testing"
+ "time"
+
+ scanner "atcr.io/scanner"
+ "atcr.io/scanner/internal/client"
+ "atcr.io/scanner/internal/config"
+ "atcr.io/scanner/internal/mockhold"
+ "atcr.io/scanner/internal/queue"
+ "atcr.io/scanner/internal/scan"
+)
+
+// --- helpers ---------------------------------------------------------------
+
+// silenceWindow is how long a test waits before concluding the scanner sent
+// nothing back. It is short on purpose: the failure mode being pinned is
+// "nothing, ever", and the hold's own ackTimeout is five minutes, so anything
+// the scanner has not said within a second here it will not say at all.
+const silenceWindow = 1 * time.Second
+
+// synthJob builds a descriptor-only container image job. The digests point at
+// nothing, which is fine for every scenario that either refuses the job before
+// downloading or deliberately stalls the download.
+func synthJob(repo string) *scanner.ScanJob {
+ return &scanner.ScanJob{
+ ManifestDigest: "sha256:" + strings.Repeat("a", 64),
+ Repository: repo,
+ Tag: "latest",
+ Tier: "deckhand",
+ HoldDID: "did:web:hold.example",
+ Config: scanner.BlobDescriptor{
+ Digest: "sha256:" + strings.Repeat("c", 64),
+ Size: 100,
+ MediaType: "application/vnd.oci.image.config.v1+json",
+ },
+ Layers: []scanner.BlobDescriptor{{
+ Digest: "sha256:" + strings.Repeat("1", 64),
+ Size: 200,
+ MediaType: "application/vnd.oci.image.layer.v1.tar+gzip",
+ }},
+ }
+}
+
+// helmJob returns a corpus Helm chart, the cheapest job that reaches a
+// terminal message: the scanner refuses it on config media type alone, before
+// any tmp dir or download work, so it is a reliable liveness probe.
+func helmJob(t *testing.T) *scanner.ScanJob {
+ t.Helper()
+ charts, err := mockhold.CorpusByShape(mockhold.ShapeHelm)
+ if err != nil {
+ t.Fatalf("load corpus: %v", err)
+ }
+ if len(charts) == 0 {
+ t.Fatal("corpus contains no helm manifests")
+ }
+ return charts[0].Job()
+}
+
+// rawJob renders a job frame from a map, so a test can omit a field, or give
+// it a shape the scanner cannot decode.
+func rawJob(t *testing.T, fields map[string]any) []byte {
+ t.Helper()
+ data, err := json.Marshal(fields)
+ if err != nil {
+ t.Fatalf("marshal raw job: %v", err)
+ }
+ return data
+}
+
+// expectSilence fails if any message matching match arrives within the window.
+func expectSilence(t *testing.T, h *Harness, match func(mockhold.Message) bool, window time.Duration, what string) {
+ t.Helper()
+ msg, err := h.Hold.WaitForMessage(match, window)
+ if err == nil {
+ t.Fatalf("%s: expected no reply, got %s for seq %d", what, msg.Type, msg.Seq)
+ }
+}
+
+// forSeq matches any message about one job.
+func forSeq(seq int64) func(mockhold.Message) bool {
+ return func(m mockhold.Message) bool { return m.Seq == seq }
+}
+
+// assertAlive proves the WebSocket survived whatever the previous step did to
+// it, by pushing a job the scanner is guaranteed to answer.
+func assertAlive(t *testing.T, h *Harness) {
+ t.Helper()
+ seq, err := h.Hold.SendJob(helmJob(t))
+ if err != nil {
+ t.Fatalf("connection did not survive: %v", err)
+ }
+ if msg := h.AwaitTerminal(t, seq, 30*time.Second); msg.Type != "skipped" {
+ t.Fatalf("liveness probe: want skipped, got %s", msg.Type)
+ }
+}
+
+// --- 1. frames the scanner cannot parse -------------------------------------
+
+// TestUnparseableFramesAreAnsweredWithSkipped is the central protocol finding,
+// and the shape of a nine-day outage.
+//
+// connectOnce decodes three things — the frame, then the config sub-document,
+// then the layers sub-document — and every failure branch used to be
+// slog.Error followed by continue, sending nothing back at all. The ack was
+// sent only after both sub-document unmarshals, so a job whose config or
+// layers did not decode was never even acknowledged.
+//
+// The hold has already written status='assigned' for that seq before it wrote
+// the frame. Its only escape was the five-minute ackTimeout in
+// reDispatchTimedOut, after which the row was re-offered — to the same
+// scanner, which dropped it again for exactly the same reason, because a
+// decoding disagreement is permanent. And because hasActiveJobs counts
+// 'assigned' rows and dispatchLoop admits one proactive candidate at a time
+// behind waitForCapacity, one such row meant no proactive scan was ever
+// dispatched again, deployment-wide.
+//
+// The scanner now answers "skipped" for any frame that carries a usable seq.
+// Skipped is the correct verdict rather than "error": the hold retries
+// failures on the rescan interval, and no retry of an undecodable frame can
+// ever succeed, whereas handleSkipped writes a terminal record and releases
+// the row and the in-flight digest for good. The hold-side half is
+// TestScanSkipped_RetiresAnUndecodableFrame in
+// pkg/hold/pds/scan_broadcaster_stuck_test.go.
+func TestUnparseableFramesAreAnsweredWithSkipped(t *testing.T) {
+ h := Start(t, mockhold.NewMemory())
+
+ base := func(seq int64) map[string]any {
+ return map[string]any{
+ "type": "job",
+ "seq": seq,
+ "manifestDigest": "sha256:" + strings.Repeat("a", 64),
+ "repository": "probe/unparseable",
+ "tag": "latest",
+ "userDid": "did:plc:probe",
+ "holdDid": "did:web:hold.example",
+ "holdEndpoint": h.Hold.URL(),
+ "tier": "deckhand",
+ "config": map[string]any{"digest": "sha256:" + strings.Repeat("c", 64), "size": 1, "mediaType": "application/vnd.oci.image.config.v1+json"},
+ "layers": []any{},
+ }
+ }
+
+ cases := []struct {
+ name string
+ frame func(seq int64) []byte
+ }{
+ {"config is a string", func(seq int64) []byte {
+ f := base(seq)
+ f["config"] = "not-an-object"
+ return rawJob(t, f)
+ }},
+ {"config field absent", func(seq int64) []byte {
+ f := base(seq)
+ delete(f, "config")
+ return rawJob(t, f)
+ }},
+ {"layers is an object", func(seq int64) []byte {
+ f := base(seq)
+ f["layers"] = map[string]any{"oops": 1}
+ return rawJob(t, f)
+ }},
+ {"layers field absent", func(seq int64) []byte {
+ f := base(seq)
+ delete(f, "layers")
+ return rawJob(t, f)
+ }},
+ // encoding/json records the first type error and keeps decoding the
+ // rest of the object, so a frame the top-level unmarshal rejects can
+ // still yield the seq that addresses the hold's row. That row must be
+ // answered too, which is why the reply is keyed on the seq rather than
+ // on which of the three decodes failed.
+ {"frame has a type-mismatched field", func(seq int64) []byte {
+ f := base(seq)
+ f["tier"] = 12345
+ return rawJob(t, f)
+ }},
+ }
+
+ for _, tc := range cases {
+ t.Run(tc.name, func(t *testing.T) {
+ seq := h.Hold.NextSeq()
+ if err := h.Hold.SendRaw(tc.frame(seq)); err != nil {
+ t.Fatalf("send raw frame: %v", err)
+ }
+ msg, err := h.Hold.WaitForMessage(func(m mockhold.Message) bool {
+ return m.Seq == seq && m.Type != "ack"
+ }, 10*time.Second)
+ if err != nil {
+ t.Fatalf("%s: no reply for seq %d; the hold's row stays "+
+ "assigned and is re-offered to this same scanner forever. "+
+ "transcript: %s", tc.name, seq, describe(h.Hold.Transcript()))
+ }
+ if msg.Type != "skipped" {
+ t.Fatalf("%s: reply for seq %d was %s (%s%s), want skipped: an "+
+ "undecodable frame is a permanent condition and the hold "+
+ "retries anything it records as a failure",
+ tc.name, seq, msg.Type, msg.Error, msg.Reason)
+ }
+ if msg.Reason == "" {
+ t.Errorf("%s: skipped with no reason; the hold stores it on the "+
+ "scan record and it is all a user ever sees", tc.name)
+ }
+ })
+ }
+
+ // The connection survives all of it, which is what made the retry loop
+ // infinite rather than self-limiting.
+ assertAlive(t, h)
+}
+
+// TestFramesWithNoUsableSeqAreDroppedSilently is the deliberate exception.
+//
+// A frame that does not decode far enough to yield a seq addresses no job:
+// there is no row to retire and no seq to put in a reply, so logging is the
+// only thing left. This is safe in a way the config/layers case never was —
+// the hold writes status='assigned' keyed by seq before it sends, so a frame
+// whose seq never made it onto the wire cannot be the frame that stranded a
+// row. The connection must survive, since the hold will keep using it.
+func TestFramesWithNoUsableSeqAreDroppedSilently(t *testing.T) {
+ h := Start(t, mockhold.NewMemory())
+
+ for _, tc := range []struct {
+ name string
+ frame []byte
+ }{
+ {"frame is not JSON", []byte("{ this is not json")},
+ {"frame is a JSON array", []byte(`[1,2,3]`)},
+ {"frame carries seq 0", rawJob(t, map[string]any{"type": "job", "seq": 0, "config": "nope"})},
+ } {
+ t.Run(tc.name, func(t *testing.T) {
+ before := len(h.Hold.Transcript())
+ if err := h.Hold.SendRaw(tc.frame); err != nil {
+ t.Fatalf("send raw frame: %v", err)
+ }
+ expectSilence(t, h, func(m mockhold.Message) bool {
+ return m.Seq == 0
+ }, silenceWindow, tc.name)
+ if got := len(h.Hold.Transcript()); got != before {
+ t.Errorf("%s: scanner sent %d messages for an unaddressable frame",
+ tc.name, got-before)
+ }
+ })
+ }
+
+ assertAlive(t, h)
+}
+
+// TestNullConfigIsAckedThenFailsRetryably covers the shape that *does* decode:
+// a JSON null unmarshals into a zero BlobDescriptor without error, so the job
+// is acked and enters the pipeline, then dies in buildOCILayout on the empty
+// config digest. That is reported as "error", which the hold treats as
+// transient and retries on the rescan interval forever, even though no retry
+// can ever succeed: nothing about a null config will change.
+func TestNullConfigIsAckedThenFailsRetryably(t *testing.T) {
+ h := Start(t, mockhold.NewMemory())
+
+ seq := h.Hold.NextSeq()
+ frame := rawJob(t, map[string]any{
+ "type": "job",
+ "seq": seq,
+ "manifestDigest": "sha256:" + strings.Repeat("b", 64),
+ "repository": "probe/null-config",
+ "tag": "latest",
+ "userDid": "did:plc:probe",
+ "holdDid": "did:web:hold.example",
+ "holdEndpoint": h.Hold.URL(),
+ "tier": "deckhand",
+ "config": nil,
+ "layers": []any{map[string]any{
+ "digest": "sha256:" + strings.Repeat("1", 64),
+ "size": 10,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip",
+ }},
+ })
+ if err := h.Hold.SendRaw(frame); err != nil {
+ t.Fatalf("send raw frame: %v", err)
+ }
+
+ if _, err := h.Hold.WaitForMessage(func(m mockhold.Message) bool {
+ return m.Seq == seq && m.Type == "ack"
+ }, 10*time.Second); err != nil {
+ t.Fatalf("null config was not acked: %v", err)
+ }
+
+ msg := h.AwaitTerminal(t, seq, 30*time.Second)
+ if msg.Type != "error" {
+ t.Fatalf("want the current retryable-error behaviour, got %s (%s%s)", msg.Type, msg.Error, msg.Reason)
+ }
+ if !strings.Contains(msg.Error, "empty digest") {
+ t.Errorf("unexpected error text %q", msg.Error)
+ }
+ t.Logf("null config is a permanent condition reported as a retryable error: %q", msg.Error)
+}
+
+// TestUnknownMessageTypeIsIgnored pins the benign half: a frame the scanner
+// does not recognise is logged and skipped, the connection survives, and
+// nothing is sent back. Harmless today because the hold only ever sends
+// "job", but it means any future message type is silently swallowed by an
+// older scanner rather than refused.
+func TestUnknownMessageTypeIsIgnored(t *testing.T) {
+ h := Start(t, mockhold.NewMemory())
+
+ seq := h.Hold.NextSeq()
+ if err := h.Hold.SendRaw(rawJob(t, map[string]any{"type": "cancel", "seq": seq})); err != nil {
+ t.Fatalf("send raw frame: %v", err)
+ }
+ expectSilence(t, h, forSeq(seq), silenceWindow, "unknown message type")
+ assertAlive(t, h)
+}
+
+// --- 2. duplicate delivery --------------------------------------------------
+
+// TestDuplicateSeqIsProcessedTwice shows the scanner has no idea it has seen a
+// job before. Nothing dedupes on seq or on manifest digest: the job is acked
+// twice, queued twice, and scanned twice.
+//
+// On the hold side the second ack is a no-op (handleAck's UPDATE is guarded by
+// status='assigned', which the first ack already cleared) but the second
+// terminal message is not: handleSkipped/handleResult/handleError re-run
+// unconditionally, writing a second scan record to the PDS for the same
+// manifest. For a real image this is also a full second download and Syft run.
+func TestDuplicateSeqIsProcessedTwice(t *testing.T) {
+ h := Start(t, mockhold.NewMemory())
+
+ job := helmJob(t)
+ job.Seq = h.Hold.NextSeq()
+ for i := 0; i < 2; i++ {
+ dup := *job
+ if _, err := h.Hold.SendJob(&dup); err != nil {
+ t.Fatalf("send job %d: %v", i, err)
+ }
+ }
+
+ deadline := time.Now().Add(30 * time.Second)
+ var acks, terminals int
+ for time.Now().Before(deadline) {
+ acks, terminals = 0, 0
+ for _, m := range h.Hold.Transcript() {
+ if m.Seq != job.Seq {
+ continue
+ }
+ switch m.Type {
+ case "ack":
+ acks++
+ case "result", "error", "skipped":
+ terminals++
+ }
+ }
+ if acks >= 2 && terminals >= 2 {
+ break
+ }
+ time.Sleep(20 * time.Millisecond)
+ }
+
+ if acks != 2 || terminals != 2 {
+ t.Fatalf("duplicate seq: got %d acks and %d terminal messages, want 2 and 2", acks, terminals)
+ }
+ t.Logf("seq %d was acked %d times and answered %d times: no dedup anywhere in the scanner",
+ job.Seq, acks, terminals)
+}
+
+// --- 3. queue capacity ------------------------------------------------------
+
+// TestQueueFullIsReportedAsRetryableError pins what happens past the queue's
+// high-water mark: the job is acked (hold: assigned -> processing) and then
+// immediately answered with "error: scanner queue full", which the hold
+// records as a *failure*. Failures are retryable, so the same job comes back
+// on the rescan interval and will overflow again for as long as the backlog
+// persists. A capacity signal is being reported through the channel reserved
+// for scan outcomes, and it lands in the user's scan history as a failed scan.
+func TestQueueFullIsReportedAsRetryableError(t *testing.T) {
+ const queueSize = 2
+ h := Start(t, mockhold.NewMemory(),
+ WithWorkers(0), // nothing drains the queue, so the Nth job is deterministic
+ func(c *config.Config) { c.Scanner.QueueSize = queueSize })
+
+ var seqs []int64
+ for i := 0; i < queueSize+2; i++ {
+ seq, err := h.Hold.SendJob(synthJob("probe/overflow"))
+ if err != nil {
+ t.Fatalf("send job %d: %v", i, err)
+ }
+ seqs = append(seqs, seq)
+ }
+
+ for _, seq := range seqs[queueSize:] {
+ msg := h.AwaitTerminal(t, seq, 10*time.Second)
+ if msg.Type != "error" || !strings.Contains(msg.Error, "queue full") {
+ t.Fatalf("seq %d: want a queue-full error, got %s (%s%s)", seq, msg.Type, msg.Error, msg.Reason)
+ }
+ }
+
+ // And the overflowed jobs were acked first, so the hold saw them go
+ // assigned -> processing -> failed for a condition that never involved the
+ // job at all.
+ for _, seq := range seqs[queueSize:] {
+ if _, err := h.Hold.WaitForMessage(func(m mockhold.Message) bool {
+ return m.Seq == seq && m.Type == "ack"
+ }, time.Second); err != nil {
+ t.Errorf("seq %d was rejected without ever being acked", seq)
+ }
+ }
+ t.Log("queue overflow is reported as a retryable per-job failure, not as backpressure")
+}
+
+// TestZeroQueueSizeRejectsEveryJob is the configuration corner of the same
+// path. scanner.queue_size = 0 passes validation, and NewJobQueue(0) then
+// refuses every Enqueue, so a scanner that looks healthy (connected, health
+// endpoint green, workers idle) fails 100% of jobs with "scanner queue full"
+// and the hold retries all of them forever.
+func TestZeroQueueSizeRejectsEveryJob(t *testing.T) {
+ h := Start(t, mockhold.NewMemory(), func(c *config.Config) { c.Scanner.QueueSize = 0 })
+
+ seq, err := h.Hold.SendJob(helmJob(t))
+ if err != nil {
+ t.Fatalf("send job: %v", err)
+ }
+ msg := h.AwaitTerminal(t, seq, 10*time.Second)
+ if msg.Type != "error" || !strings.Contains(msg.Error, "queue full") {
+ t.Fatalf("queue_size=0: want a queue-full error, got %s (%s%s)", msg.Type, msg.Error, msg.Reason)
+ }
+}
+
+// --- 4. worker configuration ------------------------------------------------
+
+// TestZeroWorkersAcksAndStrands: scanner.workers = 0 passes validation and
+// starts a pool with no workers at all. The client still acks everything it
+// receives, so the hold moves each job to 'processing' and then waits out the
+// ten-minute processing timeout in reDispatchTimedOut before failing it. The
+// scanner logs "Scanner worker pool started workers=0" once at boot and
+// nothing else; there is no health signal that distinguishes this from idle.
+func TestZeroWorkersAcksAndStrands(t *testing.T) {
+ h := Start(t, mockhold.NewMemory(), WithWorkers(0))
+
+ seq, err := h.Hold.SendJob(helmJob(t))
+ if err != nil {
+ t.Fatalf("send job: %v", err)
+ }
+
+ if _, err := h.Hold.WaitForMessage(func(m mockhold.Message) bool {
+ return m.Seq == seq && m.Type == "ack"
+ }, 10*time.Second); err != nil {
+ t.Fatalf("job was not even acked: %v", err)
+ }
+
+ expectSilence(t, h, func(m mockhold.Message) bool {
+ return m.Seq == seq && (m.Type == "result" || m.Type == "error" || m.Type == "skipped")
+ }, 2*time.Second, "workers=0")
+
+ if n := h.Queue.Len(); n != 1 {
+ t.Errorf("queue holds %d jobs, want 1 (acked and stranded)", n)
+ }
+}
+
+// TestEmptyTmpDirFailsEveryJob: vuln.tmp_dir = "" is accepted by config
+// loading, skips the TMPDIR export in WorkerPool.Start, and then fails every
+// single job in processJob's ensureDir, because os.MkdirAll("") is an error.
+// The failure is retryable, so every job in the deployment loops forever.
+func TestEmptyTmpDirFailsEveryJob(t *testing.T) {
+ h := Start(t, mockhold.NewMemory(), func(c *config.Config) { c.Vuln.TmpDir = "" })
+
+ seq, err := h.Hold.SendJob(synthJob("probe/no-tmpdir"))
+ if err != nil {
+ t.Fatalf("send job: %v", err)
+ }
+ msg := h.AwaitTerminal(t, seq, 30*time.Second)
+ if msg.Type != "error" || !strings.Contains(msg.Error, "tmp dir") {
+ t.Fatalf("empty tmp_dir: want a tmp dir error, got %s (%s%s)", msg.Type, msg.Error, msg.Reason)
+ }
+ if n := len(h.Hold.BlobRequests()); n != 0 {
+ t.Errorf("failed before download but still fetched %d blobs", n)
+ }
+ t.Logf("every job fails with %q, retryably", msg.Error)
+}
+
+// --- 5. priority ------------------------------------------------------------
+
+// gate is an HTTP stand-in for a hold whose getBlob hangs. A job pointed at it
+// occupies a worker for exactly as long as the test wants, which is how the
+// backlog scenarios below build a queue without needing real image bytes.
+type gate struct {
+ srv *httptest.Server
+ entered chan struct{}
+ release chan struct{}
+}
+
+func newGate(t *testing.T) *gate {
+ t.Helper()
+ g := &gate{entered: make(chan struct{}, 8), release: make(chan struct{})}
+ g.srv = httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
+ select {
+ case g.entered <- struct{}{}:
+ default:
+ }
+ <-g.release
+ http.Error(w, "gate released", http.StatusNotFound)
+ }))
+ t.Cleanup(g.srv.Close)
+ return g
+}
+
+func (g *gate) waitEntered(t *testing.T, timeout time.Duration) {
+ t.Helper()
+ select {
+ case <-g.entered:
+ case <-time.After(timeout):
+ t.Fatal("worker never reached the stalled blob fetch")
+ }
+}
+
+func (g *gate) open() { close(g.release) }
+
+// TestHighTierJumpsQueuedBacklog confirms the priority heap does what it
+// claims across the real client/queue/worker path: while one worker is busy,
+// a later owner-tier job overtakes an earlier deckhand-tier one.
+//
+// It also shows the limit of that guarantee. Priority is consulted only at
+// Dequeue, so a high-tier job that arrives while the single worker is inside
+// a scan waits for that scan to finish plus the full JobCooldown. With the
+// production 10s cooldown and multi-minute scans, "priority" means position in
+// a queue, not preemption, and a saturated scanner starves the low tier
+// entirely: every owner job admitted during a scan is dequeued before any
+// deckhand job, no matter how long the deckhand job has waited.
+func TestHighTierJumpsQueuedBacklog(t *testing.T) {
+ h := Start(t, mockhold.NewMemory())
+ g := newGate(t)
+
+ // Occupy the single worker with a job whose blob fetch never returns.
+ blocker := synthJob("probe/blocker")
+ blocker.HoldEndpoint = g.srv.URL
+ blockerSeq, err := h.Hold.SendJob(blocker)
+ if err != nil {
+ t.Fatalf("send blocker: %v", err)
+ }
+ g.waitEntered(t, 15*time.Second)
+
+ // Queue a deckhand job first, then an owner job.
+ low := helmJob(t)
+ low.Tier = "deckhand"
+ lowSeq, err := h.Hold.SendJob(low)
+ if err != nil {
+ t.Fatalf("send low: %v", err)
+ }
+ high := helmJob(t)
+ high.Tier = "owner"
+ highSeq, err := h.Hold.SendJob(high)
+ if err != nil {
+ t.Fatalf("send high: %v", err)
+ }
+
+ // Both must be in the queue before the worker is freed, or the test would
+ // be measuring arrival order rather than priority.
+ for _, seq := range []int64{lowSeq, highSeq} {
+ if _, err := h.Hold.WaitForMessage(func(m mockhold.Message) bool {
+ return m.Seq == seq && m.Type == "ack"
+ }, 10*time.Second); err != nil {
+ t.Fatalf("seq %d never acked: %v", seq, err)
+ }
+ }
+ if n := h.Queue.Len(); n != 2 {
+ t.Fatalf("queue depth %d, want 2 backlogged jobs", n)
+ }
+
+ g.open()
+ h.AwaitTerminal(t, blockerSeq, 30*time.Second)
+
+ highMsg := h.AwaitTerminal(t, highSeq, 30*time.Second)
+ lowMsg := h.AwaitTerminal(t, lowSeq, 30*time.Second)
+ if !highMsg.At.Before(lowMsg.At) {
+ t.Errorf("owner-tier job finished at %s, deckhand at %s: priority did not apply",
+ highMsg.At, lowMsg.At)
+ }
+}
+
+// --- 6. shutdown ------------------------------------------------------------
+
+// TestQueueCloseDrainsRatherThanCancels documents queue.Close semantics, which
+// are not what "close" suggests. Dequeue returns nil only when the queue is
+// closed *and* empty, so a shutdown with a backlog hands every remaining job
+// to a worker rather than dropping it. Combined with HoldClient.Close having
+// already severed the socket, whatever those jobs produce is written into a
+// dead connection and lost, while the hold sits on them until the ten-minute
+// processing timeout.
+func TestQueueCloseDrainsRatherThanCancels(t *testing.T) {
+ q := queue.NewJobQueue(10)
+ for i := 0; i < 3; i++ {
+ if !q.Enqueue(&scanner.ScanJob{Seq: int64(i + 1), Tier: "deckhand"}) {
+ t.Fatalf("enqueue %d refused", i)
+ }
+ }
+
+ q.Close()
+
+ var drained []int64
+ for {
+ job := q.Dequeue()
+ if job == nil {
+ break
+ }
+ drained = append(drained, job.Seq)
+ }
+ if len(drained) != 3 {
+ t.Fatalf("Close() discarded the backlog: drained %v, want 3 jobs", drained)
+ }
+ t.Logf("Close() left %d jobs to be dequeued and scanned after shutdown began", len(drained))
+}
+
+// TestShutdownDoesNotInterruptInFlightDownload proves processJob ignores
+// context cancellation everywhere that matters. Syft and Grype take ctx, but
+// the blob fetches go through client.GetBlobPresignedURL / DownloadBlob, which
+// use a package-level http.Client with no request context at all. Cancelling
+// the pool's context while a fetch is in flight changes nothing: the worker
+// stays inside the download until the client's own five-minute timeout, and
+// WorkerPool.Wait (which cmd/scanner calls on SIGTERM, after cancel) blocks
+// for just as long. Under a typical 30-second termination grace period that is
+// a SIGKILL, with the scan directory left behind because cleanup never runs.
+func TestShutdownDoesNotInterruptInFlightDownload(t *testing.T) {
+ hold := mockhold.New(mockhold.NewMemory(), mockhold.WithSecret(testSecret))
+ t.Cleanup(hold.Close)
+
+ cfg := config.DefaultConfig()
+ cfg.Hold.URL = hold.URL()
+ cfg.Hold.Secret = testSecret
+ cfg.Scanner.Workers = 1
+ cfg.Vuln.Enabled = false
+ cfg.Vuln.TmpDir = t.TempDir()
+
+ origTmpDir, hadTmpDir := os.LookupEnv("TMPDIR")
+ t.Cleanup(func() {
+ if hadTmpDir {
+ os.Setenv("TMPDIR", origTmpDir)
+ return
+ }
+ os.Unsetenv("TMPDIR")
+ })
+
+ restoreCooldown := scan.JobCooldown
+ scan.JobCooldown = 10 * time.Millisecond
+ t.Cleanup(func() { scan.JobCooldown = restoreCooldown })
+
+ q := queue.NewJobQueue(cfg.Scanner.QueueSize)
+ c := client.NewHoldClient(cfg.Hold.URL, cfg.Hold.Secret, q)
+ pool := scan.NewWorkerPool(cfg, q, c)
+
+ ctx, cancel := context.WithCancel(context.Background())
+ pool.Start(ctx)
+ go c.Connect()
+
+ // HoldClient.Close is not idempotent (see TestHoldClientCloseIsNotIdempotent),
+ // so the cleanup must not repeat the shutdown the test itself performs.
+ var closed bool
+ t.Cleanup(func() {
+ cancel()
+ if !closed {
+ c.Close()
+ }
+ q.Close()
+ })
+
+ if err := hold.WaitForScanner(10 * time.Second); err != nil {
+ t.Fatalf("scanner never connected: %v", err)
+ }
+
+ g := newGate(t)
+ job := synthJob("probe/shutdown")
+ job.HoldEndpoint = g.srv.URL
+ if _, err := hold.SendJob(job); err != nil {
+ t.Fatalf("send job: %v", err)
+ }
+ g.waitEntered(t, 15*time.Second)
+
+ // This is the shutdown sequence cmd/scanner runs on SIGTERM.
+ cancel()
+ c.Close()
+ closed = true
+ q.Close()
+
+ waited := make(chan struct{})
+ go func() { pool.Wait(); close(waited) }()
+
+ select {
+ case <-waited:
+ t.Fatal("worker exited on context cancellation; the download path now honours ctx")
+ case <-time.After(500 * time.Millisecond):
+ // Expected: the worker is still inside an uncancellable HTTP fetch.
+ }
+
+ g.open()
+ select {
+ case <-waited:
+ case <-time.After(30 * time.Second):
+ t.Fatal("worker never exited even after the download completed")
+ }
+ t.Log("context cancellation does not reach blob downloads; shutdown waits on the HTTP timeout")
+}
+
+// TestHoldClientCloseIsNotIdempotent pins a sharp edge rather than a live bug:
+// HoldClient.Close closes c.done unconditionally, so a second call panics the
+// process with "close of closed channel". cmd/scanner calls it exactly once
+// today, which is the only reason this is not already an incident, and there
+// is no guard if a future shutdown path (a health-check restart, a reconnect
+// supervisor) calls it again.
+func TestHoldClientCloseIsNotIdempotent(t *testing.T) {
+ c := client.NewHoldClient("ws://127.0.0.1:1", "secret", queue.NewJobQueue(1))
+ c.Close()
+
+ defer func() {
+ if recover() == nil {
+ t.Fatal("Close is idempotent now; this finding is fixed and the test should be inverted")
+ }
+ }()
+ c.Close()
+ t.Fatal("unreachable")
+}
diff --git a/scanner/internal/e2e/stuck_test.go b/scanner/internal/e2e/stuck_test.go
new file mode 100644
index 0000000..bbf7eeb
--- /dev/null
+++ b/scanner/internal/e2e/stuck_test.go
@@ -0,0 +1,571 @@
+package e2e
+
+import (
+ "bytes"
+ "crypto/sha256"
+ "fmt"
+ "io"
+ "sync"
+ "testing"
+ "time"
+
+ scanner "atcr.io/scanner"
+ "atcr.io/scanner/internal/mockhold"
+)
+
+// Scenarios for the "the scanner gets stuck and never finishes" report.
+//
+// Everything here is a characterization test: it asserts what the scanner does
+// today, and the comment says where that is wrong. Nothing asserts the fixed
+// behaviour, so the suite stays green for whoever is working next door.
+//
+// Every scenario in this file is built so no job ever reaches a successful
+// result: each one fails on a blob download instead. That was originally
+// forced on them (a successful scan used to panic the test binary on
+// worker.go's unconditional result.Summary dereference) and is now simply what
+// keeps them focused on the dispatch and reconnect behaviour under test.
+//
+// The hold-side counterparts are in pkg/hold/pds/scan_broadcaster_stuck_test.go.
+
+// stuckSource is a BlobSource that can stall, delay, and selectively 404, and
+// records enough about each Open for a test to tell serial downloads from
+// concurrent ones.
+//
+// Blobs it does hold are served as junk bytes. Nothing verifies that a
+// downloaded blob hashes to the digest that asked for it (buildOCILayout writes
+// whatever arrives straight to blobs/sha256/), so junk is enough to make a
+// download "succeed" and move the pipeline on to the next one.
+type stuckSource struct {
+ mu sync.Mutex
+ opens []time.Time
+ perDigest map[string]int
+ active int
+ maxActive int
+
+ // delay is slept inside Open, before answering.
+ delay time.Duration
+ // gate, when non-nil, blocks Open until it is closed.
+ gate chan struct{}
+ // have lists digests to answer with junk bytes. Anything else 404s.
+ have map[string]int
+}
+
+func newStuckSource() *stuckSource {
+ return &stuckSource{
+ perDigest: make(map[string]int),
+ have: make(map[string]int),
+ }
+}
+
+// gated makes every Open block until Release is called.
+func (s *stuckSource) gated() *stuckSource {
+ s.gate = make(chan struct{})
+ return s
+}
+
+// slow makes every Open take d before answering.
+func (s *stuckSource) slow(d time.Duration) *stuckSource {
+ s.delay = d
+ return s
+}
+
+// servingJunk registers n junk bytes for download and returns the digest that
+// names them. Anything not registered answers 404.
+//
+// The digest is derived from the content rather than from a scenario label,
+// which it has to be: the scanner hashes what arrives and refuses bytes that
+// are not what the descriptor said, so a source can no longer answer an
+// arbitrary digest with arbitrary bytes.
+func (s *stuckSource) servingJunk(n int) string {
+ digest := digestOf(bytes.Repeat([]byte("x"), n))
+ s.have[mockhold.DigestHex(digest)] = n
+ return digest
+}
+
+// Release unblocks a gated source. Safe to call more than once, and always
+// registered as a cleanup: httptest.Server.Close waits for in-flight requests,
+// so a gate that is never opened deadlocks the teardown.
+func (s *stuckSource) Release() {
+ s.mu.Lock()
+ defer s.mu.Unlock()
+ if s.gate != nil {
+ select {
+ case <-s.gate:
+ default:
+ close(s.gate)
+ }
+ }
+}
+
+func (s *stuckSource) Open(digest string) (io.ReadCloser, int64, error) {
+ hex := mockhold.DigestHex(digest)
+
+ s.mu.Lock()
+ s.opens = append(s.opens, time.Now())
+ s.perDigest[hex]++
+ s.active++
+ if s.active > s.maxActive {
+ s.maxActive = s.active
+ }
+ gate := s.gate
+ delay, size := s.delay, s.have[hex]
+ _, served := s.have[hex]
+ s.mu.Unlock()
+
+ defer func() {
+ s.mu.Lock()
+ s.active--
+ s.mu.Unlock()
+ }()
+
+ if gate != nil {
+ select {
+ case <-gate:
+ case <-time.After(90 * time.Second):
+ // Safety net: a test that forgets to Release must fail on its own
+ // assertions rather than wedging the whole package.
+ }
+ }
+ if delay > 0 {
+ time.Sleep(delay)
+ }
+
+ if !served {
+ return nil, 0, fmt.Errorf("%w: %s", mockhold.ErrBlobNotFound, digest)
+ }
+ return io.NopCloser(bytes.NewReader(bytes.Repeat([]byte("x"), size))), int64(size), nil
+}
+
+func (s *stuckSource) openCount() int {
+ s.mu.Lock()
+ defer s.mu.Unlock()
+ return len(s.opens)
+}
+
+func (s *stuckSource) countFor(digest string) int {
+ s.mu.Lock()
+ defer s.mu.Unlock()
+ return s.perDigest[mockhold.DigestHex(digest)]
+}
+
+func (s *stuckSource) peakConcurrency() int {
+ s.mu.Lock()
+ defer s.mu.Unlock()
+ return s.maxActive
+}
+
+// waitForOpens blocks until the source has been asked for at least n blobs.
+func (s *stuckSource) waitForOpens(t *testing.T, n int, timeout time.Duration) {
+ t.Helper()
+
+ deadline := time.Now().Add(timeout)
+ for time.Now().Before(deadline) {
+ if s.openCount() >= n {
+ return
+ }
+ time.Sleep(5 * time.Millisecond)
+ }
+ t.Fatalf("only %d blob opens after %s, wanted %d", s.openCount(), timeout, n)
+}
+
+// stuckDigest builds a well-formed digest from a label so each synthetic job has
+// its own blobs and the accounting is unambiguous.
+func stuckDigest(label string) string {
+ sum := sha256.Sum256([]byte(label))
+ return fmt.Sprintf("sha256:%x", sum)
+}
+
+// stuckJob builds an ordinary-looking image job: one config, n tar layers.
+// skipReason waves it through, so the pipeline runs for real.
+func stuckJob(label string, layers int) *scanner.ScanJob {
+ job := &scanner.ScanJob{
+ ManifestDigest: stuckDigest(label + "/manifest"),
+ Repository: label,
+ Tag: "latest",
+ UserDID: "did:plc:example",
+ UserHandle: "user.example.com",
+ HoldDID: "did:web:hold.example.com",
+ Tier: "deckhand",
+ Config: scanner.BlobDescriptor{
+ Digest: stuckDigest(label + "/config"),
+ Size: 64,
+ MediaType: "application/vnd.oci.image.config.v1+json",
+ },
+ }
+ for i := 0; i < layers; i++ {
+ job.Layers = append(job.Layers, scanner.BlobDescriptor{
+ Digest: stuckDigest(fmt.Sprintf("%s/layer/%d", label, i)),
+ Size: 1024,
+ MediaType: "application/vnd.oci.image.layer.v1.tar+gzip",
+ })
+ }
+ return job
+}
+
+// stuckWaitForDials blocks until the mock hold has accepted at least n connections.
+func stuckWaitForDials(t *testing.T, h *Harness, n int, timeout time.Duration) {
+ t.Helper()
+
+ deadline := time.Now().Add(timeout)
+ for time.Now().Before(deadline) {
+ if len(h.Hold.Dials()) >= n {
+ return
+ }
+ time.Sleep(20 * time.Millisecond)
+ }
+ t.Fatalf("only %d dials after %s, wanted %d", len(h.Hold.Dials()), timeout, n)
+}
+
+// stuckTerminalFor reports the first result/error/skipped message for a seq, if any.
+func stuckTerminalFor(h *Harness, seq int64) (mockhold.Message, bool) {
+ for _, m := range h.Hold.Transcript() {
+ if m.Seq == seq && (m.Type == "result" || m.Type == "error" || m.Type == "skipped") {
+ return m, true
+ }
+ }
+ return mockhold.Message{}, false
+}
+
+// TestStuckJobBlocksEveryJobBehindIt is the first and simplest way scanning
+// stops: one job that never finishes, and a worker pool with nothing to
+// interrupt it.
+//
+// processJob takes a context and honours it nowhere. buildOCILayout does not
+// take one at all, and the worker's context is the process's, cancelled only at
+// shutdown — there is no per-job deadline anywhere in the scanner. So a job
+// that hangs holds the only worker (scanner.workers defaults to 1) for as long
+// as it hangs, and every job behind it sits in the priority queue having
+// already been acked.
+//
+// Here the hang is a blob the hold never answers. That one is bounded, at
+// 5 minutes per HTTP request by client.httpClient — see
+// TestBlobDownloadsAreSerialSoTheirTimeoutsAdd for how far that bound stretches.
+// The unbounded version is the same shape with no timeout at all: Syft's
+// stereoscope extraction and Grype's matching run outside any deadline, and a
+// layer that decompresses forever holds the worker forever.
+func TestStuckJobBlocksEveryJobBehindIt(t *testing.T) {
+ src := newStuckSource().gated()
+ h := Start(t, src)
+ t.Cleanup(src.Release)
+
+ stuck, err := h.Hold.SendJob(stuckJob("stuck", 1))
+ if err != nil {
+ t.Fatalf("send stuck job: %v", err)
+ }
+ src.waitForOpens(t, 1, 10*time.Second)
+
+ behind, err := h.Hold.SendJob(stuckJob("behind", 1))
+ if err != nil {
+ t.Fatalf("send second job: %v", err)
+ }
+
+ // The hold is told the second job is under way immediately: the scanner
+ // acks on receipt, before it even reaches the queue.
+ if _, err := h.Hold.WaitForMessage(func(m mockhold.Message) bool {
+ return m.Seq == behind && m.Type == "ack"
+ }, 5*time.Second); err != nil {
+ t.Fatalf("second job was never acked: %v", err)
+ }
+
+ // And then nothing happens to it, because the only worker is wedged.
+ if _, err := h.Hold.WaitForMessage(func(m mockhold.Message) bool {
+ return m.Seq == behind && m.Type != "ack"
+ }, 2*time.Second); err == nil {
+ t.Fatal("the queued job finished while the first job was stuck; " +
+ "head-of-line blocking is gone, update this test")
+ }
+
+ if n := src.openCount(); n != 1 {
+ t.Errorf("blob opens = %d, want 1: only the stuck job should have "+
+ "started downloading", n)
+ }
+ if _, ok := stuckTerminalFor(h, stuck); ok {
+ t.Error("the stuck job reported a terminal message while its download " +
+ "was still hanging")
+ }
+ t.Logf("seq %d acked and queued, seq %d holding the only worker, no "+
+ "deadline on either", behind, stuck)
+
+ // Let both unwind so the mock's teardown does not block on the open request.
+ src.Release()
+ h.AwaitTerminal(t, stuck, 30*time.Second)
+ h.AwaitTerminal(t, behind, 30*time.Second)
+}
+
+// TestAcksLandLongBeforeTheWorkDoes quantifies the ack timing gap from the
+// scanner's side, and pins the message that closes it.
+//
+// The scanner acks in handleFrame the moment a job is decoded, before
+// queue.Enqueue. That has not changed and should not: the ack means "I have
+// it". What it never meant is "a worker is on it", and the hold used to have
+// no other signal — handleAck left assigned_at at the dispatch time and the
+// ten-minute processing deadline was measured from there, so the deadline
+// covered queueing. With scanner.workers=1 and a JobCooldown of 10 seconds,
+// job N cannot start earlier than 10*(N-1) seconds after the burst is acked
+// even if every scan were instant, so job 61 in a burst was past the hold's
+// deadline before a worker touched it. The queue is 100 deep.
+//
+// A worker now sends 'started' when it dequeues, and the hold measures the
+// scanning deadline from that. This test sends a burst and shows the shape
+// that makes the two messages different signals: every ack lands up front,
+// while the starts are spread across the whole drain.
+func TestAcksLandLongBeforeTheWorkDoes(t *testing.T) {
+ const burst = 6
+
+ src := newStuckSource().slow(300 * time.Millisecond)
+ h := Start(t, src)
+
+ var seqs []int64
+ for i := 0; i < burst; i++ {
+ seq, err := h.Hold.SendJob(stuckJob(fmt.Sprintf("burst-%d", i), 1))
+ if err != nil {
+ t.Fatalf("send job %d: %v", i, err)
+ }
+ seqs = append(seqs, seq)
+ }
+
+ // Wait for every ack.
+ for _, seq := range seqs {
+ if _, err := h.Hold.WaitForMessage(func(m mockhold.Message) bool {
+ return m.Seq == seq && m.Type == "ack"
+ }, 10*time.Second); err != nil {
+ t.Fatalf("seq %d was never acked: %v", seq, err)
+ }
+ }
+
+ // Then for every job to actually finish.
+ for _, seq := range seqs {
+ h.AwaitTerminal(t, seq, 60*time.Second)
+ }
+
+ var firstAck, lastAck, firstStart, lastStart, firstTerminal, lastTerminal time.Time
+ starts := 0
+ for _, m := range h.Hold.Transcript() {
+ switch m.Type {
+ case "ack":
+ if firstAck.IsZero() {
+ firstAck = m.At
+ }
+ lastAck = m.At
+ case "started":
+ starts++
+ if firstStart.IsZero() {
+ firstStart = m.At
+ }
+ lastStart = m.At
+ default:
+ if firstTerminal.IsZero() {
+ firstTerminal = m.At
+ }
+ lastTerminal = m.At
+ }
+ }
+
+ if starts != burst {
+ t.Fatalf("%d 'started' messages for %d jobs: without one per job the "+
+ "hold is back to measuring its scanning deadline from dispatch",
+ starts, burst)
+ }
+ if !lastAck.Before(firstTerminal) {
+ t.Fatalf("the last ack (%s) did not precede the first terminal message (%s)",
+ lastAck, firstTerminal)
+ }
+
+ ackSpread := lastAck.Sub(firstAck)
+ startSpread := lastStart.Sub(firstStart)
+ drain := lastTerminal.Sub(lastAck)
+ perJob := drain / burst
+
+ // The point of the pair: the acks are a burst, the starts are the drain.
+ // Any deadline measured from the ack is measuring the queue.
+ if startSpread <= ackSpread {
+ t.Errorf("starts spread over %s and acks over %s; the two signals are "+
+ "not distinguishable in this run, so the test proves nothing",
+ startSpread.Round(time.Millisecond), ackSpread.Round(time.Millisecond))
+ }
+ if !lastStart.After(lastAck) {
+ t.Errorf("the last start (%s) did not follow the last ack (%s)",
+ lastStart, lastAck)
+ }
+
+ t.Logf("%d jobs acked within %s of each other, started over %s, and took "+
+ "%s to drain (%s per job at a %s cooldown)",
+ burst, ackSpread.Round(time.Millisecond), startSpread.Round(time.Millisecond),
+ drain.Round(time.Millisecond), perJob.Round(time.Millisecond), testJobCooldown)
+ t.Logf("measured from dispatch, at the production cooldown of 10s this "+
+ "same burst would take %s to drain", (perJob+10*time.Second)*burst)
+
+ // The arithmetic that used to make this a bug, kept as an assertion so a
+ // change to either constant shows up here. It is now the bound on the
+ // fallback budget a hold applies to a scanner that sends no 'started', not
+ // on healthy work.
+ const holdScanningDeadline = 10 * time.Minute
+ const productionCooldown = 10 * time.Second
+ if ceiling := int(holdScanningDeadline / productionCooldown); ceiling != 60 {
+ t.Errorf("a burst of more than %d jobs cannot be drained inside the "+
+ "hold's scanning deadline even with free scans; that number moved",
+ ceiling)
+ }
+}
+
+// TestBlobDownloadsAreSerialSoTheirTimeoutsAdd shows why the one timeout the
+// download path does have is not a bound on a job.
+//
+// buildOCILayout fetches the config and then every layer in sequence, each
+// through client.httpClient, whose Timeout is 5 minutes and applies per
+// request. A manifest with 19 layers therefore has a worst case of 20 × 5
+// minutes before the job fails — twice the hold's ten-minute processing
+// deadline, so the hold gives up on a job the scanner is still legitimately
+// working on, and does so without re-dispatching it.
+func TestBlobDownloadsAreSerialSoTheirTimeoutsAdd(t *testing.T) {
+ job := stuckJob("serial", 3)
+
+ // Config and the first two layers download and verify; the third keeps its
+ // label-derived digest, which the source does not have, so it 404s and the
+ // job fails there rather than reaching Syft. The sizes differ so the three
+ // served blobs are three distinct digests.
+ src := newStuckSource().slow(100 * time.Millisecond)
+ job.Config.Digest, job.Config.Size = src.servingJunk(64), 64
+ job.Layers[0].Digest, job.Layers[0].Size = src.servingJunk(65), 65
+ job.Layers[1].Digest, job.Layers[1].Size = src.servingJunk(66), 66
+
+ h := Start(t, src)
+
+ start := time.Now()
+ seq, err := h.Hold.SendJob(job)
+ if err != nil {
+ t.Fatalf("send job: %v", err)
+ }
+ msg := h.AwaitTerminal(t, seq, 60*time.Second)
+ elapsed := time.Since(start)
+
+ if msg.Type != "error" {
+ t.Fatalf("want error, got %s (%s%s)", msg.Type, msg.Error, msg.Reason)
+ }
+ if n := src.openCount(); n != 4 {
+ t.Errorf("blob opens = %d, want 4 (config + 3 layers)", n)
+ }
+ if peak := src.peakConcurrency(); peak != 1 {
+ t.Errorf("peak concurrent downloads = %d, want 1: downloads within a "+
+ "job are serial", peak)
+ }
+ if elapsed < 4*100*time.Millisecond {
+ t.Errorf("job took %s, less than the sum of its four downloads", elapsed)
+ }
+
+ t.Logf("4 serial downloads took %s; each is bounded only by the 5-minute "+
+ "client timeout, so this job's worst case is 20 minutes against a "+
+ "10-minute hold deadline", elapsed.Round(time.Millisecond))
+}
+
+// TestReDispatchAfterDisconnectScansTheSameImageTwice proves the duplicate-scan
+// prediction.
+//
+// The hold's Unsubscribe flips the dropped scanner's assigned and processing
+// rows back to 'pending', and drainPendingJobs hands them to the next
+// connection (pinned in
+// pkg/hold/pds/scan_broadcaster_stuck_test.go:TestScanUnsubscribe_ReoffersAJobTheScannerIsStillRunning).
+// Nothing tells the scanner's worker pool the socket went away: the job it was
+// running is still running, and the re-offered copy is enqueued again with no
+// dedupe by seq or digest.
+//
+// With two workers that is two concurrent downloads of the same blob, which the
+// per-digest request count and the source's peak concurrency both show.
+func TestReDispatchAfterDisconnectScansTheSameImageTwice(t *testing.T) {
+ src := newStuckSource().gated()
+ h := Start(t, src, WithWorkers(2))
+ t.Cleanup(src.Release)
+
+ job := stuckJob("duplicated", 1)
+ seq, err := h.Hold.SendJob(job)
+ if err != nil {
+ t.Fatalf("send job: %v", err)
+ }
+ src.waitForOpens(t, 1, 10*time.Second)
+
+ // The scanner is mid-download when the socket dies.
+ h.Hold.DropConnections(mockhold.DropAbrupt)
+ stuckWaitForDials(t, h, 2, 15*time.Second)
+
+ // What drainPendingJobs does on the new connection: the same seq again.
+ again := stuckJob("duplicated", 1)
+ again.Seq = seq
+ if _, err := h.Hold.SendJob(again); err != nil {
+ t.Fatalf("re-send job: %v", err)
+ }
+
+ src.waitForOpens(t, 2, 15*time.Second)
+
+ if n := src.countFor(job.Config.Digest); n < 2 {
+ t.Errorf("config blob fetched %d times, want 2: the re-offered job was "+
+ "deduplicated somewhere", n)
+ }
+ if peak := src.peakConcurrency(); peak < 2 {
+ t.Errorf("peak concurrent downloads = %d, want 2: the duplicate ran "+
+ "after the original rather than alongside it", peak)
+ }
+ t.Logf("seq %d is being downloaded by two workers at once; neither knows "+
+ "about the other and both will report a result", seq)
+
+ src.Release()
+ h.AwaitTerminal(t, seq, 30*time.Second)
+}
+
+// TestResultComputedWhileTheSocketIsDownIsLost covers the last way a job goes
+// quiet: the scanner finishes, and its answer goes nowhere.
+//
+// client.sendJSON guards only on `c.conn == nil`, and nothing ever nils conn —
+// connectOnce sets it on dial and the read loop just returns on error. So a
+// terminal message computed between a disconnect and the next dial is written
+// to a closed connection, WriteJSON fails, sendJSON logs it and returns. Ack,
+// SendResult, SendError and SendSkipped all return no error, so the worker has
+// no way to know and no path to retry: it moves straight on to the next job.
+//
+// The hold, meanwhile, put the row back to 'pending' when the scanner dropped,
+// so the work is simply done twice — and the second copy is the one that counts.
+// The loss window is the reconnect backoff, a flat 5 seconds (Connect's comment
+// claims exponential backoff to 30s; the code sleeps 5s and the cursor variable
+// it declares is never assigned).
+func TestResultComputedWhileTheSocketIsDownIsLost(t *testing.T) {
+ src := newStuckSource().gated()
+ h := Start(t, src)
+ t.Cleanup(src.Release)
+
+ seq, err := h.Hold.SendJob(stuckJob("lost", 1))
+ if err != nil {
+ t.Fatalf("send job: %v", err)
+ }
+ src.waitForOpens(t, 1, 10*time.Second)
+
+ if _, err := h.Hold.WaitForMessage(func(m mockhold.Message) bool {
+ return m.Seq == seq && m.Type == "ack"
+ }, 5*time.Second); err != nil {
+ t.Fatalf("job was never acked: %v", err)
+ }
+
+ // Kill the socket, then let the job finish. The worker computes its answer
+ // with nowhere to send it.
+ h.Hold.DropConnections(mockhold.DropAbrupt)
+ src.Release()
+
+ // Give the worker time to finish and write into the dead conn, then wait
+ // for the reconnect so we can show nothing is re-sent afterwards either.
+ stuckWaitForDials(t, h, 2, 15*time.Second)
+ time.Sleep(500 * time.Millisecond)
+
+ if msg, ok := stuckTerminalFor(h, seq); ok {
+ t.Fatalf("the hold received a %s for seq %d after all; sendJSON now "+
+ "survives a dropped socket, update this test", msg.Type, seq)
+ }
+ if n := h.Queue.Len(); n != 0 {
+ t.Errorf("scanner queue holds %d jobs; the lost job was requeued, "+
+ "update this test", n)
+ }
+ t.Logf("seq %d was scanned to completion and its outcome discarded; the "+
+ "hold will only notice via its own ten-minute processing timeout", seq)
+
+ // The connection is healthy again — the loss was silent, not fatal.
+ if _, err := h.Hold.SendJob(stuckJob("after-reconnect", 1)); err != nil {
+ t.Fatalf("hold could not dispatch after the reconnect: %v", err)
+ }
+}
diff --git a/scanner/internal/e2e/vulnreport_test.go b/scanner/internal/e2e/vulnreport_test.go
new file mode 100644
index 0000000..8da7d13
--- /dev/null
+++ b/scanner/internal/e2e/vulnreport_test.go
@@ -0,0 +1,958 @@
+package e2e
+
+// Scan REPORT verification: what Grype actually matches, and whether the JSON
+// the scanner publishes is the JSON the appview reads back.
+//
+// Everything else in this package proves the pipeline mechanics — a job is
+// dispatched, blobs are fetched, an SBOM comes back. None of it exercises
+// Grype: the harness disables it precisely so no test has to download a
+// vulnerability database, which left the entire matching path, the matcher
+// configuration, and every vulnerability count the scanner has ever published
+// unverified (SCANNER_BUGS.md section 4). These scenarios close that.
+//
+// # Running them
+//
+// scanner/internal/mockhold/testdata/fetch-vulndb.sh # ~2 GB, once
+// cd scanner && ATCR_SCANNER_VULNDB=1 go test ./internal/e2e -run TestVulnDB -v
+//
+// Without ATCR_SCANNER_VULNDB=1 every scenario here skips, so an ordinary
+// `go test ./...` stays fast and offline.
+//
+// # Why nothing here asserts a CVE count or a CVE ID
+//
+// The upstream feed is rebuilt daily. A test that pins "alpine:3.10 has 47
+// criticals" is wrong within a week and then gets muted, which is worse than
+// not having written it. Every assertion below is an invariant that holds
+// whatever the feed says: internal consistency of the counts, agreement
+// between the SBOM and the report, the JSON shape the appview parses,
+// determinism against a fixed database, and "more than zero" on an image that
+// has been end-of-life for years.
+
+import (
+ "encoding/json"
+ "fmt"
+ "os"
+ "runtime"
+ "sort"
+ "strconv"
+ "strings"
+ "sync"
+ "testing"
+ "time"
+
+ scanner "atcr.io/scanner"
+ "atcr.io/scanner/internal/mockhold"
+)
+
+// ---------------------------------------------------------------------------
+// Fixtures
+// ---------------------------------------------------------------------------
+
+// vulnFixtureOld is an image that has been end-of-life since 2021. Its apk
+// packages are frozen at versions the feed has accumulated advisories against
+// for years, so "this scan finds vulnerabilities" is a safe assertion in a way
+// it would not be for a current tag. fetch-vulndb.sh pulls it.
+const vulnFixtureOld = "vuln-alpine310"
+
+// vulnFixtureMinimal is distroless static: a scratch-like base with a handful
+// of dpkg status.d entries and no interpreter, no shell, no libc package. It
+// is the control. A matcher that invents results has nothing here to invent
+// them from, so a large finding count against it would mean the matcher is
+// reporting on packages the image does not contain.
+const vulnFixtureMinimal = "vuln-static"
+
+// vulnFixtureDebian is oldstable Debian. It is in the set for one reason: the
+// severity strings differ by ecosystem. Alpine's secdb resolves everything to
+// an NVD severity, so an Alpine image can accidentally satisfy a
+// bucket-completeness check that a Debian image does not — Debian's security
+// tracker carries "Negligible" and "Unknown" ratings, neither of which the
+// scanner buckets. A consistency check run against one distro would report
+// whichever answer that distro happens to give.
+const vulnFixtureDebian = "vuln-debian11"
+
+// severitySweepFixtures is every image on disk worth running the
+// bucket-consistency check against, cheapest first. It deliberately spans
+// package ecosystems (apk, dpkg, go modules) because the shortfall the check
+// looks for is a property of the severity vocabulary each feed uses, not of
+// the scanner's arithmetic.
+func severitySweepFixtures(t *testing.T) []string {
+ t.Helper()
+ var out []string
+ for _, name := range []string{
+ vulnFixtureMinimal,
+ vulnFixtureOld,
+ vulnFixtureDebian,
+ "hsm-secrets-operator", // a Go binary image: go-module matches, no distro packages
+ "perf-python", // debian bookworm plus a python ecosystem
+ } {
+ if hasFixture(name) {
+ out = append(out, name)
+ }
+ }
+ if len(out) == 0 {
+ t.Skipf("no image fixtures present; run scanner/internal/mockhold/testdata/fetch-vulndb.sh")
+ }
+ return out
+}
+
+// ---------------------------------------------------------------------------
+// The report, decoded the way the appview decodes it
+// ---------------------------------------------------------------------------
+
+// appviewReport and friends are a deliberate restatement of the structs the
+// appview parses this blob with, so a change to the report's shape fails here
+// rather than silently rendering an empty vulnerabilities tab in production.
+//
+// The originals are pkg/appview/handlers/vuln_details.go:25-53 (and a verbatim
+// duplicate in cmd/image-advisor/main.go:62-86). They are restated rather than
+// imported because they are unexported, and because a copy that has to be kept
+// in sync is the point: this file is the thing that notices when it drifts.
+//
+// The one intentional difference: Metadata is a pointer here and a value there.
+// The scanner's own counting code branches on Metadata being nil
+// (scanner/internal/scan/grype.go, countVulnerabilitiesBySeverity), so the
+// tests below need to tell "no metadata" apart from "metadata with an empty
+// severity". The appview cannot: a null Metadata decodes to a zero struct and
+// arrives at the template as an empty severity string.
+type appviewReport struct {
+ Matches []appviewMatch `json:"matches"`
+}
+
+type appviewMatch struct {
+ Vulnerability appviewVuln `json:"Vulnerability"`
+ Package appviewPackage `json:"Package"`
+}
+
+type appviewVuln struct {
+ ID string `json:"ID"`
+ Metadata *appviewMetadata `json:"Metadata"`
+ Fix appviewFix `json:"Fix"`
+}
+
+type appviewMetadata struct {
+ Severity string `json:"Severity"`
+}
+
+type appviewFix struct {
+ Versions []string `json:"Versions"`
+ State string `json:"State"`
+}
+
+type appviewPackage struct {
+ Name string `json:"Name"`
+ Version string `json:"Version"`
+ Type string `json:"Type"`
+}
+
+// spdxDoc is the slice of the SPDX JSON SBOM these scenarios need: the package
+// list, so the report can be checked against it.
+type spdxDoc struct {
+ SPDXVersion string `json:"spdxVersion"`
+ Packages []struct {
+ Name string `json:"name"`
+ VersionInfo string `json:"versionInfo"`
+ } `json:"packages"`
+}
+
+// finding is one match reduced to the identity a user would recognise. It is
+// the unit the determinism scenario compares, because the full report JSON
+// carries a scan-local temp path (see TestVulnDBScanIsDeterministic).
+type finding struct {
+ ID string
+ Package string
+ Version string
+}
+
+func (f finding) String() string { return f.ID + " " + f.Package + "@" + f.Version }
+
+// ---------------------------------------------------------------------------
+// Running one scan
+// ---------------------------------------------------------------------------
+
+// vulnScan is everything one real scan produced.
+type vulnScan struct {
+ Fixture string
+ Summary scanner.VulnerabilitySummary
+ Report appviewReport
+ Raw []byte
+ SBOM spdxDoc
+ Digest string
+
+ Elapsed time.Duration
+ PeakRSS uint64 // bytes, sampled from /proc/self/statm during the scan
+ HeapEnd uint64
+}
+
+// Findings returns the report's matches as comparable identities.
+func (v vulnScan) Findings() []finding {
+ out := make([]finding, 0, len(v.Report.Matches))
+ for _, m := range v.Report.Matches {
+ out = append(out, finding{m.Vulnerability.ID, m.Package.Name, m.Package.Version})
+ }
+ sort.Slice(out, func(i, j int) bool { return out[i].String() < out[j].String() })
+ return out
+}
+
+// severityHistogram counts every match by the exact severity string Grype put
+// on it, with a distinct key for a match that carries no metadata at all.
+// Both are cases countVulnerabilitiesBySeverity adds to Total and puts in no
+// bucket, so naming them separately is what makes the shortfall diagnosable.
+func (v vulnScan) severityHistogram() map[string]int {
+ hist := map[string]int{}
+ for _, m := range v.Report.Matches {
+ switch {
+ case m.Vulnerability.Metadata == nil:
+ hist["(no metadata)"]++
+ case m.Vulnerability.Metadata.Severity == "":
+ hist["(empty)"]++
+ default:
+ hist[m.Vulnerability.Metadata.Severity]++
+ }
+ }
+ return hist
+}
+
+// requireVulnFixture skips unless the pinned image has been pulled.
+func requireVulnFixture(t *testing.T, name string) {
+ t.Helper()
+ if !hasFixture(name) {
+ t.Skipf("fixture %q not present; run scanner/internal/mockhold/testdata/fetch-vulndb.sh", name)
+ }
+}
+
+// runVulnScan drives one image through the real pipeline with Grype enabled
+// and decodes both artifacts it produced.
+//
+// The timeout is generous because the first scan in a binary also pays the
+// database load: opening a ~2 GB SQLite file and building its indexes.
+// Each scan runs inside its own subtest, which is what bounds the harness's
+// lifetime. Start writes the package-level scan.JobCooldown and every worker
+// reads it once per loop iteration, so two harnesses alive at the same moment
+// is a data race -race reports against worker.go — and several scenarios here
+// scan a list of fixtures in one test. A subtest's t.Cleanup runs when t.Run
+// returns (and Start's cleanup joins the pool before restoring the global), so
+// harness N is fully torn down before harness N+1 starts.
+func runVulnScan(t *testing.T, fixture string) vulnScan {
+ t.Helper()
+ requireVulnFixture(t, fixture)
+
+ var out vulnScan
+ ok := t.Run(fixture, func(t *testing.T) {
+ dir := fixtureDir(fixture)
+ h := Start(t, mockhold.NewOCILayout(dir), WithVulnDB(t))
+ out = sendVulnScan(t, h, jobFromLayout(t, dir), fixture)
+ })
+ if !ok {
+ t.Fatalf("scan of %s did not complete", fixture)
+ }
+ if out.Fixture == "" {
+ t.Skipf("scan of %s was skipped", fixture)
+ }
+ return out
+}
+
+// sendVulnScan sends one job to an already-running harness and decodes the
+// result. Split out from runVulnScan so a scenario can put two scans through
+// one harness (and therefore one loaded provider).
+func sendVulnScan(t *testing.T, h *Harness, job *scanner.ScanJob, fixture string) vulnScan {
+ t.Helper()
+
+ job.Seq = 0 // let the mock allocate a fresh seq per send
+ stopRSS := watchRSS()
+ start := time.Now()
+
+ seq, err := h.Hold.SendJob(job)
+ if err != nil {
+ t.Fatalf("send job: %v", err)
+ }
+ msg := h.AwaitTerminal(t, seq, 15*time.Minute)
+ elapsed := time.Since(start)
+ peakRSS := stopRSS()
+
+ if msg.Type != "result" {
+ t.Fatalf("%s: want result, got %s: %s%s", fixture, msg.Type, msg.Error, msg.Reason)
+ }
+ if msg.Summary == nil {
+ t.Fatalf("%s: vuln scanning is enabled but the result carried no summary", fixture)
+ }
+ if msg.VulnReport == "" {
+ t.Fatalf("%s: result carried no vulnerability report", fixture)
+ }
+ if msg.SBOM == "" {
+ t.Fatalf("%s: result carried no SBOM", fixture)
+ }
+
+ out := vulnScan{
+ Fixture: fixture,
+ Summary: *msg.Summary,
+ Raw: []byte(msg.VulnReport),
+ Elapsed: elapsed,
+ PeakRSS: peakRSS,
+ }
+ if err := json.Unmarshal(out.Raw, &out.Report); err != nil {
+ t.Fatalf("%s: the appview's own structs cannot parse the vulnerability report: %v", fixture, err)
+ }
+ if err := json.Unmarshal([]byte(msg.SBOM), &out.SBOM); err != nil {
+ t.Fatalf("%s: parse SBOM: %v", fixture, err)
+ }
+
+ var digest struct {
+ Descriptor struct {
+ Name string `json:"name"`
+ Version string `json:"version"`
+ } `json:"descriptor"`
+ }
+ _ = json.Unmarshal(out.Raw, &digest)
+ out.Digest = digest.Descriptor.Version
+
+ var ms runtime.MemStats
+ runtime.ReadMemStats(&ms)
+ out.HeapEnd = ms.HeapAlloc
+
+ return out
+}
+
+// watchRSS samples this process's resident set every 50ms until the returned
+// stop func is called, and reports the peak.
+//
+// The scanner's production numbers all exclude Grype: FindMatches was never
+// run, so nothing measured what matching and the mmap'd database cost. This is
+// crude (one process, no cgroup) but it is the first measurement of that cost
+// there has been, and the database is mmap-backed, so RSS is the number that
+// matters rather than heap.
+func watchRSS() (stop func() uint64) {
+ var mu sync.Mutex
+ var high uint64
+ done := make(chan struct{})
+ exited := make(chan struct{})
+ var once sync.Once
+
+ sample := func() {
+ rss := currentRSS()
+ mu.Lock()
+ if rss > high {
+ high = rss
+ }
+ mu.Unlock()
+ }
+
+ sample()
+ go func() {
+ defer close(exited)
+ tick := time.NewTicker(50 * time.Millisecond)
+ defer tick.Stop()
+ for {
+ select {
+ case <-done:
+ return
+ case <-tick.C:
+ sample()
+ }
+ }
+ }()
+
+ // stop joins the sampler before reading the peak. Returning a pointer and
+ // letting the caller dereference it races with the last sample, which
+ // -race reports.
+ return func() uint64 {
+ once.Do(func() { close(done) })
+ <-exited
+ sample()
+ mu.Lock()
+ defer mu.Unlock()
+ return high
+ }
+}
+
+// currentRSS reads this process's resident set size. Resident rather than heap
+// because the vulnerability database is a mapped file: the Go runtime's own
+// accounting, and the GOMEMLIMIT built on it, cannot see it.
+func currentRSS() uint64 {
+ data, err := os.ReadFile("/proc/self/statm")
+ if err != nil {
+ return 0
+ }
+ fields := strings.Fields(string(data))
+ if len(fields) < 2 {
+ return 0
+ }
+ pages, err := strconv.ParseUint(fields[1], 10, 64)
+ if err != nil {
+ return 0
+ }
+ return pages * uint64(os.Getpagesize())
+}
+
+func mib(n uint64) string { return fmt.Sprintf("%.0f MiB", float64(n)/(1<<20)) }
+
+// ---------------------------------------------------------------------------
+// One-time priming
+// ---------------------------------------------------------------------------
+
+// The vulnerability provider is a package-level global in internal/scan behind
+// an RWMutex, so one load serves every scenario in this binary. primeVulnDB
+// makes that ordering explicit rather than leaving it to whichever test the
+// -run filter happened to select first: the first scan pays the load, and the
+// rest measure matching alone.
+//
+// This does not reach the stubbed-loader tests in
+// internal/scan/vulndb_refresh_test.go. Those are a different package and
+// therefore a different test binary, with their own copy of these globals and
+// their own stubbed loadVulnDB, so they stay offline whatever this file does.
+var (
+ primeOnce sync.Once
+ primed vulnScan
+ primeFail string
+)
+
+// primeVulnDB loads the database once and returns the control scan it used to
+// do it. The minimal fixture is the cheapest image that still forces a real
+// match run, so priming costs a scan that a scenario wanted anyway.
+func primeVulnDB(t *testing.T) vulnScan {
+ t.Helper()
+ requireVulnFixture(t, vulnFixtureMinimal)
+
+ primeOnce.Do(func() {
+ defer func() {
+ if r := recover(); r != nil {
+ primeFail = fmt.Sprintf("panic while priming the vulnerability database: %v", r)
+ }
+ }()
+ before := currentRSS()
+ // runVulnScan already bounds its harness in a subtest, so the priming
+ // scan's worker pool is joined before the caller's own harness starts.
+ primed = runVulnScan(t, vulnFixtureMinimal)
+ t.Logf("primed vulnDB: first scan (%s) took %s; RSS %s before, %s peak, %s after — "+
+ "the database is a ~2 GB mmap'd SQLite file, so only the pages the matcher touches "+
+ "become resident",
+ vulnFixtureMinimal, primed.Elapsed.Round(time.Millisecond),
+ mib(before), mib(primed.PeakRSS), mib(currentRSS()))
+ })
+ if primeFail != "" {
+ t.Fatal(primeFail)
+ }
+ if primed.Fixture == "" {
+ // The priming scan skipped or failed inside the Once, which unwinds
+ // its goroutine rather than returning, so there is nothing to report
+ // here beyond "it did not happen". The scenario that ran it carries
+ // the reason.
+ t.Skip("vulnerability database was never primed; see the first scenario's output")
+ }
+ return primed
+}
+
+// ---------------------------------------------------------------------------
+// Scenarios
+// ---------------------------------------------------------------------------
+
+// TestVulnDBKnownVulnerableImageFindsSomething is the base case the whole file
+// rests on: an image that has been out of support for years, scanned against
+// the real feed, must produce findings. If this comes back empty the matcher
+// is not working, whatever the rest of the assertions say.
+//
+// No count is asserted. "More than zero on an EOL distro" is true of every
+// version of the feed that has ever existed; "47 criticals" is true of one
+// day's build.
+func TestVulnDBKnownVulnerableImageFindsSomething(t *testing.T) {
+ primeVulnDB(t)
+ scan := runVulnScan(t, vulnFixtureOld)
+
+ t.Logf("%s: %d packages, %d matches, summary C=%d H=%d M=%d L=%d total=%d in %s (peak RSS %s)",
+ scan.Fixture, len(scan.SBOM.Packages), len(scan.Report.Matches),
+ scan.Summary.Critical, scan.Summary.High, scan.Summary.Medium, scan.Summary.Low,
+ scan.Summary.Total, scan.Elapsed.Round(time.Millisecond), mib(scan.PeakRSS))
+
+ if scan.Summary.Total == 0 {
+ t.Fatal("an end-of-life image scanned clean; the matcher is not finding anything")
+ }
+ if len(scan.Report.Matches) == 0 {
+ t.Fatal("summary claims findings but the report carries no matches")
+ }
+
+ // The summary must describe the report it ships with. This is not a feed
+ // assertion: whatever the feed says, Total is counted from exactly the
+ // match set that gets marshalled.
+ if scan.Summary.Total != len(scan.Report.Matches) {
+ t.Errorf("summary.Total is %d but the report carries %d matches",
+ scan.Summary.Total, len(scan.Report.Matches))
+ }
+
+ // Severity is plausible rather than exact: an EOL base image always has
+ // something above Low. Asserting the pair rather than either alone keeps
+ // this true across the feed's periodic severity re-scoring.
+ if scan.Summary.Critical+scan.Summary.High == 0 {
+ t.Errorf("no critical or high findings on an end-of-life image; histogram: %v",
+ scan.severityHistogram())
+ }
+
+ // Every finding must name a vulnerability. An empty ID renders as a blank
+ // cell in the appview's table (partials/vuln-details.html), which is the
+ // shape of a report that scanned but shows nothing.
+ for i, m := range scan.Report.Matches {
+ if m.Vulnerability.ID == "" {
+ t.Errorf("match %d has no vulnerability ID", i)
+ }
+ if m.Package.Name == "" {
+ t.Errorf("match %d (%s) names no package", i, m.Vulnerability.ID)
+ }
+ }
+}
+
+// TestVulnDBSeverityBucketsAccountForEveryMatch is the internal-consistency
+// check, and it is written to pass against today's behaviour while recording
+// precisely how far off it is.
+//
+// countVulnerabilitiesBySeverity (scanner/internal/scan/grype.go) increments
+// Total for every match but buckets only matches whose Metadata is non-nil and
+// whose Severity is exactly one of Critical/High/Medium/Low. Grype also emits
+// "Negligible" and "Unknown", and neither has a bucket anywhere in the system:
+// not in the wire summary (pkg/hold/pds/scan_broadcaster.go), not in the
+// io.atcr.hold.scan record (pkg/atproto/lexicon.go), not in the CSS
+// (pkg/appview/src/css/main.css has exactly four .vuln-box-* classes).
+//
+// So the assertion is one-directional — the buckets may never exceed the total,
+// and no match may be counted twice — plus a log of the shortfall with the
+// severity strings that caused it. See the report accompanying this file.
+func TestVulnDBSeverityBucketsAccountForEveryMatch(t *testing.T) {
+ primeVulnDB(t)
+
+ unbucketedTotal := 0
+ for _, fixture := range severitySweepFixtures(t) {
+ scan := runVulnScan(t, fixture)
+
+ buckets := scan.Summary.Critical + scan.Summary.High + scan.Summary.Medium + scan.Summary.Low
+ hist := scan.severityHistogram()
+
+ t.Logf("%-22s buckets=%-5d total=%-5d unaccounted=%-4d | %s",
+ fixture, buckets, scan.Summary.Total, scan.Summary.Total-buckets, formatHistogram(hist))
+
+ if buckets > scan.Summary.Total {
+ t.Errorf("%s: severity buckets sum to %d, more than the total of %d — a match is counted twice",
+ fixture, buckets, scan.Summary.Total)
+ }
+
+ // Cross-check the counter against the report it summarises, bucket by
+ // bucket. This is what proves any shortfall is unbucketed severities
+ // rather than a miscount.
+ for _, sev := range []struct {
+ name string
+ got int
+ }{
+ {"Critical", scan.Summary.Critical},
+ {"High", scan.Summary.High},
+ {"Medium", scan.Summary.Medium},
+ {"Low", scan.Summary.Low},
+ } {
+ if hist[sev.name] != sev.got {
+ t.Errorf("%s: summary says %d %s but the report carries %d",
+ fixture, sev.got, sev.name, hist[sev.name])
+ }
+ }
+
+ // Name what fell through, so the finding is actionable rather than a
+ // number. These are the matches the appview counts in "N
+ // vulnerabilities" (partials/vuln-details.html:36) and renders as a
+ // grey "?" row (:93), while the four coloured boxes beside them sum
+ // lower.
+ var unbucketed []string
+ for sev, n := range hist {
+ switch sev {
+ case "Critical", "High", "Medium", "Low":
+ continue
+ }
+ unbucketed = append(unbucketed, fmt.Sprintf("%s=%d", sev, n))
+ }
+ if len(unbucketed) > 0 {
+ sort.Strings(unbucketed)
+ unbucketedTotal += scan.Summary.Total - buckets
+ t.Logf(" FINDING %s: %d of %d matches are in no severity bucket: %s",
+ fixture, scan.Summary.Total-buckets, scan.Summary.Total, strings.Join(unbucketed, " "))
+ }
+ }
+
+ if unbucketedTotal == 0 {
+ t.Log("every match on every fixture landed in one of the four buckets. That is a property " +
+ "of these images and this feed build, not of the code: countVulnerabilitiesBySeverity " +
+ "has no default arm, so any Negligible, Unknown, empty or metadata-less severity would " +
+ "still be counted in Total and bucketed nowhere.")
+ return
+ }
+ t.Logf("%d matches across the sweep are counted in Total and shown in no severity box", unbucketedTotal)
+}
+
+// TestVulnDBMinimalImageIsNearlyClean is the control. A distroless static base
+// has a couple of dpkg entries and nothing else, so a matcher behaving itself
+// has almost nothing to report. The assertion is proportional rather than
+// absolute: findings must not outnumber what the SBOM says is installed by
+// more than a small factor, and every one must name a package the SBOM lists.
+//
+// That second half is the real control. A count ceiling can drift with the
+// feed; "the report may only mention packages the image contains" cannot.
+func TestVulnDBMinimalImageIsNearlyClean(t *testing.T) {
+ scan := primeVulnDB(t)
+
+ t.Logf("%s: %d packages in the SBOM, %d matches, total=%d",
+ scan.Fixture, len(scan.SBOM.Packages), len(scan.Report.Matches), scan.Summary.Total)
+ if len(scan.SBOM.Packages) == 0 {
+ t.Fatal("the control image catalogued no packages at all; it is not a control, it is an empty scan")
+ }
+
+ sbomNames := map[string]bool{}
+ for _, p := range scan.SBOM.Packages {
+ sbomNames[p.Name] = true
+ }
+ for _, m := range scan.Report.Matches {
+ if !sbomNames[m.Package.Name] {
+ t.Errorf("report names package %q, which is not in the SBOM: the matcher is reporting on "+
+ "something the image does not contain", m.Package.Name)
+ }
+ }
+
+ // One package can legitimately carry several advisories, so this is a
+ // sanity ceiling, not an expected value.
+ if max := 20 * len(scan.SBOM.Packages); scan.Summary.Total > max {
+ t.Errorf("%d findings against %d packages (ceiling %d) — implausible for a distroless base",
+ scan.Summary.Total, len(scan.SBOM.Packages), max)
+ }
+}
+
+// TestVulnDBSBOMAndReportAgree pins the join between the two artifacts a scan
+// publishes. They are produced by different libraries from one catalog and
+// stored as two separate blobs, and nothing downstream reconciles them: the
+// appview renders the SBOM tab from one and the vulnerabilities tab from the
+// other. A report naming a package the SBOM does not list would show a user a
+// vulnerability in something their image demonstrably does not have.
+//
+// Feed-independent: it says nothing about which vulnerabilities were found,
+// only that whatever was found is attributed to a real package at the version
+// the SBOM recorded.
+func TestVulnDBSBOMAndReportAgree(t *testing.T) {
+ primeVulnDB(t)
+ scan := runVulnScan(t, vulnFixtureOld)
+
+ type pkgKey struct{ name, version string }
+ sbomPkgs := map[pkgKey]bool{}
+ sbomNames := map[string]bool{}
+ for _, p := range scan.SBOM.Packages {
+ sbomPkgs[pkgKey{p.Name, p.VersionInfo}] = true
+ sbomNames[p.Name] = true
+ }
+
+ unknownName := map[string]bool{}
+ versionMismatch := map[string]bool{}
+ for _, m := range scan.Report.Matches {
+ if !sbomNames[m.Package.Name] {
+ unknownName[m.Package.Name] = true
+ continue
+ }
+ if !sbomPkgs[pkgKey{m.Package.Name, m.Package.Version}] {
+ versionMismatch[m.Package.Name+"@"+m.Package.Version] = true
+ }
+ }
+
+ if len(unknownName) > 0 {
+ t.Errorf("the vulnerability report names %d packages absent from the SBOM: %s",
+ len(unknownName), strings.Join(sortedKeys(unknownName), ", "))
+ }
+ // A version disagreement is softer: Grype normalises some ecosystems'
+ // versions (epoch handling, for one) and the SPDX encoder writes the raw
+ // syft value. Report it rather than failing, so the first real instance is
+ // visible without this rotting into a false alarm.
+ if len(versionMismatch) > 0 {
+ t.Logf("note: %d matched packages carry a version the SBOM records differently: %s",
+ len(versionMismatch), strings.Join(sortedKeys(versionMismatch), ", "))
+ }
+ t.Logf("%d matches across %d SBOM packages, all attributable", len(scan.Report.Matches), len(scan.SBOM.Packages))
+}
+
+// TestVulnDBReportIsTheShapeTheAppviewReads asserts the encoder's output
+// against the parser that consumes it.
+//
+// This is the failure mode nobody would notice. A report that scans correctly
+// but decodes to nothing renders as an empty vulnerabilities tab with a
+// healthy-looking summary beside it, because the summary comes from the
+// io.atcr.hold.scan record and never from the blob. The appview would show a
+// user "31 vulnerabilities" above an empty table and log nothing.
+//
+// The shape is unusual enough to be worth pinning explicitly: the top-level
+// key is lowercase ("matches"), because grype.go marshals it into a
+// map[string]any literal, while everything nested is PascalCase, because
+// match.Match has no JSON tags at all and Go falls back to field names. Adding
+// tags upstream, or switching to Grype's own presenter model (which is
+// lowercase throughout), would silently break every stored report.
+func TestVulnDBReportIsTheShapeTheAppviewReads(t *testing.T) {
+ primeVulnDB(t)
+ scan := runVulnScan(t, vulnFixtureOld)
+
+ // Parsed via the appview's structs already, in sendVulnScan. That it
+ // parsed is necessary but not sufficient: encoding/json is happy to
+ // produce an empty slice from a document with the wrong key.
+ if len(scan.Report.Matches) == 0 {
+ t.Fatal(`the appview's structs parsed the report into zero matches; the top-level "matches" key is gone`)
+ }
+
+ // Walk the raw JSON for the exact keys the appview's field tags name, so a
+ // rename is caught at the key rather than at the value.
+ var raw struct {
+ Matches []map[string]json.RawMessage `json:"matches"`
+ }
+ if err := json.Unmarshal(scan.Raw, &raw); err != nil {
+ t.Fatalf("parse report: %v", err)
+ }
+ if len(raw.Matches) != len(scan.Report.Matches) {
+ t.Fatalf("raw report has %d matches, typed parse produced %d", len(raw.Matches), len(scan.Report.Matches))
+ }
+ for _, key := range []string{"Vulnerability", "Package"} {
+ if _, ok := raw.Matches[0][key]; !ok {
+ t.Errorf("match objects have no %q key; appview handlers/vuln_details.go cannot read this report", key)
+ }
+ }
+
+ // Fields the templates actually render. Each is named with the template
+ // that would go blank without it.
+ var withMetadata, withFix, withType, cveLinkable int
+ for _, m := range scan.Report.Matches {
+ if m.Vulnerability.Metadata != nil && m.Vulnerability.Metadata.Severity != "" {
+ withMetadata++ // vuln-details.html:84-94, the severity badge
+ }
+ if len(m.Vulnerability.Fix.Versions) > 0 {
+ withFix++ // vuln-details.html:101-105, the "Fix" column
+ }
+ if m.Package.Type != "" {
+ withType++ // vuln-details.html:97, the "(type)" suffix
+ }
+ if strings.HasPrefix(m.Vulnerability.ID, "CVE-") || strings.HasPrefix(m.Vulnerability.ID, "GHSA-") {
+ cveLinkable++ // vuln_details.go:242-246, the CVE URL synthesis
+ }
+ }
+ t.Logf("of %d matches: %d carry a severity, %d a fixed-in version, %d a package type, %d a linkable ID",
+ len(scan.Report.Matches), withMetadata, withFix, withType, cveLinkable)
+
+ if withMetadata == 0 {
+ t.Error("no match carries a severity; every row would render as the grey \"?\" badge")
+ }
+ if withType == 0 {
+ t.Error("no match carries a package type; the Package column loses its qualifier for every row")
+ }
+
+ // The other three top-level keys the encoder writes. Two are unused
+ // downstream, which is itself worth pinning: if they ever start being read,
+ // something has to guarantee they are present.
+ var top map[string]json.RawMessage
+ if err := json.Unmarshal(scan.Raw, &top); err != nil {
+ t.Fatalf("parse report: %v", err)
+ }
+ for _, key := range []string{"matches", "source", "distro", "descriptor", "summary"} {
+ if _, ok := top[key]; !ok {
+ t.Errorf("report is missing top-level key %q", key)
+ }
+ }
+
+ // The report's embedded summary must agree with the one sent over the
+ // wire. Nothing downstream reads the embedded copy — the hold takes the
+ // counts from the separate WebSocket field and the appview takes them from
+ // the record — so a divergence would be invisible in production and would
+ // mean a report that contradicts the badge rendered over it.
+ var embedded struct {
+ Summary scanner.VulnerabilitySummary `json:"summary"`
+ }
+ if err := json.Unmarshal(scan.Raw, &embedded); err != nil {
+ t.Fatalf("parse embedded summary: %v", err)
+ }
+ if embedded.Summary != scan.Summary {
+ t.Errorf("the report's embedded summary %+v disagrees with the one sent to the hold %+v",
+ embedded.Summary, scan.Summary)
+ }
+
+ // The descriptor is what a stored report claims it was produced by.
+ if scan.Digest == "" {
+ t.Error("report descriptor carries no version")
+ } else {
+ t.Logf("report descriptor claims grype %s", scan.Digest)
+ }
+
+ // Size is not an assertion, but it belongs next to the shape. The encoder
+ // marshals whole match.Match values, so each row the appview renders from
+ // five fields ships with the vulnerability's full description, every CVSS
+ // vector, every CPE and every match detail. The appview fetches and parses
+ // this blob on each render of the vulnerabilities tab.
+ t.Logf("report is %d bytes for %d matches (%d bytes per rendered row, of which the "+
+ "appview reads ID, Metadata.Severity, Fix.Versions, Package.Name/Version/Type)",
+ len(scan.Raw), len(scan.Report.Matches), len(scan.Raw)/max(1, len(scan.Report.Matches)))
+}
+
+// TestVulnDBScanIsDeterministic scans the same image twice through one
+// harness, so both scans run against the same loaded provider. The findings
+// must be identical: same summary, same set of (vulnerability, package,
+// version) triples.
+//
+// Feed-independent by construction — a feed refresh between the two scans is
+// impossible, since the provider is loaded once and held in memory.
+//
+// It deliberately does not compare the report digests. See the log line at the
+// end for why.
+func TestVulnDBScanIsDeterministic(t *testing.T) {
+ primeVulnDB(t)
+ requireVulnFixture(t, vulnFixtureOld)
+
+ dir := fixtureDir(vulnFixtureOld)
+ h := Start(t, mockhold.NewOCILayout(dir), WithVulnDB(t))
+
+ first := sendVulnScan(t, h, jobFromLayout(t, dir), vulnFixtureOld)
+ second := sendVulnScan(t, h, jobFromLayout(t, dir), vulnFixtureOld)
+
+ if first.Summary != second.Summary {
+ t.Errorf("two scans of the same image against the same database disagree:\n first: %+v\nsecond: %+v",
+ first.Summary, second.Summary)
+ }
+
+ a, b := first.Findings(), second.Findings()
+ if len(a) != len(b) {
+ t.Fatalf("scan produced %d findings, rescan produced %d", len(a), len(b))
+ }
+ for i := range a {
+ if a[i] != b[i] {
+ t.Errorf("finding %d differs: %s vs %s", i, a[i], b[i])
+ }
+ }
+
+ // Content-identical scans, byte-identical reports? Not necessarily. The
+ // report's digest is what the hold stores as vulnDigest, so a report that
+ // changes when nothing about the image did means every rescan uploads a
+ // fresh blob under a fresh digest. Rather than assert either way — this is
+ // current behaviour, not a requirement — name the top-level keys that
+ // moved, so the cause is in the log instead of in a comment's guess.
+ if string(first.Raw) == string(second.Raw) {
+ t.Logf("the two reports are byte-identical (%d bytes), so a rescan of unchanged content "+
+ "reuses the same vulnDigest", len(first.Raw))
+ } else {
+ var ka, kb map[string]json.RawMessage
+ if err := json.Unmarshal(first.Raw, &ka); err != nil {
+ t.Fatalf("parse first report: %v", err)
+ }
+ if err := json.Unmarshal(second.Raw, &kb); err != nil {
+ t.Fatalf("parse second report: %v", err)
+ }
+ var moved []string
+ for k := range ka {
+ if string(ka[k]) != string(kb[k]) {
+ moved = append(moved, k)
+ }
+ }
+ sort.Strings(moved)
+ t.Logf("the two reports differ byte-for-byte despite identical findings (%d vs %d bytes); "+
+ "the top-level keys that changed are %v, so the vulnDigest the hold records changes "+
+ "on every rescan of unchanged content",
+ len(first.Raw), len(second.Raw), moved)
+ if len(moved) == 1 && moved[0] == "source" {
+ t.Logf(" \"source\" is s.Source from the SBOM, whose Name/Reference is the per-job "+
+ "OCI layout path: %s", firstLineOf(string(ka["source"])))
+ }
+ }
+
+ t.Logf("first scan %s, rescan %s (the difference is the database load, paid once)",
+ first.Elapsed.Round(time.Millisecond), second.Elapsed.Round(time.Millisecond))
+}
+
+// TestVulnDBMatchingCost measures what Grype adds to a scan, which every
+// performance number in SCANNER_BUGS.md section 3 explicitly excludes. It
+// asserts nothing about the numbers — a threshold here would be a flake on
+// slower hardware — it records them.
+//
+// Read the RSS figure with the database in mind: it is a ~2 GB SQLite file
+// opened mmap'd, so resident pages grow as the matcher touches the parts of it
+// the image's packages index into. That interacts directly with the 512 MiB
+// GOMEMLIMIT the scanner sets in cmd/scanner/main.go, which is a soft limit
+// the Go runtime applies to the heap and which does not see mapped file pages
+// at all.
+func TestVulnDBMatchingCost(t *testing.T) {
+ if os.Getenv(perfEnv) != "1" {
+ t.Skipf("set %s=1 (with %s=1) to measure Grype's cost", perfEnv, VulnDBEnv)
+ }
+ primeVulnDB(t)
+
+ t.Log("fixture pkgs matches grypeOff grypeOn delta peakRSS report")
+ for _, fixture := range severitySweepFixtures(t) {
+ off, offRSS := runPipelineWithoutGrype(t, fixture)
+ scan := runVulnScan(t, fixture)
+ t.Logf("%-22s %-5d %-8d %-10s %-10s %-10s %-9s %s",
+ scan.Fixture, len(scan.SBOM.Packages), len(scan.Report.Matches),
+ off.Round(time.Millisecond),
+ scan.Elapsed.Round(time.Millisecond),
+ (scan.Elapsed - off).Round(time.Millisecond),
+ mib(scan.PeakRSS), mib(uint64(len(scan.Raw))))
+ // The two RSS figures are not an isolation of the database's cost: the
+ // provider is already loaded and its pages already resident by the
+ // time the Grype-off run happens, and nothing unmaps it. The number
+ // that does isolate it is the before/after pair logged by
+ // primeVulnDB, which straddles the only load in the process.
+ t.Logf(" peak RSS %s on the Grype-off run, %s on the Grype-on run (both after the "+
+ "database was loaded and its pages made resident)", mib(offRSS), mib(scan.PeakRSS))
+ }
+}
+
+// runPipelineWithoutGrype runs the same image through the same pipeline with
+// vuln.enabled false, which is the baseline every performance figure in
+// SCANNER_BUGS.md section 3 was measured against. The difference between this
+// and the Grype-enabled scan is the matching cost those figures excluded.
+func runPipelineWithoutGrype(t *testing.T, fixture string) (time.Duration, uint64) {
+ t.Helper()
+
+ var elapsed time.Duration
+ var peak uint64
+ t.Run(fixture+"/no-grype", func(t *testing.T) {
+ dir := fixtureDir(fixture)
+ h := Start(t, mockhold.NewOCILayout(dir))
+ job := jobFromLayout(t, dir)
+
+ stopRSS := watchRSS()
+ start := time.Now()
+ seq, err := h.Hold.SendJob(job)
+ if err != nil {
+ t.Fatalf("send job: %v", err)
+ }
+ msg := h.AwaitTerminal(t, seq, 15*time.Minute)
+ elapsed = time.Since(start)
+ peak = stopRSS()
+ if msg.Type != "result" {
+ t.Fatalf("%s: want result, got %s: %s%s", fixture, msg.Type, msg.Error, msg.Reason)
+ }
+ })
+ return elapsed, peak
+}
+
+// ---------------------------------------------------------------------------
+// Small helpers
+// ---------------------------------------------------------------------------
+
+func sortedKeys(m map[string]bool) []string {
+ out := make([]string, 0, len(m))
+ for k := range m {
+ out = append(out, k)
+ }
+ sort.Strings(out)
+ return out
+}
+
+func formatHistogram(hist map[string]int) string {
+ keys := make([]string, 0, len(hist))
+ for k := range hist {
+ keys = append(keys, k)
+ }
+ sort.Strings(keys)
+ parts := make([]string, 0, len(keys))
+ for _, k := range keys {
+ parts = append(parts, fmt.Sprintf("%s=%d", k, hist[k]))
+ }
+ return strings.Join(parts, " ")
+}
+
+// firstLineOf collapses a pretty-printed JSON fragment onto one line and
+// truncates it to something a log line can carry.
+func firstLineOf(s string) string {
+ s = strings.Join(strings.Fields(s), " ")
+ if len(s) > 240 {
+ return s[:240] + "…"
+ }
+ return s
+}
diff --git a/scanner/internal/mockhold/blobsource.go b/scanner/internal/mockhold/blobsource.go
new file mode 100644
index 0000000..3a97f8f
--- /dev/null
+++ b/scanner/internal/mockhold/blobsource.go
@@ -0,0 +1,144 @@
+// Package mockhold provides an in-process stand-in for the hold service so
+// scanner tests can drive the real scanner binary through failure scenarios
+// that are impractical to provoke against a live hold: mid-scan disconnects,
+// truncated blobs, stalled downloads, and artifact shapes the real hold's
+// dispatch guards currently refuse to send.
+//
+// The scanner touches exactly three hold endpoints, and this package serves
+// all three:
+//
+// GET /xrpc/io.atcr.hold.subscribeScanJobs WebSocket, jobs out / acks in
+// GET /xrpc/com.atproto.sync.getBlob returns {"url": "..."}
+// GET /blobs/{hex} the bytes that URL points at
+//
+// Blob bytes come from a BlobSource, so the same scenario can run against a
+// real image pulled with skopeo or against synthetic bytes built in the test.
+package mockhold
+
+import (
+ "bytes"
+ "errors"
+ "fmt"
+ "io"
+ "os"
+ "path/filepath"
+ "strings"
+)
+
+// ErrBlobNotFound is returned by a BlobSource that has no bytes for a digest.
+// The mock hold turns this into a 404 from getBlob, which is what a real hold
+// returns for a blob that was garbage collected out from under a queued job.
+var ErrBlobNotFound = errors.New("mockhold: blob not found")
+
+// BlobSource supplies blob bytes by digest. Implementations must be safe for
+// concurrent use: the scanner downloads a config and every layer in sequence,
+// but several workers may be scanning at once.
+//
+// Size is returned separately from the reader so the mock can set
+// Content-Length. A source that does not know the size ahead of time (a
+// generated stream, say) may return -1, and the mock will use chunked
+// encoding.
+type BlobSource interface {
+ Open(digest string) (io.ReadCloser, int64, error)
+}
+
+// DigestHex extracts the hex portion of a digest string, mirroring the
+// scanner's own digestHex so a source keyed either way resolves the same.
+func DigestHex(digest string) string {
+ if _, hex, ok := strings.Cut(digest, ":"); ok {
+ return hex
+ }
+ return digest
+}
+
+// OCILayout serves blobs out of an OCI image layout directory, which is what
+// `skopeo copy docker://[ oci::` writes. The layout stores blobs
+// at blobs/sha256/, keyed by exactly the digest the scan job references,
+// so a layout pulled from a real registry can back a scan with no rewriting.
+type OCILayout struct {
+ Dir string
+}
+
+// NewOCILayout returns a source reading from an OCI layout directory.
+func NewOCILayout(dir string) *OCILayout { return &OCILayout{Dir: dir} }
+
+// Open implements BlobSource.
+func (o *OCILayout) Open(digest string) (io.ReadCloser, int64, error) {
+ path := filepath.Join(o.Dir, "blobs", "sha256", DigestHex(digest))
+
+ // Refuse a digest that escapes the layout. Digests are attacker-controlled
+ // in the sense that a test may deliberately feed a malformed one, and a
+ // traversal would read outside the fixture rather than failing the way a
+ // real hold would.
+ if !strings.HasPrefix(filepath.Clean(path), filepath.Clean(o.Dir)+string(os.PathSeparator)) {
+ return nil, 0, fmt.Errorf("%w: %q escapes the layout", ErrBlobNotFound, digest)
+ }
+
+ f, err := os.Open(path)
+ if err != nil {
+ if os.IsNotExist(err) {
+ return nil, 0, fmt.Errorf("%w: %s", ErrBlobNotFound, digest)
+ }
+ return nil, 0, err
+ }
+
+ fi, err := f.Stat()
+ if err != nil {
+ f.Close()
+ return nil, 0, err
+ }
+
+ return f, fi.Size(), nil
+}
+
+// Memory serves blobs from an in-memory map keyed by digest. Use it for
+// synthetic content: a hand-built tar layer, a gzip bomb, bytes that do not
+// hash to the digest claiming them.
+type Memory struct {
+ blobs map[string][]byte
+}
+
+// NewMemory returns an empty in-memory source.
+func NewMemory() *Memory { return &Memory{blobs: make(map[string][]byte)} }
+
+// Add stores bytes under a digest. The digest is stored by hex, so callers may
+// pass either "sha256:abc" or "abc".
+func (m *Memory) Add(digest string, data []byte) *Memory {
+ m.blobs[DigestHex(digest)] = data
+ return m
+}
+
+// Open implements BlobSource.
+func (m *Memory) Open(digest string) (io.ReadCloser, int64, error) {
+ data, ok := m.blobs[DigestHex(digest)]
+ if !ok {
+ return nil, 0, fmt.Errorf("%w: %s", ErrBlobNotFound, digest)
+ }
+ return io.NopCloser(bytes.NewReader(data)), int64(len(data)), nil
+}
+
+// Multi tries each source in order and returns the first hit. It lets one
+// scenario span several skopeo-pulled layouts, or overlay synthetic bytes on
+// top of a real image to corrupt exactly one layer.
+type Multi struct {
+ sources []BlobSource
+}
+
+// NewMulti returns a source that consults each of sources in order.
+func NewMulti(sources ...BlobSource) *Multi { return &Multi{sources: sources} }
+
+// Open implements BlobSource.
+func (m *Multi) Open(digest string) (io.ReadCloser, int64, error) {
+ for _, s := range m.sources {
+ rc, n, err := s.Open(digest)
+ if err == nil {
+ return rc, n, nil
+ }
+ // A source that failed for a reason other than "I don't have it" is a
+ // real fault and must not be masked by a later source happening to.
+ if !errors.Is(err, ErrBlobNotFound) {
+ return nil, 0, err
+ }
+ }
+ return nil, 0, fmt.Errorf("%w: %s", ErrBlobNotFound, digest)
+}
diff --git a/scanner/internal/mockhold/corpus.go b/scanner/internal/mockhold/corpus.go
new file mode 100644
index 0000000..8875ec4
--- /dev/null
+++ b/scanner/internal/mockhold/corpus.go
@@ -0,0 +1,177 @@
+package mockhold
+
+import (
+ _ "embed"
+ "encoding/json"
+ "fmt"
+ "sort"
+ "strings"
+
+ scanner "atcr.io/scanner"
+)
+
+//go:embed testdata/corpus.json
+var corpusJSON []byte
+
+// Artifact shapes present in the corpus. The scanner treats each differently,
+// and three of the four are shapes the real hold's dispatch guards currently
+// refuse to send, which is precisely why they are worth being able to send.
+const (
+ // ShapeImage is an ordinary container image: a config plus tar layers.
+ // The only shape the scanner is expected to scan through to an SBOM.
+ ShapeImage = "image"
+
+ // ShapeAttestation is an in-toto or DSSE payload carried by a manifest
+ // with an ordinary image config. The config media type alone does not
+ // identify it, which is what let one reach production.
+ ShapeAttestation = "attestation"
+
+ // ShapeHelm is a Helm chart, identified by its config media type.
+ ShapeHelm = "helm"
+
+ // ShapeIndex is a multi-arch manifest list: no config, no layers, just a
+ // manifests array. Both of the hold's enqueue paths filter these out via
+ // HasScannableContent, so the scanner never sees one in production. Fed
+ // one directly it does not skip: skipReason returns empty (no config
+ // media type to match, and the layer check is skipped when there are no
+ // layers), then buildOCILayout fails with "config blob has empty digest",
+ // which is a retryable error rather than a permanent skip.
+ ShapeIndex = "index"
+
+ // ShapeReferrer is an artifact with a subject pointing at another
+ // manifest. Also filtered by HasScannableContent.
+ ShapeReferrer = "referrer"
+)
+
+// Manifest is one real io.atcr.manifest record, reduced to the fields a scan
+// job carries. Blob bytes are not included: descriptors are enough for every
+// scenario that does not run Syft, and the ones that do pair the corpus with
+// an OCILayout pulled by testdata/fetch-blobs.sh.
+type Manifest struct {
+ Digest string `json:"digest"`
+ Repository string `json:"repository"`
+ Shape string `json:"shape"`
+ MediaType string `json:"mediaType"`
+ HoldDID string `json:"holdDid"`
+ Config *scanner.BlobDescriptor `json:"config"`
+ Layers []scanner.BlobDescriptor `json:"layers"`
+ Manifests []scanner.BlobDescriptor `json:"manifests,omitempty"`
+ Subject *scanner.BlobDescriptor `json:"subject,omitempty"`
+}
+
+type corpusDoc struct {
+ SourceRepo string `json:"sourceRepo"`
+ SourceHandle string `json:"sourceHandle"`
+ Manifests []Manifest `json:"manifests"`
+}
+
+// Corpus returns every manifest in the embedded fixture.
+//
+// These are real records fetched anonymously from a live PDS, not
+// hand-written shapes, so the distribution reflects what a hold actually
+// holds: mostly ordinary images, a substantial minority of buildx
+// attestations, a handful of indexes and Helm charts.
+func Corpus() ([]Manifest, error) {
+ var doc corpusDoc
+ if err := json.Unmarshal(corpusJSON, &doc); err != nil {
+ return nil, fmt.Errorf("mockhold: parse corpus: %w", err)
+ }
+ return doc.Manifests, nil
+}
+
+// CorpusByShape returns the manifests matching one of the Shape constants.
+func CorpusByShape(shape string) ([]Manifest, error) {
+ all, err := Corpus()
+ if err != nil {
+ return nil, err
+ }
+ var out []Manifest
+ for _, m := range all {
+ if m.Shape == shape {
+ out = append(out, m)
+ }
+ }
+ return out, nil
+}
+
+// Job builds a scan job from a manifest record, the same conversion the hold
+// performs in dispatchCandidate. HoldEndpoint is left empty so SendJob points
+// it at the mock.
+func (m Manifest) Job() *scanner.ScanJob {
+ job := &scanner.ScanJob{
+ ManifestDigest: m.Digest,
+ Repository: m.Repository,
+ Tag: "latest",
+ HoldDID: m.HoldDID,
+ Tier: "deckhand",
+ Layers: m.Layers,
+ }
+ if m.Config != nil {
+ job.Config = *m.Config
+ }
+ return job
+}
+
+// Digests returns every blob digest a scan of this manifest will request, in
+// the order buildOCILayout asks for them: config first, then layers. Tests use
+// it to build a Memory source, or to assert which blobs were actually fetched.
+func (m Manifest) Digests() []string {
+ var out []string
+ if m.Config != nil && m.Config.Digest != "" {
+ out = append(out, m.Config.Digest)
+ }
+ for _, l := range m.Layers {
+ if l.Digest != "" {
+ out = append(out, l.Digest)
+ }
+ }
+ return out
+}
+
+// Signature identifies a manifest's media-type shape: its config media type
+// plus the distinct set of layer media types. It is what the scanner's
+// skipReason actually branches on, so two manifests sharing a signature are
+// indistinguishable to the pipeline.
+func (m Manifest) Signature() string {
+ cfg := "(none)"
+ if m.Config != nil {
+ cfg = m.Config.MediaType
+ }
+
+ seen := map[string]bool{}
+ var types []string
+ for _, l := range m.Layers {
+ if seen[l.MediaType] {
+ continue
+ }
+ seen[l.MediaType] = true
+ types = append(types, l.MediaType)
+ }
+ sort.Strings(types)
+
+ return cfg + " | " + strings.Join(types, ",")
+}
+
+// Representatives returns one manifest per distinct Signature, preserving
+// corpus order.
+//
+// The corpus is real production data, so its distribution is lopsided: 16 of
+// the 84 manifests are buildx attestations of exactly the same shape. Running
+// all 16 through the pipeline tests one code path 16 times, and at the
+// production JobCooldown that costs two and a half minutes to learn nothing
+// the first one did not already establish. Scenarios that care about breadth
+// of shape should use this; scenarios that care about volume should say so
+// explicitly and set their own cooldown.
+func Representatives(manifests []Manifest) []Manifest {
+ seen := map[string]bool{}
+ var out []Manifest
+ for _, m := range manifests {
+ sig := m.Signature()
+ if seen[sig] {
+ continue
+ }
+ seen[sig] = true
+ out = append(out, m)
+ }
+ return out
+}
diff --git a/scanner/internal/mockhold/mockhold.go b/scanner/internal/mockhold/mockhold.go
new file mode 100644
index 0000000..b2b6bab
--- /dev/null
+++ b/scanner/internal/mockhold/mockhold.go
@@ -0,0 +1,473 @@
+package mockhold
+
+import (
+ "encoding/json"
+ "errors"
+ "fmt"
+ "io"
+ "net/http"
+ "net/http/httptest"
+ "sync"
+ "sync/atomic"
+ "time"
+
+ scanner "atcr.io/scanner"
+ "github.com/gorilla/websocket"
+)
+
+// DropMode selects how DropConnections severs a scanner's WebSocket.
+//
+// The distinction is not cosmetic. The scanner's read loop passes errors
+// through websocket.IsUnexpectedCloseError, which treats CloseNormalClosure
+// and CloseGoingAway as expected and logs anything else as an error. Only
+// DropAbrupt resembles what a proxy timeout or an evicted pod does in
+// production, so a scenario that only ever closes cleanly tests the easy half.
+type DropMode int
+
+const (
+ // DropAbrupt closes the underlying TCP connection with no close frame.
+ DropAbrupt DropMode = iota
+ // DropClean sends a normal-closure frame before closing.
+ DropClean
+)
+
+// Message is a message received from the scanner. It mirrors the hold's own
+// ScannerMessage; the scanner module cannot import pkg/hold/pds without
+// dragging go-libsql into the scanner binary, so the shape is restated here.
+type Message struct {
+ Type string `json:"type"`
+ Seq int64 `json:"seq"`
+ SBOM string `json:"sbom,omitempty"`
+ VulnReport string `json:"vulnReport,omitempty"`
+ Summary *scanner.VulnerabilitySummary `json:"summary,omitempty"`
+ Error string `json:"error,omitempty"`
+ Reason string `json:"reason,omitempty"`
+
+ // At records arrival time. Several scenarios assert on timing rather than
+ // content: how long the scanner waited before redialing, or whether a
+ // result landed before or after a disconnect.
+ At time.Time `json:"-"`
+}
+
+// BlobRequest records one blob fetch. Counting these is how a test proves the
+// scanner downloaded the same image twice after a reconnect, which no message
+// in the transcript would reveal on its own.
+type BlobRequest struct {
+ Digest string
+ At time.Time
+}
+
+// Hold is an in-process stand-in for the hold service.
+type Hold struct {
+ // Secret, when non-empty, is required as the ?secret= query parameter on
+ // the subscribe endpoint, matching the real hold's ValidateScannerSecret.
+ Secret string
+
+ blobs BlobSource
+ srv *httptest.Server
+ seq atomic.Int64
+ closed atomic.Bool
+
+ mu sync.Mutex
+ conns []*websocket.Conn
+ transcript []Message
+ blobReqs []BlobRequest
+ dials []time.Time
+ msgWaiters []*msgWaiter
+ connWaiters []chan struct{}
+ presignHook func(digest string) (string, bool)
+ blobRespHook func(w http.ResponseWriter, r *http.Request, digest string) bool
+}
+
+// msgWaiter is one in-flight WaitForMessage call. Waiters stay registered
+// until the call returns, rather than being consumed by the first delivery:
+// a scanner emits an ack before its terminal message, so a waiter that
+// unregistered on the first non-matching message would sleep through the one
+// it was waiting for.
+type msgWaiter struct {
+ ch chan Message
+}
+
+// Option configures a Hold at construction.
+type Option func(*Hold)
+
+// WithSecret requires the scanner to present this shared secret.
+func WithSecret(secret string) Option {
+ return func(h *Hold) { h.Secret = secret }
+}
+
+// WithPresignHook installs a hook consulted by getBlob before the default
+// response. Returning ok==false makes getBlob answer 404, standing in for a
+// blob the hold can no longer resolve. Returning a URL redirects the scanner
+// somewhere else entirely, which is how a test points one layer at a slow or
+// broken server while the rest come from the fixture.
+func WithPresignHook(fn func(digest string) (string, bool)) Option {
+ return func(h *Hold) { h.presignHook = fn }
+}
+
+// WithBlobResponseHook installs a hook that may write the blob response
+// itself. Returning true means the hook handled the request; the default
+// BlobSource path is then skipped. Use it to truncate, stall, or return a
+// 500 partway through a body.
+func WithBlobResponseHook(fn func(w http.ResponseWriter, r *http.Request, digest string) bool) Option {
+ return func(h *Hold) { h.blobRespHook = fn }
+}
+
+// New starts a mock hold serving blobs from the given source. Call Close when
+// finished. A nil source answers every getBlob with 404, which is a valid
+// configuration for scenarios that never reach a download.
+func New(blobs BlobSource, opts ...Option) *Hold {
+ h := &Hold{blobs: blobs}
+ for _, opt := range opts {
+ opt(h)
+ }
+
+ mux := http.NewServeMux()
+ mux.HandleFunc("/xrpc/io.atcr.hold.subscribeScanJobs", h.handleSubscribe)
+ mux.HandleFunc("/xrpc/com.atproto.sync.getBlob", h.handleGetBlob)
+ mux.HandleFunc("/blobs/", h.handleBlob)
+ h.srv = httptest.NewServer(mux)
+
+ return h
+}
+
+// URL returns the base URL of the mock hold, suitable for both the scanner's
+// hold.url config and a scan job's HoldEndpoint.
+func (h *Hold) URL() string { return h.srv.URL }
+
+// Close shuts down the server and severs any live scanner connections.
+func (h *Hold) Close() {
+ if !h.closed.CompareAndSwap(false, true) {
+ return
+ }
+ h.DropConnections(DropAbrupt)
+ h.srv.Close()
+}
+
+// --- WebSocket -------------------------------------------------------------
+
+var upgrader = websocket.Upgrader{
+ CheckOrigin: func(*http.Request) bool { return true },
+}
+
+func (h *Hold) handleSubscribe(w http.ResponseWriter, r *http.Request) {
+ if h.Secret != "" && r.URL.Query().Get("secret") != h.Secret {
+ http.Error(w, "invalid scanner secret", http.StatusUnauthorized)
+ return
+ }
+
+ conn, err := upgrader.Upgrade(w, r, nil)
+ if err != nil {
+ return
+ }
+
+ h.mu.Lock()
+ h.conns = append(h.conns, conn)
+ h.dials = append(h.dials, time.Now())
+ waiters := h.connWaiters
+ h.connWaiters = nil
+ h.mu.Unlock()
+
+ for _, ch := range waiters {
+ close(ch)
+ }
+
+ go h.readLoop(conn)
+}
+
+func (h *Hold) readLoop(conn *websocket.Conn) {
+ defer func() {
+ conn.Close()
+ h.mu.Lock()
+ for i, c := range h.conns {
+ if c == conn {
+ h.conns = append(h.conns[:i], h.conns[i+1:]...)
+ break
+ }
+ }
+ h.mu.Unlock()
+ }()
+
+ for {
+ _, data, err := conn.ReadMessage()
+ if err != nil {
+ return
+ }
+
+ var msg Message
+ if err := json.Unmarshal(data, &msg); err != nil {
+ continue
+ }
+ msg.At = time.Now()
+
+ h.mu.Lock()
+ h.transcript = append(h.transcript, msg)
+ waiters := append([]*msgWaiter(nil), h.msgWaiters...)
+ h.mu.Unlock()
+
+ // Buffered, non-blocking: a waiter that has already returned must not
+ // wedge the read loop, and losing a message for such a waiter is
+ // harmless since nobody is listening.
+ for _, w := range waiters {
+ select {
+ case w.ch <- msg:
+ default:
+ }
+ }
+ }
+}
+
+// SendJob dispatches a scan job to a connected scanner and returns the seq it
+// was assigned. HoldEndpoint is filled in automatically so blob downloads come
+// back to this server; an explicitly set HoldEndpoint is left alone.
+//
+// Unlike the real hold there is no round-robin, no persistence, and no
+// assignment state: the job goes to the first live connection. Anything a test
+// needs to assert about pending/assigned/processing bookkeeping belongs in a
+// hold-side test against the real ScanBroadcaster, not here.
+func (h *Hold) SendJob(job *scanner.ScanJob) (int64, error) {
+ if job.Seq == 0 {
+ job.Seq = h.seq.Add(1)
+ }
+ if job.HoldEndpoint == "" {
+ job.HoldEndpoint = h.srv.URL
+ }
+
+ h.mu.Lock()
+ if len(h.conns) == 0 {
+ h.mu.Unlock()
+ return 0, errors.New("mockhold: no scanner connected")
+ }
+ conn := h.conns[0]
+ h.mu.Unlock()
+
+ // The scanner decodes into ScanJobRaw, where Config and Layers are raw
+ // JSON, so marshal them separately rather than relying on ScanJob's own
+ // shape lining up with the wire format.
+ configJSON, err := json.Marshal(job.Config)
+ if err != nil {
+ return 0, fmt.Errorf("marshal config: %w", err)
+ }
+ layersJSON, err := json.Marshal(job.Layers)
+ if err != nil {
+ return 0, fmt.Errorf("marshal layers: %w", err)
+ }
+
+ raw := scanner.ScanJobRaw{
+ Type: "job",
+ Seq: job.Seq,
+ ManifestDigest: job.ManifestDigest,
+ Repository: job.Repository,
+ Tag: job.Tag,
+ UserDID: job.UserDID,
+ UserHandle: job.UserHandle,
+ HoldDID: job.HoldDID,
+ HoldEndpoint: job.HoldEndpoint,
+ Tier: job.Tier,
+ Config: configJSON,
+ Layers: layersJSON,
+ }
+
+ data, err := json.Marshal(raw)
+ if err != nil {
+ return 0, fmt.Errorf("marshal job: %w", err)
+ }
+ if err := conn.WriteMessage(websocket.TextMessage, data); err != nil {
+ return 0, fmt.Errorf("write job: %w", err)
+ }
+ return job.Seq, nil
+}
+
+// NextSeq allocates a sequence number from the same counter SendJob uses, so
+// a hand-built raw frame cannot collide with a job sent the normal way.
+func (h *Hold) NextSeq() int64 { return h.seq.Add(1) }
+
+// SendRaw writes bytes to the first connected scanner verbatim, with no
+// marshaling and no validation.
+//
+// SendJob always produces a well-formed frame, which makes it useless for the
+// half of the protocol surface that matters most: what the scanner does with
+// a frame it cannot parse. Those paths (bad JSON, a config that is not an
+// object, an absent layers field) are reachable in production from any hold
+// whose stored config_json/layers_json columns disagree with what the scanner
+// expects, and the scanner's response to them is to say nothing at all.
+func (h *Hold) SendRaw(data []byte) error {
+ h.mu.Lock()
+ if len(h.conns) == 0 {
+ h.mu.Unlock()
+ return errors.New("mockhold: no scanner connected")
+ }
+ conn := h.conns[0]
+ h.mu.Unlock()
+
+ if err := conn.WriteMessage(websocket.TextMessage, data); err != nil {
+ return fmt.Errorf("write raw frame: %w", err)
+ }
+ return nil
+}
+
+// DropConnections severs every live scanner connection. The scanner's Connect
+// loop redials on its own, so a test drops and then waits for the next dial.
+func (h *Hold) DropConnections(mode DropMode) {
+ h.mu.Lock()
+ conns := append([]*websocket.Conn(nil), h.conns...)
+ h.mu.Unlock()
+
+ for _, c := range conns {
+ if mode == DropClean {
+ _ = c.WriteControl(websocket.CloseMessage,
+ websocket.FormatCloseMessage(websocket.CloseNormalClosure, ""),
+ time.Now().Add(time.Second))
+ c.Close()
+ continue
+ }
+ // Abrupt: kill the TCP connection with no close frame at all.
+ _ = c.UnderlyingConn().Close()
+ }
+}
+
+// --- Observation -----------------------------------------------------------
+
+// Transcript returns every message received so far, in arrival order.
+func (h *Hold) Transcript() []Message {
+ h.mu.Lock()
+ defer h.mu.Unlock()
+ return append([]Message(nil), h.transcript...)
+}
+
+// BlobRequests returns every blob fetch so far, in arrival order.
+func (h *Hold) BlobRequests() []BlobRequest {
+ h.mu.Lock()
+ defer h.mu.Unlock()
+ return append([]BlobRequest(nil), h.blobReqs...)
+}
+
+// Dials returns the time of each scanner connection, which is how a test
+// measures reconnect cadence.
+func (h *Hold) Dials() []time.Time {
+ h.mu.Lock()
+ defer h.mu.Unlock()
+ return append([]time.Time(nil), h.dials...)
+}
+
+// WaitForScanner blocks until a scanner is connected or the timeout elapses.
+func (h *Hold) WaitForScanner(timeout time.Duration) error {
+ h.mu.Lock()
+ if len(h.conns) > 0 {
+ h.mu.Unlock()
+ return nil
+ }
+ ch := make(chan struct{})
+ h.connWaiters = append(h.connWaiters, ch)
+ h.mu.Unlock()
+
+ select {
+ case <-ch:
+ return nil
+ case <-time.After(timeout):
+ return errors.New("mockhold: timed out waiting for scanner to connect")
+ }
+}
+
+// WaitForMessage blocks until a message satisfying match arrives, or the
+// timeout elapses. Messages already in the transcript are considered first, so
+// a test cannot lose a race against a fast scanner.
+func (h *Hold) WaitForMessage(match func(Message) bool, timeout time.Duration) (Message, error) {
+ h.mu.Lock()
+ for _, m := range h.transcript {
+ if match(m) {
+ h.mu.Unlock()
+ return m, nil
+ }
+ }
+ w := &msgWaiter{ch: make(chan Message, 256)}
+ h.msgWaiters = append(h.msgWaiters, w)
+ h.mu.Unlock()
+
+ defer func() {
+ h.mu.Lock()
+ for i, cand := range h.msgWaiters {
+ if cand == w {
+ h.msgWaiters = append(h.msgWaiters[:i], h.msgWaiters[i+1:]...)
+ break
+ }
+ }
+ h.mu.Unlock()
+ }()
+
+ deadline := time.After(timeout)
+ for {
+ select {
+ case m := <-w.ch:
+ if match(m) {
+ return m, nil
+ }
+ case <-deadline:
+ return Message{}, errors.New("mockhold: timed out waiting for message")
+ }
+ }
+}
+
+// --- Blobs -----------------------------------------------------------------
+
+// handleGetBlob answers com.atproto.sync.getBlob with the URL the scanner
+// should download from. The real hold returns a presigned S3 URL; here it
+// points back at this server, or wherever a presign hook sends it.
+func (h *Hold) handleGetBlob(w http.ResponseWriter, r *http.Request) {
+ digest := r.URL.Query().Get("cid")
+ if digest == "" {
+ http.Error(w, "missing cid", http.StatusBadRequest)
+ return
+ }
+
+ if h.presignHook != nil {
+ url, ok := h.presignHook(digest)
+ if !ok {
+ http.Error(w, "blob not found", http.StatusNotFound)
+ return
+ }
+ writeJSON(w, map[string]string{"url": url})
+ return
+ }
+
+ writeJSON(w, map[string]string{"url": h.srv.URL + "/blobs/" + DigestHex(digest)})
+}
+
+func (h *Hold) handleBlob(w http.ResponseWriter, r *http.Request) {
+ digest := r.URL.Path[len("/blobs/"):]
+
+ h.mu.Lock()
+ h.blobReqs = append(h.blobReqs, BlobRequest{Digest: digest, At: time.Now()})
+ h.mu.Unlock()
+
+ if h.blobRespHook != nil && h.blobRespHook(w, r, digest) {
+ return
+ }
+
+ if h.blobs == nil {
+ http.Error(w, "no blob source", http.StatusNotFound)
+ return
+ }
+
+ rc, size, err := h.blobs.Open(digest)
+ if err != nil {
+ if errors.Is(err, ErrBlobNotFound) {
+ http.Error(w, "blob not found", http.StatusNotFound)
+ return
+ }
+ http.Error(w, err.Error(), http.StatusInternalServerError)
+ return
+ }
+ defer rc.Close()
+
+ w.Header().Set("Content-Type", "application/octet-stream")
+ if size >= 0 {
+ w.Header().Set("Content-Length", fmt.Sprintf("%d", size))
+ }
+ _, _ = io.Copy(w, rc)
+}
+
+func writeJSON(w http.ResponseWriter, v any) {
+ w.Header().Set("Content-Type", "application/json")
+ _ = json.NewEncoder(w).Encode(v)
+}
diff --git a/scanner/internal/mockhold/testdata/.gitignore b/scanner/internal/mockhold/testdata/.gitignore
new file mode 100644
index 0000000..c799fd7
--- /dev/null
+++ b/scanner/internal/mockhold/testdata/.gitignore
@@ -0,0 +1,2 @@
+blobs/
+vulndb/
diff --git a/scanner/internal/mockhold/testdata/corpus.json b/scanner/internal/mockhold/testdata/corpus.json
new file mode 100644
index 0000000..9047052
--- /dev/null
+++ b/scanner/internal/mockhold/testdata/corpus.json
@@ -0,0 +1,4197 @@
+{
+ "_comment": "Real io.atcr.manifest records, fetched anonymously via com.atproto.repo.listRecords. Regenerate with testdata/fetch-corpus.sh. Descriptors only, no blob bytes.",
+ "sourceRepo": "did:plc:pddp4xt5lgnv2qsegbzzs4xg",
+ "sourceHandle": "evan.jarrett.net",
+ "manifests": [
+ {
+ "digest": "sha256:b2c92262f645eb07f1d4a4f7dbc709a5c168c27650af90030589a5b7bc1ca862",
+ "repository": "git-summarizer",
+ "shape": "attestation",
+ "mediaType": "application/vnd.oci.image.manifest.v1+json",
+ "holdDid": "did:plc:wnbpdx4bn5uts5yd2o2g3gnm",
+ "config": {
+ "digest": "sha256:7e834a515452e0f13c5b966af014a7d4141a22903bd8260f864ba93f9059e08d",
+ "size": 167,
+ "mediaType": "application/vnd.oci.image.config.v1+json"
+ },
+ "layers": [
+ {
+ "digest": "sha256:650dab468855379caabbaee98f3e036e13ae105409f8483a5225c5ad5a7735be",
+ "size": 1342,
+ "mediaType": "application/vnd.in-toto+json"
+ }
+ ]
+ },
+ {
+ "digest": "sha256:d2aa86dba7c6400b179e188c6f67f10620a9d9224d0d24bc607d12dbcd10ee9d",
+ "repository": "git-summarizer",
+ "shape": "attestation",
+ "mediaType": "application/vnd.oci.image.manifest.v1+json",
+ "holdDid": "did:plc:wnbpdx4bn5uts5yd2o2g3gnm",
+ "config": {
+ "digest": "sha256:9b5b4f30bf734771144b39af4fe66451ffb023ea1e2464990e0b6ba170a294e4",
+ "size": 167,
+ "mediaType": "application/vnd.oci.image.config.v1+json"
+ },
+ "layers": [
+ {
+ "digest": "sha256:fd8034389245a91dca9e37507cf1e538d91d5d08ff8fcc0087f875cf985b5eb6",
+ "size": 1342,
+ "mediaType": "application/vnd.in-toto+json"
+ }
+ ]
+ },
+ {
+ "digest": "sha256:488f2502332aa6ef82cad82136841c61d7c6458db6ab425aa6e86a06301f5cc1",
+ "repository": "hsm-secrets-operator",
+ "shape": "attestation",
+ "mediaType": "application/vnd.oci.image.manifest.v1+json",
+ "holdDid": "did:plc:wnbpdx4bn5uts5yd2o2g3gnm",
+ "config": {
+ "digest": "sha256:81efb464c0f52e242d60c45c2f5481ffa6aa95c93ebe95ac14f8d552a9143de3",
+ "size": 167,
+ "mediaType": "application/vnd.oci.image.config.v1+json"
+ },
+ "layers": [
+ {
+ "digest": "sha256:da04d5b7053971e8a1e6ac7a0db7491f0ded64212091ab759063870802548750",
+ "size": 1588,
+ "mediaType": "application/vnd.in-toto+json"
+ }
+ ]
+ },
+ {
+ "digest": "sha256:c39092f65e47ae1351f7232e17c36a16171a6987f89737f00b0a13c4f1dd71c7",
+ "repository": "hsm-secrets-operator",
+ "shape": "attestation",
+ "mediaType": "application/vnd.oci.image.manifest.v1+json",
+ "holdDid": "did:plc:wnbpdx4bn5uts5yd2o2g3gnm",
+ "config": {
+ "digest": "sha256:d670f91ae5bfc1717c542f31b78f46db8929766b3d1f1234c11349dc0972f49d",
+ "size": 167,
+ "mediaType": "application/vnd.oci.image.config.v1+json"
+ },
+ "layers": [
+ {
+ "digest": "sha256:a099d3324db8eadca3cce2a41d8224945e6e90b655ff6a3501e65c680a660188",
+ "size": 1588,
+ "mediaType": "application/vnd.in-toto+json"
+ }
+ ]
+ },
+ {
+ "digest": "sha256:13f458be65ba8c912d4e6c7b05275be4a16c0e51a2e8e36ca4983d2f2c2d0f5c",
+ "repository": "loom",
+ "shape": "attestation",
+ "mediaType": "application/vnd.oci.image.manifest.v1+json",
+ "holdDid": "did:plc:wnbpdx4bn5uts5yd2o2g3gnm",
+ "config": {
+ "digest": "sha256:6936ec3e783987ceb14f5cbf40873324cff8882944f4dfc8bc629a70b59ef58c",
+ "size": 167,
+ "mediaType": "application/vnd.oci.image.config.v1+json"
+ },
+ "layers": [
+ {
+ "digest": "sha256:7884036a1a4626c2d1eb2ee9038c759df417368ac88427a95ff2e1b9e170ced5",
+ "size": 1405,
+ "mediaType": "application/vnd.in-toto+json"
+ }
+ ]
+ },
+ {
+ "digest": "sha256:27483bcce096f7b40bad8d8e0030d308b86fde7ced52f4c6eb7f2d06e1bd2eb6",
+ "repository": "loom",
+ "shape": "attestation",
+ "mediaType": "application/vnd.oci.image.manifest.v1+json",
+ "holdDid": "did:plc:wnbpdx4bn5uts5yd2o2g3gnm",
+ "config": {
+ "digest": "sha256:f17133bfe37f5734495d11f77fcd427c740fb8a5f7c76984f8706edbfc8581bf",
+ "size": 167,
+ "mediaType": "application/vnd.oci.image.config.v1+json"
+ },
+ "layers": [
+ {
+ "digest": "sha256:40709380dbb0a9af644d4d71be3b4c93aaebeeeb60d17629ce9988ddfe7a2c7c",
+ "size": 1406,
+ "mediaType": "application/vnd.in-toto+json"
+ }
+ ]
+ },
+ {
+ "digest": "sha256:408a31b822a77e20150624ace1b2d9cc525e964e645d80f1b0cfbb9f2d79ca93",
+ "repository": "loom",
+ "shape": "attestation",
+ "mediaType": "application/vnd.oci.image.manifest.v1+json",
+ "holdDid": "did:plc:wnbpdx4bn5uts5yd2o2g3gnm",
+ "config": {
+ "digest": "sha256:4b1918a801df549cb8b11deecac158bc26547bf1b357be32a2a4ceaa7a5ecd36",
+ "size": 167,
+ "mediaType": "application/vnd.oci.image.config.v1+json"
+ },
+ "layers": [
+ {
+ "digest": "sha256:9ea975b684261448c80a696be18142bc618cd9598ab7b60e4ded8cdd2e7c464e",
+ "size": 1392,
+ "mediaType": "application/vnd.in-toto+json"
+ }
+ ]
+ },
+ {
+ "digest": "sha256:4448d5a773678767d094b56ec8a12254edb13b7069cb201cc92078097a529f21",
+ "repository": "loom",
+ "shape": "attestation",
+ "mediaType": "application/vnd.oci.image.manifest.v1+json",
+ "holdDid": "did:plc:wnbpdx4bn5uts5yd2o2g3gnm",
+ "config": {
+ "digest": "sha256:fef29f487c40a92cbcbaab2c2f87b42ec96218dd9c11bc3dce97f668f9024758",
+ "size": 167,
+ "mediaType": "application/vnd.oci.image.config.v1+json"
+ },
+ "layers": [
+ {
+ "digest": "sha256:9246f46d7e66ce25c8625ae210b4d37ac9295aad12e2e9e376ab229a2a447a95",
+ "size": 1406,
+ "mediaType": "application/vnd.in-toto+json"
+ }
+ ]
+ },
+ {
+ "digest": "sha256:4613affa7cf617c9e21c39c58996445d0d305b290320b13e6f5e4e2c8581b0ea",
+ "repository": "loom",
+ "shape": "attestation",
+ "mediaType": "application/vnd.oci.image.manifest.v1+json",
+ "holdDid": "did:plc:wnbpdx4bn5uts5yd2o2g3gnm",
+ "config": {
+ "digest": "sha256:e6149ae99ce918b544bc0fadceb492bd9fdb2634cff28841e35c79a834101b7f",
+ "size": 167,
+ "mediaType": "application/vnd.oci.image.config.v1+json"
+ },
+ "layers": [
+ {
+ "digest": "sha256:276bd968c910973858b0db48e5299f9170923703269afe3266fe33d39ec1b216",
+ "size": 1233,
+ "mediaType": "application/vnd.in-toto+json"
+ }
+ ]
+ },
+ {
+ "digest": "sha256:50d74fa5b25de41f61de45d0513ae04bad372bc20ee0c8a270a749bde833ae39",
+ "repository": "loom",
+ "shape": "attestation",
+ "mediaType": "application/vnd.oci.image.manifest.v1+json",
+ "holdDid": "did:plc:wnbpdx4bn5uts5yd2o2g3gnm",
+ "config": {
+ "digest": "sha256:67a929c00f8329fd7a72d5e09b93b69c0649dc7c36c30a9ab86cf1fc92c3b4fd",
+ "size": 167,
+ "mediaType": "application/vnd.oci.image.config.v1+json"
+ },
+ "layers": [
+ {
+ "digest": "sha256:acf2557abc08603d232d425bcb3889b9ff3f09ba0800960c7952eddc3225a661",
+ "size": 1392,
+ "mediaType": "application/vnd.in-toto+json"
+ }
+ ]
+ },
+ {
+ "digest": "sha256:73f6b098d970573e87405f052b3523c5016a06870d0466a68210fe6aea6fe24d",
+ "repository": "loom",
+ "shape": "attestation",
+ "mediaType": "application/vnd.oci.image.manifest.v1+json",
+ "holdDid": "did:plc:wnbpdx4bn5uts5yd2o2g3gnm",
+ "config": {
+ "digest": "sha256:b2da9b34c7cb182afb7f37f7580cb2f48422ddf5092e5cb5ade880a2a0bc40f9",
+ "size": 167,
+ "mediaType": "application/vnd.oci.image.config.v1+json"
+ },
+ "layers": [
+ {
+ "digest": "sha256:1783184d0e61ddb69e43bfa5f3b7393e044c5127eb209d7e72053206a31b2d08",
+ "size": 1233,
+ "mediaType": "application/vnd.in-toto+json"
+ }
+ ]
+ },
+ {
+ "digest": "sha256:b8105431d01fcd3cc56e6e38fcc3b57844d05ce99e1118f60cc8632df3f110d5",
+ "repository": "loom",
+ "shape": "attestation",
+ "mediaType": "application/vnd.oci.image.manifest.v1+json",
+ "holdDid": "did:plc:wnbpdx4bn5uts5yd2o2g3gnm",
+ "config": {
+ "digest": "sha256:7ef42846e7eff4714961e697d3def4581411feb7f6ed560993d63bb91935d698",
+ "size": 167,
+ "mediaType": "application/vnd.oci.image.config.v1+json"
+ },
+ "layers": [
+ {
+ "digest": "sha256:84b1e1379a6f62617382cd587bab67312e9a893716da38833ce9599cfefa3300",
+ "size": 1405,
+ "mediaType": "application/vnd.in-toto+json"
+ }
+ ]
+ },
+ {
+ "digest": "sha256:e3961cbe85688a825d5642eadb4e850fad94af94dc1a6e21d383322ecab84a2c",
+ "repository": "loom",
+ "shape": "attestation",
+ "mediaType": "application/vnd.oci.image.manifest.v1+json",
+ "holdDid": "did:plc:wnbpdx4bn5uts5yd2o2g3gnm",
+ "config": {
+ "digest": "sha256:7db6e28af89f307e55c7a8cecfe06c71e8f38496e5db88fb6941e07287d20faa",
+ "size": 167,
+ "mediaType": "application/vnd.oci.image.config.v1+json"
+ },
+ "layers": [
+ {
+ "digest": "sha256:70ee1534dc9f7a5e95c68182b02fc8e08d664568d08546dd8002c645106c8741",
+ "size": 1405,
+ "mediaType": "application/vnd.in-toto+json"
+ }
+ ]
+ },
+ {
+ "digest": "sha256:f330312e825cd1de8406ddc5eb09a566cd47ac1b9aa543bd94636a332b27f72e",
+ "repository": "loom",
+ "shape": "attestation",
+ "mediaType": "application/vnd.oci.image.manifest.v1+json",
+ "holdDid": "did:plc:wnbpdx4bn5uts5yd2o2g3gnm",
+ "config": {
+ "digest": "sha256:654acdcdb075cdacd5fe3ccb0586130d04fbed80b0d68c38366106d3d2483eb2",
+ "size": 167,
+ "mediaType": "application/vnd.oci.image.config.v1+json"
+ },
+ "layers": [
+ {
+ "digest": "sha256:b457880477eecd215d11cb95e00a09fba755e1b35f5aa0e6c20b1970b0926a91",
+ "size": 1405,
+ "mediaType": "application/vnd.in-toto+json"
+ }
+ ]
+ },
+ {
+ "digest": "sha256:21eb1f1fd68ca89fa06da8a054e2a366b1b4103167077f42f7258851ee38f0cd",
+ "repository": "loom-runner",
+ "shape": "attestation",
+ "mediaType": "application/vnd.oci.image.manifest.v1+json",
+ "holdDid": "did:plc:wnbpdx4bn5uts5yd2o2g3gnm",
+ "config": {
+ "digest": "sha256:47f009509e5ffc0d0176dd54846b9957a0907811c0b648919b952a6a66582cb7",
+ "size": 167,
+ "mediaType": "application/vnd.oci.image.config.v1+json"
+ },
+ "layers": [
+ {
+ "digest": "sha256:f710505196d3e262dc7ac433b9062698b36427c14a26a17296f4ef9c0b95fda9",
+ "size": 1343,
+ "mediaType": "application/vnd.in-toto+json"
+ }
+ ]
+ },
+ {
+ "digest": "sha256:93670b296ef1774f53e088b0a3cec6fb737ecf7fe6299d8cd0c2210759a38276",
+ "repository": "loom-runner",
+ "shape": "attestation",
+ "mediaType": "application/vnd.oci.image.manifest.v1+json",
+ "holdDid": "did:plc:wnbpdx4bn5uts5yd2o2g3gnm",
+ "config": {
+ "digest": "sha256:0fc1b479578081e36aa37905f939882dd7a701993a9d14f0f77ede1f6928104f",
+ "size": 167,
+ "mediaType": "application/vnd.oci.image.config.v1+json"
+ },
+ "layers": [
+ {
+ "digest": "sha256:6276d7ea2e58ab8315bab81cecfe9dab75fa7f51727ea19da12961573fd38e2a",
+ "size": 1343,
+ "mediaType": "application/vnd.in-toto+json"
+ }
+ ]
+ },
+ {
+ "digest": "sha256:09dbf8c631140aa27332ab2b3444db0b4b6fecf5e6b46ee6af0046aa5ae2b4a6",
+ "repository": "hsm-secrets-operator",
+ "shape": "helm",
+ "mediaType": "application/vnd.oci.image.manifest.v1+json",
+ "holdDid": "did:plc:wnbpdx4bn5uts5yd2o2g3gnm",
+ "config": {
+ "digest": "sha256:00d52547a21ea29e56f6672081d2aa6f20cabbdab09931aa1e9ca5076a2839b9",
+ "size": 623,
+ "mediaType": "application/vnd.cncf.helm.config.v1+json"
+ },
+ "layers": [
+ {
+ "digest": "sha256:bafc635ea0da86be0b37573e54a50f544af314513358c0a56102f8096fe4dfaa",
+ "size": 14965,
+ "mediaType": "application/vnd.cncf.helm.chart.content.v1.tar+gzip"
+ }
+ ]
+ },
+ {
+ "digest": "sha256:1754d55b3d186367b9a4f4f56082d069eef7df21f64c646509fbdd159c6d2315",
+ "repository": "hsm-secrets-operator",
+ "shape": "helm",
+ "mediaType": "application/vnd.oci.image.manifest.v1+json",
+ "holdDid": "did:plc:wnbpdx4bn5uts5yd2o2g3gnm",
+ "config": {
+ "digest": "sha256:41dcdbce442e9f9f5e45aa761e4f768384c78ceb4f0a720af092b4fe3c719e8b",
+ "size": 637,
+ "mediaType": "application/vnd.cncf.helm.config.v1+json"
+ },
+ "layers": [
+ {
+ "digest": "sha256:0cc086128079c6641e24190a3bb86cd6412c7c2d52cc979371424b3ad1484381",
+ "size": 15697,
+ "mediaType": "application/vnd.cncf.helm.chart.content.v1.tar+gzip"
+ }
+ ]
+ },
+ {
+ "digest": "sha256:68b8c226c28ee020b9499b08eae52e3e0677b2e54c283fbdfd832058834b85d9",
+ "repository": "loom",
+ "shape": "helm",
+ "mediaType": "application/vnd.oci.image.manifest.v1+json",
+ "holdDid": "did:plc:wnbpdx4bn5uts5yd2o2g3gnm",
+ "config": {
+ "digest": "sha256:840f8a7a89535ce9f75035d620eaad6e58e66397c4bddbf33116ed4d03ab4a71",
+ "size": 376,
+ "mediaType": "application/vnd.cncf.helm.config.v1+json"
+ },
+ "layers": [
+ {
+ "digest": "sha256:734447fb45c3f6356d7364adbea108a072958d55329e52c9419abeb59927fdf1",
+ "size": 13007,
+ "mediaType": "application/vnd.cncf.helm.chart.content.v1.tar+gzip"
+ }
+ ]
+ },
+ {
+ "digest": "sha256:9d97c829d17649c8f22853d73c22d8ab534a4b4c4e5eeda61b6719c268775fbe",
+ "repository": "secret-service-operator",
+ "shape": "helm",
+ "mediaType": "application/vnd.oci.image.manifest.v1+json",
+ "holdDid": "did:plc:wnbpdx4bn5uts5yd2o2g3gnm",
+ "config": {
+ "digest": "sha256:b1a5cf0ada49dd39d0c235ed9dd51028e5429cce031406d682f55e184d056da0",
+ "size": 307,
+ "mediaType": "application/vnd.cncf.helm.config.v1+json"
+ },
+ "layers": [
+ {
+ "digest": "sha256:37e8873a6eb9cbc7f91b608184e87a6b6a152f4d01f5018df5451bc375056d29",
+ "size": 3384,
+ "mediaType": "application/vnd.cncf.helm.chart.content.v1.tar+gzip"
+ }
+ ]
+ },
+ {
+ "digest": "sha256:0d4b92ead12e0e947fab489e08bae670e7e962aca6e25dd95774a0d1ef4c4a7a",
+ "repository": "agent-gateway",
+ "shape": "image",
+ "mediaType": "application/vnd.docker.distribution.manifest.v2+json",
+ "holdDid": "did:plc:wnbpdx4bn5uts5yd2o2g3gnm",
+ "config": {
+ "digest": "sha256:02667532c2098f2dc9a6f0e4354e3bdd8ea616e007f8932c5ec4abb47c81497e",
+ "size": 8975,
+ "mediaType": "application/vnd.docker.container.image.v1+json"
+ },
+ "layers": [
+ {
+ "digest": "sha256:44e5fefee415bc2f433adc1146461230101bd94ca3fa531538da96bc690ba24b",
+ "size": 29153858,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:8a8f99d8b62deb507fbd5795068231986ca9788dc403c54373a862cbbdc1090a",
+ "size": 3315,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:3de69d675a4de6c71e2311f1e547de804e6e68f051a0b9e17e6283d500894468",
+ "size": 49937593,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:6c81c36100ac3e6b034b455564d4455ecaa854542bc46cfb4a0bbca4817a6dcd",
+ "size": 1712624,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:c08735c7f0f9602a107c9b3b786ca0e9d07d35af36c3434a8c3fd9328aa4dd16",
+ "size": 446,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:52b2bda2918e4168d872ef015424b014c2630d0b94a4a192e43b8814e40857b4",
+ "size": 100849820,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:bbaad6a27eec4a15a9a052966087bd3f72a10440961d0f7a05e9ba50d4edf5a2",
+ "size": 93,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:ce1e57fdc88d123a729ddb0f855445c891d572f0f29fd9aefe99808fa5516271",
+ "size": 369,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:1ae964bbcc3ed70500512454b42f4883a61366a9fdcc0e1db60cc3fce905a729",
+ "size": 137704370,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:738321556b121e24fea2c2391457731b564344c463f9efe3124ec57d20628f3c",
+ "size": 6093,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:a4cc5904cbd30bd252b05c595652a50e39658c4cd2b02bf3ae2256b87a5498b4",
+ "size": 5017,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ }
+ ]
+ },
+ {
+ "digest": "sha256:173b0c37b79488ce35e5f1a3e272e79d3db8a0c636c9fbe0bb45ae58166c09c0",
+ "repository": "agent-gateway",
+ "shape": "image",
+ "mediaType": "application/vnd.docker.distribution.manifest.v2+json",
+ "holdDid": "did:plc:wnbpdx4bn5uts5yd2o2g3gnm",
+ "config": {
+ "digest": "sha256:5aef6a42d752af44ca1f2ab78d64422552be7fd7d8acd2231ccdc856158bc578",
+ "size": 8975,
+ "mediaType": "application/vnd.docker.container.image.v1+json"
+ },
+ "layers": [
+ {
+ "digest": "sha256:44e5fefee415bc2f433adc1146461230101bd94ca3fa531538da96bc690ba24b",
+ "size": 29153858,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:8a8f99d8b62deb507fbd5795068231986ca9788dc403c54373a862cbbdc1090a",
+ "size": 3315,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:3de69d675a4de6c71e2311f1e547de804e6e68f051a0b9e17e6283d500894468",
+ "size": 49937593,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:6c81c36100ac3e6b034b455564d4455ecaa854542bc46cfb4a0bbca4817a6dcd",
+ "size": 1712624,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:c08735c7f0f9602a107c9b3b786ca0e9d07d35af36c3434a8c3fd9328aa4dd16",
+ "size": 446,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:52b2bda2918e4168d872ef015424b014c2630d0b94a4a192e43b8814e40857b4",
+ "size": 100849820,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:bbaad6a27eec4a15a9a052966087bd3f72a10440961d0f7a05e9ba50d4edf5a2",
+ "size": 93,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:627fedbe6dff82c1b8ef3a1ee6ba1fdbcb752301f94fae360dae1579114de7f5",
+ "size": 366,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:38da3974810198d0d9ddefc196dc05892faeaa0b6eac51aa33503f57aab39c5f",
+ "size": 137893244,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:bae781eb0f1efa134c4535c2f9668ee0fbd4aeda70cd2228698ca700d806c424",
+ "size": 7819,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:f13b703148c2a3a98fce640a73f66d8a72908cc1bd94ce837c54258660b6f872",
+ "size": 6776,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ }
+ ]
+ },
+ {
+ "digest": "sha256:1a802ec63b12dc120dac2c4e0f3594d3bbe356581ff261db753fa4b3fbab4ee5",
+ "repository": "agent-gateway",
+ "shape": "image",
+ "mediaType": "application/vnd.docker.distribution.manifest.v2+json",
+ "holdDid": "did:plc:wnbpdx4bn5uts5yd2o2g3gnm",
+ "config": {
+ "digest": "sha256:512b7233698f20bd45526b3f10a5e06f2b697cb679e063f08f1d496645ca7375",
+ "size": 8976,
+ "mediaType": "application/vnd.docker.container.image.v1+json"
+ },
+ "layers": [
+ {
+ "digest": "sha256:44e5fefee415bc2f433adc1146461230101bd94ca3fa531538da96bc690ba24b",
+ "size": 29153858,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:8a8f99d8b62deb507fbd5795068231986ca9788dc403c54373a862cbbdc1090a",
+ "size": 3315,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:3de69d675a4de6c71e2311f1e547de804e6e68f051a0b9e17e6283d500894468",
+ "size": 49937593,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:6c81c36100ac3e6b034b455564d4455ecaa854542bc46cfb4a0bbca4817a6dcd",
+ "size": 1712624,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:c08735c7f0f9602a107c9b3b786ca0e9d07d35af36c3434a8c3fd9328aa4dd16",
+ "size": 446,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:52b2bda2918e4168d872ef015424b014c2630d0b94a4a192e43b8814e40857b4",
+ "size": 100849820,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:bbaad6a27eec4a15a9a052966087bd3f72a10440961d0f7a05e9ba50d4edf5a2",
+ "size": 93,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:ce1e57fdc88d123a729ddb0f855445c891d572f0f29fd9aefe99808fa5516271",
+ "size": 369,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:1ae964bbcc3ed70500512454b42f4883a61366a9fdcc0e1db60cc3fce905a729",
+ "size": 137704370,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:54deb992c6a739047ceae773aa2fd3c5c9bc11f4ca3c80a231ac14f689ac68f7",
+ "size": 6358,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:b2c281937b75cba1e49968ce411682d3d061e63851c3c230de4b7c212093a594",
+ "size": 5180,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ }
+ ]
+ },
+ {
+ "digest": "sha256:3ef993cce3bf1200a095cd39e9b8dd2dcaa7c2bde94cda178366a725fa7ac5a3",
+ "repository": "agent-gateway",
+ "shape": "image",
+ "mediaType": "application/vnd.docker.distribution.manifest.v2+json",
+ "holdDid": "did:plc:wnbpdx4bn5uts5yd2o2g3gnm",
+ "config": {
+ "digest": "sha256:0429ca5a33bb4e254c474edfd856934a48bcc643314e394cf1ced138f21c20c7",
+ "size": 8976,
+ "mediaType": "application/vnd.docker.container.image.v1+json"
+ },
+ "layers": [
+ {
+ "digest": "sha256:44e5fefee415bc2f433adc1146461230101bd94ca3fa531538da96bc690ba24b",
+ "size": 29153858,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:8a8f99d8b62deb507fbd5795068231986ca9788dc403c54373a862cbbdc1090a",
+ "size": 3315,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:3de69d675a4de6c71e2311f1e547de804e6e68f051a0b9e17e6283d500894468",
+ "size": 49937593,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:6c81c36100ac3e6b034b455564d4455ecaa854542bc46cfb4a0bbca4817a6dcd",
+ "size": 1712624,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:c08735c7f0f9602a107c9b3b786ca0e9d07d35af36c3434a8c3fd9328aa4dd16",
+ "size": 446,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:52b2bda2918e4168d872ef015424b014c2630d0b94a4a192e43b8814e40857b4",
+ "size": 100849820,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:bbaad6a27eec4a15a9a052966087bd3f72a10440961d0f7a05e9ba50d4edf5a2",
+ "size": 93,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:ce1e57fdc88d123a729ddb0f855445c891d572f0f29fd9aefe99808fa5516271",
+ "size": 369,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:1ae964bbcc3ed70500512454b42f4883a61366a9fdcc0e1db60cc3fce905a729",
+ "size": 137704370,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:2441654820ecd67553d766aaeafa0108e2cc6d04e995be7edf4b3d35e802e314",
+ "size": 5238,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:6d6829f29a1fc45ff97ee239c7c9fb36a8db267e9452fa3e4349843de0d2428d",
+ "size": 4725,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ }
+ ]
+ },
+ {
+ "digest": "sha256:41dcfbabd0f7218f05801be849569262f5a0fa5b824c9e38e7732eab942c2dfb",
+ "repository": "agent-gateway",
+ "shape": "image",
+ "mediaType": "application/vnd.docker.distribution.manifest.v2+json",
+ "holdDid": "did:plc:wnbpdx4bn5uts5yd2o2g3gnm",
+ "config": {
+ "digest": "sha256:44566249369145770d43c2191ad7a3a657b3888f3043135062428bf7339df6fd",
+ "size": 8975,
+ "mediaType": "application/vnd.docker.container.image.v1+json"
+ },
+ "layers": [
+ {
+ "digest": "sha256:44e5fefee415bc2f433adc1146461230101bd94ca3fa531538da96bc690ba24b",
+ "size": 29153858,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:8a8f99d8b62deb507fbd5795068231986ca9788dc403c54373a862cbbdc1090a",
+ "size": 3315,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:3de69d675a4de6c71e2311f1e547de804e6e68f051a0b9e17e6283d500894468",
+ "size": 49937593,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:6c81c36100ac3e6b034b455564d4455ecaa854542bc46cfb4a0bbca4817a6dcd",
+ "size": 1712624,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:c08735c7f0f9602a107c9b3b786ca0e9d07d35af36c3434a8c3fd9328aa4dd16",
+ "size": 446,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:52b2bda2918e4168d872ef015424b014c2630d0b94a4a192e43b8814e40857b4",
+ "size": 100849820,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:bbaad6a27eec4a15a9a052966087bd3f72a10440961d0f7a05e9ba50d4edf5a2",
+ "size": 93,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:ce1e57fdc88d123a729ddb0f855445c891d572f0f29fd9aefe99808fa5516271",
+ "size": 369,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:1ae964bbcc3ed70500512454b42f4883a61366a9fdcc0e1db60cc3fce905a729",
+ "size": 137704370,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:7b1424801ac912ff8781c3d56e172405b7802ad67834621f2606dee099723e7d",
+ "size": 4283,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:4718603452b48fe427e38d818d0e777ece7e9aee9a264719cf8c6436900b1988",
+ "size": 4361,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ }
+ ]
+ },
+ {
+ "digest": "sha256:653c3d7388d9d2f517ec4aac1a767568ec12bc0a1a2c29a0c05829acff203def",
+ "repository": "agent-gateway",
+ "shape": "image",
+ "mediaType": "application/vnd.docker.distribution.manifest.v2+json",
+ "holdDid": "did:plc:wnbpdx4bn5uts5yd2o2g3gnm",
+ "config": {
+ "digest": "sha256:4afd624a1d7841a59cfc1519a285fefe0b9cb22574715f8c21a564e396421909",
+ "size": 8976,
+ "mediaType": "application/vnd.docker.container.image.v1+json"
+ },
+ "layers": [
+ {
+ "digest": "sha256:68629629b516c3cd6f5e71ffbe18e32afb1ae5b4926c92d058c0f11ef1fd58a3",
+ "size": 28237639,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:8a8f99d8b62deb507fbd5795068231986ca9788dc403c54373a862cbbdc1090a",
+ "size": 3315,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:3de69d675a4de6c71e2311f1e547de804e6e68f051a0b9e17e6283d500894468",
+ "size": 49937593,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:6c81c36100ac3e6b034b455564d4455ecaa854542bc46cfb4a0bbca4817a6dcd",
+ "size": 1712624,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:c08735c7f0f9602a107c9b3b786ca0e9d07d35af36c3434a8c3fd9328aa4dd16",
+ "size": 446,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:52b2bda2918e4168d872ef015424b014c2630d0b94a4a192e43b8814e40857b4",
+ "size": 100849820,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:bbaad6a27eec4a15a9a052966087bd3f72a10440961d0f7a05e9ba50d4edf5a2",
+ "size": 93,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:7824c5f3ce3ef90bd42a5b740a7519e3c028023f6aa5fec4ce013a6f66791db6",
+ "size": 366,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:12a01f5225cd0df58771d43e42f62923f278fbef3f5ae92f228f3fc6db5165f4",
+ "size": 137893889,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:3749b1d8baff8da33a7cec8fb2609828f1e3f59146025b03e0eed0021d7907b8",
+ "size": 9293,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:8e6cce16189662fa83a30fe2bbfa7690571b744850aff363f6b7ef1b7e6871df",
+ "size": 7619,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ }
+ ]
+ },
+ {
+ "digest": "sha256:678d8ba99c186dbbfd73c18878a6cd7503aa4da38a893d295f30b06950cf92d8",
+ "repository": "agent-gateway",
+ "shape": "image",
+ "mediaType": "application/vnd.docker.distribution.manifest.v2+json",
+ "holdDid": "did:plc:wnbpdx4bn5uts5yd2o2g3gnm",
+ "config": {
+ "digest": "sha256:3348e58f41e2c96647e6120552c7706448e982c4fd65488df087f352c1515e04",
+ "size": 8975,
+ "mediaType": "application/vnd.docker.container.image.v1+json"
+ },
+ "layers": [
+ {
+ "digest": "sha256:44e5fefee415bc2f433adc1146461230101bd94ca3fa531538da96bc690ba24b",
+ "size": 29153858,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:8a8f99d8b62deb507fbd5795068231986ca9788dc403c54373a862cbbdc1090a",
+ "size": 3315,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:3de69d675a4de6c71e2311f1e547de804e6e68f051a0b9e17e6283d500894468",
+ "size": 49937593,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:6c81c36100ac3e6b034b455564d4455ecaa854542bc46cfb4a0bbca4817a6dcd",
+ "size": 1712624,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:c08735c7f0f9602a107c9b3b786ca0e9d07d35af36c3434a8c3fd9328aa4dd16",
+ "size": 446,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:52b2bda2918e4168d872ef015424b014c2630d0b94a4a192e43b8814e40857b4",
+ "size": 100849820,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:bbaad6a27eec4a15a9a052966087bd3f72a10440961d0f7a05e9ba50d4edf5a2",
+ "size": 93,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:ce1e57fdc88d123a729ddb0f855445c891d572f0f29fd9aefe99808fa5516271",
+ "size": 369,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:1ae964bbcc3ed70500512454b42f4883a61366a9fdcc0e1db60cc3fce905a729",
+ "size": 137704370,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:738321556b121e24fea2c2391457731b564344c463f9efe3124ec57d20628f3c",
+ "size": 6093,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:be5d3b62415144c7fbd77fe3f94095cbb3d3059d8bd041c9b128f30a69f55ef3",
+ "size": 5069,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ }
+ ]
+ },
+ {
+ "digest": "sha256:7b80999e5541da08df00ce1376eabf7617a6ce79f3b505bd5c5b99bf6605fe49",
+ "repository": "agent-gateway",
+ "shape": "image",
+ "mediaType": "application/vnd.docker.distribution.manifest.v2+json",
+ "holdDid": "did:plc:wnbpdx4bn5uts5yd2o2g3gnm",
+ "config": {
+ "digest": "sha256:a1dc571ddc6a9674bddbfcf716240ad413ac0ccd9e7183855e5a850271fdf64f",
+ "size": 8976,
+ "mediaType": "application/vnd.docker.container.image.v1+json"
+ },
+ "layers": [
+ {
+ "digest": "sha256:44e5fefee415bc2f433adc1146461230101bd94ca3fa531538da96bc690ba24b",
+ "size": 29153858,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:8a8f99d8b62deb507fbd5795068231986ca9788dc403c54373a862cbbdc1090a",
+ "size": 3315,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:3de69d675a4de6c71e2311f1e547de804e6e68f051a0b9e17e6283d500894468",
+ "size": 49937593,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:6c81c36100ac3e6b034b455564d4455ecaa854542bc46cfb4a0bbca4817a6dcd",
+ "size": 1712624,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:c08735c7f0f9602a107c9b3b786ca0e9d07d35af36c3434a8c3fd9328aa4dd16",
+ "size": 446,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:52b2bda2918e4168d872ef015424b014c2630d0b94a4a192e43b8814e40857b4",
+ "size": 100849820,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:bbaad6a27eec4a15a9a052966087bd3f72a10440961d0f7a05e9ba50d4edf5a2",
+ "size": 93,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:ce1e57fdc88d123a729ddb0f855445c891d572f0f29fd9aefe99808fa5516271",
+ "size": 369,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:1ae964bbcc3ed70500512454b42f4883a61366a9fdcc0e1db60cc3fce905a729",
+ "size": 137704370,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:a9a8bfe62cc9f1aeed29968a2fec92b674973cda07ff4b085f6585a74253e9be",
+ "size": 4132,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:526b7557f0f54dd3d7de5122d275ad9a8cedf6e8604e3609948f43df6225b87d",
+ "size": 3505,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ }
+ ]
+ },
+ {
+ "digest": "sha256:8e6ddaac223313f74bb108774ba96f3b65360c87ddf81154946afa2f236ffcee",
+ "repository": "agent-gateway",
+ "shape": "image",
+ "mediaType": "application/vnd.docker.distribution.manifest.v2+json",
+ "holdDid": "did:plc:wnbpdx4bn5uts5yd2o2g3gnm",
+ "config": {
+ "digest": "sha256:4b6aad6c670b8b43c33b0c5ffde703f8629ba97b7b00cc689556c4543605942c",
+ "size": 8976,
+ "mediaType": "application/vnd.docker.container.image.v1+json"
+ },
+ "layers": [
+ {
+ "digest": "sha256:44e5fefee415bc2f433adc1146461230101bd94ca3fa531538da96bc690ba24b",
+ "size": 29153858,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:8a8f99d8b62deb507fbd5795068231986ca9788dc403c54373a862cbbdc1090a",
+ "size": 3315,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:3de69d675a4de6c71e2311f1e547de804e6e68f051a0b9e17e6283d500894468",
+ "size": 49937593,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:6c81c36100ac3e6b034b455564d4455ecaa854542bc46cfb4a0bbca4817a6dcd",
+ "size": 1712624,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:c08735c7f0f9602a107c9b3b786ca0e9d07d35af36c3434a8c3fd9328aa4dd16",
+ "size": 446,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:52b2bda2918e4168d872ef015424b014c2630d0b94a4a192e43b8814e40857b4",
+ "size": 100849820,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:bbaad6a27eec4a15a9a052966087bd3f72a10440961d0f7a05e9ba50d4edf5a2",
+ "size": 93,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:185d760be57e7df097a4ca22a949e1dc8df2cc2145a674bcaa20f8390580912d",
+ "size": 362,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:faf9314e291f098bf859e323efb6bb94bcaab1f320e8783a94153e99d2230f45",
+ "size": 137770650,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:b54e0de10f3fe56056814ed26480056d0ae7866402e00cf9bf4fb3cb78999cf3",
+ "size": 7820,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:cbfb4efb85c0249f5e4dc10952029aa62fce1c807814ad7cf26703bc053b746b",
+ "size": 6782,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ }
+ ]
+ },
+ {
+ "digest": "sha256:96ae69bc80b4ef6df271e232b1152ecba27d1bcd0e5330c8b3e23a6885123337",
+ "repository": "agent-gateway",
+ "shape": "image",
+ "mediaType": "application/vnd.docker.distribution.manifest.v2+json",
+ "holdDid": "did:plc:wnbpdx4bn5uts5yd2o2g3gnm",
+ "config": {
+ "digest": "sha256:cce50edaf2ab81fd8daac1eb56c706a28b1f2eefe26363eb456cab7d5e73b7bd",
+ "size": 8976,
+ "mediaType": "application/vnd.docker.container.image.v1+json"
+ },
+ "layers": [
+ {
+ "digest": "sha256:44e5fefee415bc2f433adc1146461230101bd94ca3fa531538da96bc690ba24b",
+ "size": 29153858,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:8a8f99d8b62deb507fbd5795068231986ca9788dc403c54373a862cbbdc1090a",
+ "size": 3315,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:3de69d675a4de6c71e2311f1e547de804e6e68f051a0b9e17e6283d500894468",
+ "size": 49937593,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:6c81c36100ac3e6b034b455564d4455ecaa854542bc46cfb4a0bbca4817a6dcd",
+ "size": 1712624,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:c08735c7f0f9602a107c9b3b786ca0e9d07d35af36c3434a8c3fd9328aa4dd16",
+ "size": 446,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:52b2bda2918e4168d872ef015424b014c2630d0b94a4a192e43b8814e40857b4",
+ "size": 100849820,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:bbaad6a27eec4a15a9a052966087bd3f72a10440961d0f7a05e9ba50d4edf5a2",
+ "size": 93,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:ce1e57fdc88d123a729ddb0f855445c891d572f0f29fd9aefe99808fa5516271",
+ "size": 369,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:1ae964bbcc3ed70500512454b42f4883a61366a9fdcc0e1db60cc3fce905a729",
+ "size": 137704370,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:a9a8bfe62cc9f1aeed29968a2fec92b674973cda07ff4b085f6585a74253e9be",
+ "size": 4132,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:1d68a9d957a5ebbcc41ee1d12dee8f021f6347daf93581af8203b00d945dfd72",
+ "size": 2869,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ }
+ ]
+ },
+ {
+ "digest": "sha256:ab640024204c8c6b5c6505b81f4beb4a3ae05246afcaae02fb85d91b1b05cb03",
+ "repository": "agent-gateway",
+ "shape": "image",
+ "mediaType": "application/vnd.docker.distribution.manifest.v2+json",
+ "holdDid": "did:plc:wnbpdx4bn5uts5yd2o2g3gnm",
+ "config": {
+ "digest": "sha256:e8c580c4f80a526d820b0c8337d67bb8a944d5ed8c1a95f554ab3383547ceb0e",
+ "size": 8976,
+ "mediaType": "application/vnd.docker.container.image.v1+json"
+ },
+ "layers": [
+ {
+ "digest": "sha256:44e5fefee415bc2f433adc1146461230101bd94ca3fa531538da96bc690ba24b",
+ "size": 29153858,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:8a8f99d8b62deb507fbd5795068231986ca9788dc403c54373a862cbbdc1090a",
+ "size": 3315,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:3de69d675a4de6c71e2311f1e547de804e6e68f051a0b9e17e6283d500894468",
+ "size": 49937593,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:6c81c36100ac3e6b034b455564d4455ecaa854542bc46cfb4a0bbca4817a6dcd",
+ "size": 1712624,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:c08735c7f0f9602a107c9b3b786ca0e9d07d35af36c3434a8c3fd9328aa4dd16",
+ "size": 446,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:52b2bda2918e4168d872ef015424b014c2630d0b94a4a192e43b8814e40857b4",
+ "size": 100849820,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:bbaad6a27eec4a15a9a052966087bd3f72a10440961d0f7a05e9ba50d4edf5a2",
+ "size": 93,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:ce1e57fdc88d123a729ddb0f855445c891d572f0f29fd9aefe99808fa5516271",
+ "size": 369,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:1ae964bbcc3ed70500512454b42f4883a61366a9fdcc0e1db60cc3fce905a729",
+ "size": 137704370,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:730991d1bbe984fc7bcb152d7d9b73ef4d9ef463024c27f498db2dd2754b808c",
+ "size": 5075,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:11418cd30daf9cd9f26c5fca0baa2361f939f14b12780dfd9d520cfa65ea692b",
+ "size": 4723,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ }
+ ]
+ },
+ {
+ "digest": "sha256:b4d59c08e5bdd5321f78c16221f864d4cb183a3d236d33cd932f4656443cceb3",
+ "repository": "agent-gateway",
+ "shape": "image",
+ "mediaType": "application/vnd.docker.distribution.manifest.v2+json",
+ "holdDid": "did:plc:wnbpdx4bn5uts5yd2o2g3gnm",
+ "config": {
+ "digest": "sha256:930ad3755c9f8d61bc1e3584aecfb92efb8fab0891f9be065f8df0c7f72cbf83",
+ "size": 8976,
+ "mediaType": "application/vnd.docker.container.image.v1+json"
+ },
+ "layers": [
+ {
+ "digest": "sha256:44e5fefee415bc2f433adc1146461230101bd94ca3fa531538da96bc690ba24b",
+ "size": 29153858,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:8a8f99d8b62deb507fbd5795068231986ca9788dc403c54373a862cbbdc1090a",
+ "size": 3315,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:3de69d675a4de6c71e2311f1e547de804e6e68f051a0b9e17e6283d500894468",
+ "size": 49937593,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:6c81c36100ac3e6b034b455564d4455ecaa854542bc46cfb4a0bbca4817a6dcd",
+ "size": 1712624,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:c08735c7f0f9602a107c9b3b786ca0e9d07d35af36c3434a8c3fd9328aa4dd16",
+ "size": 446,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:52b2bda2918e4168d872ef015424b014c2630d0b94a4a192e43b8814e40857b4",
+ "size": 100849820,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:bbaad6a27eec4a15a9a052966087bd3f72a10440961d0f7a05e9ba50d4edf5a2",
+ "size": 93,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:ce1e57fdc88d123a729ddb0f855445c891d572f0f29fd9aefe99808fa5516271",
+ "size": 369,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:1ae964bbcc3ed70500512454b42f4883a61366a9fdcc0e1db60cc3fce905a729",
+ "size": 137704370,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:ec850d9bdc0252bb44a93d0c95e7abc0444ff6e9923135fd3811a36c9e29ea1a",
+ "size": 4556,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:2e3fd791268a48c79914947bd2e402e093c45cf9c0edceae275f9712d95f1d94",
+ "size": 4366,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ }
+ ]
+ },
+ {
+ "digest": "sha256:d17f9084e2b3d924a5aa8859876cd9b40d9467c19f1b6b86a82314f920e2e52a",
+ "repository": "agent-gateway",
+ "shape": "image",
+ "mediaType": "application/vnd.docker.distribution.manifest.v2+json",
+ "holdDid": "did:plc:wnbpdx4bn5uts5yd2o2g3gnm",
+ "config": {
+ "digest": "sha256:8e5446ac8bc8e44c61a2f6620ec65c7f1f2041b6383fdc22459738c179f37fe1",
+ "size": 8988,
+ "mediaType": "application/vnd.docker.container.image.v1+json"
+ },
+ "layers": [
+ {
+ "digest": "sha256:68629629b516c3cd6f5e71ffbe18e32afb1ae5b4926c92d058c0f11ef1fd58a3",
+ "size": 28237639,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:8a8f99d8b62deb507fbd5795068231986ca9788dc403c54373a862cbbdc1090a",
+ "size": 3315,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:3de69d675a4de6c71e2311f1e547de804e6e68f051a0b9e17e6283d500894468",
+ "size": 49937593,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:6c81c36100ac3e6b034b455564d4455ecaa854542bc46cfb4a0bbca4817a6dcd",
+ "size": 1712624,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:c08735c7f0f9602a107c9b3b786ca0e9d07d35af36c3434a8c3fd9328aa4dd16",
+ "size": 446,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:52b2bda2918e4168d872ef015424b014c2630d0b94a4a192e43b8814e40857b4",
+ "size": 100849820,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:bbaad6a27eec4a15a9a052966087bd3f72a10440961d0f7a05e9ba50d4edf5a2",
+ "size": 93,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:7824c5f3ce3ef90bd42a5b740a7519e3c028023f6aa5fec4ce013a6f66791db6",
+ "size": 366,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:12a01f5225cd0df58771d43e42f62923f278fbef3f5ae92f228f3fc6db5165f4",
+ "size": 137893889,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:be9bf5f507988a353ec57ae6822f5effd3d3059c9bc9b7ee88fb50f4ea25598a",
+ "size": 16654,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:771ece482f46e36fb837a3006f5812219e17705e7d7fa1a75251a635fdf574a3",
+ "size": 7623,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ }
+ ]
+ },
+ {
+ "digest": "sha256:d7fc62b44fa8c5447acfff9eee7d57442943e731c6c192046104dfe412a1346f",
+ "repository": "agent-gateway",
+ "shape": "image",
+ "mediaType": "application/vnd.docker.distribution.manifest.v2+json",
+ "holdDid": "did:plc:wnbpdx4bn5uts5yd2o2g3gnm",
+ "config": {
+ "digest": "sha256:a5ed1092845cf38aeb889c7b45fc7501eb42d674592631e39830bcb6dd6a5e08",
+ "size": 8988,
+ "mediaType": "application/vnd.docker.container.image.v1+json"
+ },
+ "layers": [
+ {
+ "digest": "sha256:68629629b516c3cd6f5e71ffbe18e32afb1ae5b4926c92d058c0f11ef1fd58a3",
+ "size": 28237639,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:8a8f99d8b62deb507fbd5795068231986ca9788dc403c54373a862cbbdc1090a",
+ "size": 3315,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:3de69d675a4de6c71e2311f1e547de804e6e68f051a0b9e17e6283d500894468",
+ "size": 49937593,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:6c81c36100ac3e6b034b455564d4455ecaa854542bc46cfb4a0bbca4817a6dcd",
+ "size": 1712624,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:c08735c7f0f9602a107c9b3b786ca0e9d07d35af36c3434a8c3fd9328aa4dd16",
+ "size": 446,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:52b2bda2918e4168d872ef015424b014c2630d0b94a4a192e43b8814e40857b4",
+ "size": 100849820,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:bbaad6a27eec4a15a9a052966087bd3f72a10440961d0f7a05e9ba50d4edf5a2",
+ "size": 93,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:7824c5f3ce3ef90bd42a5b740a7519e3c028023f6aa5fec4ce013a6f66791db6",
+ "size": 366,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:12a01f5225cd0df58771d43e42f62923f278fbef3f5ae92f228f3fc6db5165f4",
+ "size": 137893889,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:73a15447e3df7c65743d5352f2fcf18f93f8bfdf3fa4401df9c2f8c64222bac3",
+ "size": 16184,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:e37386ac763e406358522b5c1ca4575460be93b2513ed6db9152dbe0c870785a",
+ "size": 7623,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ }
+ ]
+ },
+ {
+ "digest": "sha256:741c6d92ad43532b92ddc184edca78578ea6597f9092b93e7e003214c0f24ec3",
+ "repository": "atcr-appview",
+ "shape": "image",
+ "mediaType": "application/vnd.docker.distribution.manifest.v2+json",
+ "holdDid": "did:web:172.28.0.3:8080",
+ "config": {
+ "digest": "sha256:4734bc89340dbd9cbd26faa373cc3a3ade7f3fb16655ed5a863ffd499d5a8af2",
+ "size": 3122,
+ "mediaType": "application/vnd.docker.container.image.v1+json"
+ },
+ "layers": [
+ {
+ "digest": "sha256:14e8661c5a9fe87ebecf1fbf68a9469323e390f2d2bfae0677f2f95676e079ef",
+ "size": 129498,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:19eaacb95d7b5b032648c7771d0a2e6fe1c47c8596caad3de115ea24531b297f",
+ "size": 188282,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:62b733b2937f1d896c70bb99241cf02d74d9d2a337f104d7830fc7afbe94ea6d",
+ "size": 14558626,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ }
+ ]
+ },
+ {
+ "digest": "sha256:a13463f975e38f8953ccaf9215b7473f48771c956b3ecb189a78cb63b6b1cf77",
+ "repository": "atcr-appview-dev",
+ "shape": "image",
+ "mediaType": "application/vnd.docker.distribution.manifest.v2+json",
+ "holdDid": "did:plc:wnbpdx4bn5uts5yd2o2g3gnm",
+ "config": {
+ "digest": "sha256:dc8d83f79923e0d8469362089248e095c837ce5b554066df52cc43ffe4d40c68",
+ "size": 4919,
+ "mediaType": "application/vnd.docker.container.image.v1+json"
+ },
+ "layers": [
+ {
+ "digest": "sha256:f2520f145cbbf7dc787dd7a5d1cf7c93b03a027ab39b9465f25718cdeed75deb",
+ "size": 50475329,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:954d6059ca7bdbb9ceb566ca2239e01ef312165659d656753d7dbace7771a591",
+ "size": 25614010,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:b5e2021c4c8bd1a46b34d9608a9381afdc333600ee1ef3c94306ecf7373e1956",
+ "size": 67787365,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:d3b37a8f89c93c2b56d4f2cea38e4d53f9d7f8d10f0241ae1a526dab622bdf05",
+ "size": 102138701,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:0fdf71b47847e47b44531d019e8eed7d243fd7189fe6b14cf6754724b04fbdd6",
+ "size": 60156973,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:cac7b480234dda467a567c0327f68728ccf10a679b4f797393e9b61161862427",
+ "size": 126,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:4f4fb700ef54461cfa02571ae0db9a0dc1e0cdb5577484a6d75e68dc38e8acc1",
+ "size": 32,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:0567bf1f76912a7c840c06d73e635d1e5651583b8334c0139db9bcc162c16084",
+ "size": 148312006,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:40d0c0fd7f0f9b51ea26e55556c9cf7c5c3a948f7d70fda1736bdb5d0d35e9ee",
+ "size": 93,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:a40001a35dc61defa43dcd6b0dbdaf75a63d21ec0ac70b07e5e1e804c6e4ad44",
+ "size": 28347,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:557840e1c03b2de1c4f6291f8c2215b86d413727a53759a642833eacddf96e98",
+ "size": 341972213,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ }
+ ]
+ },
+ {
+ "digest": "sha256:12aee1facf9b326c04983d2592136da37d386df7558b314a5561d6c39418dd00",
+ "repository": "atcr-hold-dev",
+ "shape": "image",
+ "mediaType": "application/vnd.docker.distribution.manifest.v2+json",
+ "holdDid": "did:plc:wnbpdx4bn5uts5yd2o2g3gnm",
+ "config": {
+ "digest": "sha256:c2afe0d3cfcea4d38baf2ff47947822e57fc573f4289c3f727a45c0679c3f41d",
+ "size": 4938,
+ "mediaType": "application/vnd.docker.container.image.v1+json"
+ },
+ "layers": [
+ {
+ "digest": "sha256:f2520f145cbbf7dc787dd7a5d1cf7c93b03a027ab39b9465f25718cdeed75deb",
+ "size": 50475329,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:954d6059ca7bdbb9ceb566ca2239e01ef312165659d656753d7dbace7771a591",
+ "size": 25614010,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:b5e2021c4c8bd1a46b34d9608a9381afdc333600ee1ef3c94306ecf7373e1956",
+ "size": 67787365,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:d3b37a8f89c93c2b56d4f2cea38e4d53f9d7f8d10f0241ae1a526dab622bdf05",
+ "size": 102138701,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:0fdf71b47847e47b44531d019e8eed7d243fd7189fe6b14cf6754724b04fbdd6",
+ "size": 60156973,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:cac7b480234dda467a567c0327f68728ccf10a679b4f797393e9b61161862427",
+ "size": 126,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:4f4fb700ef54461cfa02571ae0db9a0dc1e0cdb5577484a6d75e68dc38e8acc1",
+ "size": 32,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:5959bee8b8de7b988dec6678beede57a71b5e1d4c32de2c623083be4376eabbd",
+ "size": 148310727,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:40d0c0fd7f0f9b51ea26e55556c9cf7c5c3a948f7d70fda1736bdb5d0d35e9ee",
+ "size": 93,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:a40001a35dc61defa43dcd6b0dbdaf75a63d21ec0ac70b07e5e1e804c6e4ad44",
+ "size": 28347,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:b2dc34fc757edd6e9a4fd3a83c9e4388657fa3a050219d6f69afb328d8051cce",
+ "size": 341971807,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ }
+ ]
+ },
+ {
+ "digest": "sha256:e6b9be1d13debaf5e8536c9e6ce1b6e2f6f4e41fa6d1b7bde7b7fb9382e812bf",
+ "repository": "atcr-hold-dev",
+ "shape": "image",
+ "mediaType": "application/vnd.docker.distribution.manifest.v2+json",
+ "holdDid": "did:plc:wnbpdx4bn5uts5yd2o2g3gnm",
+ "config": {
+ "digest": "sha256:1f7ceb43a6b0b5c56729a570000ca115a0e782514f1067dd87a58e5c058847e7",
+ "size": 4938,
+ "mediaType": "application/vnd.docker.container.image.v1+json"
+ },
+ "layers": [
+ {
+ "digest": "sha256:f2520f145cbbf7dc787dd7a5d1cf7c93b03a027ab39b9465f25718cdeed75deb",
+ "size": 50475329,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:954d6059ca7bdbb9ceb566ca2239e01ef312165659d656753d7dbace7771a591",
+ "size": 25614010,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:b5e2021c4c8bd1a46b34d9608a9381afdc333600ee1ef3c94306ecf7373e1956",
+ "size": 67787365,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:d3b37a8f89c93c2b56d4f2cea38e4d53f9d7f8d10f0241ae1a526dab622bdf05",
+ "size": 102138701,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:0fdf71b47847e47b44531d019e8eed7d243fd7189fe6b14cf6754724b04fbdd6",
+ "size": 60156973,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:cac7b480234dda467a567c0327f68728ccf10a679b4f797393e9b61161862427",
+ "size": 126,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:4f4fb700ef54461cfa02571ae0db9a0dc1e0cdb5577484a6d75e68dc38e8acc1",
+ "size": 32,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:5959bee8b8de7b988dec6678beede57a71b5e1d4c32de2c623083be4376eabbd",
+ "size": 148310727,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:40d0c0fd7f0f9b51ea26e55556c9cf7c5c3a948f7d70fda1736bdb5d0d35e9ee",
+ "size": 93,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:bf519c963be918f34f5fc91d014c6ba990cd6b6040b8f9a1b3c2ed06c172fb70",
+ "size": 28480,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:0990cfa2243cbd2e67992f18a5f23142583349f378ffb07e0bfa961cd353ceb8",
+ "size": 342658113,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ }
+ ]
+ },
+ {
+ "digest": "sha256:b6df625b4f311d07d5720ccf3802877eb2a9613b45d573b3f43be237386e4324",
+ "repository": "git-summarizer",
+ "shape": "image",
+ "mediaType": "application/vnd.oci.image.manifest.v1+json",
+ "holdDid": "did:plc:wnbpdx4bn5uts5yd2o2g3gnm",
+ "config": {
+ "digest": "sha256:ffeaf5e5d4f3abc139964de8308c072af2dd612b2f24743c668ea70e7752c294",
+ "size": 2612,
+ "mediaType": "application/vnd.oci.image.config.v1+json"
+ },
+ "layers": [
+ {
+ "digest": "sha256:c765ae84869fd59a62821873e5413a3e92e36bdc1ced8fab3520334863720a49",
+ "size": 4089377,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:d48315457c1f119ed4dfa190c360c183b9782fdad9687bc1b267f74812a6705d",
+ "size": 293622,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:8d4887f643ba213b7033b52145ab4286d68f6a3763690d724db5dae73504d042",
+ "size": 4064355,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:88f653dd60b2ebc73e0aebad1fe41f0cd91fcb08d70531d16c0345d399f07840",
+ "size": 113,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ }
+ ]
+ },
+ {
+ "digest": "sha256:f5c69d77263ef0025535dd25879fb0eef2958929738ac2968787c11e7011e3d6",
+ "repository": "git-summarizer",
+ "shape": "image",
+ "mediaType": "application/vnd.oci.image.manifest.v1+json",
+ "holdDid": "did:plc:wnbpdx4bn5uts5yd2o2g3gnm",
+ "config": {
+ "digest": "sha256:996df5d65693bab229cd14fd3993e5d65d1e0a45eac11feb622ed88d8a28f3f8",
+ "size": 2613,
+ "mediaType": "application/vnd.oci.image.config.v1+json"
+ },
+ "layers": [
+ {
+ "digest": "sha256:5311e7f182d02360a7194aa2995849bcdf04795c39a0ffdcf413eae625865970",
+ "size": 3627056,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:81d36881b292171523493598d6c70e9d37c764480a981cbc04d69a6c0a1e5e57",
+ "size": 290780,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:992098569391d8e37fc6f90dddea6c285937934ab1f68b079cda831f558aa0d2",
+ "size": 4064355,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:88f653dd60b2ebc73e0aebad1fe41f0cd91fcb08d70531d16c0345d399f07840",
+ "size": 113,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ }
+ ]
+ },
+ {
+ "digest": "sha256:6eaa5ddb7d063629b6bf5a85a37b4c44e79e083753ede99fdadc1e5d6a528cc6",
+ "repository": "go-tnyclick",
+ "shape": "image",
+ "mediaType": "application/vnd.docker.distribution.manifest.v2+json",
+ "holdDid": "did:plc:wnbpdx4bn5uts5yd2o2g3gnm",
+ "config": {
+ "digest": "sha256:e6284f76f49d5d45aadef0f83b7688bf058ab79211c02aa54ef7a4cc9df61dc7",
+ "size": 2009,
+ "mediaType": "application/vnd.docker.container.image.v1+json"
+ },
+ "layers": [
+ {
+ "digest": "sha256:b9eb00512c281ce6a04b4cf4fad72e8845f604e83030e4c9499de76bfca0becb",
+ "size": 3792402,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ }
+ ]
+ },
+ {
+ "digest": "sha256:bf0327b436432089259af16c31ce7c5206239aa2b1eeeaa9cb0e9d3380741e66",
+ "repository": "go-tnyclick",
+ "shape": "image",
+ "mediaType": "application/vnd.docker.distribution.manifest.v2+json",
+ "holdDid": "did:plc:wnbpdx4bn5uts5yd2o2g3gnm",
+ "config": {
+ "digest": "sha256:13f42406750753b53ce0f3ac74e5be2f4510cca9ce7d5a9ce6668a6ee3e73fac",
+ "size": 2009,
+ "mediaType": "application/vnd.docker.container.image.v1+json"
+ },
+ "layers": [
+ {
+ "digest": "sha256:53dfc70716b9866e748f80f912325f2f656aa86523cb603a1f911700ddec126c",
+ "size": 4086521,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ }
+ ]
+ },
+ {
+ "digest": "sha256:019996323ee44afa7e275318bd570577341a761fb3905be43c3a00dfdf258f17",
+ "repository": "hsm-secrets-operator",
+ "shape": "image",
+ "mediaType": "application/vnd.oci.image.manifest.v1+json",
+ "holdDid": "did:plc:wnbpdx4bn5uts5yd2o2g3gnm",
+ "config": {
+ "digest": "sha256:b3e7f6da12a6a4b776144117b514977ce59849a896e7cabbc6ecfd951095bb9c",
+ "size": 2281,
+ "mediaType": "application/vnd.oci.image.config.v1+json"
+ },
+ "layers": [
+ {
+ "digest": "sha256:0796e1cc2018078808565ec8304b770268210b9fa43184b002dd095059190b4f",
+ "size": 29450232,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:97ec2aa92dcaaf035c06a4faaf6349107f2f8cdfc2cd8313a18627f9dc195f45",
+ "size": 14739679,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:c6eb795b1daff476d0ee3736ec316bbbdd81547b890cff7500369abe64aa49c6",
+ "size": 156,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:04e4d700a6bda1c8a1c159f565257ded715977a8e1d04a08cd4b6754c0da25cc",
+ "size": 150,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:d96e0f0408cef8a10e1cf56207c66b8220822f5ee353ad21c584527446b1fc8d",
+ "size": 151,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:81187c0098b3c25f403e5b4b29966175afaa9c41e7efacd91a92d71360c61b02",
+ "size": 19948399,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:8dc3048a36b6934a2cfebf37af6d4cc849c83401abf0eff199d0cf044b5467b3",
+ "size": 13321805,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ }
+ ]
+ },
+ {
+ "digest": "sha256:1cfa4e2b09e127b9c4ed43578d3f3c18e7d44ea47b9ea98475c0cbe9086525f8",
+ "repository": "hsm-secrets-operator",
+ "shape": "image",
+ "mediaType": "application/vnd.oci.image.manifest.v1+json",
+ "holdDid": "did:plc:wnbpdx4bn5uts5yd2o2g3gnm",
+ "config": {
+ "digest": "sha256:c6348fa86ba0fb2108c9334f5fe913ddc6d853313e655891f133a0127c30099f",
+ "size": 459,
+ "mediaType": "application/vnd.oci.image.config.v1+json"
+ },
+ "layers": [
+ {
+ "digest": "sha256:b05093807bb0294152bb9cf86d64da722732dddaf7f8882fa1f120477dbc4db3",
+ "size": 2226327,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ }
+ ]
+ },
+ {
+ "digest": "sha256:40c29c380d7dc132fc8f8daf37560b0c967adc1c4cd829ed7adbfd7001a56016",
+ "repository": "hsm-secrets-operator",
+ "shape": "image",
+ "mediaType": "application/vnd.docker.distribution.manifest.v2+json",
+ "holdDid": "did:plc:wnbpdx4bn5uts5yd2o2g3gnm",
+ "config": {
+ "digest": "sha256:8d8cce2250a092d0c7a3a1965ca3ef315b254a144705fa848c952d1d4b4b0bf7",
+ "size": 3472,
+ "mediaType": "application/vnd.docker.container.image.v1+json"
+ },
+ "layers": [
+ {
+ "digest": "sha256:afe736d376e9c6f29a746f80fd63c04493abd337a9f1588a84a7a8dbbcb24568",
+ "size": 157,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:07a4694810ed2e5b777447eff2a0813b2e032f926089cfae3e901397998655ee",
+ "size": 150,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:7dc71da1753c4256fdeca0fdff86482f0ee4d5e4b0e05db1fc3d3ed7714bba76",
+ "size": 20829699,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:73b61e704382664d858440d61428a061157b772575ae58ef783c0f82f1909428",
+ "size": 108003,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:c4d696455b7d9f3824ef9902a1a7a9f189c628f298e25e503bbea6728ad993f3",
+ "size": 61054,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:733bce2e8f60456b2bcacb2fea80fad20315cb77fc46862df2feb3c72e89ff8e",
+ "size": 78676,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:a4239a8820e149d0f9cca9a9e45d351edc4dc665fb4ed933be3680679b6f34f9",
+ "size": 765,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:a08d35e7b2d0c83e026a5238874108537a85ae99134e70661c88e9744c104493",
+ "size": 108613,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:1176f24e4646d27e5b7fd0f9f3fb512be50ad34289bc22f6cfbd0f52b87edc94",
+ "size": 10417,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:1c7c87da293263096dbafd2b5fef13cd4048041aa8ff00134de89c5c7c200141",
+ "size": 129498,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:4614cc5a0bde544678f3a55637532faf145f966402126c115cb208a11f634d7f",
+ "size": 34278001,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ }
+ ]
+ },
+ {
+ "digest": "sha256:5924c964be6b28cda9aa7d3ade479586b9ce4bc33a9cadebb3c63322cc0ecc15",
+ "repository": "hsm-secrets-operator",
+ "shape": "image",
+ "mediaType": "application/vnd.docker.distribution.manifest.v2+json",
+ "holdDid": "did:plc:wnbpdx4bn5uts5yd2o2g3gnm",
+ "config": {
+ "digest": "sha256:ab7f6758ea744cf87175b5ed3d94423c0b97f3cc5d2591154e4b8fd883887bf5",
+ "size": 2788,
+ "mediaType": "application/vnd.docker.container.image.v1+json"
+ },
+ "layers": [
+ {
+ "digest": "sha256:ec07bd474c56bb4ee91db32b561f76d83ec03bdb43fa019a6c2158868e31a001",
+ "size": 31372207,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:7efbd2fe75316fdc489541769ed370c95835c46446c38c3537295e6001efc92a",
+ "size": 11287948,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:71048e1a9fbfd0ddef79b2e6bf186e38f3601b25558a1ba2a9b3d1054e403ac6",
+ "size": 157,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:e8e931102a5273289e017f7fd806d89ac21a011ada51c371f64cc7d2bb551098",
+ "size": 151,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:2cbf2d5b0c389b8e73a267fcf1a1873d3cb82f66dec9a1203e3096296f8987fa",
+ "size": 148,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:e9d06f66a63467e9c84135347ccb8643c838d28c303f92ce2ef2b8340ca6e102",
+ "size": 17677172,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:42f48a19a923f9f2dd6ce8cd80aba078e1e3ed548d2ccf37d5b8292fbecd8872",
+ "size": 11630118,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ }
+ ]
+ },
+ {
+ "digest": "sha256:5c6f5a0a13b944e2047768c539901d4f38e3bdec2ed1f7119750689235285f40",
+ "repository": "hsm-secrets-operator",
+ "shape": "image",
+ "mediaType": "application/vnd.docker.distribution.manifest.v2+json",
+ "holdDid": "did:plc:wnbpdx4bn5uts5yd2o2g3gnm",
+ "config": {
+ "digest": "sha256:360df0afd224dc681149d5eb74e7216ca0ee2234b3de255c8b7f98468f5f0702",
+ "size": 3472,
+ "mediaType": "application/vnd.docker.container.image.v1+json"
+ },
+ "layers": [
+ {
+ "digest": "sha256:f39702867cca53997b9cea80fa2c7f542a352555befd06bc6ad5817458fa39b3",
+ "size": 157,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:312d5fbc8554d954af8b5789e2ffc3e670a2f4ba6aafe6963da1afe3da840bb9",
+ "size": 150,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:5e597da3e66d40effd133559dddff9da6df236683a87a78c88ab287cc4b9fef2",
+ "size": 21000982,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:8f5031d272c60e73400622822964e76389cd74d8ae39836447fa1e94f3b944d9",
+ "size": 113840,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:a1112f3685b9e1f4f6bee98c33fcf3e521d5d75fc66074f47d11c18d0430d3d2",
+ "size": 62191,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:25b03f021dbb1cada326765dc7b345dc1eb7a61047a90847553b191a3578fb49",
+ "size": 78850,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:0cecdbda5a246eb8679164e482f66230634f630146f22b739411a9d838bd11f0",
+ "size": 766,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:84a94798eeeb8d22341a433324d9920d31f80dc222e64e557f43c60f7647def8",
+ "size": 112912,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:7741e1dc8b7e7db9d71a4443fe03e8517d2c8df2f7c88bc1227772089ea78dec",
+ "size": 10416,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:fa8d16efd6649e6768c4ad4ce2e1fe2b1dd3a8db3b12f978d68768f8cdf187d7",
+ "size": 129499,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:2b35bd2bf600d90b95d57233d1f311bb24b964c77b393e1e422db704d40a3035",
+ "size": 36961762,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ }
+ ]
+ },
+ {
+ "digest": "sha256:80140045ec657898f0c5f07a192c632bafd6999ff208d22cded6bbd1b93110d3",
+ "repository": "hsm-secrets-operator",
+ "shape": "image",
+ "mediaType": "application/vnd.docker.distribution.manifest.v2+json",
+ "holdDid": "did:plc:wnbpdx4bn5uts5yd2o2g3gnm",
+ "config": {
+ "digest": "sha256:83bdb57f52ae92c3e0ef463b2667aa08acf35d1e50eac0540d98c99a2dbcf988",
+ "size": 2789,
+ "mediaType": "application/vnd.docker.container.image.v1+json"
+ },
+ "layers": [
+ {
+ "digest": "sha256:47d2daa5f3238cad1d0d197d938b6cb597b5f923b201413ad2590f68f7f2836c",
+ "size": 30779985,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:376bfaca5113e470060bab0fa19696aea0cf1253fc3d0b9404fb3b4a30f054f4",
+ "size": 11471717,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:53193b1aa047a45371920171f3ca1c45e9d5fedaa6c805259eb4de9ef7bf4b63",
+ "size": 156,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:8f8a5e217023bd017d44044ba63d37a0fb69c273d4cea527baaead98709b03e4",
+ "size": 151,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:0ea5e490716ff93b13f4b3273f5c213a8aca12c0a0095783b84eeef687055c44",
+ "size": 150,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:2bf0bc7d9659bc05eb072820bf60b736c295262652311ec49e34617408693555",
+ "size": 19479203,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:0f7d4fbfe079aa23bc372166312598bfadb818f5cbd3278b64c2f164399fc5a4",
+ "size": 12933552,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ }
+ ]
+ },
+ {
+ "digest": "sha256:a48100dadb92511cdf7e818446125548276c0ed30328ab630648b292b046fb2c",
+ "repository": "hsm-secrets-operator",
+ "shape": "image",
+ "mediaType": "application/vnd.docker.distribution.manifest.v2+json",
+ "holdDid": "did:plc:wnbpdx4bn5uts5yd2o2g3gnm",
+ "config": {
+ "digest": "sha256:5ef4909cec6141fbc8b6b29f59ada64a69421cbb52862a6c9bb0a4c829e89da6",
+ "size": 2784,
+ "mediaType": "application/vnd.docker.container.image.v1+json"
+ },
+ "layers": [
+ {
+ "digest": "sha256:5ba766340b3db1900c6b6602eff2e0b604affbd43c1bba793dc8ec490335a0cb",
+ "size": 30779955,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:fbbfa6f715ce9407a78a8fe88b1796398e816ef369c778f6e44c8c13daa8c0f0",
+ "size": 11471728,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:162e00451e1de1a327f34eaff888e56b661f32d8074c0a8d77993b4588dd4387",
+ "size": 157,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:12b018e6512fd943cee2eb93386fc04873babf058d94ba58c85ff57ddef711ae",
+ "size": 151,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:7dc131491b322280c17b42c6c15cca5f5a2ab5d54b51270723efc851d52de3ce",
+ "size": 148,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:aba9be74255473e36c2b0341c689c96a8296432cff03c8b55133a27a33de564a",
+ "size": 19476009,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:6a417aad020cf866b77a33f26729dbd4ab8651fc09d79c6821dc0a61e7b2cd40",
+ "size": 12933108,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ }
+ ]
+ },
+ {
+ "digest": "sha256:c1805ebbe44cdd9ac6d284b7ef8205b901a2571ef55a8c0c60ee8fafd75dc139",
+ "repository": "hsm-secrets-operator",
+ "shape": "image",
+ "mediaType": "application/vnd.oci.image.manifest.v1+json",
+ "holdDid": "did:plc:wnbpdx4bn5uts5yd2o2g3gnm",
+ "config": {
+ "digest": "sha256:c8083573efd5178e0b8f4a09aca514944f2334b7658a77da8bdcfbe231ab3cea",
+ "size": 2283,
+ "mediaType": "application/vnd.oci.image.config.v1+json"
+ },
+ "layers": [
+ {
+ "digest": "sha256:7561b06b57afccbfb4f85f302be5d27615f848aacaaa6e35503fbb3363db439d",
+ "size": 29463745,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:2ce796edcb8eae467fbc905198a06e863d02ab23907c5186c52e776590ab4cb8",
+ "size": 14151578,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:cdeebb3923fb75e6e9a079a09d3f002dfa3c1fe0898c26a67e32eb51dcf96720",
+ "size": 156,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:e5f4b1537631605512acf90c89576243da2f424d41665f73c85efaa0b4fa8073",
+ "size": 150,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:1c316ce18122c9f412df2b00f16f9d2002be3551c6e75be29c52fcbb82a241c2",
+ "size": 149,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:30326deb94110e15979fdcbb802ef3209633fbd8e730f896f35aa3929b222c6b",
+ "size": 17825172,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:a8edd23161dc3e28a69215b1f741a052d32ce1f2d3af2d15e3cf9b527f57cc37",
+ "size": 11796674,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ }
+ ]
+ },
+ {
+ "digest": "sha256:e4ac9574659da87cbc42ea2ecbce36b8a8c5536eeb48565cc63e9e075580d082",
+ "repository": "hsm-secrets-operator",
+ "shape": "image",
+ "mediaType": "application/vnd.docker.distribution.manifest.v2+json",
+ "holdDid": "did:plc:wnbpdx4bn5uts5yd2o2g3gnm",
+ "config": {
+ "digest": "sha256:f360963679f3727aad839e9ce51a711dfd43372ab859431b092fe1e0060d1f8e",
+ "size": 2790,
+ "mediaType": "application/vnd.docker.container.image.v1+json"
+ },
+ "layers": [
+ {
+ "digest": "sha256:dbdbb61f7502917ae218891ceab1add371ebd16d338badc0714798ddac32fc79",
+ "size": 31372225,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:0bb2e19027478f3e8c12eb17ca5c707c76bded62fee8cd070561a4ef536c0cee",
+ "size": 11287980,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:8abdca0e050416f19bec6319653177796bf0ae48c1ecc1751668ab4de1fac4d9",
+ "size": 158,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:06b23b741f25369d57be07665f359c44d408baf33687e938ba23a0f7de3ff96a",
+ "size": 150,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:d3c4b6bc19005a4830f636a2d869f38278a91a52f6584eb9fe0b6d84d6a7c19e",
+ "size": 149,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:0fa56467a3c796265d7e726b98f9c6eed28da2e3ac993fce1c96ef6fcdfb1d27",
+ "size": 17684365,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:56b9161a85499f5a0562c28400e0dd7cdd660fca131762ec11be2f2beb3bc546",
+ "size": 11629033,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ }
+ ]
+ },
+ {
+ "digest": "sha256:24248cbc10e181c6f2ca27de0e445ff7eac5087f4eb3af432a676219896a2d97",
+ "repository": "letta-code",
+ "shape": "image",
+ "mediaType": "application/vnd.docker.distribution.manifest.v2+json",
+ "holdDid": "did:plc:wnbpdx4bn5uts5yd2o2g3gnm",
+ "config": {
+ "digest": "sha256:147b1077a7c6d5de5727299c39ae425e26255e382ea43eecbe3d3a207b54ccaa",
+ "size": 9471,
+ "mediaType": "application/vnd.docker.container.image.v1+json"
+ },
+ "layers": [
+ {
+ "digest": "sha256:44e5fefee415bc2f433adc1146461230101bd94ca3fa531538da96bc690ba24b",
+ "size": 29153858,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:8a8f99d8b62deb507fbd5795068231986ca9788dc403c54373a862cbbdc1090a",
+ "size": 3315,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:3de69d675a4de6c71e2311f1e547de804e6e68f051a0b9e17e6283d500894468",
+ "size": 49937593,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:6c81c36100ac3e6b034b455564d4455ecaa854542bc46cfb4a0bbca4817a6dcd",
+ "size": 1712624,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:c08735c7f0f9602a107c9b3b786ca0e9d07d35af36c3434a8c3fd9328aa4dd16",
+ "size": 446,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:bbd21183dc1d80f5591a989805c2ba4ec8f7ceacf2c2f52b6657f3acb3d4e9d2",
+ "size": 147965429,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:951da7c2fd1989bf9f261e5c26390ba4c74d1c97260646a844a9853cba3d7aa4",
+ "size": 137379306,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:ec78f8d0cf270bb9da2dee815c0e9c1460333f35cd5d32503ffeac061d4dd83e",
+ "size": 14356832,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:32a4ac5d734b5d361c2d7997afdfb34236cdb1a19048fa132a4553c22d2831ef",
+ "size": 93,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ }
+ ]
+ },
+ {
+ "digest": "sha256:24b684f43389f11385b132cb532e68a74cb211166825331af96b1320819a4d7b",
+ "repository": "letta-code",
+ "shape": "image",
+ "mediaType": "application/vnd.docker.distribution.manifest.v2+json",
+ "holdDid": "did:plc:wnbpdx4bn5uts5yd2o2g3gnm",
+ "config": {
+ "digest": "sha256:6dd1da2ffbacc3e65285517b0d5a8c96490b7f58ef08be6c9ae2a52db968b57a",
+ "size": 8654,
+ "mediaType": "application/vnd.docker.container.image.v1+json"
+ },
+ "layers": [
+ {
+ "digest": "sha256:44e5fefee415bc2f433adc1146461230101bd94ca3fa531538da96bc690ba24b",
+ "size": 29153858,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:8a8f99d8b62deb507fbd5795068231986ca9788dc403c54373a862cbbdc1090a",
+ "size": 3315,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:3de69d675a4de6c71e2311f1e547de804e6e68f051a0b9e17e6283d500894468",
+ "size": 49937593,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:6c81c36100ac3e6b034b455564d4455ecaa854542bc46cfb4a0bbca4817a6dcd",
+ "size": 1712624,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:c08735c7f0f9602a107c9b3b786ca0e9d07d35af36c3434a8c3fd9328aa4dd16",
+ "size": 446,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:58582cf2ee4ad12220418b6dc5a4fac69bc2602f1f0dca3b411dd08d3a316518",
+ "size": 130038934,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:f606cbdf28a35d2ad3e6a27852667512d0286aa613e959ab5c50aae122911358",
+ "size": 137295077,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:7277b9489a4739d96c3447809dec4587fbf938351530f7f6e184cb3b357f6a16",
+ "size": 94,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ }
+ ]
+ },
+ {
+ "digest": "sha256:0f2dd0cc1c8604cc980ff922d5f72106900644120325a1a3f5dce64af0ee81d3",
+ "repository": "loom",
+ "shape": "image",
+ "mediaType": "application/vnd.oci.image.manifest.v1+json",
+ "holdDid": "did:plc:wnbpdx4bn5uts5yd2o2g3gnm",
+ "config": {
+ "digest": "sha256:d51b873217871ec4490b810ae85605820f03459cfe75c855a4ab2a1771226632",
+ "size": 4676,
+ "mediaType": "application/vnd.oci.image.config.v1+json"
+ },
+ "layers": [
+ {
+ "digest": "sha256:0e7df641b9650fff27219d95215561d9505c24acf5bb919c101091aaf45c9b6b",
+ "size": 85813,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:c172f21841dff4c8cf45cde46589c1c2616cefe7e819965e92e6d3475c428aa0",
+ "size": 12675,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:b4e6f1bfce0a1fba2b5421041552f4a897aada9cd5680926580f9e2c6247a7ae",
+ "size": 288209,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:b4242723c53fe4e094eb78569a2c15b6aafb8eb42aa9c3c2666130654a316ae2",
+ "size": 254104,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:d6b1b89eccacc15c2420b2776d72c1dae334a00805ed9af54bf2f71e4d536f28",
+ "size": 32093,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:2780920e5dbfbe103d03a583ed75345306e572ec5a48cb10361f046767d9f29a",
+ "size": 67,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:7c12895b777bcaa8ccae0605b4de635b68fc32d60fa08f421dc3818bf55ee212",
+ "size": 188,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:3214acf345c0cc6bbdb56b698a41ccdefc624a09d6beb0d38b5de0b2303ecaf4",
+ "size": 123,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:52630fc75a18675c530ed9eba5f55eca09b03e91bd5bc15307918bbc1a7e7296",
+ "size": 162,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:dd64bf2dd177757451a98fcdc999a339c35dee5d9872d8f4dc69c8f3c4dd0112",
+ "size": 80,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:b839dfae01f66e15c6a8b63520557ed315bdfe036342fa7a0c537259f10d7a9a",
+ "size": 351,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:ebddc55facdc6b1f7e0f30816a5fc7cc62f38abdf76c0a8b0a0ce52085754795",
+ "size": 311,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:bdfd7f7e5bf6fc27e70b59101db21c3d8284d283884419dd5fe7020583bb79ca",
+ "size": 136993,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:f3e519d19e1f1f55031eb74473e730dea946f61217ec7d49b5f3e951bb2ca4cb",
+ "size": 5140955,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:25f4fd3c8ebe47ad2dfcc86e29bf139a68b552e04f5966b8aa537e793c9c06fd",
+ "size": 3408462,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:52b2cf548ae5e56fa607ae06498752a8ed05fb4ce56128e90b2380945552c54b",
+ "size": 346965,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:8072b5a0f7957662b5097bbaec05c482fb580118820d83849421d879854b12be",
+ "size": 95467,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:f2dad492a6263bedb536ec1922bb2ba621fe0b019a564e232767515439d5a1ad",
+ "size": 20389882,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:ce95a2c6a9ddd9487978e4ad98c8f5f1f7b453707076c9087a45a97ceafaf616",
+ "size": 6447448,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ }
+ ]
+ },
+ {
+ "digest": "sha256:30d7f33c7f15ff3c6a1e4302575dcebfe31f3e2403486a2902a6bf30d44c7cdd",
+ "repository": "loom",
+ "shape": "image",
+ "mediaType": "application/vnd.oci.image.manifest.v1+json",
+ "holdDid": "did:plc:wnbpdx4bn5uts5yd2o2g3gnm",
+ "config": {
+ "digest": "sha256:3c6aa261908f7fd979f64e19a95a6b921564ed6197c13536fb73a9a34d2f6f68",
+ "size": 4683,
+ "mediaType": "application/vnd.oci.image.config.v1+json"
+ },
+ "layers": [
+ {
+ "digest": "sha256:0e7df641b9650fff27219d95215561d9505c24acf5bb919c101091aaf45c9b6b",
+ "size": 85813,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:c172f21841dff4c8cf45cde46589c1c2616cefe7e819965e92e6d3475c428aa0",
+ "size": 12675,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:b4e6f1bfce0a1fba2b5421041552f4a897aada9cd5680926580f9e2c6247a7ae",
+ "size": 288209,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:b4242723c53fe4e094eb78569a2c15b6aafb8eb42aa9c3c2666130654a316ae2",
+ "size": 254104,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:d6b1b89eccacc15c2420b2776d72c1dae334a00805ed9af54bf2f71e4d536f28",
+ "size": 32093,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:2780920e5dbfbe103d03a583ed75345306e572ec5a48cb10361f046767d9f29a",
+ "size": 67,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:7c12895b777bcaa8ccae0605b4de635b68fc32d60fa08f421dc3818bf55ee212",
+ "size": 188,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:3214acf345c0cc6bbdb56b698a41ccdefc624a09d6beb0d38b5de0b2303ecaf4",
+ "size": 123,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:52630fc75a18675c530ed9eba5f55eca09b03e91bd5bc15307918bbc1a7e7296",
+ "size": 162,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:dd64bf2dd177757451a98fcdc999a339c35dee5d9872d8f4dc69c8f3c4dd0112",
+ "size": 80,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:b839dfae01f66e15c6a8b63520557ed315bdfe036342fa7a0c537259f10d7a9a",
+ "size": 351,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:ebddc55facdc6b1f7e0f30816a5fc7cc62f38abdf76c0a8b0a0ce52085754795",
+ "size": 311,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:bdfd7f7e5bf6fc27e70b59101db21c3d8284d283884419dd5fe7020583bb79ca",
+ "size": 136993,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:f3e519d19e1f1f55031eb74473e730dea946f61217ec7d49b5f3e951bb2ca4cb",
+ "size": 5140955,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:cc446582cc7ca5c34713b585da8c0bc9322efd90b70541523c3309e011dda165",
+ "size": 3407812,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:52b2cf548ae5e56fa607ae06498752a8ed05fb4ce56128e90b2380945552c54b",
+ "size": 346965,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:8072b5a0f7957662b5097bbaec05c482fb580118820d83849421d879854b12be",
+ "size": 95467,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:3d7d9b0879c36eff0609d90cfc254a7760ba614cfa64c5732ebb29e174b6d763",
+ "size": 20344755,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:d8c45399be3765b09c6211943b179709ad03e5083a933eff2a6e1682d5b5747b",
+ "size": 6442020,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ }
+ ]
+ },
+ {
+ "digest": "sha256:450c330c1dc051d27015308b93925f74142d7736d5298fdd71a78a08563a6a09",
+ "repository": "loom",
+ "shape": "image",
+ "mediaType": "application/vnd.oci.image.manifest.v1+json",
+ "holdDid": "did:plc:wnbpdx4bn5uts5yd2o2g3gnm",
+ "config": {
+ "digest": "sha256:a796c001dae0e905c930e032f39b5419f78ef178976928694e04285482b3dd35",
+ "size": 4677,
+ "mediaType": "application/vnd.oci.image.config.v1+json"
+ },
+ "layers": [
+ {
+ "digest": "sha256:fa8ae93e2b3a7478248483e942ff665efa7219c6cd72d7a03c775372076e98dc",
+ "size": 85841,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:c172f21841dff4c8cf45cde46589c1c2616cefe7e819965e92e6d3475c428aa0",
+ "size": 12675,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:b4e6f1bfce0a1fba2b5421041552f4a897aada9cd5680926580f9e2c6247a7ae",
+ "size": 288209,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:b4242723c53fe4e094eb78569a2c15b6aafb8eb42aa9c3c2666130654a316ae2",
+ "size": 254104,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:d6b1b89eccacc15c2420b2776d72c1dae334a00805ed9af54bf2f71e4d536f28",
+ "size": 32093,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:2780920e5dbfbe103d03a583ed75345306e572ec5a48cb10361f046767d9f29a",
+ "size": 67,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:7c12895b777bcaa8ccae0605b4de635b68fc32d60fa08f421dc3818bf55ee212",
+ "size": 188,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:3214acf345c0cc6bbdb56b698a41ccdefc624a09d6beb0d38b5de0b2303ecaf4",
+ "size": 123,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:52630fc75a18675c530ed9eba5f55eca09b03e91bd5bc15307918bbc1a7e7296",
+ "size": 162,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:dd64bf2dd177757451a98fcdc999a339c35dee5d9872d8f4dc69c8f3c4dd0112",
+ "size": 80,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:b839dfae01f66e15c6a8b63520557ed315bdfe036342fa7a0c537259f10d7a9a",
+ "size": 351,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:ebddc55facdc6b1f7e0f30816a5fc7cc62f38abdf76c0a8b0a0ce52085754795",
+ "size": 311,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:bdfd7f7e5bf6fc27e70b59101db21c3d8284d283884419dd5fe7020583bb79ca",
+ "size": 136993,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:4ec36cb7292a44a45129867ef6d6ae1715956938fbca32c926ec0f600d884101",
+ "size": 5075464,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:b73b1097e3a18c388a239e24aee83f0b69599a806a6cf62d08fa0f72e3240c28",
+ "size": 3063845,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:bd8962e292918c50c90edebd9684c053bc386b9c5503acd8fec75d0c6f93a0b7",
+ "size": 372772,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:cac2ae0193cb073e7492050b05fc342a888651b4bbc966908c8793f889c299ba",
+ "size": 96148,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:1db78c3af4269be77f4b745c4d3b182902604a616ada82079efbdc810f480612",
+ "size": 23121091,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:85d6e923d2eb8b300fb96d5ace7cfa703c0d7c34a45fc0a004f5f880688c2871",
+ "size": 7219332,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ }
+ ]
+ },
+ {
+ "digest": "sha256:7c63a98bbe9f6428c9446e10022c48123a7f851a3ee8b5ebce4917c646703812",
+ "repository": "loom",
+ "shape": "image",
+ "mediaType": "application/vnd.oci.image.manifest.v1+json",
+ "holdDid": "did:plc:wnbpdx4bn5uts5yd2o2g3gnm",
+ "config": {
+ "digest": "sha256:0a057bdacc0ecef655d7a95d3616f0fb45b2df32fe487e0da9e16230cc1c77d8",
+ "size": 4673,
+ "mediaType": "application/vnd.oci.image.config.v1+json"
+ },
+ "layers": [
+ {
+ "digest": "sha256:fa8ae93e2b3a7478248483e942ff665efa7219c6cd72d7a03c775372076e98dc",
+ "size": 85841,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:c172f21841dff4c8cf45cde46589c1c2616cefe7e819965e92e6d3475c428aa0",
+ "size": 12675,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:b4e6f1bfce0a1fba2b5421041552f4a897aada9cd5680926580f9e2c6247a7ae",
+ "size": 288209,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:b4242723c53fe4e094eb78569a2c15b6aafb8eb42aa9c3c2666130654a316ae2",
+ "size": 254104,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:d6b1b89eccacc15c2420b2776d72c1dae334a00805ed9af54bf2f71e4d536f28",
+ "size": 32093,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:2780920e5dbfbe103d03a583ed75345306e572ec5a48cb10361f046767d9f29a",
+ "size": 67,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:7c12895b777bcaa8ccae0605b4de635b68fc32d60fa08f421dc3818bf55ee212",
+ "size": 188,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:3214acf345c0cc6bbdb56b698a41ccdefc624a09d6beb0d38b5de0b2303ecaf4",
+ "size": 123,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:52630fc75a18675c530ed9eba5f55eca09b03e91bd5bc15307918bbc1a7e7296",
+ "size": 162,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:dd64bf2dd177757451a98fcdc999a339c35dee5d9872d8f4dc69c8f3c4dd0112",
+ "size": 80,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:b839dfae01f66e15c6a8b63520557ed315bdfe036342fa7a0c537259f10d7a9a",
+ "size": 351,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:ebddc55facdc6b1f7e0f30816a5fc7cc62f38abdf76c0a8b0a0ce52085754795",
+ "size": 311,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:bdfd7f7e5bf6fc27e70b59101db21c3d8284d283884419dd5fe7020583bb79ca",
+ "size": 136993,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:4ec36cb7292a44a45129867ef6d6ae1715956938fbca32c926ec0f600d884101",
+ "size": 5075464,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:b73b1097e3a18c388a239e24aee83f0b69599a806a6cf62d08fa0f72e3240c28",
+ "size": 3063845,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:bd8962e292918c50c90edebd9684c053bc386b9c5503acd8fec75d0c6f93a0b7",
+ "size": 372772,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:cac2ae0193cb073e7492050b05fc342a888651b4bbc966908c8793f889c299ba",
+ "size": 96148,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:d6ffcdceacc5cc7572ecae6e886b40fba780bc600c7c746cf4c5ff4ae9e38a2a",
+ "size": 23119498,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:1f9e08fdc48a0f2a9f7d5db34ea0a35fe73f5e017eb1235c41647b4001a06901",
+ "size": 7219331,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ }
+ ]
+ },
+ {
+ "digest": "sha256:881dc3357bb92d4455f0f9b2989858e522c46d395c91feb002a3c0bebda4387c",
+ "repository": "loom",
+ "shape": "image",
+ "mediaType": "application/vnd.oci.image.manifest.v1+json",
+ "holdDid": "did:plc:wnbpdx4bn5uts5yd2o2g3gnm",
+ "config": {
+ "digest": "sha256:ba11065835ac8c9ce54f871d9222dfcbcf05dd8c18d53f6870ead97482d957b2",
+ "size": 4677,
+ "mediaType": "application/vnd.oci.image.config.v1+json"
+ },
+ "layers": [
+ {
+ "digest": "sha256:0e7df641b9650fff27219d95215561d9505c24acf5bb919c101091aaf45c9b6b",
+ "size": 85813,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:c172f21841dff4c8cf45cde46589c1c2616cefe7e819965e92e6d3475c428aa0",
+ "size": 12675,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:b4e6f1bfce0a1fba2b5421041552f4a897aada9cd5680926580f9e2c6247a7ae",
+ "size": 288209,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:b4242723c53fe4e094eb78569a2c15b6aafb8eb42aa9c3c2666130654a316ae2",
+ "size": 254104,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:d6b1b89eccacc15c2420b2776d72c1dae334a00805ed9af54bf2f71e4d536f28",
+ "size": 32093,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:2780920e5dbfbe103d03a583ed75345306e572ec5a48cb10361f046767d9f29a",
+ "size": 67,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:7c12895b777bcaa8ccae0605b4de635b68fc32d60fa08f421dc3818bf55ee212",
+ "size": 188,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:3214acf345c0cc6bbdb56b698a41ccdefc624a09d6beb0d38b5de0b2303ecaf4",
+ "size": 123,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:52630fc75a18675c530ed9eba5f55eca09b03e91bd5bc15307918bbc1a7e7296",
+ "size": 162,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:dd64bf2dd177757451a98fcdc999a339c35dee5d9872d8f4dc69c8f3c4dd0112",
+ "size": 80,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:b839dfae01f66e15c6a8b63520557ed315bdfe036342fa7a0c537259f10d7a9a",
+ "size": 351,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:ebddc55facdc6b1f7e0f30816a5fc7cc62f38abdf76c0a8b0a0ce52085754795",
+ "size": 311,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:bdfd7f7e5bf6fc27e70b59101db21c3d8284d283884419dd5fe7020583bb79ca",
+ "size": 136993,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:f3e519d19e1f1f55031eb74473e730dea946f61217ec7d49b5f3e951bb2ca4cb",
+ "size": 5140955,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:25f4fd3c8ebe47ad2dfcc86e29bf139a68b552e04f5966b8aa537e793c9c06fd",
+ "size": 3408462,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:52b2cf548ae5e56fa607ae06498752a8ed05fb4ce56128e90b2380945552c54b",
+ "size": 346965,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:8072b5a0f7957662b5097bbaec05c482fb580118820d83849421d879854b12be",
+ "size": 95467,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:48b71642fdd370d88677062a51ec1b3461ced4f784963c7b1caa1dbe05e30643",
+ "size": 20689773,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:c4a57e87b85f26024f5c0bc762dc6f8afc1cc829202807c303ac68136d3bfee3",
+ "size": 6451311,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ }
+ ]
+ },
+ {
+ "digest": "sha256:a4297a7dd6f5874778046372147ff530d1b620302f2186eb3d554c61ef0a2daa",
+ "repository": "loom",
+ "shape": "image",
+ "mediaType": "application/vnd.oci.image.manifest.v1+json",
+ "holdDid": "did:plc:wnbpdx4bn5uts5yd2o2g3gnm",
+ "config": {
+ "digest": "sha256:37cd9fc3239fe226acf03bce7d395daa1770933d9fb7b82eed044cae79e08664",
+ "size": 4673,
+ "mediaType": "application/vnd.oci.image.config.v1+json"
+ },
+ "layers": [
+ {
+ "digest": "sha256:0e7df641b9650fff27219d95215561d9505c24acf5bb919c101091aaf45c9b6b",
+ "size": 85813,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:c172f21841dff4c8cf45cde46589c1c2616cefe7e819965e92e6d3475c428aa0",
+ "size": 12675,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:b4e6f1bfce0a1fba2b5421041552f4a897aada9cd5680926580f9e2c6247a7ae",
+ "size": 288209,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:b4242723c53fe4e094eb78569a2c15b6aafb8eb42aa9c3c2666130654a316ae2",
+ "size": 254104,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:d6b1b89eccacc15c2420b2776d72c1dae334a00805ed9af54bf2f71e4d536f28",
+ "size": 32093,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:2780920e5dbfbe103d03a583ed75345306e572ec5a48cb10361f046767d9f29a",
+ "size": 67,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:7c12895b777bcaa8ccae0605b4de635b68fc32d60fa08f421dc3818bf55ee212",
+ "size": 188,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:3214acf345c0cc6bbdb56b698a41ccdefc624a09d6beb0d38b5de0b2303ecaf4",
+ "size": 123,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:52630fc75a18675c530ed9eba5f55eca09b03e91bd5bc15307918bbc1a7e7296",
+ "size": 162,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:dd64bf2dd177757451a98fcdc999a339c35dee5d9872d8f4dc69c8f3c4dd0112",
+ "size": 80,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:b839dfae01f66e15c6a8b63520557ed315bdfe036342fa7a0c537259f10d7a9a",
+ "size": 351,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:ebddc55facdc6b1f7e0f30816a5fc7cc62f38abdf76c0a8b0a0ce52085754795",
+ "size": 311,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:bdfd7f7e5bf6fc27e70b59101db21c3d8284d283884419dd5fe7020583bb79ca",
+ "size": 136993,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:f3e519d19e1f1f55031eb74473e730dea946f61217ec7d49b5f3e951bb2ca4cb",
+ "size": 5140955,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:25f4fd3c8ebe47ad2dfcc86e29bf139a68b552e04f5966b8aa537e793c9c06fd",
+ "size": 3408462,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:52b2cf548ae5e56fa607ae06498752a8ed05fb4ce56128e90b2380945552c54b",
+ "size": 346965,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:8072b5a0f7957662b5097bbaec05c482fb580118820d83849421d879854b12be",
+ "size": 95467,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:45b5001bd65806926478dfefbb5b083668a6cac48405f7864001bca5c0f8de22",
+ "size": 20689780,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:bd3d6299a3ca6ddaf3faf056bdc24c4b8357704035321ef0a052a9f5e3c54e14",
+ "size": 6451311,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ }
+ ]
+ },
+ {
+ "digest": "sha256:abbcf87344eaf7751067f1df58f93697cedbfc1b59b1f3f638a2d1eaa3a69d27",
+ "repository": "loom",
+ "shape": "image",
+ "mediaType": "application/vnd.oci.image.manifest.v1+json",
+ "holdDid": "did:plc:wnbpdx4bn5uts5yd2o2g3gnm",
+ "config": {
+ "digest": "sha256:8964f8b0e8db48b0cf3b675dfccd79e695f87c95b4479522484c79a96e7024c0",
+ "size": 4687,
+ "mediaType": "application/vnd.oci.image.config.v1+json"
+ },
+ "layers": [
+ {
+ "digest": "sha256:fa8ae93e2b3a7478248483e942ff665efa7219c6cd72d7a03c775372076e98dc",
+ "size": 85841,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:c172f21841dff4c8cf45cde46589c1c2616cefe7e819965e92e6d3475c428aa0",
+ "size": 12675,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:b4e6f1bfce0a1fba2b5421041552f4a897aada9cd5680926580f9e2c6247a7ae",
+ "size": 288209,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:b4242723c53fe4e094eb78569a2c15b6aafb8eb42aa9c3c2666130654a316ae2",
+ "size": 254104,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:d6b1b89eccacc15c2420b2776d72c1dae334a00805ed9af54bf2f71e4d536f28",
+ "size": 32093,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:2780920e5dbfbe103d03a583ed75345306e572ec5a48cb10361f046767d9f29a",
+ "size": 67,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:7c12895b777bcaa8ccae0605b4de635b68fc32d60fa08f421dc3818bf55ee212",
+ "size": 188,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:3214acf345c0cc6bbdb56b698a41ccdefc624a09d6beb0d38b5de0b2303ecaf4",
+ "size": 123,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:52630fc75a18675c530ed9eba5f55eca09b03e91bd5bc15307918bbc1a7e7296",
+ "size": 162,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:dd64bf2dd177757451a98fcdc999a339c35dee5d9872d8f4dc69c8f3c4dd0112",
+ "size": 80,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:b839dfae01f66e15c6a8b63520557ed315bdfe036342fa7a0c537259f10d7a9a",
+ "size": 351,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:ebddc55facdc6b1f7e0f30816a5fc7cc62f38abdf76c0a8b0a0ce52085754795",
+ "size": 311,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:bdfd7f7e5bf6fc27e70b59101db21c3d8284d283884419dd5fe7020583bb79ca",
+ "size": 136993,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:4ec36cb7292a44a45129867ef6d6ae1715956938fbca32c926ec0f600d884101",
+ "size": 5075464,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:e65d8d69ea295fa49afb39cf9867e5da8f3e26bb9a9ee9ba0bf926ecd675326b",
+ "size": 3063236,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:bd8962e292918c50c90edebd9684c053bc386b9c5503acd8fec75d0c6f93a0b7",
+ "size": 372772,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:cac2ae0193cb073e7492050b05fc342a888651b4bbc966908c8793f889c299ba",
+ "size": 96148,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:b995f0f2de80fcbd1ced49fd2b5a40dbb7a0c00694e5861e140c3ed14db34af0",
+ "size": 22734235,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:4a82940ac1c411086523f93ee771882afefeacae1ab7573b8d1a759277bc45d9",
+ "size": 7207004,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ }
+ ]
+ },
+ {
+ "digest": "sha256:ac29ffa1161f0c2a9097da0e1b3f6bc9940d1941287152c4f0093859dc61d991",
+ "repository": "loom",
+ "shape": "image",
+ "mediaType": "application/vnd.oci.image.manifest.v1+json",
+ "holdDid": "did:plc:wnbpdx4bn5uts5yd2o2g3gnm",
+ "config": {
+ "digest": "sha256:e5299ee7fc470b409143257658e579e273df494f836c03ed7dc8777da29d627d",
+ "size": 4673,
+ "mediaType": "application/vnd.oci.image.config.v1+json"
+ },
+ "layers": [
+ {
+ "digest": "sha256:fa8ae93e2b3a7478248483e942ff665efa7219c6cd72d7a03c775372076e98dc",
+ "size": 85841,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:c172f21841dff4c8cf45cde46589c1c2616cefe7e819965e92e6d3475c428aa0",
+ "size": 12675,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:b4e6f1bfce0a1fba2b5421041552f4a897aada9cd5680926580f9e2c6247a7ae",
+ "size": 288209,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:b4242723c53fe4e094eb78569a2c15b6aafb8eb42aa9c3c2666130654a316ae2",
+ "size": 254104,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:d6b1b89eccacc15c2420b2776d72c1dae334a00805ed9af54bf2f71e4d536f28",
+ "size": 32093,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:2780920e5dbfbe103d03a583ed75345306e572ec5a48cb10361f046767d9f29a",
+ "size": 67,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:7c12895b777bcaa8ccae0605b4de635b68fc32d60fa08f421dc3818bf55ee212",
+ "size": 188,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:3214acf345c0cc6bbdb56b698a41ccdefc624a09d6beb0d38b5de0b2303ecaf4",
+ "size": 123,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:52630fc75a18675c530ed9eba5f55eca09b03e91bd5bc15307918bbc1a7e7296",
+ "size": 162,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:dd64bf2dd177757451a98fcdc999a339c35dee5d9872d8f4dc69c8f3c4dd0112",
+ "size": 80,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:b839dfae01f66e15c6a8b63520557ed315bdfe036342fa7a0c537259f10d7a9a",
+ "size": 351,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:ebddc55facdc6b1f7e0f30816a5fc7cc62f38abdf76c0a8b0a0ce52085754795",
+ "size": 311,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:bdfd7f7e5bf6fc27e70b59101db21c3d8284d283884419dd5fe7020583bb79ca",
+ "size": 136993,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:4ec36cb7292a44a45129867ef6d6ae1715956938fbca32c926ec0f600d884101",
+ "size": 5075464,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:b73b1097e3a18c388a239e24aee83f0b69599a806a6cf62d08fa0f72e3240c28",
+ "size": 3063845,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:bd8962e292918c50c90edebd9684c053bc386b9c5503acd8fec75d0c6f93a0b7",
+ "size": 372772,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:cac2ae0193cb073e7492050b05fc342a888651b4bbc966908c8793f889c299ba",
+ "size": 96148,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:a836c46a8b8ef8cd7eb62fb97f22806e776e73e9783ef1cdaa39708d5c58285a",
+ "size": 23119500,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:5e93bcc1583ed1dc01776b2511e8b29acff5d72f7587d6658eac3ef98643559d",
+ "size": 7219331,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ }
+ ]
+ },
+ {
+ "digest": "sha256:b495605bcafd1e7d3e911c49df242265f4d8f68379d29b48524e460354cbb4bc",
+ "repository": "loom",
+ "shape": "image",
+ "mediaType": "application/vnd.oci.image.manifest.v1+json",
+ "holdDid": "did:plc:wnbpdx4bn5uts5yd2o2g3gnm",
+ "config": {
+ "digest": "sha256:20a4ce80cc10b4e9588d4ab9c95bfe9913471d7a78eda16e80c82e547d653113",
+ "size": 4676,
+ "mediaType": "application/vnd.oci.image.config.v1+json"
+ },
+ "layers": [
+ {
+ "digest": "sha256:0e7df641b9650fff27219d95215561d9505c24acf5bb919c101091aaf45c9b6b",
+ "size": 85813,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:c172f21841dff4c8cf45cde46589c1c2616cefe7e819965e92e6d3475c428aa0",
+ "size": 12675,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:b4e6f1bfce0a1fba2b5421041552f4a897aada9cd5680926580f9e2c6247a7ae",
+ "size": 288209,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:b4242723c53fe4e094eb78569a2c15b6aafb8eb42aa9c3c2666130654a316ae2",
+ "size": 254104,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:d6b1b89eccacc15c2420b2776d72c1dae334a00805ed9af54bf2f71e4d536f28",
+ "size": 32093,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:2780920e5dbfbe103d03a583ed75345306e572ec5a48cb10361f046767d9f29a",
+ "size": 67,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:7c12895b777bcaa8ccae0605b4de635b68fc32d60fa08f421dc3818bf55ee212",
+ "size": 188,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:3214acf345c0cc6bbdb56b698a41ccdefc624a09d6beb0d38b5de0b2303ecaf4",
+ "size": 123,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:52630fc75a18675c530ed9eba5f55eca09b03e91bd5bc15307918bbc1a7e7296",
+ "size": 162,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:dd64bf2dd177757451a98fcdc999a339c35dee5d9872d8f4dc69c8f3c4dd0112",
+ "size": 80,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:b839dfae01f66e15c6a8b63520557ed315bdfe036342fa7a0c537259f10d7a9a",
+ "size": 351,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:ebddc55facdc6b1f7e0f30816a5fc7cc62f38abdf76c0a8b0a0ce52085754795",
+ "size": 311,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:bdfd7f7e5bf6fc27e70b59101db21c3d8284d283884419dd5fe7020583bb79ca",
+ "size": 136993,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:f3e519d19e1f1f55031eb74473e730dea946f61217ec7d49b5f3e951bb2ca4cb",
+ "size": 5140955,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:25f4fd3c8ebe47ad2dfcc86e29bf139a68b552e04f5966b8aa537e793c9c06fd",
+ "size": 3408462,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:52b2cf548ae5e56fa607ae06498752a8ed05fb4ce56128e90b2380945552c54b",
+ "size": 346965,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:8072b5a0f7957662b5097bbaec05c482fb580118820d83849421d879854b12be",
+ "size": 95467,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:4610b4de65b8b1bc4f971af3ebde6805b57f2ab723a1ec7ea8be8201db66c3df",
+ "size": 20690973,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:ddcd883cc0e81c216181f3f997867cc9548d05bac98700e437575a7937fa7c6a",
+ "size": 6451310,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ }
+ ]
+ },
+ {
+ "digest": "sha256:cbeaeb1ec371925960b1fc8e8d7090ea670835a9559a70cd4bf80be17c66f8aa",
+ "repository": "loom",
+ "shape": "image",
+ "mediaType": "application/vnd.oci.image.manifest.v1+json",
+ "holdDid": "did:plc:wnbpdx4bn5uts5yd2o2g3gnm",
+ "config": {
+ "digest": "sha256:a00d99082fe5e571284d159ba964db9eab0344ee3e6ea11ad42f0f0246520c59",
+ "size": 4676,
+ "mediaType": "application/vnd.oci.image.config.v1+json"
+ },
+ "layers": [
+ {
+ "digest": "sha256:fa8ae93e2b3a7478248483e942ff665efa7219c6cd72d7a03c775372076e98dc",
+ "size": 85841,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:c172f21841dff4c8cf45cde46589c1c2616cefe7e819965e92e6d3475c428aa0",
+ "size": 12675,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:b4e6f1bfce0a1fba2b5421041552f4a897aada9cd5680926580f9e2c6247a7ae",
+ "size": 288209,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:b4242723c53fe4e094eb78569a2c15b6aafb8eb42aa9c3c2666130654a316ae2",
+ "size": 254104,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:d6b1b89eccacc15c2420b2776d72c1dae334a00805ed9af54bf2f71e4d536f28",
+ "size": 32093,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:2780920e5dbfbe103d03a583ed75345306e572ec5a48cb10361f046767d9f29a",
+ "size": 67,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:7c12895b777bcaa8ccae0605b4de635b68fc32d60fa08f421dc3818bf55ee212",
+ "size": 188,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:3214acf345c0cc6bbdb56b698a41ccdefc624a09d6beb0d38b5de0b2303ecaf4",
+ "size": 123,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:52630fc75a18675c530ed9eba5f55eca09b03e91bd5bc15307918bbc1a7e7296",
+ "size": 162,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:dd64bf2dd177757451a98fcdc999a339c35dee5d9872d8f4dc69c8f3c4dd0112",
+ "size": 80,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:b839dfae01f66e15c6a8b63520557ed315bdfe036342fa7a0c537259f10d7a9a",
+ "size": 351,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:ebddc55facdc6b1f7e0f30816a5fc7cc62f38abdf76c0a8b0a0ce52085754795",
+ "size": 311,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:bdfd7f7e5bf6fc27e70b59101db21c3d8284d283884419dd5fe7020583bb79ca",
+ "size": 136993,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:4ec36cb7292a44a45129867ef6d6ae1715956938fbca32c926ec0f600d884101",
+ "size": 5075464,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:b73b1097e3a18c388a239e24aee83f0b69599a806a6cf62d08fa0f72e3240c28",
+ "size": 3063845,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:bd8962e292918c50c90edebd9684c053bc386b9c5503acd8fec75d0c6f93a0b7",
+ "size": 372772,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:cac2ae0193cb073e7492050b05fc342a888651b4bbc966908c8793f889c299ba",
+ "size": 96148,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:52da34de16fe4617a65a0fa5a7acb7e94a16aaff0fca28c4c0204d296fd97a67",
+ "size": 22786243,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:f843f4e19552da7a8b51c8e1e2cb4891d9fb6a3261e8f3668495895399207727",
+ "size": 7215250,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ }
+ ]
+ },
+ {
+ "digest": "sha256:1178a001d60a5a6c7ac4cba5a8fad190405510982ce75dcea1ff44be39e004a9",
+ "repository": "loom-runner",
+ "shape": "image",
+ "mediaType": "application/vnd.oci.image.manifest.v1+json",
+ "holdDid": "did:plc:wnbpdx4bn5uts5yd2o2g3gnm",
+ "config": {
+ "digest": "sha256:178ec9e1482996813d8ad8f16397e2b16f9cc772dc298c723233805ca1fa9be1",
+ "size": 1203,
+ "mediaType": "application/vnd.oci.image.config.v1+json"
+ },
+ "layers": [
+ {
+ "digest": "sha256:2d35ebdb57d9971fea0cac1582aa78935adf8058b2cc32db163c98822e5dfa1b",
+ "size": 3802452,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:c6bb25765dcc98a74d58df44516fa5ee29c004c876caf87b8b8482686ea23be6",
+ "size": 291164,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:5c76186593304accbdad9ed22c30450f47c101fb8bc7e73d2106fb38848abd55",
+ "size": 6763426,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ }
+ ]
+ },
+ {
+ "digest": "sha256:1fd4b800e2605202f4844b38732cf6b0732a2233d067d5b0d7249ff3a37f02f9",
+ "repository": "loom-runner",
+ "shape": "image",
+ "mediaType": "application/vnd.oci.image.manifest.v1+json",
+ "holdDid": "did:plc:wnbpdx4bn5uts5yd2o2g3gnm",
+ "config": {
+ "digest": "sha256:bc601b56904ec267b1042e9cbd62cf192734e113c004a53dd499de978ec1e9f1",
+ "size": 1201,
+ "mediaType": "application/vnd.oci.image.config.v1+json"
+ },
+ "layers": [
+ {
+ "digest": "sha256:6b59a28fa20117e6048ad0616b8d8c901877ef15ff4c7f18db04e4f01f43bc39",
+ "size": 4138069,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:61d7d39eb37b8732f8590faffe13c58e5f6a2c619776aaf1ad0de1f79df0a9af",
+ "size": 294109,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:e06db7f10d82e9e5f8a3daf5a52b0c8f0067a44bf6a55f371fb3146435eccf24",
+ "size": 6171626,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ }
+ ]
+ },
+ {
+ "digest": "sha256:f7d3e846b6e66188b94dc317ab162a7a81deaac212f4a9aa91917fb609b050dc",
+ "repository": "multi",
+ "shape": "image",
+ "mediaType": "application/vnd.docker.distribution.manifest.v2+json",
+ "holdDid": "did:web:localhost%3A8080",
+ "config": {
+ "digest": "sha256:7ba4384e714c8ff6b48dc9a9335c4bc14bcd557ff63e47fe7192863f20f8e149",
+ "size": 560,
+ "mediaType": "application/vnd.docker.container.image.v1+json"
+ },
+ "layers": [
+ {
+ "digest": "sha256:9616a727d776944e0fc02e7ab68788d96023a6069408aebbf7051607dc985a0b",
+ "size": 152,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:41c771bb4f506653379d38763dee0a699a7d86a47e7bf840ef59f3d4e3aae1cc",
+ "size": 153,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:b58fba4a75277e861e92c48e953aac19d65ee33a8c3fde772f0273d3039572b3",
+ "size": 152,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:5cb5a6d2eb42aa0e1caab9f66f2ccebb3851a4580a02651a3144cf7e18959874",
+ "size": 151,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ }
+ ]
+ },
+ {
+ "digest": "sha256:1c9d101e18faf7a0133ac28f228a4305878b0f361e98066428d88915c569e2bc",
+ "repository": "secret-service-operator",
+ "shape": "image",
+ "mediaType": "application/vnd.docker.distribution.manifest.v2+json",
+ "holdDid": "did:plc:wnbpdx4bn5uts5yd2o2g3gnm",
+ "config": {
+ "digest": "sha256:fcd06383d1942cb274276b8111646cafa0cc84325cac700f85d51325346a7060",
+ "size": 2134,
+ "mediaType": "application/vnd.docker.container.image.v1+json"
+ },
+ "layers": [
+ {
+ "digest": "sha256:51c1b6699f435b7ccff149db8fdfc0479d802406fea5712271fac54f97eb3b8f",
+ "size": 84670,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:2e4cf50eeb92ac3a7afe75e15d96a26dee99449f86b46c75b5d95f4418a5bca0",
+ "size": 12579,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:4e9f20d26c878a5db592123720f66b04bddf045879f6c0ad45e069a991543fa9",
+ "size": 458279,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:0f8b424aa0b96c1c388a5fd4d90735604459256336853082afb61733438872b5",
+ "size": 75,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:d557676654e572af3e3173c90e7874644207fda32cd87e9d3d66b5d7b98a7b21",
+ "size": 193,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:d82bc7a76a838c9a4a6025192429c2fed58f73742ef1fb9c8bb7b995fc3b7213",
+ "size": 130,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:d858cbc252ade14879807ff8dbc3043a26bbdb92087da98cda831ee040b172b3",
+ "size": 173,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:1069fc2daed1aceff7232f4b8ab21200dd3d8b04f61be9da86977a34a105dfdc",
+ "size": 97,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:b40161cd83fc5d470d6abe50e87aa288481b6b89137012881d74187cfbf9f502",
+ "size": 382,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:3f4e2c5863480125882d92060440a5250766bce764fee10acdbac18c872e4dc7",
+ "size": 326,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:80a8c047508ae5cd6a591060fc43422cb8e3aea1bd908d913e8f0146e2297fea",
+ "size": 129107,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:daafffdb23e69fe061034d719891849ef57a641acd13fd79d2e3507cc9c50807",
+ "size": 31201686,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ }
+ ]
+ },
+ {
+ "digest": "sha256:2046e88531a108d5428ce082a1495da72dada4a03e38eb7e8d3a7594c587ebff",
+ "repository": "secret-service-operator",
+ "shape": "image",
+ "mediaType": "application/vnd.docker.distribution.manifest.v2+json",
+ "holdDid": "did:plc:wnbpdx4bn5uts5yd2o2g3gnm",
+ "config": {
+ "digest": "sha256:9bbfefd88fcadaf2ee57eb80921b6483ff5d3103b6ced797ac15f6309e3a7d97",
+ "size": 2134,
+ "mediaType": "application/vnd.docker.container.image.v1+json"
+ },
+ "layers": [
+ {
+ "digest": "sha256:2ae251fec02fb291b816530bdcf7100d568a00cf07a17962297fc48f43198368",
+ "size": 84670,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:2e4cf50eeb92ac3a7afe75e15d96a26dee99449f86b46c75b5d95f4418a5bca0",
+ "size": 12579,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:4e9f20d26c878a5db592123720f66b04bddf045879f6c0ad45e069a991543fa9",
+ "size": 458279,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:0f8b424aa0b96c1c388a5fd4d90735604459256336853082afb61733438872b5",
+ "size": 75,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:d557676654e572af3e3173c90e7874644207fda32cd87e9d3d66b5d7b98a7b21",
+ "size": 193,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:d82bc7a76a838c9a4a6025192429c2fed58f73742ef1fb9c8bb7b995fc3b7213",
+ "size": 130,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:d858cbc252ade14879807ff8dbc3043a26bbdb92087da98cda831ee040b172b3",
+ "size": 173,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:1069fc2daed1aceff7232f4b8ab21200dd3d8b04f61be9da86977a34a105dfdc",
+ "size": 97,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:b40161cd83fc5d470d6abe50e87aa288481b6b89137012881d74187cfbf9f502",
+ "size": 382,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:3f4e2c5863480125882d92060440a5250766bce764fee10acdbac18c872e4dc7",
+ "size": 326,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:80a8c047508ae5cd6a591060fc43422cb8e3aea1bd908d913e8f0146e2297fea",
+ "size": 129107,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:cfca3a1d06eab2fe0bcfa4198f7ba02c4bfdb082c666937acbb35b4799316fc1",
+ "size": 29119239,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ }
+ ]
+ },
+ {
+ "digest": "sha256:3acc2f3e9fd5a23fb0dea0c586ccd4eb1031276e3532259a396c629d111f4fdd",
+ "repository": "secret-service-operator",
+ "shape": "image",
+ "mediaType": "application/vnd.docker.distribution.manifest.v2+json",
+ "holdDid": "did:plc:wnbpdx4bn5uts5yd2o2g3gnm",
+ "config": {
+ "digest": "sha256:ea916da70cced3e1487a734052e798f75b3711020388923b4f2187745b32ae8e",
+ "size": 3068,
+ "mediaType": "application/vnd.docker.container.image.v1+json"
+ },
+ "layers": [
+ {
+ "digest": "sha256:7af6fd0da1f0a6ed90049962f5559afd1664115ea92c960fb06e9cbc55097314",
+ "size": 86391,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:5b3013706a463509f05fabda3a5f854b1c5f3429561cf9374db6e6681da5493f",
+ "size": 12777,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:b5531e57aa11a76b1b7ab455627c2d102dd9cba3e7a834be3238376de5de8fb9",
+ "size": 294955,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:6916de9450eeaa0fc896dffc5c752e971afc163f2f74f9c1d43460a753e6b424",
+ "size": 263494,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:9d2753d762cc58b055d94d102989775221df5156b7bec5861353454c308ad197",
+ "size": 31908,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:2780920e5dbfbe103d03a583ed75345306e572ec5a48cb10361f046767d9f29a",
+ "size": 67,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:7c12895b777bcaa8ccae0605b4de635b68fc32d60fa08f421dc3818bf55ee212",
+ "size": 188,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:3214acf345c0cc6bbdb56b698a41ccdefc624a09d6beb0d38b5de0b2303ecaf4",
+ "size": 123,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:52630fc75a18675c530ed9eba5f55eca09b03e91bd5bc15307918bbc1a7e7296",
+ "size": 162,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:dd64bf2dd177757451a98fcdc999a339c35dee5d9872d8f4dc69c8f3c4dd0112",
+ "size": 80,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:b839dfae01f66e15c6a8b63520557ed315bdfe036342fa7a0c537259f10d7a9a",
+ "size": 351,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:411e1c946a02ee44c9adc542515ae29ff58d65149a4a201cbd688297c9900ee4",
+ "size": 316,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:77964957095d7872770920c8f55d28613e85f8e27aa0b157a440e0d635193fa9",
+ "size": 136427,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:b37c8d4f50f2118c229c7804e9c2e8ab79c7179f460aaa1d6cdabeffc5d5aa13",
+ "size": 31192983,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ }
+ ]
+ },
+ {
+ "digest": "sha256:8bb7f653dcd5acd6819279835387077be242ad3782fbdb36b9d9e4c3110b6cdd",
+ "repository": "secret-service-operator",
+ "shape": "image",
+ "mediaType": "application/vnd.docker.distribution.manifest.v2+json",
+ "holdDid": "did:plc:wnbpdx4bn5uts5yd2o2g3gnm",
+ "config": {
+ "digest": "sha256:136e7eea345df79e61bb0b500c38f029c565c0e04c712a5f59c5c4a2c8aadb90",
+ "size": 2243,
+ "mediaType": "application/vnd.docker.container.image.v1+json"
+ },
+ "layers": [
+ {
+ "digest": "sha256:259db2ee6b876afac49b298219529e5ff850e8efb940856f8397c17d1bab100e",
+ "size": 84719,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:2e4cf50eeb92ac3a7afe75e15d96a26dee99449f86b46c75b5d95f4418a5bca0",
+ "size": 12579,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:56ce5a7a0a8cc3aded9e8cab00fd85f0a1c50376aa9f2318c4e66beb03eadf8f",
+ "size": 458603,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:e1089d61b200106053b1717881bf0c1c47551478f01569d9224d33cccf3e4692",
+ "size": 28389,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:0f8b424aa0b96c1c388a5fd4d90735604459256336853082afb61733438872b5",
+ "size": 75,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:d557676654e572af3e3173c90e7874644207fda32cd87e9d3d66b5d7b98a7b21",
+ "size": 193,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:d82bc7a76a838c9a4a6025192429c2fed58f73742ef1fb9c8bb7b995fc3b7213",
+ "size": 130,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:68e4cd60e60fed4486d5b4b40b079e57781e828d93d80a7411e8760dcccfbbf7",
+ "size": 168,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:0783a66ae7492bb16937c7723c534e431404aa3ebbdcaaed65d264a95a266b9e",
+ "size": 84,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:b40161cd83fc5d470d6abe50e87aa288481b6b89137012881d74187cfbf9f502",
+ "size": 382,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:6b72b81ed96620b324f40ade069bbd3fcee69d62182bb03681c3976b6909eec4",
+ "size": 320,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:6cdd517ba90571b01009637268b5fddb4e3eea409bb6dae0bf40ed8ad9e33ef4",
+ "size": 130874,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:6125dae3d024e5d3763c136d01109b0c70dd08795f85e08b7e4b5636985ac22c",
+ "size": 31201687,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ }
+ ]
+ },
+ {
+ "digest": "sha256:99e282ec2639504bd734f94f6958f785a99292a9fbeb3f8170442114f8471a62",
+ "repository": "secret-service-operator",
+ "shape": "image",
+ "mediaType": "application/vnd.docker.distribution.manifest.v2+json",
+ "holdDid": "did:plc:wnbpdx4bn5uts5yd2o2g3gnm",
+ "config": {
+ "digest": "sha256:0775706e96f1614406a987aaaf84ae2b6db563ae2ce379409c9b47db811d83d9",
+ "size": 2243,
+ "mediaType": "application/vnd.docker.container.image.v1+json"
+ },
+ "layers": [
+ {
+ "digest": "sha256:44d654bc6e9919c7ea77a18a4b7c8cb114f0a82fee9f9a1a010be958a6b62ee0",
+ "size": 104272,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:2e4cf50eeb92ac3a7afe75e15d96a26dee99449f86b46c75b5d95f4418a5bca0",
+ "size": 12579,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:56ce5a7a0a8cc3aded9e8cab00fd85f0a1c50376aa9f2318c4e66beb03eadf8f",
+ "size": 458603,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:e1089d61b200106053b1717881bf0c1c47551478f01569d9224d33cccf3e4692",
+ "size": 28389,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:0f8b424aa0b96c1c388a5fd4d90735604459256336853082afb61733438872b5",
+ "size": 75,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:d557676654e572af3e3173c90e7874644207fda32cd87e9d3d66b5d7b98a7b21",
+ "size": 193,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:d82bc7a76a838c9a4a6025192429c2fed58f73742ef1fb9c8bb7b995fc3b7213",
+ "size": 130,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:68e4cd60e60fed4486d5b4b40b079e57781e828d93d80a7411e8760dcccfbbf7",
+ "size": 168,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:0783a66ae7492bb16937c7723c534e431404aa3ebbdcaaed65d264a95a266b9e",
+ "size": 84,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:b40161cd83fc5d470d6abe50e87aa288481b6b89137012881d74187cfbf9f502",
+ "size": 382,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:6b72b81ed96620b324f40ade069bbd3fcee69d62182bb03681c3976b6909eec4",
+ "size": 320,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:6cdd517ba90571b01009637268b5fddb4e3eea409bb6dae0bf40ed8ad9e33ef4",
+ "size": 130874,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ },
+ {
+ "digest": "sha256:8a1c3a3e097ec05a93cc9805a706953afb71bfb0797e4bb88183fb8ef08e4c92",
+ "size": 29119237,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ }
+ ]
+ },
+ {
+ "digest": "sha256:e22b7288f5327e75298c3fd47e78541964d8b021669806413d1cca4b7521f870",
+ "repository": "secret-service-operator",
+ "shape": "image",
+ "mediaType": "application/vnd.oci.image.manifest.v1+json",
+ "holdDid": "did:plc:wnbpdx4bn5uts5yd2o2g3gnm",
+ "config": {
+ "digest": "sha256:554d2a6a4920a238f7eb62d393621a22a64aef2d306e6e06a6824180422fe993",
+ "size": 3119,
+ "mediaType": "application/vnd.oci.image.config.v1+json"
+ },
+ "layers": [
+ {
+ "digest": "sha256:fd594acce1faf1270eec97c48c1ddc192acdeb0d380d27b5870dac55ae3d739f",
+ "size": 89187,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:7750c1a1cd8e356f717610d15552bd8a4af1801b72b5855c9e8494cc07636e77",
+ "size": 12963,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:04d5287996342f0af5f1e876ba86230b48508d292c62cf451d8ba3432e03e355",
+ "size": 302815,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:bc7797c9e96cd983f9b836a89bcca43ee0cab9e39c505804f6a19d16b9573dac",
+ "size": 270719,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:97ae0a42274e126686192f377a2b4a112583630eadde1920a761ed910cdbde34",
+ "size": 32972,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:2780920e5dbfbe103d03a583ed75345306e572ec5a48cb10361f046767d9f29a",
+ "size": 67,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:7c12895b777bcaa8ccae0605b4de635b68fc32d60fa08f421dc3818bf55ee212",
+ "size": 188,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:3214acf345c0cc6bbdb56b698a41ccdefc624a09d6beb0d38b5de0b2303ecaf4",
+ "size": 123,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:52630fc75a18675c530ed9eba5f55eca09b03e91bd5bc15307918bbc1a7e7296",
+ "size": 162,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:dd64bf2dd177757451a98fcdc999a339c35dee5d9872d8f4dc69c8f3c4dd0112",
+ "size": 80,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:b839dfae01f66e15c6a8b63520557ed315bdfe036342fa7a0c537259f10d7a9a",
+ "size": 351,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:7f53453a06ce9b3dccd47b64b7573be15e40c8f1307713500732d75e2f52a3dc",
+ "size": 318,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:674938dfff19b862d50c13a3f151fbbbabfebe2fffdf70da406be9a184d5e983",
+ "size": 138230,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ },
+ {
+ "digest": "sha256:cfb4377670da14c41f816defdf1f4eb3fe532777d8eb3e7db07ce3737a7dddd6",
+ "size": 32029671,
+ "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip"
+ }
+ ]
+ },
+ {
+ "digest": "sha256:8637808eae0e7c2a2875ab58a1c7f74999823afdb67b4570e3bf778eeccc119b",
+ "repository": "testimage",
+ "shape": "image",
+ "mediaType": "application/vnd.docker.distribution.manifest.v2+json",
+ "holdDid": "did:plc:wnbpdx4bn5uts5yd2o2g3gnm",
+ "config": {
+ "digest": "sha256:a40c03cbb81c59bfb0e0887ab0b1859727075da7b9cc576a1cec2c771f38c5fb",
+ "size": 611,
+ "mediaType": "application/vnd.docker.container.image.v1+json"
+ },
+ "layers": [
+ {
+ "digest": "sha256:589002ba0eaed121a1dbf42f6648f29e5be55d5c8a6ee0f8eaa0285cc21ac153",
+ "size": 3861821,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ }
+ ]
+ },
+ {
+ "digest": "sha256:36de2e53ed95fb570ad7f03d9bd6885dfc808e48ee645e5c265a428bcfc3ead2",
+ "repository": "valtest",
+ "shape": "image",
+ "mediaType": "application/vnd.docker.distribution.manifest.v2+json",
+ "holdDid": "did:web:localhost%3A8080",
+ "config": {
+ "digest": "sha256:a54cc65492b33e4532c416a3ae673b9c7ac285a2cb8d41d1748b2b1a58a0812c",
+ "size": 233,
+ "mediaType": "application/vnd.docker.container.image.v1+json"
+ },
+ "layers": [
+ {
+ "digest": "sha256:6094473b2d3638a77d8fdb8fa2909efae92a6a455b72d9938ff579ce839fe234",
+ "size": 165,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ }
+ ]
+ },
+ {
+ "digest": "sha256:49dcaac57bcc1c5b0a72bd26d038706c83c0b05b21325b569f5355394cfd6db0",
+ "repository": "valtest",
+ "shape": "image",
+ "mediaType": "application/vnd.docker.distribution.manifest.v2+json",
+ "holdDid": "did:plc:wnbpdx4bn5uts5yd2o2g3gnm",
+ "config": {
+ "digest": "sha256:84cfb18a07c653ab56ecc2a1e2fdb8c5423ccd50f4a3f05082a478dbc76009ce",
+ "size": 233,
+ "mediaType": "application/vnd.docker.container.image.v1+json"
+ },
+ "layers": [
+ {
+ "digest": "sha256:a32661cd10bc0d2bda45efe0274d422d3f6c6b4daf31d5a7366557a7cf38c601",
+ "size": 173,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ }
+ ]
+ },
+ {
+ "digest": "sha256:b06d1b4aff11971b46797857b3d42236eddac51db1d28d5065c4e7c4b6d8af70",
+ "repository": "valtest",
+ "shape": "image",
+ "mediaType": "application/vnd.docker.distribution.manifest.v2+json",
+ "holdDid": "did:web:localhost%3A8080",
+ "config": {
+ "digest": "sha256:029cee3cd5cc89b70546bad89337a1546032b8495b71757a6864262a82a06aed",
+ "size": 233,
+ "mediaType": "application/vnd.docker.container.image.v1+json"
+ },
+ "layers": [
+ {
+ "digest": "sha256:1c4f419eb23a758557406222757ca3e4a0829ff8fc64ddb092e8e2e2931b1b3a",
+ "size": 170,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ }
+ ]
+ },
+ {
+ "digest": "sha256:b4fd607b1a22a8296df5fedc145ffd89cd97553149772b9ecacbd72199e0a3a7",
+ "repository": "valtest",
+ "shape": "image",
+ "mediaType": "application/vnd.docker.distribution.manifest.v2+json",
+ "holdDid": "did:web:localhost%3A8080",
+ "config": {
+ "digest": "sha256:348315485d6b1c33a9e1016bb8f6d26f14db894924f853ad4d547f5d87970f11",
+ "size": 233,
+ "mediaType": "application/vnd.docker.container.image.v1+json"
+ },
+ "layers": [
+ {
+ "digest": "sha256:26b3ee688b8f836f81d357b1d1a8e71ac2e99a677170e4a7b633c4df5b075510",
+ "size": 165,
+ "mediaType": "application/vnd.docker.image.rootfs.diff.tar.gzip"
+ }
+ ]
+ },
+ {
+ "digest": "sha256:8fbad0ee3b91b66af5428e0674dfbfff9606fca8273013c20512ee06c8fb6334",
+ "repository": "git-summarizer",
+ "shape": "index",
+ "mediaType": "application/vnd.oci.image.index.v1+json",
+ "holdDid": "did:plc:wnbpdx4bn5uts5yd2o2g3gnm",
+ "config": null,
+ "layers": [],
+ "manifests": [
+ {
+ "digest": "sha256:f5c69d77263ef0025535dd25879fb0eef2958929738ac2968787c11e7011e3d6",
+ "size": 1052,
+ "mediaType": "application/vnd.oci.image.manifest.v1+json"
+ },
+ {
+ "digest": "sha256:b6df625b4f311d07d5720ccf3802877eb2a9613b45d573b3f43be237386e4324",
+ "size": 1052,
+ "mediaType": "application/vnd.oci.image.manifest.v1+json"
+ },
+ {
+ "digest": "sha256:b2c92262f645eb07f1d4a4f7dbc709a5c168c27650af90030589a5b7bc1ca862",
+ "size": 566,
+ "mediaType": "application/vnd.oci.image.manifest.v1+json"
+ },
+ {
+ "digest": "sha256:d2aa86dba7c6400b179e188c6f67f10620a9d9224d0d24bc607d12dbcd10ee9d",
+ "size": 566,
+ "mediaType": "application/vnd.oci.image.manifest.v1+json"
+ }
+ ]
+ },
+ {
+ "digest": "sha256:1c0650159575ba04690e57ea928c80e8d04dcaae1fc5a5f53db412e88749a841",
+ "repository": "go-tnyclick",
+ "shape": "index",
+ "mediaType": "application/vnd.docker.distribution.manifest.list.v2+json",
+ "holdDid": "did:plc:wnbpdx4bn5uts5yd2o2g3gnm",
+ "config": null,
+ "layers": [],
+ "manifests": [
+ {
+ "digest": "sha256:bf0327b436432089259af16c31ce7c5206239aa2b1eeeaa9cb0e9d3380741e66",
+ "size": 528,
+ "mediaType": "application/vnd.docker.distribution.manifest.v2+json"
+ },
+ {
+ "digest": "sha256:6eaa5ddb7d063629b6bf5a85a37b4c44e79e083753ede99fdadc1e5d6a528cc6",
+ "size": 528,
+ "mediaType": "application/vnd.docker.distribution.manifest.v2+json"
+ }
+ ]
+ },
+ {
+ "digest": "sha256:086f13321684221f397c7583530a6199d7470b0cbb7ee839588d8b16bd8d9719",
+ "repository": "hsm-secrets-operator",
+ "shape": "index",
+ "mediaType": "application/vnd.docker.distribution.manifest.list.v2+json",
+ "holdDid": "did:plc:wnbpdx4bn5uts5yd2o2g3gnm",
+ "config": null,
+ "layers": [],
+ "manifests": [
+ {
+ "digest": "sha256:a48100dadb92511cdf7e818446125548276c0ed30328ab630648b292b046fb2c",
+ "size": 1786,
+ "mediaType": "application/vnd.docker.distribution.manifest.v2+json"
+ },
+ {
+ "digest": "sha256:5924c964be6b28cda9aa7d3ade479586b9ce4bc33a9cadebb3c63322cc0ecc15",
+ "size": 1786,
+ "mediaType": "application/vnd.docker.distribution.manifest.v2+json"
+ }
+ ]
+ },
+ {
+ "digest": "sha256:5e6a18ae5c6f9f199313fcbe6231d46a74acaa9e24fa9713e623bf54a3afce86",
+ "repository": "hsm-secrets-operator",
+ "shape": "index",
+ "mediaType": "application/vnd.oci.image.index.v1+json",
+ "holdDid": "did:plc:wnbpdx4bn5uts5yd2o2g3gnm",
+ "config": null,
+ "layers": [],
+ "manifests": [
+ {
+ "digest": "sha256:019996323ee44afa7e275318bd570577341a761fb3905be43c3a00dfdf258f17",
+ "size": 1625,
+ "mediaType": "application/vnd.oci.image.manifest.v1+json"
+ },
+ {
+ "digest": "sha256:c1805ebbe44cdd9ac6d284b7ef8205b901a2571ef55a8c0c60ee8fafd75dc139",
+ "size": 1625,
+ "mediaType": "application/vnd.oci.image.manifest.v1+json"
+ },
+ {
+ "digest": "sha256:488f2502332aa6ef82cad82136841c61d7c6458db6ab425aa6e86a06301f5cc1",
+ "size": 564,
+ "mediaType": "application/vnd.oci.image.manifest.v1+json"
+ },
+ {
+ "digest": "sha256:c39092f65e47ae1351f7232e17c36a16171a6987f89737f00b0a13c4f1dd71c7",
+ "size": 564,
+ "mediaType": "application/vnd.oci.image.manifest.v1+json"
+ }
+ ]
+ },
+ {
+ "digest": "sha256:b9d7882fda6c813775eac1864a0b27910673f7e7f99e7b82576287fd5e78eaf1",
+ "repository": "hsm-secrets-operator",
+ "shape": "index",
+ "mediaType": "application/vnd.docker.distribution.manifest.list.v2+json",
+ "holdDid": "did:plc:wnbpdx4bn5uts5yd2o2g3gnm",
+ "config": null,
+ "layers": [],
+ "manifests": [
+ {
+ "digest": "sha256:80140045ec657898f0c5f07a192c632bafd6999ff208d22cded6bbd1b93110d3",
+ "size": 1786,
+ "mediaType": "application/vnd.docker.distribution.manifest.v2+json"
+ },
+ {
+ "digest": "sha256:e4ac9574659da87cbc42ea2ecbce36b8a8c5536eeb48565cc63e9e075580d082",
+ "size": 1786,
+ "mediaType": "application/vnd.docker.distribution.manifest.v2+json"
+ }
+ ]
+ },
+ {
+ "digest": "sha256:9bda98644f897dc6e051177d1477a7d8da106b3ac4c0991e1d8f019270a4e09e",
+ "repository": "loom",
+ "shape": "index",
+ "mediaType": "application/vnd.oci.image.index.v1+json",
+ "holdDid": "did:plc:wnbpdx4bn5uts5yd2o2g3gnm",
+ "config": null,
+ "layers": [],
+ "manifests": [
+ {
+ "digest": "sha256:450c330c1dc051d27015308b93925f74142d7736d5298fdd71a78a08563a6a09",
+ "size": 3896,
+ "mediaType": "application/vnd.oci.image.manifest.v1+json"
+ },
+ {
+ "digest": "sha256:b495605bcafd1e7d3e911c49df242265f4d8f68379d29b48524e460354cbb4bc",
+ "size": 3896,
+ "mediaType": "application/vnd.oci.image.manifest.v1+json"
+ },
+ {
+ "digest": "sha256:f330312e825cd1de8406ddc5eb09a566cd47ac1b9aa543bd94636a332b27f72e",
+ "size": 564,
+ "mediaType": "application/vnd.oci.image.manifest.v1+json"
+ },
+ {
+ "digest": "sha256:13f458be65ba8c912d4e6c7b05275be4a16c0e51a2e8e36ca4983d2f2c2d0f5c",
+ "size": 564,
+ "mediaType": "application/vnd.oci.image.manifest.v1+json"
+ }
+ ]
+ },
+ {
+ "digest": "sha256:13bdf01e8d94a37512d7fc6310b1360b29aed005b401436145633b9557fbf6a8",
+ "repository": "loom-runner",
+ "shape": "index",
+ "mediaType": "application/vnd.oci.image.index.v1+json",
+ "holdDid": "did:plc:wnbpdx4bn5uts5yd2o2g3gnm",
+ "config": null,
+ "layers": [],
+ "manifests": [
+ {
+ "digest": "sha256:1178a001d60a5a6c7ac4cba5a8fad190405510982ce75dcea1ff44be39e004a9",
+ "size": 864,
+ "mediaType": "application/vnd.oci.image.manifest.v1+json"
+ },
+ {
+ "digest": "sha256:1fd4b800e2605202f4844b38732cf6b0732a2233d067d5b0d7249ff3a37f02f9",
+ "size": 864,
+ "mediaType": "application/vnd.oci.image.manifest.v1+json"
+ },
+ {
+ "digest": "sha256:93670b296ef1774f53e088b0a3cec6fb737ecf7fe6299d8cd0c2210759a38276",
+ "size": 566,
+ "mediaType": "application/vnd.oci.image.manifest.v1+json"
+ },
+ {
+ "digest": "sha256:21eb1f1fd68ca89fa06da8a054e2a366b1b4103167077f42f7258851ee38f0cd",
+ "size": 566,
+ "mediaType": "application/vnd.oci.image.manifest.v1+json"
+ }
+ ]
+ }
+ ]
+}
\ No newline at end of file
diff --git a/scanner/internal/mockhold/testdata/fetch-blobs.sh b/scanner/internal/mockhold/testdata/fetch-blobs.sh
new file mode 100755
index 0000000..3f06f7b
--- /dev/null
+++ b/scanner/internal/mockhold/testdata/fetch-blobs.sh
@@ -0,0 +1,66 @@
+#!/usr/bin/env bash
+# Pull real image blobs into OCI layouts that mockhold.OCILayout can serve.
+#
+# The layouts land in testdata/blobs// and are gitignored: they are
+# megabytes of container layers, and every test that needs them skips cleanly
+# when they are absent. Descriptor-only scenarios (disconnect, reconnect,
+# skip classification, oversize) need nothing from here.
+#
+# Why skopeo rather than the scanner's own download path: `skopeo copy ...
+# oci:` writes blobs/sha256/, keyed by exactly the digest the scan
+# job references. That is the same layout buildOCILayout reconstructs at scan
+# time, so a pulled image drops straight in with no rewriting.
+#
+# Auth: the registry host must be mapped to the credential helper. seamark.cr
+# is NOT a registry you can log into by default (its handle appears under the
+# atcr.io appview, which is a different thing); buoy.cr works out of the box
+# if `docker-credential-atcr status` lists an account for it. The script writes
+# its own authfile so it never touches ~/.docker/config.json.
+#
+# Usage:
+# ./fetch-blobs.sh # pull the default fixture set
+# REGISTRY=atcr.io ./fetch-blobs.sh
+set -euo pipefail
+
+REGISTRY="${REGISTRY:-buoy.cr}"
+ACCOUNT="${ACCOUNT:-evan.jarrett.net}"
+DIR="$(cd "$(dirname "$0")" && pwd)/blobs"
+
+# name|repository|manifest digest
+# Digests come from corpus.json. Keep these small: the point is real layer
+# bytes for Syft to catalog, not coverage of every image on the hold.
+FIXTURES=(
+ "hsm-secrets-operator|hsm-secrets-operator|sha256:1cfa4e2b09e127b9c4ed43578d3f3c18e7d44ea47b9ea98475c0cbe9086525f8"
+)
+
+command -v skopeo >/dev/null || { echo "skopeo not found"; exit 1; }
+
+AUTHFILE=$(mktemp); trap 'rm -f "$AUTHFILE"' EXIT
+printf '{"credHelpers":{"%s":"atcr"}}\n' "$REGISTRY" > "$AUTHFILE"
+
+mkdir -p "$DIR"
+for entry in "${FIXTURES[@]}"; do
+ IFS='|' read -r name repo digest <<< "$entry"
+ dest="$DIR/$name"
+
+ if [ -f "$dest/oci-layout" ]; then
+ echo "→ $name already present, skipping"
+ continue
+ fi
+
+ echo "→ pulling ${REGISTRY}/${ACCOUNT}/${repo}@${digest:0:19}..."
+ rm -rf "$dest"
+ if ! skopeo copy --authfile "$AUTHFILE" \
+ "docker://${REGISTRY}/${ACCOUNT}/${repo}@${digest}" \
+ "oci:${dest}:img"; then
+ echo " FAILED. If this is an auth error, check:"
+ echo " docker-credential-atcr status"
+ echo " and confirm an account is configured for ${REGISTRY}."
+ rm -rf "$dest"
+ exit 1
+ fi
+done
+
+echo
+echo "→ layouts in $DIR:"
+du -sh "$DIR"/*/ 2>/dev/null || true
diff --git a/scanner/internal/mockhold/testdata/fetch-corpus.sh b/scanner/internal/mockhold/testdata/fetch-corpus.sh
new file mode 100755
index 0000000..0f17aa9
--- /dev/null
+++ b/scanner/internal/mockhold/testdata/fetch-corpus.sh
@@ -0,0 +1,104 @@
+#!/usr/bin/env bash
+# Regenerate corpus.json: real io.atcr.manifest records from a live PDS.
+#
+# Manifest records ARE the scan job payload. The config and layer descriptors
+# the hold sends to the scanner are stored verbatim in the user's PDS, and
+# com.atproto.repo.listRecords serves them without authentication. So the
+# corpus needs no registry credentials and no blob bytes, only a DID.
+#
+# Blob bytes are a separate concern; see fetch-blobs.sh.
+#
+# Usage:
+# ./fetch-corpus.sh # default DID below
+# ./fetch-corpus.sh did:plc:xxxx # some other repo
+set -euo pipefail
+
+DID="${1:-did:plc:pddp4xt5lgnv2qsegbzzs4xg}"
+OUT="$(cd "$(dirname "$0")" && pwd)/corpus.json"
+
+# Resolve the DID's PDS from its DID document.
+PDS=$(curl -sS --max-time 20 "https://plc.directory/${DID}" \
+ | python3 -c 'import json,sys; d=json.load(sys.stdin); print(next(s["serviceEndpoint"] for s in d["service"] if s["id"]=="#atproto_pds"))')
+HANDLE=$(curl -sS --max-time 20 "https://plc.directory/${DID}" \
+ | python3 -c 'import json,sys; d=json.load(sys.stdin); print(d["alsoKnownAs"][0].replace("at://",""))')
+
+echo "→ repo ${DID} (${HANDLE})"
+echo "→ pds ${PDS}"
+
+TMP=$(mktemp -d); trap 'rm -rf "$TMP"' EXIT
+: > "$TMP/all.jsonl"
+cursor=""
+while :; do
+ url="${PDS}/xrpc/com.atproto.repo.listRecords?repo=${DID}&collection=io.atcr.manifest&limit=100"
+ [ -n "$cursor" ] && url="${url}&cursor=${cursor}"
+ curl -sS --max-time 30 "$url" > "$TMP/page.json"
+ cursor=$(python3 -c '
+import json,sys
+d=json.load(open(sys.argv[1]))
+with open(sys.argv[2],"a") as f:
+ for r in d.get("records",[]): f.write(json.dumps(r["value"])+"\n")
+print(d.get("cursor",""))' "$TMP/page.json" "$TMP/all.jsonl")
+ [ -z "$cursor" ] && break
+done
+
+python3 - "$TMP/all.jsonl" "$OUT" "$DID" "$HANDLE" <<'PY'
+import json, sys, collections
+src, out, did, handle = sys.argv[1:5]
+
+records = [json.loads(l) for l in open(src)]
+seen, uniq = set(), []
+for v in records:
+ if v.get("digest") in seen:
+ continue
+ seen.add(v.get("digest"))
+ uniq.append(v)
+
+def shape(v):
+ layers = v.get("layers") or []
+ if v.get("subject"):
+ return "referrer"
+ if not layers and v.get("manifests"):
+ return "index"
+ if "helm" in (v.get("config") or {}).get("mediaType", ""):
+ return "helm"
+ media = {l.get("mediaType", "") for l in layers}
+ if any("in-toto" in m or "dsse" in m for m in media):
+ return "attestation"
+ if any("tar" in m for m in media):
+ return "image"
+ return "other"
+
+def desc(d):
+ return None if not d else {
+ "digest": d.get("digest"), "size": d.get("size"), "mediaType": d.get("mediaType")}
+
+entries = []
+for v in sorted(uniq, key=lambda v: (shape(v), v.get("repository") or "", v["digest"])):
+ e = {
+ "digest": v["digest"],
+ "repository": v.get("repository"),
+ "shape": shape(v),
+ "mediaType": v.get("mediaType"),
+ "holdDid": v.get("holdDid"),
+ "config": desc(v.get("config")),
+ "layers": [desc(l) for l in (v.get("layers") or [])],
+ }
+ if v.get("manifests"):
+ e["manifests"] = [desc(m) for m in v["manifests"]]
+ if v.get("subject"):
+ e["subject"] = desc(v["subject"])
+ entries.append(e)
+
+json.dump({
+ "_comment": ("Real io.atcr.manifest records, fetched anonymously via "
+ "com.atproto.repo.listRecords. Regenerate with "
+ "testdata/fetch-corpus.sh. Descriptors only, no blob bytes."),
+ "sourceRepo": did,
+ "sourceHandle": handle,
+ "manifests": entries,
+}, open(out, "w"), indent=2)
+
+print(f"→ wrote {out}: {len(entries)} manifests")
+for s, n in sorted(collections.Counter(e["shape"] for e in entries).items()):
+ print(f" {s:<12} {n}")
+PY
diff --git a/scanner/internal/mockhold/testdata/fetch-vulndb.sh b/scanner/internal/mockhold/testdata/fetch-vulndb.sh
new file mode 100755
index 0000000..165138b
--- /dev/null
+++ b/scanner/internal/mockhold/testdata/fetch-vulndb.sh
@@ -0,0 +1,225 @@
+#!/usr/bin/env bash
+# Download the real Grype vulnerability database, plus the pinned images it is
+# matched against, so scan REPORTS can be verified and not just the pipeline
+# mechanics.
+#
+# Everything else in this testdata directory is descriptors and layer bytes.
+# This is the other half: without a real database, Grype's matcher, its
+# configuration, and every vulnerability count the scanner publishes are
+# entirely unexercised. The e2e scenarios in vulnreport_test.go are the ones
+# that need it, and they skip cleanly when it is absent.
+#
+# DISK COST: ~2.0 GB in testdata/vulndb/6/vulnerability.db (uncompressed
+# SQLite, mmap'd at match time), plus a transient ~350 MB archive and its
+# decompression scratch under testdata/vulndb/.tmp during the download. The
+# image fixtures add ~32 MB. Budget 3 GB free before running. Both directories
+# are gitignored.
+#
+# TIME: about 30-60s on a fast link; the download itself is a few hundred MB.
+#
+# FRESHNESS: grypeDBConfig (scanner/internal/scan/grype.go) sets
+# ValidateAge:true with MaxAllowedBuiltAge of 14 days, so a cached database
+# goes stale and then fails to load outright. Re-running this script refreshes
+# it: the same Grype curator the scanner uses checks the upstream listing and
+# downloads a newer build when the local one is superseded. Re-running on a
+# fresh database is a cheap no-op (one listing fetch).
+#
+# Usage:
+# ./fetch-vulndb.sh # download/refresh the database and the images
+# ./fetch-vulndb.sh --force # delete and re-download the database
+# ./fetch-vulndb.sh --status # report the cached database's build age only
+# ./fetch-vulndb.sh --db-only # skip the image pulls
+#
+# Why this shells into Go rather than curling the archive: the curator writes
+# import.json alongside the database during activation, and the scanner's
+# installation.Config sets ValidateChecksum:true. A hand-extracted archive is
+# missing that file and fails to load with "no import metadata file at ...",
+# which is a real production failure mode (SCANNER_BUGS.md section 6, finding
+# 6) and not something a fixture should reproduce by accident. Calling
+# grype.LoadVulnerabilityDB with the scanner's exact config is the only way to
+# land bytes on disk in the state the scanner expects to find them in.
+set -euo pipefail
+
+HERE="$(cd "$(dirname "$0")" && pwd)"
+MODULE_ROOT="$(cd "$HERE/../../.." && pwd)" # scanner/
+DB_DIR="${GRYPE_DB_DIR:-$HERE/vulndb}"
+TMP_DIR="$DB_DIR/.tmp"
+
+BLOB_DIR="$HERE/blobs"
+
+# name|reference. Pinned, public, and small. Each earns its place:
+#
+# vuln-alpine310 alpine:3.10, end of life since 2021. The guarantee that a
+# scan finds something, whatever the feed says today.
+# vuln-debian11 debian:11-slim. Debian's security tracker rates advisories
+# "Negligible" and "Unknown", which Alpine's does not, and
+# the severity-bucket consistency check gives a different
+# answer on each. One distro is not a sample.
+# vuln-static distroless static. The control: a handful of dpkg entries,
+# no shell, no libc. A matcher inventing results has nothing
+# here to invent them from.
+#
+# Tags rather than digests: these are the upstream's own frozen tags for
+# released images, and pinning a digest here would mean an unpullable fixture
+# the first time a registry re-pushes a manifest list.
+IMAGE_FIXTURES=(
+ "vuln-alpine310|docker://docker.io/library/alpine:3.10"
+ "vuln-debian11|docker://docker.io/library/debian:11-slim"
+ "vuln-static|docker://gcr.io/distroless/static-debian12:latest"
+)
+
+FORCE=0
+STATUS_ONLY=0
+DB_ONLY=0
+for arg in "$@"; do
+ case "$arg" in
+ --force) FORCE=1 ;;
+ --status) STATUS_ONLY=1 ;;
+ --db-only) DB_ONLY=1 ;;
+ *) echo "unknown argument: $arg" >&2; exit 2 ;;
+ esac
+done
+
+command -v go >/dev/null || { echo "go not found" >&2; exit 1; }
+
+if [ "$STATUS_ONLY" = 1 ]; then
+ if [ -f "$DB_DIR/6/import.json" ]; then
+ echo "→ cached database at $DB_DIR"
+ # import.json carries no build timestamp of its own; the archive name in
+ # its "source" field does, and that is the value Grype's ValidateAge check
+ # runs against.
+ built=$(grep -o '_[0-9]\{4\}-[0-9-]*T[0-9:]*Z_' "$DB_DIR/6/import.json" | head -1 | tr -d _)
+ echo " built: ${built:-unknown}"
+ if [ -n "$built" ]; then
+ age=$(( ( $(date -u +%s) - $(date -u -d "$built" +%s) ) / 86400 ))
+ echo " age: ${age}d (Grype refuses to load one built more than 14d ago)"
+ fi
+ du -sh "$DB_DIR"
+ else
+ echo "→ no database at $DB_DIR; run this script with no arguments"
+ fi
+ echo "→ image fixtures:"
+ for entry in "${IMAGE_FIXTURES[@]}"; do
+ name="${entry%%|*}"
+ if [ -f "$BLOB_DIR/$name/oci-layout" ]; then
+ echo " present: $name ($(du -sh "$BLOB_DIR/$name" | cut -f1))"
+ else
+ echo " missing: $name"
+ fi
+ done
+ exit 0
+fi
+
+if [ "$FORCE" = 1 ]; then
+ echo "→ removing $DB_DIR"
+ rm -rf "$DB_DIR"
+fi
+
+mkdir -p "$TMP_DIR"
+
+# The loader is written to a scratch package under the module so `go run`
+# resolves grype from scanner/go.mod — the same version the scanner links, so
+# the on-disk schema is by construction the one it can open.
+RUNNER="$MODULE_ROOT/.fetch-vulndb.$$"
+mkdir -p "$RUNNER"
+trap 'rm -rf "$RUNNER"' EXIT
+
+cat > "$RUNNER/main.go" <<'GO'
+package main
+
+import (
+ "fmt"
+ "os"
+ "time"
+
+ "github.com/anchore/grype/grype"
+ "github.com/anchore/grype/grype/db/v6/distribution"
+ "github.com/anchore/grype/grype/db/v6/installation"
+)
+
+// Mirrors grypeDBConfig in scanner/internal/scan/grype.go. Keep in sync: a
+// database fetched under different settings is not the fixture the scanner
+// would have produced for itself.
+func main() {
+ root := os.Args[1]
+ if err := os.MkdirAll(root, 0o755); err != nil {
+ fmt.Fprintln(os.Stderr, err)
+ os.Exit(1)
+ }
+ start := time.Now()
+ provider, status, err := grype.LoadVulnerabilityDB(
+ distribution.DefaultConfig(),
+ installation.Config{
+ DBRootDir: root,
+ ValidateAge: true,
+ ValidateChecksum: true,
+ MaxAllowedBuiltAge: 14 * 24 * time.Hour,
+ },
+ true, // update: check upstream, download when superseded, then open
+ )
+ if err != nil {
+ fmt.Fprintf(os.Stderr, "grype database load failed: %v\n", err)
+ os.Exit(1)
+ }
+ defer provider.Close()
+ fmt.Printf("built=%s age=%s schema=%s\npath=%s\nelapsed=%s\n",
+ status.Built.Format(time.RFC3339),
+ time.Since(status.Built).Round(time.Minute),
+ status.SchemaVersion,
+ status.Path,
+ time.Since(start).Round(time.Second))
+}
+GO
+
+echo "→ fetching Grype database into $DB_DIR"
+(
+ cd "$MODULE_ROOT"
+ # Same reason WorkerPool.Start exports TMPDIR in production: go-getter's zstd
+ # decompression is 1 GB+ and must not land on a small tmpfs.
+ TMPDIR="$TMP_DIR" go run "./$(basename "$RUNNER")" "$DB_DIR"
+)
+
+rm -rf "$TMP_DIR"
+
+if [ "$DB_ONLY" = 0 ]; then
+ if ! command -v skopeo >/dev/null; then
+ echo "→ skopeo not found; skipping the image fixtures. Install it, or pull them by hand:" >&2
+ for entry in "${IMAGE_FIXTURES[@]}"; do
+ echo " skopeo copy --override-os linux --override-arch amd64 ${entry#*|} oci:$BLOB_DIR/${entry%%|*}:img" >&2
+ done
+ else
+ mkdir -p "$BLOB_DIR"
+ for entry in "${IMAGE_FIXTURES[@]}"; do
+ name="${entry%%|*}"
+ ref="${entry#*|}"
+ dest="$BLOB_DIR/$name"
+ if [ -f "$dest/oci-layout" ]; then
+ echo "→ $name already present, skipping"
+ continue
+ fi
+ echo "→ pulling $ref"
+ rm -rf "$dest"
+ # Pin the platform: an unqualified pull of a manifest list on a non-amd64
+ # host would produce a fixture whose findings differ from everyone
+ # else's, and the scenarios compare against no golden file precisely so
+ # that stays a non-issue.
+ if ! skopeo copy --override-os linux --override-arch amd64 "$ref" "oci:${dest}:img"; then
+ echo " FAILED to pull $ref" >&2
+ rm -rf "$dest"
+ exit 1
+ fi
+ done
+ fi
+fi
+
+echo
+echo "→ on disk:"
+du -sh "$DB_DIR"
+for entry in "${IMAGE_FIXTURES[@]}"; do
+ d="$BLOB_DIR/${entry%%|*}"
+ [ -d "$d" ] && du -sh "$d"
+done
+echo
+echo "Run the scenarios that use it with:"
+echo " cd $MODULE_ROOT && ATCR_SCANNER_VULNDB=1 go test ./internal/e2e -run TestVulnDB -v"
+echo "Add ATCR_SCANNER_PERF=1 for the Grype cost measurement."
diff --git a/scanner/internal/scan/extractor.go b/scanner/internal/scan/extractor.go
index b36d281..19597a1 100644
--- a/scanner/internal/scan/extractor.go
+++ b/scanner/internal/scan/extractor.go
@@ -3,6 +3,7 @@ package scan
import (
"crypto/sha256"
"encoding/json"
+ "errors"
"fmt"
"log/slog"
"os"
@@ -46,7 +47,15 @@ type ociIndex struct {
// ├──
// ├──
// └── ...
-func buildOCILayout(job *scanner.ScanJob, tmpDir, secret string) (string, func(), error) {
+//
+// Every blob is verified against its descriptor as it streams, and every
+// digest is validated before it becomes a path, so the layout describes bytes
+// that are on disk and hash to the names they are filed under.
+//
+// maxBytes is the ceiling on the total transferred for this job; zero or less
+// means unlimited. It is spent down blob by blob, so it bounds the real bytes
+// on disk rather than the numbers the manifest record claims.
+func buildOCILayout(job *scanner.ScanJob, tmpDir, secret string, maxBytes int64) (string, func(), error) {
scanDir, err := os.MkdirTemp(tmpDir, "scan-*")
if err != nil {
return "", nil, fmt.Errorf("failed to create temp directory: %w", err)
@@ -64,58 +73,54 @@ func buildOCILayout(job *scanner.ScanJob, tmpDir, secret string) (string, func()
return "", nil, fmt.Errorf("failed to create blobs directory: %w", err)
}
- // Download config blob
if job.Config.Digest == "" {
cleanup()
return "", nil, fmt.Errorf("config blob has empty digest, cannot download")
}
- slog.Info("Downloading config blob", "digest", job.Config.Digest)
- if err := downloadBlob(job, job.Config.Digest, blobsDir, secret); err != nil {
- cleanup()
- return "", nil, fmt.Errorf("failed to download config blob: %w", err)
- }
- // Download layer blobs (no extraction — kept compressed)
- for i, layer := range job.Layers {
- if layer.Digest == "" {
- slog.Warn("Skipping layer with empty digest", "index", i)
- continue
- }
- // Skip non-tar layers (cosign signatures, in-toto attestations, etc.)
- if layer.MediaType != "" && !strings.Contains(layer.MediaType, "tar") {
- slog.Info("Skipping non-tar layer", "index", i, "digest", layer.Digest, "mediaType", layer.MediaType)
- continue
- }
- slog.Info("Downloading layer", "index", i, "digest", layer.Digest, "size", layer.Size, "mediaType", layer.MediaType)
- if err := downloadBlob(job, layer.Digest, blobsDir, secret); err != nil {
- cleanup()
- return "", nil, fmt.Errorf("failed to download layer %d: %w", i, err)
- }
+ // Download every referenced blob and describe it in the same pass, so the
+ // layout can only ever list a blob that arrived and was verified, at the
+ // size it actually arrived in. Verification has already established that a
+ // declared size, where the record gave one, matches; recording the
+ // measured figure keeps the layout honest for the descriptors that
+ // declared none.
+ remaining := int64(-1) // -1 is unbounded
+ if maxBytes > 0 {
+ remaining = maxBytes
}
-
- // Build OCI manifest from job descriptors
manifest := ociManifest{
SchemaVersion: 2,
MediaType: "application/vnd.oci.image.manifest.v1+json",
- Config: ociDescriptor{
- MediaType: defaultMediaType(job.Config.MediaType, "application/vnd.oci.image.config.v1+json"),
- Digest: job.Config.Digest,
- Size: job.Config.Size,
- },
- Layers: make([]ociDescriptor, 0, len(job.Layers)),
+ Layers: make([]ociDescriptor, 0, len(job.Layers)),
}
- for _, layer := range job.Layers {
- if layer.Digest == "" {
+ for _, ref := range referencedBlobs(job) {
+ digest, err := scanner.ParseDigest(ref.Descriptor.Digest)
+ if err != nil {
+ cleanup()
+ return "", nil, &SkipError{Reason: fmt.Sprintf("%s: %v", ref.what(), err)}
+ }
+ slog.Info("Downloading blob", "blob", ref.what(), "digest", digest,
+ "declaredSize", ref.Descriptor.Size, "mediaType", ref.Descriptor.MediaType)
+
+ n, err := downloadBlob(job, digest, ref.Descriptor.Size, remaining, blobsDir, secret)
+ if err != nil {
+ cleanup()
+ return "", nil, blobFailure(ref.what(), err)
+ }
+ if remaining >= 0 {
+ remaining -= n
+ }
+
+ d := ociDescriptor{
+ MediaType: defaultMediaType(ref.Descriptor.MediaType, ref.defaultMediaType()),
+ Digest: digest.String(),
+ Size: n,
+ }
+ if ref.isConfig() {
+ manifest.Config = d
continue
}
- if layer.MediaType != "" && !strings.Contains(layer.MediaType, "tar") {
- continue
- }
- manifest.Layers = append(manifest.Layers, ociDescriptor{
- MediaType: defaultMediaType(layer.MediaType, "application/vnd.oci.image.layer.v1.tar+gzip"),
- Digest: layer.Digest,
- Size: layer.Size,
- })
+ manifest.Layers = append(manifest.Layers, d)
}
// Write manifest blob
@@ -168,24 +173,95 @@ func buildOCILayout(job *scanner.ScanJob, tmpDir, secret string) (string, func()
return scanDir, cleanup, nil
}
-// downloadBlob downloads a blob by digest to the blobs directory using presigned URLs.
-func downloadBlob(job *scanner.ScanJob, digest, blobsDir, secret string) error {
- hex := digestHex(digest)
- destPath := filepath.Join(blobsDir, hex)
+// blobRef is one blob the layout will contain: the descriptor that named it,
+// plus where it sits in the record so an error can say which blob it means.
+type blobRef struct {
+ Index int // position in job.Layers; configIndex for the config blob
+ Descriptor scanner.BlobDescriptor
+}
+
+const configIndex = -1
+
+func (r blobRef) isConfig() bool { return r.Index == configIndex }
+
+func (r blobRef) what() string {
+ if r.isConfig() {
+ return "config blob"
+ }
+ return fmt.Sprintf("layer %d", r.Index)
+}
+
+func (r blobRef) defaultMediaType() string {
+ if r.isConfig() {
+ return "application/vnd.oci.image.config.v1+json"
+ }
+ return "application/vnd.oci.image.layer.v1.tar+gzip"
+}
+
+// referencedBlobs returns every blob a scan of this job touches, in the order
+// it touches them: the config first, then the layers that are not dropped.
+//
+// This is the single definition of "which blobs does this job reference".
+// Digest validation, the download loop, the byte budget and the layout
+// manifest all walk this list, so a layer dropped here is dropped everywhere
+// and a layer kept here is one that has been checked.
+func referencedBlobs(job *scanner.ScanJob) []blobRef {
+ refs := make([]blobRef, 0, len(job.Layers)+1)
+ if job.Config.Digest != "" {
+ refs = append(refs, blobRef{Index: configIndex, Descriptor: job.Config})
+ }
+ for i, layer := range job.Layers {
+ if layer.Digest == "" {
+ continue
+ }
+ // Non-tar layers (cosign signatures, in-toto attestations) are not
+ // something Syft can read, so they are never fetched or listed.
+ if layer.MediaType != "" && !strings.Contains(layer.MediaType, "tar") {
+ continue
+ }
+ refs = append(refs, blobRef{Index: i, Descriptor: layer})
+ }
+ return refs
+}
+
+// blobFailure labels a download error, converting the ones that can never
+// succeed on a retry into a SkipError.
+//
+// The hold's stale-scan loop re-offers "error" records on every pass and never
+// re-offers "skipped" ones, so a failure decided by the record itself or by
+// bytes already stored belongs on the skip side. A transport fault (a 5xx, a
+// dropped connection, an expired presigned URL) stays an error, because those
+// really can succeed next time.
+func blobFailure(what string, err error) error {
+ if errors.Is(err, client.ErrBlobCorrupt) || errors.Is(err, client.ErrBlobTooLarge) {
+ return &SkipError{Reason: fmt.Sprintf("%s: %v", what, err)}
+ }
+ return fmt.Errorf("failed to download %s: %w", what, err)
+}
+
+// downloadBlob fetches one validated blob into the blobs directory and returns
+// how many bytes arrived. maxBytes is the remaining job budget, negative for
+// unbounded.
+func downloadBlob(job *scanner.ScanJob, digest scanner.Digest, declaredSize, maxBytes int64, blobsDir, secret string) (int64, error) {
+ destPath := filepath.Join(blobsDir, digest.Hex)
+
+ // The invariant, asserted rather than assumed. ParseDigest has already
+ // constrained Hex to [0-9a-f], so this cannot fire; it is here so that any
+ // future path that reaches os.Create with something less constrained fails
+ // loudly instead of writing outside the scan directory.
+ if filepath.Dir(filepath.Clean(destPath)) != filepath.Clean(blobsDir) {
+ return 0, fmt.Errorf("refusing to write blob %s outside %s", digest, blobsDir)
+ }
presignedURL, err := client.GetBlobPresignedURL(job.HoldEndpoint, job.HoldDID, digest, secret)
if err != nil {
- return fmt.Errorf("failed to get presigned URL for %s: %w", digest, err)
+ return 0, fmt.Errorf("failed to get presigned URL for %s: %w", digest, err)
}
- return client.DownloadBlob(presignedURL, destPath)
-}
-
-// digestHex extracts the hex portion from a digest string (e.g., "sha256:abc123" → "abc123").
-func digestHex(digest string) string {
- if _, hex, ok := strings.Cut(digest, ":"); ok {
- return hex
- }
- return digest
+ return client.DownloadBlob(presignedURL, destPath, client.BlobExpectation{
+ Digest: digest,
+ DeclaredSize: declaredSize,
+ MaxBytes: maxBytes,
+ })
}
func defaultMediaType(mediaType, fallback string) string {
diff --git a/scanner/internal/scan/grype.go b/scanner/internal/scan/grype.go
index 956b60f..f9570b5 100644
--- a/scanner/internal/scan/grype.go
+++ b/scanner/internal/scan/grype.go
@@ -4,9 +4,11 @@ import (
"context"
"crypto/sha256"
"encoding/json"
+ "errors"
"fmt"
"log/slog"
"os"
+ "strings"
"sync"
"sync/atomic"
"time"
@@ -36,6 +38,7 @@ var (
vulnDBLock sync.RWMutex
vulnDBBuilt time.Time // build timestamp of the current vulnDB (not load time)
vulnDBAttempt time.Time // last time we attempted a (re)load, success or fail
+ vulnDBLastErr error // why the last attempt failed, replayed to throttled callers
vulnDBScans atomic.Int64 // scan counter for periodic reload
)
@@ -51,6 +54,24 @@ const vulnDBRefreshAge = 7 * 24 * time.Hour
// every worker serializing through its own download timeout.
const vulnDBRetryBackoff = 30 * time.Minute
+// vulnDBColdRetryBackoff throttles reload attempts when there is no usable
+// provider at all. Nothing can be scanned in that state, so it is worth
+// escaping quickly, but a scanner in it must not run one full download attempt
+// per queued scan while the hold's stale loop keeps re-queueing the failures.
+const vulnDBColdRetryBackoff = 2 * time.Minute
+
+// vulnDBMaxServeAge is the ceiling on the serve-the-old-DB fallback. Past it
+// the provider is refused rather than served, because a scan against a database
+// this old reports "0 critical" with a confidence it has not earned. It matches
+// the MaxAllowedBuiltAge that grypeDBConfig asks Grype to enforce on disk; the
+// in-memory provider never re-checks its own age, so the ceiling has to live
+// here.
+const vulnDBMaxServeAge = 14 * 24 * time.Hour
+
+// vulnDBReloadEvery is the periodic close-and-reopen interval, in scans, that
+// flushes SQLite's page cache and mmap region.
+const vulnDBReloadEvery = 50
+
// scanVulnerabilities scans an SBOM for vulnerabilities using Grype
func scanVulnerabilities(ctx context.Context, s *sbom.SBOM, vulnDBPath string) ([]byte, string, scanner.VulnerabilitySummary, error) {
slog.Info("Scanning for vulnerabilities with Grype")
@@ -166,51 +187,150 @@ func grypeDBConfig(vulnDBPath string) (distribution.Config, installation.Config)
// real download.
var loadVulnDB = grype.LoadVulnerabilityDB
-// loadVulnDatabase loads the Grype vulnerability database with caching and
-// automatic refresh. The cached DB is returned if loaded less than
-// vulnDBRefreshAge ago. On a stale or missing DB, Grype downloads a fresh copy
-// in the same call (update=true) — a single curator handles everything so
-// there is no chance of a double-curator update+load seeing different state.
+// vulnDBDecision is what a snapshot of the database state says the current
+// caller should do.
+type vulnDBDecision int
+
+const (
+ vulnDBServe vulnDBDecision = iota // the provider in hand is usable and no reload is due
+ vulnDBReload // run a (re)load attempt now
+ vulnDBUnusable // nothing usable to serve, and the retry backoff has not elapsed
+)
+
+// vulnDBUsable reports whether a loaded provider may still be scanned against.
+// The provider itself cannot answer this: v6.NewVulnerabilityProvider closes
+// over a reader and nothing else, with no build timestamp and no age check on
+// any query path, which is why a stale-but-loaded DB keeps working and why the
+// ceiling has to be imposed out here.
+func vulnDBUsable(db vulnerability.Provider, built time.Time, now time.Time) bool {
+ return db != nil && now.Sub(built) < vulnDBMaxServeAge
+}
+
+// vulnDBDecide is the whole freshness/backoff policy, as a pure function of a
+// state snapshot. Both the read-lock fast path and the write-lock double-check
+// run it, so they cannot disagree — the previous version had a fast path that
+// returned on exactly the condition the write-lock branch required, which is
+// what made the periodic reload unreachable.
+func vulnDBDecide(db vulnerability.Provider, built, attempt time.Time, now time.Time) vulnDBDecision {
+ if vulnDBUsable(db, built, now) {
+ // Fresh, or stale but inside the retry backoff: either way this
+ // provider is what the scan will use.
+ if now.Sub(built) < vulnDBRefreshAge || now.Sub(attempt) < vulnDBRetryBackoff {
+ return vulnDBServe
+ }
+ return vulnDBReload
+ }
+ // No usable provider. There is no fallback to protect, so the throttle is
+ // the much shorter cold backoff — but it is still a throttle. Without one,
+ // every queued scan runs a complete load attempt (a 30s upstream check plus
+ // up to a 300s download) under the exclusive lock, and the hold's stale-scan
+ // loop feeds the failures straight back in.
+ if !attempt.IsZero() && now.Sub(attempt) < vulnDBColdRetryBackoff {
+ return vulnDBUnusable
+ }
+ return vulnDBReload
+}
+
+// vulnDBUnavailableErr explains a refusal to a caller that was throttled out of
+// attempting its own load, so the hold sees why rather than a bare "no database".
+func vulnDBUnavailableErr(db vulnerability.Provider, built time.Time, lastErr error) error {
+ if lastErr == nil {
+ lastErr = errors.New("no database loaded")
+ }
+ if db != nil {
+ // A provider is in hand but past the ceiling. This is the case worth
+ // naming precisely: it is the difference between "we cannot scan" and
+ // "we scanned against a four-month-old database and found nothing".
+ return fmt.Errorf("vulnerability database is too old to trust (built %s ago, max %s) and the refresh failed: %w",
+ time.Since(built).Round(time.Minute), vulnDBMaxServeAge, lastErr)
+ }
+ return fmt.Errorf("vulnerability database unavailable, retrying in at most %s: %w", vulnDBColdRetryBackoff, lastErr)
+}
+
+// vulnDBCorrupt reports whether a load failure is one that deleting the
+// database directory can fix. An expired database self-heals on its own
+// (curator.Update nils the current description when validateAge fails, so
+// isSupersededBy is unconditionally true and it always re-downloads). A
+// checksum or import-metadata failure does not: the on-disk description stays
+// valid, so an update only lands if the upstream has something strictly newer,
+// and until then every scan fails against a database nothing ever deletes.
+func vulnDBCorrupt(err error) bool {
+ if err == nil {
+ return false
+ }
+ msg := err.Error()
+ return strings.Contains(msg, "import metadata") || strings.Contains(msg, "bad db checksum")
+}
+
+// loadVulnDatabase returns the vulnerability database to scan against,
+// refreshing it when due.
+//
+// The state machine, in one place:
+//
+// - a usable provider that is fresh, or stale but inside the retry backoff,
+// is served from the read lock;
+// - every vulnDBReloadEvery scans that path instead falls through to a
+// periodic reload, which loads the replacement first and swaps second, so a
+// failed reload leaves the working provider exactly where it was;
+// - a stale provider past the backoff triggers a refresh, and if that refresh
+// fails the provider keeps serving until it crosses vulnDBMaxServeAge;
+// - past that ceiling, and on a cold start with nothing loaded, the caller
+// gets an error rather than a silent success — but attempts are throttled
+// on vulnDBColdRetryBackoff, so a scanner in that state probes occasionally
+// instead of once per scan, and recovers on its own when the upstream
+// returns.
func loadVulnDatabase(ctx context.Context, vulnDBPath string) (vulnerability.Provider, error) {
vulnDBLock.RLock()
- // Fresh, or stale but inside the retry backoff: either way this provider is
- // what the scan will use, so take the shared lock only. Testing the backoff
- // here as well as under the write lock matters — once the DB is stale the
- // freshness test never passes again, and without this every scan would
- // serialize through the exclusive lock just to reach the same conclusion.
- if vulnDB != nil && (time.Since(vulnDBBuilt) < vulnDBRefreshAge || time.Since(vulnDBAttempt) < vulnDBRetryBackoff) {
- vulnDBLock.RUnlock()
- return vulnDB, nil
- }
+ db, built, attempt, lastErr := vulnDB, vulnDBBuilt, vulnDBAttempt, vulnDBLastErr
vulnDBLock.RUnlock()
+ periodic := false
+ switch vulnDBDecide(db, built, attempt, time.Now()) {
+ case vulnDBServe:
+ // Count the scan here, on the path scans actually take. It used to be
+ // counted under the write lock, which a serviceable database never
+ // reaches, so the periodic reload never ran once in production.
+ n := vulnDBScans.Add(1)
+ // Only recycle a database that is genuinely fresh. When the DB is stale
+ // and we are serving it inside the retry backoff, the backoff is the
+ // authority on when to touch the upstream again; a scan counter must not
+ // smuggle an extra download attempt past it during an outage.
+ if n%vulnDBReloadEvery != 0 || time.Since(built) >= vulnDBRefreshAge {
+ return db, nil
+ }
+ periodic = true
+ case vulnDBUnusable:
+ return nil, vulnDBUnavailableErr(db, built, lastErr)
+ }
+
+ // Don't start a database download on behalf of a caller that has already
+ // gone away: the load runs under the exclusive lock and Grype bounds only
+ // its HTTP portion, so an abandoned refresh stalls the whole pool.
+ if err := ctx.Err(); err != nil {
+ if vulnDBUsable(db, built, time.Now()) {
+ return db, nil
+ }
+ return nil, err
+ }
+
vulnDBLock.Lock()
defer vulnDBLock.Unlock()
- // Double-check after acquiring write lock. Freshness is measured from the
- // DB's build timestamp, so a load that fell back to a stale-but-valid DB
- // doesn't earn a fresh cache lease — it stays stale and keeps retrying.
- if vulnDB != nil && time.Since(vulnDBBuilt) < vulnDBRefreshAge {
- // Periodic reload: close and reopen DB every 50 scans to flush
- // SQLite's page cache and mmap region.
- n := vulnDBScans.Add(1)
- if n%50 == 0 {
- slog.Info("Periodic vulnDB reload to release memory", "scans", n)
- vulnDB.Close()
- vulnDB = nil
- // Fall through to reload below
- } else {
+ if periodic {
+ // A periodic reload must survive the double-check — re-deciding here is
+ // what made it unreachable before. Skip it only if another goroutine
+ // reloaded while we waited for the lock, which flushed the caches anyway.
+ if !vulnDBAttempt.Equal(attempt) {
return vulnDB, nil
}
- }
-
- // DB is stale (or absent). Throttle reload attempts: if we probed recently
- // and still hold a usable provider, keep serving it rather than having every
- // worker serialize through its own upstream timeout. The in-memory provider
- // doesn't re-validate build age on queries, so a stale-but-loaded DB still
- // scans fine until the upstream recovers.
- if vulnDB != nil && time.Since(vulnDBAttempt) < vulnDBRetryBackoff {
- return vulnDB, nil
+ slog.Info("Periodic vulnDB reload to release memory", "scans", vulnDBScans.Load())
+ } else {
+ switch vulnDBDecide(vulnDB, vulnDBBuilt, vulnDBAttempt, time.Now()) {
+ case vulnDBServe:
+ return vulnDB, nil
+ case vulnDBUnusable:
+ return nil, vulnDBUnavailableErr(vulnDB, vulnDBBuilt, vulnDBLastErr)
+ }
}
slog.Info("Loading Grype vulnerability database", "path", vulnDBPath, "tmpdir", os.Getenv("TMPDIR"))
@@ -227,35 +347,69 @@ func loadVulnDatabase(ctx context.Context, vulnDBPath string) (vulnerability.Pro
// back to serving the existing DB.
vulnDBAttempt = time.Now()
store, status, err := loadVulnDB(distConfig, installConfig, true)
+
+ if err != nil && vulnDBCorrupt(err) && !vulnDBUsable(vulnDB, vulnDBBuilt, time.Now()) {
+ // `grype db delete && grype db update`, in-process. Once per attempt
+ // window, so the self-heal cannot become its own retry storm, and only
+ // when there is no usable provider — nothing holds the read lock while
+ // we hold the write lock, but there is no reason to pull files out from
+ // under a database that is still scanning.
+ slog.Warn("Vulnerability database is corrupt; deleting it and re-downloading",
+ "error", err, "path", vulnDBPath)
+ if rmErr := os.RemoveAll(vulnDBPath); rmErr != nil {
+ slog.Error("Failed to delete corrupt vulnerability database", "error", rmErr, "path", vulnDBPath)
+ } else if mkErr := os.MkdirAll(vulnDBPath, 0o755); mkErr != nil {
+ slog.Error("Failed to recreate vulnerability database directory", "error", mkErr, "path", vulnDBPath)
+ } else {
+ store, status, err = loadVulnDB(distConfig, installConfig, true)
+ }
+ }
+
if err != nil {
- // Reload failed (e.g. upstream down and on-disk DB past the 14-day max
- // age). If we still hold a usable provider, keep serving it until the
- // backoff elapses; only surface the error on a cold start with no DB.
- if vulnDB != nil {
+ vulnDBLastErr = err
+ // Load-then-swap: the previous provider is still in place, so a failed
+ // refresh (transient or not, periodic or not) costs nothing.
+ if vulnDBUsable(vulnDB, vulnDBBuilt, time.Now()) {
slog.Warn("Vulnerability database reload failed; serving previously loaded DB",
"error", err,
"built", vulnDBBuilt,
"age", time.Since(vulnDBBuilt).Round(time.Minute).String())
return vulnDB, nil
}
- return nil, fmt.Errorf("failed to load vulnerability database: %w", err)
+ if vulnDB != nil {
+ slog.Error("Vulnerability database is past the maximum age and could not be refreshed",
+ "error", err,
+ "built", vulnDBBuilt,
+ "age", time.Since(vulnDBBuilt).Round(time.Minute).String(),
+ "max", vulnDBMaxServeAge.String())
+ }
+ return nil, vulnDBUnavailableErr(vulnDB, vulnDBBuilt, err)
}
+ built = time.Time{}
+ if status != nil {
+ built = status.Built
+ }
age := "unknown"
- if !status.Built.IsZero() {
- age = time.Since(status.Built).Round(time.Minute).String()
+ if !built.IsZero() {
+ age = time.Since(built).Round(time.Minute).String()
+ }
+ if status != nil {
+ slog.Info("Vulnerability database loaded",
+ "built", built,
+ "age", age,
+ "schemaVersion", status.SchemaVersion,
+ "path", status.Path)
+ } else {
+ slog.Info("Vulnerability database loaded", "built", built, "age", age)
}
- slog.Info("Vulnerability database loaded",
- "built", status.Built,
- "age", age,
- "schemaVersion", status.SchemaVersion,
- "path", status.Path)
- if vulnDB != nil {
+ if vulnDB != nil && vulnDB != store {
vulnDB.Close()
}
vulnDB = store
- vulnDBBuilt = status.Built
+ vulnDBBuilt = built
+ vulnDBLastErr = nil
return vulnDB, nil
}
diff --git a/scanner/internal/scan/grype_test.go b/scanner/internal/scan/grype_test.go
index cf4bccf..6282d6a 100644
--- a/scanner/internal/scan/grype_test.go
+++ b/scanner/internal/scan/grype_test.go
@@ -33,14 +33,14 @@ func resetVulnDBState(t *testing.T) {
original := loadVulnDB
t.Cleanup(func() {
vulnDBLock.Lock()
- vulnDB, vulnDBBuilt, vulnDBAttempt = nil, time.Time{}, time.Time{}
+ vulnDB, vulnDBBuilt, vulnDBAttempt, vulnDBLastErr = nil, time.Time{}, time.Time{}, nil
vulnDBLock.Unlock()
vulnDBScans.Store(0)
loadVulnDB = original
})
vulnDBLock.Lock()
- vulnDB, vulnDBBuilt, vulnDBAttempt = nil, time.Time{}, time.Time{}
+ vulnDB, vulnDBBuilt, vulnDBAttempt, vulnDBLastErr = nil, time.Time{}, time.Time{}, nil
vulnDBLock.Unlock()
vulnDBScans.Store(0)
}
@@ -125,11 +125,10 @@ func TestLoadVulnDatabase_FreshDBIsServedWithoutReloading(t *testing.T) {
// serving, untouched.
//
// What this pins, precisely: the behaviour, not either check that implements
-// it. fa1dfb0 tests the backoff twice, once on the read-lock fast path and
-// again under the write lock, and mutation confirms they are redundant for
-// correctness — deleting either one alone leaves this test passing, and only
-// deleting both fails it. That redundancy is deliberate. The fast-path copy
-// exists so a stale DB does not push every scan through the exclusive lock, a
+// it. The policy is one function, vulnDBDecide, consulted twice — once on the
+// read-lock fast path so a stale DB does not push every scan through the
+// exclusive lock, and again under the write lock as the double-check. Deleting
+// either call alone leaves this test passing; the fast-path copy exists for a
// contention property no unit test can assert without being flaky.
func TestLoadVulnDatabase_StaleDBIsThrottled(t *testing.T) {
resetVulnDBState(t)
@@ -141,16 +140,25 @@ func TestLoadVulnDatabase_StaleDBIsThrottled(t *testing.T) {
_, calls := stubLoader(t, time.Now(), nil)
- got, err := loadVulnDatabase(context.Background(), t.TempDir())
- if err != nil {
- t.Fatalf("loadVulnDatabase: %v", err)
- }
- if got != cached {
- t.Error("the stale provider was replaced during the backoff window")
+ // Enough scans to cross the periodic-reload interval several times. The
+ // backoff outranks it: recycling the provider is memory hygiene, and it
+ // must not smuggle an extra upstream attempt past the throttle during an
+ // outage.
+ for i := 0; i < 3*vulnDBReloadEvery; i++ {
+ got, err := loadVulnDatabase(context.Background(), t.TempDir())
+ if err != nil {
+ t.Fatalf("scan %d: %v", i, err)
+ }
+ if got != cached {
+ t.Fatalf("scan %d: the stale provider was replaced during the backoff window", i)
+ }
}
if *calls != 0 {
t.Errorf("loader called %d times inside the retry backoff, want 0", *calls)
}
+ if cached.closed {
+ t.Error("the provider still in use was closed")
+ }
}
// TestLoadVulnDatabase_RetriesOnceTheBackoffElapses is the other half of the
@@ -225,6 +233,9 @@ func TestLoadVulnDatabase_FailedReloadKeepsServingTheOldDB(t *testing.T) {
// TestLoadVulnDatabase_ColdStartFailureIsAnError is the one case that must fail
// loudly. With no provider in hand there is nothing to scan against, and
// returning success would report every image as clean.
+//
+// The first attempt is the one that reaches the loader; what the ones behind it
+// do is TestColdStart_ThrottlesRepeatedAttempts, in vulndb_refresh_test.go.
func TestLoadVulnDatabase_ColdStartFailureIsAnError(t *testing.T) {
resetVulnDBState(t)
diff --git a/scanner/internal/scan/vulndb_refresh_test.go b/scanner/internal/scan/vulndb_refresh_test.go
new file mode 100644
index 0000000..83ef1ad
--- /dev/null
+++ b/scanner/internal/scan/vulndb_refresh_test.go
@@ -0,0 +1,764 @@
+package scan
+
+import (
+ "context"
+ "errors"
+ "os"
+ "path/filepath"
+ "strconv"
+ "strings"
+ "sync"
+ "testing"
+ "time"
+
+ "atcr.io/scanner/internal/client"
+ "atcr.io/scanner/internal/config"
+ "atcr.io/scanner/internal/mockhold"
+ "atcr.io/scanner/internal/queue"
+ "github.com/anchore/grype/grype"
+ v6 "github.com/anchore/grype/grype/db/v6"
+ v6dist "github.com/anchore/grype/grype/db/v6/distribution"
+ v6inst "github.com/anchore/grype/grype/db/v6/installation"
+ "github.com/anchore/grype/grype/vulnerability"
+)
+
+// This file covers the production complaint "grype is out of date, please
+// update the database": the vulnerability database stops refreshing and scans
+// start failing, with no way back to health short of a restart.
+//
+// grype_test.go pins the ordinary lifecycle of loadVulnDatabase. What follows
+// pins the failure lifecycle — cold start, upstream outage, an on-disk database
+// that is corrupt rather than merely old, the periodic reload, and the ceiling
+// on how old a database may get before the scanner refuses to vouch for its
+// results.
+//
+// Every test that reaches a real Grype code path does so without network
+// access. The only calls into grype.LoadVulnerabilityDB use update=false, so
+// no listing is fetched and no archive is downloaded; the on-disk database is
+// synthesised locally with the exported v6 low-level writer.
+//
+// These tests share the package-level vulnDB globals with grype_test.go and so
+// must not be run in parallel with anything.
+
+// slowLoader installs a loader that blocks for d before returning. It is how
+// the tests below stand in for a real refresh, which in production is a 30s
+// listing check plus an up-to-300s archive download plus hydration, all of it
+// inside loadVulnDatabase's exclusive lock.
+func slowLoader(t *testing.T, d time.Duration, built time.Time, err error) (*fakeProvider, *int) {
+ t.Helper()
+ loaded := &fakeProvider{name: "slow"}
+ calls := 0
+ var mu sync.Mutex
+ loadVulnDB = func(v6dist.Config, v6inst.Config, bool) (vulnerability.Provider, *vulnerability.ProviderStatus, error) {
+ mu.Lock()
+ calls++
+ mu.Unlock()
+ time.Sleep(d)
+ if err != nil {
+ return nil, nil, err
+ }
+ return loaded, &vulnerability.ProviderStatus{Built: built}, nil
+ }
+ return loaded, &calls
+}
+
+// scriptedLoader installs a loader that returns the given responses in order,
+// repeating the last one once the script runs out. It is how the tests below
+// express "the upstream fails and then recovers" without a network.
+type loaderResponse struct {
+ built time.Time
+ err error
+}
+
+func scriptedLoader(t *testing.T, responses ...loaderResponse) (*fakeProvider, *int) {
+ t.Helper()
+ loaded := &fakeProvider{name: "scripted"}
+ calls := 0
+ var mu sync.Mutex
+ loadVulnDB = func(v6dist.Config, v6inst.Config, bool) (vulnerability.Provider, *vulnerability.ProviderStatus, error) {
+ mu.Lock()
+ i := calls
+ calls++
+ mu.Unlock()
+ if i >= len(responses) {
+ i = len(responses) - 1
+ }
+ r := responses[i]
+ if r.err != nil {
+ return nil, nil, r.err
+ }
+ return loaded, &vulnerability.ProviderStatus{Built: r.built}, nil
+ }
+ return loaded, &calls
+}
+
+// --- the database lifecycle under failure ---
+
+// TestColdStart_ThrottlesRepeatedAttempts. With no provider in hand the
+// retry backoff must still apply, or every queued scan runs a complete download
+// attempt (30s listing check plus up to a 300s archive fetch) under the
+// exclusive lock while the hold's stale loop keeps re-queueing the failures.
+//
+// The throttled calls must still report an error — the scan genuinely cannot
+// run — but they must not repeat the download.
+func TestColdStart_ThrottlesRepeatedAttempts(t *testing.T) {
+ resetVulnDBState(t)
+
+ _, calls := stubLoader(t, time.Time{}, errors.New("upstream unreachable"))
+
+ const attempts = 5
+ for i := 0; i < attempts; i++ {
+ got, err := loadVulnDatabase(context.Background(), t.TempDir())
+ if err == nil {
+ t.Fatalf("attempt %d: a cold start with no database returned success", i)
+ }
+ if got != nil {
+ t.Fatalf("attempt %d: a provider was returned alongside the error", i)
+ }
+ }
+
+ if *calls != 1 {
+ t.Errorf("loader called %d times across %d cold-start scans, want 1: the retry "+
+ "backoff must apply with no provider in hand, not only with one", *calls, attempts)
+ }
+}
+
+// TestColdStart_ConcurrentWorkersShareOneAttempt is the pool-level consequence
+// of B1. Four workers arriving together must produce one download attempt, not
+// four serialized ones, and must finish in about one attempt's time.
+func TestColdStart_ConcurrentWorkersShareOneAttempt(t *testing.T) {
+ resetVulnDBState(t)
+
+ const workers = 4
+ const loadDelay = 40 * time.Millisecond
+ _, calls := slowLoader(t, loadDelay, time.Time{}, errors.New("upstream unreachable"))
+
+ start := time.Now()
+ var wg sync.WaitGroup
+ for i := 0; i < workers; i++ {
+ wg.Add(1)
+ go func() {
+ defer wg.Done()
+ _, _ = loadVulnDatabase(context.Background(), t.TempDir())
+ }()
+ }
+ wg.Wait()
+ elapsed := time.Since(start)
+
+ if *calls != 1 {
+ t.Errorf("loader called %d times for %d concurrent workers, want 1", *calls, workers)
+ }
+ if max := time.Duration(workers-1) * loadDelay; elapsed >= max {
+ t.Errorf("elapsed %s for %d workers at %s each: the attempts are still being "+
+ "serialized through the exclusive lock", elapsed, workers, loadDelay)
+ }
+}
+
+// TestColdStart_RecoversWhenUpstreamReturns is the other half of the backoff:
+// it has to let go. A scanner that failed its cold start must load the database
+// on its own once the upstream comes back, with no process restart.
+func TestColdStart_RecoversWhenUpstreamReturns(t *testing.T) {
+ resetVulnDBState(t)
+
+ freshBuild := time.Now().Add(-1 * time.Hour)
+ loaded, calls := scriptedLoader(t,
+ loaderResponse{err: errors.New("upstream unreachable")},
+ loaderResponse{built: freshBuild},
+ )
+
+ if _, err := loadVulnDatabase(context.Background(), t.TempDir()); err == nil {
+ t.Fatal("the first cold-start attempt was expected to fail")
+ }
+
+ // Age out the cold-start backoff, standing in for the passage of time.
+ vulnDBLock.Lock()
+ vulnDBAttempt = time.Now().Add(-vulnDBColdRetryBackoff - time.Minute)
+ vulnDBLock.Unlock()
+
+ got, err := loadVulnDatabase(context.Background(), t.TempDir())
+ if err != nil {
+ t.Fatalf("the scanner did not recover once the upstream returned: %v", err)
+ }
+ if got != loaded {
+ t.Error("the newly loaded provider was not adopted")
+ }
+ if *calls != 2 {
+ t.Errorf("loader called %d times, want 2", *calls)
+ }
+ if !vulnDBBuilt.Equal(freshBuild) {
+ t.Errorf("vulnDBBuilt = %v, want %v", vulnDBBuilt, freshBuild)
+ }
+}
+
+// TestPeriodicReload_FiresOnSchedule. The "close and reopen the DB every N
+// scans to flush SQLite's page cache and mmap region" mitigation has to
+// actually run; for most of this file's history it never did. The counter therefore has to live on the path scans take — the
+// read-lock fast path — because a fresh database never reaches the write lock.
+func TestPeriodicReload_FiresOnSchedule(t *testing.T) {
+ resetVulnDBState(t)
+
+ cached := &fakeProvider{name: "cached"}
+ vulnDB = cached
+ vulnDBBuilt = time.Now().Add(-1 * time.Hour) // fresh
+
+ loaded, calls := stubLoader(t, time.Now().Add(-1*time.Hour), nil)
+
+ const scans = 200
+ for i := 0; i < scans; i++ {
+ got, err := loadVulnDatabase(context.Background(), t.TempDir())
+ if err != nil {
+ t.Fatalf("scan %d: %v", i, err)
+ }
+ if got == nil {
+ t.Fatalf("scan %d returned no provider", i)
+ }
+ }
+
+ if want := scans / vulnDBReloadEvery; *calls != want {
+ t.Errorf("loader called %d times in %d fresh scans, want %d (one every %d)",
+ *calls, scans, want, vulnDBReloadEvery)
+ }
+ if !cached.closed {
+ t.Error("the replaced provider was not closed: every periodic reload would leak its predecessor")
+ }
+ if vulnDB != loaded {
+ t.Error("the reloaded provider was not adopted")
+ }
+ if n := vulnDBScans.Load(); n != scans {
+ t.Errorf("vulnDBScans = %d after %d scans, want %d: the counter is not on the path scans take", n, scans, scans)
+ }
+}
+
+// TestPeriodicReload_FailureKeepsTheWorkingProvider is the interaction that
+// makes the reload dangerous to enable on its own. It used to null the provider
+// before loading its replacement, so one transient failure would drop a scanner
+// that was working perfectly into the no-provider state above.
+//
+// Load first, swap second: a failed periodic reload must be a no-op.
+func TestPeriodicReload_FailureKeepsTheWorkingProvider(t *testing.T) {
+ resetVulnDBState(t)
+
+ cached := &fakeProvider{name: "cached"}
+ vulnDB = cached
+ vulnDBBuilt = time.Now().Add(-1 * time.Hour) // fresh
+
+ _, calls := stubLoader(t, time.Time{}, errors.New("upstream blip"))
+
+ for i := 0; i < vulnDBReloadEvery+5; i++ {
+ got, err := loadVulnDatabase(context.Background(), t.TempDir())
+ if err != nil {
+ t.Fatalf("scan %d: a failed periodic reload must not fail the scan: %v", i, err)
+ }
+ if got != cached {
+ t.Fatalf("scan %d: the working provider was dropped for a failed reload", i)
+ }
+ }
+
+ if cached.closed {
+ t.Error("the working provider was closed before its replacement was in hand")
+ }
+ if *calls != 1 {
+ t.Errorf("loader called %d times, want 1", *calls)
+ }
+}
+
+// TestStaleFallback_RefusesPastTheCeiling. Serving a slightly old database
+// beats refusing to scan, which is why the fallback exists. Serving one of
+// unbounded age while reporting success is a different thing: a scanner whose
+// egress is blocked would publish confident "0 critical" verdicts for months.
+//
+// Grype's own MaxAllowedBuiltAge is the natural line, since that is the
+// guarantee grypeDBConfig already asks for.
+func TestStaleFallback_RefusesPastTheCeiling(t *testing.T) {
+ resetVulnDBState(t)
+
+ ancient := &fakeProvider{name: "ancient"}
+ vulnDB = ancient
+ vulnDBBuilt = time.Now().Add(-120 * 24 * time.Hour) // four months old
+
+ _, calls := stubLoader(t, time.Time{}, errors.New("upstream unreachable"))
+
+ got, err := loadVulnDatabase(context.Background(), t.TempDir())
+ if err == nil {
+ t.Fatal("a database built 120 days ago was served with a nil error")
+ }
+ if got != nil {
+ t.Error("a provider past the ceiling was returned alongside the error")
+ }
+ if !strings.Contains(err.Error(), "too old") {
+ t.Errorf("error %q does not say the database is too old to trust", err)
+ }
+ if *calls != 1 {
+ t.Errorf("loader called %d times, want 1", *calls)
+ }
+
+ // And the refusal is throttled like any other failure, rather than running
+ // a download per scan.
+ if _, err := loadVulnDatabase(context.Background(), t.TempDir()); err == nil {
+ t.Fatal("the second scan was served from the same ancient database")
+ }
+ if *calls != 1 {
+ t.Errorf("loader called %d times across two scans, want 1", *calls)
+ }
+}
+
+// TestStaleFallback_ServesInsideTheCeiling pins the other side of the trade: a
+// database that is stale but still inside Grype's max allowed age keeps
+// scanning through an upstream outage.
+func TestStaleFallback_ServesInsideTheCeiling(t *testing.T) {
+ resetVulnDBState(t)
+
+ old := &fakeProvider{name: "old"}
+ vulnDB = old
+ vulnDBBuilt = time.Now().Add(-10 * 24 * time.Hour) // stale, inside the 14-day ceiling
+
+ _, calls := stubLoader(t, time.Time{}, errors.New("upstream unreachable"))
+
+ got, err := loadVulnDatabase(context.Background(), t.TempDir())
+ if err != nil {
+ t.Fatalf("a stale but usable database must keep scanning: %v", err)
+ }
+ if got != old {
+ t.Error("the previously loaded provider was not served")
+ }
+ if old.closed {
+ t.Error("the provider still in use was closed")
+ }
+ if *calls != 1 {
+ t.Errorf("loader called %d times, want 1", *calls)
+ }
+}
+
+// TestCorruptDB_IsDeletedAndRetriedOnce covers the database that is current but
+// whose import metadata was lost — an interrupted activate, a truncated write,
+// a restored volume snapshot. curator.Update only installs something strictly
+// newer, and a checksum failure leaves the on-disk description in place, so
+// such a database never re-downloads: it fails every scan forever and nothing
+// deletes it. Delete it and retry once, which is what `grype db delete && grype
+// db update` does for CLI users.
+func TestCorruptDB_IsDeletedAndRetriedOnce(t *testing.T) {
+ resetVulnDBState(t)
+
+ root := t.TempDir()
+ marker := filepath.Join(root, "6", "vulnerability.db")
+ if err := os.MkdirAll(filepath.Dir(marker), 0o755); err != nil {
+ t.Fatalf("mkdir: %v", err)
+ }
+ if err := os.WriteFile(marker, []byte("corrupt"), 0o644); err != nil {
+ t.Fatalf("write: %v", err)
+ }
+
+ freshBuild := time.Now().Add(-1 * time.Hour)
+ loaded, calls := scriptedLoader(t,
+ loaderResponse{err: errors.New("no import metadata file at: " + filepath.Join(root, "6", "import.json"))},
+ loaderResponse{built: freshBuild},
+ )
+
+ got, err := loadVulnDatabase(context.Background(), root)
+ if err != nil {
+ t.Fatalf("a corrupt database must be deleted and re-downloaded, not served as a failure: %v", err)
+ }
+ if got != loaded {
+ t.Error("the re-downloaded provider was not adopted")
+ }
+ if *calls != 2 {
+ t.Fatalf("loader called %d times, want 2 (the load and one retry after the delete)", *calls)
+ }
+ if _, err := os.Stat(marker); !os.IsNotExist(err) {
+ t.Errorf("the corrupt database file still exists at %s (stat err %v)", marker, err)
+ }
+}
+
+// TestCorruptDB_RetriesOnlyOncePerAttempt keeps the self-heal from becoming its
+// own storm: if the retry fails too, the attempt ends there and the backoff
+// takes over.
+func TestCorruptDB_RetriesOnlyOncePerAttempt(t *testing.T) {
+ resetVulnDBState(t)
+
+ root := t.TempDir()
+ _, calls := stubLoader(t, time.Time{}, errors.New("no import metadata file at: "+root+"/6/import.json"))
+
+ if _, err := loadVulnDatabase(context.Background(), root); err == nil {
+ t.Fatal("a database that stays corrupt must still fail the scan")
+ }
+ if *calls != 2 {
+ t.Errorf("loader called %d times, want 2: the delete-and-retry must happen once per attempt", *calls)
+ }
+
+ // And the next scan is throttled rather than repeating the delete-and-retry.
+ if _, err := loadVulnDatabase(context.Background(), root); err == nil {
+ t.Fatal("the second scan unexpectedly succeeded")
+ }
+ if *calls != 2 {
+ t.Errorf("loader called %d times across two scans, want 2", *calls)
+ }
+}
+
+// TestLoadVulnDatabase_HonoursContextCancellation. A shutdown, or a hold that
+// has abandoned the job, must not start a fresh database download that then
+// holds the exclusive lock until Grype's own timeouts expire — up to 300s for
+// the archive alone, unbounded for hydration.
+func TestLoadVulnDatabase_HonoursContextCancellation(t *testing.T) {
+ resetVulnDBState(t)
+
+ _, calls := stubLoader(t, time.Now(), nil)
+
+ ctx, cancel := context.WithCancel(context.Background())
+ cancel()
+
+ if _, err := loadVulnDatabase(ctx, t.TempDir()); !errors.Is(err, context.Canceled) {
+ t.Errorf("loadVulnDatabase err = %v, want context.Canceled", err)
+ }
+ if *calls != 0 {
+ t.Errorf("loader called %d times with an already-cancelled context, want 0", *calls)
+ }
+
+ // A cancelled context must not take a usable database away from a caller
+ // that needs no load at all.
+ cached := &fakeProvider{name: "cached"}
+ vulnDBLock.Lock()
+ vulnDB = cached
+ vulnDBBuilt = time.Now().Add(-1 * time.Hour)
+ vulnDBLock.Unlock()
+
+ got, err := loadVulnDatabase(ctx, t.TempDir())
+ if err != nil {
+ t.Fatalf("a cached provider needs no load and must not be refused: %v", err)
+ }
+ if got != cached {
+ t.Error("the cached provider was not served")
+ }
+}
+
+// TestRefresh_BlocksAllInFlightScans quantifies the third finding: a refresh
+// holds the write lock for the entire download, and scanVulnerabilities holds
+// the read lock across FindMatches. A scan that arrives during a refresh
+// therefore waits out the whole download before it can even begin matching.
+//
+// The wait is measured on vulnDBLock directly rather than through
+// scanVulnerabilities, which would need a real SBOM and a real provider. It is
+// the same lock and the same acquisition scanVulnerabilities performs at
+// grype.go:105.
+func TestRefresh_BlocksAllInFlightScans(t *testing.T) {
+ resetVulnDBState(t)
+
+ old := &fakeProvider{name: "old"}
+ vulnDB = old
+ vulnDBBuilt = time.Now().Add(-10 * 24 * time.Hour) // stale, so a reload is due
+ vulnDBAttempt = time.Now().Add(-vulnDBRetryBackoff - time.Minute) // backoff expired
+
+ const loadDelay = 150 * time.Millisecond
+ _, calls := slowLoader(t, loadDelay, time.Now(), nil)
+
+ refreshDone := make(chan struct{})
+ go func() {
+ defer close(refreshDone)
+ if _, err := loadVulnDatabase(context.Background(), t.TempDir()); err != nil {
+ t.Errorf("refresh: %v", err)
+ }
+ }()
+
+ // Give the refresher time to take the write lock before the scan tries for
+ // the read lock. Any interleaving still produces a valid measurement; this
+ // only makes the intended one likely.
+ time.Sleep(20 * time.Millisecond)
+
+ start := time.Now()
+ vulnDBLock.RLock()
+ blocked := time.Since(start)
+ vulnDBLock.RUnlock()
+
+ <-refreshDone
+
+ if *calls != 1 {
+ t.Fatalf("loader called %d times, want 1", *calls)
+ }
+ if blocked < loadDelay/2 {
+ t.Skipf("scan acquired the read lock in %s: the refresher had not taken the "+
+ "write lock yet, so this run measured nothing", blocked)
+ }
+ t.Logf("a scan arriving during a refresh waited %s for the read lock; in production "+
+ "the same wait is the full listing check plus archive download plus hydration, "+
+ "which grype bounds only by its 300s UpdateTimeout", blocked)
+}
+
+// --- what Grype itself does with the configuration this package passes it ---
+
+// writeSyntheticDB creates a minimal but genuine Grype v6 database on disk with
+// the given build timestamp, at the layout grypeDBConfig expects
+// (//vulnerability.db) and returns that directory.
+//
+// No import.json is written. That file is produced by the curator's activate
+// step, and its absence is deliberate: writing a valid one needs grype's
+// xxhash digest helper (and a new direct module dependency), and the tests
+// that call into the curator here are about Status(), which reports the age
+// failure whether or not the checksum check also fires.
+func writeSyntheticDB(t *testing.T, root string, built time.Time) string {
+ t.Helper()
+
+ dir := filepath.Join(root, strconv.Itoa(v6.ModelVersion))
+ if err := os.MkdirAll(dir, 0o755); err != nil {
+ t.Fatalf("mkdir: %v", err)
+ }
+
+ gdb, err := v6.NewLowLevelDB(filepath.Join(dir, v6.VulnerabilityDBFileName), true, true, false)
+ if err != nil {
+ t.Fatalf("create db: %v", err)
+ }
+ ts := built.UTC().Round(time.Second)
+ err = gdb.Create(&v6.DBMetadata{
+ BuildTimestamp: &ts,
+ Model: v6.ModelVersion,
+ Revision: v6.Revision,
+ Addition: v6.Addition,
+ }).Error
+ if err != nil {
+ t.Fatalf("write metadata: %v", err)
+ }
+ sqlDB, err := gdb.DB()
+ if err != nil {
+ t.Fatalf("unwrap db: %v", err)
+ }
+ if err := sqlDB.Close(); err != nil {
+ t.Fatalf("close db: %v", err)
+ }
+ return dir
+}
+
+// TestGrype_ExpiredOnDiskDBIsAHardError is the production error message,
+// reproduced with no network.
+//
+// grypeDBConfig sets ValidateAge with a 14-day MaxAllowedBuiltAge. Once the
+// on-disk database crosses that line, curator.Status() fails, and
+// grype.LoadVulnerabilityDB returns that failure instead of a provider — after
+// the update attempt, not before it. So an expired database plus an upstream
+// that cannot be reached is a hard load failure, not a degraded scan.
+//
+// The call here uses update=false purely to keep the test off the network.
+// Production calls with update=true, and on the path that matters (the update
+// fails, or finds nothing to install) reaches the identical Status() check.
+func TestGrype_ExpiredOnDiskDBIsAHardError(t *testing.T) {
+ root := t.TempDir()
+ writeSyntheticDB(t, root, time.Now().Add(-20*24*time.Hour))
+
+ distCfg, instCfg := grypeDBConfig(root)
+ provider, _, err := grype.LoadVulnerabilityDB(distCfg, instCfg, false)
+ if err == nil {
+ if provider != nil {
+ provider.Close()
+ }
+ t.Fatal("a database built 20 days ago loaded successfully; MaxAllowedBuiltAge is not being enforced")
+ }
+ if provider != nil {
+ t.Error("a provider was returned alongside the error")
+ }
+
+ // This is Grype's own wording, from curator.validateAge
+ // (grype/db/v6/installation/curator.go:613). It is what reaches the hold,
+ // wrapped twice by grype.go, and what users report as "grype is out of
+ // date, please update the database".
+ if !strings.Contains(err.Error(), "the vulnerability database was built") ||
+ !strings.Contains(err.Error(), "max allowed age") {
+ t.Errorf("unexpected error text %q; the age check may have moved", err)
+ }
+ t.Logf("grype: %v", err)
+}
+
+// TestGrype_FreshDBWithoutImportMetadataIsAlsoAHardError is a second, quieter
+// way into the same symptom, and unlike the age cliff it does not self-heal.
+//
+// grypeDBConfig also sets ValidateChecksum. Status() joins the checksum failure
+// into the same error, so a database with a missing or corrupt import.json
+// fails to load even when its build timestamp is minutes old.
+//
+// The asymmetry is what makes this dangerous. curator.Update() nils out the
+// current description when validateAge fails, and isSupersededBy(nil, ...) is
+// unconditionally true, so an expired database is always re-downloaded. A
+// checksum failure leaves the description in place, so a download only happens
+// if the upstream has something strictly newer. A database that is current but
+// whose import metadata was lost (an interrupted activate, a truncated write, a
+// restored volume snapshot) therefore fails every scan and no update fixes it.
+func TestGrype_FreshDBWithoutImportMetadataIsAlsoAHardError(t *testing.T) {
+ root := t.TempDir()
+ writeSyntheticDB(t, root, time.Now().Add(-1*time.Hour))
+
+ distCfg, instCfg := grypeDBConfig(root)
+ provider, _, err := grype.LoadVulnerabilityDB(distCfg, instCfg, false)
+ if err == nil {
+ if provider != nil {
+ provider.Close()
+ }
+ t.Fatal("a database with no import metadata loaded successfully; ValidateChecksum is not being enforced")
+ }
+ if !strings.Contains(err.Error(), "import metadata") {
+ t.Errorf("unexpected error text %q; expected the checksum/import-metadata failure", err)
+ }
+ t.Logf("grype: %v", err)
+}
+
+// TestGrype_InMemoryProviderDoesNotRevalidateAge checks the claim grype.go
+// leans on at line 210: "the in-memory provider doesn't re-validate build age
+// on queries, so a stale-but-loaded DB still scans fine".
+//
+// The claim holds, and this is the mechanism the serve-the-old-DB fallback
+// depends on. v6.NewVulnerabilityProvider closes over a Reader and an
+// architecture-alias map and nothing else: no Config, no MaxAllowedBuiltAge, no
+// build timestamp. Age is checked by the curator when opening the file on disk,
+// never by the provider that results. Which is also why the fallback has no
+// ceiling: there is nothing in the object that could impose one.
+//
+// The provider is built the way LoadVulnerabilityDB builds it, but without the
+// curator, so no age or checksum gate runs at all.
+func TestGrype_InMemoryProviderDoesNotRevalidateAge(t *testing.T) {
+ root := t.TempDir()
+ // A year old: far past MaxAllowedBuiltAge, which the curator would refuse
+ // and the provider cannot see.
+ built := time.Now().Add(-365 * 24 * time.Hour)
+ dir := writeSyntheticDB(t, root, built)
+
+ rdr, err := v6.NewReader(v6.Config{DBDirPath: dir})
+ if err != nil {
+ t.Fatalf("open reader: %v", err)
+ }
+ defer rdr.Close()
+
+ meta, err := rdr.GetDBMetadata()
+ if err != nil {
+ t.Fatalf("read metadata: %v", err)
+ }
+ if age := time.Since(*meta.BuildTimestamp); age < 300*24*time.Hour {
+ t.Fatalf("test setup drifted: build age is %s", age)
+ }
+
+ provider := v6.NewVulnerabilityProvider(rdr)
+
+ // A query against a year-old database. It answers, rather than reporting
+ // that the database is out of date, and would answer identically at any
+ // age. That is the whole content of the comment at grype.go:210.
+ vulns, err := provider.FindVulnerabilities()
+ if err != nil {
+ t.Fatalf("query: %v", err)
+ }
+ if len(vulns) != 0 {
+ t.Fatalf("synthetic database returned %d vulnerabilities", len(vulns))
+ }
+}
+
+// --- the symptom, through the real worker pool ---
+
+// TestWorkerPool_EveryScanFailsWhileTheDBIsUnloadable drives the whole scan
+// pipeline against a mock hold with the database loader failing, which is what
+// production looks like during the outage. It is here rather than in
+// internal/e2e because loadVulnDB is unexported and the harness cannot stub it.
+//
+// Two things are asserted that the unit tests cannot: the failure reaches the
+// hold as an "error" message (which the hold's stale-scan loop retries, unlike
+// a "skipped"), and the cold-start backoff holds across jobs arriving from the
+// hold — so the hold's retry loop no longer drives one full download attempt
+// per queued scan for as long as the upstream is down.
+func TestWorkerPool_EveryScanFailsWhileTheDBIsUnloadable(t *testing.T) {
+ fixture := filepath.Join("..", "mockhold", "testdata", "blobs", "hsm-secrets-operator")
+ if _, err := os.Stat(filepath.Join(fixture, "oci-layout")); err != nil {
+ t.Skipf("fixture not present; run scanner/internal/mockhold/testdata/fetch-blobs.sh")
+ }
+
+ resetVulnDBState(t)
+ _, calls := stubLoader(t, time.Time{}, errors.New("upstream unreachable"))
+
+ const secret = "test-scanner-secret"
+ hold := mockhold.New(mockhold.NewOCILayout(fixture), mockhold.WithSecret(secret))
+ t.Cleanup(hold.Close)
+
+ cfg := config.DefaultConfig()
+ cfg.Hold.URL = hold.URL()
+ cfg.Hold.Secret = secret
+ cfg.Scanner.Workers = 1
+ cfg.Vuln.Enabled = true // the point of the test
+ cfg.Vuln.DBPath = t.TempDir()
+ cfg.Vuln.TmpDir = t.TempDir()
+
+ // WorkerPool.Start exports TMPDIR process-wide and never restores it.
+ origTmpDir, hadTmpDir := os.LookupEnv("TMPDIR")
+ t.Cleanup(func() {
+ if hadTmpDir {
+ os.Setenv("TMPDIR", origTmpDir)
+ return
+ }
+ os.Unsetenv("TMPDIR")
+ })
+
+ restoreCooldown := JobCooldown
+ JobCooldown = 10 * time.Millisecond
+
+ q := queue.NewJobQueue(cfg.Scanner.QueueSize)
+ c := client.NewHoldClient(cfg.Hold.URL, cfg.Hold.Secret, q)
+ pool := NewWorkerPool(cfg, q, c)
+
+ ctx, cancel := context.WithCancel(context.Background())
+ pool.Start(ctx)
+ go c.Connect()
+ t.Cleanup(func() {
+ cancel()
+ c.Close()
+ q.Close()
+ // Wait for the workers before restoring JobCooldown. The worker loop
+ // reads it on every iteration, so restoring it while a worker is still
+ // running is a genuine data race that -race reports. (The e2e harness
+ // restores it without waiting, and has the same latent race.)
+ pool.Wait()
+ JobCooldown = restoreCooldown
+ })
+
+ if err := hold.WaitForScanner(10 * time.Second); err != nil {
+ t.Fatalf("scanner never connected: %v", err)
+ }
+
+ const target = "sha256:1cfa4e2b09e127b9c4ed43578d3f3c18e7d44ea47b9ea98475c0cbe9086525f8"
+ all, err := mockhold.Corpus()
+ if err != nil {
+ t.Fatalf("load corpus: %v", err)
+ }
+ var manifest mockhold.Manifest
+ for _, m := range all {
+ if m.Digest == target {
+ manifest = m
+ break
+ }
+ }
+ if manifest.Digest == "" {
+ t.Fatalf("digest %s not in corpus", target)
+ }
+
+ // The startup goroutine in WorkerPool.Start also calls the loader once.
+ // Take a baseline after connecting so the per-job count is unambiguous.
+ baseline := *calls
+
+ const jobs = 2
+ for i := 0; i < jobs; i++ {
+ seq, err := hold.SendJob(manifest.Job())
+ if err != nil {
+ t.Fatalf("send job %d: %v", i, err)
+ }
+ msg, err := hold.WaitForMessage(func(m mockhold.Message) bool {
+ return m.Seq == seq && (m.Type == "result" || m.Type == "error" || m.Type == "skipped")
+ }, 3*time.Minute)
+ if err != nil {
+ t.Fatalf("no terminal message for seq %d: %v", seq, err)
+ }
+ if msg.Type != "error" {
+ t.Fatalf("job %d: want error while the vulnerability DB is unloadable, got %s", i, msg.Type)
+ }
+ if !strings.Contains(msg.Error, "failed to load vulnerability database") {
+ t.Errorf("job %d: error %q does not name the database failure", i, msg.Error)
+ }
+ t.Logf("job %d: %q", i, msg.Error)
+ }
+
+ // The jobs still fail — there is genuinely nothing to scan against — but
+ // they fail cheaply. Both arrive inside vulnDBColdRetryBackoff, so at most
+ // one of them reaches the loader.
+ if extra := *calls - baseline; extra > 1 {
+ t.Errorf("loader ran %d times across %d jobs, want at most 1: the cold-start "+
+ "backoff must throttle the hold's retry loop", extra, jobs)
+ }
+}
diff --git a/scanner/internal/scan/worker.go b/scanner/internal/scan/worker.go
index 4b140e4..1754040 100644
--- a/scanner/internal/scan/worker.go
+++ b/scanner/internal/scan/worker.go
@@ -9,7 +9,6 @@ import (
"log/slog"
"os"
"runtime"
- "strings"
"sync"
"time"
@@ -104,6 +103,11 @@ func (wp *WorkerPool) worker(ctx context.Context, id int) {
"digest", job.ManifestDigest,
"tier", job.Tier)
+ // Tell the hold the scan is actually starting. The ack it already has
+ // was sent on receipt, before this job joined the queue, so it cannot
+ // tell queueing from scanning without this.
+ wp.client.SendStarted(job.Seq)
+
result, err := wp.processJob(ctx, job)
if err != nil {
var skipErr *SkipError
@@ -123,10 +127,21 @@ func (wp *WorkerPool) worker(ctx context.Context, id int) {
} else {
wp.client.SendResult(job.Seq, result)
- slog.Info("Scan job completed",
- "worker_id", id,
- "repository", job.Repository,
- "vulnerabilities", result.Summary.Total)
+ // A nil Summary means Grype never ran (vuln.enabled=false), which
+ // is not the same as "scanned, found nothing". Log the completion
+ // without a count rather than printing a zero the scan never
+ // established.
+ if result.Summary != nil {
+ slog.Info("Scan job completed",
+ "worker_id", id,
+ "repository", job.Repository,
+ "vulnerabilities", result.Summary.Total)
+ } else {
+ slog.Info("Scan job completed",
+ "worker_id", id,
+ "repository", job.Repository,
+ "vulnerabilities", "not scanned")
+ }
}
// Free large scan artifacts and trigger GC before the cooldown
@@ -139,11 +154,20 @@ func (wp *WorkerPool) worker(ctx context.Context, id int) {
select {
case <-ctx.Done():
return
- case <-time.After(10 * time.Second):
+ case <-time.After(JobCooldown):
}
}
}
+// JobCooldown is the pause a worker takes after each job so Go's GC can
+// reclaim what Syft and Grype allocated before the next scan starts.
+//
+// It is a variable, and exported, solely so tests in other packages can
+// shorten it: at the production value a scenario that runs a handful of jobs
+// through a single worker spends nearly all its runtime asleep. Production
+// code must not change it.
+var JobCooldown = 10 * time.Second
+
// unscannable config media types — these are OCI artifacts that aren't
// container images so Syft/Grype can't analyze their layers.
var unscannableConfigTypes = map[string]bool{
@@ -153,6 +177,10 @@ var unscannableConfigTypes = map[string]bool{
}
// skipReason reports why a job cannot be scanned, or "" when it can be.
+//
+// Everything decided here is a permanent property of the manifest record, so
+// everything here is a skip rather than an error: the hold records a skip once
+// and re-offers a failure on every stale pass, forever.
func skipReason(job *scanner.ScanJob) string {
if unscannableConfigTypes[job.Config.MediaType] {
return fmt.Sprintf("unscannable artifact type %s", job.Config.MediaType)
@@ -162,21 +190,31 @@ func skipReason(job *scanner.ScanJob) string {
// single in-toto or DSSE payload as its layer, so the config media type
// alone does not identify it. buildOCILayout drops every non-tar layer,
// which would hand Syft an image with nothing in it.
- if len(job.Layers) > 0 && !hasScannableLayer(job.Layers) {
+ if len(job.Layers) > 0 && !hasScannableLayer(job) {
return fmt.Sprintf("no scannable layers (%s)", job.Layers[0].MediaType)
}
+ // Every digest the scan will use has to be a digest. Each one names a blob
+ // to ask the hold for and a file to write into the layout, and a string
+ // that is neither cannot start being one on a later attempt. Checking here
+ // rejects the job before a single request goes out; buildOCILayout parses
+ // the same digests again because it is what turns them into paths, and
+ // that boundary must hold on its own.
+ for _, ref := range referencedBlobs(job) {
+ if _, err := scanner.ParseDigest(ref.Descriptor.Digest); err != nil {
+ return fmt.Sprintf("%s: %v", ref.what(), err)
+ }
+ }
+
return ""
}
-// hasScannableLayer mirrors the layer filter in buildOCILayout: anything that
-// is not a tar of some flavour is not something Syft can read.
-func hasScannableLayer(layers []scanner.BlobDescriptor) bool {
- for _, layer := range layers {
- if layer.Digest == "" {
- continue
- }
- if layer.MediaType == "" || strings.Contains(layer.MediaType, "tar") {
+// hasScannableLayer reports whether any layer survives the filter
+// buildOCILayout applies. It asks referencedBlobs rather than repeating the
+// media-type rule, so this answer and the layout can never disagree.
+func hasScannableLayer(job *scanner.ScanJob) bool {
+ for _, ref := range referencedBlobs(job) {
+ if !ref.isConfig() {
return true
}
}
@@ -199,7 +237,13 @@ func (wp *WorkerPool) processJob(ctx context.Context, job *scanner.ScanJob) (*sc
return nil, fmt.Errorf("failed to create tmp dir: %w", err)
}
- // Check total compressed image size before downloading
+ // The cheap guard: refuse an image that admits to being over the ceiling
+ // before a byte moves. It is only a pre-check, because the sizes it adds up
+ // come from the same user-writable record as the digests; buildOCILayout
+ // enforces the same ceiling against the bytes that actually arrive.
+ //
+ // Either way the verdict is permanent — an image does not shrink — so it is
+ // a skip, not a failure the stale loop will offer back forever.
if wp.cfg.Vuln.MaxImageSize > 0 {
var totalSize int64
for _, layer := range job.Layers {
@@ -207,13 +251,14 @@ func (wp *WorkerPool) processJob(ctx context.Context, job *scanner.ScanJob) (*sc
}
totalSize += job.Config.Size
if totalSize > wp.cfg.Vuln.MaxImageSize {
- return nil, fmt.Errorf("image too large: %d bytes compressed (limit %d bytes)", totalSize, wp.cfg.Vuln.MaxImageSize)
+ return nil, &SkipError{Reason: fmt.Sprintf(
+ "image too large: %d bytes compressed (limit %d bytes)", totalSize, wp.cfg.Vuln.MaxImageSize)}
}
}
// Step 1: Build OCI image layout from hold via presigned URLs
slog.Info("Building OCI layout", "repository", job.Repository)
- ociLayoutDir, cleanup, err := buildOCILayout(job, wp.cfg.Vuln.TmpDir, wp.cfg.Hold.Secret)
+ ociLayoutDir, cleanup, err := buildOCILayout(job, wp.cfg.Vuln.TmpDir, wp.cfg.Hold.Secret, wp.cfg.Vuln.MaxImageSize)
if err != nil {
return nil, fmt.Errorf("failed to build OCI layout: %w", err)
}
diff --git a/scanner/internal/scan/worker_skip_test.go b/scanner/internal/scan/worker_skip_test.go
index b6e3e14..e2edea5 100644
--- a/scanner/internal/scan/worker_skip_test.go
+++ b/scanner/internal/scan/worker_skip_test.go
@@ -1,11 +1,19 @@
package scan
import (
+ "strings"
"testing"
scanner "atcr.io/scanner"
)
+// hexDigest builds a well-formed sha256 digest from a one-character seed.
+// skipReason validates every digest it will use, so a placeholder like
+// "sha256:config" is now itself a skip and would mask what these cases test.
+func hexDigest(seed rune) string {
+ return "sha256:" + strings.Repeat(string(seed), scanner.HexLen)
+}
+
// TestSkipReason covers the artifact shapes the scanner must refuse before it
// spends a download on them. The attestation case is the one that reached
// production: an in-toto SLSA provenance manifest carries an ordinary image
@@ -65,12 +73,12 @@ func TestSkipReason(t *testing.T) {
for _, tt := range tests {
t.Run(tt.name, func(t *testing.T) {
job := &scanner.ScanJob{
- Config: scanner.BlobDescriptor{MediaType: tt.configType, Digest: "sha256:config"},
+ Config: scanner.BlobDescriptor{MediaType: tt.configType, Digest: hexDigest('c')},
}
for i, mt := range tt.layerTypes {
job.Layers = append(job.Layers, scanner.BlobDescriptor{
MediaType: mt,
- Digest: "sha256:layer" + string(rune('a'+i)),
+ Digest: hexDigest(rune('a' + i)),
})
}
@@ -81,3 +89,83 @@ func TestSkipReason(t *testing.T) {
})
}
}
+
+// TestSkipReasonRejectsMalformedDigests pins the other half of skipReason: a
+// digest that is not a digest is a permanent property of the record, so the
+// job is refused here rather than failing later as a retryable error.
+//
+// A malformed digest on a layer the layout would drop anyway is not a reason
+// to refuse the image, and the last case pins that: the check walks exactly the
+// blobs referencedBlobs will fetch.
+func TestSkipReasonRejectsMalformedDigests(t *testing.T) {
+ const layerType = "application/vnd.oci.image.layer.v1.tar+gzip"
+ const configType = "application/vnd.oci.image.config.v1+json"
+
+ tests := []struct {
+ name string
+ job *scanner.ScanJob
+ wantSkip bool
+ }{
+ {
+ name: "path traversal in a layer digest",
+ job: &scanner.ScanJob{
+ Config: scanner.BlobDescriptor{MediaType: configType, Digest: hexDigest('c')},
+ Layers: []scanner.BlobDescriptor{{MediaType: layerType, Digest: "sha256:../../../escaped"}},
+ },
+ wantSkip: true,
+ },
+ {
+ name: "path traversal in the config digest",
+ job: &scanner.ScanJob{
+ Config: scanner.BlobDescriptor{MediaType: configType, Digest: "sha256:../../../escaped"},
+ Layers: []scanner.BlobDescriptor{{MediaType: layerType, Digest: hexDigest('a')}},
+ },
+ wantSkip: true,
+ },
+ {
+ name: "no algorithm prefix",
+ job: &scanner.ScanJob{
+ Config: scanner.BlobDescriptor{MediaType: configType, Digest: hexDigest('c')},
+ Layers: []scanner.BlobDescriptor{{MediaType: layerType, Digest: strings.Repeat("ef", 32)}},
+ },
+ wantSkip: true,
+ },
+ {
+ name: "unsupported algorithm",
+ job: &scanner.ScanJob{
+ Config: scanner.BlobDescriptor{MediaType: configType, Digest: hexDigest('c')},
+ Layers: []scanner.BlobDescriptor{{MediaType: layerType, Digest: "sha512:" + strings.Repeat("cd", 64)}},
+ },
+ wantSkip: true,
+ },
+ {
+ name: "an empty layer digest is dropped, not refused",
+ job: &scanner.ScanJob{
+ Config: scanner.BlobDescriptor{MediaType: configType, Digest: hexDigest('c')},
+ Layers: []scanner.BlobDescriptor{
+ {MediaType: layerType, Digest: ""},
+ {MediaType: layerType, Digest: hexDigest('a')},
+ },
+ },
+ },
+ {
+ name: "a dropped non-tar layer's digest is not checked",
+ job: &scanner.ScanJob{
+ Config: scanner.BlobDescriptor{MediaType: configType, Digest: hexDigest('c')},
+ Layers: []scanner.BlobDescriptor{
+ {MediaType: "application/vnd.oci.image.layer.v1.zstd+odd", Digest: "not a digest"},
+ {MediaType: layerType, Digest: hexDigest('a')},
+ },
+ },
+ },
+ }
+
+ for _, tt := range tests {
+ t.Run(tt.name, func(t *testing.T) {
+ reason := skipReason(tt.job)
+ if got := reason != ""; got != tt.wantSkip {
+ t.Errorf("skipReason = %q, wantSkip=%v", reason, tt.wantSkip)
+ }
+ })
+ }
+}
diff --git a/scanner/types.go b/scanner/types.go
index 4bf9000..e1c4453 100644
--- a/scanner/types.go
+++ b/scanner/types.go
@@ -67,6 +67,20 @@ type AckMessage struct {
Seq int64 `json:"seq"`
}
+// StartedMessage is sent from scanner to hold when a worker actually begins a
+// scan, as distinct from AckMessage, which is sent from the WebSocket reader
+// the moment the job frame arrives.
+//
+// The gap between the two is the depth of this scanner's own queue multiplied
+// by the per-image scan time, and the hold cannot see into that queue. Without
+// this message its only clock is dispatch, so its scanning deadline budgets
+// queueing as well as scanning and cancels healthy work out from under a
+// backlogged scanner. A hold that does not know the message ignores it.
+type StartedMessage struct {
+ Type string `json:"type"` // "started"
+ Seq int64 `json:"seq"`
+}
+
// ResultMessage is sent from scanner to hold with scan results
type ResultMessage struct {
Type string `json:"type"` // "result"
]