Files
at-container-registry/pkg
Evan JarrettandClaude Opus 5 017755c6d4 hold/gc: stop the aux-record sweep deleting live records
Follow-up to f12d5e0. Extending the orphan sweep to io.atcr.hold.scan and
io.atcr.hold.image.config brought those collections under a delete path built
for layer records, and the two are not interchangeable. Layer rkeys are
generated per write, so a remembered rkey is a stable handle on one record.
Scan and image-config rkeys are atproto.ScanRecordKey(manifestDigest) — the bare
digest hex, with no DID — which is both reused across re-pushes and shared
between users who push identical images. "Delete rkey K" therefore stopped
meaning "delete the record I judged", and the sweep could destroy live data two
ways.

Both were confirmed by probe before being fixed, and each guard below has a
regression test that fails when that guard alone is reverted.

Shared records: two users pushing the same image collapse onto ONE record whose
body names only the last writer. When that user deleted their manifest the
record looked orphaned, and deleting it stripped the vulnerability scan and the
layer history/env/entrypoint from every co-owner's still-live image.

  - knownDigests keeps a record while any successfully fetched user still holds
    a manifest at that digest, not just the one the record names.

  - knownDigests alone is not enough: it is built only from users we reached, so
    a co-owner whose PDS was down this run is indistinguishable from one who
    deleted their image. digestOwners closes that. It maps each digest to every
    DID that pushed it here, derived from the hold's own layer records during
    the walk analyzeRecords already performs, so it costs nothing extra and
    needs no network. If any owner was unreachable, the record stays. This is
    the sweep's existing principle — an unreachable PDS never causes a deletion —
    extended from the record's named user to everyone the record serves.

    Layer records are the right source because they are load-bearing for storage
    accounting and billing, so they exist for anything a user is charged for.

Reused rkeys: a re-push upserts into the same digest-derived slot, so a record
marked orphaned by an earlier scan can be replaced by a LIVE one before the
delete runs. The admin delete button makes this wide — lastPreview is in-memory
for the life of the process, so an open tab keeps a stale scan actionable — and
doRun has the same race across its analyze-to-delete gap.

  - Records now carry the CID they had when judged, and the delete re-reads the
    slot to confirm it still holds that revision. Anything else (rewritten,
    missing, unreadable, or a ref with no recorded CID) is skipped, not deleted.

  - The CID check is record identity, not orphanhood, and identity alone is not
    enough either. A re-push does not necessarily rewrite the SCAN record:
    scan-on-push is tier-gated, and the broadcaster's discovery loop skips any
    manifest that already has a scan record, so its CID survives a re-push
    unchanged. manifestsWithLayerRecords covers that — a push writes layer
    records to this hold, so their presence proves the manifest is back
    regardless of what the scan concluded. Local, no PDS round trip. A manifest
    whose layer records are themselves orphaned but not yet collected also lands
    in the set, which only defers its aux record one cycle; one more round of a
    leak beats deleting a live user's records.

  - maxPreviewAgeForDelete (30m) refuses a stale preview outright. The
    per-record checks are the real safety net; this stops the admin acting on a
    picture that is hours old. Refusal surfaces through startBackground ->
    lastError -> the polling progress fragment, so it is visible rather than a
    silent no-op.

The four state maps auxRecordOrphaned consults are grouped into auxOrphanState.
Three are same-shaped maps that were trivial to transpose positionally, and
swapping knownDigests with fetchedUsers would have silently widened what the
sweep deletes.

Not addressed, all failure-to-collect rather than data loss: maxPreviewItems
caps the admin button's list while doRun uses the uncapped one, so a hold with
10000+ orphaned layer records surfaces no aux orphans in the preview;
DeleteManifestAuxRecord cannot distinguish "already gone" from "write failed";
and discoverUserDIDs never consults the aux collections, so a record whose owner
has no remaining layer records is kept indefinitely.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-08 21:24:00 -05:00
..
2026-05-11 19:53:13 -05:00