Files
at-container-registry/docs/PERFORMANCE_BACKLOG.md
T
Evan JarrettandClaude Fable 5.1 8b5f195f09 docs: record the remaining performance findings
The September 2026 read of the appview and hold request paths produced
more findings than were worth doing immediately. This records what
shipped, what remains with problem, location, fix and expected impact,
one non-performance bug found along the way, and the production size
distribution that informed the 16MB direct-PUT threshold.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018Yf1ZVA7sXYhQNb9tCo1m5
2026-09-09 20:42:58 -05:00

9.5 KiB

Performance Backlog

Findings from a September 2026 read of the appview and hold request paths, focused on the handshakes between the two services and on image uploads. The items that were worth doing immediately are listed first for context. The rest are recorded here so they are not lost. None of them is urgent; each is a known cost with a known fix.

All counts below come from reading the code, not from measurements, except the size distribution in the appendix, which was queried from the production appview database.

Done

Commit Change
9228579 One shared HTTP transport for the blob proxy instead of a new transport per registry request.
61a934d Hold-side auth looks a crew member up by record key instead of walking the whole crew collection.
034ea59 Hold reports blob size on read presigns; appview Stat no longer HEADs S3 for it.
f4343d7 Blobs under 16MB go up as one presigned PUT; every upload is hashed and verified against its digest.
47a1070 Process-wide budget on upload buffer memory; sweep for abandoned uploads; flush boundary pinned at exactly 16MB.

In flight as of 2026-09-09: the registry reads the sailor profile from the local users row instead of the owner's PDS on every request, and part uploads are pipelined so one part is in flight while the next fills.

Remaining

Presign memoization within a request

Problem. Distribution calls Stat then ServeBlob for every blob GET and HEAD. Each asks the hold for a presigned URL, so a client fetch costs two hold calls, each doing full token validation and crew lookup.

Where. ProxyBlobStore.Stat and ServeBlob in pkg/appview/storage/proxy_blob_store.go.

Fix. The store is built per request, so a small map keyed by digest and method inside it is safe. Stat requests the URL and size once; ServeBlob reuses the URL if the method matches. S3 signs the HTTP method, so a HEAD from Docker cannot reuse a GET-signed URL; either key the cache by method or have the hold return both URLs in one response.

Impact. One hold call per blob fetch instead of two.

Stat from the appview's own layers table

Problem. For a layer of an already indexed manifest, the appview's layers table (pkg/appview/db/schema.sql) knows the digest and size, with an index on digest, yet Stat still asks the hold.

Fix. Answer Stat from the local row when the digest is known for that owner. Fall through to the hold on a miss: a freshly uploaded blob, or a blob whose manifest lives on a different hold, will not be in the table, and a miss must not become blob-unknown.

Impact. Zero network for Stat on known layers, including Docker's pre-upload existence checks for layers that were pushed before.

Caveat. Sizes in the table come from manifest descriptors the pusher declared, not from the stored object. The hold's records index has the same property since 034ea59.

Batch presigned part URLs from the hold

Problem. Each multipart part costs a hold round trip to fetch its presigned URL. Part numbers are predictable and the URLs are valid for 15 minutes.

Where. getPartUploadInfo in proxy_blob_store.go; HandleGetPartUploadURL in pkg/hold/oci/xrpc.go.

Fix. Let the hold return URLs for the next N parts from one call, either from initiateUpload or from a batched part URL endpoint. The writer consumes them in order and asks again when it runs out.

Impact. One hold round trip per 16MB removed. Does not change the shape of the transfer; the pipeline does that.

Hold complete handler: redundant HeadObject, and the copy

Problem. CompleteMultipartUploadWithManager in pkg/hold/oci/multipart.go does CompleteMultipartUpload, then HeadObject on the temp key, then CopyObject to the final key, then DeleteObject. The HeadObject buys nothing: CopyObject fails on a missing source. The copy scales with object size on most S3-compatible backends, and plain CopyObject fails above 5GB.

Fix. Drop the HeadObject. For the copy, either use multipart copy (UploadPartCopy) above the single-copy limit, or find a way to know the final key before the first part lands. The digest is only known at the final PUT, so the second option needs Docker's monolithic upload path or a client hint and is not a small change.

Impact. One S3 call per multipart upload now; the copy cost only matters for very large layers, which are about 1% of distinct layers but 11% of bytes.

One signed repo commit per layer record on notify

Problem. HandleNotifyManifest in pkg/hold/oci/xrpc.go loops CreateLayerRecord once per layer, each a full repo commit under the user lock with its own signature and firehose event. Stats and daily stats add two or three more commits. A 10-layer push is about 13 sequential commits.

Fix. BatchCreateLayerRecords in pkg/hold/pds/layer.go already exists and is used only by GC. Use it here. Consider folding the stats updates into the same batch write.

Impact. Two or three commits per push instead of one per layer. Also fewer firehose events for relays to ingest.

Layer-record dedup scans every layer record on the hold

Problem. ListLayerRecordsForManifest in pkg/hold/pds/layer.go pages the entire layer collection through the records index, fetches and CBOR-decodes each record from the CAR store, and filters by manifest URI in memory. It runs once per push notify. Cost is proportional to the total number of layer records on the hold, not to the manifest.

Fix. The records index (pkg/hold/pds/records.go) already has a did column with an index, so ListRecordsByDID would narrow the scan to the pushing user's records. Adding a manifest column with an index makes it a single indexed query.

Impact. Proportional to one manifest instead of the whole hold. Matters most on the shared hold with many crew.

Config blob fetched twice per push

Problem. On manifest PUT the appview fetches the config blob through the hold to extract labels (extractConfigLabels in pkg/appview/storage/manifest_store.go). The hold then fetches the same blob from S3 in notify to create the image config record.

Fix. Include the config bytes, or just the parsed labels, in the notify payload, or have the hold cache what it fetched. Small either way.

Successor drain on every push

Problem. MigrateManifestsForSuccessor in pkg/appview/storage/drain.go runs in the background on every push. It is guarded per DID but still fetches the profile and queries the database before discovering that no hold has a successor.

Fix. Check the local captain record cache for any successor first, which is a cheap query, and skip the rest when there is none.

Hold multipart sessions expire without an S3 abort

Problem. MultipartManager.cleanupExpiredSessions in pkg/hold/oci/multipart.go drops its in-memory session after 24 hours of inactivity but does not abort the S3 multipart. Parts from a session the appview never reaped, for example across an appview restart, stay in the bucket.

Fix. Abort in S3 when the session expires, or add a bucket lifecycle rule for incomplete multipart uploads. The appview's sweep (47a1070) covers the common case; this is the backstop.

Pipeline depth

Problem. Once part uploads are pipelined with one part in flight, a second in-flight part could overlap more on high-latency links.

Fix. Generalize the two-buffer swap to N buffers with ordered completion. Budget cost is N buffers per large upload.

Impact. Uncertain; measure with one in flight first.

Not performance, found along the way

Crew edit handler deletes the member it just updated

pkg/hold/admin/handlers_crew.go, around the role and permissions change: the handler calls AddCrewMember, which upserts at the deterministic record key for the member DID, then RemoveCrewMember with the record key from the URL. For any member created since the deterministic key scheme, those keys are the same, so the delete removes the record that was just written and the member disappears. The comment above it about a transient duplicate describes the earlier TID-keyed world. Fix: skip the delete when the URL key equals the deterministic key.

Single Write larger than the budget

Put in proxy_blob_store.go hands the writer a whole blob in one slice. Since 47a1070 the writer splits it into threshold-sized parts, so this no longer buffers the whole slice; noted here only because the earlier behavior was referenced in commit messages.

Appendix: production size distribution

Queried read-only from the appview database on 2026-09-09. Sizes come from manifest layer descriptors.

Distinct layer digests: 9,286. Manifests: 4,203, of which 3,357 have layers of their own (the rest are manifest lists and indexes).

Threshold Distinct layers under it Bytes under it Manifests with every layer under it
5 MB 72% 4%
10 MB 79% 10% 43%
16 MB 86% 21% 49%
32 MB 93% 38% 70%
64 MB 99% 64% 96%

Config blobs average about 2 KB; the largest is under 30 KB.

Reading: blob count drives the fixed per-blob overhead, and 86% of blobs skip multipart at 16MB. Bytes drive transfer time, and 79% of bytes are in layers over 16MB, so large-layer improvements (pipelining, part URL batching, the copy on complete) are what move push time for big images. The largest layer of an image clusters at 16 to 64 MB, which is where base image layers land, so raising the threshold to 32MB would make 70% of images all-direct-PUT at the cost of 32MB of buffer per in-flight large upload.