Files
at-container-registry/docs/PERFORMANCE_BACKLOG.md
T

8.9 KiB

Performance Backlog

Findings from a September 2026 read of the appview and hold request paths, focused on the handshakes between the two services and on image uploads. The items that were worth doing immediately are listed first for context. The rest are recorded here so they are not lost. None of them is urgent; each is a known cost with a known fix.

All counts below come from reading the code, not from measurements, except the size distribution in the appendix, which was queried from the production appview database.

Done

Commit Change
9228579 One shared HTTP transport for the blob proxy instead of a new transport per registry request.
61a934d Hold-side auth looks a crew member up by record key instead of walking the whole crew collection.
034ea59 Hold reports blob size on read presigns; appview Stat no longer HEADs S3 for it.
f4343d7 Blobs under 16MB go up as one presigned PUT; every upload is hashed and verified against its digest.
47a1070 Process-wide budget on upload buffer memory; sweep for abandoned uploads; flush boundary pinned at exactly 16MB.
2a94f92 Per-request memo of blob presigns; a blob GET is one hold call, not three (Stat, ServeBlob and the notification listener's Stat shared it).

In flight as of 2026-09-09: the registry reads the sailor profile from the local users row instead of the owner's PDS on every request, and part uploads are pipelined so one part is in flight while the next fills.

Remaining

Stat from the appview's own layers table

Problem. For a layer of an already indexed manifest, the appview's layers table (pkg/appview/db/schema.sql) knows the digest and size, with an index on digest, yet Stat still asks the hold.

Fix. Answer Stat from the local row when the digest is known for that owner. Fall through to the hold on a miss: a freshly uploaded blob, or a blob whose manifest lives on a different hold, will not be in the table, and a miss must not become blob-unknown.

Impact. Zero network for Stat on known layers, including Docker's pre-upload existence checks for layers that were pushed before.

Caveat. Sizes in the table come from manifest descriptors the pusher declared, not from the stored object. The hold's records index has the same property since 034ea59.

Batch presigned part URLs from the hold

Problem. Each multipart part costs a hold round trip to fetch its presigned URL. Part numbers are predictable and the URLs are valid for 15 minutes.

Where. getPartUploadInfo in proxy_blob_store.go; HandleGetPartUploadURL in pkg/hold/oci/xrpc.go.

Fix. Let the hold return URLs for the next N parts from one call, either from initiateUpload or from a batched part URL endpoint. The writer consumes them in order and asks again when it runs out.

Impact. One hold round trip per 16MB removed. Does not change the shape of the transfer; the pipeline does that.

Hold complete handler: redundant HeadObject, and the copy

Problem. CompleteMultipartUploadWithManager in pkg/hold/oci/multipart.go does CompleteMultipartUpload, then HeadObject on the temp key, then CopyObject to the final key, then DeleteObject. The HeadObject buys nothing: CopyObject fails on a missing source. The copy scales with object size on most S3-compatible backends, and plain CopyObject fails above 5GB.

Fix. Drop the HeadObject. For the copy, either use multipart copy (UploadPartCopy) above the single-copy limit, or find a way to know the final key before the first part lands. The digest is only known at the final PUT, so the second option needs Docker's monolithic upload path or a client hint and is not a small change.

Impact. One S3 call per multipart upload now; the copy cost only matters for very large layers, which are about 1% of distinct layers but 11% of bytes.

One signed repo commit per layer record on notify

Problem. HandleNotifyManifest in pkg/hold/oci/xrpc.go loops CreateLayerRecord once per layer, each a full repo commit under the user lock with its own signature and firehose event. Stats and daily stats add two or three more commits. A 10-layer push is about 13 sequential commits.

Fix. BatchCreateLayerRecords in pkg/hold/pds/layer.go already exists and is used only by GC. Use it here. Consider folding the stats updates into the same batch write.

Impact. Two or three commits per push instead of one per layer. Also fewer firehose events for relays to ingest.

Layer-record dedup scans every layer record on the hold

Problem. ListLayerRecordsForManifest in pkg/hold/pds/layer.go pages the entire layer collection through the records index, fetches and CBOR-decodes each record from the CAR store, and filters by manifest URI in memory. It runs once per push notify. Cost is proportional to the total number of layer records on the hold, not to the manifest.

Fix. The records index (pkg/hold/pds/records.go) already has a did column with an index, so ListRecordsByDID would narrow the scan to the pushing user's records. Adding a manifest column with an index makes it a single indexed query.

Impact. Proportional to one manifest instead of the whole hold. Matters most on the shared hold with many crew.

Config blob fetched twice per push

Problem. On manifest PUT the appview fetches the config blob through the hold to extract labels (extractConfigLabels in pkg/appview/storage/manifest_store.go). The hold then fetches the same blob from S3 in notify to create the image config record.

Fix. Include the config bytes, or just the parsed labels, in the notify payload, or have the hold cache what it fetched. Small either way.

Successor drain on every push

Problem. MigrateManifestsForSuccessor in pkg/appview/storage/drain.go runs in the background on every push. It is guarded per DID but still fetches the profile and queries the database before discovering that no hold has a successor.

Fix. Check the local captain record cache for any successor first, which is a cheap query, and skip the rest when there is none.

Hold multipart sessions expire without an S3 abort

Problem. MultipartManager.cleanupExpiredSessions in pkg/hold/oci/multipart.go drops its in-memory session after 24 hours of inactivity but does not abort the S3 multipart. Parts from a session the appview never reaped, for example across an appview restart, stay in the bucket.

Fix. Abort in S3 when the session expires, or add a bucket lifecycle rule for incomplete multipart uploads. The appview's sweep (47a1070) covers the common case; this is the backstop.

Pipeline depth

Problem. Once part uploads are pipelined with one part in flight, a second in-flight part could overlap more on high-latency links.

Fix. Generalize the two-buffer swap to N buffers with ordered completion. Budget cost is N buffers per large upload.

Impact. Uncertain; measure with one in flight first.

Not performance, found along the way

Crew edit handler deletes the member it just updated

pkg/hold/admin/handlers_crew.go, around the role and permissions change: the handler calls AddCrewMember, which upserts at the deterministic record key for the member DID, then RemoveCrewMember with the record key from the URL. For any member created since the deterministic key scheme, those keys are the same, so the delete removes the record that was just written and the member disappears. The comment above it about a transient duplicate describes the earlier TID-keyed world. Fix: skip the delete when the URL key equals the deterministic key.

Single Write larger than the budget

Put in proxy_blob_store.go hands the writer a whole blob in one slice. Since 47a1070 the writer splits it into threshold-sized parts, so this no longer buffers the whole slice; noted here only because the earlier behavior was referenced in commit messages.

Appendix: production size distribution

Queried read-only from the appview database on 2026-09-09. Sizes come from manifest layer descriptors.

Distinct layer digests: 9,286. Manifests: 4,203, of which 3,357 have layers of their own (the rest are manifest lists and indexes).

Threshold Distinct layers under it Bytes under it Manifests with every layer under it
5 MB 72% 4%
10 MB 79% 10% 43%
16 MB 86% 21% 49%
32 MB 93% 38% 70%
64 MB 99% 64% 96%

Config blobs average about 2 KB; the largest is under 30 KB.

Reading: blob count drives the fixed per-blob overhead, and 86% of blobs skip multipart at 16MB. Bytes drive transfer time, and 79% of bytes are in layers over 16MB, so large-layer improvements (pipelining, part URL batching, the copy on complete) are what move push time for big images. The largest layer of an image clusters at 16 to 64 MB, which is where base image layers land, so raising the threshold to 32MB would make 70% of images all-direct-PUT at the cost of 32MB of buffer per in-flight large upload.