Files
at-container-registry/docs/PERFORMANCE_BACKLOG.md
T
Evan JarrettandClaude Fable 5.1 8b5f195f09 docs: record the remaining performance findings
The September 2026 read of the appview and hold request paths produced
more findings than were worth doing immediately. This records what
shipped, what remains with problem, location, fix and expected impact,
one non-performance bug found along the way, and the production size
distribution that informed the 16MB direct-PUT threshold.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018Yf1ZVA7sXYhQNb9tCo1m5
2026-09-09 20:42:58 -05:00

214 lines
9.5 KiB
Markdown

# Performance Backlog
Findings from a September 2026 read of the appview and hold request paths,
focused on the handshakes between the two services and on image uploads. The
items that were worth doing immediately are listed first for context. The rest
are recorded here so they are not lost. None of them is urgent; each is a known
cost with a known fix.
All counts below come from reading the code, not from measurements, except the
size distribution in the appendix, which was queried from the production
appview database.
## Done
| Commit | Change |
|---|---|
| `9228579` | One shared HTTP transport for the blob proxy instead of a new transport per registry request. |
| `61a934d` | Hold-side auth looks a crew member up by record key instead of walking the whole crew collection. |
| `034ea59` | Hold reports blob size on read presigns; appview Stat no longer HEADs S3 for it. |
| `f4343d7` | Blobs under 16MB go up as one presigned PUT; every upload is hashed and verified against its digest. |
| `47a1070` | Process-wide budget on upload buffer memory; sweep for abandoned uploads; flush boundary pinned at exactly 16MB. |
In flight as of 2026-09-09: the registry reads the sailor profile from the
local users row instead of the owner's PDS on every request, and part uploads
are pipelined so one part is in flight while the next fills.
## Remaining
### Presign memoization within a request
**Problem.** Distribution calls `Stat` then `ServeBlob` for every blob GET and
HEAD. Each asks the hold for a presigned URL, so a client fetch costs two hold
calls, each doing full token validation and crew lookup.
**Where.** `ProxyBlobStore.Stat` and `ServeBlob` in
`pkg/appview/storage/proxy_blob_store.go`.
**Fix.** The store is built per request, so a small map keyed by digest and
method inside it is safe. Stat requests the URL and size once; ServeBlob reuses
the URL if the method matches. S3 signs the HTTP method, so a HEAD from Docker
cannot reuse a GET-signed URL; either key the cache by method or have the hold
return both URLs in one response.
**Impact.** One hold call per blob fetch instead of two.
### Stat from the appview's own layers table
**Problem.** For a layer of an already indexed manifest, the appview's `layers`
table (`pkg/appview/db/schema.sql`) knows the digest and size, with an index on
digest, yet Stat still asks the hold.
**Fix.** Answer Stat from the local row when the digest is known for that
owner. Fall through to the hold on a miss: a freshly uploaded blob, or a blob
whose manifest lives on a different hold, will not be in the table, and a miss
must not become blob-unknown.
**Impact.** Zero network for Stat on known layers, including Docker's
pre-upload existence checks for layers that were pushed before.
**Caveat.** Sizes in the table come from manifest descriptors the pusher
declared, not from the stored object. The hold's records index has the same
property since `034ea59`.
### Batch presigned part URLs from the hold
**Problem.** Each multipart part costs a hold round trip to fetch its presigned
URL. Part numbers are predictable and the URLs are valid for 15 minutes.
**Where.** `getPartUploadInfo` in `proxy_blob_store.go`;
`HandleGetPartUploadURL` in `pkg/hold/oci/xrpc.go`.
**Fix.** Let the hold return URLs for the next N parts from one call, either
from initiateUpload or from a batched part URL endpoint. The writer consumes
them in order and asks again when it runs out.
**Impact.** One hold round trip per 16MB removed. Does not change the shape of
the transfer; the pipeline does that.
### Hold complete handler: redundant HeadObject, and the copy
**Problem.** `CompleteMultipartUploadWithManager` in
`pkg/hold/oci/multipart.go` does CompleteMultipartUpload, then HeadObject on
the temp key, then CopyObject to the final key, then DeleteObject. The
HeadObject buys nothing: CopyObject fails on a missing source. The copy scales
with object size on most S3-compatible backends, and plain CopyObject fails
above 5GB.
**Fix.** Drop the HeadObject. For the copy, either use multipart copy
(UploadPartCopy) above the single-copy limit, or find a way to know the final
key before the first part lands. The digest is only known at the final PUT, so
the second option needs Docker's monolithic upload path or a client hint and is
not a small change.
**Impact.** One S3 call per multipart upload now; the copy cost only matters
for very large layers, which are about 1% of distinct layers but 11% of bytes.
### One signed repo commit per layer record on notify
**Problem.** `HandleNotifyManifest` in `pkg/hold/oci/xrpc.go` loops
`CreateLayerRecord` once per layer, each a full repo commit under the user lock
with its own signature and firehose event. Stats and daily stats add two or
three more commits. A 10-layer push is about 13 sequential commits.
**Fix.** `BatchCreateLayerRecords` in `pkg/hold/pds/layer.go` already exists
and is used only by GC. Use it here. Consider folding the stats updates into
the same batch write.
**Impact.** Two or three commits per push instead of one per layer. Also fewer
firehose events for relays to ingest.
### Layer-record dedup scans every layer record on the hold
**Problem.** `ListLayerRecordsForManifest` in `pkg/hold/pds/layer.go` pages the
entire layer collection through the records index, fetches and CBOR-decodes
each record from the CAR store, and filters by manifest URI in memory. It runs
once per push notify. Cost is proportional to the total number of layer records
on the hold, not to the manifest.
**Fix.** The records index (`pkg/hold/pds/records.go`) already has a `did`
column with an index, so `ListRecordsByDID` would narrow the scan to the
pushing user's records. Adding a `manifest` column with an index makes it a
single indexed query.
**Impact.** Proportional to one manifest instead of the whole hold. Matters
most on the shared hold with many crew.
### Config blob fetched twice per push
**Problem.** On manifest PUT the appview fetches the config blob through the
hold to extract labels (`extractConfigLabels` in
`pkg/appview/storage/manifest_store.go`). The hold then fetches the same blob
from S3 in notify to create the image config record.
**Fix.** Include the config bytes, or just the parsed labels, in the notify
payload, or have the hold cache what it fetched. Small either way.
### Successor drain on every push
**Problem.** `MigrateManifestsForSuccessor` in
`pkg/appview/storage/drain.go` runs in the background on every push. It is
guarded per DID but still fetches the profile and queries the database before
discovering that no hold has a successor.
**Fix.** Check the local captain record cache for any successor first, which
is a cheap query, and skip the rest when there is none.
### Hold multipart sessions expire without an S3 abort
**Problem.** `MultipartManager.cleanupExpiredSessions` in
`pkg/hold/oci/multipart.go` drops its in-memory session after 24 hours of
inactivity but does not abort the S3 multipart. Parts from a session the
appview never reaped, for example across an appview restart, stay in the
bucket.
**Fix.** Abort in S3 when the session expires, or add a bucket lifecycle rule
for incomplete multipart uploads. The appview's sweep (`47a1070`) covers the
common case; this is the backstop.
### Pipeline depth
**Problem.** Once part uploads are pipelined with one part in flight, a second
in-flight part could overlap more on high-latency links.
**Fix.** Generalize the two-buffer swap to N buffers with ordered completion.
Budget cost is N buffers per large upload.
**Impact.** Uncertain; measure with one in flight first.
## Not performance, found along the way
### Crew edit handler deletes the member it just updated
`pkg/hold/admin/handlers_crew.go`, around the role and permissions change:
the handler calls `AddCrewMember`, which upserts at the deterministic record
key for the member DID, then `RemoveCrewMember` with the record key from the
URL. For any member created since the deterministic key scheme, those keys are
the same, so the delete removes the record that was just written and the
member disappears. The comment above it about a transient duplicate describes
the earlier TID-keyed world. Fix: skip the delete when the URL key equals the
deterministic key.
### Single Write larger than the budget
`Put` in `proxy_blob_store.go` hands the writer a whole blob in one slice.
Since `47a1070` the writer splits it into threshold-sized parts, so this no
longer buffers the whole slice; noted here only because the earlier behavior
was referenced in commit messages.
## Appendix: production size distribution
Queried read-only from the appview database on 2026-09-09. Sizes come from
manifest layer descriptors.
Distinct layer digests: 9,286. Manifests: 4,203, of which 3,357 have layers of
their own (the rest are manifest lists and indexes).
| Threshold | Distinct layers under it | Bytes under it | Manifests with every layer under it |
|---|---|---|---|
| 5 MB | 72% | 4% | |
| 10 MB | 79% | 10% | 43% |
| 16 MB | 86% | 21% | 49% |
| 32 MB | 93% | 38% | 70% |
| 64 MB | 99% | 64% | 96% |
Config blobs average about 2 KB; the largest is under 30 KB.
Reading: blob count drives the fixed per-blob overhead, and 86% of blobs skip
multipart at 16MB. Bytes drive transfer time, and 79% of bytes are in layers
over 16MB, so large-layer improvements (pipelining, part URL batching, the
copy on complete) are what move push time for big images. The largest layer of
an image clusters at 16 to 64 MB, which is where base image layers land, so
raising the threshold to 32MB would make 70% of images all-direct-PUT at the
cost of 32MB of buffer per in-flight large upload.