test: production-shaped push/pull benchmark with per-backend request counts

TestBenchRealImages pushes and pulls three images whose layer sizes are
copied from real manifests in the production appview database (the median,
p75 and p90 images by layer count) and reports, per operation, wall time and
the number of requests to the registry, the fake PDS, the hold and S3, broken
down by endpoint. Skipped unless BENCH_PROFILES is set, so the integration
target does not run it. BENCH_LAT_{PDS,HOLD,S3} inject per-request latency,
which is what makes byte-path changes visible in-process; request counts are
the reliable signal either way.

internal/reqcount counts and delays requests through a handler wrapper and a
client-side RoundTripper. testharness.WithBackendTap wraps the PDS and S3
handlers and puts a counting reverse proxy in front of the hold;
testpds.WithMiddleware is the hook that makes the PDS side possible.

The bench showed a pull costs three hold calls per blob, not two: distribution
installs its notifications listener unconditionally and it re-Stats every blob
after ServeBlob to build the pull event. The backlog's presign memoization
item is rewritten with the measured numbers.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WTdBxLFU5TpwmqVdVsN1wq
This commit is contained in:
Evan Jarrett
2026-09-11 17:05:19 -05:00
co-authored by Claude Fable 5.1
parent 65db7945b2
commit bf4e63e810
5 changed files with 438 additions and 13 deletions
+18 -9
View File
@@ -28,20 +28,29 @@ are pipelined so one part is in flight while the next fills.
### Presign memoization within a request
**Problem.** Distribution calls `Stat` then `ServeBlob` for every blob GET and
HEAD. Each asks the hold for a presigned URL, so a client fetch costs two hold
calls, each doing full token validation and crew lookup.
**Problem.** Every blob GET costs three hold calls, each doing full token
validation, a captain record read and a presign. Distribution's blob handler
calls `Stat` then `ServeBlob`, and distribution's notifications listener
(`notifications.Listen` wraps every repository unconditionally, whether or not
any endpoint is configured) calls `Stat` a third time after `ServeBlob` to
build the pull event. Measured 2026-09-11 with `TestBenchRealImages`: a p90
pull of 22 layers is 22 registry GETs and 66 hold calls, 44 of them HEAD
presigns whose URL is never used.
**Where.** `ProxyBlobStore.Stat` and `ServeBlob` in
`pkg/appview/storage/proxy_blob_store.go`.
**Fix.** The store is built per request, so a small map keyed by digest and
method inside it is safe. Stat requests the URL and size once; ServeBlob reuses
the URL if the method matches. S3 signs the HTTP method, so a HEAD from Docker
cannot reuse a GET-signed URL; either key the cache by method or have the hold
return both URLs in one response.
**Fix.** The store is built per request, so a small memo keyed by digest and
method inside it is safe. Stat reads the request method from the context
(`storage.HTTPRequestMethod`, already set by the auth middleware), presigns
for that method when it is GET or HEAD, and remembers the URL and size.
ServeBlob reuses the URL when the digest and method match; the listener's
second Stat is answered from the memo. S3 signs the HTTP method, so a HEAD
URL cannot serve a GET; anything other than GET or HEAD must still presign
HEAD, since the hold refuses `method=PUT` on the read path.
**Impact.** One hold call per blob fetch instead of two.
**Impact.** One hold call per blob fetch instead of three. On the p90 pull
that is 66 hold calls down to 22.
### Stat from the appview's own layers table