Design for moving push bandwidth off the appview and onto the hold, mirroring the pull path's 307 redirect. Proposal only, not implemented. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
21 KiB
Push Offload: Shifting Blob Uploads to the Hold
Status: Proposal / design. Not yet implemented.
This document describes shifting OCI blob upload (push) bandwidth off the AppView
and onto the hold, mirroring what the pull path already does for reads. It is the
write-side counterpart to the existing pull 307 redirect.
Motivation
The asymmetry
The hold design is deliberately lightweight: it handshakes a few HTTP requests and
hands back presigned S3 URLs. On the read path this works perfectly. A pull is a
bodyless GET that AppView answers with a 307 to a presigned S3 URL
(proxy_blob_store.go:267-281 ServeBlob). Neither AppView nor the hold ever carries
a blob byte, and reads are 10-100x more frequent than writes, so this is where the
property matters most.
The write path inverts this. Today a push streams:
client --PATCH/PUT--> AppView (buffers 10MB chunks in RAM) --presigned PUT--> S3
AppView ingests every layer and re-uploads it to S3 (proxy_blob_store.go:586-665
Write/flushPart). It carries the full upload bandwidth even though the hold and S3
are the actual storage.
Why this is economically backwards for BYOS
The entire point of BYOS is that a user runs their own hold and storage so the AppView operator is not on the hook for their data. But on push, a BYOS upload goes:
client --bytes--> AppView (operator's bandwidth) --> user's own S3
The operator pays ingress + egress to shuttle bytes into a bucket they don't own and will never bill for. The people most motivated to self-host (to avoid paying the operator) are exactly the ones costing the operator bandwidth on every push. The BYOS bargain runs backwards on writes.
Why pure client -> S3 direct is not possible on push
It is tempting to do for push what pull does: redirect the client straight to a
presigned S3 URL. This cannot work with stock OCI clients (docker, containerd,
podman), for reasons independent of each other:
- Method preservation kills PATCH. A
307/308preserves method and body. The chunked path usesPATCH; S3 has noPATCH(object API isPUTor the multipartPOST+UploadIdflow). A redirectedPATCH405s. - Streamed bodies cannot be replayed. Following a redirect on a write requires
re-sending the body; Go's
http.Clientonly does so whenRequest.GetBodyis set (rewindable). Docker streams the layer tar, which generally is not rewindable. - The finalize response contract. The closing
PUT <location>?digest=sha256:...expects201+Docker-Content-Digest+Location. S3 returns 200/XML; the client chokes. The client also appends?digest=to the Location, invalidating any SigV4 presigned URL used as the Location. - Multipart coordination cannot ride a redirect. Blobs >5GB need S3 multipart, which needs an interactive init -> per-part-URL -> complete loop. A redirect cannot express that handshake.
The signature/size question ("you need to know the size before you sign") is a red
herring: S3 lets you presign a single PUT with UNSIGNED-PAYLOAD (no size needed,
5GB cap), and POST Object policies support a content-length-range (min/max). The
real blocker is that the thing terminating an OCI upload must speak OCI and return
OCI responses -- which S3 cannot do.
The "pick two" on push
On push you cannot have all three at once:
- Lightweight hold (never touches bytes)
- Stock
docker push(no custom client) - AppView free of the byte path
- (1)+(2) -> AppView eats the bandwidth. (Where we are today.)
- (2)+(3) -> the hold terminates the upload and carries the bytes. <- this proposal
- (1)+(3) -> a custom client streams straight to S3, no
docker push.
Pull gets all three only because a bodyless GET redirect satisfies them simultaneously. Since the whole reason for BYOS is "don't pay the operator," the correct trade is (2)+(3): make pushes terminate on the hold, where the bandwidth belongs.
Goal and non-goals
Goal: On push, the layer bytes go client -> hold -> S3 and never transit AppView.
The handshake (POST .../blobs/uploads/) and the manifest PUT stay on AppView (cheap,
no bytes).
Non-goals:
- Client -> S3 direct on push (impossible with stock clients, see above).
- Changing the pull path (already optimal).
- A custom push client (out of scope; would be the only way to get (1)+(3)).
The hold "carrying bytes" costs bandwidth, not its lightweight character:
- Monolithic push (docker sends
Content-Length): the hold does a streamingPutObject(io.Copypassthrough) -- no disk, trivial RAM. - Large/chunked push (>5GB or chunked): multipart with ~5MB part buffers -- still trivial memory.
The hold stays memory-light and ops-simple; it just becomes bandwidth-heavy on writes, on the BYOS owner's box.
Current state (for reference)
| Concern | Today |
|---|---|
| Push handshake | AppView POST .../blobs/uploads/ -> distribution lib calls ProxyBlobStore.Create -> XRPC io.atcr.hold.initiateUpload (proxy_blob_store.go:287-325) |
| Push bytes | Client -> AppView RAM (10MB chunks) -> presigned S3 PUT (proxy_blob_store.go:605-665) |
| Push finalize | ProxyBlobWriter.Commit -> XRPC io.atcr.hold.completeUpload (proxy_blob_store.go:703-741) |
| Hold upload API | Custom XRPC only: initiateUpload, getPartUploadUrl, completeUpload, abortUpload, notifyManifest (pkg/hold/oci/xrpc.go:50-61). No standard OCI /v2 upload surface. |
Upload Location |
AppView-relative /v2/<name>/blobs/uploads/<id>, generated by the distribution library from BlobWriter.ID(), not by ATCR code. |
| Hold auth | Service token (Bearer, aud=hold DID, signed by user's PDS) or DPoP. Validated on every op (pkg/hold/pds/auth.go:377-429 ValidateBlobWriteAccess, :507-608 ValidateServiceToken). The hold cannot validate AppView's registry JWT. |
| Pull (contrast) | ServeBlob -> 307 to presigned S3 URL; client fetches direct (proxy_blob_store.go:267-281). |
Note: an unused HoldUploadPart = "/xrpc/io.atcr.hold.uploadPart" constant already
exists (pkg/atproto/endpoints.go:29) for "direct buffered part uploads" -- prior art
for the hold carrying bytes, never registered.
Proposed design
The flow the operator described:
AppView looks at what hold it needs to go to, checks if that hold is anonymous-push or not, does a serviceAuth request if necessary, then lets the hold take it from there.
1. docker: POST https://<appview>/v2/<identity>/<image>/blobs/uploads/
2. AppView: resolve hold DID (findHoldDIDAndProfile + resolveSuccessor)
check hold push policy (anonymous push allowed?)
if auth required: mint/fetch service token (aud=hold DID) via user OAuth
-> 202 Accepted
Location: https://<hold-public-url>/v2/<identity>/<image>/blobs/uploads/<session>?_t=<serviceToken>
3. docker: PATCH/PUT the layer bytes to that Location (client -> hold)
4. hold: validate service token, bind session to user
stream bytes to its S3 (passthrough or 5MB multipart)
finalize -> 201 Docker-Content-Digest: sha256:... (hold -> S3, OCI response)
5. docker: PUT https://<appview>/v2/<identity>/<image>/manifests/<ref>
6. AppView: verify referenced blobs exist on the hold, store manifest in user's PDS
AppView is out of the byte path entirely (steps 3-4). It keeps only the cheap metadata operations (steps 1-2, 5-6).
Why the handshake stays on AppView
Hold resolution and the serviceAuth request both require AppView's OAuth session for the
user (only AppView holds it). So the POST .../blobs/uploads/ must be answered by
AppView. Everything after the 202 lives on the hold.
Cross-host Location is spec-legal and opaque to the client
The OCI distribution spec permits an absolute, cross-host Location, and the client
treats the returned session URL as opaque -- it PATCH/PUTs to it verbatim. So:
- The hold's session URL does not need to be a literal
/v2/...path; the hold may namespace it however it likes, as long as its responses are OCI-compliant (202+LocationforPATCHcontinuation,201+Docker-Content-Digestfor the finalizePUT). - Only the finalize
PUTappends?digest=-- the hold parses this normally (it is not S3, so the appended query breaks nothing).
Authentication model
This is the crux, and the operator's proposed flow resolves it cleanly.
Service token in the Location URL (chosen approach)
AppView mints a service token during the handshake (the same machinery used today
for AppView's own XRPC calls: GetOrFetchServiceToken, pkg/auth/servicetoken.go:53,
getServiceAuth?aud=<holdDID>&lxm=<method>&exp=<ts>) and embeds it in the upload
Location.
Why this works and the registry JWT does not:
- The hold already validates service tokens (
ValidateServiceToken,pkg/hold/pds/auth.go:507): checksaud== hold DID,exp, and the signature against the user's DID document. No new trust infrastructure. - The docker client only holds an AppView-issued registry JWT, which the hold cannot validate. By putting the credential in the Location URL, the client does not need to know anything about hold auth -- it just uses the URL.
- Cross-host auth headers: docker generally does not forward its
Authorizationbearer to a different host on a redirected upload (credential-leak avoidance). The token-in-URL avoids depending on header forwarding entirely.
Trade-off: the token appears in the URL (and thus potentially in hold access logs). It is short-lived; mitigations below.
Session-bound authorization (handles token expiry mid-upload)
Service tokens are short-lived (~5 min typical; reference PDSes grant up to 1h). A large push could outlive the token. To avoid re-validating on every chunk:
- The hold validates the token once, when the upload session is created (or first byte-bearing request arrives), and binds the session to the authorized user.
- Subsequent
PATCH/PUTon that session are authorized by session ownership, not by re-checking the token. The session ID is the capability. - Request a longer
expfor upload tokens where the PDS allows it.
This means token expiry mid-upload does not abort an in-flight push.
Anonymous / managed-push policy
The handshake checks the hold's push policy before deciding whether to mint a token:
- Auth required (default): mint a service token; embed in Location. The hold enforces
captain/crew
blob:writeexactly as today (ValidateBlobWriteAccess). - Anonymous push allowed: return a tokenless Location; the hold accepts the upload without a service token.
Note the captain record's existing Public flag governs anonymous reads only
(pkg/hold/pds/captain.go:24, auth.go:445); writes always require auth today.
"Anonymous push" is a new policy bit (see Open Questions). Most holds will keep
auth-required; anonymous push is opt-in for, e.g., open CI mirrors.
Rejected alternative: JWKS on AppView
AppView could publish a JWKS so the hold validates AppView's registry JWT directly. This needs new trust plumbing (JWKS endpoint, hold-side fetch + trust config) and still requires the client to send the registry JWT cross-host (which it avoids doing). The service-token approach reuses validation the hold already performs, so JWKS is not pursued.
AppView changes
| Change | Where |
|---|---|
POST .../blobs/uploads/ returns a cross-host hold Location instead of driving ProxyBlobStore.Create |
new handler ahead of the distribution /v2 handler; see "Distribution interaction" |
| Hold push-policy check during handshake | findHoldDIDAndProfile result + captain push policy (Jetstream-fed local table, keep it local-fast) |
| Mint service token for upload, embed in Location | reuse GetOrFetchServiceToken (pkg/auth/servicetoken.go) |
HEAD .../blobs/<digest> (existence / layer skip) stays on AppView |
answer from local layer metadata / hold query; keeps docker's skip-existing fast |
Manifest PUT stays on AppView; verify referenced blobs exist on the hold before storing |
manifest_store.go + a hold existence check |
ProxyBlobWriter byte path becomes unused for offloaded pushes |
proxy_blob_store.go (see below) |
Distribution library interaction
The distribution library currently generates the upload Location itself and owns the
PATCH/PUT session. Two ways to take over the handshake:
- Intercept the upload-init route ahead of distribution: register
POST /v2/<name>/blobs/uploads/on ATCR's router so it short-circuits distribution, returning the cross-host hold Location. The subsequentPATCH/PUTgo to the hold (different host), so distribution's blob-upload machinery is never invoked for these pushes. Cleanest incremental path. - Do it as part of
docs/REMOVING_DISTRIBUTION.md. Push-offload makesproxy_blob_store.go'sBlobWriter-- the doc's single gnarliest distribution impl -- obsolete for BYOS. The two efforts reinforce each other: owning the/v2HTTP layer makes emitting a cross-host upload Location trivial, and push-offload removes the hardest reason theBlobWriterexists.
Recommendation: do (1) first behind a capability flag, fold into (2) when distribution
is removed. After offload, AppView's push surface is just three cheap operations:
POST uploads/ (handshake), HEAD blob (existence), PUT manifest.
Hold changes
The hold grows a real, client-facing OCI blob-upload surface (it speaks only custom XRPC today):
| Capability | Notes |
|---|---|
| Accept the upload session created/referenced by the Location | session state on the hold (existing multipart manager is the basis: pkg/hold/oci/multipart.go) |
PATCH (chunked) and PUT (monolithic + finalize) byte handlers returning OCI-compliant responses |
202+Location for continuation; 201+Docker-Content-Digest+Location for finalize |
Stream to S3: monolithic PutObject passthrough; multipart (~5MB parts) for chunked/large |
reuse StartMultipartUploadWithManager / CompleteMultipartUploadWithManager (multipart.go:134-282) but feed bytes from the client request instead of presigning back to AppView |
| Validate the service token once, bind session to user | ValidateServiceToken (auth.go:507) + new session ownership |
| Enforce push policy (auth-required vs anonymous push) | captain record policy |
| Existing post-upload side effects still fire | layer records, quota, Bluesky status, scan dispatch -- today via notifyManifest (pkg/hold/oci/xrpc.go:189-484); fold into finalize or keep as the AppView->hold call after manifest PUT |
The presigned-URL-to-AppView path (getPartUploadUrl) is no longer needed for offloaded
pushes -- the hold writes to its own S3 directly with its own credentials (presigning
buys nothing when you hold the keys).
Capability detection and rollout
Not every hold will support offload immediately (older holds, third-party holds). AppView must detect support and fall back to today's proxy path:
- Per-hold capability flag in the captain record (e.g.
supportsPushOffload: true) or a hold/v2-probe. AppView reads it during hold discovery. - Offload-capable hold -> return cross-host Location (this proposal).
- Legacy hold -> keep today's
ProxyBlobStorepath (AppView buffers).
Suggested rollout:
- Phase 0: implement, default off; managed holds opt in first (operator's own infra, safe to test client compatibility).
- Phase 1: enable for managed holds; validate
docker/containerd/podman/buildkitpush compatibility (cross-host upload Location is the main compatibility risk). - Phase 2: enable for BYOS holds that advertise the capability.
- Phase 3: fold into
REMOVING_DISTRIBUTIONand drop theProxyBlobWriterbyte path.
Economics after this change
| Path | Today | After offload |
|---|---|---|
| Pull (any hold) | client <- S3 (307), operator carries 0 | unchanged |
| Push, managed hold (operator storage) | client -> AppView -> operator S3 (operator pays both legs) | client -> hold -> operator S3 (operator still pays, but one leg; or keep proxy) |
| Push, BYOS hold (user storage) | client -> AppView -> user S3 (operator pays) | client -> user's hold -> user S3 (user pays) |
The BYOS bargain is restored: the operator pays for reads it serves; self-hosters pay for writes to their own storage.
Edge cases and risks
- Client compatibility with cross-host upload Location. The primary risk. Some clients historically mishandle absolute/cross-host upload locations or auth across hosts. Token-in-URL mitigates the auth half. Must validate docker, containerd, podman, buildkit before BYOS rollout.
- Token in URL leaks to hold logs. Short-lived; bound to a single hold; consider
redaction in hold access logs and a dedicated short
expfor upload tokens. - Token expiry mid-upload. Handled by session-bound authorization (validate once,
bind session) plus a longer
expwhere the PDS grants it. - Blobs > 5GB. Require multipart on the hold (single
PutObjectcaps at 5GB). The hold's existing multipart manager covers this. - Resumable / chunked uploads,
GETupload status. The hold must implement the OCI upload-statusGETandRangesemantics if clients use them. - Cross-repo blob mount (
POST .../uploads/?mount=<digest>&from=<repo>). Resolve at AppView (metadata) or proxy a mount hint to the hold; document which. - Abort / cleanup (
DELETEupload, client disconnect). The hold owns session cleanup and S3 multipart abort (existingabortUploadlogic). - Quota enforcement timing. Quotas are enforced at the hold; ensure the hold checks quota at session creation and/or finalize, not only via the old AppView path.
- Manifest references unflushed blobs. Manifest
PUTon AppView must verify all referenced blobs exist on the hold (the upload happened out-of-band on the hold), and fail the manifest if any are missing. - Two TLS endpoints / hostnames. The hold needs a valid public TLS cert at
server.public_url(already required:pkg/hold/config.go:155,371).
Implementation checklist
Per repo conventions (CLAUDE.md):
- Lexicons / endpoints: add the hold's OCI upload endpoints (or document the
/v2passthrough) inlexicons/,pkg/atproto/endpoints.go, anddocs/HOLD_XRPC_ENDPOINTS.md; runmake lex-lint. - Captain record: add the push-policy /
supportsPushOffloadfield; regenerate cbor (go generate ./pkg/atproto/...) and register inpkg/hold/pds/server.goif the record type changes. - Config: any new hold/appview config (e.g. offload toggle) -> regenerate example
configs and sync
deploy/upcloud/configs/*.yaml.tmpl. - AppView: new upload-init handler ahead of distribution; push-policy check; service token mint into Location; manifest-time blob existence check; capability detection + fallback.
- Hold: OCI
PATCH/PUT/finalize/DELETE/status handlers; streaming-to-S3 + multipart; service-token validation + session binding; push policy; post-upload side effects (layer records, quota, scan dispatch). - Jetstream: if the captain push-policy field is firehose-fed, update
pkg/appview/jetstream/backfill.goandprocessor.go. - Docs: update
BYOS.md,hold.md,HOLD_XRPC_ENDPOINTS.md; cross-linkREMOVING_DISTRIBUTION.md. - Tests + lint:
make lint,make test; add client-compat integration coverage.
Open questions
- Push-policy shape. New
supportsPushOffload+anonymousPushbits on the captain record, or reuse/extendPublic? (Publiccurrently = anonymous reads only.) - Upload session URL namespace. Real
/v2/...on the hold, or a hold-specific path (opaque to the client)?/v2is most compatible; opaque is cleaner internally. - Service token
lxmbinding.getServiceAuthbinds a token to a lexicon method; OCI/v2routes are not XRPC. Bind to a representative method (e.g.io.atcr.hold.initiateUpload) and have the hold accept it for the upload routes, or relaxlxmchecking forblob:write? - Managed holds: keep proxying through AppView (operator pays anyway and avoids client-compat risk), or also offload for consistency?
- Sequencing vs
REMOVING_DISTRIBUTION. Ship the intercept-handler version first, or wait and build it into the post-distribution/v2layer?
See also
docs/BYOS.md-- BYOS architecture and authorization modeldocs/REMOVING_DISTRIBUTION.md-- owning the/v2HTTP layer (reinforces this work)docs/HOLD_XRPC_ENDPOINTS.md-- current hold upload endpointsdocs/DIRECT_HOLD_ACCESS.md-- service token acquisition flow (same credential reused here)docs/HOLD_DISCOVERY.md-- how AppView resolves the target hold