# Push Offload: Shifting Blob Uploads to the Hold **Status:** Proposal / design. Not yet implemented. This document describes shifting OCI blob *upload* (push) bandwidth off the AppView and onto the hold, mirroring what the pull path already does for reads. It is the write-side counterpart to the existing pull `307` redirect. ## Motivation ### The asymmetry The hold design is deliberately lightweight: it handshakes a few HTTP requests and hands back presigned S3 URLs. On the **read path** this works perfectly. A pull is a bodyless `GET` that AppView answers with a `307` to a presigned S3 URL (`proxy_blob_store.go:267-281` `ServeBlob`). Neither AppView nor the hold ever carries a blob byte, and reads are 10-100x more frequent than writes, so this is where the property matters most. The **write path** inverts this. Today a push streams: ``` client --PATCH/PUT--> AppView (buffers 10MB chunks in RAM) --presigned PUT--> S3 ``` AppView ingests every layer and re-uploads it to S3 (`proxy_blob_store.go:586-665` `Write`/`flushPart`). It carries the full upload bandwidth even though the hold and S3 are the actual storage. ### Why this is economically backwards for BYOS The entire point of BYOS is that a user runs their own hold and storage so the AppView operator is **not** on the hook for their data. But on push, a BYOS upload goes: ``` client --bytes--> AppView (operator's bandwidth) --> user's own S3 ``` The operator pays ingress + egress to shuttle bytes into a bucket they don't own and will never bill for. The people most motivated to self-host (to avoid paying the operator) are exactly the ones costing the operator bandwidth on every push. The BYOS bargain runs backwards on writes. ### Why pure client -> S3 direct is not possible on push It is tempting to do for push what pull does: redirect the client straight to a presigned S3 URL. This cannot work with stock OCI clients (`docker`, `containerd`, `podman`), for reasons independent of each other: 1. **Method preservation kills PATCH.** A `307`/`308` preserves method and body. The chunked path uses `PATCH`; S3 has no `PATCH` (object API is `PUT` or the multipart `POST`+`UploadId` flow). A redirected `PATCH` 405s. 2. **Streamed bodies cannot be replayed.** Following a redirect on a write requires re-sending the body; Go's `http.Client` only does so when `Request.GetBody` is set (rewindable). Docker streams the layer tar, which generally is not rewindable. 3. **The finalize response contract.** The closing `PUT ?digest=sha256:...` expects `201` + `Docker-Content-Digest` + `Location`. S3 returns 200/XML; the client chokes. The client also *appends* `?digest=` to the Location, invalidating any SigV4 presigned URL used as the Location. 4. **Multipart coordination cannot ride a redirect.** Blobs >5GB need S3 multipart, which needs an interactive init -> per-part-URL -> complete loop. A redirect cannot express that handshake. The signature/size question ("you need to know the size before you sign") is a red herring: S3 lets you presign a single `PUT` with `UNSIGNED-PAYLOAD` (no size needed, 5GB cap), and `POST Object` policies support a `content-length-range` (min/max). The real blocker is that **the thing terminating an OCI upload must speak OCI and return OCI responses** -- which S3 cannot do. ### The "pick two" on push On push you cannot have all three at once: 1. Lightweight hold (never touches bytes) 2. Stock `docker push` (no custom client) 3. AppView free of the byte path - **(1)+(2)** -> AppView eats the bandwidth. *(Where we are today.)* - **(2)+(3)** -> the hold terminates the upload and carries the bytes. **<- this proposal** - **(1)+(3)** -> a custom client streams straight to S3, no `docker push`. Pull gets all three only because a bodyless GET redirect satisfies them simultaneously. Since the whole reason for BYOS is "don't pay the operator," the correct trade is **(2)+(3)**: make pushes terminate on the hold, where the bandwidth belongs. ## Goal and non-goals **Goal:** On push, the layer bytes go `client -> hold -> S3` and never transit AppView. The handshake (`POST .../blobs/uploads/`) and the manifest `PUT` stay on AppView (cheap, no bytes). **Non-goals:** - Client -> S3 direct on push (impossible with stock clients, see above). - Changing the pull path (already optimal). - A custom push client (out of scope; would be the only way to get (1)+(3)). The hold "carrying bytes" costs **bandwidth**, not its lightweight character: - **Monolithic push** (docker sends `Content-Length`): the hold does a streaming `PutObject` (`io.Copy` passthrough) -- no disk, trivial RAM. - **Large/chunked push** (>5GB or chunked): multipart with ~5MB part buffers -- still trivial memory. The hold stays memory-light and ops-simple; it just becomes bandwidth-heavy on writes, on the BYOS owner's box. ## Current state (for reference) | Concern | Today | |---|---| | Push handshake | AppView `POST .../blobs/uploads/` -> distribution lib calls `ProxyBlobStore.Create` -> XRPC `io.atcr.hold.initiateUpload` (`proxy_blob_store.go:287-325`) | | Push bytes | Client -> AppView RAM (10MB chunks) -> presigned S3 PUT (`proxy_blob_store.go:605-665`) | | Push finalize | `ProxyBlobWriter.Commit` -> XRPC `io.atcr.hold.completeUpload` (`proxy_blob_store.go:703-741`) | | Hold upload API | Custom XRPC only: `initiateUpload`, `getPartUploadUrl`, `completeUpload`, `abortUpload`, `notifyManifest` (`pkg/hold/oci/xrpc.go:50-61`). No standard OCI `/v2` upload surface. | | Upload `Location` | AppView-relative `/v2//blobs/uploads/`, generated by the distribution library from `BlobWriter.ID()`, not by ATCR code. | | Hold auth | Service token (Bearer, `aud`=hold DID, signed by user's PDS) or DPoP. Validated on every op (`pkg/hold/pds/auth.go:377-429` `ValidateBlobWriteAccess`, `:507-608` `ValidateServiceToken`). The hold **cannot** validate AppView's registry JWT. | | Pull (contrast) | `ServeBlob` -> `307` to presigned S3 URL; client fetches direct (`proxy_blob_store.go:267-281`). | Note: an unused `HoldUploadPart = "/xrpc/io.atcr.hold.uploadPart"` constant already exists (`pkg/atproto/endpoints.go:29`) for "direct buffered part uploads" -- prior art for the hold carrying bytes, never registered. ## Proposed design The flow the operator described: > AppView looks at what hold it needs to go to, checks if that hold is anonymous-push or > not, does a serviceAuth request if necessary, then lets the hold take it from there. ``` 1. docker: POST https:///v2///blobs/uploads/ 2. AppView: resolve hold DID (findHoldDIDAndProfile + resolveSuccessor) check hold push policy (anonymous push allowed?) if auth required: mint/fetch service token (aud=hold DID) via user OAuth -> 202 Accepted Location: https:///v2///blobs/uploads/?_t= 3. docker: PATCH/PUT the layer bytes to that Location (client -> hold) 4. hold: validate service token, bind session to user stream bytes to its S3 (passthrough or 5MB multipart) finalize -> 201 Docker-Content-Digest: sha256:... (hold -> S3, OCI response) 5. docker: PUT https:///v2///manifests/ 6. AppView: verify referenced blobs exist on the hold, store manifest in user's PDS ``` AppView is out of the byte path entirely (steps 3-4). It keeps only the cheap metadata operations (steps 1-2, 5-6). ### Why the handshake stays on AppView Hold resolution and the serviceAuth request both require AppView's OAuth session for the user (only AppView holds it). So the `POST .../blobs/uploads/` must be answered by AppView. Everything after the `202` lives on the hold. ### Cross-host Location is spec-legal and opaque to the client The OCI distribution spec permits an absolute, cross-host `Location`, and the client treats the returned session URL as **opaque** -- it `PATCH`/`PUT`s to it verbatim. So: - The hold's session URL does **not** need to be a literal `/v2/...` path; the hold may namespace it however it likes, as long as its responses are OCI-compliant (`202` + `Location` for `PATCH` continuation, `201` + `Docker-Content-Digest` for the finalize `PUT`). - Only the finalize `PUT` appends `?digest=` -- the hold parses this normally (it is not S3, so the appended query breaks nothing). ## Authentication model This is the crux, and the operator's proposed flow resolves it cleanly. ### Service token in the Location URL (chosen approach) AppView mints a **service token** during the handshake (the same machinery used today for AppView's own XRPC calls: `GetOrFetchServiceToken`, `pkg/auth/servicetoken.go:53`, `getServiceAuth?aud=&lxm=&exp=`) and embeds it in the upload Location. Why this works and the registry JWT does not: - The hold **already** validates service tokens (`ValidateServiceToken`, `pkg/hold/pds/auth.go:507`): checks `aud` == hold DID, `exp`, and the signature against the user's DID document. No new trust infrastructure. - The docker client only *holds* an AppView-issued registry JWT, which the hold cannot validate. By putting the credential in the Location URL, the client does not need to know anything about hold auth -- it just uses the URL. - **Cross-host auth headers:** docker generally does not forward its `Authorization` bearer to a *different* host on a redirected upload (credential-leak avoidance). The token-in-URL avoids depending on header forwarding entirely. Trade-off: the token appears in the URL (and thus potentially in hold access logs). It is short-lived; mitigations below. ### Session-bound authorization (handles token expiry mid-upload) Service tokens are short-lived (~5 min typical; reference PDSes grant up to 1h). A large push could outlive the token. To avoid re-validating on every chunk: - The hold validates the token **once**, when the upload session is created (or first byte-bearing request arrives), and **binds the session to the authorized user**. - Subsequent `PATCH`/`PUT` on that session are authorized by **session ownership**, not by re-checking the token. The session ID is the capability. - Request a longer `exp` for upload tokens where the PDS allows it. This means token expiry mid-upload does not abort an in-flight push. ### Anonymous / managed-push policy The handshake checks the hold's push policy before deciding whether to mint a token: - **Auth required (default):** mint a service token; embed in Location. The hold enforces captain/crew `blob:write` exactly as today (`ValidateBlobWriteAccess`). - **Anonymous push allowed:** return a tokenless Location; the hold accepts the upload without a service token. Note the captain record's existing `Public` flag governs anonymous *reads* only (`pkg/hold/pds/captain.go:24`, `auth.go:445`); writes always require auth today. "Anonymous push" is a **new** policy bit (see Open Questions). Most holds will keep auth-required; anonymous push is opt-in for, e.g., open CI mirrors. ### Rejected alternative: JWKS on AppView AppView could publish a JWKS so the hold validates AppView's registry JWT directly. This needs new trust plumbing (JWKS endpoint, hold-side fetch + trust config) and still requires the client to send the registry JWT cross-host (which it avoids doing). The service-token approach reuses validation the hold already performs, so JWKS is not pursued. ## AppView changes | Change | Where | |---|---| | `POST .../blobs/uploads/` returns a cross-host hold Location instead of driving `ProxyBlobStore.Create` | new handler ahead of the distribution `/v2` handler; see "Distribution interaction" | | Hold push-policy check during handshake | `findHoldDIDAndProfile` result + captain push policy (Jetstream-fed local table, keep it local-fast) | | Mint service token for upload, embed in Location | reuse `GetOrFetchServiceToken` (`pkg/auth/servicetoken.go`) | | `HEAD .../blobs/` (existence / layer skip) stays on AppView | answer from local layer metadata / hold query; keeps docker's skip-existing fast | | Manifest `PUT` stays on AppView; verify referenced blobs exist on the hold before storing | `manifest_store.go` + a hold existence check | | `ProxyBlobWriter` byte path becomes unused for offloaded pushes | `proxy_blob_store.go` (see below) | ### Distribution library interaction The distribution library currently generates the upload Location itself and owns the `PATCH`/`PUT` session. Two ways to take over the handshake: 1. **Intercept the upload-init route** ahead of distribution: register `POST /v2//blobs/uploads/` on ATCR's router so it short-circuits distribution, returning the cross-host hold Location. The subsequent `PATCH`/`PUT` go to the hold (different host), so distribution's blob-upload machinery is never invoked for these pushes. Cleanest incremental path. 2. **Do it as part of `docs/REMOVING_DISTRIBUTION.md`.** Push-offload makes `proxy_blob_store.go`'s `BlobWriter` -- the doc's single gnarliest distribution impl -- obsolete for BYOS. The two efforts reinforce each other: owning the `/v2` HTTP layer makes emitting a cross-host upload Location trivial, and push-offload removes the hardest reason the `BlobWriter` exists. Recommendation: do (1) first behind a capability flag, fold into (2) when distribution is removed. After offload, AppView's push surface is just three cheap operations: `POST uploads/` (handshake), `HEAD blob` (existence), `PUT manifest`. ## Hold changes The hold grows a real, client-facing OCI blob-upload surface (it speaks only custom XRPC today): | Capability | Notes | |---|---| | Accept the upload session created/referenced by the Location | session state on the hold (existing multipart manager is the basis: `pkg/hold/oci/multipart.go`) | | `PATCH` (chunked) and `PUT` (monolithic + finalize) byte handlers returning OCI-compliant responses | `202`+`Location` for continuation; `201`+`Docker-Content-Digest`+`Location` for finalize | | Stream to S3: monolithic `PutObject` passthrough; multipart (~5MB parts) for chunked/large | reuse `StartMultipartUploadWithManager` / `CompleteMultipartUploadWithManager` (`multipart.go:134-282`) but feed bytes from the client request instead of presigning back to AppView | | Validate the service token once, bind session to user | `ValidateServiceToken` (`auth.go:507`) + new session ownership | | Enforce push policy (auth-required vs anonymous push) | captain record policy | | Existing post-upload side effects still fire | layer records, quota, Bluesky status, scan dispatch -- today via `notifyManifest` (`pkg/hold/oci/xrpc.go:189-484`); fold into finalize or keep as the AppView->hold call after manifest PUT | The presigned-URL-to-AppView path (`getPartUploadUrl`) is no longer needed for offloaded pushes -- the hold writes to its own S3 directly with its own credentials (presigning buys nothing when you hold the keys). ## Capability detection and rollout Not every hold will support offload immediately (older holds, third-party holds). AppView must detect support and fall back to today's proxy path: 1. **Per-hold capability flag** in the captain record (e.g. `supportsPushOffload: true`) or a hold `/v2`-probe. AppView reads it during hold discovery. 2. **Offload-capable hold** -> return cross-host Location (this proposal). 3. **Legacy hold** -> keep today's `ProxyBlobStore` path (AppView buffers). Suggested rollout: - Phase 0: implement, default **off**; managed holds opt in first (operator's own infra, safe to test client compatibility). - Phase 1: enable for managed holds; validate `docker` / `containerd` / `podman` / `buildkit` push compatibility (cross-host upload Location is the main compatibility risk). - Phase 2: enable for BYOS holds that advertise the capability. - Phase 3: fold into `REMOVING_DISTRIBUTION` and drop the `ProxyBlobWriter` byte path. ## Economics after this change | Path | Today | After offload | |---|---|---| | Pull (any hold) | client <- S3 (307), operator carries 0 | unchanged | | Push, managed hold (operator storage) | client -> AppView -> operator S3 (operator pays both legs) | client -> hold -> operator S3 (operator still pays, but one leg; or keep proxy) | | Push, BYOS hold (user storage) | client -> **AppView** -> user S3 (**operator pays**) | client -> **user's hold** -> user S3 (**user pays**) | The BYOS bargain is restored: the operator pays for reads it serves; self-hosters pay for writes to their own storage. ## Edge cases and risks - **Client compatibility with cross-host upload Location.** The primary risk. Some clients historically mishandle absolute/cross-host upload locations or auth across hosts. Token-in-URL mitigates the auth half. Must validate docker, containerd, podman, buildkit before BYOS rollout. - **Token in URL leaks to hold logs.** Short-lived; bound to a single hold; consider redaction in hold access logs and a dedicated short `exp` for upload tokens. - **Token expiry mid-upload.** Handled by session-bound authorization (validate once, bind session) plus a longer `exp` where the PDS grants it. - **Blobs > 5GB.** Require multipart on the hold (single `PutObject` caps at 5GB). The hold's existing multipart manager covers this. - **Resumable / chunked uploads, `GET` upload status.** The hold must implement the OCI upload-status `GET` and `Range` semantics if clients use them. - **Cross-repo blob mount** (`POST .../uploads/?mount=&from=`). Resolve at AppView (metadata) or proxy a mount hint to the hold; document which. - **Abort / cleanup** (`DELETE` upload, client disconnect). The hold owns session cleanup and S3 multipart abort (existing `abortUpload` logic). - **Quota enforcement timing.** Quotas are enforced at the hold; ensure the hold checks quota at session creation and/or finalize, not only via the old AppView path. - **Manifest references unflushed blobs.** Manifest `PUT` on AppView must verify all referenced blobs exist on the hold (the upload happened out-of-band on the hold), and fail the manifest if any are missing. - **Two TLS endpoints / hostnames.** The hold needs a valid public TLS cert at `server.public_url` (already required: `pkg/hold/config.go:155,371`). ## Implementation checklist Per repo conventions (`CLAUDE.md`): 1. **Lexicons / endpoints:** add the hold's OCI upload endpoints (or document the `/v2` passthrough) in `lexicons/`, `pkg/atproto/endpoints.go`, and `docs/HOLD_XRPC_ENDPOINTS.md`; run `make lex-lint`. 2. **Captain record:** add the push-policy / `supportsPushOffload` field; regenerate cbor (`go generate ./pkg/atproto/...`) and register in `pkg/hold/pds/server.go` if the record type changes. 3. **Config:** any new hold/appview config (e.g. offload toggle) -> regenerate example configs and sync `deploy/upcloud/configs/*.yaml.tmpl`. 4. **AppView:** new upload-init handler ahead of distribution; push-policy check; service token mint into Location; manifest-time blob existence check; capability detection + fallback. 5. **Hold:** OCI `PATCH`/`PUT`/finalize/`DELETE`/status handlers; streaming-to-S3 + multipart; service-token validation + session binding; push policy; post-upload side effects (layer records, quota, scan dispatch). 6. **Jetstream:** if the captain push-policy field is firehose-fed, update `pkg/appview/jetstream/backfill.go` and `processor.go`. 7. **Docs:** update `BYOS.md`, `hold.md`, `HOLD_XRPC_ENDPOINTS.md`; cross-link `REMOVING_DISTRIBUTION.md`. 8. **Tests + lint:** `make lint`, `make test`; add client-compat integration coverage. ## Open questions 1. **Push-policy shape.** New `supportsPushOffload` + `anonymousPush` bits on the captain record, or reuse/extend `Public`? (`Public` currently = anonymous reads only.) 2. **Upload session URL namespace.** Real `/v2/...` on the hold, or a hold-specific path (opaque to the client)? `/v2` is most compatible; opaque is cleaner internally. 3. **Service token `lxm` binding.** `getServiceAuth` binds a token to a lexicon method; OCI `/v2` routes are not XRPC. Bind to a representative method (e.g. `io.atcr.hold.initiateUpload`) and have the hold accept it for the upload routes, or relax `lxm` checking for `blob:write`? 4. **Managed holds:** keep proxying through AppView (operator pays anyway and avoids client-compat risk), or also offload for consistency? 5. **Sequencing vs `REMOVING_DISTRIBUTION`.** Ship the intercept-handler version first, or wait and build it into the post-distribution `/v2` layer? ## See also - `docs/BYOS.md` -- BYOS architecture and authorization model - `docs/REMOVING_DISTRIBUTION.md` -- owning the `/v2` HTTP layer (reinforces this work) - `docs/HOLD_XRPC_ENDPOINTS.md` -- current hold upload endpoints - `docs/DIRECT_HOLD_ACCESS.md` -- service token acquisition flow (same credential reused here) - `docs/HOLD_DISCOVERY.md` -- how AppView resolves the target hold