diff --git a/docs/HOLD_PUSH_OFFLOAD.md b/docs/HOLD_PUSH_OFFLOAD.md new file mode 100644 index 0000000..5655714 --- /dev/null +++ b/docs/HOLD_PUSH_OFFLOAD.md @@ -0,0 +1,376 @@ +# Push Offload: Shifting Blob Uploads to the Hold + +**Status:** Proposal / design. Not yet implemented. + +This document describes shifting OCI blob *upload* (push) bandwidth off the AppView +and onto the hold, mirroring what the pull path already does for reads. It is the +write-side counterpart to the existing pull `307` redirect. + +## Motivation + +### The asymmetry + +The hold design is deliberately lightweight: it handshakes a few HTTP requests and +hands back presigned S3 URLs. On the **read path** this works perfectly. A pull is a +bodyless `GET` that AppView answers with a `307` to a presigned S3 URL +(`proxy_blob_store.go:267-281` `ServeBlob`). Neither AppView nor the hold ever carries +a blob byte, and reads are 10-100x more frequent than writes, so this is where the +property matters most. + +The **write path** inverts this. Today a push streams: + +``` +client --PATCH/PUT--> AppView (buffers 10MB chunks in RAM) --presigned PUT--> S3 +``` + +AppView ingests every layer and re-uploads it to S3 (`proxy_blob_store.go:586-665` +`Write`/`flushPart`). It carries the full upload bandwidth even though the hold and S3 +are the actual storage. + +### Why this is economically backwards for BYOS + +The entire point of BYOS is that a user runs their own hold and storage so the AppView +operator is **not** on the hook for their data. But on push, a BYOS upload goes: + +``` +client --bytes--> AppView (operator's bandwidth) --> user's own S3 +``` + +The operator pays ingress + egress to shuttle bytes into a bucket they don't own and +will never bill for. The people most motivated to self-host (to avoid paying the +operator) are exactly the ones costing the operator bandwidth on every push. The BYOS +bargain runs backwards on writes. + +### Why pure client -> S3 direct is not possible on push + +It is tempting to do for push what pull does: redirect the client straight to a +presigned S3 URL. This cannot work with stock OCI clients (`docker`, `containerd`, +`podman`), for reasons independent of each other: + +1. **Method preservation kills PATCH.** A `307`/`308` preserves method and body. The + chunked path uses `PATCH`; S3 has no `PATCH` (object API is `PUT` or the multipart + `POST`+`UploadId` flow). A redirected `PATCH` 405s. +2. **Streamed bodies cannot be replayed.** Following a redirect on a write requires + re-sending the body; Go's `http.Client` only does so when `Request.GetBody` is set + (rewindable). Docker streams the layer tar, which generally is not rewindable. +3. **The finalize response contract.** The closing `PUT ?digest=sha256:...` + expects `201` + `Docker-Content-Digest` + `Location`. S3 returns 200/XML; the client + chokes. The client also *appends* `?digest=` to the Location, invalidating any SigV4 + presigned URL used as the Location. +4. **Multipart coordination cannot ride a redirect.** Blobs >5GB need S3 multipart, + which needs an interactive init -> per-part-URL -> complete loop. A redirect cannot + express that handshake. + +The signature/size question ("you need to know the size before you sign") is a red +herring: S3 lets you presign a single `PUT` with `UNSIGNED-PAYLOAD` (no size needed, +5GB cap), and `POST Object` policies support a `content-length-range` (min/max). The +real blocker is that **the thing terminating an OCI upload must speak OCI and return +OCI responses** -- which S3 cannot do. + +### The "pick two" on push + +On push you cannot have all three at once: + +1. Lightweight hold (never touches bytes) +2. Stock `docker push` (no custom client) +3. AppView free of the byte path + +- **(1)+(2)** -> AppView eats the bandwidth. *(Where we are today.)* +- **(2)+(3)** -> the hold terminates the upload and carries the bytes. **<- this proposal** +- **(1)+(3)** -> a custom client streams straight to S3, no `docker push`. + +Pull gets all three only because a bodyless GET redirect satisfies them simultaneously. +Since the whole reason for BYOS is "don't pay the operator," the correct trade is +**(2)+(3)**: make pushes terminate on the hold, where the bandwidth belongs. + +## Goal and non-goals + +**Goal:** On push, the layer bytes go `client -> hold -> S3` and never transit AppView. +The handshake (`POST .../blobs/uploads/`) and the manifest `PUT` stay on AppView (cheap, +no bytes). + +**Non-goals:** +- Client -> S3 direct on push (impossible with stock clients, see above). +- Changing the pull path (already optimal). +- A custom push client (out of scope; would be the only way to get (1)+(3)). + +The hold "carrying bytes" costs **bandwidth**, not its lightweight character: +- **Monolithic push** (docker sends `Content-Length`): the hold does a streaming + `PutObject` (`io.Copy` passthrough) -- no disk, trivial RAM. +- **Large/chunked push** (>5GB or chunked): multipart with ~5MB part buffers -- still + trivial memory. + +The hold stays memory-light and ops-simple; it just becomes bandwidth-heavy on writes, +on the BYOS owner's box. + +## Current state (for reference) + +| Concern | Today | +|---|---| +| Push handshake | AppView `POST .../blobs/uploads/` -> distribution lib calls `ProxyBlobStore.Create` -> XRPC `io.atcr.hold.initiateUpload` (`proxy_blob_store.go:287-325`) | +| Push bytes | Client -> AppView RAM (10MB chunks) -> presigned S3 PUT (`proxy_blob_store.go:605-665`) | +| Push finalize | `ProxyBlobWriter.Commit` -> XRPC `io.atcr.hold.completeUpload` (`proxy_blob_store.go:703-741`) | +| Hold upload API | Custom XRPC only: `initiateUpload`, `getPartUploadUrl`, `completeUpload`, `abortUpload`, `notifyManifest` (`pkg/hold/oci/xrpc.go:50-61`). No standard OCI `/v2` upload surface. | +| Upload `Location` | AppView-relative `/v2//blobs/uploads/`, generated by the distribution library from `BlobWriter.ID()`, not by ATCR code. | +| Hold auth | Service token (Bearer, `aud`=hold DID, signed by user's PDS) or DPoP. Validated on every op (`pkg/hold/pds/auth.go:377-429` `ValidateBlobWriteAccess`, `:507-608` `ValidateServiceToken`). The hold **cannot** validate AppView's registry JWT. | +| Pull (contrast) | `ServeBlob` -> `307` to presigned S3 URL; client fetches direct (`proxy_blob_store.go:267-281`). | + +Note: an unused `HoldUploadPart = "/xrpc/io.atcr.hold.uploadPart"` constant already +exists (`pkg/atproto/endpoints.go:29`) for "direct buffered part uploads" -- prior art +for the hold carrying bytes, never registered. + +## Proposed design + +The flow the operator described: + +> AppView looks at what hold it needs to go to, checks if that hold is anonymous-push or +> not, does a serviceAuth request if necessary, then lets the hold take it from there. + +``` +1. docker: POST https:///v2///blobs/uploads/ +2. AppView: resolve hold DID (findHoldDIDAndProfile + resolveSuccessor) + check hold push policy (anonymous push allowed?) + if auth required: mint/fetch service token (aud=hold DID) via user OAuth + -> 202 Accepted + Location: https:///v2///blobs/uploads/?_t= +3. docker: PATCH/PUT the layer bytes to that Location (client -> hold) +4. hold: validate service token, bind session to user + stream bytes to its S3 (passthrough or 5MB multipart) + finalize -> 201 Docker-Content-Digest: sha256:... (hold -> S3, OCI response) +5. docker: PUT https:///v2///manifests/ +6. AppView: verify referenced blobs exist on the hold, store manifest in user's PDS +``` + +AppView is out of the byte path entirely (steps 3-4). It keeps only the cheap +metadata operations (steps 1-2, 5-6). + +### Why the handshake stays on AppView + +Hold resolution and the serviceAuth request both require AppView's OAuth session for the +user (only AppView holds it). So the `POST .../blobs/uploads/` must be answered by +AppView. Everything after the `202` lives on the hold. + +### Cross-host Location is spec-legal and opaque to the client + +The OCI distribution spec permits an absolute, cross-host `Location`, and the client +treats the returned session URL as **opaque** -- it `PATCH`/`PUT`s to it verbatim. So: + +- The hold's session URL does **not** need to be a literal `/v2/...` path; the hold may + namespace it however it likes, as long as its responses are OCI-compliant + (`202` + `Location` for `PATCH` continuation, `201` + `Docker-Content-Digest` for the + finalize `PUT`). +- Only the finalize `PUT` appends `?digest=` -- the hold parses this normally (it is not + S3, so the appended query breaks nothing). + +## Authentication model + +This is the crux, and the operator's proposed flow resolves it cleanly. + +### Service token in the Location URL (chosen approach) + +AppView mints a **service token** during the handshake (the same machinery used today +for AppView's own XRPC calls: `GetOrFetchServiceToken`, `pkg/auth/servicetoken.go:53`, +`getServiceAuth?aud=&lxm=&exp=`) and embeds it in the upload +Location. + +Why this works and the registry JWT does not: + +- The hold **already** validates service tokens (`ValidateServiceToken`, + `pkg/hold/pds/auth.go:507`): checks `aud` == hold DID, `exp`, and the signature + against the user's DID document. No new trust infrastructure. +- The docker client only *holds* an AppView-issued registry JWT, which the hold cannot + validate. By putting the credential in the Location URL, the client does not need to + know anything about hold auth -- it just uses the URL. +- **Cross-host auth headers:** docker generally does not forward its `Authorization` + bearer to a *different* host on a redirected upload (credential-leak avoidance). The + token-in-URL avoids depending on header forwarding entirely. + +Trade-off: the token appears in the URL (and thus potentially in hold access logs). It +is short-lived; mitigations below. + +### Session-bound authorization (handles token expiry mid-upload) + +Service tokens are short-lived (~5 min typical; reference PDSes grant up to 1h). A large +push could outlive the token. To avoid re-validating on every chunk: + +- The hold validates the token **once**, when the upload session is created (or first + byte-bearing request arrives), and **binds the session to the authorized user**. +- Subsequent `PATCH`/`PUT` on that session are authorized by **session ownership**, not + by re-checking the token. The session ID is the capability. +- Request a longer `exp` for upload tokens where the PDS allows it. + +This means token expiry mid-upload does not abort an in-flight push. + +### Anonymous / managed-push policy + +The handshake checks the hold's push policy before deciding whether to mint a token: + +- **Auth required (default):** mint a service token; embed in Location. The hold enforces + captain/crew `blob:write` exactly as today (`ValidateBlobWriteAccess`). +- **Anonymous push allowed:** return a tokenless Location; the hold accepts the upload + without a service token. + +Note the captain record's existing `Public` flag governs anonymous *reads* only +(`pkg/hold/pds/captain.go:24`, `auth.go:445`); writes always require auth today. +"Anonymous push" is a **new** policy bit (see Open Questions). Most holds will keep +auth-required; anonymous push is opt-in for, e.g., open CI mirrors. + +### Rejected alternative: JWKS on AppView + +AppView could publish a JWKS so the hold validates AppView's registry JWT directly. This +needs new trust plumbing (JWKS endpoint, hold-side fetch + trust config) and still +requires the client to send the registry JWT cross-host (which it avoids doing). The +service-token approach reuses validation the hold already performs, so JWKS is not +pursued. + +## AppView changes + +| Change | Where | +|---|---| +| `POST .../blobs/uploads/` returns a cross-host hold Location instead of driving `ProxyBlobStore.Create` | new handler ahead of the distribution `/v2` handler; see "Distribution interaction" | +| Hold push-policy check during handshake | `findHoldDIDAndProfile` result + captain push policy (Jetstream-fed local table, keep it local-fast) | +| Mint service token for upload, embed in Location | reuse `GetOrFetchServiceToken` (`pkg/auth/servicetoken.go`) | +| `HEAD .../blobs/` (existence / layer skip) stays on AppView | answer from local layer metadata / hold query; keeps docker's skip-existing fast | +| Manifest `PUT` stays on AppView; verify referenced blobs exist on the hold before storing | `manifest_store.go` + a hold existence check | +| `ProxyBlobWriter` byte path becomes unused for offloaded pushes | `proxy_blob_store.go` (see below) | + +### Distribution library interaction + +The distribution library currently generates the upload Location itself and owns the +`PATCH`/`PUT` session. Two ways to take over the handshake: + +1. **Intercept the upload-init route** ahead of distribution: register + `POST /v2//blobs/uploads/` on ATCR's router so it short-circuits distribution, + returning the cross-host hold Location. The subsequent `PATCH`/`PUT` go to the hold + (different host), so distribution's blob-upload machinery is never invoked for these + pushes. Cleanest incremental path. +2. **Do it as part of `docs/REMOVING_DISTRIBUTION.md`.** Push-offload makes + `proxy_blob_store.go`'s `BlobWriter` -- the doc's single gnarliest distribution impl + -- obsolete for BYOS. The two efforts reinforce each other: owning the `/v2` HTTP + layer makes emitting a cross-host upload Location trivial, and push-offload removes + the hardest reason the `BlobWriter` exists. + +Recommendation: do (1) first behind a capability flag, fold into (2) when distribution +is removed. After offload, AppView's push surface is just three cheap operations: +`POST uploads/` (handshake), `HEAD blob` (existence), `PUT manifest`. + +## Hold changes + +The hold grows a real, client-facing OCI blob-upload surface (it speaks only custom +XRPC today): + +| Capability | Notes | +|---|---| +| Accept the upload session created/referenced by the Location | session state on the hold (existing multipart manager is the basis: `pkg/hold/oci/multipart.go`) | +| `PATCH` (chunked) and `PUT` (monolithic + finalize) byte handlers returning OCI-compliant responses | `202`+`Location` for continuation; `201`+`Docker-Content-Digest`+`Location` for finalize | +| Stream to S3: monolithic `PutObject` passthrough; multipart (~5MB parts) for chunked/large | reuse `StartMultipartUploadWithManager` / `CompleteMultipartUploadWithManager` (`multipart.go:134-282`) but feed bytes from the client request instead of presigning back to AppView | +| Validate the service token once, bind session to user | `ValidateServiceToken` (`auth.go:507`) + new session ownership | +| Enforce push policy (auth-required vs anonymous push) | captain record policy | +| Existing post-upload side effects still fire | layer records, quota, Bluesky status, scan dispatch -- today via `notifyManifest` (`pkg/hold/oci/xrpc.go:189-484`); fold into finalize or keep as the AppView->hold call after manifest PUT | + +The presigned-URL-to-AppView path (`getPartUploadUrl`) is no longer needed for offloaded +pushes -- the hold writes to its own S3 directly with its own credentials (presigning +buys nothing when you hold the keys). + +## Capability detection and rollout + +Not every hold will support offload immediately (older holds, third-party holds). AppView +must detect support and fall back to today's proxy path: + +1. **Per-hold capability flag** in the captain record (e.g. `supportsPushOffload: true`) + or a hold `/v2`-probe. AppView reads it during hold discovery. +2. **Offload-capable hold** -> return cross-host Location (this proposal). +3. **Legacy hold** -> keep today's `ProxyBlobStore` path (AppView buffers). + +Suggested rollout: +- Phase 0: implement, default **off**; managed holds opt in first (operator's own infra, + safe to test client compatibility). +- Phase 1: enable for managed holds; validate `docker` / `containerd` / `podman` / + `buildkit` push compatibility (cross-host upload Location is the main compatibility + risk). +- Phase 2: enable for BYOS holds that advertise the capability. +- Phase 3: fold into `REMOVING_DISTRIBUTION` and drop the `ProxyBlobWriter` byte path. + +## Economics after this change + +| Path | Today | After offload | +|---|---|---| +| Pull (any hold) | client <- S3 (307), operator carries 0 | unchanged | +| Push, managed hold (operator storage) | client -> AppView -> operator S3 (operator pays both legs) | client -> hold -> operator S3 (operator still pays, but one leg; or keep proxy) | +| Push, BYOS hold (user storage) | client -> **AppView** -> user S3 (**operator pays**) | client -> **user's hold** -> user S3 (**user pays**) | + +The BYOS bargain is restored: the operator pays for reads it serves; self-hosters pay +for writes to their own storage. + +## Edge cases and risks + +- **Client compatibility with cross-host upload Location.** The primary risk. Some + clients historically mishandle absolute/cross-host upload locations or auth across + hosts. Token-in-URL mitigates the auth half. Must validate docker, containerd, podman, + buildkit before BYOS rollout. +- **Token in URL leaks to hold logs.** Short-lived; bound to a single hold; consider + redaction in hold access logs and a dedicated short `exp` for upload tokens. +- **Token expiry mid-upload.** Handled by session-bound authorization (validate once, + bind session) plus a longer `exp` where the PDS grants it. +- **Blobs > 5GB.** Require multipart on the hold (single `PutObject` caps at 5GB). The + hold's existing multipart manager covers this. +- **Resumable / chunked uploads, `GET` upload status.** The hold must implement the OCI + upload-status `GET` and `Range` semantics if clients use them. +- **Cross-repo blob mount** (`POST .../uploads/?mount=&from=`). Resolve at + AppView (metadata) or proxy a mount hint to the hold; document which. +- **Abort / cleanup** (`DELETE` upload, client disconnect). The hold owns session + cleanup and S3 multipart abort (existing `abortUpload` logic). +- **Quota enforcement timing.** Quotas are enforced at the hold; ensure the hold checks + quota at session creation and/or finalize, not only via the old AppView path. +- **Manifest references unflushed blobs.** Manifest `PUT` on AppView must verify all + referenced blobs exist on the hold (the upload happened out-of-band on the hold), and + fail the manifest if any are missing. +- **Two TLS endpoints / hostnames.** The hold needs a valid public TLS cert at + `server.public_url` (already required: `pkg/hold/config.go:155,371`). + +## Implementation checklist + +Per repo conventions (`CLAUDE.md`): + +1. **Lexicons / endpoints:** add the hold's OCI upload endpoints (or document the `/v2` + passthrough) in `lexicons/`, `pkg/atproto/endpoints.go`, and + `docs/HOLD_XRPC_ENDPOINTS.md`; run `make lex-lint`. +2. **Captain record:** add the push-policy / `supportsPushOffload` field; regenerate cbor + (`go generate ./pkg/atproto/...`) and register in `pkg/hold/pds/server.go` if the + record type changes. +3. **Config:** any new hold/appview config (e.g. offload toggle) -> regenerate example + configs and sync `deploy/upcloud/configs/*.yaml.tmpl`. +4. **AppView:** new upload-init handler ahead of distribution; push-policy check; service + token mint into Location; manifest-time blob existence check; capability detection + + fallback. +5. **Hold:** OCI `PATCH`/`PUT`/finalize/`DELETE`/status handlers; streaming-to-S3 + + multipart; service-token validation + session binding; push policy; post-upload side + effects (layer records, quota, scan dispatch). +6. **Jetstream:** if the captain push-policy field is firehose-fed, update + `pkg/appview/jetstream/backfill.go` and `processor.go`. +7. **Docs:** update `BYOS.md`, `hold.md`, `HOLD_XRPC_ENDPOINTS.md`; cross-link + `REMOVING_DISTRIBUTION.md`. +8. **Tests + lint:** `make lint`, `make test`; add client-compat integration coverage. + +## Open questions + +1. **Push-policy shape.** New `supportsPushOffload` + `anonymousPush` bits on the captain + record, or reuse/extend `Public`? (`Public` currently = anonymous reads only.) +2. **Upload session URL namespace.** Real `/v2/...` on the hold, or a hold-specific path + (opaque to the client)? `/v2` is most compatible; opaque is cleaner internally. +3. **Service token `lxm` binding.** `getServiceAuth` binds a token to a lexicon method; + OCI `/v2` routes are not XRPC. Bind to a representative method (e.g. + `io.atcr.hold.initiateUpload`) and have the hold accept it for the upload routes, or + relax `lxm` checking for `blob:write`? +4. **Managed holds:** keep proxying through AppView (operator pays anyway and avoids + client-compat risk), or also offload for consistency? +5. **Sequencing vs `REMOVING_DISTRIBUTION`.** Ship the intercept-handler version first, + or wait and build it into the post-distribution `/v2` layer? + +## See also + +- `docs/BYOS.md` -- BYOS architecture and authorization model +- `docs/REMOVING_DISTRIBUTION.md` -- owning the `/v2` HTTP layer (reinforces this work) +- `docs/HOLD_XRPC_ENDPOINTS.md` -- current hold upload endpoints +- `docs/DIRECT_HOLD_ACCESS.md` -- service token acquisition flow (same credential reused here) +- `docs/HOLD_DISCOVERY.md` -- how AppView resolves the target hold