docs: add hold push-offload design proposal

Design for moving push bandwidth off the appview and onto the hold,
mirroring the pull path's 307 redirect. Proposal only, not implemented.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
Evan Jarrett
2026-08-09 20:49:24 -05:00
co-authored by Claude Opus 5
parent fa34da0f26
commit 3e73da059a
+376
View File
@@ -0,0 +1,376 @@
# Push Offload: Shifting Blob Uploads to the Hold
**Status:** Proposal / design. Not yet implemented.
This document describes shifting OCI blob *upload* (push) bandwidth off the AppView
and onto the hold, mirroring what the pull path already does for reads. It is the
write-side counterpart to the existing pull `307` redirect.
## Motivation
### The asymmetry
The hold design is deliberately lightweight: it handshakes a few HTTP requests and
hands back presigned S3 URLs. On the **read path** this works perfectly. A pull is a
bodyless `GET` that AppView answers with a `307` to a presigned S3 URL
(`proxy_blob_store.go:267-281` `ServeBlob`). Neither AppView nor the hold ever carries
a blob byte, and reads are 10-100x more frequent than writes, so this is where the
property matters most.
The **write path** inverts this. Today a push streams:
```
client --PATCH/PUT--> AppView (buffers 10MB chunks in RAM) --presigned PUT--> S3
```
AppView ingests every layer and re-uploads it to S3 (`proxy_blob_store.go:586-665`
`Write`/`flushPart`). It carries the full upload bandwidth even though the hold and S3
are the actual storage.
### Why this is economically backwards for BYOS
The entire point of BYOS is that a user runs their own hold and storage so the AppView
operator is **not** on the hook for their data. But on push, a BYOS upload goes:
```
client --bytes--> AppView (operator's bandwidth) --> user's own S3
```
The operator pays ingress + egress to shuttle bytes into a bucket they don't own and
will never bill for. The people most motivated to self-host (to avoid paying the
operator) are exactly the ones costing the operator bandwidth on every push. The BYOS
bargain runs backwards on writes.
### Why pure client -> S3 direct is not possible on push
It is tempting to do for push what pull does: redirect the client straight to a
presigned S3 URL. This cannot work with stock OCI clients (`docker`, `containerd`,
`podman`), for reasons independent of each other:
1. **Method preservation kills PATCH.** A `307`/`308` preserves method and body. The
chunked path uses `PATCH`; S3 has no `PATCH` (object API is `PUT` or the multipart
`POST`+`UploadId` flow). A redirected `PATCH` 405s.
2. **Streamed bodies cannot be replayed.** Following a redirect on a write requires
re-sending the body; Go's `http.Client` only does so when `Request.GetBody` is set
(rewindable). Docker streams the layer tar, which generally is not rewindable.
3. **The finalize response contract.** The closing `PUT <location>?digest=sha256:...`
expects `201` + `Docker-Content-Digest` + `Location`. S3 returns 200/XML; the client
chokes. The client also *appends* `?digest=` to the Location, invalidating any SigV4
presigned URL used as the Location.
4. **Multipart coordination cannot ride a redirect.** Blobs >5GB need S3 multipart,
which needs an interactive init -> per-part-URL -> complete loop. A redirect cannot
express that handshake.
The signature/size question ("you need to know the size before you sign") is a red
herring: S3 lets you presign a single `PUT` with `UNSIGNED-PAYLOAD` (no size needed,
5GB cap), and `POST Object` policies support a `content-length-range` (min/max). The
real blocker is that **the thing terminating an OCI upload must speak OCI and return
OCI responses** -- which S3 cannot do.
### The "pick two" on push
On push you cannot have all three at once:
1. Lightweight hold (never touches bytes)
2. Stock `docker push` (no custom client)
3. AppView free of the byte path
- **(1)+(2)** -> AppView eats the bandwidth. *(Where we are today.)*
- **(2)+(3)** -> the hold terminates the upload and carries the bytes. **<- this proposal**
- **(1)+(3)** -> a custom client streams straight to S3, no `docker push`.
Pull gets all three only because a bodyless GET redirect satisfies them simultaneously.
Since the whole reason for BYOS is "don't pay the operator," the correct trade is
**(2)+(3)**: make pushes terminate on the hold, where the bandwidth belongs.
## Goal and non-goals
**Goal:** On push, the layer bytes go `client -> hold -> S3` and never transit AppView.
The handshake (`POST .../blobs/uploads/`) and the manifest `PUT` stay on AppView (cheap,
no bytes).
**Non-goals:**
- Client -> S3 direct on push (impossible with stock clients, see above).
- Changing the pull path (already optimal).
- A custom push client (out of scope; would be the only way to get (1)+(3)).
The hold "carrying bytes" costs **bandwidth**, not its lightweight character:
- **Monolithic push** (docker sends `Content-Length`): the hold does a streaming
`PutObject` (`io.Copy` passthrough) -- no disk, trivial RAM.
- **Large/chunked push** (>5GB or chunked): multipart with ~5MB part buffers -- still
trivial memory.
The hold stays memory-light and ops-simple; it just becomes bandwidth-heavy on writes,
on the BYOS owner's box.
## Current state (for reference)
| Concern | Today |
|---|---|
| Push handshake | AppView `POST .../blobs/uploads/` -> distribution lib calls `ProxyBlobStore.Create` -> XRPC `io.atcr.hold.initiateUpload` (`proxy_blob_store.go:287-325`) |
| Push bytes | Client -> AppView RAM (10MB chunks) -> presigned S3 PUT (`proxy_blob_store.go:605-665`) |
| Push finalize | `ProxyBlobWriter.Commit` -> XRPC `io.atcr.hold.completeUpload` (`proxy_blob_store.go:703-741`) |
| Hold upload API | Custom XRPC only: `initiateUpload`, `getPartUploadUrl`, `completeUpload`, `abortUpload`, `notifyManifest` (`pkg/hold/oci/xrpc.go:50-61`). No standard OCI `/v2` upload surface. |
| Upload `Location` | AppView-relative `/v2/<name>/blobs/uploads/<id>`, generated by the distribution library from `BlobWriter.ID()`, not by ATCR code. |
| Hold auth | Service token (Bearer, `aud`=hold DID, signed by user's PDS) or DPoP. Validated on every op (`pkg/hold/pds/auth.go:377-429` `ValidateBlobWriteAccess`, `:507-608` `ValidateServiceToken`). The hold **cannot** validate AppView's registry JWT. |
| Pull (contrast) | `ServeBlob` -> `307` to presigned S3 URL; client fetches direct (`proxy_blob_store.go:267-281`). |
Note: an unused `HoldUploadPart = "/xrpc/io.atcr.hold.uploadPart"` constant already
exists (`pkg/atproto/endpoints.go:29`) for "direct buffered part uploads" -- prior art
for the hold carrying bytes, never registered.
## Proposed design
The flow the operator described:
> AppView looks at what hold it needs to go to, checks if that hold is anonymous-push or
> not, does a serviceAuth request if necessary, then lets the hold take it from there.
```
1. docker: POST https://<appview>/v2/<identity>/<image>/blobs/uploads/
2. AppView: resolve hold DID (findHoldDIDAndProfile + resolveSuccessor)
check hold push policy (anonymous push allowed?)
if auth required: mint/fetch service token (aud=hold DID) via user OAuth
-> 202 Accepted
Location: https://<hold-public-url>/v2/<identity>/<image>/blobs/uploads/<session>?_t=<serviceToken>
3. docker: PATCH/PUT the layer bytes to that Location (client -> hold)
4. hold: validate service token, bind session to user
stream bytes to its S3 (passthrough or 5MB multipart)
finalize -> 201 Docker-Content-Digest: sha256:... (hold -> S3, OCI response)
5. docker: PUT https://<appview>/v2/<identity>/<image>/manifests/<ref>
6. AppView: verify referenced blobs exist on the hold, store manifest in user's PDS
```
AppView is out of the byte path entirely (steps 3-4). It keeps only the cheap
metadata operations (steps 1-2, 5-6).
### Why the handshake stays on AppView
Hold resolution and the serviceAuth request both require AppView's OAuth session for the
user (only AppView holds it). So the `POST .../blobs/uploads/` must be answered by
AppView. Everything after the `202` lives on the hold.
### Cross-host Location is spec-legal and opaque to the client
The OCI distribution spec permits an absolute, cross-host `Location`, and the client
treats the returned session URL as **opaque** -- it `PATCH`/`PUT`s to it verbatim. So:
- The hold's session URL does **not** need to be a literal `/v2/...` path; the hold may
namespace it however it likes, as long as its responses are OCI-compliant
(`202` + `Location` for `PATCH` continuation, `201` + `Docker-Content-Digest` for the
finalize `PUT`).
- Only the finalize `PUT` appends `?digest=` -- the hold parses this normally (it is not
S3, so the appended query breaks nothing).
## Authentication model
This is the crux, and the operator's proposed flow resolves it cleanly.
### Service token in the Location URL (chosen approach)
AppView mints a **service token** during the handshake (the same machinery used today
for AppView's own XRPC calls: `GetOrFetchServiceToken`, `pkg/auth/servicetoken.go:53`,
`getServiceAuth?aud=<holdDID>&lxm=<method>&exp=<ts>`) and embeds it in the upload
Location.
Why this works and the registry JWT does not:
- The hold **already** validates service tokens (`ValidateServiceToken`,
`pkg/hold/pds/auth.go:507`): checks `aud` == hold DID, `exp`, and the signature
against the user's DID document. No new trust infrastructure.
- The docker client only *holds* an AppView-issued registry JWT, which the hold cannot
validate. By putting the credential in the Location URL, the client does not need to
know anything about hold auth -- it just uses the URL.
- **Cross-host auth headers:** docker generally does not forward its `Authorization`
bearer to a *different* host on a redirected upload (credential-leak avoidance). The
token-in-URL avoids depending on header forwarding entirely.
Trade-off: the token appears in the URL (and thus potentially in hold access logs). It
is short-lived; mitigations below.
### Session-bound authorization (handles token expiry mid-upload)
Service tokens are short-lived (~5 min typical; reference PDSes grant up to 1h). A large
push could outlive the token. To avoid re-validating on every chunk:
- The hold validates the token **once**, when the upload session is created (or first
byte-bearing request arrives), and **binds the session to the authorized user**.
- Subsequent `PATCH`/`PUT` on that session are authorized by **session ownership**, not
by re-checking the token. The session ID is the capability.
- Request a longer `exp` for upload tokens where the PDS allows it.
This means token expiry mid-upload does not abort an in-flight push.
### Anonymous / managed-push policy
The handshake checks the hold's push policy before deciding whether to mint a token:
- **Auth required (default):** mint a service token; embed in Location. The hold enforces
captain/crew `blob:write` exactly as today (`ValidateBlobWriteAccess`).
- **Anonymous push allowed:** return a tokenless Location; the hold accepts the upload
without a service token.
Note the captain record's existing `Public` flag governs anonymous *reads* only
(`pkg/hold/pds/captain.go:24`, `auth.go:445`); writes always require auth today.
"Anonymous push" is a **new** policy bit (see Open Questions). Most holds will keep
auth-required; anonymous push is opt-in for, e.g., open CI mirrors.
### Rejected alternative: JWKS on AppView
AppView could publish a JWKS so the hold validates AppView's registry JWT directly. This
needs new trust plumbing (JWKS endpoint, hold-side fetch + trust config) and still
requires the client to send the registry JWT cross-host (which it avoids doing). The
service-token approach reuses validation the hold already performs, so JWKS is not
pursued.
## AppView changes
| Change | Where |
|---|---|
| `POST .../blobs/uploads/` returns a cross-host hold Location instead of driving `ProxyBlobStore.Create` | new handler ahead of the distribution `/v2` handler; see "Distribution interaction" |
| Hold push-policy check during handshake | `findHoldDIDAndProfile` result + captain push policy (Jetstream-fed local table, keep it local-fast) |
| Mint service token for upload, embed in Location | reuse `GetOrFetchServiceToken` (`pkg/auth/servicetoken.go`) |
| `HEAD .../blobs/<digest>` (existence / layer skip) stays on AppView | answer from local layer metadata / hold query; keeps docker's skip-existing fast |
| Manifest `PUT` stays on AppView; verify referenced blobs exist on the hold before storing | `manifest_store.go` + a hold existence check |
| `ProxyBlobWriter` byte path becomes unused for offloaded pushes | `proxy_blob_store.go` (see below) |
### Distribution library interaction
The distribution library currently generates the upload Location itself and owns the
`PATCH`/`PUT` session. Two ways to take over the handshake:
1. **Intercept the upload-init route** ahead of distribution: register
`POST /v2/<name>/blobs/uploads/` on ATCR's router so it short-circuits distribution,
returning the cross-host hold Location. The subsequent `PATCH`/`PUT` go to the hold
(different host), so distribution's blob-upload machinery is never invoked for these
pushes. Cleanest incremental path.
2. **Do it as part of `docs/REMOVING_DISTRIBUTION.md`.** Push-offload makes
`proxy_blob_store.go`'s `BlobWriter` -- the doc's single gnarliest distribution impl
-- obsolete for BYOS. The two efforts reinforce each other: owning the `/v2` HTTP
layer makes emitting a cross-host upload Location trivial, and push-offload removes
the hardest reason the `BlobWriter` exists.
Recommendation: do (1) first behind a capability flag, fold into (2) when distribution
is removed. After offload, AppView's push surface is just three cheap operations:
`POST uploads/` (handshake), `HEAD blob` (existence), `PUT manifest`.
## Hold changes
The hold grows a real, client-facing OCI blob-upload surface (it speaks only custom
XRPC today):
| Capability | Notes |
|---|---|
| Accept the upload session created/referenced by the Location | session state on the hold (existing multipart manager is the basis: `pkg/hold/oci/multipart.go`) |
| `PATCH` (chunked) and `PUT` (monolithic + finalize) byte handlers returning OCI-compliant responses | `202`+`Location` for continuation; `201`+`Docker-Content-Digest`+`Location` for finalize |
| Stream to S3: monolithic `PutObject` passthrough; multipart (~5MB parts) for chunked/large | reuse `StartMultipartUploadWithManager` / `CompleteMultipartUploadWithManager` (`multipart.go:134-282`) but feed bytes from the client request instead of presigning back to AppView |
| Validate the service token once, bind session to user | `ValidateServiceToken` (`auth.go:507`) + new session ownership |
| Enforce push policy (auth-required vs anonymous push) | captain record policy |
| Existing post-upload side effects still fire | layer records, quota, Bluesky status, scan dispatch -- today via `notifyManifest` (`pkg/hold/oci/xrpc.go:189-484`); fold into finalize or keep as the AppView->hold call after manifest PUT |
The presigned-URL-to-AppView path (`getPartUploadUrl`) is no longer needed for offloaded
pushes -- the hold writes to its own S3 directly with its own credentials (presigning
buys nothing when you hold the keys).
## Capability detection and rollout
Not every hold will support offload immediately (older holds, third-party holds). AppView
must detect support and fall back to today's proxy path:
1. **Per-hold capability flag** in the captain record (e.g. `supportsPushOffload: true`)
or a hold `/v2`-probe. AppView reads it during hold discovery.
2. **Offload-capable hold** -> return cross-host Location (this proposal).
3. **Legacy hold** -> keep today's `ProxyBlobStore` path (AppView buffers).
Suggested rollout:
- Phase 0: implement, default **off**; managed holds opt in first (operator's own infra,
safe to test client compatibility).
- Phase 1: enable for managed holds; validate `docker` / `containerd` / `podman` /
`buildkit` push compatibility (cross-host upload Location is the main compatibility
risk).
- Phase 2: enable for BYOS holds that advertise the capability.
- Phase 3: fold into `REMOVING_DISTRIBUTION` and drop the `ProxyBlobWriter` byte path.
## Economics after this change
| Path | Today | After offload |
|---|---|---|
| Pull (any hold) | client <- S3 (307), operator carries 0 | unchanged |
| Push, managed hold (operator storage) | client -> AppView -> operator S3 (operator pays both legs) | client -> hold -> operator S3 (operator still pays, but one leg; or keep proxy) |
| Push, BYOS hold (user storage) | client -> **AppView** -> user S3 (**operator pays**) | client -> **user's hold** -> user S3 (**user pays**) |
The BYOS bargain is restored: the operator pays for reads it serves; self-hosters pay
for writes to their own storage.
## Edge cases and risks
- **Client compatibility with cross-host upload Location.** The primary risk. Some
clients historically mishandle absolute/cross-host upload locations or auth across
hosts. Token-in-URL mitigates the auth half. Must validate docker, containerd, podman,
buildkit before BYOS rollout.
- **Token in URL leaks to hold logs.** Short-lived; bound to a single hold; consider
redaction in hold access logs and a dedicated short `exp` for upload tokens.
- **Token expiry mid-upload.** Handled by session-bound authorization (validate once,
bind session) plus a longer `exp` where the PDS grants it.
- **Blobs > 5GB.** Require multipart on the hold (single `PutObject` caps at 5GB). The
hold's existing multipart manager covers this.
- **Resumable / chunked uploads, `GET` upload status.** The hold must implement the OCI
upload-status `GET` and `Range` semantics if clients use them.
- **Cross-repo blob mount** (`POST .../uploads/?mount=<digest>&from=<repo>`). Resolve at
AppView (metadata) or proxy a mount hint to the hold; document which.
- **Abort / cleanup** (`DELETE` upload, client disconnect). The hold owns session
cleanup and S3 multipart abort (existing `abortUpload` logic).
- **Quota enforcement timing.** Quotas are enforced at the hold; ensure the hold checks
quota at session creation and/or finalize, not only via the old AppView path.
- **Manifest references unflushed blobs.** Manifest `PUT` on AppView must verify all
referenced blobs exist on the hold (the upload happened out-of-band on the hold), and
fail the manifest if any are missing.
- **Two TLS endpoints / hostnames.** The hold needs a valid public TLS cert at
`server.public_url` (already required: `pkg/hold/config.go:155,371`).
## Implementation checklist
Per repo conventions (`CLAUDE.md`):
1. **Lexicons / endpoints:** add the hold's OCI upload endpoints (or document the `/v2`
passthrough) in `lexicons/`, `pkg/atproto/endpoints.go`, and
`docs/HOLD_XRPC_ENDPOINTS.md`; run `make lex-lint`.
2. **Captain record:** add the push-policy / `supportsPushOffload` field; regenerate cbor
(`go generate ./pkg/atproto/...`) and register in `pkg/hold/pds/server.go` if the
record type changes.
3. **Config:** any new hold/appview config (e.g. offload toggle) -> regenerate example
configs and sync `deploy/upcloud/configs/*.yaml.tmpl`.
4. **AppView:** new upload-init handler ahead of distribution; push-policy check; service
token mint into Location; manifest-time blob existence check; capability detection +
fallback.
5. **Hold:** OCI `PATCH`/`PUT`/finalize/`DELETE`/status handlers; streaming-to-S3 +
multipart; service-token validation + session binding; push policy; post-upload side
effects (layer records, quota, scan dispatch).
6. **Jetstream:** if the captain push-policy field is firehose-fed, update
`pkg/appview/jetstream/backfill.go` and `processor.go`.
7. **Docs:** update `BYOS.md`, `hold.md`, `HOLD_XRPC_ENDPOINTS.md`; cross-link
`REMOVING_DISTRIBUTION.md`.
8. **Tests + lint:** `make lint`, `make test`; add client-compat integration coverage.
## Open questions
1. **Push-policy shape.** New `supportsPushOffload` + `anonymousPush` bits on the captain
record, or reuse/extend `Public`? (`Public` currently = anonymous reads only.)
2. **Upload session URL namespace.** Real `/v2/...` on the hold, or a hold-specific path
(opaque to the client)? `/v2` is most compatible; opaque is cleaner internally.
3. **Service token `lxm` binding.** `getServiceAuth` binds a token to a lexicon method;
OCI `/v2` routes are not XRPC. Bind to a representative method (e.g.
`io.atcr.hold.initiateUpload`) and have the hold accept it for the upload routes, or
relax `lxm` checking for `blob:write`?
4. **Managed holds:** keep proxying through AppView (operator pays anyway and avoids
client-compat risk), or also offload for consistency?
5. **Sequencing vs `REMOVING_DISTRIBUTION`.** Ship the intercept-handler version first,
or wait and build it into the post-distribution `/v2` layer?
## See also
- `docs/BYOS.md` -- BYOS architecture and authorization model
- `docs/REMOVING_DISTRIBUTION.md` -- owning the `/v2` HTTP layer (reinforces this work)
- `docs/HOLD_XRPC_ENDPOINTS.md` -- current hold upload endpoints
- `docs/DIRECT_HOLD_ACCESS.md` -- service token acquisition flow (same credential reused here)
- `docs/HOLD_DISCOVERY.md` -- how AppView resolves the target hold