appview: upload small blobs with one presigned PUT, and verify every digest

Every blob went through the multipart machinery: an S3 multipart started
on Docker's initial POST, a hold round trip per part, and a complete on
the hold that finished the multipart, HEADed the temp object, copied it
to its final key, and deleted the temp. For a 2KB config blob that was
three hold calls and six S3 operations. On production data 86% of
distinct layers and every config blob fit in a 16MB buffer, and 49% of
image manifests have no layer larger than that.

The writer now buffers up to 16MB (also the multipart part size) and
makes no hold call until it has to. A blob that never overflows the
buffer is written at Commit with a single presigned PUT to its final
key, via the hold's existing method=PUT presign; the multipart only
starts on the first flush. The hold's completeUpload does nothing the
direct path skips: quota, layer records, stats and scan dispatch all
hang off notifyManifest, which is unchanged.

The buffer starts empty and grows on demand, with the doubling capped so
capacity never overshoots 16MB: a config blob costs kilobytes, and only
layers that approach the threshold fill it.

Bytes are hashed as they arrive. Commit compares the computed sha256 to
the digest the client claimed before any network call, and returns
DIGEST_INVALID on mismatch, aborting a multipart if one was started.
Previously nothing verified the content, so a pusher could store wrong
bytes under a digest in the shared content-addressed space.

Tests observe request counts on a fake hold and fake S3 rather than
return values. The growth test streams in 24KB chunks because
power-of-two chunks land on 16MB by luck and hid an earlier weaker guard.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018Yf1ZVA7sXYhQNb9tCo1m5
This commit is contained in:
Evan Jarrett
2026-09-09 09:50:08 -05:00
co-authored by Claude Fable 5.1
parent 034ea5988b
commit f4343d7956
7 changed files with 862 additions and 103 deletions
+19 -5
View File
@@ -222,23 +222,37 @@ fly secrets set HOLD_REGISTRATION_OWNER_DID=did:plc:your-did-here
GET /xrpc/com.atproto.server.getServiceAuth?aud=did:web:alice-storage.fly.dev
Response: { "token": "eyJ..." }
5. AppView initiates multipart upload to hold:
5. AppView buffers the blob (16MB limit) and verifies the bytes it received
against the digest the client claimed. A mismatch is rejected here, before
anything reaches storage.
5a. Small blob (fits in the buffer, which is every config blob and most layers):
AppView asks for one write capability and PUTs the whole blob to its final
content-addressed key. No multipart session, no temp object, no copy.
GET /xrpc/com.atproto.sync.getBlob?did=...&cid=sha256:abc...&method=PUT
Authorization: Bearer {serviceToken}
Response: { "url": "https://s3.../presigned" }
AppView: PUT that URL with Content-Type: application/octet-stream
5b. Large blob (outgrew the buffer): multipart, as below.
6. AppView initiates multipart upload to hold, on the first flush:
POST https://alice-storage.fly.dev/xrpc/io.atcr.hold.initiateUpload
Authorization: Bearer {serviceToken}
Body: { "digest": "sha256:abc..." }
Response: { "uploadId": "xyz" }
6. For each part:
7. For each part:
- AppView: POST /xrpc/io.atcr.hold.getPartUploadUrl
- Hold validates service token, checks crew membership
- Hold returns: { "url": "https://s3.../presigned" }
- Client uploads directly to S3 presigned URL
- AppView uploads the part to the S3 presigned URL
7. AppView completes upload:
8. AppView completes upload:
POST /xrpc/io.atcr.hold.completeUpload
Body: { "uploadId": "xyz", "digest": "sha256:abc...", "parts": [...] }
8. Manifest stored in alice's PDS:
9. Manifest stored in alice's PDS:
- holdDid: "did:web:alice-storage.fly.dev"
- holdEndpoint: "https://alice-storage.fly.dev" (backward compat)
```
+4 -4
View File
@@ -20,7 +20,7 @@ property matters most.
The **write path** inverts this. Today a push streams:
```
client --PATCH/PUT--> AppView (buffers 10MB chunks in RAM) --presigned PUT--> S3
client --PATCH/PUT--> AppView (buffers 16MB chunks in RAM) --presigned PUT--> S3
```
AppView ingests every layer and re-uploads it to S3 (`proxy_blob_store.go:586-665`
@@ -107,9 +107,9 @@ on the BYOS owner's box.
| Concern | Today |
|---|---|
| Push handshake | AppView `POST .../blobs/uploads/` -> distribution lib calls `ProxyBlobStore.Create` -> XRPC `io.atcr.hold.initiateUpload` (`proxy_blob_store.go:287-325`) |
| Push bytes | Client -> AppView RAM (10MB chunks) -> presigned S3 PUT (`proxy_blob_store.go:605-665`) |
| Push finalize | `ProxyBlobWriter.Commit` -> XRPC `io.atcr.hold.completeUpload` (`proxy_blob_store.go:703-741`) |
| Push handshake | AppView `POST .../blobs/uploads/` -> distribution lib calls `ProxyBlobStore.Create`, which makes no hold call. `io.atcr.hold.initiateUpload` is only reached if the blob outgrows the 16MB buffer. |
| Push bytes | Client -> AppView RAM (16MB buffer) -> presigned S3 PUT |
| Push finalize | `ProxyBlobWriter.Commit` verifies the received bytes against the client's digest, then either PUTs the whole buffered blob to its final key (a `com.atproto.sync.getBlob` PUT presign) or, for a blob that went multipart, calls XRPC `io.atcr.hold.completeUpload` |
| Hold upload API | Custom XRPC only: `initiateUpload`, `getPartUploadUrl`, `completeUpload`, `abortUpload`, `notifyManifest` (`pkg/hold/oci/xrpc.go:50-61`). No standard OCI `/v2` upload surface. |
| Upload `Location` | AppView-relative `/v2/<name>/blobs/uploads/<id>`, generated by the distribution library from `BlobWriter.ID()`, not by ATCR code. |
| Hold auth | Service token (Bearer, `aud`=hold DID, signed by user's PDS) or DPoP. Validated on every op (`pkg/hold/pds/auth.go:377-429` `ValidateBlobWriteAccess`, `:507-608` `ValidateServiceToken`). The hold **cannot** validate AppView's registry JWT. |
+2
View File
@@ -57,6 +57,8 @@ The endpoint routes on the `cid` parameter and answers differently per branch.
The field is additive. A client that reads only `url` behaves exactly as before, and an AppView that gets no `size` falls back to HEADing the presigned URL. `size` is never sent for a PUT presign: the object does not exist yet.
A `method=PUT` presign is a write capability (gated on `blob:write`, same as the multipart endpoints) and is how the AppView uploads a blob small enough to fit in its 16MB buffer: one presigned PUT to the blob's final `sha256:` key, instead of initiateUpload plus part URLs plus completeUpload plus the server side copy out of the temp key. The URL is signed with `Content-Type: application/octet-stream`, so the PUT has to carry that header or S3 rejects the signature.
On GET and HEAD, a blob that is not in storage is answered with **404** and a JSON error body instead of a presigned URL:
```json