Files
at-container-registry/docs/BYOS.md
T
Evan JarrettandClaude Fable 5.1 f4343d7956 appview: upload small blobs with one presigned PUT, and verify every digest
Every blob went through the multipart machinery: an S3 multipart started
on Docker's initial POST, a hold round trip per part, and a complete on
the hold that finished the multipart, HEADed the temp object, copied it
to its final key, and deleted the temp. For a 2KB config blob that was
three hold calls and six S3 operations. On production data 86% of
distinct layers and every config blob fit in a 16MB buffer, and 49% of
image manifests have no layer larger than that.

The writer now buffers up to 16MB (also the multipart part size) and
makes no hold call until it has to. A blob that never overflows the
buffer is written at Commit with a single presigned PUT to its final
key, via the hold's existing method=PUT presign; the multipart only
starts on the first flush. The hold's completeUpload does nothing the
direct path skips: quota, layer records, stats and scan dispatch all
hang off notifyManifest, which is unchanged.

The buffer starts empty and grows on demand, with the doubling capped so
capacity never overshoots 16MB: a config blob costs kilobytes, and only
layers that approach the threshold fill it.

Bytes are hashed as they arrive. Commit compares the computed sha256 to
the digest the client claimed before any network call, and returns
DIGEST_INVALID on mismatch, aborting a multipart if one was started.
Previously nothing verified the content, so a pusher could store wrong
bytes under a digest in the shared content-addressed space.

Tests observe request counts on a fake hold and fake S3 rather than
return values. The growth test streams in 24KB chunks because
power-of-two chunks land on 16MB by luck and hid an earlier weaker guard.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018Yf1ZVA7sXYhQNb9tCo1m5
2026-09-09 09:50:08 -05:00

13 KiB

Bring Your Own Storage (BYOS)

Overview

ATCR supports "Bring Your Own Storage" (BYOS) for blob storage. Users can:

  • Deploy their own hold service with embedded PDS
  • Control access via crew membership in the hold's PDS
  • Keep blob data in their own S3-compatible storage (AWS S3, Storj, Minio, UpCloud, etc.) while manifests stay in their user PDS

Architecture

┌──────────────────────────────────────────┐
│ ATCR AppView (API)                       │
│ - Manifests → User's PDS                 │
│ - Auth & service token management        │
│ - Blob routing via XRPC                  │
│ - Profile management                     │
└────────────┬─────────────────────────────┘
             │
             │ Hold discovery (findHoldDIDAndProfile):
             │ 1. io.atcr.sailor.profile.defaultHold (DID)
             │ 2. AppView default hold (server.managed_holds[0])
             │
             │ Then resolveSuccessor: if the chosen hold's
             │ captain record sets a successor DID, apply a
             │ single-hop redirect to it (hold migration).
             ▼
┌──────────────────────────────────────────┐
│ User's PDS                               │
│ - io.atcr.sailor.profile (hold DID)      │
│ - io.atcr.manifest (with holdDid)        │
└────────────┬─────────────────────────────┘
             │
             │ Service token from user's PDS
             ▼
┌──────────────────────────────────────────┐
│ Hold Service (did:web:hold.example.com)  │
│ ├── Embedded PDS                         │
│ │   ├── Captain record (ownership)       │
│ │   └── Crew records (access control)    │
│ ├── XRPC multipart upload endpoints      │
│ └── Storage driver (S3/Storj/etc.)       │
└──────────────────────────────────────────┘

Hold Service Components

Each hold is a full ATProto actor with:

  • DID: did:web:hold.example.com (hold's identity)
  • Embedded PDS: Stores captain + crew records (shared data)
  • Storage backend: S3-compatible (AWS S3, Storj, Minio, UpCloud, etc.)
  • XRPC endpoints: Standard ATProto + custom OCI multipart upload

Records in Hold's PDS

Captain record (io.atcr.hold.captain/self):

{
  "$type": "io.atcr.hold.captain",
  "owner": "did:plc:alice123",
  "public": false,
  "allowAllCrew": false,
  "enableBlueskyPosts": false,
  "deployedAt": "2025-10-14T...",
  "region": "iad",
  "successor": ""
}

region and successor are optional. successor holds the DID of a replacement hold; when set, the AppView applies a single-hop redirect to it during hold discovery (see the Architecture diagram above).

Crew records (io.atcr.hold.crew/{rkey}):

{
  "$type": "io.atcr.hold.crew",
  "member": "did:plc:bob456",
  "role": "captain",
  "permissions": ["blob:read", "blob:write"],
  "tier": "bosun",
  "plankowner": false,
  "addedAt": "2025-10-14T..."
}

Authorization is driven by the permissions array (blob:read, blob:write, crew:admin), not the role string. blob:write implicitly grants blob:read (you can't push without being able to pull). tier and plankowner are optional and feed quota limits.

Sailor Profile (User's PDS)

Users set their preferred hold in their sailor profile:

{
  "$type": "io.atcr.sailor.profile",
  "defaultHold": "did:web:hold.example.com",
  "createdAt": "2025-10-02T...",
  "updatedAt": "2025-10-02T..."
}

Deployment

Configuration

The hold service is configured with Viper: a YAML file is the primary source, and environment variables override individual fields. Env var names are HOLD_ plus the YAML path with _ separators (e.g. server.public_url → HOLD_SERVER_PUBLIC_URL). S3 credentials use the standard AWS names.

Generate a fully commented config and run with it:

./bin/atcr-hold config init config-hold.yaml
# edit config-hold.yaml, then:
./bin/atcr-hold serve --config config-hold.yaml

Key fields (YAML on the left, env override on the right):

server:
  public_url: https://hold.example.com   # HOLD_SERVER_PUBLIC_URL (REQUIRED)
  public: false                          # HOLD_SERVER_PUBLIC (allow anonymous reads)

registration:
  owner_did: did:plc:your-did-here       # HOLD_REGISTRATION_OWNER_DID
  allow_all_crew: false                  # HOLD_REGISTRATION_ALLOW_ALL_CREW

database:
  path: /var/lib/atcr-hold               # HOLD_DATABASE_PATH (carstore + SQLite)
  key_path: ""                           # HOLD_DATABASE_KEY_PATH (defaults to {path}/signing.key)

storage:
  bucket: my-blobs                       # S3_BUCKET (REQUIRED)
  region: us-east-1                      # AWS_REGION
  endpoint: ""                           # S3_ENDPOINT (for non-AWS providers)

S3 credentials are read from the standard AWS env vars (AWS_ACCESS_KEY_ID, AWS_SECRET_ACCESS_KEY).

Running Locally

For local development, use Minio as an S3-compatible storage:

# Start Minio (in separate terminal)
docker run -p 9000:9000 -p 9001:9001 minio/minio server /data --console-address ":9001"

# Build
go build -o bin/atcr-hold ./cmd/hold

# Run (env overrides shown; a YAML config works too)
export HOLD_SERVER_PUBLIC_URL=http://localhost:8080
export HOLD_REGISTRATION_OWNER_DID=did:plc:your-did-here
export AWS_ACCESS_KEY_ID=minioadmin
export AWS_SECRET_ACCESS_KEY=minioadmin
export S3_BUCKET=test
export S3_ENDPOINT=http://localhost:9000
export HOLD_DATABASE_PATH=/tmp/atcr-hold

./bin/atcr-hold serve

On first run, the hold service creates:

  • Captain record in embedded PDS (making you the owner)
  • Crew record for owner with all permissions
  • DID document at /.well-known/did.json

Deploy to Fly.io

# Create fly.toml
cat > fly.toml <<EOF
app = "my-atcr-hold"
primary_region = "ord"

[env]
  HOLD_SERVER_PUBLIC_URL = "https://my-atcr-hold.fly.dev"
  AWS_REGION = "us-east-1"
  S3_BUCKET = "my-blobs"
  HOLD_SERVER_PUBLIC = "false"
  HOLD_REGISTRATION_ALLOW_ALL_CREW = "false"

[http_service]
  internal_port = 8080
  force_https = true
  auto_stop_machines = true
  auto_start_machines = true
  min_machines_running = 0

[[vm]]
  cpu_kind = "shared"
  cpus = 1
  memory_mb = 256
EOF

# Deploy
fly launch
fly deploy

# Set secrets
fly secrets set AWS_ACCESS_KEY_ID=...
fly secrets set AWS_SECRET_ACCESS_KEY=...
fly secrets set HOLD_REGISTRATION_OWNER_DID=did:plc:your-did-here

Request Flow

Push with BYOS

1. Client: docker push atcr.io/alice/myapp:latest

2. AppView resolves alice → did:plc:alice123

3. AppView discovers hold DID:
   - Check alice's sailor profile for defaultHold
   - Returns: "did:web:alice-storage.fly.dev"

4. AppView gets service token from alice's PDS:
   GET /xrpc/com.atproto.server.getServiceAuth?aud=did:web:alice-storage.fly.dev
   Response: { "token": "eyJ..." }

5. AppView buffers the blob (16MB limit) and verifies the bytes it received
   against the digest the client claimed. A mismatch is rejected here, before
   anything reaches storage.

5a. Small blob (fits in the buffer, which is every config blob and most layers):
   AppView asks for one write capability and PUTs the whole blob to its final
   content-addressed key. No multipart session, no temp object, no copy.
   GET /xrpc/com.atproto.sync.getBlob?did=...&cid=sha256:abc...&method=PUT
   Authorization: Bearer {serviceToken}
   Response: { "url": "https://s3.../presigned" }
   AppView: PUT that URL with Content-Type: application/octet-stream

5b. Large blob (outgrew the buffer): multipart, as below.

6. AppView initiates multipart upload to hold, on the first flush:
   POST https://alice-storage.fly.dev/xrpc/io.atcr.hold.initiateUpload
   Authorization: Bearer {serviceToken}
   Body: { "digest": "sha256:abc..." }
   Response: { "uploadId": "xyz" }

7. For each part:
   - AppView: POST /xrpc/io.atcr.hold.getPartUploadUrl
   - Hold validates service token, checks crew membership
   - Hold returns: { "url": "https://s3.../presigned" }
   - AppView uploads the part to the S3 presigned URL

8. AppView completes upload:
   POST /xrpc/io.atcr.hold.completeUpload
   Body: { "uploadId": "xyz", "digest": "sha256:abc...", "parts": [...] }

9. Manifest stored in alice's PDS:
   - holdDid: "did:web:alice-storage.fly.dev"
   - holdEndpoint: "https://alice-storage.fly.dev" (backward compat)

Pull with BYOS

1. Client: docker pull atcr.io/alice/myapp:latest

2. AppView fetches manifest from alice's PDS

3. Manifest contains:
   - holdDid: "did:web:alice-storage.fly.dev"

4. Client requests blob: GET /v2/alice/myapp/blobs/sha256:abc123

5. AppView reads the hold DID from the manifest's holdDid field (per request)

6. AppView gets service token from alice's PDS
   (validated service tokens are cached ~45s to absorb a burst of blob requests)

7. AppView calls hold XRPC:
   GET /xrpc/com.atproto.sync.getBlob?did={userDID}&cid=sha256:abc123
   Authorization: Bearer {serviceToken}
   Response: { "url": "https://s3.../presigned-download" }

8. AppView redirects client to presigned S3 URL

9. Client downloads directly from S3

Key insight: Pull uses the holdDid stored in the manifest, ensuring blobs are fetched from where they were originally pushed.

Access Control

Read Access

  • Public hold (server.public: true): Anonymous + authenticated users
  • Private hold (server.public: false): Authenticated users with crew membership

Write Access

  • Hold owner (captain) OR crew members only
  • Verified via io.atcr.hold.crew records in hold's embedded PDS
  • Service token proves user identity (from user's PDS)

Authorization Flow

1. AppView gets service token from user's PDS
2. AppView sends request to hold with service token
3. Hold validates service token (checks it's from user's PDS)
4. Hold extracts user's DID from token
5. Hold checks crew records in its embedded PDS
6. If crew member found → allow, else → deny

Managing Crew Members

Add Crew Member

Use ATProto client to create crew record in hold's PDS:

# Via XRPC (if hold supports it)
POST https://hold.example.com/xrpc/io.atcr.hold.requestCrew
Authorization: Bearer {userOAuthToken}

# Or manually via captain's OAuth to hold's PDS
atproto put-record \
  --pds https://hold.example.com \
  --collection io.atcr.hold.crew \
  --rkey "{memberDID}" \
  --value '{
    "$type": "io.atcr.hold.crew",
    "member": "did:plc:bob456",
    "role": "crew",
    "permissions": ["blob:read", "blob:write"]
  }'

Remove Crew Member

atproto delete-record \
  --pds https://hold.example.com \
  --collection io.atcr.hold.crew \
  --rkey "{memberDID}"

Storage Backends

Hold service requires S3-compatible storage. Supported providers:

  • AWS S3 - Amazon Simple Storage Service
  • Storj - Decentralized cloud storage (via S3 gateway)
  • Minio - High-performance object storage (great for local development)
  • UpCloud - European cloud provider
  • Azure - Azure Blob Storage (via S3-compatible API)
  • GCS - Google Cloud Storage (via S3-compatible API)

Example: Team Hold

# 1. Deploy hold service
export HOLD_SERVER_PUBLIC_URL=https://team-hold.fly.dev
export HOLD_REGISTRATION_OWNER_DID=did:plc:admin
export HOLD_SERVER_PUBLIC=false  # Private
export AWS_ACCESS_KEY_ID=...
export AWS_SECRET_ACCESS_KEY=...
export S3_BUCKET=team-blobs

fly deploy

# 2. Hold auto-creates captain + crew records on first run

# 3. Admin adds team members via hold's PDS (requires OAuth)
# (TODO: Implement crew management UI/CLI)

# 4. Team members set their sailor profile:
atproto put-record \
  --collection io.atcr.sailor.profile \
  --rkey "self" \
  --value '{
    "$type": "io.atcr.sailor.profile",
    "defaultHold": "did:web:team-hold.fly.dev"
  }'

# 5. Team members can now push/pull using team hold

Limitations

Current IAM Challenges

See EMBEDDED_PDS.md for detailed discussion.

Known issues:

  1. RPC permission format: Service tokens don't work with IP-based DIDs in local dev
  2. Dynamic hold discovery: AppView can't dynamically OAuth arbitrary holds from sailor profiles
  3. Manual profile management: No UI for updating sailor profile (must use ATProto client)

Workaround: Use hostname-based DIDs (did:web:hold.example.com) and public holds for now.

Future Improvements

  1. Crew management UI - Web interface for adding/removing crew members
  2. Dynamic OAuth - Support for arbitrary BYOS holds without pre-configuration
  3. Hold migration - Tools for moving blobs between holds
  4. Storage analytics - Track usage per user/repository
  5. Distributed cache - Redis for hold DID cache in multi-instance deployments

References