Files
at-container-registry/docs/HOLD_DISCOVERY.md
T
Evan JarrettandClaude Fable 5.1 0080957a21 remove the runtime test_mode switch; the testmode build tag is the only one
server.test_mode survived the build-tag refactor only to feed five
behavioral branches: the registry's fall-back to the default hold when
the user's hold is unreachable, backfill warning suppression for
external holds, the appview listener close on shutdown, the hold's
relay-crawl skip, and the hold's appview-issuer tolerance. Every one of
them is a "this is a local development build" decision, which is what
the tag already says, and local development has to build with the tag
or nothing resolves. So they read atproto.TestModeBuild now, and the
flag, SetTestMode, IsTestMode, the middleware option, the backfill
constructor parameter, the never-read field on RemoteHoldAuthorizer,
the example and template YAML lines, and the docker-compose env vars
are gone. The registry keeps the fallback as a field seeded from the
constant so the production-path tests can pin it off under the tag.

The 24 SetTestMode calls in tests were dead already: stripping them and
running the affected packages tagged changed nothing.

Tests that resolve a loopback did:web used to t.Fatal naming the tag,
which left a bare `go test ./...` permanently red in five packages.
They now live under `//go:build testmode`: whole-file constraints where
every test needs it, and sibling *_testmode_test.go files holding the
moved tests plus their fixtures where a file mixed. The harness carries
the constraint too, with its package doc in an untagged doc.go so the
package still exists without it. An untagged run compiles those tests
out and passes; make test keeps the tag and runs everything.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UwYzaG3Yy7uA8FbZ5qk3tQ
2026-09-11 11:09:44 -05:00

12 KiB

Hold Discovery

Status: implemented. This document describes the hold discovery system as built. It was originally written as a design proposal; the design has since shipped with some deliberate divergences, noted in Divergences from the original proposal. Remaining gaps are listed in Remaining work.

TL;DR

AppView discovers holds by consuming io.atcr.hold.captain and io.atcr.hold.crew records from the ATProto network (Jetstream live tail + relay backfill), caches them in SQLite, and presents them to users as a grouped dropdown in Settings → Storage. Users select a default hold by DID; the selection is validated against the cache, written to the sailor profile, and mirrored into the local users table.

┌─────────────────┐     ┌─────────────────┐     ┌─────────────────┐
│  Hold Service   │────▶│     Relay       │────▶│   Jetstream     │
│ (embedded PDS)  │     │ (BGS/bigsky)    │     │                 │
└─────────────────┘     └─────────────────┘     └────────┬────────┘
                                                         │
                                                         ▼
                                                ┌─────────────────┐
                                                │    AppView      │
                                                │  (subscriber)   │
                                                └────────┬────────┘
                                                         │
                                                         ▼
                                                ┌─────────────────┐
                                                │    SQLite       │
                                                │  (cache)        │
                                                └─────────────────┘

Record Types

Both records live in the hold's embedded PDS, so the repo DID of the event is the hold DID. Go structs are in pkg/atproto/lexicon.go; lexicon JSON in lexicons/io/atcr/hold/captain.json and lexicons/io/atcr/hold/crew.json.

io.atcr.hold.captain

Singleton record (rkey self) describing the hold:

{
  "$type": "io.atcr.hold.captain",
  "owner": "did:plc:abc123",
  "public": false,
  "allowAllCrew": true,
  "enableBlueskyPosts": false,
  "deployedAt": "2025-01-07T12:00:00Z",
  "region": "us-east-1",
  "successor": "did:web:hold02.example.com"
}
Field Type Description
owner string (DID) DID of the hold owner (captain)
public boolean Anyone can read (pull) blobs without authentication
allowAllCrew boolean Any authenticated user can self-register as crew
enableBlueskyPosts boolean Post to Bluesky when manifests are pushed
deployedAt string RFC3339 deployment timestamp
region string Optional S3 region where blobs are stored
successor string (DID) Optional successor hold (migration redirect, single-hop)

io.atcr.hold.crew

One record per crew member:

{
  "$type": "io.atcr.hold.crew",
  "member": "did:plc:xyz789",
  "role": "write",
  "permissions": ["blob:read", "blob:write"],
  "tier": "deckhand",
  "addedAt": "2025-01-07T12:00:00Z"
}
Field Type Description
member string (DID) DID of the crew member
role string Role name (owner, admin, write, read)
permissions string[] blob:read, blob:write, crew:admin (blob:write implies read)
tier string Optional quota tier (deckhand, bosun, quartermaster)
plankowner boolean Optional early-adopter flag
addedAt string RFC3339 timestamp when added

Record key: deterministic hash of the member DID, enabling O(1) lookup (CrewRecordKey() in pkg/atproto/lexicon.go — SHA-256, first 16 bytes, lowercase base32 without padding).

Hold Identity

Every hold publishes two services in its DID document (HoldServices() in pkg/hold/pds/hold_pds.go):

Service ID Type Endpoint
#atproto_pds AtprotoPersonalDataServer hold public URL
#atcr_hold AtcrHoldService hold public URL

The #atcr_hold service is what distinguishes a real hold from an ordinary ATProto account. DID → endpoint resolution goes through the shared indigo identity directory (24h TTL cache) via ResolveHoldURL() / ResolveHoldDIDToURL() in pkg/atproto/resolver.go; there is no endpoint column in the cache tables.

Captain record verification

Because any ATProto account can publish an io.atcr.hold.captain record, the discovery pipeline verifies the publishing DID before caching: HasHoldService() (pkg/atproto/resolver.go) resolves the DID and requires an #atcr_hold service endpoint. Records from non-hold DIDs are skipped (logged, not errored); unresolvable DIDs are skipped and retried by the periodic backfill. This gating happens in both the live processor (ProcessCaptain, pkg/appview/jetstream/processor.go) and relay backfill (batchCaptains, pkg/appview/jetstream/backfill_batch.go).

Crew records are not independently verified: they only become visible through joins against a verified captain record, so gating captains gates both. This prevents two abuse cases: spam holds appearing in every user's picker (allowAllCrew: true), and forged crew records placing a fake hold in targeted users' member lists. It does not defend against a malicious actor running a real hold service — that is inherent to open federation, same as any open-registration hold.

In a -tags testmode build, local did:web identifiers (HTTP, IP:port, localhost) resolve through the directory wrapper in pkg/atproto/indigo_local.go, so the same verification applies to them. A production build cannot resolve them at all.

Data Model

Cache tables in pkg/appview/db/schema.sql (SQLite/libsql, source of truth is the network — the cache is rebuilt by backfill):

CREATE TABLE IF NOT EXISTS hold_captain_records (
    hold_did TEXT PRIMARY KEY,
    owner_did TEXT NOT NULL,
    public BOOLEAN NOT NULL,
    allow_all_crew BOOLEAN NOT NULL,
    deployed_at TEXT,
    region TEXT,
    successor TEXT,
    updated_at TIMESTAMP DEFAULT CURRENT_TIMESTAMP
);

CREATE TABLE IF NOT EXISTS hold_crew_members (
    hold_did TEXT NOT NULL,
    member_did TEXT NOT NULL,
    rkey TEXT NOT NULL,          -- for delete-event handling
    role TEXT,
    permissions TEXT,            -- JSON array
    tier TEXT,
    added_at TEXT,
    created_at TIMESTAMP DEFAULT CURRENT_TIMESTAMP,
    updated_at TIMESTAMP DEFAULT CURRENT_TIMESTAMP,
    PRIMARY KEY (hold_did, member_did)
);

Related tables hold_crew_approvals / hold_crew_denials support the crew self-registration flow and are not part of discovery itself.

The user's selected default hold is mirrored in users.default_hold_did (the authoritative copy lives in the sailor profile record on the user's PDS).

Ingestion

Live tail (Jetstream)

The worker (pkg/appview/jetstream/worker.go) subscribes with a wildcard:

wantedCollections: []string{
    "io.atcr.*",              // includes io.atcr.hold.captain / io.atcr.hold.crew
    "app.bsky.actor.profile", // avatar sync
}

ProcessRecord (pkg/appview/jetstream/processor.go) routes captain/crew events to ProcessCaptain / ProcessCrew, which upsert into the cache tables; delete events call DeleteCaptainRecord / DeleteCrewMemberByRkey (crew deletes are matched by the stored rkey, since delete events carry no record body).

Relay backfill

BackfillWorker (pkg/appview/jetstream/backfill.go) runs on startup and periodically when jetstream.backfill_enabled is set. It lists repos per collection from the configured relay endpoints and batch-upserts records (batchCaptains / batchCrew in backfill_batch.go — single batched writes to avoid long transactions against remote libsql). Captain and crew are in the standard backfill collection list.

Direct bootstrap (primary hold)

Holds not yet crawled by a relay are invisible to both paths above, so backfill also queries the AppView's primary managed hold (server.managed_holds[0]) directly via queryCaptainRecord / queryCrewRecords (XRPC getRecord/listRecords against the hold's own PDS, with startup retries and a freshness memo).

Queries

pkg/appview/db/hold_store.go:

Function Purpose
GetAvailableHolds(db, userDID) Holds the user can select, with Membership: owner / crew / eligible
GetAccessibleHoldDIDs(db, viewerDID) Hold DIDs whose content the viewer may see in listings (includes public)
GetCrewMemberships(db, memberDID) Reverse lookup: holds where the user is crew
GetCrewHoldDID(db, memberDID) Most recent crew hold, fallback when no default is cached
GetCaptainRecord(db, holdDID) / UpsertCaptainRecord / ListHoldDIDs Cache plumbing

GetAvailableHolds excludes holds that declare a successor (they are mid-migration) and excludes pure-public holds (see divergences below).

Settings UI

Settings → Storage (pkg/appview/handlers/settings.go):

  • buildHoldsData() calls GetAvailableHolds, resolves display names, attaches health status, and computes ReadOnly (crew without blob:write — pushes would be rejected).
  • hold_selector.html renders a <select name="hold_did"> with optgroups Your Holds (owner + crew, with (Crew, read-only) markers) and Available Holds (eligible, allowAllCrew holds the user can join). hold_card.html shows the active hold with membership badge, online/offline status, and storage stats.
  • The form posts the DID to POST /api/profile/default-hold (UpdateDefaultHoldHandler), which validates the user actually has access via GetAvailableHolds, writes defaultHold to the sailor profile on the user's PDS, mirrors it into users.default_hold_did, and kicks a background refresh of the captain/crew cache for the chosen hold.

Cache freshness: Jetstream gives near-real-time updates; the periodic backfill reconciles anything missed. There is no manual refresh button — saving a selection refreshes the relevant hold's records.

Divergences from the original proposal

This doc previously proposed a design that differs from what shipped:

  • Field names: the proposal used ownerDid / memberDid; the real fields are owner / member. It also omitted successor, enableBlueskyPosts, and plankowner.
  • No "Public Holds" group in the picker: a default hold is a push target and pushes require crew membership, so listing public-read-only holds would offer broken options. Public holds still matter for read-side visibility (GetAccessibleHoldDIDs).
  • No provider field / column: never added to the lexicon or schema.
  • No cached endpoint column: DID → URL resolution goes through the identity directory cache instead.

Remaining work

  • Bootstrap beyond the primary hold: direct bootstrap only queries managed_holds[0]. A deployment with multiple managed holds (or a need to seed un-crawled third-party holds, e.g. an ATCR_BOOTSTRAP_HOLDS list) would need the bootstrap loop extended. Deferred until such a deployment exists; relay backfill covers crawled holds.