mirror of
https://tangled.org/evan.jarrett.net/at-container-registry
synced 2026-09-29 05:25:35 +00:00
Three defects that all turn on whether the hold knows what a scanner is doing. Neither end had a keepalive, a read deadline, or a read limit, so a half-open connection was invisible until some other timeout fired. That got worse with capacity-aware dispatch: a dead-but-connected scanner holds its advertised worker count out of the budget and keeps winning jobs. Both ends now ping every 30s against a 90s read deadline, so three unanswered pings condemn a connection. Detection takes about 90 seconds, after which the existing reconnect grace reclaims the rows, against the 60-minute queueing timeout that was previously the only escape. Liveness decides when a scanner is gone; the grace window decides when its work is reassignable. Read limits are asymmetric and deliberately generous, because exceeding one closes the connection rather than truncating, which would turn a large but legitimate result into a permanent retry loop. Write deadlines were absent everywhere; the scanner in particular held a mutex across an unbounded write, so a wedged write silenced it without disconnecting it. handleResult did two S3 uploads and a CAR commit inline on the reader goroutine, so a slow S3 looked exactly like a dead scanner. Terminal messages now go to a per-subscriber storage goroutine while acks and starts stay on the reader. One goroutine, not a pool: every path ends in CreateScanRecord, which serialises on the repo lock anyway, and ordering is worth more than parallelism that cannot be used. Uploads and the record write get separate budgets, so a stalled upload cannot spend the time the record needs and the record write stays unconditional. checkPredecessor cached an inconclusive answer as a definitive negative in a map that is never reset, so one unreachable hold meant its manifests were never scanned again for the life of the process. gc.go already carried the corrected logic for the same problem; this follows it rather than inventing a second approach, and also stops treating an unparseable captain record as definitive. The test PDS had to move off ":memory:", which go-libsql scopes per connection: writing a scan record from any goroutine but the caller's got a connection with no tables. That was invisible while every record write happened inline. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01U1Km3N3uUmeGaj7VbaM8PF