Commit Graph
3 Commits
Author SHA1 Message Date
Evan JarrettandClaude Opus 5 8d7ccd7cb7 apply go fix modernizations across the workspace
`go fix` carries the modernize analyzers now, and the tree had drifted behind
them. This is the mechanical result, reviewed rather than trusted: the tool is
capable of rewriting code into something that no longer tests or does what it
did, so every non-test change was read individually and the concurrency-bearing
packages were re-run under -race.

Production code, four changes, all semantics-preserving:

  - leases/manager.go: wg.Add(1) + go + defer wg.Done() becomes wg.Go. The
    comment above that function turns on Add happening before the goroutine
    starts, so that a Wait cannot return before the worker has run. wg.Go does
    the Add synchronously on the calling goroutine, so the invariant it
    describes still holds.
  - auth/token/handler.go: strings.Fields -> strings.FieldsSeq, same splitting,
    iterated rather than allocated.
  - hold/gc/gc.go: a hand-written map copy -> maps.Copy.
  - hold/pds/scan_broadcaster.go: three-clause loop -> range over int.

The rest are tests. The one worth naming is carstore_contention_test.go, where a
careless rewrite could have quietly stopped exercising contention: go fix
converted the reader and side-table goroutines to loopWG.Go but correctly
declined to touch the writer loop, which passes its index as a parameter. The
writer/reader/side-table shape and the stop channel are unchanged, so the test
still contends over the same carstore transactions.

Verified: go build for hold and appview, `make lint` 0 issues, the deploy and
credential-helper modules 0 issues, `make test` green across all 43 packages,
and -race green on leases, hold/pds, hold/gc and auth/token. The scanner module's
two lint findings are unchanged from HEAD and are in files go fix never touched.

Kept separate from the HTTP/2 commit so that one stays readable, and so this can
be reverted on its own if a modernization turns out to matter.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TA9D4DjaLZTvzQ7dJbu4eg
2026-09-08 22:38:01 -05:00
Evan JarrettandClaude Opus 5 6a7ddb819b appview: run background workers under a lease
The Jetstream consumer, backfill, labeler subscriber, cleanup sweep and billing
tier refresh all started unconditionally in every process. That is correct for
one instance and wrong for two.

The consumer is the case with teeth. StatsCache is per-process in-memory state,
and the aggregate it produces is written to repository_stats as an absolute
value rather than an increment, so two consumers each hold a partial view of the
holds and each write their partial sum as though it were the whole truth,
overwriting one another indefinitely. The webhook dispatcher hangs off the same
processor, so a second consumer also doubles every delivery.

Each now runs under a named lease, so exactly one instance runs it and a
replacement takes over when that instance goes away. The health worker is
deliberately not leased: it refreshes a cache each instance needs locally, so
running it everywhere is correct.

Two structural changes came with it. The cleanup loop moved out of
InitializeDatabase, where it was a bare goroutine with no way to reach the lease
manager, into RunPeriodicCleanup called from the server. And backfill's startup
run and periodic schedule became one leased worker instead of two goroutines on
context.Background(), so shutdown actually stops a backfill in flight rather
than letting it run on against a closing database. With interval=0 that worker
holds its lease instead of returning, since releasing would let another instance
acquire and run its own startup backfill, turning "once" into "once per
instance".

Verified with two instances against one database: exactly one acquired, the
other contended without starting a worker; SIGTERM handed over in 13ms via the
release, SIGKILL handed over in ~12s via TTL expiry.

That first number only holds because of Manager.Go and Manager.Wait, which this
commit adds. The first cut used `go m.Run(...)` and cancelled the worker context
during shutdown without waiting, so the process exited before the release landed
and the lease survived to its TTL — a rolling deploy would have paused indexing
for a minute rather than a second. Nothing in the unit tests caught it; the
two-instance run did. TestWaitBlocksUntilLeaseReleased covers it now.

leases.enabled defaults to true. A single instance is unaffected, since it
always wins its own leases, while an operator who scales out without reading the
docs still gets correct behavior instead of silent stats corruption.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-11 21:26:46 -05:00
Evan JarrettandClaude Opus 5 f84e8ffa27 db: create instance_leases and add the lease manager
Groundwork for running more than one AppView instance. Nothing is wired to this
yet; the next commit moves the background workers onto it.

Several workers must run on exactly one instance. The Jetstream consumer is the
sharpest case: StatsCache is per-process in-memory state, and the aggregate it
produces is written to repository_stats as an absolute value rather than an
increment, so two consumers would each hold a partial view of the holds and each
write its partial sum as the whole truth, overwriting one another indefinitely.
The webhook dispatcher hangs off the same processor, so a second consumer also
means every webhook fires twice.

Instances contend for a named lease; only the holder runs the worker. Acquire is
a single INSERT ... ON CONFLICT ... WHERE, so two instances racing for the same
expired lease cannot both win: the loser's update matches no rows. The fence
token increments on every change of custody, so a process that stalled past its
TTL discovers on its next renewal that it was superseded, rather than continuing
to act as the holder.

A renewal blackout is treated as a loss. If the database has been unreachable
for longer than the TTL, another instance is entitled to steal the lease and we
must assume it has, even though we cannot ask. Continuing to work in that state
is the one outcome the lease exists to prevent.

Clean shutdown expires the lease in place rather than deleting the row, so a
replacement starts in seconds instead of waiting out the TTL, while the fence
token survives to keep a stalled former holder from matching again.

Timestamps are Unix milliseconds, not TIMESTAMP text. libSQL normalizes
date-like TEXT on the way in, and Go's driver and CURRENT_TIMESTAMP disagree on
format, so a stored expiry and a literal would compare as strings that sort
differently. That comparison is the whole safety property, so it does not get to
be subtle. The cost is a dependency on roughly-synced clocks, the same
assumption Kubernetes leases make; keep the TTL well above any plausible skew.

The lease tests are file-backed rather than :memory:. go-libsql gives every
connection to an in-memory DSN its own private database, so with MaxOpenConns of
8 a second goroutine lands on a connection where the schema was never applied
("no such table"). Every existing test in the package is sequential and reuses
one pooled connection, which is why this has stayed invisible.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-11 21:17:55 -05:00