Pins the architecture principle that V3 egress components (per-peer
shipper / session pump / flusher / barrier driver) must be modeled
as a single decision core with single-queue / single serializable
worker shape. Monotonic pointers (cursor, applied LSN, emit profile,
bound conn) advance only via the core's internal transitions; external
callers deliver commands/events; direct external mutation is a
design-debt side door.
Grounds the principle in three hardware-validated incidents on m01/M02
during 2026-05-02, all of which turn out to be the same shape:
§3.1 executor.Ship overwrites the WalShipper's emit context mid-
session (g7 #5: 498 LBA mismatches at concurrent-write range)
§3.2 PrimaryBridge onStart/onClose dropped ReplicaID; engine and
runtime peer state diverged (g7 #5/#6 dispatch never fires)
§3.3 A-class sender→coord RecordBarrierWalLegOk side-write
(reverted 2026-05-02 working tree per this principle)
Companion to v3-rebuild-from-lsn-pin-clarification.md. Anchors in
consensus: §I P1, §I P7, §6.8, INV-SINGLE.
No new wire field, predicate, or invariant; clarifies a rule that
earlier docs imply but don't articulate. Provides judgment criterion
for upcoming Ship/PushLiveWrite collapse and peer-state ownership
work.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Pins the rebuild path's `fromLSN=0` sentinel semantic for the current
`StartRebuild` signature. Closes the gap §I P7 calls out ("transport
silently overwriting `fromLSN := 0` violates parity") in the absence of
an engine-published `fromLSN` for rebuild.
Hardware-validated by seaweed_block@bc4286e g7-dual-lane on m01/M02:
G7-#2 PASS (dispatch=1s, complete=1s, total=2s, 1000 LBAs byte-equal).
Sentinel rule:
- Caller passes 0 ⇒ "rebuild — primary picks the pin"
- Transport translates: sessionFromLSN := targetLSN
- Receiver-visible fromLSN = targetLSN
- Future catch-up (engine surfaces fromLSN := replicaLSN): passthrough
Three transport mechanics that satisfy the rule without violating any
consensus invariant: sentinel translation in startRebuildDualLane,
cursor-caught-up shortcut in WalShipper.DrainBacklog (preserves the
recycle gate's <= strictness verbatim), SeedWalApplied at SessionStart
so base-only rebuilds satisfy the A-class TryComplete conjunct.
Anti-discipline: no new wire fields, predicates, or invariants. Memo
clarifies the meaning of an existing transport-layer constant.
Anchors: v3-recovery-algorithm-consensus.md §I P7 / §6.9 / §6.10;
recover-semantics-adjustment-plan.md §1.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Adds the WalShipper implementation mini-plan that bridges
v3-recovery-wal-shipper-spec.md to the seaweed_block layout (phased
PR rollout P0..P4, INV ↔ test mapping, reviewer checklist).
§10 P2d decision request — the architect-gated handoff:
P2c is closed (slice A / B-1 / B-2 merged on g7-redo/wal-shipper-impl).
The bridging senderBacklogSink owns the live-write buffer + flushAndSeal
under sinkMu; Sender.Run barriers as soon as sink.DrainBacklog returns;
Close/closeCh/liveQueue/drainAndSeal are deleted from Sender. Atomic-seal
contract migrates intact (capture-vs-reject from queueMu → sinkMu).
P2d is gated on a three-axis decision the architect must make before a
real transport.WalShipper sink can replace the bridging path:
1. Body format on the dual-lane port:
(A) MsgShipEntry payload (unify on legacy steady encoding), OR
(B) frameWALEntry payloads (teach WalShipper.Emit to encode), OR
(C) documented third (e.g. envelope byte).
2. Single applier owner:
recovery.Receiver vs transport replica handler.
3. Replay source of truth:
which encoding the on-disk WAL playback decoder reads.
§10 also lists pre-decision deliverables that can land in parallel:
adapter scaffolding (transport-side struct satisfying recovery.WalShipperSink
by duck typing) + integration tests for architect rules 1+2 (emit context
before StartSession; restore steady lineage after EndSession).
V2 wire-compat is gated separately per feedback_porting_discipline.md.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Per repository policy: dev/design docs live in
seaweedfs/sw-block/design/, not in seaweed_block/docs/. Formal
product docs come later. This commit relocates the 7 recovery
design markdown docs (4 trunk-merged in seaweed_block phase-15;
3 in-flight on g7-redo branches) plus the 1 hardware canonical
YAML to sw-block/design/ with v3-recovery-* prefix to match the
existing naming pattern (v3-recovery-live-line-backlog-spec.md).
Companion cleanup: a follow-on PR on seaweed_block removes the
docs from docs/ (and the YAML from testrunner/scenarios/) — that
PR is the seaweed_block side of the relocation.
Files added:
v3-recovery-pin-floor-wire.md — was docs/recovery-pin-floor-wire.md
on seaweed_block phase-15 (PR #11+#16)
v3-recovery-wiring-plan.md — was docs/recovery-wiring-plan.md
(PR #13)
v3-recovery-execution-institution.md — was docs/recovery-execution-institution.md
v3-recovery-inv-test-map.md — was docs/recovery-inv-test-map.md
(PR #11/#14/#15)
v3-recovery-unified-wal-stream-kickoff.md — was docs/recovery-unified-wal-stream-kickoff.md
g7-redo/unified-wal-kickoff (v0.3)
v3-recovery-unified-wal-stream-mini-plan.md — was docs/recovery-unified-wal-stream-mini-plan.md
g7-redo/unified-wal-mini-plan (v0.2)
v3-recovery-dual-lane-canonical-runbook.md — was docs/recovery-dual-lane-canonical-runbook.md
g7-redo/hardware-canonical-paper
v3-recovery-dual-lane-canonical.yaml — was testrunner/scenarios/recovery-dual-lane-canonical.yaml
g7-redo/hardware-canonical-paper
Internal cross-references updated in-place via sed:
- docs/recovery-inv-test-map.md → v3-recovery-inv-test-map.md
- docs/recovery-pin-floor-wire.md → v3-recovery-pin-floor-wire.md
- docs/recovery-wiring-plan.md → v3-recovery-wiring-plan.md
- testrunner/scenarios/recovery-dual-lane-canonical.yaml →
v3-recovery-dual-lane-canonical.yaml
Hand-edits:
- runbook §1 companion-YAML link: was
"../v3-recovery-dual-lane-canonical.yaml" (parent dir from
seaweed_block/docs); now same-directory link in design/.
- runbook §8 §3.2 #3 reference: was relative to seaweed_block
memory file (../../.claude/...); rewritten to point to
v3-recovery-unified-wal-stream-kickoff.md §4 directly.
- mini-plan Q15: docs/archive/ wording updated to
sw-block/design/archive/.
Stages-of-evidence still readable from the docs themselves
(kickoff §11, mini-plan §10 resolution logs, inv-test-map row
versions). Original seaweed_block branches preserve git
history for the in-flight content; the cleanup PR closes them
once this lands.
NOTE: this commit does NOT include the user's unrelated
ongoing edits in feature/sw-block (M v3-batch-process.md,
M v3-dev-roadmap.md, M v3-phase-15-g6-mini-plan.md, etc.).
Those stay uncommitted for the user to handle separately.
QA's G7 pre-work surfaced a discrepancy: the v0.1 §harness-notes
pointed at `exec_rebuild_started` / `exec_rebuild_completed` as
the harness markers. Those are RecoveryLog event names (internal
Orchestrator.Log ring buffer, process-local) — NOT visible in
primary.log on hardware. Hardware harnesses can't scrape them
without a /recovery-log HTTP surface (G5-3 forward-carry).
Hardware-visible markers (corrected):
- START: `executor: rebuild start replica=<id> sessionID=<n>
epoch=<n> EV=<n> targetLSN=<n>`
from core/transport/rebuild_sender.go:41
(added at G6 #1, seaweed_block@85475cd)
- COMPLETE: `executor: rebuild complete, sent <n> blocks
(targetLSN=<n>)`
from core/transport/rebuild_sender.go:120
(pre-existing T4d-4 part B / earlier)
Both produced via log.Printf in rebuild_sender.go and routed to
the daemon's stdout/stderr stream (which iterate harness captures
to ${REMOTE_RUN_DIR}/logs/primary.log). Both are sessionID-
correlatable for chained-scenario filtering. The G6 hardware run
already proved the START marker pattern; COMPLETE follows the
same shape.
Files corrected:
- §2 #7 acceptance row (harness helper text)
- §2 entry-marker table row
- §3 risks "Ambiguous rebuild done vs peer healthy" row
- §harness-notes (full rewrite with v0.1 correction note +
marker table + RecoveryLog clarification + recommended helper
shape with sessionID filter)
Negative-references to RecoveryLog event names retained in
explanatory context (so future readers don't re-introduce the
mistake by reading the engine code in isolation).
QA pre-work artifact V:\share\g5-test\scenarios\g7-helpers.sh is
already written against the corrected literals; this commit
brings the §harness-notes source-of-truth into alignment.
Standing by for architect §1.A ratification (Q1 topology / Q2
fold-G6 / Q3 deadline / etc.) before §1.H code-start audit.
Per architect ruling 2026-04-28 + sw §close.appendix: D's WALRecycled
boundary finding is G6 territory, not a G5-5C reopener. Adding the
backlog ticket here so it doesn't get lost between G5-5C close and
G6 kickoff.
Ticket text + evidence pointer + cross-references all preserved
from the §close.appendix; this is the dev-roadmap-side mirror so
the ticket surfaces when planning G6 scope.
Standing by for architect final §close single-sign on G5-5C.
Per architect ruling 2026-04-28 on QA's expanded scenario report:
- A (capacity): 🐛 → ✅ already-fixed at seaweed_block@a250b52, INV inscribed.
- B (500 random LBAs over 65536-LBA volume): ✅ GREEN. Confidence
bump on dirty-map skew + ship order under random write pattern.
- C (kill replica mid-write-storm + restart + 200 LBAs converge):
✅ GREEN. Highest-signal recovery scenario in the expansion;
validates G5-5C peer-recovery trigger under load.
- D (5000-LBA sustained write → WALRecycled past replica LSN):
🐛 boundary finding. Architect: G6 territory, not G5-5C reopener.
Catch-up requires WAL retention; rebuild path is for gap-beyond-
WAL. Engine has dispatch-branch tests (Batch 4); runtime
escalation path under sustained pressure is G6 acceptance scope.
Doc updates:
- New §close.appendix table with all 4 scenario rows + dispositions.
- Semantic clarification on D — catch-up vs WAL recycle vs rebuild.
- §close.forward-carries gets a NEW G6 entry with backlog ticket
text, evidence pointer, cross-reference to INV-G5-5C-PROBE-BEFORE-
CATCHUP, and explicit non-reopener rationale.
- Logs + scenario script paths recorded for QA continuity.
§close substance unchanged: G5-5C gate (verify_restart_catchup
GREEN within 30 s) was met on the canonical case at
seaweed_block@712cbc47 + capacity addendum at a250b52. B/C are
strengthening, not gating; D is forward-carry.
Awaiting architect final §close single-sign on this tree.
Per architect ruling 2026-04-28 + sw addendum landing at
seaweed_block@a250b52: inscribe new INV in the ledger.
Statement: iSCSI/NVMe externally-visible volume capacity and block
size MUST derive from --durable-blocks × --durable-blocksize when
--durable-root is set, not silently fall back to frontend defaults
(DefaultVolumeBlocks=2048 × DefaultBlockSize=512 = 1 MiB). Without
this plumb-through, a daemon configured for N MiB durable storage
advertises a 1 MiB iSCSI/NVMe LUN and any workload above LBA 256
fails.
Test pointers: cmd/blockvolume/frontend_capacity_test.go (6 tests:
ProductOfBlocksAndBlockSize, RejectsZero, OverflowGuard,
IscsiHandlerCapacity, NvmeHandlerCapacity, FrontendDefaults_
StillReturn1MiB). Source-side: cmd/blockvolume/main.go::
computeFrontendVolumeSize flows into both iscsi.TargetConfig and
nvme.TargetConfig handler.
First introduced: P15 G5-5C addendum (P0 product fix).
Owner layer: host (binary, frontend wiring).
Last verified: 2026-04-28 (G5-5C addendum P0; m01 hardware re-
verification pending QA).
Status: ACTIVE.
Awaiting m01 hardware re-run for full §close ledger update.
Architect approved Option B 2026-04-27: absorb the hardware-revealed
gap into G5-5C as Batch #7 instead of carrying to G5-5D.
§1.I scope:
- core/host/volume/peer_command_executor.go (NEW, ~120 LOC)
- core/host/volume/peer_adapter_registry.go (NEW, ~100 LOC)
- core/replication/volume.go ConfigurePeerLifecycleHook (~30 LOC)
- core/host/volume/probe_loop_wiring.go router signature (~20 net)
- cmd/blockvolume/main.go registry wire-up (~20 net)
- ~10 new tests, ~250 LOC test code
INV INV-G5-5C-PER-PEER-ADAPTER-PER-PEER-ENGINE absorbed back
in-batch (was previously deferred to G5-5D in pre-architect-ruling
draft).
Pass criterion unchanged: m01 verify_restart_catchup GREEN within
30s deadline; #1-#3 regression GREEN in the same run.
§close updated: ceremony waits for Batch #7 land + hardware re-run;
G5-5C closes at full L4 in one shot.
m01 hardware run 3 at seaweed_block@ac9392d:
- #1 verify_cluster_ready ✅ GREEN
- #2 verify_byte_equal ✅ GREEN
- #3 verify_network_catchup ✅ GREEN (9s)
- #4 verify_restart_catchup ❌ RED (30s timeout)
Root cause (verified in code + log):
Primary log shows probe loop fired correctly post-restart and the
wire probe SUCCEEDED twice (R=2 S=1 H=3), but no StartCatchUp ever
dispatched. Engine apply.go:117-128 checkReplicaID drops events
whose ReplicaID doesn't match the adapter's tracked Identity —
cmd/blockvolume's host adapter tracks the PRIMARY'S OWN slot
(ReplicaID=r1), not peer r2. Probe results for r2 are correctly
dropped as wrong_replica.
Component test (Batch #6) passed because cluster.go's
WithEngineDrivenRecovery constructs c.primary.adapters[] — one
per peer. cmd/blockvolume only constructs ONE adapter for the
host's own slot. The component test exercised a different
(architecturally-correct) wiring than production has.
§1.H audit verdict was correct on engine SEMANTICS; it did not
extend to whether the production binary CONSTRUCTS per-peer engine
state. That layer was assumed; hardware revealed the assumption.
§close decision:
- G5-5C software pieces all sound, stay landed (50 unit + integ
tests PASS; full ./... regression PASS).
- Hardware finding carries to G5-5D — Per-peer adapter wiring for
primary-side recovery dispatch.
- G5-5D pass criterion = exact verify_restart_catchup case from
this run; seed evidence = sw-block/design/g5-artifacts/primary-fail.log.
- New INV to inscribe at G5-5D close:
INV-G5-5D-PER-PEER-ADAPTER-PER-PEER-ENGINE.
Doc updates:
- §close.evidence: hardware-pin row table filled with run 3 results.
- §close.deltas: 3 implicit assumptions surfaced.
- §close.findings: 2 findings (#1 per-peer adapter gap; #2 script
port-release race already fixed).
- §close.forward-carries: G5-5D added as named carry.
- architect-review-checklist: scope/audit/engine-impact/product
level all updated to reflect actual reached state (L3+, not L4).
Awaiting architect ratification of G5-5D scope at single-sign or
earlier; sw drafts G5-5D mini-plan once architect rules.
Per v3-batch-process.md §2: §close drafted as soon as software is
ready. Hardware row table left as TBD; sw fills evidence pointers
once iterate-m01-replicated-write.sh completes. Forward-carries +
deferred ledger pointers + architect-review-checklist all populated
based on G5-5C scope already in-batch.
Awaiting:
1. m01 hardware run completion → fill #1-#4 evidence rows
2. QA evidence verification → §close.deltas / findings if needed
3. architect single-sign per v3-batch-process.md §5 + §8C.2
Per v0.5 §1.H step 3, sw publishes audit findings as a commit note
before any G5-5C production code change.
AUDIT METHOD: greped seaweed_block/core/{engine,replication,adapter}
for the structural backing of each in-scope INV; cited apply.go +
state.go + replication/volume.go + adapter/adapter.go line numbers
as evidence.
PER-INV FINDINGS:
[1] INV-G5-5C-PRIMARY-RECOVERY-AUTHORITY-BOUNDED
Owner: core/replication/volume.go (ReplicationVolume.peers map)
Status: ✅ PASS. peers map is sole probe target collection;
UpdateReplicaSet is sole mutator and is master-fact-driven only.
Halt-cond cleared.
[2] INV-G5-5C-GENERATION-FENCE
Owner: core/engine/apply.go:132-166 (stale event rejection) +
state.go:24-32 (IdentityTruth.{Epoch, EndpointVersion} carrier)
Status: ✅ PASS. Engine rejects events with epoch < Identity.Epoch
or (epoch == AND ev < Identity.EndpointVersion). identityChanged
triggers wholesale Recovery reset (line 166-169). Fence is
carried on engine state, not re-derived per call site.
Halt-cond cleared.
[3] INV-G5-5C-SINGLE-INFLIGHT-PER-PEER
Owner: core/engine/state.go:144-151 (SessionTruth single-slot) +
apply.go phase-guards at 183/236/364/417/442/455/472/507/536
Status: ✅ PASS. ReplicaState.Session is one slot per peer.
Engine FSM handlers explicitly skip / reject when Phase is
PhaseStarting or PhaseRunning. apply.go:536 "Skip if a rebuild
session already exists" pinned. In-flight is engine-explicit,
not implicit. Halt-cond cleared.
[4] INV-G5-5C-PROBE-BEFORE-CATCHUP
Owner: core/engine/state.go:84-121 (RecoveryTruth) +
decide() probe-driven decision path
Status: ✅ PASS. RecoveryTruth.Decision is derived from R/S/H
(boundaries from probe), NOT from transport reachability.
Engine's RebuildPinned guard prevents stale auto-probe from
downgrading Rebuild back to CatchUp mid-flight (line 105-120).
Halt-cond cleared.
[5] INV-G5-5C-RECOVERY-BACKOFF
Owner: engine retry budget (state.go:91-103
RecoveryTruth.Attempts + RuntimePolicy.MaxRetries from T4c-3) +
NEW G5-5C runtime cooldown (5s base → 10s → 20s → 40s → 60s cap;
reset on success)
Status: ⚠ PARTIAL — engine has retry budget but no exponential
cooldown. G5-5C adds the cooldown as a primary-runtime policy on
top of engine retry budget. NOT an engine FSM change. Acceptable
under §1.H "minimum evolution" criterion. Halt-cond cleared.
[6] INV-G5-5C-STALE-ACK-NO-HEALTH-PROMOTION
Owner: core/engine/apply.go:766-789 (Healthy gate)
Status: ✅ PASS. Healthy = true requires three conjuncts:
(a) Recovery.Decision == DecisionNone, (b) Reachability.Status
== ProbeReachable, (c) Identity.Epoch <= Reachability.FencedEpoch.
A barrier ack with AchievedLSN < TargetLSN does not transition
SessionTruth, decide() does not flip Decision to None on
insufficient achieved LSN — Healthy stays false. Halt-cond
cleared.
OVERALL VERDICT: PROCEED.
All six in-scope INVs have their backing infrastructure in engine
(state.go + apply.go) or replication (volume.go). G5-5C is a runtime
wiring batch + small policy extension (backoff). No engine FSM
rewrite needed. No halt-condition fires; no engine-evolution
mini-plan required.
NEXT STEP: implement primary-side probe loop +
ReplicaPeer.ProbeIfDegraded() + lifecycle/cooldown/dispatch tests +
component test, all under core/replication/. Probe loop owned by
ReplicationVolume lifecycle per architect binding. Test method
names to be concretized as code-start commit-note addendum to §2.
This audit commit fulfills §1.H step 3 (audit findings published) +
§2 #15 (audit commit note before production code).
Architect single-signed §1-§6 at seaweedfs@ba7bd0ba4 2026-04-27 with:
- Option B trigger source (primary-side degraded-peer probe loop)
- Probe loop placement = core/replication/ owned by ReplicationVolume
- Master protocol unchanged
- §1.H code-start audit gate before code
This commit:
1. Records the single-sign in the doc header.
2. Adds a §1 scope-rule one-liner near the top so future readers find
the architect-bound boundary without re-reading the v0.1→v0.5 trail:
"master owns identity/topology; primary+engine own data recovery;
the protocol aligns the two via (PeerSetGeneration, epoch,
EndpointVersion) fences."
§1.A already bound Option B in v0.4; no flip needed there. No design
change. §1.H audit is the next sw step before any production code.
Architect framing 2026-04-27: enumerate ten protocol boundary rules
and address engine-evolution question.
Engine vs primary runtime vs master split:
- Engine owns: recovery FSM, single in-flight per peer,
generation/epoch fence, probe→decision, backoff/cooldown policy,
stale-ack-cannot-promote-health rule, recovery reason / projection
- Primary runtime/adapter owns: timer / degraded-peer loop, transport
probe execution, feeding probe result into engine, executing
engine-emitted commands, ReplicationVolume / ReplicaPeer connection
lifecycle
- Master owns: identity / topology / assignment / health observation
ONLY. No runtime recovery. No epoch bumps for short up/down.
Six in-scope boundary rules (#1, #2, #3, #4, #7, #8):
- #1 Admitted Peer Rule — already INV-G5-5C-PRIMARY-RECOVERY-AUTHORITY-BOUNDED
- #2 Generation Fence — NEW INV-G5-5C-GENERATION-FENCE
- #3 Single In-Flight Per Peer — NEW INV-G5-5C-SINGLE-INFLIGHT-PER-PEER
- #4 Probe Before Catch-Up — NEW INV-G5-5C-PROBE-BEFORE-CATCHUP
- #7 Backoff/Cooldown — NEW INV-G5-5C-RECOVERY-BACKOFF (extends v0.4
fixed-5s into 5s→10s→20s→40s→60s cap, reset on success)
- #8 Stale Ack Guard — NEW INV-G5-5C-STALE-ACK-NO-HEALTH-PROMOTION
(cross-refs G5-5 round-14 gate-degraded artifact)
Three forward-carries OUT of G5-5C (per §5):
- #5 Durability Mode Explicit → G5-2 / G5-6
- #6 RF Health Reporting Separate From Recovery → future master
observability batch
- #10 Status Surface (recovery reason, effective RF, last probe) →
G5-3 metrics/backpressure
One citation (#9 Replica-side lineage check): already enforced by T4
acceptMutationLineage gate; G5-5C cites, no new code.
§1.H code-start audit gate: sw audits per-INV current owner location
BEFORE writing any code. Halt-condition: if recovery FSM is embedded
in ReplicationVolume, fence is re-derived per call site, in-flight is
implicit, or stale-ack guard is missing — sw stops and re-scopes as
engine-evolution batch instead of layering ifs in core/replication/.
Audit findings published as commit note pre-code; PR includes
audit-summary.
§2 acceptance criteria: add #13 (stale-ack guard), #14 (backoff
progression), #15 (code-start audit). Acceptance count now 15
covering 7 INVs (6 new + reconnect orthogonality from v0.4.3).
Standing by for architect single-sign of v0.4.4.
Architect framing 2026-04-27 (sharpening v0.4.2): reconnect splits
along two orthogonal dimensions — connection recovery vs identity /
lineage change. Each axis has different protocol semantics; G5-5C
must handle both correctly.
Architect's protocol judgment points:
1. PeerSetGeneration only changes for identity / address / lineage
change. Brief disconnects / restarts / freshness flapping do NOT
bump generation.
2. Primary's degraded-peer loop only acts on currently-admitted peers
(§1.E reaffirmed).
3. After reconnect, primary still probes R/S/H — reconnect alone is
not assumed sufficient.
4. If a higher PeerSetGeneration arrives during reconnect / probe,
the in-flight recovery must stop or invalidate.
Changes:
- New §1.F with two cases:
Case 1 (identity unchanged): primary retries existing peer
descriptor; new sessionID minted (sessions are session-scoped, not
peer-scoped); probe R/S/H; catch-up / rebuild as needed; no master
re-emit needed. This is G5-5C's core path.
Case 2 (identity changed): existing UpdateReplicaSet T4a-5 path
(volume.go:229-246) tears down + recreates; in-flight aborts via
Close(); new peer with new lineage takes over.
- Misread guards documented: "primary keeps retrying old address
forever" rejected by Case 2 + §1.E (c); "master must bump on every
blip" rejected by Case 1 + §1.D.
- New INV-G5-5C-RECONNECT-ORTHOGONAL-AXES in §3.
- New §2 #11 (reconnect Case 1 — identity unchanged, no re-emit) and
§2 #12 (reconnect Case 2 — lineage bump mid-flight).
This is structural reaffirmation: the V3 code already does Case 2
correctly (T4a-5 teardown). Case 1 is what the probe loop adds. The
new tests pin both axes against future drift.
Standing by for architect single-sign of v0.4.3.
Architect framing 2026-04-27 (sharpening v0.4.1): §1.D ordering-
independence must NOT be misread as "primary may self-discover and
connect to any replica it sees on the network." Tighten with a
second protocol invariant.
Rule (architect verbatim): "Primary recovery loop may retry only peers
that were previously admitted by a master-issued assignment fact for
the current authority lineage."
Layering: master establishes identity once; primary owns retry /
recovery for that admitted peer until master revokes or changes the
assignment.
This is structurally true in V3 today (probe loop reads
ReplicationVolume.peers, which UpdateReplicaSet populates from master
facts) but v0.4.2 promotes it from implementation detail to protocol
invariant so future contributors don't widen the probe surface.
Changes:
- New §1.E with three scenarios:
(a) first-time replica join — disallowed without master fact
(b) brief outage + recovery (G5-5C core case) — allowed without
master re-emit
(c) epoch / assignment change — probe must stop; in-flight aborts
- Implementation requirement made explicit: ReplicaPeer.Close() must
abort in-flight probe synchronously.
- Authority alignment surface table: replicaID/epoch/EV → identity;
AssignmentFact.Peers → only legal probe targets;
PeerSetGeneration → existing lastAppliedGeneration guard preserved.
- New INV-G5-5C-PRIMARY-RECOVERY-AUTHORITY-BOUNDED in §3.
- New §2 #9 (authority-bounded targets test) and §2 #10 (lineage-
change-during-probe test).
- §1.A bound-shape Master-interaction row references §1.E.
- §1 Files peer.go row notes Close() must abort in-flight probe.
Standing by for architect single-sign of v0.4.2.
Architect framing 2026-04-27: when a replica goes down or recovers,
both the control-plane identity/health loop and the data-plane
governance loop receive feedback. Protocol must treat them as two
independent loops with no ordering dependency, alignment via durable
identity facts (replicaID/epoch/EV/peer address), and idempotency on
primary-side dispatch absorbing duplicate triggers.
This is a sharpening of v0.4, not a re-bind. Design unchanged:
Option B primary-side probe loop, no master protocol change.
Changes:
- New §1.D: explicit two-loop table, five ordering scenarios all
ending safe, anti-requirements (master re-emit NOT prerequisite,
primary recovery NOT blocked on master), idempotency guarantees,
future RF-health observability noted as different-batch scope.
- New INV-G5-5C-TWO-LOOPS-ORDERING-INDEPENDENT in §3 with test
pointer (peer_test.go simultaneous-fire test).
- New §2 #8 acceptance criterion: unit test exercising the
"simultaneous-fire" case (concurrent fact replay + concurrent
ProbeIfDegraded on same degraded peer; idempotent absorption).
Standing by for architect single-sign of v0.4.1.
Architect REVISE ruling 2026-04-27: bind trigger source to Option A
with both halves in scope (no split into G5-5B). Reject B and C.
QA review v0.1 flagged: master-side scope must be explicit; pin §5
evidence path.
Changes:
- §1.A: collapse three-option proposal to bound Option A. Make A1
(master-side observation-driven re-emission) and A2 (primary-side
recovery dispatch) explicit as two halves of one causal chain.
Record B/C rejection rationale for future reference.
- §1 Files: revise table with Side column (master/primary). Add
master-side rows (A1 re-emit logic + ObservationStore freshness
helper). Total estimate ~360 prod + ~150 test, split master ~90 /
primary ~120 / tests ~150.
- §2: rewrite criteria #1-#5 around bound Option A (drop per-Option
deadline language). Split #2/#3 into A1 master-side + A2
primary-side criteria. Hardware deadline at #5 stays 30s.
- §2 verifier note: file paths + test names pinned at code-start
(acceptable for v0.2 per QA review).
- §5: pin G5-5 seed evidence to actual artifact path
V:\share\g5-test\logs\artifacts-20260427T092858Z\primary-fail.log
(no future task — fact-pointer).
- §7: trigger-source binding row marked done (architect REVISE);
single-sign of v0.2 still pending.
- Header: v0.1 → v0.2 status note updated.
Standing by for architect single-sign of v0.2.
No code starts until single-sign.
Per architect single-sign of G5-5 §close (`seaweedfs@c78116fd2`):
(a) v3-dev-roadmap.md
- §3: G5 line note now mentions G5-5 closed at L3 + G5-5C carry-forward
- §4: G5-5 row → CLOSED (link to seaweedfs@c78116fd2); G5-5C row added
as next active gate with bound pass criterion
- §7: G5-5 close commit appended to recently-closed table
(seaweed_block@5c4718f + seaweedfs@c78116fd2, L3 reached, #4 carry)
(b) v3-phase-15-g5-5c-mini-plan.md (new) v0.1 kickoff
- §1 scope: peer recovery trigger after replica restart; reuse T4d-4
primitives (architect binding); no engine logic change
- §1.A: three trigger source options (A master observation, B periodic
probe, C transport reconnect) with tradeoffs; sw recommends A; final
pick deferred to architect ratification
- §2: 6 acceptance criteria, hardware step is exactly G5-5 #4
(verify_restart_catchup → GREEN with no harness changes)
- §3: 2 new INVs proposed (REPL-PEER-RECOVERY-TRIGGER-001 +
-NO-RETRIGGER-LOOP) + 2 deferred ledger pointers from G5-5 close
- §4: G-1 N/A (new build, no V2 PORT)
- §5: forward-carries from G5-5 §close all addressed
- §6: 5 risks tabled
- §7: sign table awaiting architect §1-§6 ratification including
trigger source pick
Standing by for architect ratification of trigger source binding.
No code starts until §1-§6 signed.
Architect's round-15 hygiene callout: §7 sign table still had
three pre-code 'blocked' rows after the real close-state rows
landed in the prior doc-fix commit. Pure leftover from before
the close-state update overwrote earlier rows but didn't delete
the trailing pre-code rows.
Removed:
- 'Code start (script + Go helper) ... blocked on ratification'
- 'm01 hardware verification run ... blocked'
- '§close append + close sign ... blocked'
Sign table now ends cleanly at the §close architect single-sign
pending row. Ready for sign.
§close summary per v3-batch-process.md §12 template:
Done:
- #1 verify_cluster_ready
- #2 verify_byte_equal — live iSCSI replicated write, byte-equal
verified via storage-aware m01verify (LBA[0]=0xab on cross-host
hardware)
- #3 verify_network_catchup — iptables disconnect+heal, replica
converges to LBA[1]=0xcd byte-equal in 8s via engine-driven
catch-up
- 14 bugs surfaced+fixed across 14 m01 self-iteration rounds
- 5 INV-BIN-WIRING-* invariants in v3-invariant-ledger.md from
G5-4 still load-bearing; G5-5 hardware run is Integration backstop
Not done:
- #4 verify_restart_catchup — kill replica + write while down +
restart: replica's LBA[2]=0xef does NOT converge in 30s. Per
architect ruling 2 (round 13): real recovery-path finding,
surface as G5-5C carry-forward.
- #5 verify_race_stress + #6 verify_full_suite — gated on #4 fix
or test sequencing rework.
Product level reached: L3 (Replicated IO) per v3-architecture.md §13.
Falls short of full L4 (Failure/recovery under IO) — process-restart
recovery is the gap, scoped as G5-5C.
Next gate that makes it usable: G5-5C Peer Recovery Trigger After
Replica Restart — fix engine-driven catch-up re-trigger when a
degraded peer becomes reachable again. After G5-5C: re-run #4#5#6
in this same harness; full L4 reached.
Forward-carries to G5-5C (architect-bound 2026-04-27):
- Reuse existing engine-driven recovery primitives (T4d-4); no
ad-hoc re-ship from replication layer.
- Define trigger source first: observation reappearance, periodic
probe loop, or stream/transport reconnect signal.
- Pass criterion: exactly the failed hardware case from G5-5 #4.
- Seed evidence: seaweed_block@5c4718f primary-fail.log shows the
gate-degraded + stale-barrier-ack pattern.
Forward-carries to opportunistic future hardening:
- Unit test for EnsureStorage→assignment-arrives→first-Open
Identity-latch path (would have caught round-10/11 bug pre-m01).
- Generalize start_cluster() pre-flight stale-state cleanup pattern
for future hardware harnesses.
Forward-carries to G5-6:
- G5-DECISION-001 (Path A vs Path B) — engine-state serializability
pinned in T4d still holds; G5-5 doesn't change posture.
Pending: architect single-sign on §close per v3-batch-process.md §5
+ §8C.2.
Refs: 24 commits in seaweed_block@phase-15 spanning rounds 1-14
(documented in §close.evidence.commits table).
Architect's v0.2 review caught that §1 absorbed the 3 binding revisions
but §2 (the close contract per v3-batch-process.md §2) stayed stale:
- §2 #2 still said "byte-equal on replica's walstore extent"
- §2 had old #4 (race stress) instead of new #4 (process restart)
- §2 #3 didn't name /status/recovery as the R/H source
v0.3 rewrites §2 to match §1, with explicit verifier names:
#1 verify_cluster_ready
#2 verify_byte_equal — m01verify Go helper using walstore.OpenReadOnly
+ storage.LogicalStorage.Read(lba) + SHA-256 (NO raw extent peek)
#3 verify_network_catchup — iptables disconnect + polls
/status/recovery?volume=v1 for R/H; asserts RecoveryDecision="catch_up"
#4 verify_restart_catchup — SIGTERM replica + restart same binary +
same --durable-root; polls /status/recovery same as #3#5 verify_race_stress — 10x -race on G5-4.5 integration test
#6 verify_full_suite — go test ./... clean from m01
#7 v3-dev-roadmap.md updated at gate-close per v3-batch-process.md §8
§1 file map and §5 forward-carry table already match v0.3 numbering
(grep confirmed no stale references). Implementation scope unchanged
from v0.2 (~310 prod LOC + ~30 unit tests).
v3-batch-process.md §2 single-source-of-truth discipline preserved:
§2 acceptance criteria IS the close contract; §1 scope description
stays in sync but is not load-bearing for close evidence.
Addresses 3 architect revision requirements (round 51):
REVISION 1 — process restart distinct from network disconnect:
Split G5-4 #4 forward-carry into TWO scenarios:
§2 #3 network disconnect (iptables) — proves live TCP interrupt
+ recovery without process restart
§2 #4 replica process stop/restart — proves durable reopen +
master resubscribe + recovery reconstruction
G5-4 #4 is now FULLY consumed (was: only network proxy in v0.1).
REVISION 2 — storage-aware byte verifier:
Replace raw walstore .extent peek with storage-abstraction Read(lba):
helper opens replica's walstore via core/storage/walstore (or
equivalent OpenReadOnly path), invokes Read(lba) per LBA in the
range, SHA-256 vs primary's known payload. Raw extent peek
REJECTED — walstore on-disk includes WAL frames + checkpoints
+ sparse regions + potentially-stale-but-valid blocks; only
Read(lba) returns the authoritative current value.
Risk added: if walstore.OpenReadOnly is missing, sw adds it as
part of this batch (small scope expansion contained in
core/storage/walstore; read-only opener for verification only,
NOT a substrate semantic change).
REVISION 3 — named R/H observation source:
/status?volume=v1 returns frontend.Projection (no R/S/H). G5-5
adds /status/recovery?volume=v1 returning engine.ReplicaProjection
(Mode, R, S, H, RecoveryDecision); gated by new --status-recovery
daemon flag (default off; production binaries don't enable).
Loopback-only via existing isLoopbackRemote guard. ~30 prod LOC
+ ~30 unit tests. Engine/adapter logic unchanged — surfaces
already-computed projection through HTTP.
Updated §1 file map, §1.4 truth-domain check, §5 forward-carry
table, §6 risks (3 new rows), §close template unchanged.
Re-submitted for architect §1-§6 ratification. After ratify, sw
codes per §1 file map; estimate ~310 prod LOC + ~30 unit tests.
Codifies the lessons from T4 + G5-4 retrospective:
KEEP — earned its keep on T4:
- G-1 V2 PORT read (saved 5 hidden invariants on T4b-4, probe non-
mutation pin on T4c-2, 3 placement decisions on T4d-3)
- Mini-plan acceptance criteria (single source of truth for close)
- Invariant ledger discipline ("claim without test = wish")
- m01 -race verification (caught 2 engine bugs at T4d-4 part C)
- Architect single-sign at close (caught 4 stale refs at G5-4 close)
DROP — overhead without payoff:
- Separate kickoff PROPOSAL doc (mini-plan §1-§6 = same thing)
- Separate G-1 doc (inline §4 of mini-plan)
- Separate closure report doc (§close section of mini-plan)
- Separate forward-carry checklist (§5 of next-batch mini-plan)
- Separate QA scenario catalogue (write tests directly when ready)
- Multi-version doc churn (v0.1→v0.5)
- Cross-doc invariant restatement (ledger is sole source)
- Mixed T-track + G-N naming for same gate
Compressed sign cycles: was 4-5 architect signs per batch; now 2
(scope ratify + close sign).
Per-batch artifact count: was 5+ (kickoff + mini-plan + G-1 +
closure + checklist + scenario catalogue); now 1 (mini-plan with
§close appended).
Decision rules codified:
§6.1 G-1 yes/no (V2 PORT yes; V3-native no)
§6.2 T-track vs G-N naming (architect picks at kickoff)
§6.3 When to skip mini-plan (1-line hotfix-class)
§8 names the 6 first-order control docs to keep current
(v3-dev-roadmap, v3-phase-15-mvp-scope-gates, v3-invariant-ledger,
v3-block-behavior-contract-index, v3-product-placement-authority-
rationale, v2-v3-contract-bridge-catalogue).
§9 first trial: G5-5 m01 hardware first-light.
§11 honesty principle: documentation that catches bugs is
discipline; documentation that doesn't is ceremony. Drop ceremony,
keep discipline.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
PR-atomic with seaweed_block@c820e17 per architect binding round 50
(mini-plan v0.4 §4 #7): ledger inscription required at G5-4 close.
- INV-BIN-WIRING-ROLE-FROM-ASSIGNMENT
- INV-BIN-WIRING-PEER-SET-FROM-ASSIGNMENT-FACT
- INV-BIN-WIRING-LISTENER-LIFECYCLE-LIFO
- INV-BIN-WIRING-ASSIGNMENT-DRIVES-MEMBERPRESENT
- INV-BIN-WIRING-SESSIONID-VIA-ADAPTER
All 5 are ACTIVE with test pointers to
cmd/blockvolume/g5_4_l2_replication_test.go (subprocess integration)
+ source-side checks in cmd/blockvolume/main.go and core/host/volume.
Last verified 2026-04-26 (G5-4 close).
Architect ratification 2026-04-26:
"Role inference, in-process acceptance, G5-DECISION-001 seam, and
sessionID discipline are architecturally correct. G-1 must clarify
replica readiness semantics and confirm ctrl-addr reuse or introduce
repl-addr before code."
2 binding clarifications baked into v0.3:
#1 — §4 #2 acceptance criterion split by role:
- Primary: Healthy=true per existing frontend/write-ready projection
- Replica: replication-ready / listener-bound + ApplyEntry byte-equal
verified — MUST NOT report Healthy=true if existing field implies
frontend-primary-write-ready
- If existing status field is too coarse, G5-4.5 uses precise
assertion names (assertReplicaReplicationReady,
assertPrimaryFrontendReady) instead of unified assertHealthy
#2 — §4 #7 acceptance criterion strengthened:
- Catalogue inscription ALONE insufficient at G5-4 close
- 5 INV-BIN-WIRING-* invariants MUST land in v3-invariant-ledger.md
- Per v3-quality-system.md §6 "an invariant without a test is a wish"
- Ledger updated as PR atomic with code (not after-the-fact)
§7.1 G-1 deliverable extended (G-1-blocking subitems):
- Replica readiness semantics — what existing volume.Status /
ProjectionView field expresses replication-ready (vs Healthy)?
G-1 either proposes new field OR specifies precise assertion names
- --ctrl-addr reuse confirmation — verify NO conflict with NVMe/iSCSI
control-plane traffic on same port. If conflict, G-1 introduces
--repl-addr flag (small scope expansion, contained in this batch)
§3 #4 predicate flipped to ✅ DONE (architect round 50).
§8 sign table updated with explicit ledger requirement at close.
Architect-pre-baked: ratification stays valid; no further mini-plan
revisions needed before G-1.
Sw next: produce G-1 V3-native PORT read deliverable per §7.1
(includes the 2 binding subitems). Code stays blocked until
architect ratifies G-1.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Addresses QA's 3 notes + 1 clarification ask:
Note 1 (role inference):
§1.3 rewritten — fact.ReplicaID is master-minted (proto verified
at control.proto:128-148 + mint site at services.go:198-205).
Binary reads `fact.ReplicaID == self.ReplicaID` directly. No
lex-smallest fallback (master always names exactly one bound
replica per volume per line). Removes the binary-side authority
inference that violated the master-authority rule.
Note 2 (acceptance circular):
§4 #2 verifier reframed to G5-4.5 in-process test. m01 hardware
verification belongs to G5-5; G5-4 closes on the in-process pin.
Note 3 (G5-DECISION-001 contradiction):
§5 rewritten — G5-4 ships Path B runtime AND keeps Path A
serializability seam open. T4d-4 part B's RoundTripJSON test
already pins serializability; G5-4 preserves it. G5-6 architect
ratification can promote to Path A by adding persistence on top
of the existing struct, with no engine-state-shape change.
Clarification ask (sessionID minting):
§6 added INV-BIN-WIRING-SESSIONID-VIA-ADAPTER. Adapter mints
unique sessionIDs via process-wide atomic counter at
adapter.go:70; binary inherits for free as long as it dispatches
via the adapter (never via framework shortcuts that hardcode
sessionID=1, which is the known T4c §I + QA G5-1 round 1 SKIP
gap). Pinning this invariant keeps the gap test-side.
Re-submitted for QA re-review per parent kickoff §7 governance loop.
Mirror cmd/blockvolume to T4d-4 part B's WithEngineDrivenRecovery()
framework binding. Single batch (~250 prod + ~150 tests), 5 ordered
subtasks. Design decisions (a-d per kickoff §3 G5-4 row):
(a) Role inference: assignment-driven, no new CLI flag
(b) Peer discovery: AssignmentFact.Peers per T4a-5 P-refined
(c) Listener lifecycle: --ctrl-addr reuse + LIFO Stop in host.Close()
(d) Engine instantiation: one engine per volume, single --volume-id
Pre-merge gates require G-1 V3-native PORT read of cluster.go:357-369
+ V2 lesson check on weed/storage/blockvol/blockvol.go before code.
4 new invariants to inscribe at close (INV-BIN-WIRING-*).
Submitted for QA + architect ratification per parent kickoff §7
governance loop. No code until ratify.
Hand-off doc v0.3 + G5 kickoff v0.2: m01+M02 bring-up smoke surfaced
that cmd/blockvolume binary lacks T4 replication wiring entirely.
Sw-confirmed root cause:
- --t1-readiness HealthyPathExecutor is primary-only by design
- volume.Config.ReplicationVolume slot exists (host.go:73) with godoc
"T4a-5 production wiring sets this" — but T4a-5 only added the
field; the wiring NEVER landed
- T4d-4 part B wired WithEngineDrivenRecovery() for component test
framework (cluster.go:357-369), NOT for the binary
- Result: V3 components compose end-to-end (proven by T4d HARD GATE
#3); the production binary still constructs a primary-only data
plane
Sw confirmed this is real implementation work (150-300 LOC + design),
not a 50-LOC quick patch. Four design decisions needed:
1. Role inference (assignment vs CLI flag vs topology)
2. Peer discovery (from AssignmentFact.Peers)
3. Listener lifecycle (--data-addr reuse + Stop)
4. Engine instantiation (one engine per volume)
G5 kickoff revised to v0.2:
- 5 batches → 6 batches (binary wiring promoted to G5-4)
- G5-1/2/3 are NOT blocked by G5-4 (component framework already
binds T4d-4 part B; QA scenarios + walstore cadence at
component/primary-only scope can run in parallel)
- G5-4 binary wiring: needs full governance loop (kickoff →
architect ratify → mini-plan → architect ratify → G-1 → code).
G-1 source: T4d-4 part B component framework as V3-native PORT
- G5-5 m01 hardware first-light DEPENDS on G5-4 (script can't
drive replica scenarios until binary supports replicas)
- G5-6 G5-DECISION-001 resolution at G5 close (was G5-5 in v0.1)
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Root cause for "volume not ready" gate: missing
--expected-slots-per-volume 2 flag on blockmaster.
Default is 3; QA's 2-node topology had 2 slots; controller
silently rejected observation snapshot (cmd/blockmaster/main.go:39).
Fix verified locally on Windows (single-node, no m01/M02 needed):
- Add --expected-slots-per-volume 2 to blockmaster command
- Primary reaches Healthy=true with epoch=1
- assignment-received fires; durable storage opens; status
endpoint serves {"Healthy":true}
Lesson learned (process improvement): for V3-internal bring-up
debug, try single-node local reproduction FIRST. The cluster
bring-up gate is V3 logic, not network topology. Reproduces in
seconds locally with full source-code access; m01/M02 only needed
for cross-node-specific scenarios (real network conditions,
iptables, multi-host wire).
Secondary finding: replica r2 sees primary r1's assignment but
records "supersede, not applying to adapter" because T1
HealthyPathExecutor only handles primary case. For G5-4 replica
bring-up, sw needs to wire T4a-T4d ReplicationVolume + ReplicaPeer
+ ReplicaListener stack (not just --t1-readiness flag). This is
the actual next gap for G5-4.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Root cause: cmd/blockmaster/main.go hardcoded ExpectedSlotsPerVolume=3.
QA's 2-slot topology silently failed validateVolumeTopology in the
controller, so no assignments were minted, no master-log lines,
and volumes timed out at durable open.
Fix landed in seaweed_block@f5de7c5: --expected-slots-per-volume
CLI flag, default 3, set 2 for the 2-node smoke.
QA next: rebuild blockmaster, pass --expected-slots-per-volume 2
in §3.4 of the handoff command sequence; rest unchanged.
Records QA's cross-node smoke attempt 2026-04-26: infrastructure
fully verified READY (m01+M02 reachability, SMB share for binary
distribution, master cross-node listen, network OK), but cluster
bring-up blocked at V3-internal gate.
Symptom: blockvolume on both nodes connects to master but logs
"durable open: frontend: volume not ready" — never reaches steady
state, status endpoint never binds, master log shows no heartbeat
or assignment-mint events.
Hand-off contents:
- §1 specific questions for sw (5 gaps to fill)
- §2 infrastructure verified READY (no action needed)
- §3 copy-pasteable commands sw can run/debug
(build → topology → master → primary → replica → cleanup)
- §4 QA's hypothesis on the gap (assignment-from-master flow)
- §5 debug suggestions for sw (log levels, integration test
references)
- §6 G5-4 script skeleton current state
- §7 QA's next steps once sw answers
Working dirs reproducible:
- Binaries: /mnt/smb/work/share/g5-binaries/{blockmaster,blockvolume}
- Run state: /tmp/g5sm/ on both nodes
- Logs: /tmp/g5sm/logs/{master,primary,replica}.log
Blocks: G5-4 implementation work (script scenario bodies, hardware
first-light scenarios). Does NOT block QA scenario authoring at
component scope (Cluster framework already covers that).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Per QA infra-check round 2026-04-26, surfaces real readiness gaps
before architect ratifies G5-4 schedule:
m01 (192.168.1.181 — primary node):
✅ 32-day uptime; sudo password-less; 16 cores; 19 GiB RAM
✅ 177 GiB free disk; Go 1.26.2 installed
✅ iptables / netns / multi-process tools all available
✅ T2 m01 NVMe script template available as pattern reference
M02 (192.168.1.184 — replica node):
✅ Reachable from m01 (0.92ms); same kernel; 178 GiB free disk
❌ Go NOT installed — must scp binaries from m01
Implication for G5-4:
Build binaries on m01, scp to M02. Same cross-node binary pattern
T2 already uses for its iSCSI target deployment. G5-4 skeleton at
seaweed_block/scripts/iterate-m01-replicated-write.sh implements
this build-then-scp flow.
No infrastructure blockers. Architecture ready as soon as G5 mini-plan
ratifies scenario list.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Two artifacts landing together to close T4 batch series:
1. v3-phase-15-t4d-closure-report.md (NEW)
QA single-sign artifact for T4d batch close per §8C.2; architect
T-end three-sign per §8C.1 (T4d IS final T4 batch — confirmed at
round-48 review). Round-48 + round-49 corrections incorporated:
- Part C commit hash bound to e642ae8 throughout
- CARRY-T4D-LANE-CONTEXT-001 bind point = post-G5 hardening
backlog (not T4e — consistent with "T-end at this close")
- §H Finding #1 reworded — walstore HAS background flusher
(walstore.go:189-190); QA's earlier "caller-driven" was wrong
- §H Finding #3 RESOLVED at a0be6d5 (T2A NVMe race fixed +
m01 -race ×50 PASS)
- 16 invariants pinned (added 2 named for part C bug fixes:
INV-REPL-FAILED-SESSION-KIND-DRIVES-ESCALATION +
INV-REPL-REBUILD-ESCALATION-STICKY-UNTIL-TERMINAL)
- 22/22 packages green under -race on m01 (post-a0be6d5)
2. v3-phase-15-t4d-mini-plan.md (NEW — was uncommitted across
v0.1 → v0.5 evolution)
Final v0.5 incorporates: architect Path B fold; round-47
rebuild path engine-driven HARD GATE expansion; G5-DECISION-001
named decision record; 4-batch shape ratified; T4d-3 G-1 binding.
Active forward-carries (post-G5 hardening backlog):
- CARRY-T4D-LANE-CONTEXT-001 — replace TargetLSN==1 caller shim
with true handler/session-context lane signal
- G5-DECISION-001 — engine recovery state behavior across
primary restart (Path A persist vs Path B rebuild-from-probe)
G5 collective close items (NOT post-G5):
- m01 hardware first-light for replicated write path
- Multi-replica concurrent live + recovery scenarios
- walstore flusher cadence verification + tuning policy
- Minimal metrics/backpressure assessment
- G5-DECISION-001 architect resolution
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Architect sign by pingqiu 2026-04-25:
"T4d v0.2 scope accepted as one batch series; Option C for appliedLSN
source; BlockStore walHead hotfix may land pre-T4d; substrate defense-
in-depth included where practical; 4-batch order approved; T4d-3 G-1
required; T4d-2 no G-1; T-end three-sign at T4d close if T4d remains
final T4 batch."
All open architect-decision points (§2 scope, §2.5 Option/hotfix/
substrate, §3 batch shape, §4 acceptance bar) resolved. §6 open
issues all closed. §8 inscribes the verbatim ratification record.
Sw clearances effective immediately:
- Land BlockStore walHead one-liner as pre-T4d hotfix (single PR with
un-skipped regression test)
- Produce T4d mini-plan (4-batch shape per §3)
- Produce T4d-3 G-1 V2 read on wal_shipper.go runCatchUpTo
- T4d-2 spec is round-43/44 architect text (no G-1 needed)
T-end horizon: §8C.1 T-end three-sign lands at T4d close IF T4d
remains final T4 batch (per architect's criterion #10 wording tweak).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
QA single-sign artifact for T4c batch close per §8C.2; architect
acceptance of §B scope deltas signed 2026-04-25 by pingqiu.
Scope deltas accepted:
- T4c closes as mid-T4 batch under §8C.2, not T4 T-end
- L2/L3 mini-plan bar narrowed to muscle-level L2 + component evidence
- L3 m01 first-light deferred to T4d / G5 final close
- Substring "WAL recycled" matching accepted as TEMPORARY, replacement
bound to T4d (preferred) or G5 final sign (latest)
- INV-REPL-CATCHUP-WITHIN-RETENTION-001 downgraded to T4d blocker
(catch-up sender hardcodes ScanLBAs(1); replica's R+1 not threaded)
Doc-hygiene fixes per PM round-2 review (this commit):
- Drop INV-REPL-CATCHUP-DONE-MARKER-EMITTED (non-existent: V2 marker
collapsed into barrier-as-terminator per catchup_sender.go:48,187)
- §B/#2 + #5 reword "green at HEAD" to acknowledge architect Windows
cleanup-only repro failures (tracked as next-batch carry)
- Active formal-INV count 8 -> 6
Forward-carries to T4d (BLOCKERS):
- R+1 catch-up threading (StartCatchUp signature + adapter wire)
- Full engine→adapter→executor recovery wiring
- Structured RecoveryFailureKind replacing substring sentinel
- LastSentMonotonic_AcrossRetries cross-call form scenario
- Windows TempDir cleanup race investigation
Forward-carry to G5 final close:
- m01 hardware first-light for replicated write path
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Closes QA round-2 feedback loop. Three concerns resolved and one L2-blocker
hazard added.
## Q1-Q3 resolution (sw-verifiable per QA concern; V2 source check)
Q1 scope completeness: VERIFIED complete. V2 grep shows sync_all_* are
three test files only — `sync_all_adversarial_test.go`, `sync_all_bug_test.go`,
`sync_all_protocol_test.go`. Zero production files for sync_all / split_brain /
takeover / arbiter. These are cross-entity invariants, not distinct types.
10-entity set stands.
Q2 ReplicaReceiver scope: VERIFIED per-volume, not per-assignment.
`v.replRecv = recv` at `blockvol.go:1515` is the only write site; zero
`replRecv = nil` assignments in codebase. Receiver is constructed-once per
BlockVol instance. L1 §2.3 wording stands.
Q3 RebuildSession/Bitmap durability: VERIFIED no sidecar. Grep
`rebuild_bitmap.go` + `rebuild_session.go` for `os.Open / os.Create /
WriteFile / ReadFile / persist / sidecar` → empty. Recovery is WAL
hydration only (`hydrateBitmapFromRecoveredWAL` at `rebuild_session.go:102`).
L1 §2.10 invariant #3 CORRECTED — earlier draft incorrectly called out a
"sidecar schema" that doesn't exist.
## QA concern #3 resolution: §3.14 new hazard
`AllBlocks()` semantic divergence: V3 `walstore.go:565` and
`smartwal/store.go:367` both call `s.Read(lba)` which reads through the
dirty map (includes unflushed WAL bytes). V2 `rebuild.go:handleExtentStream`
uses `readBlockFromExtent` which BYPASSES dirty map (flushed-only).
Concrete impact: V3 base stream can contain bytes the primary hasn't fsynced.
If primary crashes pre-fsync, replica's copy is "newer" than primary's
recovered state. Epoch fencing + WAL-wins bitmap still prevent corruption,
but the invariant chain is "eventually consistent via epoch churn" instead
of V2's "base stream never contains unflushed bytes". Different contracts,
same end state.
Two L2 options proposed: (a) keep AllBlocks semantics + document non-claim
in §2.7 bridge; (b) add `LogicalStorage.AllBlocksFlushed()` preserving V2
invariant. H5 architect-line decision affects which path is safer.
## QA concern #2 resolution: §3.a locked-pairs section (new)
Documents pre-coupled L2 decisions driven by V3 existing shape:
H6 Option C → H7b locks automatically (Provider intercepts at LogicalStorage
layer; Backend.Write stays host-facing, doesn't carry LSN)
§3.14 + H5 → AllBlocks safety rationale depends on which H5 shape wins
Per BUG-005 documentation-discipline lesson: record coupled pairs explicitly
rather than leaving them as "implied". Saves L2 cycles and gives future
readers visible intent for why Backend.Write excludes LSN.
## QA concern #1 deferred to L2
Volumes map extension (single-map with role discrimination vs two separate
primaryHandles + replicaHandles maps) is a legitimate L2 design concern.
L1 appropriately hedges with "likely needs to grow" (§3.11 Option C); L2
picks shape. QA's BUG-005-adjacent concern (role-discriminated handle
callers forgetting to check role) is the right frame for the L2 decision.
No L1 edit needed; flagged for L2 attention.
## §4 open questions status
Q1-Q3 ✓ resolved
Q4 DistGroupCommit residence → effectively answered by §3.11 C
Q5 protocol-frame wire-compat stance → still architect-line (pairs with H5)
Blocking L2 start now: only H5 + Q5, both architect-line. QA to draft
one-page arch memo per round-2 offer.
## Change log
§5 feedback-round log gains round-3 entry
§6 change log gains full round-3 detail with V2 line citations
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
§8C.8 specifies exactly one three-sign per T-boundary — at T-start,
covering the bundled L1+L2+L3 package. I had proposed a separate
L1 three-sign in §5 that isn't in the rule. Architect correctly
pushed back.
§5 rewritten as lightweight cadence:
1. sw V3 pre-scan (~5 min, inline reply, prerequisite to L2 not a
sign gate) — same grep checklist retained, same BUG-005 rationale
2. sw + QA iterate on L2 (catalogue §3 filled) informally
3. sw + QA draft L3 (T4 port plan sketch)
4. T4 T-start three-sign on bundled L1+L2+L3 (only governance event)
Informal feedback-round log hook added so architect/PM inputs are
tracked without per-round sign ceremony.
Change log updated.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
All 5 feedback items accepted; no subsetting.
F1 — RebuildBitmap split into standalone §2.10 entity (10 total,
was 9). Rationale: bitmap has independent on-disk schema (~84 LOC
rebuild_bitmap.go) + independent conflict-resolution invariant
(WAL-wins-over-base). Collapsing into §2.6 RebuildSession at L1
would lose granularity for L2 — bitmap and session may have
different PRESERVE/REBUILD verdicts. §2.6 now explicitly
cross-references §2.10.
F2 — ShipperGroup §2.2 gains "External deps" row: N = RF comes
from master assignment via BlockVol.SetReplicaAddrs, not from
shipper-internal decision. Cross-entity contract (master assignment
↔ ShipperGroup size ↔ ReplicaReceiver expected-connection-count
↔ DistGroupCommit quorum arithmetic) made explicit so L2 split
can't silently drift sync_quorum.
F3 — ReplicaBarrier §2.4 scope rewritten from "per-request
ephemeral" to "per-request call-closure, BUT queue-state shared
per-volume via cond.Wait". Prior wording risked 1:1-porting into
a V3 stateless function, losing multi-watcher cond.Broadcast
semantics.
H5 added to §3 observations — cross-node epoch consistency
observation window for sync_quorum. V2 implicit via ack frame
carrying epoch; V3 L2 must pick "ack frame carries epoch" vs
"primary maintains per-replica epoch cache" before locking.
Different choices → different failover + rebuild-trigger semantics.
H6 added to §3 observations — write-path vs replication-path
concurrency residence. Three L2 options documented:
A) StorageBackend.Write triggers shipper (violates T3a layering)
B) ReplicatedBackend wraps StorageBackend+shipper (clean; +1 entity)
C) Replication inside DurableProvider (extends BUG-005 lesson)
L1 makes no recommendation; L2 LOCKS the decision before L3.
§5 restructured into 5 gated steps; step 1 is a mandatory sw V3
pre-scan of core/frontend/durable/ + core/frontend/*.go for
pre-baked replication-adjacent assumptions. Rationale cited per
architect: BUG-005 latent drift came from implicit V3 convention;
L1 must surface any such convention before L2 verdicts lock.
Concrete grep checklist included so the scan is 5 min, not open-ended.
§2 header + §4 open question #1 updated for 10-entity count.
Scope block references rebuild_bitmap.go explicitly.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Predecessor docs for the T3 batch, retained here for audit trail.
The closure report (`v3-phase-15-t3-closure-report.md`), contract
bridge catalogue, and BUG-005/006 artifacts already landed in
commits `4127e5136` + `6e196885e`; this commit fills the docs
those closure artifacts reference back to.
Landed:
v3-phase-15-t3-port-plan-sketch.md T3 umbrella sketch (rev-2.1, three-signed)
v3-phase-15-t3-port-audit.md T3.0 port audit + Addendum A (QA-signed)
v3-phase-15-t3a-mini-plan.md T3a scope + sign-off (CLOSED 0e1595c)
v3-phase-15-t3b-mini-plan.md T3b scope + sign-off (CLOSED 72d0d40)
v3-phase-15-t3c-mini-plan.md T3c scope + sign-off (CLOSED 829c6a9)
Total 1,346 lines of doc; no code impact.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Covers 6 areas based on CockroachDB/Ceph/etcd/Longhorn research:
1. Structured logging: zap + JSON + channel model (OPS/STORAGE/REPL/ISCSI/AUDIT/HEALTH)
2. Distributed tracing: OpenTelemetry spans across write/rebuild/failover paths
3. Metrics: 40+ must-have Prometheus metrics with histogram latency buckets
4. Debug tools: debug zip (logs+pprof+state), log merge, live tail
5. Audit logging: every admin mutation with actor/target/operation/result
6. Alert design: 3 tiers (page/ticket/log), anti-patterns to avoid
Identifies existing gaps: no I/O latency histogram, no rebuild duration
metric, no audit trail, no structured logging, no distributed tracing.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Covers three personas (developer/operator/platform engineer) with:
- One-command setup: weed server -block (10 seconds to first volume)
- Shell commands: block.list, block.status, block.health, block.create, etc.
- REST API: /block/volumes CRUD, /block/health
- Observability: Prometheus metrics, alerting rules, Grafana dashboard
- Actionable error messages (every error tells you what to do next)
- Dry-run by default for all destructive operations
Competitive comparison: 10s setup vs Ceph 30min, 13.5x write IOPS,
single binary for object + block storage.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
P1 feature updated: replace generic "structured results" with concrete
runs.db design (newline-delimited JSON, one line per run). Leverages
existing RunBundle system (manifest.json, result.json already exist).
New CLI commands: list, trend, gc, reindex, diff.
Regression detection via stddev comparison against rolling baseline.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
testrunner-roadmap.md: P0-P3 feature plan for multi-version comparison,
Ceph adapter, result tracking, cluster templates, debug mode.
dm-stripe-two-server.yaml: proven Linux dm-stripe across 2 sw-block
volumes on 2 servers. Results: single=42K IOPS → striped=79K IOPS (1.87x).
Data integrity verified via md5. Zero sw-block code changes needed.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
The architect's refactor correctly routes remote rebuild acks through
the shared observation path (pins, watchdog, deferred terminal success).
But requireReplicaSession fails with "sender not found" when the
orchestrator registry is reconciled between installSession and the
first ack arrival.
Fix: when emitTerminal=false (remote path), treat sender-not-found as
non-fatal. The remote coordinator already validated the session — the
sender lookup is for local observation only. Pins and watchdog handle
nil snap gracefully (updateRebuildProgressPin line 296 already checks
snap != nil).
This preserves the architect's design (shared observation + deferred
terminal success) while tolerating the sender registry race that only
affects the remote rebuild path.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
After RemoteRebuildIO.TransferFullBase returns, the OnAck callback has
already emitted SessionCompleted and stored achievedLSN. But
RebuildExecutor.Execute() continues calling sender methods which fail
("sender stopped") because the completion event already cleaned up the
sender. This error propagated to ExecutePendingRebuild which emitted a
spurious SessionFailed, knocking the mode back to degraded.
Fix: check remoteRebuildAchieved before emitting SessionFailed. If the
rebuild already completed via the ack path, log the post-completion
error but suppress the SessionFailed event.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Three fixes for the remote rebuild path:
1. Base-only completion: when BaseLSN == TargetLSN, the base image covers
all data — no WAL tail needed. MarkBaseComplete now auto-satisfies the
WAL condition and calls TryComplete so the session completes immediately
after the base transfer finishes.
2. Base lane protocol handshake: runBaseLaneClient now sends MsgRebuildReq
{Type: RebuildSessionBase} before reading. The RebuildServer requires
this handshake to dispatch to ServeBaseBlocks. Without it, the server
received raw frames it couldn't understand.
3. Direct ack events: OnAck emits engine events directly (SessionCompleted,
SessionProgressObserved, SessionFailed) instead of routing through
ObserveReplicaRebuildSessionAck which requires the sender in the
orchestrator registry. The remote coordinator owns the session — no
registry lookup needed.
Also adds diagnostic logging on both sides:
- Replica: logs parsed RebuildAddr and base lane client start
- Primary: logs sender state after installSession
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
The accepted ack from the replica is rejected with "sender not found"
even though installSession succeeds. Add diagnostic logging to verify
the sender exists in the orchestrator registry immediately after
installSession, and dump all registry IDs if not found.
This will reveal whether the sender is removed between installSession
and the ack arrival (by syncProtocolExecutionState, evaluateActivationGate,
or another ProcessAssignment that reconciles with a stale replica list).
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
When CommittedLSN=0 (sync_all mode, replica degraded), snapshot-tail
rebuild was chosen because IsRecoverable(checkpoint, 0) is vacuously
true (0 <= HeadLSN always). But snapshot-tail requires a valid committed
endpoint for tail-replay. Without it, ExecuteRebuildPlan calls
TransferSnapshot which RemoteRebuildIO doesn't support → immediate fail.
Fix: if CommittedLSN=0, force RebuildFullBase. This is the correct
source when the primary has data but no replica has confirmed durability.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Root cause: StatusSnapshot().CommittedLSN reports 0 in sync_all mode when
the replica shipper has no flushed progress (NeedsRebuild state). This is
correct for lineage-safe committed boundary, but PlanRebuild uses
CommittedLSN as RebuildTargetLSN. With target=0, shouldStartSessionCommand
rejects the StartRebuildCommand, and the rebuild IO never executes.
Fix: PlanRebuild falls back to HeadLSN when CommittedLSN is 0. The
primary's WAL head IS the data boundary the replica needs to reach.
The fact that no replica has confirmed durability is exactly why we're
rebuilding.
Also adds command type logging to coreApplyAndLog so tester can verify
which commands are actually emitted vs silently dropped.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Three correctness fixes for the remote rebuild path:
1. No double completion: for remote rebuilds, OnRebuildCompleted skips
RebuildCommitted since ObserveReplicaRebuildSessionAck already emitted
SessionCompleted on the accepted ack. One rebuild = one completion event.
2. SessionAckFailed with rejected observation: if OnAck rejects the failed
ack (stale session), don't use the sentinel errRebuildAckFailed. Return
a regular error so ExecutePendingRebuild emits the fallback SessionFailed.
No path leaves the engine session hanging.
3. Diagnostic logging in ExecutePendingRebuild: log the replicaID and
targetLSN on both nil-return (TakeRebuild mismatch) and successful take
paths. Also log the pending store in runRebuild with replicaID, targetLSN,
and IO type. This makes the TakeRebuild seam diagnosable on hardware
without rebuilding the engine package.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Replace the broken primary-local rebuild executor with RemoteRebuildIO,
a server-side engine.RebuildIO implementation that coordinates remotely.
The primary sends SessionControlV2 (with RebuildAddr trailer) to the
replica's control channel; the replica starts a local rebuild session
and auto-connects to the primary's rebuild server for the base lane.
Single rebuild route: ALL core-present rebuilds use RemoteRebuildIO.
The entire command chain is preserved unchanged:
PlanRebuild → pending → RebuildStarted → StartRebuildCommand
→ ExecutePendingRebuild → RemoteRebuildIO.TransferFullBase
Key changes:
- SessionControlMsg v2: optional RebuildAddr trailer (len-based decode)
- ReplicaRebuilding shipper state: session-gated live WAL lane
- RemoteRebuildIO: dials replica ctrl, sends session control, reads acks
- Ack forwarding through ObserveReplicaRebuildSessionAck (pins/watchdog)
- Completion proof from replica's achievedLSN, not primary's local vol
- Transport failures emit SessionFailed (no double-emit on ack failures)
- Progress ack rejection fails closed (stale session = abort)
- Replica auto-starts base lane client on v2 session control
State transitions:
NeedsRebuild → [accepted ack] → Rebuilding → [completed] → InSync
Rebuilding → [failed/EOF] → NeedsRebuild → [next probe] → retry
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Replace three bypass mechanisms with one unified model. When the
probe returns ProbeRebuildRequired, the host now starts the rebuild
through the existing recovery manager (StartRecoveryTask), which
resolves the rebuild address, plans the rebuild, and executes via
the v2bridge executor — the same path as master-driven RoleRebuilding.
New per-replica probe API:
- WALShipper.ProbeReconnect() → ReplicaProbeResult with typed outcome
- ShipperGroup.ProbeReconnectAll() → []ReplicaProbeResult
- BlockVol.ProbeReplicaOnboarding() / IsClosed()
Host-side wiring:
- handleReplicaProbeResult routes outcomes:
KeepUp → ShipperConnectedObserved
CatchUp → ShipperConnectedObserved (recovery manager handles session)
Rebuild → NeedsRebuildObserved + StartRecoveryTask (executes rebuild)
TemporaryFailure → no-op
- lastAssignmentsForPath reconstructs assignment for recovery manager
- onPrimaryRosterChanged probes all replicas (defined, called from watchdog)
- observePrimaryShipperConnectivity uses probe API
Probe fires via syncProtocolExecutionState immediately after assignment
processing — same heartbeat cycle, no timer delay.
Deleted: startDirectRebuild, resolveCtrlAddrForShipper,
TryReconnect/TryReconnectAll/TryReconnectShippers.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
When proactive reconnect finds WAL gap exceeds retained range:
1. Emit per-replica NeedsRebuildObserved to engine (with ReplicaID)
2. Resolve replica ctrl address from shipper group
3. Start direct rebuild session: send sessionControl(start_rebuild)
to replica's ctrl channel, stream base blocks, emit RebuildStarted
The primary drives the rebuild directly without master round-trip.
The master sees the result via heartbeat projection (needs_rebuild →
rebuilding → healthy). This matches V2 authority: master owns identity,
primary owns data-control recovery.
Added WALShipper.CtrlAddr() getter for address resolution.
resolveCtrlAddrForShipper maps data address to ctrl address via
shipper group (works for RF=2 and RF=3+).
startDirectRebuild runs in a goroutine: dials replica ctrl, sends
start_rebuild, waits for accepted ack, serves base blocks, emits
RebuildStarted to engine on success.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Revert detectAndEnqueueRebuildFromHeartbeat (Bridge 2) — master
should not drive rebuild assignments from heartbeat. The primary
owns data-control recovery per the V2 authority split.
Fix Bridge 1: NeedsRebuildObserved now carries per-replica identity.
resolveReplicaIDForShipper maps shipper DataAddr to ReplicaID via
the shipper group (works for RF=2 and RF=3+). The engine receives
the specific replica that needs rebuild, not a volume-level broadcast.
Primary-direct rebuild: the primary detects which replica needs
rebuild and will drive the session directly. The master learns about
it via subsequent heartbeat projection (needs_rebuild → rebuilding →
healthy). No master round-trip needed for the rebuild decision.
Added WALShipper.DataAddr() getter for address resolution.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
After rejoin, the shipper is configured but no I/O triggers Ship(),
so the shipper stays Disconnected and the core stays at
awaiting_shipper_connected indefinitely.
Fix: observePrimaryShipperConnectivity now calls TryReconnectShippers
when ShipperConfigured=true but ShipperConnected=false. This triggers
the full reconnect protocol (dial + handshake + bounded catch-up)
proactively, bringing the replica current without waiting for I/O.
Option B approach: uses the same reconnect path as Barrier() — not a
fake write or bare dial probe. CatchUpTo(headLSN) replays any retained
WAL entries, bringing the replica fully current.
New methods:
- WALShipper.TryReconnect(): full reconnect without foreground I/O
- ShipperGroup.TryReconnectAll(): probes all disconnected shippers
- BlockVol.TryReconnectShippers(): volume-level entry point
Also fix pre-existing test expectation: engine now emits
start_recovery_task on primary assignment with replicas.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Fix recover path TOCTOU: re-Lookup after AddReplica so the primary
refresh assignment includes the freshly added replica addresses.
Previously, Lookup (copy) was called before AddReplica modified the
registry, so entry.Replicas was empty → primary got replicas=0 →
shipper never configured.
Add 2 WAL pressure edge case tests:
- ShipperCatchUpOrEscalate: 64KB WAL, 200 writes, aggressive flusher.
Proves no hang/deadlock/corruption. Shipper either keeps up or
correctly escalates to NeedsRebuild.
- RebuildWithPinWhilePrimaryWrites: rebuild session active while
primary writes 7600+ blocks in 2s. Proves primary never freezes
— rebuild pin is on replica only, primary WAL recycles freely.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
All 43 actions pass on m01/m02 hardware. Auto-failover PASS.
dd_write: 30s → 123ms. Post-failover write: 33,621 IOPS.
1. WAL retention: remove keepup retention floor (MinShippedLSN).
WAL cannot be pinned during sustained async writes — any pin
strategy either fills WAL (blocking writes) or over-recycles
(breaking catch-up). Flusher recycles freely. Future LBA map
will provide catch-up without WAL retention.
MinShippedLSN on ShipperGroup retained as diagnostic surface.
2. Registry stale-cleanup race: add RegisteredAt grace period.
Race: master registers volume → next VS heartbeat arrives before
VS discovers the volume → stale cleanup deletes the entry →
failover finds 0 entries. Fix: skip stale cleanup for entries
registered within 30s (> 2 heartbeat intervals).
2 new tests: grace protects new entry, old entry still cleaned.
3. Shutdown heartbeat: VS disconnect heartbeat no longer claims
block inventory authority. Previously, the shutdown beat's
empty inventory triggered stale cleanup, deleting the entry
before failover could use it.
Scenario fix: recovery-baseline-failover.yaml now kills the
correct node (discovered primary, not hardcoded), connects to
the correct new primary for post-failover verification.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Wire protocol messages and transport handlers for the rebuild MVP:
Protocol messages (rebuild_transport.go):
- SessionControlMsg: epoch, sessionID, command, baseLSN, targetLSN,
snapshotID. Encode/Decode with fixed 37-byte wire format.
- SessionAckMsg: epoch, sessionID, phase, walAppliedLSN, baseComplete,
achievedLSN. Encode/Decode with fixed 34-byte wire format.
- MsgSessionControl (0x10) and MsgSessionAck (0x11) on control channel.
- SendSessionControl/SendSessionAck convenience functions.
Transport handlers:
- RebuildTransportServer: primary-side, streams all extent blocks as
MsgRebuildExtent frames (reusing existing rebuild message type),
ends with MsgRebuildDone.
- RebuildTransportClient: replica-side, receives base blocks and
routes through vol.ApplyRebuildSessionBaseBlock, marks base
complete on MsgRebuildDone.
4 transport tests:
- SessionControl wire round-trip
- SessionAck wire round-trip
- BaseBlockStreaming: full TCP loop, 1024 blocks streamed and verified
- SessionControlOverTCP: real TCP send/receive with accepted ack
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Add BlockService replica-side rebuild routing API that bridges
transport/host layer to BlockVol session surface:
StartReplicaRebuildSession(path, config)
ApplyReplicaRebuildWALEntry(path, sessionID, entry)
ApplyReplicaRebuildBaseBlock(path, sessionID, lba, data)
MarkReplicaRebuildBaseComplete(path, sessionID, totalBlocks)
TryCompleteReplicaRebuildSession(path, sessionID)
CancelReplicaRebuildSession(path, sessionID, reason)
ReplicaRebuildSession(path) → snapshot
Each method does one thing: validate → WithVolume → delegate to BlockVol.
No wire decoding, no protocol decisions, no state invention. Transport
wiring (sessionControl/walData/sessionData handlers) is the next step.
2 focused tests: skeleton routes correctly, stale session ID rejected.
Updated v2-rebuild-mvp-session-protocol.md with server skeleton section.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Tighten acceptance matrix with explicit per-boundary rows, signoff
reading split into hard blockers vs product hardening, and clear
rule: architecture-complete ≠ product-complete.
6 hard blockers before T6/T7:
1. WriteLBA/SyncCache/sync_all contract closure
2. Fresh replica bounded catch-up before live tail
3. Timeout/retention-loss classification for catch-up
4. publish_healthy alignment with one protocol contract
5. RF=2 stable identity on all shipping paths
6. Test audit for incorrect WriteLBA==commit assumptions
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
7-area acceptance matrix mapping current state vs product requirements:
write/durability contract, fresh replica bootstrap, host observation
completeness, serving/publish alignment, snapshot/rebuild convergence,
adapter consistency, test contract alignment.
Each item marked with: current state, required for product, blocks
T6/T7, best test level. Priority ordered into must-close-before-Stage-1,
should-close-before-Stage-2, and can-close-after-T6/T7.
Key diagnosis: architecture-complete, execution-incomplete. The engine
thinks like a product; the data plane still behaves partly like a
prototype. The gap is end-to-end contract closure.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Add host-side protocol state seam that derives per-replica execution
state from V2 sender/session snapshots and blocks live-tail WAL
shipping while an active recovery session is in progress.
New file: weed/server/block_protocol_state.go
- replicaProtocolExecutionState derived from engine snapshots
- LiveEligible=false during active catch-up/rebuild sessions
- bindProtocolExecutionPolicy wires policy into BlockVol
- syncProtocolExecutionState called after assignments + core events
Data plane changes:
- WALShipper.Ship() checks liveShippingPolicy before dial/send
- BlockVol.SetLiveShippingPolicy persists across shipper group rebuilds
- ShipperGroup propagates policy to all shippers
Design contract: sw-block/design/v2-protocol-aware-execution.md
Scope: WAL-first rollout only. Prevents illegal live-tail delivery
during active recovery. Does not change snapshot/build behavior or
move backlog. Next wave: bounded WAL catch-up under same contract.
Tests: 4 unit/component tests for phase gate behavior, plus bootstrap
seam tests that confirmed the two pre-existing bugs locally.
13 files changed, 900 insertions, 69 deletions.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Update coverage reading to reflect 49 tests (6 new component tests).
Add full roster status table with per-item strong/bounded/missing
marking and mapped test function names.
Unit+component: 32 of 33 items strong (T4-C7 NVMe bounded).
Integration: 6 of 10 missing (Tier 2 next).
Hardware: 4 of 4 missing (T6/T7 staged plan).
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Add detailed coverage mapping of 43 existing tests against the test
roster. Identify 7 missing component tests and 3 missing integration
tests with concrete scenarios, file placement, and must-prove criteria.
Key finding: every tester-found bug during T1-T5 was a wiring bug caught
by reviewing the production path, not by unit tests on pure logic. This
confirms component tests are the highest-value gap for CI/CD protection.
Priority order: Tier 1 (7 component tests, do now), Tier 2 (3 integration
tests, do before hardware), Tier 3 (4 hardware scenarios, T6/T7).
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Add ClusterReplicationMode and EngineProjectionMode to
FailoverVolumeState so each volume in the failover diagnostic
carries its cluster/engine mode at diagnosis time.
FailoverDiagnosticSnapshot() enriches volume entries by looking up
the registry entry for each volume. This covers both the block
volume API (GET /block/volume/{name}) and the failover diagnostic
snapshot surface.
Update phase doc to reflect actual exposure paths.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Fix three tester findings on T5:
1. RF2 with missing replicas now reports "degraded" instead of
"no_replicas". Only RF=1 with no replicas returns "no_replicas".
Missing replica in an RF2 set is a degraded cluster state.
2. TransportDegraded signal now incorporated: if master-observed
transport is degraded, ClusterReplicationMode is at least
"degraded" regardless of individual replica health.
3. API surface exposure: EngineProjectionMode and
ClusterReplicationMode now appear on blockapi.VolumeInfo and are
populated in entryToVolumeInfo(). Operators can consume both
through GET /block/volume/{name} with distinct JSON field names.
12 tests: keepup, catching_up, stale degraded, LSN gap needs_rebuild,
rebuilding role, RF1 no_replicas, RF2 missing degraded, transport
degraded, distinctness, heartbeat update, worst dominates, API
surface distinct naming.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Add ClusterReplicationMode as a distinct master-owned cluster-level
replication health judgment, computed from multi-replica facts:
replica LSN lag, heartbeat freshness, role state. Monotonic: worst
replica state dominates.
Modes: "no_replicas" (RF=1), "keepup" (all healthy), "catching_up"
(replica behind but recoverable), "degraded" (stale heartbeat or
barrier failure), "needs_rebuild" (unrecoverable gap or rebuilding
role).
Distinct from EngineProjectionMode (VS-local engine truth) and
VolumeMode (legacy). They answer different questions, live in
different fields, have different names. Tests explicitly prove the
two can differ without conflict.
Computed in recomputeReplicaState() alongside existing VolumeMode.
Updated on every heartbeat that touches the entry.
9 tests: keepup, catching_up, stale degraded, LSN gap needs_rebuild,
rebuilding role, no_replicas, distinctness from EngineProjectionMode,
heartbeat-driven update, worst-replica-dominates (RF3).
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Fix two tester findings:
1. Missing engine projection now fails closed: if v2Core is active but
CoreProjection(path) is missing, gate locally with reason
"missing_engine_projection". Mirrors T2's fail-closed posture.
Only skips enforcement when V2 core is entirely absent.
2. NVMe/TCP now gated alongside iSCSI: gateServing() calls both
targetServer.DisconnectVolume() and nvmeServer.RemoveVolume().
ungateServing() re-registers with both iSCSI and NVMe. A gated
volume is unreachable through all frontend paths.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Fix three tester findings on T4 activation gate:
1. Real serving enforcement: evaluateActivationGate now calls
gateServing() → DisconnectVolume(iqn) on gate (terminates active
iSCSI sessions, removes volume from target). ungateServing() →
AddVolume(iqn, adapter) on clear (re-registers volume). This is
actual serving enforcement, not just bookkeeping.
2. Wire propagation: add activation_gated (field 25) and
activation_gate_reason (field 26) to proto BlockVolumeInfoMessage.
Add generated Go fields + getters. Add proto conversion in
InfoMessageToProto/InfoMessageFromProto. Gate state now rides the
real VS→master heartbeat wire.
3. Runtime ungate: evaluateActivationGate() now also runs in
applyCoreEvent() (the observation-driven path), not just
applyCoreAssignmentEvent(). Recovery/catch-up completion that
transitions the projection to publish_healthy/replica_ready now
clears the gate and re-registers the volume automatically.
ClearActivationGate() remains as an explicit override for edge cases
but is no longer the primary ungate path.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
After assignment executes through V2 core, evaluateActivationGate()
checks the resulting projection locally. If mode is degraded,
needs_rebuild, bootstrap_pending, or allocated_only, the volume is
gated from serving. Gate is enforced immediately after assignment,
before the next heartbeat round-trip.
Gate cleared only when projection reaches publish_healthy or
replica_ready. IsActivationGated() provides the query surface for
iSCSI/NVMe adapter enforcement. Heartbeat carries ActivationGated
and ActivationGateReason fields so master can observe the gated state
(report path, not enforcement path).
activationGated map on BlockService tracks per-volume gate state.
Initialized in constructor. Test helper updated to include it.
6 tests: degraded gates, needs_rebuild gates, healthy clears gate,
gate enforced before heartbeat, recovery re-enables, assignment with
degraded projection triggers gate.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Replace misleading V2PromotionEnabled/V2PromotionReady booleans with
single V2PromotionMode string: "disabled", "placeholder_fail_closed",
or "transport_ready".
Previous V2PromotionReady was true whenever any querier was installed,
including the placeholder that always returns error. Now the diagnostic
accurately distinguishes placeholder (fail-closed until proto regen)
from real gRPC transport.
blockV2EvidenceTransport bool on MasterServer tracks whether the real
transport querier is installed. Currently always false (placeholder).
Set to true only when real gRPC querier replaces the placeholder after
proto regen.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
FailoverDiagnostic now carries V2PromotionEnabled and V2PromotionReady
fields. MasterServer.FailoverDiagnosticSnapshot() enriches the failover
state diagnostic with rollout gate visibility so operators can confirm
whether the master is on V1, V2, or V2-fail-closed-placeholder mode.
Update phase-20.md: document default=false rollout policy (safe default
until proto regen enables evidence RPC, then flip to default true).
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Wire V2 promotion into production binary:
- Add --block.v2Promotion CLI flag on weed master (default false)
- MasterOption.BlockV2Promotion → NewMasterServer wires flag + querier
- defaultBlockVSQueryEvidence placeholder (returns explicit error until
proto regen on M01 enables gRPC evidence RPC)
Fix three fail-closed violations found by tester:
1. blockV2Promotion=true + nil querier now fails closed with explicit
log instead of silently falling back to V1
2. Partial evidence (any candidate query failed) now fails closed —
unreachable candidate may be the most durable, promoting from
incomplete evidence violates durability-first ordering
3. Clear EngineProjectionMode in applyPromotionLocked (already in
previous commit, verified in tests here)
2 new tests: NilQuerier_FailsClosed, PartialEvidenceFailure_FailsClosed.
Total T3 tests: 7, all pass. Existing V1 failover tests unaffected.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Wire V2 promotion into the real master failover decision path:
promoteReplica() now dispatches to promoteReplicaV2() when
blockV2Promotion flag is true. V2 path queries each candidate for
fresh evidence via pluggable BlockPromotionEvidenceQuerier, selects
by CommittedLSN (durability-first), and fail-closes when no eligible
candidate exists. No silent fallback to V1.
Feature flag: blockV2Promotion bool on MasterServer. When false,
existing promoteReplicaV1() (health-score-first) is used unchanged.
Flag is explicit and observable, not a hidden rescue path.
Registry: add PromoteReplicaByServer() for V2 path where master
already knows the winner. Clear stale EngineProjectionMode in
applyPromotionLocked (complements T1 turnover fix).
T2 fix: fail-closed when V2 core projection is absent —
Eligible=false with reason "missing_engine_projection". CommittedLSN
from core used unconditionally (no WALHeadLSN overstatement).
5 T3 integration tests: higher CommittedLSN wins, all-ineligible
fail-closed, evidence-failure fail-closed, flag-off uses legacy,
epoch bump + assignment enqueue only after selection.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
VS-side evidence handler (QueryBlockPromotionEvidence) reads live
blockvol.Status() + V2 core projection at call time. Fail-closed:
no core projection → ineligible with reason "missing_engine_projection".
Engine CommittedLSN used unconditionally when core present (no WALHeadLSN
overstatement). Eligibility owned by local V2 engine, not master.
Master-side selection (selectDurabilityFirstCandidate): durability-first
ordering by CommittedLSN, tie-break WALHeadLSN then HealthScore. All
ineligible → fail-closed, no promotion. Pluggable querier
(BlockPromotionEvidenceQuerier) for T3 wiring.
Proto messages added to volume_server.proto. gRPC transport binding
pending proto regen on M01 — this commit delivers evidence semantics
and selection substrate, not full end-to-end RPC closure.
Phase 20 doc updated with T2-T5 reviewer packs and cross-task guardrails.
13 tests: live facts, core projection mode, fail-closed no-core, 4 gated
modes, missing volume, epoch mismatch, CommittedLSN ordering, WALHeadLSN
tie-break, HealthScore tie-break, all-ineligible, mixed collection.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Add engine_projection_mode as a distinct proto/wire/registry field
that carries pure V2 engine-derived local projection mode from VS
to master. Reads ONLY from CoreProjection — no ad-hoc fallback.
Separate from existing VolumeMode: EngineProjectionMode is VS-local
V2 engine truth, VolumeMode is the existing field that conflates V2
and V1 paths. Both exist during transition; only EngineProjectionMode
is V2-authoritative.
Clears stale value on primary turnover: when a newly promoted primary
heartbeats without the field, the old primary's projection is not
preserved (prevents synthetic master-side truth).
5 focused tests: propagation, distinctness (hard assertion), backward
compat preservation, turnover-clears, turnover-with-field.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Live HTTP evidence transport, continuous Loop2 service, bounded auto
failover trigger, runtime-managed frontend export, bounded replica
repair, end-to-end RF2 handoff with continued I/O on new primary,
bounded operator HTTP surface, and CSI V2 runtime backend adapter.
11 new proof tests covering the full M6-M10 chain plus CSI create/
lookup/publish through the V2 runtime path.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Freeze the first bounded pilot/preflight/stop/rollout-review artifact set and sync the global product ledgers so productionization can start from an explicit chosen-envelope discipline instead of ad hoc rollout judgment.
Made-with: Cursor
Freeze the first Phase 17 branch/contract/policy/envelope package, add review and supported-matrix artifacts, and sync the product-completion and claim-evidence ledgers to the new bounded post-Phase-16 checkpoint.
Made-with: Cursor
Bind non-authoritative inventory, restart primary-truth rebasing, and sparse replica readiness retention into the heartbeat/master seam, and package the bounded finish-line checkpoint with explicit claims, non-claims, and proof commands.
Made-with: Cursor
Carry explicit volume_mode_reason across the heartbeat/master/API seam so outward surfaces retain the bounded core-owned explanation behind mode transitions.
Made-with: Cursor
Use ReplicaEligible instead of PublishHealthy in the heartbeat collector test now that publish health is rebound to publication truth rather than receiver readiness.
Made-with: Cursor
Make the heartbeat/master boundary preserve explicit volume_mode truth so master consume no longer reconstructs outward mode only from secondary heartbeat signals. Keep backward compatibility by falling back to the previous reconstruction when older heartbeats do not send the field.
Made-with: Cursor
Make the heartbeat/master boundary preserve explicit publish_healthy truth so master consume no longer reconstructs healthy publication only from secondary readiness and degraded heuristics. Keep backward compatibility by falling back to the previous reconstruction when older heartbeats do not send the field.
Made-with: Cursor
Make the heartbeat/master boundary preserve explicit needs_rebuild truth so primary heartbeat consume no longer collapses that stronger mode into a generic degraded signal. Keep backward compatibility by falling back to the previous heuristic when older heartbeats do not send the field.
Made-with: Cursor
Make the heartbeat/master boundary carry explicit replica readiness truth so the registry no longer depends only on replica transport-address presence as a readiness proxy. Keep backward compatibility by falling back to the old address heuristic when older heartbeats do not send the field.
Made-with: Cursor
Move removed-replica drain and replica-scoped invalidation onto explicit core-command paths so the widened multi-replica runtime no longer depends on coarse host-side recovery handling.
Made-with: Cursor
Emit one core-owned start_recovery_task per primary catch-up replica so the bounded multi-replica startup path no longer depends on a single-replica assumption.
Made-with: Cursor
Track catch-up observations per replica so the volume-level recovery view stays in catching_up until all bounded replicas complete. This preserves the current bounded semantics while removing an overclaim that would block later multi-replica startup ownership work.
Made-with: Cursor
Carry replica-scoped addressing through bounded recovery planning and completion events so the core no longer depends on a volume-only observation seam. This preserves the current single-replica catch-up and rebuilding behavior while aligning the observation side with the replica-scoped command path.
Made-with: Cursor
Replace the remaining volume-scoped recovery command and pending slot
with replica-scoped addressing on the bounded core-present path. This
preserves the current single-replica catch-up and rebuilding behavior
while removing the structural blocker for later multi-replica startup
ownership.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Move dispatcher-facing host effects out of volume_server_block.go into
blockcmd while keeping server-owned cache/state semantics in weed/server.
Document Batch 10 delivery and Batch 11 stop-line review so the
separation line closes without over-extracting readiness-state mutation.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Move BlockVol-backed command bindings into v2bridge and move non-BlockVol
command operations into weed/server/blockcmd. This keeps dispatch and host
effects in weed/server, keeps backend binding in v2bridge, and further
shrinks volume_server_block.go toward a host shell while preserving
current command-driven proofs.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
resolveRecoveryContext now also derives rebuildAddr from assignments,
so the full host-side recovery context is resolved in one call:
- volPath (from replicaID)
- rebuildAddr (from assignments via deriveRebuildAddr)
- recovery bindings (driver + executor via BuildRecoveryBundle)
- replicaFlushedLSN (from sender session)
startTask/runRecovery/runCatchUp/runRebuild now pass assignments
instead of rebuildAddr. No separate rebuildAddr resolution remains
outside the resolver.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
New recoveryContext type + resolveRecoveryContext method consolidates:
- volumePathForReplica (volPath from replicaID)
- v2bridge.BuildRecoveryBundle (driver + executor from BlockVol)
- sender/session lookup (replicaFlushedLSN for catch-up start)
runCatchUp and runRebuild now read as:
resolve → plan → branch (legacy or core-present)
Removed buildRecoveryBundle (inlined into resolveRecoveryContext).
block_recovery.go no longer has any inline context assembly —
it is now a pure orchestration shell.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
New v2bridge.BuildRecoveryBundle(vol, rebuildAddr) assembles all
recovery bindings (Reader + Pinner + StorageAdapter + Executor) from
a real BlockVol instance in one call.
block_recovery.go changes:
- Removed local recoveryBundle type
- buildRecoveryBundle now delegates to v2bridge.BuildRecoveryBundle
inside WithVolume, returns (driver, executor, err)
- Removed direct v2bridge.NewReader/NewPinner/NewExecutor construction
- Removed bridge import (no longer needed)
- runCatchUp/runRebuild use (driver, executor, err) directly
block_recovery.go no longer knows how to construct Reader, Pinner,
StorageAdapter, or Executor. It only knows: resolve volPath, ask the
factory for bindings, plan, branch to legacy or core-present path.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Step 2: Rebuild completion status port
- New runtime.RebuildCompletionStatus + DeriveRebuildCommitted:
reusable shaping logic for post-rebuild snapshot → RebuildCommitted event
- block_recovery.go OnRebuildCompleted: delegates to DeriveRebuildCommitted,
host only reads raw snapshot via readRebuildStatus (thin binding)
- Removed 15 lines of inline flushedLSN/checkpointLSN/achievedLSN computation
Step 3: Recovery bundle factory
- New buildRecoveryBundle: shared host-side setup for both catch-up and rebuild
(creates Reader + Pinner + StorageAdapter + Executor + RecoveryDriver)
- runCatchUp and runRebuild both use buildRecoveryBundle instead of
duplicating the WithVolume → NewReader → NewPinner → NewStorageAdapter →
NewExecutor → RecoveryDriver chain
- runCatchUp/runRebuild are now thin host-shell methods
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Replace interface{} fields in runtime.PendingExecution with typed handles:
- Driver: *engine.RecoveryDriver (was interface{})
- Plan: *engine.RecoveryPlan (was interface{})
- CatchUpIO: engine.CatchUpIO (was interface{})
- RebuildIO: engine.RebuildIO (was interface{})
block_recovery.go:
- ExecutePendingCatchUp/Rebuild: direct field access (pe.Driver, pe.Plan)
instead of type assertions (pe.Driver.(*engine.RecoveryDriver))
- CancelFunc: pe.Driver.CancelPlan(pe.Plan, reason) — no casts
- 6 type assertions removed from production path
Test files: remove Plan type assertions — fields are typed end-to-end.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
H wiring: block_recovery.go now uses runtime.PendingCoordinator
- Removed local pendingRecoveryExecution type + store/take/peek/has/cancel
- ExecutePendingCatchUp/Rebuild delegate to coord.TakeCatchUp/TakeRebuild
- Shutdown uses coord.CancelAll
- Added CancelAll to PendingCoordinator
I wiring: executeCatchUpPlan/executeRebuildPlan replaced
- ExecutePendingCatchUp now calls rt.ExecuteCatchUpPlan with RecoveryManager
as RecoveryCallbacks (OnCatchUpCompleted/OnRebuildCompleted)
- ExecutePendingRebuild follows same pattern
- Local executeCatchUpPlan/executeRebuildPlan methods removed
J structural: legacy no-core branches extracted
- executeLegacyCatchUp: wraps rt.ExecuteCatchUpPlan for v2Core==nil path
- executeLegacyRebuild: wraps rt.ExecuteRebuildPlan for v2Core==nil path
- Clear "LEGACY NO-CORE COMPATIBILITY" section with structural separation
- runCatchUp/runRebuild now branch cleanly: legacy helper vs core coordinator
Test updates: pendingRecoveryExecution → rt.PendingExecution, field casing,
Plan type assertions.
Validation: all P4, P16B, and ApplyAssignments tests pass.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Add explicit "LEGACY NO-CORE COMPATIBILITY" section header in
block_recovery.go marking HandleAssignmentResult and
HandleRemovedAssignments as compatibility-only entry points.
The comment block explicitly states:
- These are for pre-Phase-16 no-core paths and older tests
- Core-present paths use StartRecoveryTask + ExecutePending*
- These should NOT be strengthened into semantic-authority proofs
No behavioral change — structural labeling only. All validation passes.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
New reusable pending-execution coordinator with fail-closed command matching:
- Store/TakeCatchUp/TakeRebuild/Cancel/Has/Peek
- TakeCatchUp: fail-closed on target LSN mismatch (cancel + return nil)
- TakeRebuild: same fail-closed semantics
- Cancel callback invoked on mismatch or explicit cancellation
9 tests prove boundary behavior:
- match succeeds, mismatch cancels, explicit cancel, noop on empty,
peek non-destructive, store replaces, take from empty
No weed/ imports. Pure coordination logic reusable by any adapter shell.
weed/server/block_recovery.go rebinding deferred to Task I.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Task F (Pinner):
- block_recovery.go: removed pinnerShimForRecovery (11 lines of pure
pass-through). v2bridge.Pinner structurally satisfies bridge.BlockVolPinner
(same method signatures), so it's passed directly.
Task G (Executor):
- Already clean. v2bridge.Executor is used directly without any shim —
structurally satisfies engine.CatchUpIO and engine.RebuildIO.
No code changes needed.
After Task E+F+G: zero shim types remain in block_recovery.go.
v2bridge Reader/Pinner/Executor all satisfy sw-block contracts directly.
Validation:
- go test ./weed/storage/blockvol/v2bridge/ -run "TestPinner_|TestExecutor_|TestBridge_" → PASS
- go test ./weed/server/ -run "TestP4_|TestP16B_" → PASS (8 tests)
- go test ./sw-block/bridge/blockvol/... → PASS
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Reader backend-binding extraction:
- v2bridge/reader.go: Reader.ReadState() now returns bridge.BlockVolState
directly instead of a local v2bridge.BlockVolState mirror type.
Removed the local BlockVolState type entirely.
- block_recovery.go: removed readerShimForRecovery (12 lines of 1:1
field copying). Reader is now passed directly as bridge.BlockVolReader.
Before: v2bridge.Reader → v2bridge.BlockVolState → readerShim → bridge.BlockVolState
After: v2bridge.Reader → bridge.BlockVolState (direct)
v2bridge now imports sw-block/bridge/blockvol for the contract type
(control.go already did this, reader.go now follows the same pattern).
Validation:
- go test ./sw-block/bridge/blockvol/... → PASS
- go test ./weed/storage/blockvol/v2bridge/ -run "TestReader_" → PASS
- go test ./weed/server/ -run "TestP4_|TestP16B_" → PASS (8 tests)
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Remove direct fmt.Sprintf identity construction from v2bridge/control.go.
Both convertReplicaAssignment and convertRebuildAssignment now use:
- bridge.ReplicaAssignmentForServer (canonical ReplicaID derivation)
- bridge.RecoveryTargetForRole (canonical role → SessionKind mapping)
Before: 3 call sites with inline fmt.Sprintf("%s/%s", vol, server)
After: 0 — all identity construction goes through sw-block canonical helpers
volume_server_block.go already used bridge helpers (no change needed).
Validation:
- go test ./sw-block/bridge/blockvol/... → PASS (10 tests)
- go test ./weed/storage/blockvol/v2bridge/ -run "TestControl_|TestBridge_" → PASS (7 tests)
- go test ./weed/server/ -run "TestBlockService_ApplyAssignments_RebuildingRole_" → PASS
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
16B widened from catch-up-only to catch-up + rebuild:
- StartRebuildCommand: core emits rebuild command, adapter executes
- Fail-closed: pending rebuild does not run without fresh command
- Recovery observations close back into core projection
New proofs:
- StartRebuildCommand_ConsumesPendingPlanAndUpdatesProjection
- RunRebuild_FailClosedWithoutFreshStartRebuildCommand
Review docs:
- phase-16-rev3-review.md: widened 16B review object
- phase-16-rev3-manager-rereview.md: manager challenge response
- phase-16-checkpoint-review.md: updated
Non-claims: not full recovery-loop closure, not end-to-end
failover/publication, not launch readiness.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Make the first V2 core owner explicit in sw-block by freezing Phase 14 docs, mode/readiness/publication semantics, and bounded command emission rules. This turns accepted Phase 13 constraints into executable core behavior without overclaiming live runtime cutover.
Made-with: Cursor
Add computed VolumeMode to BlockVolumeEntry with 5 normalized modes:
- allocated_only: RF=1, no replicas (standalone)
- bootstrap_pending: RF>1 but replicas not yet ready (first-write pending)
- publish_healthy: all replicas ready, no transport degradation
- degraded: replication impaired but recoverable
- needs_rebuild: unrecoverable gap, rebuild required
Code changes:
- master_block_registry.go: computeVolumeMode() called from
recomputeReplicaState(), VolumeMode field on BlockVolumeEntry
- master_server_handlers_block.go: VolumeMode exposed in REST API
- blockapi/types.go: VolumeMode field in VolumeInfo
- testrunner types: VolumeMode for scenario assertions
7 tests prove mode normalization:
- AllocatedOnly, BootstrapPending (2 cases), PublishHealthy,
Degraded, NeedsRebuild, SurfaceConsistency (transition proof)
Interpretation rule: current integrated tests validate V1 runtime
under V2 constraints, not a completed V2 runtime (Phase 14 scope).
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Bug: After failover promotes a replica to primary, the old primary
re-registers via heartbeat as a replica (lower epoch). But the master
never sent an updated Primary assignment to the new primary with the
re-registered replica's addresses. The new primary had 0 shippers →
replication dead. sync_all barrier passed vacuously.
Root cause: upsertServerAsReplica (heartbeat reconciliation) added the
re-registered server to Replicas[] but didn't (a) populate DataAddr/
CtrlAddr from heartbeat info, or (b) trigger a primary assignment
refresh.
Fix:
- master_block_registry.go: upsertServerAsReplica now copies DataAddr/
CtrlAddr from heartbeat info and sets NeedsPrimaryRefresh flag.
UpdateFullHeartbeat returns HeartbeatResult with PrimaryRefreshNeeded
entries. DrainPrimaryRefreshNeeded collects and clears the flag.
- master_block_failover.go: add enqueuePrimaryRefresh — builds a
Primary assignment with all current replica addresses and enqueues it.
- master_grpc_server.go: heartbeat handler processes PrimaryRefreshNeeded
entries after UpdateFullHeartbeat.
Gate test: TestPromote_AssignmentHasReplicaAddrs now PASSES —
after promote + re-register, the new primary gets an assignment with
replicaDataAddr=vs1:14260 and replicaAddrs=1.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
P0 bug on real hardware: assignments are re-delivered every heartbeat
cycle (5s). First setupReplicaReceiver succeeds (receiver starts on
deterministic port). Second call fails with "bind: address already in
use" because the listener is already bound. The volume stays permanently
degraded, blocking all RF=2 sync_all replication.
Fix: skip StartReplicaReceiver if v.replRecv is already set. The
receiver only needs to start once per volume lifetime.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Investigation result:
- Dual-BlockVol hypothesis: DISPROVEN (one instance per path, correct wiring)
- Root cause: adapter wiring bug in test allocator
soak_test.go blockVSAllocate returned ReplicaDataAddr = "vs2:9333:14260"
(server + ":port" where server already has a port → three colons, invalid)
This caused setupReplicaReceiver to fail silently → no data replicated
Root cause classification: adapter/test-harness bug
- NOT a backend data visibility bug
- NOT a core-rule gap
- The engine read path works correctly (TestSyncAll_FullRoundTrip passes)
Code changes:
- qa_block_soak_test.go: fix allocator to use host:port (not server:port),
use deterministic FNV-hashed ports matching production ReplicationPorts
- qa_block_cp13_8a_test.go: 2 new integration tests proving replica reads
work through both ReadLBA and adapter.ReadAt, before and after promotion
Remaining contradiction for CP13-8 scenario on real hardware:
- The production weed cluster uses ReplicationPorts (deterministic) which
should not have this bug. If CP13-8 still fails on m01/M02, the cause
is different from this test-harness issue and needs a separate investigation.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
1. assert_contains: change actual/expected to value/contains (matches
the action implementation in system.go)
2. Add assert_greater for pgbench TPS > 0 after pgbench_run (closes
the pgbench durability pass criterion in the doc)
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
- master_block_registry.go: minor role-handling fixes
- qa_failover_role_test.go: new failover role test
- testrunner/actions/devops.go: new devops action helpers
- recovery-baseline-failover.yaml: scenario alignment
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
- phase-13.md: CP13-1 through CP13-6 accepted, CP13-7 active
- phase-13-log.md: full technical + delivery packs for CP13-2..CP13-7
- phase-13-cp4-state-eligibility.md: refined barrier behavior table
(Disconnected/Degraded as recovery entry points, not eligibility)
- phase-12.md: minor cross-reference updates
- Older phase docs: minor wording alignment
- Design docs: V2 development plan and completion overview updated
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Tighten TestReconnect_GapBeyondRetainedWal_NeedsRebuild assertion from
"NeedsRebuild or Degraded" to strictly "NeedsRebuild". The handshake
R < S path returns NeedsRebuild directly — tolerating Degraded weakened
the proof.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Two fixes:
1. TestReconnect_GapBeyondRetainedWal_NeedsRebuild: rewritten to test the
real reconnect handshake gap detection path (R < S in
reconnectWithHandshake). Sequence: establish sync → disconnect →
release retention hold via timeout → write + flush to advance WAL past
replica position → reconnect → handshake detects R=0 < S=9 → NeedsRebuild.
Log proves: "reconnect: gap too large R=0 H=8 S=9"
2. TestReplicaState_RebuildComplete_ReentersInSync: reclassified from
primary proof to support evidence (does not start from live NeedsRebuild
shipper state, but proves rebuild mechanics work end-to-end).
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
1. TestWalRetention_TimeoutTriggersNeedsRebuild: add hard assertion that
checkpoint advances past replicaFlushedLSN after NeedsRebuild (proves
hold is actually released, not just state transition)
2. TestWalRetention_RequiredReplicaBlocksReclaim: remove stale "EXPECTED
TO FAIL" / duplicate comment block
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Three fixes:
1. TestWalRetention_RequiredReplicaBlocksReclaim: rewritten from log-only
placeholder to hard assertion (checkpointLSN <= replicaFlushedLSN)
2. TestWalRetention_TimeoutTriggersNeedsRebuild: rewritten from log-only
to hard assertion (State() == NeedsRebuild after 1ns timeout)
3. EvaluateRetentionBudgets: uses RetentionBudgetParams struct with
actual BlockSize from volume config instead of hardcoded 4096
All 3 retention tests now have real state/progress assertions.
No placeholder or log-only evidence remains in CP13-6 proof package.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Add max-bytes retention budget alongside existing timeout budget:
- shipper_group.go: EvaluateRetentionBudgets now checks both timeout
(last contact time) and max-bytes (entry lag * 4KB > maxBytes).
Either exceeding budget → NeedsRebuild state transition.
- blockvol.go: add walRetentionMaxBytes (64MB default), pass to
EvaluateRetentionBudgets with primaryHeadLSN.
TestWalRetention_MaxBytesTriggersNeedsRebuild upgraded from PASS*
(log-only placeholder) to real PASS: asserts State()==NeedsRebuild
after lag exceeds configured max-bytes budget.
Retention contract: hold-back blocks reclaim for recoverable replicas,
timeout and max-bytes budgets escalate to NeedsRebuild and release hold.
Full rebuild lifecycle remains CP13-7 scope.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Replace "observable CatchingUp state transition" with the actual 3
signals the test asserts: seeded hasFlushedProgress, receivedLSN
advance, non-zero replicaFlushedLSN.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Findings fixed:
1. TestAdversarial_ReconnectUsesHandshakeNotBootstrap now has 3 observable
proof points instead of just "SyncCache succeeded":
- new shipper HasFlushedProgress=true (seeded from old group)
- replica receivedLSN advances during SyncCache (catch-up delivered entries)
- shipper replicaFlushedLSN > 0 after barrier (durable progress established)
Bootstrap alone would not advance receivedLSN — it only sends the barrier.
2. TestBug2 stale comment removed: "must NOT call SetReplicaAddr" replaced
with accurate CP13-5 explanation that SetReplicaAddrs now preserves
hasFlushedProgress across shipper replacement.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Bug: SetReplicaAddrs created fresh shippers (hasFlushedProgress=false),
so after disconnect, the new shipper used bootstrap instead of reconnect
handshake. Bootstrap doesn't replay missed WAL entries — barrier hung.
Fix:
- blockvol.go: SetReplicaAddrs checks if old shipper group had durable
progress (AnyHasFlushedProgress). If so, seeds new shippers with
hasFlushedProgress=true → they use reconnect handshake + catch-up.
- shipper_group.go: add AnyHasFlushedProgress() helper.
3 baseline FAILs now PASS:
- ReconnectUsesHandshakeNotBootstrap: reconnect path used, not bootstrap
- CatchupMultipleDisconnects: repeated disconnect/reconnect recovers
- CatchupDoesNotOverwriteNewerData: catch-up completes, safety exercised
7 tests promoted to CP13-5 primary proof.
TestAdversarial_NeedsRebuildBlocksAllPaths still FAIL (CP13-7 scope).
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Contract review: 6-state set (Disconnected, Connecting, CatchingUp,
InSync, Degraded, NeedsRebuild). Only InSync proceeds to barrier
request path. All other states either fail immediately or attempt
reconnect (must succeed before reaching barrier).
New test: TestBarrier_NonEligibleStates_FailClosed — systematically
verifies each non-eligible state (Connecting, CatchingUp, NeedsRebuild,
Disconnected) is rejected by Barrier(), and InSync is the only state
that enters the barrier request path.
5 baseline tests promoted to CP13-4 primary proof.
No production code changed — contract review + new focused test only.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
The previous test only checked wire decode + fresh shipper state, never
calling shipper.Barrier() against a legacy response source.
New test runs a fake TCP control server that responds with a 1-byte
BarrierOK (no FlushedLSN). Shipper.Barrier() is called against it and
must return an error containing "no FlushedLSN". Verifies the real
rejection path at wal_shipper.go:229-231.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Bug: BarrierOK with FlushedLSN == 0 (legacy 1-byte response) was counted
as successful sync_all durability even though no authoritative durable
progress was established. This allowed a legacy replica to silently pass
through the sync_all barrier without proving any LSN was fsynced.
Fix (wal_shipper.go): BarrierOK with FlushedLSN == 0 now returns an
error instead of nil. Barrier success requires the replica to report a
non-zero FlushedLSN proving which LSN was durably persisted. This makes
the code match the CP13-3 contract: replicaFlushedLSN is the sole
authority for sync_all durability.
New test: TestBarrier_LegacyResponseRejectedBySyncAll — proves legacy
1-byte responses don't establish durable authority.
Contract review doc updated to reflect the code fix.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Contract review (no code changed):
- replicaFlushedLSN is the sole authority for replica durability
- flushedLSN advanced only after fd.Sync() on replica (not on receive)
- shippedLSN/sentLSN are explicitly diagnostic (comment at line 268)
- barrier response carries flushedLSN; shipper updates via monotonic CAS
- sync_all gates on ALL barriers succeeding (fail-closed)
8 baseline tests promoted to CP13-3 primary proof:
- BarrierUsesFlushedLSN, FlushedLSNMonotonicWithinEpoch
- FlushedLSN_OnlyAfterSync, FlushedLSN_NotOnReceive
- ShipperReplicaFlushedLSN_UpdatedOnBarrier, _Monotonic
- BarrierResp_FlushedLSN_Roundtrip, BackwardCompat_1Byte
6 tests classified as support evidence (not primary proof).
Reconnect/retention/rebuild tests explicitly out of scope (CP13-4+).
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Two fixes:
1. Rename advertisedIP → advertisedHost throughout, relax contract from
"always a real IP" to "routable host from -ip flag (IP or resolvable
hostname)". This matches the actual -ip flag semantics which accepts
both IP addresses and server names.
2. Add TestCP13_2_BlockService_AdvertisedHost_NotOpaqueID that hits the
actual production wiring: BlockService with opaque localServerID +
routable advertisedHost → setupReplicaReceiver → verify exported
addresses use the routable host, not the opaque ID.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Bug: setupReplicaReceiver derived the advertised host from localServerID,
which can be an opaque string (from -id flag, e.g., "my-custom-server-id").
This would publish unusable endpoints like "my-custom-server-id:14260".
Fix:
- volume_server_block.go: add advertisedIP field (always a real IP from
-ip flag), use it instead of localServerID for replica canonicalization
- volume.go: wire *v.ip → blockService.SetAdvertisedIP() at startup
- blockvol.go: StartReplicaReceiver variadic advertisedHost unchanged
Proof (sync_all_bug_test.go TestBug3, 4 sub-cases):
- fallback: wildcard bind without advertisedHost → outbound-IP
- advertisedHost: explicit IP appears in exported addresses
- StartReplicaReceiver_API: public API forwards host correctly
- opaque_identity_not_routable: proves opaque string produces
non-routable address, confirming production must use advertisedIP
Identity vs transport separation preserved:
- localServerID: stable identity for V2 control (may be opaque)
- advertisedIP: routable IP for transport endpoints (always real IP)
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Problem: StartReplicaReceiver didn't forward advertisedHost to
NewReplicaReceiver, so wildcard-bind listeners relied on outbound-IP
fallback for canonicalization. On multi-NIC hosts this could select
the wrong interface, leaking non-routable addresses into replication
truth.
Fix:
- blockvol.go: StartReplicaReceiver now accepts optional advertisedHost
variadic param and forwards it to NewReplicaReceiver
- volume_server_block.go: setupReplicaReceiver extracts host from
localServerID (the canonical VS identity) and passes it as
advertisedHost — wildcard-bind addresses now resolve to the
authoritative server IP, not outbound-IP fallback
Proof (sync_all_bug_test.go TestBug3, upgraded from PASS* to PASS):
- fallback: wildcard bind without advertisedHost still produces ip:port
- advertisedHost: explicit host appears in exported DataAddr/CtrlAddr
- StartReplicaReceiver_API: public API forwards advertisedHost correctly
What CP13-2 does NOT change:
- No reconnect handshake changes (CP13-5)
- No retention policy changes (CP13-6)
- No rebuild behavior changes (CP13-7)
- No barrier protocol changes (CP13-3)
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Change "CP13-3/4/5/6 behavior already implemented in earlier phases" to
"current code already passes tests associated with later checkpoint themes"
— baseline evidence only, not implementation closure.
No .go files changed in CP13-1. All 44 baseline tests already existed.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
- phase-13-log.md: mark pre-baseline inventory table as superseded,
point to phase-13-cp1-baseline.md for authoritative results
- phase-13-cp1-baseline.md: replace "CP13-X done" language with neutral
"current code passes this test; suggests behavior may already exist"
— checkpoint closure still requires dedicated review
- Expand remaining-open-checkpoints section: CP13-2/5/6/7 all still
require review, main fails cluster around CP13-5 but CP13-7 and
part of CP13-6 also remain open
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Copies of design docs removed in Phase 09, preserved in sw-block/docs/archive/
for historical reference.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
HIGH: Changed-address now requires OutcomeCatchUp and fails if not.
No more conditional execution — must go through full catch-up chain.
MED: Overlapping retention is now true simultaneous overlap:
- Hold 1 at LSN T+1, Hold 2 at LSN T+2 — both coexist
- MinWALRetentionFloor = T+1 (minimum of two)
- Release hold 1 → floor moves to T+2
- Release hold 2 → ActiveHoldCount=0, no floor
MED: NeedsRebuild now asserts escalated event in logs.
PostCheckpoint now asserts handshake + catch-up execution events.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
HIGH: renamed TestP2_RebuildClosure_FullBase_OneChain → TestP2_RebuildClosure_OneChain.
Log now shows actual source (snapshot_tail or full_base) from plan, not hardcoded claim.
MED: catch-up test uses t.Skipf when V1 interim prevents OutcomeCatchUp.
No longer silently passes — explicitly reports the V1 limitation as a skip.
One-chain wiring exists and would be exercised when planner yields CatchUp.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Engine executors now have IO interfaces for real bridge I/O:
- CatchUpExecutor.IO (CatchUpIO): StreamWALEntries
- RebuildExecutor.IO (RebuildIO): TransferFullBase, TransferSnapshot,
StreamWALEntries (for tail replay)
When IO is set, executor calls real bridge I/O during execution.
When IO is nil, executor uses caller-supplied progress (test mode).
RecoveryPlan.CatchUpStartLSN: bound at plan time for IO bridge.
v2bridge.Executor now implements both interfaces:
- StreamWALEntries: real ScanFrom
- TransferFullBase: validates extent accessible
- TransferSnapshot: validates checkpoint accessible
Chain tests wire IO:
- CatchUpClosure: exec.IO = executor → real WAL scan through engine
- RebuildClosure: exec.IO = executor → real transfer through engine
This closes the engine → executor → v2bridge → blockvol chain.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Finding 1: ProcessAssignments now calls v2Orchestrator.ProcessAssignment
- BlockService.v2Orchestrator field (RecoveryOrchestrator)
- ProcessAssignment result logged at glog V(1)
- No more `_ = intent` — engine state actually changes
Finding 2: localServerID documented as interim
- BlockService.localServerID = listenAddr (transport-shaped)
- Field doc explicitly states: INTERIM, should be registry-assigned
- Used only for replica/rebuild local identity
3 integration tests (qa_block_v2bridge_test.go):
- CreatesEngineSender: ProcessAssignment → engine has sender + session
- EpochBump: epoch 1 → invalidate → epoch 2 → new session
- AddressChange: same ServerID, different IP → sender preserved,
endpoint updated
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Finding 1: Identity no longer address-derived
- ReplicaAddr.ServerID field added (stable server identity from registry)
- BlockVolumeAssignment.ReplicaServerID field added (scalar RF=2 path)
- ControlBridge uses ServerID, NOT address, for ReplicaID
- Missing ServerID → replica skipped (fail closed), logged
Finding 2: Wired into real ProcessAssignments
- BlockService.v2Bridge field initialized in StartBlockService
- ProcessAssignments converts each assignment via v2Bridge.ConvertAssignment
BEFORE existing V1 processing (parallel, not replacing yet)
- Logged at glog V(1)
Finding 3: Fail-closed on missing identity
- Empty ServerID in ReplicaAddrs → replica skipped with log
- Empty ReplicaServerID in scalar path → no replica created
- Test: MissingServerID_FailsClosed verifies both paths
7 tests: StableServerID, AddressChange_IdentityPreserved,
MultiReplica_StableServerIDs, MissingServerID_FailsClosed,
EpochFencing_IntegratedPath, RebuildAssignment, ReplicaAssignment
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
ControlBridge converts real BlockVolumeAssignment (from master heartbeat)
into V2 engine AssignmentIntent:
- Identity: ReplicaID = <volume-path>/<replica-server-id>
- Epoch from real assignment
- Role → SessionKind mapping (primary/replica/rebuilding)
- Multi-replica support (ReplicaAddrs) with scalar RF=2 fallback
Known limitation (documented in test):
- extractServerID currently uses address as server ID (matches
master registry ReplicaInfo.Server format)
- IP change = different server ID in current model
- Registry-backed stable server ID deferred
6 new tests:
- PrimaryAssignment_StableIdentity: real assignment → stable ID
- PrimaryAssignment_MultiReplica: RF=3 multi-replica mapping
- AddressChange_SameServerID: documents current identity boundary
- EpochFencing_IntegratedPath: epoch 1 → bump → epoch 2 through
real assignment conversion + engine
- RebuildAssignment: rebuilding role → SessionRebuild
- ReplicaAssignment: replica role with local server ID
Delivery template:
Changed contracts: real BlockVolumeAssignment → engine intent
Fail-closed: unknown role returns empty intent
Carry-forward: address-based server ID, not registry-backed
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
FC1: now asserts HasActiveSession() after address change AND
verifies session_created in log (not just plan_cancelled).
FC4: escalation event detail must be >15 chars (contains proof
reason with LSN values, not just "needs_rebuild").
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
P2 tests now force conditions instead of observing them:
FC3: Real WAL scan verified directly — StreamWALEntries transfers
real entries from disk (head=5, transferred=5). Engine planning also
verified (ZeroGap in V1 interim documented).
FC4: ForceFlush advances checkpoint/tail to 20. Replica at 0 is
below tail → NeedsRebuild with proof: "gap_beyond_retention: need
LSN 1 but tail=20". No early return.
FC5: ForceFlush advances checkpoint to 10. Assertive:
- replica at checkpoint=10 → ZeroGap (V1 interim)
- replica at 0 → NeedsRebuild (below tail, not CatchUp)
FC1/FC2: Labeled as integrated engine/storage (control simulated).
New: BlockVol.ForceFlush() — triggers synchronous flusher cycle for
test use. Advances checkpoint + WAL tail deterministically.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
5 failure-class replay tests against real file-backed BlockVol,
exercising the full integrated path:
bridge adapter → v2bridge reader/pinner → engine planner/executor
FC1: Changed-address restart — identity preserved, old plan cancelled,
new session created. Log shows plan_cancelled + session_created.
FC2: Stale epoch after failover — sessions invalidated at old epoch,
new assignment at epoch 2 creates fresh session. Log shows
per-replica invalidation.
FC3: Real catch-up (pre-checkpoint) — engine classifies from real
RetainedHistory, zero-gap in V1 interim (committed=0 before flush).
Documents the V1 limitation explicitly.
FC4: Unrecoverable gap — after flush, if checkpoint advances, replica
behind tail gets NeedsRebuild. Documents that V1 unit test may
not advance checkpoint (flusher timing).
FC5: Post-checkpoint boundary — replica at checkpoint = zero-gap in
V1 interim. Explicitly documents the catch-up collapse boundary.
go.mod: added replace directives for sw-block engine + bridge modules.
Carry-forward (explicit):
- CommittedLSN = CheckpointLSN (V1 interim)
- FC3/FC4/FC5 limited by flusher not advancing checkpoint in unit tests
- Executor snapshot/full-base/truncate still stubs
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
7 tests in weed/storage/blockvol/v2bridge/bridge_test.go:
Reader (2 tests):
- StatusSnapshot reads real nextLSN, WALCheckpointLSN, flusher state
- HeadLSN advances with real writes
Pinner (2 tests):
- HoldWALRetention: hold tracked, MinWALRetentionFloor reports position,
release clears hold
- HoldRejectsRecycled: validates against real WAL tail
Executor (2 tests):
- StreamWALEntries: real ScanFrom reads WAL entries from disk
- StreamPartialRange: partial range scan works
Stubs (1 test):
- TransferSnapshot/TransferFullBase/TruncateWAL return not-implemented
All tests use createTestVol (1MB file-backed BlockVol with 256KB WAL).
No mock/push adapters — direct real blockvol instances.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Finding 1: WALTailLSN semantic fix
- StatusSnapshot().WALTailLSN now reads super.WALCheckpointLSN (an LSN)
- Was: wal.Tail() which returns a physical byte offset
- Entries with LSN > WALTailLSN are guaranteed in the WAL
Finding 2: ScanWALEntries replay-source fix
- ScanWALEntries passes super.WALCheckpointLSN as the recycled boundary
- Was: flusher.CheckpointLSN() which in V1 equals CommittedLSN
- The flusher's live checkpoint may advance in memory, but entries above
the durable superblock checkpoint are still physically in the WAL
- Normal catch-up (replica at 70, committed at 100) now works because
fromLSN=71 > super.WALCheckpointLSN (which is the last persisted
checkpoint, not the live flusher state)
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Pinner (pinner.go):
- HoldWALRetention: validates startLSN >= current tail, tracks hold
- HoldSnapshot: validates checkpoint exists + trusted
- HoldFullBase: tracks hold by ID
- MinWALRetentionFloor: returns minimum held position across all
WAL/snapshot holds — designed for flusher RetentionFloorFn hookup
- Release functions remove holds from tracking map
Executor (executor.go):
- StreamWALEntries: validates range against real WAL tail/head
(actual ScanFrom integration deferred to network-layer wiring)
- TransferSnapshot/TransferFullBase/TruncateWAL: stubs for P1
Key integration points:
- Pinner reads real StatusSnapshot for validation
- Pinner.MinWALRetentionFloor can wire into flusher.RetentionFloorFn
- Executor validates WAL range availability from real state
Carry-forward:
- Real ScanFrom wiring needs WAL fd + offset (network layer)
- TransferSnapshot/TransferFullBase need extent I/O
- Control intent from confirmed failover (master-side)
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
CatchUpExecutor.OnStep: optional callback fired between executor-managed
progress steps. Enables deterministic fault injection (epoch bump)
between steps without racing or manual sender calls.
E2_EpochBump_MidExecutorLoop:
- Executor runs 5 progress steps
- OnStep hook bumps epoch after step 1 (after 2 successful steps)
- Executor's own loop detects invalidation at step 2's check
- Resources released by executor's release path (not manual cancel)
- Log shows session_invalidated + exec_resources_released
This closes the remaining FC2 gap: invalidation is now detected
and cleaned up by the executor itself, not by external code.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Planner/executor contract:
- RebuildExecutor.Execute() takes no arguments — consumes plan-bound
RebuildSource, RebuildSnapshotLSN, RebuildTargetLSN
- RecoveryPlan binds all rebuild targets at plan time
- Executor cannot re-derive policy from caller-supplied history
Catch-up timing:
- Removed unused completeTick parameter from CatchUpExecutor.Execute
- Per-step ticks synthesized as startTick + stepIndex + 1
- API shape matches implementation
New test: PlanExecuteConsistency_RebuildCannotSwitchSource
- Plans snapshot+tail, then mutates storage history
- Executor succeeds using plan-bound values (not re-derived)
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Full-base rebuild resource:
- StorageAdapter.PinFullBase/ReleaseFullBase for full-extent base image
- PlanRebuild full_base branch now acquires FullBasePin
- RecoveryPlan.FullBasePin field, released by ReleasePlan
Session cleanup on resource failure:
- PlanRecovery invalidates session when WAL pin fails
(no dangling live session after failed resource acquisition)
3 new tests:
- PlanRebuild_FullBase_PinsBaseImage: pin acquired + released
- PlanRebuild_FullBase_PinFailure: logged + error
- PlanRecovery_WALPinFailure_CleansUpSession: session invalidated,
sender disconnected (no dangling state)
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
ProcessAssignment now compares pre/post endpoint state before
logging session_invalidated with "endpoint_changed" reason.
Normal session supersede (same endpoint, assignment_intent) no
longer mislabeled as endpoint change.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Zero-gap completion:
- ExecuteRecovery auto-completes zero-gap sessions (no sender call needed)
- RecoveryResult.FinalState = StateInSync for zero-gap
Epoch transition:
- UpdateSenderEpoch: orchestrator-owned epoch advancement with auto-log
- InvalidateEpoch: per-replica session_invalidated events (not aggregate)
Endpoint-change invalidation:
- ProcessAssignment detects session ID change from endpoint update
- Logs per-replica session_invalidated with "endpoint_changed" reason
All integration tests now use orchestrator exclusively for core lifecycle.
No direct sender API calls for recovery execution in integration tests.
1 new test: EndpointChange_LogsInvalidation
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
RecordHandshakeFromHistory and SelectRebuildFromHistory now
return an error instead of panicking on nil history input.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
New file: history.go — RetainedHistory connects recovery decisions
to actual WAL retention state:
- IsRecoverable: checks gap against tail/head boundaries
- MakeHandshakeResult: generates HandshakeResult from retention state
- RebuildSourceDecision: chooses snapshot+tail vs full base from
checkpoint state (trusted vs untrusted)
- ProveRecoverability: generates explicit proof explaining why
recovery is or is not allowed
14 new tests (recoverability_test.go):
- Recoverable/unrecoverable gap (exact boundary, beyond head)
- Trusted/untrusted/no checkpoint → rebuild source selection
- Handshake from retained history → outcome classification
- Recoverability proofs (zero-gap, ahead, within retention, beyond)
- E2E: two replicas driven by retained history (catch-up + rebuild)
- Truncation required for replica ahead of committed
Engine module at 44 tests (12 + 18 + 14).
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Entry counting:
- Session.setRange now initializes recoveredTo = startLSN
- RecordCatchUpProgress delta counts only actual catch-up work
(recoveredTo - startLSN), not the replica's pre-existing prefix
Rebuild transfer gate:
- BeginTailReplay requires TransferredTo >= SnapshotLSN
- Prevents tail replay on incomplete base transfer
3 new regression tests:
- BudgetEntries_NonZeroStart_CountsOnlyDelta (30 entries within 50 budget)
- BudgetEntries_NonZeroStart_ExceedsBudget (30 entries exceeds 20 budget)
- Rebuild_PartialTransfer_BlocksTailReplay
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Registry is now keyed by stable ReplicaID, not by address.
DataAddr changes preserve sender identity — the core V2 invariant.
Changes:
- ReplicaAssignment{ReplicaID, Endpoint} replaces map[string]Endpoint
- AssignmentIntent.Replicas uses []ReplicaAssignment
- Registry.Reconcile takes []ReplicaAssignment
- Tests use stable IDs ("replica-1", "r1") independent of addresses
New test: ChangedDataAddr_PreservesSenderIdentity
- Same ReplicaID, different DataAddr (10.0.0.1 → 10.0.0.2)
- Sender pointer preserved, session invalidated, new session attached
- This is the exact V1/V1.5 regression that V2 must fix
doc.go: clarified Slice 1 core vs carried-forward files
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
All mutable state on Sender and Session is now unexported:
- Sender.state, .epoch, .endpoint, .session, .stopped → accessors
- Session.id, .phase, .kind, etc. → read-only accessors
- Session() replaced by SessionSnapshot() (returns disconnected copy)
- SessionID() and HasActiveSession() for common queries
- AttachSession returns (sessionID, error) not (*Session, error)
- SupersedeSession returns sessionID not *Session
Budget configuration via SessionOption:
- WithBudget(CatchUpBudget) passed to AttachSession
- No direct field mutation on session from external code
New test: Encapsulation_SnapshotIsReadOnly proves snapshot
mutation does not leak back to sender state.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Frozen target is now unconditional:
- FrozenTargetLSN field on RecoverySession, set by BeginCatchUp
- RecordCatchUpProgress enforces FrozenTargetLSN regardless of Budget
- Catch-up is always a bounded (R, H0] contract
Rebuild completion exclusivity:
- CompleteSessionByID explicitly rejects SessionRebuild by kind
- Rebuild sessions can ONLY complete via CompleteRebuild
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
IsRecoverable now verifies three conditions:
- startExclusive >= tailLSN (not recycled)
- endInclusive <= headLSN (within WAL)
- all LSNs in range exist contiguously (no holes)
StateAt now uses base snapshot captured during AdvanceTail:
- returns nil for LSNs before snapshot boundary (unreconstructable)
- correctly includes block state from recycled entries via snapshot
5 new tests: end-beyond-head, missing entries, state after tail
advance, nil before snapshot, block last written before tail.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
AllocateBlockVolumeResponse used bs.ListenAddr() to derive replica
addresses. When the VS binds to ":port" (no explicit IP), host
resolved to empty string, producing ":dataPort" as the replica
address. This ":port" propagated through master assignments to both
primary and replica sides.
Now canonicalizes empty/wildcard host using PreferredOutboundIP()
before constructing replication addresses. Also exported
PreferredOutboundIP for use by the server package.
This is the source fix — all downstream paths (heartbeat, API
response, assignment) inherit the canonical address.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
setupReplicaReceiver now reads back canonical addresses from
the ReplicaReceiver (which applies CP13-2 canonicalization)
instead of storing raw assignment addresses in replStates.
This fixes the API-level leak where replica_data_addr showed
":port" instead of "ip:port" in /block/volumes responses,
even though the engine-level CP13-2 fix was working.
New BlockVol.ReplicaReceiverAddr() returns canonical addresses
from the running receiver. Falls back to assignment addresses
if receiver didn't report.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
rebuildFullExtent updated superblock.WALCheckpointLSN but not the
flusher's internal checkpointLSN. NewReplicaReceiver then read
stale 0 from flusher.CheckpointLSN(), causing post-rebuild
flushedLSN to be wrong.
Added Flusher.SetCheckpointLSN() and call it after rebuild
superblock persist. TestRebuild_PostRebuild_FlushedLSN_IsCheckpoint
flips FAIL→PASS.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
The test used createSyncAllPair(t) but discarded the replica
return value, leaving the volume file open. On Windows this
caused TempDir cleanup failure. All 7 CP13-1 baseline FAILs
now PASS.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Adds per-replica state reporting in heartbeat so master can identify
which specific replica needs rebuild, not just a volume-level boolean.
New ReplicaShipperStatus{DataAddr, State, FlushedLSN} type reported
via ReplicaShipperStates field on BlockVolumeInfoMessage. Populated
from ShipperGroup.ShipperStates() on each heartbeat. Scales to RF=3+.
V1 constraints (explicit):
- NeedsRebuild cleared only by control-plane reassignment (no local exit)
- Post-rebuild replica re-enters as Disconnected/bootstrap, not InSync
- flushedLSN = checkpointLSN after rebuild (durable baseline only)
4 new tests: heartbeat per-replica state, NeedsRebuild reporting,
rebuild-complete-reenters-InSync (full cycle), epoch mismatch abort.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Flusher now holds WAL entries needed by recoverable replicas.
Both AdvanceTail (physical space) and checkpointLSN (scan gate)
are gated by the minimum flushed LSN across catch-up-eligible
replicas.
New methods on ShipperGroup:
- MinRecoverableFlushedLSN() (uint64, bool): pure read, returns
min flushed LSN across InSync/Degraded/Disconnected/CatchingUp
replicas with known progress. Excludes NeedsRebuild.
- EvaluateRetentionBudgets(timeout): separate mutation step,
escalates replicas that exceed walRetentionTimeout (5m default)
to NeedsRebuild, releasing their WAL hold.
Flusher integration: evaluates budgets then queries floor on each
flush cycle. If floor < maxLSN, holds both checkpoint and tail.
Extent writes proceed normally (reads work), only WAL reclaim
is deferred.
LastContactTime on WALShipper: updated on barrier success,
handshake success, and catch-up completion. Not on Ship (TCP
write only). Avoids misclassifying idle-but-healthy replicas.
CP13-6 ships with timeout budget only. walRetentionMaxBytes
is deferred (documented as partial slice).
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Concurrent WriteLBA/Trim calls could deliver WAL entries to replicas
out of LSN order: two goroutines allocate LSN 4 and 5 concurrently,
but LSN 5 could reach the replica first via ShipAll, causing the
replica to reject it as an LSN gap.
shipMu now wraps nextLSN.Add + wal.Append + ShipAll in both
WriteLBA and Trim, guaranteeing LSN-ordered delivery to replicas
under concurrent writers.
The dirty map update and WAL pressure check happen after shipMu
is released — they don't need ordering guarantees.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
doReconnectAndCatchUp() now uses the replicaFlushedLSN returned by
the reconnect handshake as the catch-up start point, not the
shipper's stale cached value. The replica may have less durable
progress than the shipper last knew.
ReplicaReceiver initialization: flushedLSN now set from the
volume's checkpoint LSN (durable by definition), not nextLSN
(which includes unflushed entries). receivedLSN still uses
nextLSN-1 since those entries are in the WAL buffer even if
not yet synced.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Updated 3 reconnect tests to stop/restart the ReplicaReceiver on
the same addresses WITHOUT calling SetReplicaAddr. This preserves
the shipper object, its ReplicaFlushedLSN, HasFlushedProgress flag,
and catch-up state across the disconnect/reconnect cycle.
All 3 tests now PASS:
- TestReconnect_CatchupFromRetainedWal
- CatchupReplay_DataIntegrity_AllBlocksMatch
- CatchupReplay_DuplicateEntry_Idempotent
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Adds the sync_all reconnect protocol: when a degraded shipper
reconnects, it performs a handshake (ResumeShipReq/Resp) to
determine the replica's durable progress, then streams missed
WAL entries to close the gap before resuming live shipping.
New wire messages:
- MsgResumeShipReq (0x03): primary sends epoch, headLSN, retainStart
- MsgResumeShipResp (0x04): replica returns status + flushedLSN
- MsgCatchupDone (0x05): marks end of catch-up stream
Decision matrix after handshake:
- R == H: already caught up → InSync
- S <= R+1 <= H: recoverable gap → CatchingUp → stream → InSync
- R+1 < S: gap exceeds retained WAL → NeedsRebuild
- R > H: impossible progress → NeedsRebuild
WALAccess interface: narrow abstraction (RetainedRange + StreamEntries)
avoids coupling shipper to raw WAL internals.
Bootstrap vs reconnect split: fresh shippers (HasFlushedProgress=false)
use CP13-4 bootstrap path. Previously-synced shippers use handshake.
Catch-up retry budget: maxCatchupRetries=3 before NeedsRebuild.
ReplicaReceiver now initializes receivedLSN/flushedLSN from volume's
nextLSN on construction (handles receiver restart on existing volume).
TestBug2_SyncAll_SyncCache_AfterDegradedShipperRecovers flips FAIL→PASS.
All previously-passing baseline tests remain green.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Replaces binary degraded flag with ReplicaState type:
Disconnected, Connecting, CatchingUp, InSync, Degraded, NeedsRebuild.
Ship() allowed from Disconnected (bootstrap: data must flow before
first barrier) and InSync (steady state). Ship does NOT change state.
Barrier() gating:
- InSync: proceed normally
- Disconnected: bootstrap path (connect + barrier)
- Degraded: reconnect both data+ctrl connections, then barrier
- Connecting/CatchingUp/NeedsRebuild: rejected immediately
Only barrier success grants InSync. Reconnect alone does not.
IsDegraded() now means "not sync-eligible" (any non-InSync state).
InSyncCount() added to ShipperGroup.
dist_group_commit.go: removed AllDegraded short-circuit that
prevented bootstrap. Barrier attempts always run — individual
shippers handle their own state-based gating.
8 CP13-4 tests + TestBarrier_RejectsReplicaNotInSync flips FAIL→PASS.
All previously-passing baseline tests remain green.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Barrier response extended from 1-byte status to 9-byte payload
carrying the replica's durable WAL progress (FlushedLSN). Updated
only after successful fd.Sync(), never on receive/append/send.
Replica side: new flushedLSN field on ReplicaReceiver, advanced
only in handleBarrier after proven contiguous receipt + sync.
max() guard prevents regression.
Shipper side: new replicaFlushedLSN (authoritative) replacing
ShippedLSN (diagnostic only). Monotonic CAS update from barrier
response. hasFlushedProgress flag tracks whether replica supports
the extended protocol.
ShipperGroup: MinReplicaFlushedLSN() returns (uint64, bool) —
minimum across shippers with known progress. (0, false) for empty
groups or legacy replicas.
Backward compat: 1-byte legacy responses decoded as FlushedLSN=0.
Legacy replicas explicitly excluded from sync_all correctness.
7 new tests: roundtrip, backward compat, flush-only-after-sync,
not-on-receive, shipper update, monotonicity, group minimum.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
ReplicaReceiver.DataAddr()/CtrlAddr() now return canonical ip:port
instead of raw listener addresses that may be wildcard (:port,
0.0.0.0:port, [::]:port).
New canonicalizeListenerAddr() resolves wildcard IPs using the
provided advertised host (from VS listen address). Falls back to
outbound-IP detection when no advertised host is available.
NewReplicaReceiver accepts optional advertisedHost parameter for
multi-NIC correctness. In production, the assignment path already
provides canonical addresses; this fix ensures test patterns with
:0 bind also produce routable addresses.
7 new tests. TestBug3_ReplicaAddr_MustBeIPPort_WildcardBind flips
from FAIL to PASS.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Same-epoch reconciliation now trusts reported roles first:
- one claims primary, other replica → trust roles
- both claim primary → WALHeadLSN heuristic tiebreak
- both claim replica → keep existing, log ambiguity
Replaced addServerAsReplica with upsertServerAsReplica: checks
for existing replica entry by server name before appending.
Prevents duplicate ReplicaInfo rows during restart/replay windows.
2 new tests: role-trusted same-epoch, duplicate replica prevention.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
When a second server reports the same volume during master restart,
UpdateFullHeartbeat now uses epoch-based tie-breaking instead of
first-heartbeat-wins:
1. Higher epoch wins as primary — old entry demoted to replica
2. Same epoch — higher WALHeadLSN wins (heuristic, warning logged)
3. Lower epoch — added as replica
Applied in both code paths: the auto-register branch (no entry
exists yet for this name) and the unlinked-server branch (entry
exists but this server is not in it).
This is a deterministic reconstruction improvement, not ground
truth. The long-term fix is persisting authoritative volume state.
5 new tests covering all reconciliation scenarios.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Lookup() and ListAll() now return value copies (not pointers to
internal registry state). Callers can no longer mutate registry
entries without holding a lock.
Added clone() on BlockVolumeEntry with deep-copied Replicas slice.
Added UpdateEntry(name, func(*BlockVolumeEntry)) for locked mutation.
ListByServer() also returns copies.
Migrated 1 production mutation (ReplicaPlacement + Preset in create
handler) and ~20 test mutations to use UpdateEntry.
5 new copy-correctness tests: Lookup returns copy, Replicas slice
isolated, ListAll returns copies, UpdateEntry mutates, UpdateEntry
not-found error.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
superMu is mandatory for correctness — all superblock mutation+persist
must be serialized. Remove the nil guard in updateSuperblockCheckpoint
and add SuperMu to all 7 test FlusherConfig sites.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Adds sync.Mutex (superMu) to BlockVol, shared between group commit's
syncWithWALProgress() and flusher's updateSuperblockCheckpoint().
Both paths now serialize superblock mutation + persist, preventing
WALTail/WALCheckpointLSN regression when flusher and group commit
write the full superblock concurrently.
persistSuperblock() also guarded for consistency.
Removes temporary log.Printf lines in the open/recovery path that
were added during BUG-RESTART-ZEROS investigation.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Adds sync.RWMutex (ioMu) to BlockVol enforcing mutual exclusion
between normal I/O and destructive state operations.
Shared (RLock): WriteLBA, ReadLBA, Trim, SyncCache, replica
applyEntry, rebuild applyRebuildEntry — concurrent I/O safe.
Exclusive (Lock): RestoreSnapshot, ImportSnapshot, Expand,
PrepareExpand, CommitExpand, CancelExpand — drains all in-flight
I/O before modifying extent/WAL/dirtyMap.
Scope rule: RLock covers local data-structure mutation only.
Replication shipping is asynchronous and outside the lock, so
exclusive holders block only behind local I/O, not network stalls.
Lock ordering: ioMu > snapMu > assignMu > mu.
Closes the critical ER item: restore/import vs concurrent WriteLBA
silent data corruption gap.
3 new tests: concurrent writes allowed, real restore-vs-write
contention with data integrity check, close coordination.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
New POST /block/volume/plan endpoint returns full placement preview:
resolved policy, ordered candidate list, selected primary/replicas,
and per-server rejection reasons with stable string constants.
Core design: evaluateBlockPlacement() is a pure function with no
registry/topology dependency. gatherPlacementCandidates() is the
single topology bridge point. Plan and create share the same planner —
parity contract is same ordered candidate list for same cluster state.
Create path refactored: uses evaluateBlockPlacement() instead of
PickServer(), iterates all candidates (no 3-retry cap), recomputes
replica order after primary fallback. rf_not_satisfiable severity
is durability-mode-aware (warning for best_effort, error for strict).
15 unit tests + 20 QA adversarial tests.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Preset system: ResolvePolicy resolves named presets (database, general,
throughput) with per-field overrides into concrete volume parameters.
Create path now uses resolved policy instead of ad-hoc validation.
New /block/volume/resolve diagnostic endpoint for dry-run resolution.
Review fix 1 (MED): HasNVMeCapableServer now derives NVMe capability
from server-level heartbeat attribute (block_nvme_addr proto field)
instead of scanning volume entries. Fixes false "no NVMe" warning on
fresh clusters with NVMe-capable servers but no volumes yet.
Review fix 2 (LOW): /block/volume/resolve no longer proxied to leader —
read-only diagnostic endpoint can be served by any master.
Engine fix: ReadLBA retry loop closes stale dirty-map race when WAL
entry is recycled between lookup and read.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Six-task checkpoint hardening the promotion and failover paths:
T1: 4-gate candidate evaluation (heartbeat freshness, WAL lag, role,
server liveness) with structured rejection reasons.
T2: Orphaned-primary re-evaluation on replica reconnect (B-06/B-08).
T3: Deferred timer safety — epoch validation prevents stale timers
from firing on recreated/changed volumes (B-07).
T4: Rebuild addr cleanup on promotion (B-11), NVMe publication
refresh on heartbeat, and preflight endpoint wiring.
T5: Manual promote API — POST /block/volume/{name}/promote with
force flag, target server selection, and structured rejection
response. Shared applyPromotionLocked/finalizePromotion helpers
eliminate duplication between auto and manual paths.
T6: Read-only preflight endpoint (GET /block/volume/{name}/preflight)
and blockapi client wrappers (Preflight, Promote).
BUG-T5-1: PromotionsTotal counter moved to finalizePromotion (shared
by both auto and manual paths) to prevent metrics divergence.
24 files changed, ~6500 lines added. 42 new QA adversarial tests.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
BUG-CP11A4-1 (HIGH): ImportSnapshot now rejects when active snapshots
exist. Import overwrites the extent region that non-CoW'd snapshot blocks
read from, which would silently return import data instead of snapshot-time
data. New ErrImportActiveSnapshots error and snapMu-guarded check.
BUG-CP11A4-2 (HIGH): Double import without AllowOverwrite now correctly
rejected. Import bypasses WAL so nextLSN stays at 1; added FlagImported
(Superblock.Flags bit 0) set after successful import and checked alongside
nextLSN in the non-empty gate.
BUG-CP11A4-3 (MED): Replaced fixed exportTempSnapID (0xFFFFFFFE) with
atomic sequence counter (exportTempSnapBase + exportTempSnapSeq). Each
auto-export gets a unique temp snapshot ID, preventing concurrent export
races and user snapshot ID collisions.
Also added beginOp()/endOp() lifecycle guards to both ExportSnapshot and
ImportSnapshot, and documented the non-atomic import failure semantics.
5 new regression tests + QA-EX-3 rewritten for rejection behavior.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Add PressureState() and writer wait tracking to WALAdmission, WALStatus
snapshot API on BlockVol, WAL sizing guidance pure functions, Prometheus
histogram/gauge/counter exports, and admin /status WAL fields. 23 new
tests (7 admission, 10 guidance, 6 QA adversarial).
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
B-09: ExpandBlockVolume re-reads the registry entry after acquiring
the expand inflight lock. Previously it used the entry from the
initial Lookup, which could be stale if failover changed VolumeServer
or Replicas between Lookup and PREPARE.
B-10: UpdateFullHeartbeat stale-cleanup now skips entries with
ExpandInProgress=true. Previously a primary VS restart during
coordinated expand would delete the entry (path not in heartbeat),
orphaning the volume and stranding the expand coordinator.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Two-phase prepare/commit/cancel protocol ensures all replicas expand
atomically. Standalone volumes use direct-commit (unchanged behavior).
Engine: PrepareExpand/CommitExpand/CancelExpand with on-disk
PreparedSize+ExpandEpoch in superblock, crash recovery clears stale
prepare state on open, v.mu serializes concurrent expand operations.
Proto: 3 new RPCs (PrepareExpand/CommitExpand/CancelExpandBlockVolume).
Coordinator: expandClean flag pattern — ReleaseExpandInflight only on
clean success or full cancel. Partial replica commit failure calls
MarkExpandFailed (keeps ExpandInProgress=true, suppresses heartbeat
size updates). ClearExpandFailed for manual reconciliation.
Registry: AcquireExpandInflight records PendingExpandSize+ExpandEpoch.
ExpandFailed state blocks new expands until cleared.
Tests: 15 engine + 4 VS + 10 coordinator + heartbeat suppression
regression + updated QA CP82/durability tests with prepare/commit mocks.
Also includes CP11A-1 remaining: QA storage profile tests, QA
io_backend config tests, testrunner perf-baseline scenarios and
coordinated-expand actions.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
ReplicaInfo now carries NvmeAddr/NQN. Fields are populated during
replica allocation (tryCreateOneReplica), updated from replica
heartbeats, and copied in PromoteBestReplica. This ensures master
lookup returns correct NVMe endpoints immediately after failover,
without waiting for the first post-promotion heartbeat.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Add nvme_addr and nqn fields to proto messages (AllocateBlockVolume,
CreateBlockVolume, LookupBlockVolume, BlockVolumeInfoMessage), wire
through volume server → master registry → CSI driver. Volume servers
report NVMe address in heartbeats when NVMe target is running. CSI
MasterVolumeClient now populates NvmeAddr/NQN from master responses,
enabling NVMe/TCP via the master-backend path.
Proto files regenerated with protoc 29.5.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Finding 1: IOBackend=io_uring was accepted and logged as resolved but
had no runtime effect. Now rejected by Validate() until actually wired,
preventing user confusion.
Finding 2: wal_admit_wait_seconds_total was exported as GaugeFunc but
is monotonically increasing. Changed to CounterFunc to match _total
naming convention.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Add counters (total, soft, hard, timeout) and wait-time histogram to
WALAdmission, wired through EngineMetrics and exported as Prometheus
metrics. Six new tests verify all code paths. Nil-safe for backwards
compatibility.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
All three io_uring backends (iceber, giouring, raw) now require explicit
build tags — no tag means standard-only. Each backend registers its name
via IOUringImpl so startup logs show compiled implementation alongside
requested/selected backend mode.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Split iouring_linux.go into three build-tagged implementations:
1. iouring_iceber_linux.go (-tags iouring_iceber)
iceber/iouring-go library. Goroutine-based completion model.
Known -72% write regression due to per-op channel overhead.
2. iouring_giouring_linux.go (-tags iouring_giouring)
pawelgaczynski/giouring — direct liburing port. No goroutines,
no channels. Direct SQE/CQE ring manipulation. Kernel 6.0+.
3. iouring_raw_linux.go (default on Linux, no tags needed)
Raw syscall wrappers — io_uring_setup/io_uring_enter + mmap.
Zero dependencies. ~300 LOC. Kernel 5.6+.
Build commands for benchmarking:
go build -tags iouring_iceber ./... # option A
go build -tags iouring_giouring ./... # option B
go build ./... # option C (raw, default)
go build -tags no_iouring ./... # disable all io_uring
All variants implement the same BatchIO interface. Cross-compile
verified for all four tag combinations.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
The iceber/iouring-go SubmitRequests returns a RequestSet interface
which cannot be ranged over directly. Use resultSet.Done() to wait
for all completions, then iterate resultSet.Requests().
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Replace UseIOUring bool with IOBackend IOBackendMode (tri-state):
- "standard" (default): sequential pread/pwrite/fdatasync
- "auto": try io_uring, fall back to standard with warning log
- "io_uring": require io_uring, fail startup if unavailable
NewIOUring now returns ErrIOUringUnavailable instead of silently
falling back — callers decide whether to fail or fall back based
on the requested mode. All mode transitions are logged:
io backend: requested=auto selected=standard reason=...
io backend: requested=io_uring selected=io_uring
CLI: --io-backend=standard|auto|io_uring added to iscsi-target.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
1. HIGH: LinkedWriteFsync now uses SubmitLinkRequests (IOSQE_IO_LINK)
instead of SubmitRequests, ensuring write+fdatasync execute as a
linked chain in the kernel. Falls back to sequential on error.
2. HIGH: PreadBatch/PwriteBatch chunk ops by ring capacity to prevent
"too many requests" rejection when dirty map exceeds ring size (256).
3. MED: CloseBatchIO() added to Flusher, called in BlockVol.Close()
after final flush to release io_uring ring / kernel resources.
4. MED: Sync parity — both standard and io_uring paths now use
fdatasync (via platform-specific fdatasync_linux.go / fdatasync_other.go).
Standard path previously used fsync; now matches io_uring semantics.
On non-Linux, fdatasync falls back to fsync (only option available).
10 batchio tests, all blockvol tests pass.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Add iouring_linux.go (build-tagged linux && !no_iouring) using
iceber/iouring-go for batched pread/pwrite/fdatasync. Includes
linked write+fsync chain for group commit optimization.
iouring_other.go provides silent fallback to standard on non-Linux.
blockvol.go wires UseIOUring config flag through to flusher BatchIO.
NewIOUring gracefully falls back if kernel lacks io_uring support.
10 batchio tests, all blockvol tests pass unchanged.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
New package batchio/ with BatchIO interface (PreadBatch, PwriteBatch,
Fsync, LinkedWriteFsync) and standard sequential implementation.
Flusher refactored to use BatchIO: WAL header reads, WAL entry reads,
and extent writes are now batched through the interface. With the
default NewStandard() backend, behavior is identical to before.
UseIOUring config field added for future io_uring opt-in (Linux 5.6+).
9 interface tests, all existing blockvol tests pass unchanged.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Sparse delta-file snapshots with copy-on-write in the flusher.
Zero write-path overhead when no snapshot is active.
New: snapshot.go (SnapshotBitmap, SnapshotHeader, delta file I/O)
Modified: flusher.go (flushMu, CoW phase in FlushOnce, PauseAndFlush)
Modified: blockvol.go (Create/Read/Delete/Restore/ListSnapshots, recovery)
Modified: wal_writer.go (Reset for snapshot restore)
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Add ALUA (Asymmetric Logical Unit Access) support to the iSCSI target,
enabling dm-multipath on Linux to automatically detect path state changes
and reroute I/O during HA failover without initiator-side intervention.
- ALUAProvider interface with implicit ALUA (TPGS=0x01)
- INQUIRY byte 5 TPGS bits, VPD 0x83 with NAA+TPG+RTP descriptors
- REPORT TARGET PORT GROUPS handler (MAINTENANCE IN SA=0x0A)
- MAINTENANCE OUT rejection (implicit-only, no SET TPG)
- Standby write rejection (NOT_READY ASC=04h ASCQ=0Bh)
- RoleNone maps to Active/Optimized (standalone single-node compatibility)
- NAA-6 device identifier derived from volume UUID
- -tpg-id flag with [1,65535] validation
- dm-multipath config + setup script (group_by_tpg, ALUA prio)
- 12 unit tests + 16 QA adversarial tests + 4 integration tests
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Test harness for running blockvol iSCSI tests on WSL2 and remote nodes
(m01/M02). Includes Node (SSH/local exec), ISCSIClient (discover/login/
logout), WeedTarget (weed volume server lifecycle), and test suites for
smoke, stress, crash recovery, chaos, perf benchmarks, and apps (fio/dd).
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Add ProcessBlockVolumeAssignments to BlockVolumeStore and wire
AssignmentSource/AssignmentCallback into the heartbeat collector's
Run() loop. Assignments are fetched and applied each tick after
status collection.
Bug fixes:
- BUG-CP4B3-1: TOCTOU between GetBlockVolume and HandleAssignment.
Added withVolume() helper that holds RLock across lookup+operation,
preventing RemoveBlockVolume from closing the volume mid-assignment.
- BUG-CP4B3-2: Data race on callback fields read by Run() goroutine.
Made StatusCallback/AssignmentSource/AssignmentCallback private,
added cbMu mutex and SetXxx() setter methods. Lock held only for
load/store, not during callback execution.
7 dev tests + 13 QA adversarial tests = 20 new tests.
972 total unit tests, all passing.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
BlockVolumeHeartbeatCollector periodically collects block volume status
via callback (standalone, no gRPC wiring yet). Store() accessor on
BlockService. Three bugs found by QA and fixed: Stop-before-Run deadlock
(BUG-CP4B2-1), zero interval panic (BUG-CP4B2-2), callback panic crashes
goroutine (BUG-CP4B2-3). 12 new tests (3 dev + 9 QA adversarial).
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Boundary tests for RoleFromWire, LeaseTTLToWire overflow/clamp/negative,
ToBlockVolumeInfoMessage with primary/stale/closed/concurrent volumes,
BlockVolumeAssignment roundtrip, and heartbeat collection edge cases.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Add SimulatedMaster test helper + 20 assignment sequence tests (8 sequence,
5 failover, 5 adversarial, 2 status). Add BlockVolumeStatus struct and
Status() method. Includes QA test files for CP1-CP4a. 940 total unit tests.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Add master-driven lifecycle operations: promotion, demotion, rebuild,
and split-brain prevention. All testable on Windows with mock TCP.
New files:
- promotion.go: HandleAssignment (single entry point for role changes),
promote (Replica/None -> Primary with durable epoch), demote
(Primary -> Draining -> Stale with drain timeout)
- rebuild.go: RebuildServer (WAL catch-up + full extent streaming),
StartRebuild client (WAL catch-up with full extent fallback,
two-phase rebuild with second catch-up for concurrent writes)
Modified:
- wal_writer.go: ScanFrom() method, ErrWALRecycled sentinel
- repl_proto.go: rebuild message types + RebuildRequest encode/decode
- blockvol.go: assignMu, drainTimeout, rebuildServer fields;
HandleAssignment/StartRebuildServer/StopRebuildServer methods;
rebuild server stop in Close()
- dirty_map.go: Clear() method for full extent rebuild
32 new tests covering WAL scan, promotion/demotion, rebuild server,
rebuild client, split-brain prevention, and full lifecycle scenarios.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Primary ships WAL entries to replica over TCP (data channel), confirms
durability via barrier RPC (control channel). SyncCache runs local fsync
and replica barrier in parallel via MakeDistributedSync. When replica is
unreachable, shipper enters permanent degraded mode and falls back to
local-only sync (Phase 3 behavior).
Key design: two separate TCP ports (data+control), contiguous LSN
enforcement, epoch equality check, WAL-full retry on replica,
cond.Wait-based barrier with configurable timeout, BarrierFsyncFailed
status code. Close lifecycle: shipper → receiver → drain → committer →
flusher → fd.
New files: repl_proto.go, wal_shipper.go, replica_apply.go,
replica_barrier.go, dist_group_commit.go
Modified: blockvol.go, blockvol_test.go
27 dev tests + 21 QA tests = 48 new tests; 889 total (609 engine + 280
iSCSI), all passing.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
9 categories: PDU, Params, Login, Discovery, SCSI, DataIO, Session,
Target, Integration. 2,183 lines. All 229 tests pass (164 dev + 55 QA).
No new production bugs found.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
1. Discovery session nil handler crash: reject SCSI commands with
Reject PDU when s.scsi is nil (discovery sessions have no target).
2. CmdSN window enforcement: validate incoming CmdSN against
[ExpCmdSN, MaxCmdSN] using serial arithmetic. Drop out-of-window
commands per RFC 7143 section 4.2.2.1.
3. Data-Out buffer offset validation: enforce BufferOffset == received
for ordered data (DataPDUInOrder=Yes). Prevents silent corruption
from out-of-order or overlapping data.
4. ImmediateData enforcement: reject immediate data in SCSI command
PDU when negotiated ImmediateData=No.
5. UNMAP descriptor length alignment: reject blockDescLen not a
multiple of 16 bytes.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
The Linux kernel iSCSI initiator pipelines multiple SCSI commands on
the same TCP connection (command queuing). When a write needs R2T for
data beyond the immediate portion, collectDataOut may read a pipelined
SCSI command instead of the expected Data-Out PDU.
Fix: queue non-Data-Out PDUs received during collectDataOut into a
pending buffer. The main dispatch loop drains pending PDUs before
reading from the connection. This correctly handles interleaved
commands during multi-PDU write transfers.
Bug found during WSL2 smoke test: mkfs.ext4 hangs at "Writing
superblocks" because inode table zeroing sends large writes that
exceed FirstBurstLength, triggering R2T while the kernel has already
queued the next command.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
- Skip InitiatorAlias in negotiation (was returning NotUnderstood)
- Capture TargetName in StageLoginOp direct-jump path (iscsiadm skips
security stage, sends CSG=LoginOp directly -- nil SCSIHandler crash)
- Add portalAddr to TargetServer for discovery responses (listener on
[::] is not routable from WSL2 clients)
- Add -portal flag to iscsi-target binary
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
-`NewCSIDriver(DriverConfig{Mode: "invalid"})` returns nil error. Driver runs with only identity server — no controller, no node. K8s reports capabilities but all operations fail `Unimplemented`.
- Fix: Added `switch` validation after mode defaulting. Returns `"csi: invalid mode %q, must be controller/node/all"`.
- Test: `TestQA_ModeInvalid`.
**Final CP6-2 test count: 118 dev/review + 54 QA = 172 CP6-2 tests, all PASS.**
- **Task 0: Proto Extension + Wire Type Updates** — Added replica_data_addr, replica_ctrl_addr to BlockVolumeInfoMessage/BlockVolumeAssignment; rebuild_addr to BlockVolumeAssignment; replica_server to Create/LookupBlockVolumeResponse; replica fields to AllocateBlockVolumeResponse. Updated wire types and converters. 8 tests.
- **Task 1: Master Assignment Queue + Delivery** — BlockAssignmentQueue with Enqueue/Peek/Confirm/ConfirmFromHeartbeat. Retain-until-confirmed pattern (F1): assignments resent on every heartbeat until VS confirms via matching (path, epoch, role). Stale epoch pruning during Peek. Wired into HeartbeatResponse delivery. 11 tests.
- **Task 2: VS Assignment Receiver Wiring** — VS extracts block_volume_assignments from HeartbeatResponse and calls BlockService.ProcessAssignments.
- **Task 4: Registry Replica Tracking + CreateVolume** — Added SetReplica/ClearReplica/SwapPrimaryReplica to registry. CreateBlockVolume creates on 2 servers (primary + replica), enqueues assignments. Single-copy mode if only 1 server or replica fails (F4). LookupBlockVolume returns ReplicaServer. 10 tests.
- **Task 5: Master Failover Detection** — failoverBlockVolumes on VS disconnect. Lease-aware promotion (F2): promote only after LastLeaseGrant + LeaseTTL expires. Deferred promotion via time.AfterFunc for unexpired leases. promoteReplica swaps primary/replica, bumps epoch, enqueues new primary assignment. 11 tests.
- **Task 6: ControllerPublishVolume/UnpublishVolume** — ControllerPublishVolume calls backend.LookupVolume, returns publish_context{iscsiAddr, iqn}. ControllerUnpublishVolume is no-op. Added PUBLISH_UNPUBLISH_VOLUME capability. NodeStageVolume prefers publish_context over volume_context (reflects current primary after failover). 8 tests.
- **Task 7: Rebuild on Recovery** — recoverBlockVolumes on VS reconnect drains pendingRebuilds, sets reconnected server as replica, enqueues Rebuilding assignments. 10 tests (shared with Task 5 test file).
### Design Review Findings Addressed
| # | Finding | Severity | Resolution |
|---|---------|----------|------------|
| F1 | Assignment delivery can be dropped | Critical | Retain-until-confirmed: Peek+Confirm pattern, assignments resent every heartbeat |
| F2 | Failover without lease check → split-brain | Critical | Gate promotion on `now > lastLeaseGrant + leaseTTL`; deferred promotion for unexpired leases |
| F3 | Replication ports change on VS restart | Critical | Deterministic port = FNV hash of path, offset from base iSCSI port |
| F4 | Partial create (replica fails) | Medium | Single-copy mode with ReplicaServer="", skip replica assignments |
| F5 | UpdateFullHeartbeat ignores replica addresses | Medium | VS includes replica_data/ctrl in InfoMessage; registry updates on heartbeat |
### Code Review 1 Findings Addressed
| # | Finding | Severity | Resolution |
|---|---------|----------|------------|
| R1-1 | AllocateBlockVolume missing repl addrs | High | AllocateBlockVolume now returns ReplicaDataAddr/CtrlAddr/RebuildListenAddr from ReplicationPorts() |
| R1-2 | Primary never starts rebuild server | High | setupPrimaryReplication now calls vol.StartRebuildServer(rebuildAddr) |
| R1-3 | Assignment queue never confirms after startup | High | VS sends periodic full block heartbeat (5×sleepInterval tick) enabling master confirmation |
| R1-4 | Replica addresses not reported in heartbeat | Medium | BlockService.CollectBlockVolumeHeartbeat wraps store's collector, fills ReplicaDataAddr/CtrlAddr from replStates |
| R1-5 | Lease never refreshed after create | Medium | UpdateFullHeartbeat refreshes LastLeaseGrant on every heartbeat; periodic block heartbeats keep it current |
### Code Review 2 Findings Addressed
| # | Finding | Severity | Resolution |
|---|---------|----------|------------|
| R2-F1 | LastLeaseGrant set AFTER Register → stale-lease race | High | Moved to entry initializer BEFORE Register |
| R2-F2 | Deferred promotion timer has no cancellation | Medium | Timers stored in blockFailoverState.deferredTimers; cancelled in recoverBlockVolumes on reconnect |
| R2-F3 | SwapPrimaryReplica hardcodes uint32(1) | Medium | Changed to blockvol.RoleToWire(blockvol.RolePrimary) |
| R2-F4 | DeleteBlockVolume doesn't delete replica | Medium | Added best-effort replica delete (non-fatal if replica VS is down) |
| R2-F5 | promoteReplica reads epoch without lock | Medium | SwapPrimaryReplica now computes epoch+1 atomically inside lock, returns newEpoch |
Purpose: extend the V2 simulator from final-state safety checking into protocol-state simulation that can reproduce `V1`, `V1.5`, and `V2` behavior on the same scenarios
## Goal
Make the simulator model enough node-local replication state and message-level behavior to:
1. reproduce `V1` / `V1.5` failure modes
2. show why those failures are structural
3. close the current `partial` V2 scenarios with stronger protocol assertions
Purpose: define the next simulator tier after Phase 02, focused on timeout semantics, timer races, and a cleaner split between protocol simulation and event/interleaving simulation
## Goal
Phase 03 exists to cover behavior that current `distsim` still abstracts away:
1. timeout semantics
2. timer races
3. event ordering under competing triggers
4. clearer separation between:
- protocol / lineage simulation
- event / race simulation
This phase should not reopen already-closed Phase 02 protocol scope unless a clear bug is found.
## Why A New Phase
Phase 02 already delivered:
- protocol-state assertions
- V1 / V1.5 / V2 comparison scenarios
- endpoint identity modeling
- control-plane assignment-update flow
- committed-prefix-aware promotion eligibility
What remains is different in character:
- timers
- delayed events racing with each other
- timeout-triggered state changes
- more explicit event scheduling
That deserves a new phase boundary.
## Source Of Truth
Design/source-of-truth:
-`sw-block/design/v2_scenarios.md`
-`sw-block/design/v2-dist-fsm.md`
-`sw-block/design/v2-scenario-sources-from-v1.md`
-`sw-block/design/v1-v15-v2-comparison.md`
Current prototype base:
-`sw-block/prototype/distsim/`
-`sw-block/prototype/distsim/simulator.go`
## Scope
### In scope
1. timeout semantics
- barrier timeout
- catch-up timeout
- reservation expiry timeout
- rebuild timeout
2. timer races
- delayed ack vs timeout
- timeout vs promotion
- reconnect vs timeout
- catch-up completion vs expiry
- rebuild completion vs epoch bump
3. simulator split clarification
-`distsim` keeps:
- protocol correctness
- lineage
- recoverability
- reference-state checking
-`eventsim` grows into:
- event scheduling
- timer firing
- same-time interleavings
- race exploration
### Out of scope
- production integration
- real transport
- real disk timings
- SPDK
- raw allocator
## Assigned Tasks For `sw`
### P0
1. Write a concrete `eventsim` scope note in code/docs
- define what stays in `distsim`
- define what moves to `eventsim`
- avoid overlap and duplicated semantics
2. Add minimal timeout event model
- first-class timeout event type(s)
- at minimum:
- barrier timeout
- catch-up timeout
- reservation expiry
3. Add timeout-backed scenarios
- stale delayed ack vs timeout
- catch-up timeout before convergence
- reservation expiry during active recovery
### P1
4. Add race-focused tests
- promotion vs delayed stale ack
- rebuild completion vs epoch bump
- reconnect success vs timeout firing
5. Keep traces debuggable
- failing runs must dump:
- seed
- event order
- timer events
- node states
- committed prefix
### P2
6. Decide whether selected `distsim` scenarios should also exist in `eventsim`
- only when timer/event ordering is the real point
- do not duplicate every scenario blindly
## Current Progress
Delivered in this phase so far:
-`eventsim` scope note added in code
- explicit timeout model added:
- barrier timeout
- catch-up timeout
- reservation timeout
- timeout-backed scenarios added and reviewed
- same-tick rule made explicit:
- data before timers
- recovery timeout cancellation is now model-driven, not test-driven
- stale barrier ack after timeout is explicitly rejected
- stale timeouts are separated from authoritative timeouts:
-`FiredTimeouts`
-`IgnoredTimeouts`
- race-focused scenarios added and reviewed:
- promotion vs stale catch-up timeout
- promotion vs stale barrier timeout
- rebuild completion vs epoch bump
- epoch bump vs stale catch-up timeout
- reusable trace builder added for replay/debug support
- current `distsim` suite at latest review:
- 86 tests passing
Remaining focus for `sw`:
- Phase 03 P0 and P1 are effectively complete
- Phase 03 P2 is also effectively complete after review
- any further simulator work should now be narrow and evidence-driven
- recommended next simulator additions only:
- control-plane latency parameter
- sustained-write convergence / tail-chasing load test
- one multi-promotion lineage extension
## Invariants To Preserve
1. committed data remains durable per policy
2. uncommitted data is never revived as committed
3. stale epoch traffic never mutates current lineage
4. committed prefix remains contiguous
5. timeout-triggered transitions are explicit and explainable
6. races do not silently bypass fencing or rebuild boundaries
## Required Updates Per Task
For each completed task:
1. add or update tests
2. update `sw-block/design/v2_scenarios.md` if scenario coverage changed
3. add a short note to:
-`sw-block/.private/phase/phase-03-log.md`
4. if the simulator boundary changed, record it in:
-`sw-block/.private/phase/phase-03-decisions.md`
## Exit Criteria
Phase 03 is done when:
1. timeout semantics exist as explicit simulator behavior
2. at least three important timer-race scenarios are modeled and tested
3.`distsim` vs `eventsim` responsibilities are clearly separated
4. failure traces from race/timeout scenarios are replayable enough to debug
Purpose: start the first standalone V2 implementation slice under `sw-block/`, centered on per-replica sender ownership and explicit recovery-session ownership
## Goal
Build the first real V2 implementation slice without destabilizing V1.
Purpose: close the critical V2 ownership-validation gap by making sender/session ownership explicit in both simulation and the standalone `enginev2` slice
## Goal
Validate the core V2 claim more deeply:
1. one stable sender identity per replica
2. one active recovery session per replica
3. endpoint change, epoch bump, and supersede rules invalidate stale work
4. stale late results from old sessions cannot mutate current state
This phase is not about adding broad new simulator surface.
It is about proving the ownership model that is supposed to make V2 better than V1.5.
## Why This Phase Exists
Current simulation is already strong on:
- quorum / commit rules
- stale epoch rejection
- catch-up vs rebuild
- timeout / race ordering
- changed-address recovery at the policy level
The remaining critical risk is narrower:
- the simulator still validates V2 strongly as policy
- but not yet strongly enough as owned sender/session protocol state
That is the highest-value validation gap to close before trusting V2 too much.
3. first rebuild execution path for the chosen product path
## Decision 3: Carry-forward limitations remain explicit until closed
Phase 08 must keep explicit:
1. committed truth is still not separated from checkpoint truth
2. rebuild execution is still incomplete
3. current control delivery is still simulated
## Decision 4: Phase 08 P0 is accepted
The hardening plan is sufficiently specified to begin implementation work.
In particular, `P0` now fixes:
1. the committed-truth gate decision requirement
2. the unified replay requirement after control and execution closure
3. the need for at least one real failover / reassignment validation target
## Decision 5: The committed-truth limitation must become a hardening gate
Phase 08 must explicitly decide one of:
1.`CommittedLSN != CheckpointLSN` separation is mandatory before a production-candidate phase
2. the first candidate path is intentionally bounded to the currently proven pre-checkpoint replay behavior
It must not remain only a documented carry-forward.
## Decision 6: Unified-path replay is required after control and execution closure
Once real control delivery and integrated execution closure land, `Phase 08` must replay the accepted failure-class set again on the unified live path.
This prevents independent closure of:
1. control delivery
2. execution closure
without proving that they behave correctly together.
## Decision 7: Real failover / reassignment validation is mandatory for the chosen path
Because the chosen product path depends on the existing master / volume-server heartbeat path, at least one real failover / promotion / reassignment cycle must be a named hardening target in `Phase 08`.
## Decision 8: Phase 08 should reuse the existing Seaweed control/runtime path, not invent a new one
For the first hardening path, implementation should preferentially reuse:
3. existing `blockvol` runtime and `v2bridge` storage/runtime hooks
This reuse is about:
1. control-plane reality
2. storage/runtime reality
3. execution-path reality
It is not permission to inherit old policy semantics as V2 truth.
The hard rule remains:
1. engine owns recovery policy
2. bridge translates confirmed control/storage truth
3.`blockvol` executes I/O
## Decision 9: Phase 08 P1 is accepted with explicit scope limits
Accepted `P1` coverage is:
1. real `ProcessAssignments()` path drives V2 engine sender/session state change
2. stable remote `ReplicaID` is derived from `ServerID`, not address
3. address change preserves sender identity through the live control path
4. stale epoch/session invalidation occurs through the live control path
5. missing `ServerID` fails closed
Not accepted as part of `P1`:
1. full end-to-end gRPC heartbeat delivery proof
2. integrated catch-up execution through the live path
3. rebuild execution through the live path
4. final local stable identity beyond transport-shaped `listenAddr`
## Decision 10: Phase 08 P2 is accepted as real execution closure
Accepted `P2` coverage is:
1.`CommittedLSN` is separated from `CheckpointLSN` on the chosen `sync_all` path
2. catch-up is proven as one live chain:
- engine plan
- engine executor
-`v2bridge`
- real `blockvol` I/O
- completion
- cleanup
3. rebuild is proven as one live chain for the delivered path
4. cleanup/pin release is asserted after execution
Residual non-blocking scope notes:
1.`CatchUpStartLSN` is not directly asserted in tests
2. rebuild source variants are not all forced and individually asserted
## Decision 11: Phase 08 now moves to unified hardening validation
With `P1` and `P2` accepted, the next required step is:
1. replay the accepted failure-class set again on the unified live path
2. validate at least one real failover / reassignment cycle
3. validate concurrent retention/pinner behavior
4. make the committed-truth gate decision explicit for the chosen candidate path
## Decision 12: Phase 08 P3 is accepted as unified hardening validation
Accepted `P3` coverage is:
1. replay of the accepted failure-class set on the unified `P1` + `P2` live path
2. at least one real failover / reassignment cycle through the live control path
3. one true simultaneous-overlap retention/pinner safety proof
4. stronger causality assertions for invalidation, escalation, catch-up, and completion
## Decision 13: The committed-truth gate is decided for the chosen candidate path
For the chosen `RF=2 sync_all` candidate path:
1.`CommittedLSN = WALHeadLSN`
2.`CheckpointLSN` remains the durable base-image boundary
3. this separation is accepted as sufficient for the candidate-path hardening boundary
This decision is intentionally scoped:
1. it is accepted for the chosen candidate path
2. it is not yet a blanket truth for every future path or durability mode
## Decision 14: Phase 08 P4 is candidate-path judgment, not broad new engineering expansion
`P4` should close `Phase 08` by producing one explicit candidate-path judgment.
Its main output is not more isolated engineering progress, but:
1. a bounded candidate-path statement
2. an evidence-to-claim mapping from accepted `P1` / `P2` / `P3` results
3. an explicit list of accepted bounds, remaining deferrals, and production blockers
`P4` may include small closure work if needed to make the candidate statement coherent, but it should not reopen protocol design or grow into another broad hardening slice.
## Decision 15: Phase 08 P4 is accepted as candidate package closure
Accepted `P4` coverage is:
1. one explicit candidate package for the chosen `RF=2 sync_all` path
Purpose: convert the accepted Phase 07 product path into a pre-production-hardening program without reopening accepted V2 protocol shape
## Why This Phase Exists
`Phase 07` completed:
1. a real service-slice integration around the V2 engine
2. real storage-truth bridge evidence through `v2bridge`
3. selected real-system failure replay
4. the first explicit product-path decision
What still does not exist is a pre-production-ready system path. The remaining work is no longer protocol discovery. It is closing the operational and integration gaps between the accepted product path and a hardened deployment candidate.
## Phase Goal
Harden the first accepted V2 product path until the remaining gap to a production candidate is explicit, bounded, and implementation-driven.
This phase doc is the canonical hardening contract for `sw` and `tester`.
Use `phase-08-log.md` for deeper engineering process, alternatives, and implementation detail.
Algorithm note:
- the accepted V2 algorithm / protocol shape is treated as fixed for this phase
- remaining work is engineering closure over real Seaweed/V1 runtime paths under V2 boundaries
- do not reopen protocol design unless a live contradiction is found
## Scope
### In scope
1. real master/control delivery into the engine service path
- the committed-truth carry-forward is now a required hardening gate, not just a note:
- either separate `CommittedLSN` from `CheckpointLSN` before a production-candidate phase
- or explicitly bound the first candidate path to the currently proven pre-checkpoint replay behavior
- at least one real failover / promotion / reassignment cycle is a required hardening target
- once `P1` and `P2` land, the accepted failure-class set must be replayed again on the newly unified live path
- the validation oracle for `Phase 08` is expected to reject overclaiming around:
- catch-up semantics
- rebuild execution
- master/control delivery
- candidate-path readiness vs production readiness
- accepted
Reference:
- `sw-block/docs/archive/design/phase-08-engine-skeleton-map.md` is the implementation-side skeleton map for this phase
- it is subordinate to `sw-block/design/v2-protocol-truths.md` and this `phase-08.md`; use it for module layout, execution order, interim fields, hard gates, and reuse guidance
### P1: Real Control Delivery
1. connect real master/heartbeat assignment delivery into the bridge
2. replace direct `AssignmentIntent` construction for the first live path
3. preserve stable identity and fenced authority through the real control path
4. include at least one real failover / promotion / reassignment validation target on the chosen `sync_all` path
Technical focus:
- keep the control-path split explicit:
- master confirms assignment / epoch / role
- bridge translates confirmed control truth into engine intent
- engine owns sender/session/recovery policy
- `blockvol` does not re-decide recovery policy
- preserve the identity rule through the live path:
- `ReplicaID = <volume>/<server>`
- endpoint change updates location but must not recreate logical identity
- preserve the fencing rule through the live path:
- stale epoch must invalidate old authority
- stale session must not mutate current lineage
- address change must invalidate the old live session before the new path proceeds
- treat failover / promotion / reassignment as control-truth events first, not storage-side heuristics
Implementation route (`reuse map`):
- reuse directly as the first hardening carrier:
- `weed/server/master_grpc_server.go`
- `weed/server/volume_grpc_client_to_master.go`
- `weed/server/volume_server_block.go`
- `weed/server/master_block_registry.go`
- `weed/server/master_block_failover.go`
- reuse as storage/runtime execution reality:
- `weed/storage/blockvol/blockvol.go`
- `weed/storage/blockvol/replica_apply.go`
- `weed/storage/blockvol/replica_barrier.go`
- `weed/storage/blockvol/v2bridge/`
- preserve the V2 boundary while reusing these files:
- reuse transport/control/runtime reality
- do not inherit old policy semantics as V2 truth
- keep engine as the recovery-policy owner
- keep `blockvol` as the I/O executor
Validation focus:
- prove live assignment delivery into the bridge/engine path
- prove stable `ReplicaID` across address refresh on the live path
- prove stale epoch / stale session invalidation through the live path
- prove at least one real failover / promotion / reassignment cycle on the chosen `sync_all` path
- prove the resulting logs explain:
- why reassignment happened
- why a session was invalidated
- which epoch / identity / endpoint drove the transition
Reject if:
- address-shaped identity reappears anywhere in the control path
- bridge starts re-deriving catch-up vs rebuild policy from convenience inputs
- old epoch or old session can still mutate after the new control truth arrives
- failover / reassignment is claimed without a real replay target
- delivery claims general production readiness rather than control-path closure
Status:
- accepted
- real assignment delivery into the V2 path is now proven through `ProcessAssignments()`
- accepted evidence includes:
- live assignment -> engine sender/session creation
- stable remote `ReplicaID = <volume>/<ServerID>`
- address-change identity preservation through the live path
- stale epoch/session invalidation through the live path
- fail-closed skip on missing `ServerID`
- accepted with explicit carry-forwards:
- `localServerID = listenAddr` remains transport-shaped for local identity
- heartbeat -> `ProcessAssignments()` is proven, but not full end-to-end gRPC delivery
- integrated catch-up execution is not yet proven through the live path
- rebuild coverage is limited to the first chosen executable path if that is all that lands
Status:
- accepted
- real one-chain execution is now proven for:
- catch-up
- rebuild
- accepted evidence includes:
- `CommittedLSN` separated from `CheckpointLSN` on the chosen `sync_all` path
- live engine plan -> executor -> `v2bridge` -> `blockvol` catch-up chain
- live engine plan -> executor -> `v2bridge` -> `blockvol` rebuild chain
- explicit pin cleanup assertions after execution
- accepted with explicit residual scope:
- `CatchUpStartLSN` is not directly asserted in tests
- rebuild source is not yet forced/verified per source variant
- broader rebuild-source coverage can remain follow-up work
Review checklist:
- is there one accepted catch-up proof from real `P1` control path to real session completion, using `CatchUpExecutor`
- is there one accepted first rebuild proof on the chosen path, using `RebuildExecutor`
- do live-path assertions prove pin/hold release on success, cancel, invalidation, and failure
- do logs/status explain start, cancel, failure, and completion without hidden transitions
- does the delivery avoid overclaiming general post-checkpoint catch-up, broad rebuild coverage, or production readiness
### P3: Hardening Validation
1. replay the accepted failure-class set again on the unified live path after `P1` + `P2`
2. validate at least one real failover / promotion / reassignment cycle through the live control path
3. validate concurrent retention/pinner behavior under overlapping recovery activity
4. make the committed-truth gate decision explicit for the chosen candidate path
Slice adjustment note:
- if `P2` lands only partially, `P3` should first close the missing execution outcome:
- real catch-up closure if still missing
- real first rebuild closure if still missing
- only after both are real should `P3` spend most of its weight on unified replay, failover / reassignment validation, and concurrent retention / cleanup hardening
Efficiency note:
- `P3` is a hardening-validation slice, not another execution-closure slice
- reuse the accepted `P1` / `P2` live path as the base; do not re-prove already accepted chain mechanics in isolation
- prefer one compact replay matrix over many near-duplicate tests
- prefer one real failover cycle and one true simultaneous-overlap retention case over broad scenario expansion
- the required new outputs are:
- unified replay evidence
- one real failover / reassignment replay
- one concurrent retention/pinner safety result
- one explicit committed-truth gate decision
Validation focus:
- unified replay for:
- changed-address restart
- stale epoch / stale session
- unrecoverable gap / needs-rebuild
- post-checkpoint boundary behavior
- at least one real failover / promotion / reassignment cycle
- concurrent retention/pinner safety under at least one true simultaneous-overlap hold case
- logs explain:
- why control truth changed
- why a session was invalidated
- why catch-up vs rebuild was chosen
- why execution completed, failed, or was cancelled
Reject if:
- accepted failure classes are still only partially replayed on the unified path
- failover / reassignment is claimed without a real live-path replay
- concurrent retention/pinner behavior leaks pins or violates recovery safety
- logs are too weak to replay causality offline
- the committed-truth gate is still just a note instead of an explicit decision
Status:
- accepted
- unified hardening replay is now proven on the accepted live path
- accepted evidence includes:
- replay of the accepted failure-class set on the unified `P1` + `P2` path
- at least one real failover / reassignment cycle through the live control path
- one true simultaneous-overlap retention/pinner safety proof
- stronger causality assertions for invalidation, escalation, catch-up, and completion
- committed-truth gate decision for the chosen candidate path:
- for the chosen `RF=2 sync_all` candidate path, `CommittedLSN = WALHeadLSN` with `CheckpointLSN` kept separate is accepted as sufficient for the candidate-path hardening boundary
- this is not yet a blanket truth for every future path or durability mode
### P4: Candidate Package Closure
1. classify what is truly ready for a first candidate path
2. package the accepted `P1` / `P2` / `P3` evidence into one bounded candidate package
3. turn carry-forwards into explicit candidate bounds or hard gates
4. state clearly what still remains before production readiness
Goal:
- finish `Phase 08` with one explicit candidate package, not just a collection of accepted slices
Verification mechanism:
- evidence map:
- every candidate claim must point to accepted evidence from `P1` / `P2` / `P3`
- tester validation:
- verify each candidate claim is supported by accepted evidence
- reject any claim that exceeds the proven boundary
- manager validation:
- verify the candidate statement is explicit, bounded, and not confused with production readiness
Output artifacts:
1. candidate-path statement in `phase-08.md`
2. candidate/gate decision record in `phase-08-decisions.md`
3. concise candidate package summary:
- candidate-safe capabilities
- explicit bounds
- deferred / blocking items
4. concise residual-gap summary:
- candidate-safe
- intentionally bounded
- still deferred / still blocking
5. short module/package boundary summary for later phases:
- what is already strong enough
- what moves to the next heavy engineering phase
Efficiency note:
- `P4` should mostly consume already accepted evidence, not create broad new engineering work
- only add implementation work if a small remaining blocker must be closed to make the candidate statement coherent
- if a gap is real but not worth closing in `Phase 08`, classify it explicitly rather than expanding scope implicitly
- `P4` exists inside `Phase 08` so the next phase can begin with substantial engineering work, not a light packaging-only round
Validation focus:
- make the candidate-path boundary explicit:
- what is proven
- what is intentionally bounded
- what is still deferred
- make the candidate package explicit:
- candidate-safe capability list
- evidence-to-claim mapping
- short module/package boundary summary
- make the committed-truth decision explicit:
- accepted for the chosen `RF=2 sync_all` candidate path
- still unclassified for future paths / durability modes unless separately proven
- prove the accepted product path can be described as an engineering candidate, not only as a set of slice-local proofs
- provide one explicit residual-gap list that separates:
- candidate-safe bounds
- future hardening work
- production blockers
Reject if:
- `P4` reopens protocol design instead of closing engineering gaps
- candidate claims are broader than the proven path
- carry-forwards remain informal notes rather than bounds or gates
- production readiness is implied from candidate readiness
- `P4` produces only prose summary without an evidence-to-claim mapping
- `P4` is too thin to leave the next phase with substantial engineering closure work
Status:
- accepted
- the first candidate package is now explicit for the chosen path
- newer checkpoint is rejected rather than silently accepted
4. convergence proof:
- post-install local runtime converges to `snapshotLSN`
- post-replay engine/runtime converge to `targetLSN`
5. cleanup proof:
- temporary snapshot ownership released on success/failure
Status:
- accepted
Carry-forward from `P2`:
1. `TruncateWAL` still not real
2. stronger live runtime ownership still not closed
### P3: Truncation Execution Closure
Goal:
- make `TruncateWAL` a real production-grade execution path for the chosen `RF=2 sync_all` candidate path
Required scope:
1. real truncation execution closure for the truncation-safe replica-ahead case
2. explicit rebuild escalation for replica-ahead cases that are not truncation-safe
3. one-chain proof through the catch-up executor path
4. fail-closed / no-overclaim behavior when local truncation is unsafe
5. no overclaim of broader runtime-ownership closure
Status:
- accepted
Carry-forward from `P3`:
1. truncation-safe vs rebuild-required replica-ahead split still happens at execution time, not planning time
2. stronger live runtime ownership still not closed
### P4: Stronger Live Runtime Ownership
Goal:
- move the accepted execution logic from bounded test/adapter ownership into a stronger live runtime path on the chosen `RF=2 sync_all` volume-server path
Required scope:
1. stronger volume-server/runtime ownership of recovery execution
- start `Phase 10` with one substantial control-plane closure package, not a loose collection of follow-up fixes
Must prove:
1. the phase is centered on real control-path closure rather than backend execution rework
2. the required closure targets are explicit:
- heartbeat / gRPC delivery
- reassignment / result convergence
- identity cleanup
- bounded repeated-assignment/idempotence cleanup
3. the chosen-path bound remains explicit
Verification mechanism:
1. architect review:
- control-plane scope is explicit and bounded
- proposed slices do not reopen accepted backend execution semantics
2. tester review:
- required end-to-end proofs are explicit
3. manager review:
- the package is concrete enough to assign the first implementation slice
Output artifacts:
1. explicit control-plane closure targets
2. explicit reject shapes
3. initial slice order inside `Phase 10`
Execution note:
- use `phase-10-log.md` as the technical pack for:
- semantic scope
- execution scope
- proof shapes
- assignment templates for `sw` and `tester`
Reject if:
1. `Phase 10` is framed as a vague "polish/control" phase without concrete closure targets
2. accepted `Phase 09` execution semantics are quietly reopened
3. product surfaces or unrelated hardening work are absorbed into this phase
4. no explicit end-to-end proof shape is defined
Status:
- accepted
### P1: Identity And Control-Truth Closure
Goal:
- close stable identity on the real chosen-path control wire so assignment truth, local ingest truth, and `ReplicaID` construction no longer depend on transport-shaped fallback
Accepted scope:
1. stable server identity preserved on the block assignment proto wire
2. master assignment generation preserves stable identity on the chosen path
3. volume-server local identity uses the same canonical server identity as the main volume server
4. real ingress proof:
- proto/decode
- `ProcessAssignments()`
- `ControlBridge`
- engine sender identity
5. fail-closed behavior for missing stable identity
Status:
- accepted
Carry-forward from `P1`:
1. fuller reassignment / failover result convergence is still open
2. broader control-plane reporting closure is still open
### P2: Reassignment / Result Convergence
Goal:
- prove that reassignment and failover converge through the real control path without stale local ownership or stale reported truth lingering after control truth changes
Accepted scope:
1. real failover / reassignment convergence through the chosen control path
2. no stale local runtime owner after control truth changes
3. no stale control/reporting truth after reassignment
4. one-chain proof through the real control path, not only local helper logic
5. no overclaim of broader hardening or product-surface closure
Status:
- accepted
Carry-forward from `P2`:
1. `P2` proves stale owner removal and no stale residue after control truth changes
2. bounded repeated-assignment/idempotence cleanup is still open where repeated primary assignment can still emit rebuild-server relisten warnings
3. `P2` does not claim broad master-driven failover infrastructure closure beyond the accepted volume-server-side chosen-path ingress
- close the remaining low-severity repeated-assignment/runtime-idempotence gap on the chosen path so duplicate or replacement primary assignments do not leave avoidable relisten/restart noise or ambiguous live-control ownership
Accepted scope:
1. repeated primary assignment on the same chosen-path volume should converge idempotently
2. rebuild-server/runtime side effects should not relaunch noisily when the authoritative control truth is unchanged or already active
3. bounded proof that repeated-assignment cleanup does not reopen accepted `P2` convergence or accepted `Phase 09` execution semantics
4. no expansion into broad runtime polish, product surfaces, or unrelated restart hardening
Status:
- accepted
Carry-forward from `P3`:
1. chosen-path repeated unchanged assignment is now absorbed idempotently across the accepted V2 + V1 live path
2. `P3` remains bounded cleanup; it does not itself close the remaining master-driven heartbeat/gRPC control-loop gap
3. fuller master-originated control delivery proof is still open
### P4: Master-Driven Control-Loop Closure
Goal:
- close the remaining chosen-path control-plane gap by proving that master-originated assignment truth delivered through the real heartbeat / gRPC control loop reaches the live volume-server path and converges without split truth
Required scope:
1. one bounded end-to-end proof from real master-produced chosen-path assignment truth into the live volume-server control path
2. proof that the real heartbeat / gRPC delivery path preserves the already accepted identity and convergence properties
3. proof that externally visible post-delivery state reflects the same new truth after the real master-driven path runs
4. no reopening of accepted `P1` / `P2` / `P3` semantics except for narrow bugs directly exposed by the fuller control-loop proof
5. no expansion into product surfaces, `RF>2`, or broad cluster-hardening work
Status:
- accepted
Carry-forward from `P4`:
1. bounded chosen-path master-driven heartbeat / gRPC control-loop closure is now accepted
2. `P4` does not claim full live transport-stream deployment proof or broad product hardening
3. the next phase should move to `Phase 11` product-surface rebinding
### Planned slice direction after `P0`
1. `P1`:
- identity and control-truth closure on the live control path
2. `P2`:
- reassignment / failover result convergence through the real control path
3. broad multi-surface product completion in one slice
4. `RF>2`, new durability modes, or broad cluster hardening
5. full production readiness / soak / rollout gates
## Phase 11 Items
### P0: First Surface Selection
Goal:
- choose the first product-facing surface that gives real product completion movement without turning the phase into a multi-system rewrite
Accepted decision:
1. the first bounded slice is `snapshot product path`
2. `CSI` is deferred to a later `Phase 11` slice because it pulls controller/node lifecycle, staging/publish, and broader cluster contract surface
3. `NVMe` / `iSCSI` rebinding are also deferred because they are transport/front-end adapters whose useful proof should come after one simpler product surface is already closed
Why this first:
1. snapshot is closest to already accepted backend truth
2. it exercises a real product-facing contract without immediately absorbing node/attach orchestration
3. it keeps the first `Phase 11` slice bounded to metadata/visibility/restore-contract correctness rather than transport and lifecycle breadth
Status:
- accepted
### P1: Snapshot Product-Path Rebinding
Goal:
- prove that the snapshot product path can be rebound onto the accepted V2-backed chosen path without semantic drift between snapshot-visible behavior and the accepted backend snapshot truth
Execution steps:
1. Step 1: contract freeze
- define exactly what the first slice claims:
- snapshot create
- snapshot list
- snapshot delete
- explicitly exclude clone/restore unless a later slice accepts them
- keep master/volume-server state and visible metadata coherent
3. Step 3: proof package
- prove create/list/delete on the chosen path
- prove fail-closed behavior for unsupported/invalid inputs
- prove no-overclaim around broader snapshot workflows
Required scope:
1. snapshot create/list/delete product-visible behavior on the chosen path
2. proof that snapshot metadata and visible snapshot set reflect the same accepted backend truth
3. proof that snapshot claims do not exceed the accepted V2 snapshot contract
4. explicit boundedness around restore/clone if they are not part of the first slice
Must prove:
1. snapshot creation on the product path maps to the accepted backend snapshot boundary rather than an implicit V1 truth
2. listing and deletion reflect the real volume-server/master state coherently
3. fail-closed behavior is preserved when snapshot prerequisites are missing or the volume is not eligible
4. the slice does not silently imply clone/restore/product workflow support that is not yet proven
Reuse discipline:
1. V1/master-facing snapshot RPC surface may be reused only as a product wrapper:
- `CreateBlockSnapshot`
- `DeleteBlockSnapshot`
- `ListBlockSnapshots`
2. V1/volume-server-facing snapshot surface may be reused only as the bounded execution adapter:
- `SnapshotBlockVol`
- `DeleteBlockSnapshot`
- `ListBlockSnapshots`
3. underlying `blockvol` snapshot implementation may be reused as execution reality, not as product truth ownership
4. every reused V1 surface must be called out explicitly in `phase-11-log.md` with one of:
- `update in place`
- `reference only`
- `reuse as bounded adapter`
5. no reused V1 surface may silently redefine snapshot semantics, placement truth, or product support claims
Verification mechanism:
1. focused integration tests for create/list/delete on the chosen path
2. contract checks that visible snapshot metadata matches the accepted backend snapshot truth
3. no-overclaim review on what user-visible snapshot behavior is actually supported after the slice
Hard indicators:
1. one accepted create proof:
- product-visible create succeeds on the chosen path
- created snapshot is observable through list/readback metadata
2. one accepted delete proof:
- deleted snapshot disappears from the visible snapshot set
- repeated delete is either idempotent-success or explicitly fail-closed as designed
3. one accepted list coherence proof:
- listed snapshot IDs/metadata match the real backend snapshot state
4. one accepted fail-closed proof:
- invalid volume / missing snapshot / unsupported preconditions do not imply false success
5. one accepted boundedness proof:
- docs/tests do not imply clone/restore/full snapshot workflow readiness unless separately proven
6. one accepted reuse-boundary proof:
- all V1 reuse surfaces touched by the slice are explicitly listed and their role is bounded
Reject if:
1. the slice proves only local helper behavior rather than product-visible snapshot behavior
2. visible snapshot metadata can drift from backend truth
3. the first slice quietly absorbs clone/restore or broader workflow work
4. the slice claims product readiness beyond create/list/delete on the chosen path
5. reuse of V1 surfaces is implicit or lets V1 semantics become the source of truth
Status:
- accepted
Carry-forward from `P1`:
1. bounded snapshot create/list/delete product rebinding is now accepted on the chosen path
2. `P1` does not claim restore/clone/full snapshot workflow readiness
3. `CSI` rebinding is now the next active `Phase 11` slice
### Later candidate slices inside `Phase 11`
1. `P2`: `CSI` rebinding after snapshot product-path closure
2. `P3`: `NVMe` / `iSCSI` front-end rebinding after one simpler product-visible surface is already accepted
3. `P4`: broader snapshot workflow closure (`restore` / `clone`) or other residual product workflow work only after earlier slices are bounded and proven
### P2: CSI Rebinding
Goal:
- bind the accepted V2-backed chosen path to the `CSI` controller/node product surface without reintroducing V1 recovery truth
Execution steps:
1. Step 1: contract freeze
- define the first bounded `CSI` surface claims:
- `CreateVolume`
- `DeleteVolume`
- `ControllerPublishVolume`
- `NodeStageVolume`
- `NodePublishVolume`
- `NodeUnpublishVolume`
- `NodeUnstageVolume`
- explicitly exclude CSI snapshot, expand, and NVMe-specific transport work unless a later slice accepts them
2. Step 2: backend rebinding
- bind CSI controller operations to the accepted master-backed chosen-path volume surface
- bind CSI node operations to the accepted chosen-path access contract for remote attach/stage/publish
3. Step 3: proof package
- prove bounded create/publish/stage/use/delete lifecycle on the chosen path
- prove fail-closed behavior for unsupported or invalid cases
- prove no-overclaim around broader CSI/product workflow breadth
Required scope:
1. bounded CSI controller/node lifecycle on the chosen path
2. explicit separation between accepted backend/control truth and CSI orchestration wrappers
3. remote target publication/staging behavior for the chosen path
4. no-overclaim around snapshots via CSI, expand, NVMe transport preference, multi-node topology breadth, or broad K8s readiness
Must prove:
1. CSI controller create/delete/publish map to the accepted master-backed chosen-path truth rather than a local V1 shortcut
2. CSI node stage/publish/unstage/unpublish consume the same chosen-path access truth without redefining recovery semantics
3. product-visible CSI lifecycle behavior is coherent across controller and node surfaces
4. fail-closed behavior is preserved when required publish/volume context or target information is missing
Reuse discipline:
1. V1/CSI-facing controller and node RPC surfaces may be reused only as bounded product adapters:
- `controller.go`
- `node.go`
- `server.go`
2. `volume_backend.go` may be reused only as the bounded bridge between CSI and accepted master/local surfaces
3. `volume_manager.go` may be reused only as bounded local execution reality where the slice explicitly proves that local manager behavior does not become semantic owner
4. accepted master block RPC surfaces may be reused only as bounded control/product adapters underneath the CSI backend bridge:
- `CreateBlockVolume`
- `DeleteBlockVolume`
- `LookupBlockVolume`
5. every reused V1 surface must be called out explicitly in `phase-11-log.md` with one of:
- `update in place`
- `reference only`
- `reuse as bounded adapter`
- `reuse as bounded bridge`
- `reuse as execution reality only`
6. no reused V1 surface may silently redefine lifecycle semantics, placement truth, or product support claims
Verification mechanism:
1. focused CSI controller/node integration tests on the chosen path
2. contract checks that controller-visible and node-visible truth match accepted backend/control truth
3. no-overclaim review on what CSI behavior is actually supported after the slice
Hard indicators:
1. one accepted controller create/publish proof:
- CSI create returns coherent volume/publish context on the chosen path
2. one accepted node stage/publish proof:
- node consumes the published target info and stages/publishes coherently on the chosen path
3. one accepted unpublish/unstage/delete proof:
- teardown/deletion complete without leaving false-visible ownership
4. one accepted fail-closed proof:
- missing or partial transport/context information does not imply false success
5. one accepted reuse-boundary proof:
- all CSI/V1 reuse surfaces touched by the slice are explicitly listed and bounded
6. one accepted boundedness proof:
- docs/tests do not imply CSI snapshot, expand, NVMe transport preference, or broad K8s/product readiness unless separately proven
Reject if:
1. the slice proves only CSI wrapper-local behavior without chosen-path backend/control coherence
2. controller truth and node truth can drift from accepted master-backed volume truth
3. the first CSI slice quietly absorbs snapshot, expand, NVMe, or broad multi-node/K8s readiness work
4. reuse of V1 surfaces is implicit or lets V1 semantics become the source of truth
Status:
- accepted
Carry-forward from `P2`:
1. bounded CSI controller/node lifecycle rebinding is now accepted on the chosen path
2. accepted proof uses the real master-backed create/lookup/delete path plus `mgr=nil` node consumption of published target truth
3. `P2` does not claim CSI snapshot, CSI expand, NVMe preference/failover closure, or broad Kubernetes readiness
### P3: NVMe / iSCSI Front-End Rebinding
Goal:
- bind transport/front-end publication surfaces onto the accepted V2-backed chosen path so the product-visible access path matches accepted backend/control truth
Execution steps:
1. Step 1: contract freeze
- define the first bounded front-end publication claims:
- create returns coherent front-end publication data
- lookup returns coherent front-end publication data
- heartbeat refresh preserves and updates publication truth
- failover switches publication truth to the new primary coherently
- explicitly exclude broad transport-performance claims, real initiator benchmarking, and broad cluster rollout readiness
2. Step 2: publication rebinding
- bind `iSCSI` and `NVMe` publication fields onto the accepted master-backed chosen-path truth
- keep registry-visible, lookup-visible, and CSI-visible publication truth coherent
3. Step 3: proof package
- prove bounded create/lookup/failover/restart publication truth on the chosen path
- prove fallback behavior is explicit where `NVMe` is absent
- prove no-overclaim around full transport runtime/performance closure
Required scope:
1. publication/address/naming truth for front-end adapters on the chosen path
2. bounded integration proof that master-visible and product-visible access metadata stay coherent
3. `NVMe` primary publication and `iSCSI` fallback publication where supported by the chosen path
4. explicit boundedness around real initiator behavior, transport performance, and broad cluster hardening
Must prove:
1. create/lookup publication fields map to accepted chosen-path truth rather than ad hoc wrapper-local construction
2. heartbeat refresh and failover preserve or update front-end publication truth coherently
3. `NVMe` and `iSCSI` publication fields do not drift between registry, lookup, and product-facing responses
4. mixed-capability or fallback behavior is explicit rather than silently overclaimed
Reuse discipline:
1. master-facing product/control publication surfaces may be reused only as bounded adapters:
- `CreateBlockVolume`
- `LookupBlockVolume`
2. registry publication fields may be reused only as bounded truth carriers, not independent semantic owners:
- `ISCSIAddr`
- `IQN`
- `NvmeAddr`
- `NQN`
3. volume-server allocation/publication surfaces may be reused only as bounded front-end publication sources:
- `AllocateBlockVolume`
- block heartbeat publication of `NvmeAddr` / `NQN`
4. existing `CSI` controller consumption of publication fields may be reused only as a bounded downstream consumer, not as the source of truth for `P3`
5. every reused V1 surface must be called out explicitly in `phase-11-log.md` with one of:
- `update in place`
- `reference only`
- `reuse as bounded adapter`
- `reuse as bounded truth carrier`
- `reuse as publication source only`
6. no reused V1 surface may silently redefine publication truth, failover truth, or supported transport claims
Verification mechanism:
1. focused integration tests for create/lookup publication truth on the chosen path
2. contract checks that registry-visible, lookup-visible, and consumer-visible publication fields match
3. failover/restart checks that front-end publication truth is reconstructed or updated coherently
4. no-overclaim review on what transport/front-end behavior is actually supported after the slice
| U1 | V2 engine accepts stale-epoch assignments at orchestrator level | V2 idempotence check skips only same-epoch; lower epoch creates new sender | Engine ApplyAssignment does not check epoch monotonicity on Reconcile | No — V1 HandleAssignment rejects epoch regression; V2 is secondary |
| U2 | Single-process test cannot exercise Primary→Rebuilding role transition | HandleAssignment rejects transition in shared store | Test harness limitation, not production bug | No — production VS has separate stores |
| U3 | gRPC stream transport not exercised in control-loop tests | All logic above/below stream is real; stream itself bypassed | Would require live master+VS gRPC servers in test | Blocks full integration test, not correctness |
2. If `Coherent=true`: lookup matches registry authority — no mismatch.
3. If `Coherent=false`: read `Reason` for explanation. Compare `LookupVolumeServer` vs `AuthorityVolumeServer` and `LookupIscsiAddr` vs `AuthorityIscsiAddr`.
4. Cross-check with `LookupBlockVolume` directly: repeated lookups should be self-consistent.
**Conclusion classes (from surfaces only):**
- **Coherent:**`PublicationDiagnostic.Coherent=true` — no mismatch.
- **Stale client:** Coherent but client sees old value — bounded by client re-query.
- **Unresolved:**`PublicationDiagnostic.Coherent=false` with no transient cause — escalate.
## S3: Leftover Runtime Work After Convergence
**Visible symptom:** After volume deletion or steady-state convergence, recovery tasks should have drained.
**Diagnosis surfaces:**
- `RecoveryDiagnostic` — `ActiveTasks` list (replicaIDs with active recovery work)
Purpose: move the accepted chosen-path implementation from candidate-safe product closure toward production-safe behavior under restart, disturbance, and operational reality
## Why This Phase Exists
`Phase 09` accepted production-grade execution closure on the chosen path.
`Phase 10` accepted bounded control-plane closure on that same path.
`Phase 11` accepted bounded product-surface rebinding on that same path.
What remains is no longer:
1. whether the chosen backend path works
2. whether selected product surfaces can be rebound onto it
It is now:
1. whether the chosen path stays correct under restart, failover, rejoin, and repeated disturbance
2. whether long-run behavior is stable enough for serious production use
3. whether operators can diagnose, bound, and reason about failures in practice
4. whether remaining production blockers are explicit and finite
## Phase Goal
Move from candidate-safe chosen-path closure to explicit production-hardening closure planning and execution.
Execution note:
1. treat `P0` as real planning work, not placeholder prose
2. use `phase-12-log.md` as the technical pack for:
- step breakdown
- hard indicators
- reject shapes
- assignment text for `sw` and `tester`
## Scope
### In scope
1. restart/recovery stability under repeated disturbance
2. long-run / soak viability planning and evidence design
3. operational diagnosability and blocker accounting
4. bounded hardening slices that do not reopen accepted earlier semantics casually
4. `P4`: performance floor and rollout-gate hardening
### P1: Restart / Recovery Disturbance Hardening
Goal:
- prove the accepted chosen path remains correct under restart, rejoin, repeated failover, and disturbance ordering
Acceptance object:
1. `P1` accepts correctness under restart/disturbance on the chosen path
2. it does not accept merely that recovery-related code paths exist
3. it does not accept merely that the system eventually seems to recover in a loose or approximate sense
Execution steps:
1. Step 1: disturbance contract freeze
- define the bounded disturbance classes for the first hardening slice:
- restart with same lineage
- restart with changed address / refreshed publication
- repeated failover / rejoin cycles
- delayed or stale signal arrival after restart/failover
2. Step 2: implementation hardening
- harden ownership/control reconstruction on the already accepted chosen path
- keep identity, epoch, session, and publication truth coherent across disturbance
3. Step 3: proof package
- prove repeated disturbance correctness on the chosen path
- prove stale or delayed signals fail closed rather than silently corrupting ownership truth
- prove no-overclaim around soak, perf, or broader production readiness
Required scope:
1. restart/rejoin correctness for the accepted chosen path
2. publication/address refresh correctness without identity drift
3. repeated ownership/control transitions under failover and rejoin
4. bounded reject behavior for stale heartbeat/control signals after disturbance
Must prove:
1. post-restart chosen-path ownership is reconstructed from accepted truth rather than accidental local leftovers
2. stale or delayed signals after restart/failover are rejected or explicitly bounded
3. repeated failover/rejoin cycles preserve identity, epoch/session monotonicity, and convergence on the chosen path
4. acceptance wording stays bounded to disturbance correctness rather than broad production-readiness claims
Reuse discipline:
1. `weed/server/block_recovery.go` and related tests may be updated in place as the primary restart/recovery ownership surface
2. `weed/server/master_block_failover.go`, `weed/server/master_block_registry.go`, and `weed/server/volume_server_block.go` may be updated in place as the accepted control/runtime disturbance surfaces
3. `weed/server/block_recovery_test.go`, `weed/server/block_recovery_adversarial_test.go`, and focused `qa_block_*` tests should carry the main proof burden
4. `weed/storage/blockvol/*` and `weed/storage/blockvol/v2bridge/*` are reference only unless disturbance hardening exposes a real bug in accepted earlier closure
5. no reused V1 surface may silently redefine chosen-path ownership truth, recovery choice, or disturbance acceptance wording
Verification mechanism:
1. focused restart/rejoin/failover integration tests on the chosen path
2. adversarial checks for stale or delayed control/heartbeat arrival after disturbance
3. explicit no-overclaim review so `P1` does not absorb soak/perf/product-expansion work
Hard indicators:
1. one accepted restart correctness proof:
- restart on the chosen path reconstructs valid ownership/control state
- post-restart behavior does not depend on accidental pre-restart leftovers
2. one accepted rejoin/publication-refresh proof:
- changed address or publication refresh does not break identity truth or visibility
3. one accepted repeated-disturbance proof:
- repeated failover/rejoin cycles converge without epoch/session regression
4. one accepted stale-signal proof:
- delayed heartbeat/control signals after disturbance do not re-authorize stale ownership
5. one accepted boundedness proof:
- `P1` claims correctness under disturbance, not soak, perf, or rollout readiness
Reject if:
1. evidence only shows that recovery code paths execute, rather than that correctness is preserved under disturbance
2. tests prove only one happy restart path and skip stale/delayed signal shapes
3. identity, epoch/session, or publication truth can drift across restart/rejoin
4. `P1` quietly absorbs soak, diagnosability, perf, or new product-surface work
Status:
- accepted
Carry-forward from `P0`:
1. the hardening object is the accepted chosen path from `Phase 09` + `Phase 10` + `Phase 11`
2. `P1` is the first correctness-hardening slice because disturbance threatens correctness before soak or perf
3. later `P2` / `P3` / `P4` remain distinct acceptance objects and should not be absorbed into `P1`
### P2: Soak / Long-Run Stability Hardening
Goal:
- prove the accepted chosen path remains viable over longer duration and repeated operation without hidden state drift
Acceptance object:
1. `P2` accepts bounded long-run stability on the chosen path under repeated operation or soak-like repetition
2. it does not accept merely that one disturbance test can be repeated many times manually
3. it does not accept diagnosability, performance floor, or rollout readiness by implication
Execution steps:
1. Step 1: soak contract freeze
- define one bounded repeated-operation envelope for the chosen path:
- define what counts as state drift versus expected bounded churn
2. Step 2: harness and evidence path
- build or adapt one repeatable soak/repeated-cycle harness on the accepted chosen path
- collect stable end-of-cycle truth rather than only transient pass/fail output
3. Step 3: proof package
- prove no hidden state drift across repeated cycles
- prove no unbounded growth/leak in the bounded chosen-path runtime state
- prove no-overclaim around diagnosability, perf, or production rollout
Required scope:
1. repeated-cycle correctness on the accepted chosen path
2. stable end-of-cycle ownership/control/publication truth after many cycles
3. bounded runtime-state hygiene across repeated operation
4. explicit distinction between acceptance evidence and support telemetry
Must prove:
1. repeated chosen-path cycles converge to the same bounded truth rather than accumulating semantic drift
2. registry / VS-visible / product-visible state remain mutually coherent after repeated cycles
3. repeated operation does not leave unbounded leftover tasks, sessions, or stale runtime ownership artifacts within the tested envelope
4. acceptance wording stays bounded to long-run stability rather than diagnosability/perf/launch claims
Reuse discipline:
1. `weed/server/qa_block_*test.go`, `block_recovery_test.go`, and related hardening tests may be updated in place as the primary repeated-cycle proof surface
2. testrunner / infra / metrics helpers may be reused as support instrumentation, but support telemetry must not replace acceptance assertions
3. `weed/server/master_block_failover.go`, `master_block_registry.go`, `volume_server_block.go`, and `block_recovery.go` may be updated in place only if repeated-cycle hardening exposes a real bug
4. `weed/storage/blockvol/*` and `weed/storage/blockvol/v2bridge/*` remain reference only unless soak evidence exposes a real accepted-path mismatch
5. no reused V1 surface may silently redefine the chosen-path steady-state truth, drift criteria, or soak acceptance wording
Verification mechanism:
1. one bounded repeated-cycle or soak harness on the chosen path
2. explicit end-of-cycle assertions for ownership/control/publication truth
3. explicit checks for bounded runtime-state hygiene after repeated cycles
4. no-overclaim review so `P2` does not absorb `P3` diagnosability or `P4` perf/rollout work
Hard indicators:
1. one accepted repeated-cycle proof:
- the chosen path completes many bounded cycles without semantic drift
- end-of-cycle truth remains coherent after each cycle
2. one accepted state-hygiene proof:
- no unbounded leftover runtime artifacts accumulate within the tested envelope
3. one accepted long-run stability proof:
- stability claims are based on repeated evidence, not one-shot reruns
4. one accepted boundedness proof:
- `P2` claims soak/long-run stability only, not diagnosability, perf, or rollout readiness
Reject if:
1. evidence is only a renamed rerun of `P1` disturbance tests
2. the slice counts iterations but never checks end-of-cycle truth for drift
3. support telemetry is presented without a hard acceptance assertion
4. `P2` quietly absorbs diagnosability, perf, or launch-readiness claims
Status:
- accepted
Carry-forward from `P1`:
1. bounded restart/disturbance correctness is now accepted on the chosen path
2. `P2` now asks whether that accepted path stays stable across repeated operation without hidden drift
3. later `P3` / `P4` remain distinct acceptance objects and should not be absorbed into `P2`
- make failures, residual blockers, and operator-visible diagnosis quality explicit and reviewable on the accepted chosen path
Acceptance object:
1. `P3` accepts bounded diagnosability / blocker accounting on the chosen path
2. it does not accept merely that some logs or debug strings exist
3. it does not accept performance floor or rollout readiness by implication
Execution steps:
1. Step 1: diagnosability contract freeze
- define one bounded diagnosis envelope for the accepted chosen path:
- failover / recovery does not converge in time
- publication / lookup truth does not match authority truth
- residual runtime work or stale ownership artifacts remain after an operation
- known production blockers remain open and must be made explicit
- define what counts as operator-visible diagnosis versus engineer-only source spelunking
2. Step 2: evidence-surface and blocker-ledger hardening
- identify or harden the minimum operator-visible surfaces needed to classify the bounded failure classes
- make residual blockers explicit, finite, and reviewable rather than implicit tribal knowledge
3. Step 3: proof package
- prove at least one bounded diagnosis loop closes from symptom to owning truth/blocker
- prove blocker accounting is explicit and does not hide unknown gaps behind “hardening later” language
- prove no-overclaim around perf, launch readiness, or broad topology support
Required scope:
1. operator-visible symptoms/logs/status for bounded chosen-path failure classes
2. one explicit mapping from symptom to ownership/control/runtime/publication truth
3. one explicit blocker ledger for unresolved production-hardening gaps
4. bounded runbook guidance for diagnosis of the accepted chosen path
Must prove:
1. bounded chosen-path failures can be distinguished with explicit operator-visible evidence rather than debugger-only knowledge
2. at least one diagnosis loop closes from visible symptom to the relevant authority/runtime truth without semantic ambiguity
3. residual blockers are explicit, finite, and named with a clear boundary rather than scattered across chats or memory
4. acceptance wording stays bounded to diagnosability / blocker accounting rather than perf or rollout claims
Reuse discipline:
1. `weed/server/qa_block_*test.go`, `block_recovery*_test.go`, and focused hardening tests may be updated in place where they can prove a bounded diagnosis loop on the accepted path
2. `weed/server/master_block_registry.go`, `master_block_failover.go`, `volume_server_block.go`, and `block_recovery.go` may be updated in place only if diagnosability work exposes a real visibility gap in accepted-path behavior
3. lightweight status/logging surfaces and bounded runbook docs may be updated in place as support artifacts, but support artifacts must not replace acceptance assertions
4. `weed/storage/blockvol/*` and `weed/storage/blockvol/v2bridge/*` remain reference only unless diagnosability work exposes a real accepted-path mismatch
5. no reused V1 surface may silently redefine chosen-path truth, blocker boundaries, or diagnosis acceptance wording
Verification mechanism:
1. one bounded diagnosis-loop proof on the accepted chosen path
2. one explicit blocker ledger or equivalent review artifact with finite named items
3. one explicit check that operator-visible evidence matches the underlying accepted truth being diagnosed
4. no-overclaim review so `P3` does not absorb `P4` perf/rollout work
Hard indicators:
1. one accepted symptom-classification proof:
- bounded failure classes can be told apart by explicit operator-visible evidence
2. one accepted diagnosis-loop proof:
- a visible symptom can be traced to the relevant ownership/control/runtime/publication truth
3. one accepted blocker-accounting proof:
- unresolved blockers are explicit, finite, and reviewable
4. one accepted boundedness proof:
- `P3` claims diagnosability / blockers only, not perf floor or rollout readiness
Reject if:
1. the slice merely adds logs or debug strings without proving diagnostic usefulness
2. blockers remain implicit, scattered, or dependent on private memory of prior chats
3. diagnosis requires debugger/source-level spelunking instead of bounded operator-visible evidence
1. bounded restart/disturbance correctness and bounded long-run stability are now accepted on the chosen path
2. `P3` now asks whether bounded failures and residual gaps are explicit and diagnosable in operator-facing terms
3. later `P4` remains a distinct acceptance object and should not be absorbed into `P3`
### P4: Performance Floor / Rollout Gates
Goal:
- define explicit performance floor, cost characterization, and rollout-gate criteria without letting perf claims replace correctness hardening
Acceptance object:
1. `P4` accepts a bounded performance floor and a bounded rollout-gate package for the accepted chosen path
2. it does not accept generic “performance is good” prose or one-off fast runs
3. it does not accept broad production rollout readiness outside the explicitly named launch envelope
Execution steps:
1. Step 1: performance-floor contract freeze
- define one bounded workload envelope for the accepted chosen path
- define which metrics count as acceptance evidence:
- throughput / latency floor
- resource-cost envelope
- disturbance-free steady-state behavior
- define which metrics are support-only telemetry
2. Step 2: benchmark and cost characterization
- run one repeatable benchmark package against the accepted chosen path
- record measured floor values and cost trade-offs rather than “fast enough” wording
3. Step 3: rollout-gate package
- translate accepted correctness, soak, diagnosability, and perf evidence into one bounded launch envelope
- make explicit which blockers are cleared, which remain, and what the first supported rollout shape is
Required scope:
1. one bounded benchmark matrix on the accepted chosen path
2. one explicit performance floor statement backed by measured evidence
3. one explicit resource-cost characterization
4. one rollout-gate / launch-envelope artifact with finite named requirements and exclusions
Must prove:
1. performance claims are tied to a named workload envelope rather than generic optimism
2. the chosen path has a measurable minimum acceptable floor within that envelope
3. rollout discussion is bounded by explicit gates and supported scope, not implied from prior slice acceptance
4. acceptance wording stays bounded to performance floor / rollout gates rather than broad production success claims
Reuse discipline:
1. `weed/server/qa_block_*test.go`, testrunner scenarios, and focused perf/support harnesses may be updated in place as the primary measurement surface
2. `weed/server/*`, `weed/storage/blockvol/*`, and `weed/storage/blockvol/v2bridge/*` may be updated in place only if performance-floor work exposes a real bug or a measurement-surface gap
3. `sw-block/.private/phase/` docs may be updated in place for the rollout-gate artifact and measured envelope
4. support telemetry may help characterize cost, but support telemetry must not replace the explicit floor/gate assertions
5. no reused V1 surface may silently redefine chosen-path truth, launch envelope, or rollout-gate wording
Verification mechanism:
1. one repeatable bounded benchmark package on the accepted chosen path
2. one explicit measured floor summary with named workload and cost envelope
3. one explicit rollout-gate artifact naming:
- supported launch envelope
- cleared blockers
- remaining blockers
- reject conditions for rollout
4. no-overclaim review so `P4` does not turn into generic launch optimism
Hard indicators:
1. one accepted performance-floor proof:
- measured floor values exist for the named workload envelope
2. one accepted cost-characterization proof:
- resource/replication tax or similar bounded cost is explicit
3. one accepted rollout-gate proof:
- the first supported launch envelope is explicit and finite
4. one accepted boundedness proof:
- `P4` claims only the bounded floor/gates it actually measures
Reject if:
1. the slice presents isolated benchmark numbers without a named workload contract
2. rollout gates are replaced by vague “looks ready” wording
3. support telemetry is presented without an explicit acceptance threshold or gate
4. `P4` quietly absorbs broad new topology, product-surface, or generic ops-tooling expansion
Status:
- accepted
Carry-forward from `P3`:
1. bounded disturbance correctness, bounded soak stability, and bounded diagnosability / blocker accounting are now accepted on the chosen path
2. `P4` now asks whether that accepted path has an explicit measured floor and an explicit first-launch envelope
3. later work after `Phase 12` should be a productionization program, not another hidden hardening slice
## Phase Close-Out Note
`Phase 12` is now accepted as bounded production hardening on the chosen path:
| PASS | `TestCanonicalizeAddr_NoAdvertised_FallsBackToOutbound` | fallback path works |
| PASS* | `TestBug3_ReplicaAddr_MustBeIPPort_WildcardBind` | documents gap: ReplicaReceiver may return `:port` not `ip:port` on wildcard bind; test passes as documentation, not as proof of fix → CP13-2 |
## Category 2: Durable Progress Truth
| Result | Test | Reason |
|--------|------|--------|
| PASS | `TestReplicaProgress_BarrierUsesFlushedLSN` | current code passes this test; suggests CP13-3 behavior may already exist |
| PASS | `TestReplicaProgress_FlushedLSNMonotonicWithinEpoch` | current code passes this test; suggests CP13-3 behavior may already exist |
| PASS | `TestReconnect_CatchupFromRetainedWal` | current code passes this test; suggests CP13-5 catch-up behavior may already exist |
| PASS* | `TestReconnect_GapBeyondRetainedWal_NeedsRebuild` | correctly fails SyncCache after large gap, but does NOT assert NeedsRebuild state transition — asserts barrier failure only → CP13-5+CP13-7 |
| FAIL | `TestAdversarial_ReconnectUsesHandshakeNotBootstrap` | **gap: degraded shipper with prior flushed progress reconnects but barrier fails** — shipper does not catch up before attempting barrier → CP13-5 |
| `TestAdversarial_NeedsRebuildBlocksAllPaths` | shipper stays Degraded after large gap, should be NeedsRebuild | CP13-5 (gap detection) + CP13-7 (rebuild fallback) |
| `TestAdversarial_CatchupDoesNotOverwriteNewerData` | catch-up fails at barrier, newer-data safety not exercised | CP13-5 (catch-up protocol) |
Main remaining failures cluster around CP13-5 (reconnect/catch-up), but CP13-7 (rebuild fallback) and part of CP13-6 (max-bytes retention) also remain open.
## PASS* → Checkpoint Mapping
| PASS* Test | Why Not Full Proof | Expected to close in |
This baseline does NOT close any checkpoint. Checkpoint closure requires dedicated review per checkpoint. The baseline only records which tests pass or fail on current code.
Tests passing on current code **suggests** the behavior may already exist, but does not constitute checkpoint acceptance. The following checkpoints still require dedicated review:
- **CP13-2** (canonical addressing): 1 PASS* test documents the gap
Commit: ac962fc83 → updated with legacy-response rejection fix
## Durable Progress Contract
### Definition
`replicaFlushedLSN` is the **sole authority** for replica durability in the sync_all path.
It means: the replica has called `fd.Sync()` (WAL fdatasync) for all entries through this LSN, and the barrier response carrying this value has reached the primary.
### What is NOT durable authority
| Variable | Location | Role | Why NOT authority |
| `shippedLSN` | `wal_shipper.go:269` | Diagnostic | Tracks last LSN sent over TCP; receipt not confirmed |
| `receivedLSN` | `replica_apply.go:362` | Intermediate | Entry applied to WAL buffer; not yet fsynced |
| `sentLSN` / transport progress | shipper send loop | Diagnostic | TCP write completed; no durability guarantee |
### Where durable authority lives
| Component | File | How it works |
|-----------|------|-------------|
| Replica: barrier handler | `replica_barrier.go:53-110` | Waits for `receivedLSN >= req.LSN`, calls `fd.Sync()`, advances `flushedLSN` only after sync succeeds, returns `BarrierResponse{FlushedLSN: flushed}` |
| Shipper: barrier consumer | `wal_shipper.go:220-238` | Reads `resp.FlushedLSN`, updates `replicaFlushedLSN` via monotonic CAS (never decreases) |
| Shipper: explicit API | `wal_shipper.go:273-278` | `ReplicaFlushedLSN()` is the authoritative API; `ShippedLSN()` has explicit "NOT authoritative" comment at line 268 |
| Group commit: sync_all | `dist_group_commit.go:15-83` | `BarrierAll(lsnMax)` called in parallel with local WAL sync; sync_all fails if any barrier fails |
- **Only `InSync` can complete barrier successfully.** It is the only state that proceeds
directly to the barrier request (ensureCtrlConn → MsgBarrierReq → wait for BarrierOK).
- **`Disconnected` and `Degraded` use Barrier() as a recovery entry point.** They attempt
bootstrap/reconnect inside the Barrier() call. If recovery succeeds and transitions to
InSync, the barrier request proceeds. If recovery fails, the barrier fails.
- **`Connecting`, `CatchingUp`, `NeedsRebuild` are rejected immediately** with `ErrReplicaDegraded`.
The key distinction: Barrier() can be *invoked* from Disconnected/Degraded (as a recovery
trigger), but only InSync can *satisfy* barrier success. The Disconnected/Degraded paths
are recovery attempts, not barrier eligibility.
## sync_all Gate
`dist_group_commit.go:59-66`: sync_all counts barrier failures. Any shipper that returns an error from `Barrier()` increments `failCount`. If `failCount > 0`, sync_all returns `ErrDurabilityBarrierFailed`.
Combined with the CP13-3 fix (FlushedLSN=0 rejected), the full chain is:
1. Only `InSync` shippers proceed to the barrier request
2. Disconnected/Degraded may recover inside Barrier(), transitioning to InSync before requesting
3. Only `BarrierOK` with `FlushedLSN > 0` counts as success
| `TestAdversarial_NeedsRebuildBlocksAllPaths` | FAIL | PASS | NeedsRebuild blocks Ship (drops) + Barrier (rejects) + is sticky across retries |
| `TestReconnect_GapBeyondRetainedWal_NeedsRebuild` | PASS* | PASS | Real reconnect handshake gap detection (R < S path), not budget trigger |
## Proof Promotion
### Primary proofs
| Test | What it proves for CP13-7 |
|------|--------------------------|
| `TestAdversarial_NeedsRebuildBlocksAllPaths` | 5 assertions: NeedsRebuild state, Ship drops, Barrier rejects, state sticky after barrier, second SyncCache still fails |
| `TestReconnect_GapBeyondRetainedWal_NeedsRebuild` | Real reconnect handshake detects R < S (gap beyond retained WAL) → SyncCache fails |
| `TestHeartbeat_ReportsNeedsRebuild` | Heartbeat carries per-replica `needs_rebuild` state |
| `TestRebuild_AbortOnEpochChange` | Epoch mismatch during rebuild → abort |
| `TestReplicaState_RebuildComplete_ReentersInSync` | Rebuild completion flow (reopen volume → RoleRebuilding → StartRebuild → fresh shipper → InSync). Support evidence: does not start from live NeedsRebuild shipper state, but proves the rebuild mechanics work end-to-end. |
## Updated Baseline Summary
| | PASS | FAIL | PASS* |
|---|---|---|---|
| CP13-1 (original) | 37 | 4 | 3 |
| After CP13-2..CP13-7 | **43** | **0** | **1** |
Remaining PASS*: `TestBug3_ReplicaAddr_MustBeIPPort_WildcardBind` (CP13-2 address witness — already upgraded to real proof in test but baseline doc still lists it as PASS*).
Before an explicit `V2 core` exists as a real code structure and live
event/command owner, current integrated tests are interpreted as:
1. validation of current `V1` runtime behavior under `V2` constraints
2. not proof that a completed `V2 runtime` already exists
`CP13-9` keeps that rule explicit.
It does not try to rewrite current constrained-runtime evidence into a claim that
the pure `V2 core` has already landed.
## Bounded Contract
`CP13-9` accepts one bounded thing:
1. explicit mode/publication normalization for the accepted chosen path
Scope remains bounded to:
1. `RF=2`
2. `sync_all`
3. current master / volume-server heartbeat path
4. `blockvol` as execution backend
It does not accept:
1. `Phase 14` pure `V2 core` extraction
2. broad launch approval
3. broad transport/product expansion
## Why This Checkpoint Exists
`CP13-8` and `CP13-8A` now prove:
1. one bounded real-workload package passes on the chosen path
2. assignment/readiness/publication closure is explicit enough for that path
What still needs freezing is the external mode meaning of the current path.
In particular:
1. a fresh volume before the first real replicated durability proof is not yet the
same as replicated-healthy
2. `degraded` and `NeedsRebuild` are not interchangeable
3. lookup / heartbeat / tester / debug surfaces should not silently use different
meanings of "healthy"
## Recommended Mode Contract
The semantic split below is the first-cut target.
Exact mode names may change, but the distinctions should remain explicit.
| Mode | Meaning | What it is allowed to claim |
|------|---------|-----------------------------|
| `allocated_only` | volume exists locally but runtime closure has not begun | existence only; not ready, not healthy |
| `bootstrap_pending` | assignment exists and the pair may need the first real replicated write/connect proof | not replicated-healthy; may be publishable only under bounded non-healthy wording |
| `replica_ready` | receiver / readiness closure exists on replica side | replica wiring is ready; not by itself proof of end-to-end healthy publication |
| `publish_healthy` | chosen-path publication conditions are closed | allowed to surface healthy publication on bounded chosen path |
| `degraded` | the bounded healthy path is not currently satisfied, but rebuild is not yet required | fail-closed for healthy replication claims |
| `needs_rebuild` | unrecoverable gap or equivalent fail-closed state | explicitly not healthy; normal replication path blocked |
## First-Write Bootstrap Rule
`CP13-9` should freeze this rule explicitly:
1. a freshly created `RF=2 sync_all` volume before the first real replicated write
or equivalent bounded durability proof must not be overclaimed as
replicated-healthy
2. if the current runtime needs the first replicated write to establish the first
real sync/connect proof, that is a mode-policy fact that must be surfaced
explicitly rather than hidden inside ambiguous degraded/healthy output
## Proof Shape
`CP13-9` should close with a bounded proof package:
| Proof | What it must show |
|-------|-------------------|
| Interpretation proof | current integrated evidence is described as constrained `V1` under `V2` constraints |
| Bootstrap proof | fresh volume before first replicated write is surfaced as bootstrap-pending or equivalent bounded non-healthy mode |
| Surface-consistency proof | lookup / heartbeat / tester / debug surfaces use one bounded mode meaning |
| Fail-closed proof | `publish_healthy`, `degraded`, and `needs_rebuild` remain distinct and do not overclaim health |
## Accepted Validation Summary
Tester verdict: `ACCEPT`
| Proof | Claim | Evidence |
|------|-------|----------|
| `AllocatedOnly` | `RF=1` maps to `allocated_only` | focused mode test |
| `BootstrapPending` (`Replicas` empty) | `RF=2` before replica set closure maps to `bootstrap_pending` | focused mode test |
| `BootstrapPending` (replica not ready) | `RF=2` with replica not ready maps to `bootstrap_pending` | focused mode test |
| `PublishHealthy` | ready + not transport degraded maps to `publish_healthy` | focused mode test |
| `Degraded` | transport degraded maps to `degraded` | focused mode test |
| `NeedsRebuild` | rebuilding role maps to `needs_rebuild` | focused mode test |
| `SurfaceConsistency` | mode / ready / degraded meaning stays aligned across transitions | focused transition checks |
| `InterpretationRule` | current integrated tests are constrained `V1` under `V2` constraints | explicit wording in contract + design docs |
| `NoOverclaim` | checkpoint does not claim pure `V2 core`, launch, or broad transport expansion | explicit boundedness wording |
Minor note kept bounded:
1. `assert_block_field` in the testrunner does not yet expose `volume_mode` as a first-class assert case
2. this does not block checkpoint acceptance because the bounded unit and API-surface proofs are already direct
## Relation to Earlier Checkpoints
| Prior checkpoint | What CP13-9 reuses |
|------------------|--------------------|
| `CP13-1..7` | accepted replication contract and fail-closed semantics |
| `CP13-8` | bounded real-workload pass on the chosen path |
Purpose: carry one explicit engineering gap beyond accepted `Phase 12` hardening into a bounded implementation phase so `RF=2 sync_all` becomes a correct, test-backed replicated durability mode under real reconnect, catch-up, retention, and rebuild conditions
`Phase 12` accepted bounded hardening, diagnosability, and first-launch envelope evidence.
What still remains is not broad protocol discovery.
It is one concrete engineering problem:
1. `sync_all` still needs a cleaner replicated-durability contract under cross-machine reconnect and replica recovery reality
2. that contract must be expressed in code and tests so later feature work can reuse it rather than reopen replication semantics repeatedly
## Phase Goal
Turn `RF=2 sync_all` from a bounded chosen-path mode with accepted launch-hardening evidence into a correct, reusable replicated-durability model for reconnect, catch-up, retention, and rebuild on real workloads.
Execution note:
1. use `phase-13-log.md` as the technical pack for:
- checkpoint breakdown
- acceptance objects
- reject shapes
- assignment text for `sw` and `tester`
2. prefer test-first baseline plus checkpointed implementation
3. keep the goal narrow: replication correctness first, not broad optimization or new transport work
4. unrelated control-plane or product-surface expansion
5. reopening accepted `Phase 09` / `Phase 10` / `Phase 11` / `Phase 12` semantics unless this phase exposes a real bug
## Phase 13 Items
### `CP13-1`: Test-First Baseline
Goal:
- freeze a failing/passing baseline that exposes the current replication gaps before protocol work begins
Acceptance object:
1. the focused sync-replication gap tests exist
2. they are run on current code before major implementation work
3. the fail/pass split is captured explicitly so later checkpoint claims are grounded
Status:
- accepted
Carry-forward:
1. the baseline report is frozen in `phase-13-cp1-baseline.md`
2. no protocol code was changed in `CP13-1`
3. `CP13-2` and later checkpoints must treat the baseline as the starting truth, not redefine it after implementation
### `CP13-2`: Canonical Replica Addressing
Goal:
- make replica endpoint truth canonical and routable so cross-machine replication never depends on wildcard listener strings, incomplete `:port` forms, or other non-authoritative address leakage
Acceptance object:
1. `CP13-2` accepts canonical replica address truth for the replication path
2. it does not accept durable-progress truth, reconnect protocol, WAL retention, or rebuild fallback by implication
3. it does not accept broad networking redesign beyond endpoint canonicalization
Execution steps:
1. Step 1: address truth contract freeze
- define the canonical replica endpoint form for replication surfaces as routable `host:port`
- define which forms are invalid for exported/registered truth:
- bare `:port`
- wildcard listener strings such as `[::]:port`
- accidental loopback when cross-machine routing is intended
2. Step 2: implementation hardening
- canonicalize replica listener addresses at the source where receiver/registration surfaces expose them
- keep authoritative endpoint truth aligned across local listener state, registration/heartbeat publication, and any registry copies
3. Step 3: proof package
- prove canonical `host:port` truth is emitted under wildcard-bind cases
- prove no wildcard or incomplete address string leaks into exported replication truth
- prove no-overclaim around reconnect, retention, or rebuild semantics
3. one focused wildcard-bind proof plus bounded cross-machine truth checks
4. explicit distinction between address canonicalization and later reconnect protocol work
Must prove:
1. cross-machine replica addresses exported for replication are canonical routable `host:port`
2. wildcard bind strings do not escape into replication truth
3. local canonicalization does not silently rewrite intentionally loopback-only cases into incorrect external truth
4. acceptance wording stays bounded to endpoint truth rather than later replication recovery semantics
Reuse discipline:
1. `weed/storage/blockvol/replica_receiver.go`, `replica_meta.go`, and nearby address helpers may be updated in place as the primary endpoint-truth surface
2. `weed/server/master_block_registry.go` and heartbeat/registration paths may be updated in place only if needed to keep authoritative endpoint truth aligned
3. focused unit/protocol tests should carry the main proof burden; component tests are support-only unless they prove an otherwise unreachable leak
4. no checkpoint work may silently introduce reconnect protocol, retention policy, or rebuild logic
Verification mechanism:
1. one focused wildcard-bind canonicalization proof
2. explicit checks that exported/registered replica endpoints are routable `host:port`
3. no-overclaim review so `CP13-2` does not absorb `CP13-3+`
Hard indicators:
1. one accepted canonical-endpoint proof:
- wildcard-bind listener state resolves to canonical exported `host:port`
2. one accepted no-leak proof:
- bare `:port` / wildcard listener strings no longer escape into replication truth
3. one accepted boundedness proof:
- `CP13-2` claims endpoint truth only, not reconnect or durability semantics
Reject if:
1. the checkpoint fixes only one test string shape but leaves other exported endpoint paths unchanged
2. canonicalization happens only in tests rather than at the production truth surface
3. the checkpoint quietly broadens into reconnect, retention, or rebuild protocol work
Status:
- accepted
Carry-forward:
1. `localServerID` remains stable control identity and may be opaque
2. `advertisedHost` is now the transport-facing canonicalization input for wildcard-bind replica endpoints
3. `CP13-3` and later checkpoints must not reopen identity-vs-transport separation unless a new concrete bug is exposed
### `CP13-3`: Durable Progress Truth
Goal:
- make durable replication progress explicit and authoritative so sync correctness is grounded in replica flushed durability rather than sender-side send progress or loosely inferred health
Acceptance object:
1. `CP13-3` accepts durable progress truth for the replication path
2. it does not accept reconnect/catch-up protocol, retention policy, rebuild fallback, or broader state-machine closure by implication
3. it does not accept generic “tests pass” reasoning without an explicit durable-progress contract review
Execution steps:
1. Step 1: durable-progress contract freeze
- define `replicaFlushedLSN` as replica-side WAL durability confirmed at barrier time
- define sender-side shipped/sent progress as diagnostic only, not authority for sync correctness
- define what barrier responses must expose as explicit durable progress truth
2. Step 2: implementation hardening or proof confirmation
- update the durable-progress path only where current code fails to meet the contract
- if current code already satisfies the contract, keep changes minimal and make the proof package explicit instead of broadening scope
3. Step 3: proof package
- prove barrier success is grounded in replica flushed durability
- prove flushed progress is monotonic within epoch and not updated on mere receive
- prove no-overclaim around `CP13-4+`
Required scope:
1. replica receiver durable-progress state
2. barrier request/response path
3. sender/group tracking of replica durable progress
4. explicit separation between durable-progress truth and later reconnect / retention semantics
Must prove:
1. `replicaFlushedLSN` means replica durability, not sender transmission progress
3. sender-side progress such as shipped/sent LSN is diagnostic only and cannot authorize sync success
4. acceptance wording stays bounded to durable-progress truth rather than broader recovery/state-machine closure
Reuse discipline:
1. `weed/storage/blockvol/replica_apply.go`, `wal_shipper.go`, `dist_group_commit.go`, and related protocol message code may be updated in place as the primary durable-progress surfaces
2. focused unit/protocol tests should carry the main proof burden
3. `weed/server/*` should remain reference only unless durable-progress truth requires an exposed wiring change
4. no checkpoint work may silently introduce reconnect protocol, retention policy, rebuild policy, or broader transport redesign
Verification mechanism:
1. one focused proof set around barrier/flushed progress truth
2. explicit checks that receive progress alone does not advance durable authority
3. no-overclaim review so `CP13-3` does not absorb `CP13-4+`
Hard indicators:
1. one accepted barrier-truth proof:
- barrier success is tied to replica flushed durability
2. one accepted monotonicity proof:
- `replicaFlushedLSN` is monotonic within epoch
3. one accepted no-false-authority proof:
- sender-side shipped/sent progress is diagnostic only
4. one accepted boundedness proof:
- `CP13-3` claims durable-progress truth only
Reject if:
1. the checkpoint treats passing baseline tests as automatic closure without reviewing the durable-progress contract
2. durable-progress truth is still mixed with sender-side transmission progress
3. the checkpoint quietly broadens into reconnect, retention, rebuild, or general replication redesign
Status:
- accepted
Carry-forward:
1. `replicaFlushedLSN` is now the authoritative durable-progress variable for `sync_all`
2. legacy `BarrierOK` responses without `FlushedLSN` are rejected and cannot count as durable authority
3. `CP13-4` and later checkpoints must treat sender-side send progress as diagnostic only, not as sync-correctness authority
### `CP13-4`: Replica State Machine / Barrier Eligibility
Goal:
- make replica state and barrier eligibility explicit so only `InSync` replicas can satisfy sync durability while non-eligible states fail closed instead of drifting into accidental success
Acceptance object:
1. `CP13-4` accepts the replica state machine and barrier-eligibility contract
2. it does not accept reconnect/catch-up protocol, retention policy, rebuild fallback, or broader rollout claims by implication
3. it does not accept vague “state seems fine” reasoning without an explicit eligibility contract
Execution steps:
1. Step 1: state contract freeze
- define the bounded state set used by the replication path:
- `Disconnected`
- `Connecting`
- `CatchingUp`
- `InSync`
- `Degraded`
- `NeedsRebuild`
- define barrier eligibility:
- only `InSync` replicas count toward sync durability
- non-eligible states must pre-reject or fail closed
2. Step 2: implementation hardening or proof confirmation
- update the state/eligibility path only where current code fails the contract
- if current code already satisfies much of the contract, keep code changes minimal and make the proof package explicit
3. Step 3: proof package
- prove barrier rejects replicas not eligible for sync durability
- prove degraded or catching-up replicas do not silently count toward `sync_all`
- prove no-overclaim around `CP13-5+`
Required scope:
1. replica shipper state transitions and eligibility checks
2. barrier admission path
3. `sync_all` failure semantics when replicas are non-eligible
4. explicit separation between state eligibility and later reconnect/rebuild protocol work
Must prove:
1. only `InSync` replicas count toward sync durability
2. `Disconnected`, `Connecting`, `CatchingUp`, `Degraded`, and `NeedsRebuild` do not silently satisfy barrier eligibility
3. degraded/non-eligible replicas fail closed for `sync_all` rather than producing false durability success
4. acceptance wording stays bounded to state/eligibility truth rather than reconnect, retention, or rebuild closure
Reuse discipline:
1. `weed/storage/blockvol/wal_shipper.go`, `dist_group_commit.go`, `shipper_group.go`, and nearby replication coordination code may be updated in place as the primary state/eligibility surfaces
2. focused unit/protocol/adversarial tests should carry the main proof burden
3. `weed/server/*` should remain reference only unless state eligibility requires a surfaced wiring correction
4. no checkpoint work may silently introduce reconnect handshake, retention policy, rebuild flow, or broader transport redesign
Verification mechanism:
1. one focused proof set around replica state and barrier eligibility
2. explicit checks that non-`InSync` states cannot satisfy `sync_all`
3. no-overclaim review so `CP13-4` does not absorb `CP13-5+`
Hard indicators:
1. one accepted eligibility proof:
- only `InSync` replicas count toward sync durability
2. one accepted fail-closed proof:
- non-eligible replicas cause bounded failure rather than false success
3. one accepted state-boundary proof:
- barrier rejects or excludes disallowed states explicitly
4. one accepted boundedness proof:
- `CP13-4` claims state/eligibility truth only
Reject if:
1. the checkpoint treats passing baseline tests as automatic closure without restating the state/eligibility contract
2. non-eligible replica states can still satisfy sync durability
3. the checkpoint quietly broadens into reconnect, retention, rebuild, or general replication redesign
Status:
- accepted
Carry-forward:
1. the replica state set and barrier-eligibility contract are now explicit
2. only `InSync` may satisfy sync durability; `Disconnected`/`Degraded` may invoke `Barrier()` only as bounded recovery entry paths
3. `CP13-5` and later checkpoints must preserve this eligibility boundary rather than reopening it implicitly
### `CP13-5`: Reconnect Handshake + WAL Catch-up
Goal:
- make reconnect after replica disturbance explicit and correct so a replica with known durable progress can resume from retained WAL, catch up, and re-enter `InSync` without false bootstrap success or barrier hangs
Acceptance object:
1. `CP13-5` accepts the reconnect handshake and WAL catch-up contract for recoverable gaps on the replication path
2. it does not accept replica-aware WAL retention policy, full rebuild fallback lifecycle, or broader rollout claims by implication
3. it does not accept vague “reconnect seems to work” reasoning without an explicit resume/catch-up contract
Execution steps:
1. Step 1: reconnect contract freeze
- define when a replica must use bootstrap versus reconnect:
- fresh replica with no prior durable progress may bootstrap
- replica with prior flushed progress must reconnect via explicit resume truth
- define reconnect decision outcomes:
- already caught up
- recoverable gap within retained WAL
- unrecoverable gap that must fail closed and defer full rebuild handling to `CP13-7`
2. Step 2: implementation hardening
- update the reconnect path only where current code still fails the resume/catch-up contract
- ensure catch-up replays retained WAL before barrier success is allowed
- ensure repeated disconnect/reconnect cycles remain bounded and do not silently fall back to unsafe bootstrap
3. Step 3: proof package
- prove degraded replicas with prior durable progress use handshake/reconnect rather than bootstrap
- prove retained-WAL catch-up completes and re-enters `InSync` on recoverable gaps
- prove reconnect fails closed on unrecoverable or incomplete recovery cases
- prove no-overclaim around `CP13-6+`
Required scope:
1. `wal_shipper` reconnect discriminator and resume handshake
4. bounded failure semantics for gaps that cannot be recovered within this checkpoint
5. explicit separation between reconnect/catch-up closure and later retention/rebuild policy work
Must prove:
1. fresh shippers bootstrap, but previously-synced shippers reconnect using resume truth
2. barrier success after disturbance is allowed only after reconnect/catch-up has re-established `InSync`
3. repeated disconnect/reconnect cycles do not strand the replica in false degraded recovery
4. recoverable gaps replay retained WAL correctly without overwriting newer replica data
5. acceptance wording stays bounded to reconnect/catch-up truth rather than retention or rebuild closure
Reuse discipline:
1. `weed/storage/blockvol/wal_shipper.go`, reconnect/catch-up helpers, and nearby replication protocol code may be updated in place as the primary reconnect surface
2. focused protocol/adversarial tests should carry the main proof burden; component tests are support-only unless a protocol gap is otherwise unreachable
3. `weed/server/*` should remain reference only unless reconnect correctness requires surfaced wiring changes
4. no checkpoint work may silently broaden into retention policy, rebuild orchestration, or performance tuning
Verification mechanism:
1. one focused proof set around reconnect discriminator, catch-up replay, and post-reconnect barrier behavior
2. explicit checks for repeated disconnect/reconnect recovery
3. explicit checks that recoverable gaps replay retained WAL before sync success
4. no-overclaim review so `CP13-5` does not absorb `CP13-6+`
Hard indicators:
1. one accepted reconnect-discriminator proof:
- prior durable progress uses handshake/reconnect rather than bootstrap
2. one accepted catch-up proof:
- recoverable retained-WAL gap replays and returns to `InSync`
3. one accepted repeated-recovery proof:
- multiple disconnect/reconnect cycles recover without hanging or drifting
4. one accepted fail-closed proof:
- reconnect does not falsely succeed when recovery is incomplete or impossible within retained WAL
5. one accepted boundedness proof:
- `CP13-5` claims reconnect/catch-up truth only
Reject if:
1. a previously-synced replica can still skip resume truth and succeed via unsafe bootstrap
2. barrier success can occur before reconnect/catch-up has restored `InSync`
3. repeated reconnect cycles still hang, strand, or silently degrade correctness
4. the checkpoint quietly broadens into retention, explicit `NeedsRebuild` lifecycle closure, rebuild execution, or general replication redesign
Status:
- accepted
Carry-forward:
1. replacement shippers now preserve prior durable-progress intent across `SetReplicaAddrs`
2. previously-synced replicas must reconnect through resume truth and retained-WAL catch-up rather than unsafe bootstrap
3. `CP13-6` and later checkpoints must preserve the reconnect/catch-up contract rather than weakening it through reclaim or rebuild shortcuts
### `CP13-6`: Replica-Aware WAL Retention
Goal:
- make WAL retention explicit and replica-aware so reclaim is gated by recoverable replica progress and bounded retention budgets rather than silently discarding catch-up-critical WAL
Acceptance object:
1. `CP13-6` accepts replica-aware WAL retention and retention-budget truth on the replication path
2. it does not accept full rebuild fallback lifecycle, rebuild execution, or broader rollout claims by implication
3. it does not accept vague “reclaim seems safe” reasoning without an explicit retention contract
Execution steps:
1. Step 1: retention contract freeze
- define which replica progress is authoritative for WAL retention:
- only replicas with prior durable progress and still recoverable state may hold WAL
- define bounded retention outcomes:
- reclaim blocked while a recoverable replica still needs retained WAL
- timeout / max-bytes budgets may escalate boundedly and release the WAL hold
- full rebuild handling after escalation remains `CP13-7`
2. Step 2: implementation hardening
- update the retention path only where current code still fails the bounded retention contract
- ensure retention decisions use replica-aware progress rather than primary-local heuristics alone
- ensure budget-triggered escalation is explicit and fail-closed rather than silent reclaim
3. Step 3: proof package
- prove recoverable replicas block reclaim of needed WAL
4. explicit separation between retention truth and full rebuild lifecycle closure
Must prove:
1. reclaim does not drop WAL still required by a recoverable replica
2. retention inputs come from replica-aware durable progress, not sender-side guesses
3. timeout / max-bytes budgets trigger bounded escalation when WAL cannot be held indefinitely
4. acceptance wording stays bounded to retention truth rather than full rebuild closure
Reuse discipline:
1. `weed/storage/blockvol` WAL-retention, flusher, shipper-group, and adjacent replication coordination code may be updated in place as the primary retention surface
2. focused unit/protocol tests should carry the main proof burden; component tests are support-only unless a retention gap is otherwise unreachable
3. `weed/server/*` should remain reference only unless retention truth requires surfaced reporting changes
4. no checkpoint work may silently broaden into rebuild execution, broad control-plane redesign, or performance tuning
Verification mechanism:
1. one focused proof set around retention hold, reclaim gating, and budget-triggered escalation
2. explicit checks that max-bytes and timeout paths are real production behaviors, not just comments/logs
3. explicit checks that retention stays compatible with `CP13-5` recoverable catch-up
4. no-overclaim review so `CP13-6` does not absorb `CP13-7+`
Hard indicators:
1. one accepted hold-back proof:
- recoverable replicas block reclaim of required WAL
2. one accepted timeout-budget proof:
- timeout can escalate a stalled recoverable replica into bounded fail-closed behavior
3. one accepted max-bytes-budget proof:
- max-bytes pressure triggers explicit bounded escalation rather than silent reclaim or TODO-only behavior
4. one accepted boundedness proof:
- `CP13-6` claims retention truth only
Reject if:
1. reclaim can still silently discard WAL needed for a recoverable replica
2. max-bytes behavior is still only log text / placeholder behavior without real state effect
3. the checkpoint quietly broadens into full `NeedsRebuild` lifecycle closure, rebuild execution, or general replication redesign
Status:
- accepted
Carry-forward:
1. retention inputs and bounded retention budgets are now replica-aware
2. timeout and max-bytes escalation can move a stalled recoverable replica into `NeedsRebuild`
3. `CP13-7` must turn that escalation into a real fail-closed rebuild lifecycle rather than leaving `NeedsRebuild` as a partially-signaled state
### `CP13-7`: Rebuild Fallback
Goal:
- make `NeedsRebuild` a real fail-closed recovery state so unrecoverable replicas stop participating in normal replication paths, surface rebuild intent clearly, and re-enter the replication contract only through bounded rebuild handoff
Acceptance object:
1. `CP13-7` accepts the `NeedsRebuild` fallback and bounded rebuild handoff lifecycle on the replication path
2. it does not accept broad rollout claims or real-workload validation by implication
3. it does not accept vague “rebuild eventually works” reasoning without an explicit fail-closed lifecycle contract
Execution steps:
1. Step 1: rebuild-fallback contract freeze
- define what `NeedsRebuild` means:
- unrecoverable via retained WAL catch-up
- excluded from normal ship/barrier success
- visible to rebuild orchestration and observability surfaces
- define lifecycle boundaries:
- detection/escalation into `NeedsRebuild`
- fail-closed behavior while in `NeedsRebuild`
- bounded rebuild handoff and post-rebuild re-entry
2. Step 2: implementation hardening
- update the rebuild-fallback path only where current code still leaves `NeedsRebuild` partial, leaky, or inconsistent
- ensure ship/barrier paths block correctly while `NeedsRebuild`
- ensure successful rebuild resets progress/state in a way compatible with later re-entry
3. Step 3: proof package
- prove unrecoverable gaps transition to `NeedsRebuild`
- prove `NeedsRebuild` blocks normal replication participation
- prove rebuild handoff can re-establish a bounded healthy starting point
- prove no-overclaim around `CP13-8+`
Required scope:
1. `NeedsRebuild` detection and state ownership on the primary shipper side
2. fail-closed behavior for ship/barrier and related replication paths while `NeedsRebuild`
4. post-rebuild progress/state initialization needed for safe re-entry
5. explicit separation between rebuild fallback closure and later real-workload validation
Must prove:
1. unrecoverable gaps do not remain merely degraded; they transition to `NeedsRebuild`
2. a shipper in `NeedsRebuild` cannot silently participate in ship/barrier success
3. rebuild completion restores a bounded re-entry point without faking immediate `InSync`
4. acceptance wording stays bounded to rebuild fallback truth rather than `CP13-8` rollout/workload claims
Reuse discipline:
1. `weed/storage/blockvol` rebuild, shipper-group, wal-shipper, and adjacent replication coordination code may be updated in place as the primary rebuild-fallback surface
2. focused unit/protocol/adversarial tests should carry the main proof burden; component tests are support-only unless a rebuild gap is otherwise unreachable
3. `weed/server/*` should remain reference only unless rebuild fallback requires surfaced status/reporting changes
4. no checkpoint work may silently broaden into real-workload benchmarking, performance tuning, or new protocol discovery
Verification mechanism:
1. one focused proof set around `NeedsRebuild` transition, blocking semantics, and rebuild re-entry
2. explicit checks that `NeedsRebuild` blocks normal replication paths rather than merely logging/marking degraded
3. explicit checks that post-rebuild progress initializes from bounded truth such as checkpoint state
4. no-overclaim review so `CP13-7` does not absorb `CP13-8+`
Hard indicators:
1. one accepted transition proof:
- unrecoverable retained-WAL gap transitions to `NeedsRebuild`
- rebuild start/complete path restores a bounded re-entry state
4. one accepted post-rebuild-progress proof:
- replica progress after rebuild is initialized from checkpoint truth, not stale/zeroed state
5. one accepted boundedness proof:
- `CP13-7` claims rebuild fallback only
Reject if:
1. an unrecoverable gap can still linger in `Degraded` without escalating to `NeedsRebuild`
2. a `NeedsRebuild` shipper can still satisfy normal ship/barrier paths
3. rebuild completion jumps directly to misleading healthy semantics without bounded re-entry proof
4. the checkpoint quietly broadens into `CP13-8` real-workload validation or general replication redesign
Status:
- accepted
Carry-forward:
1. `NeedsRebuild` is now a real fail-closed fallback state
2. rebuild handoff and post-rebuild progress are bounded by checkpoint truth rather than implicit recovery assumptions
3. `CP13-8` must validate the accepted replication contract on named real workloads without reopening protocol semantics or quietly broadening into mode policy work
### `CP13-8`: Real-Workload Validation
Goal:
- validate the accepted `RF=2 sync_all` replication contract on one bounded set of real workloads so the engineering proof is no longer only protocol/unit-level but also demonstrated on named real block-device consumers
Acceptance object:
1. `CP13-8` accepts one bounded real-workload validation package for the accepted `RF=2 sync_all` path
2. it does not accept broad rollout claims, broad benchmark positioning, or mode normalization by implication
3. it does not accept vague “worked in a manual run” reasoning without named workloads, bounded envelope, and replayable evidence
Execution steps:
1. Step 1: workload envelope freeze
- name one bounded validation matrix:
- workload(s)
- topology
- transport/frontend
- filesystem/application surface
- disturbance shapes included and excluded
- recommended first-cut surfaces:
- real filesystem behavior such as `ext4`
- one database/application surface such as `PostgreSQL`
2. Step 2: harness and evidence hardening
- wire the workload run through real block-device consumers on the accepted path
- keep the environment reproducible and bounded enough that failures are attributable
- collect evidence at the same semantic layer as accepted prior checkpoints
3. Step 3: proof package
- prove the named real workloads complete correctly on the accepted path
- prove disturbance/failover behavior is bounded inside the named envelope if included
- prove no-overclaim around `CP13-9+`
Required scope:
1. one bounded workload matrix on the accepted `RF=2 sync_all` path
2. real block-device consumer validation (not only protocol/unit tests)
3. bounded disturbance cases only if explicitly named in the envelope
4. explicit separation between real-workload proof and later mode normalization / rollout claims
Must prove:
1. the accepted replication contract survives contact with named real workloads
2. evidence is tied to a bounded environment and workload envelope, not generic “production ready” rhetoric
3. failures, if any, are attributable to explicit workload-envelope gaps rather than ambiguous harness drift
4. acceptance wording stays bounded to real-workload validation rather than `CP13-9` policy/mode closure
Reuse discipline:
1. prefer existing `testrunner`, bounded component scenarios, and real-device harnesses where possible
2. update `weed/storage/blockvol/*` only when the real workload exposes a concrete bug in accepted semantics
3. `weed/server/*` should remain reference only unless workload validation exposes a surfaced control/runtime issue
4. no checkpoint work may silently broaden into generic benchmark marketing, launch approval, or mode policy redesign
Verification mechanism:
1. one named workload matrix with explicit environment description
2. replayable runs or artifacts for the chosen workload package
3. explicit pass/fail conditions tied back to accepted `CP13-1..7` semantics
4. no-overclaim review so `CP13-8` does not absorb `CP13-9+`
Hard indicators:
1. one accepted filesystem proof:
- a named real filesystem workload completes correctly on the accepted path
2. one accepted application proof:
- a named real application/database workload completes correctly on the accepted path
3. one accepted envelope proof:
- the validation matrix is explicit about topology, frontend, workload, and exclusions
4. one accepted boundedness proof:
- `CP13-8` claims real-workload validation only
Reject if:
1. the checkpoint relies on ad hoc manual runs with no bounded envelope
2. a claimed real-workload proof is actually only a synthetic benchmark or unit test
3. delivery wording quietly broadens into mode normalization, launch approval, or general production-readiness claims
Status:
- accepted
Carry-forward:
1. one bounded real-workload package now passes on the chosen path:
- `RF=2`
- `sync_all`
- iSCSI
- `ext4 + pgbench`
- one failover
2. this checkpoint validates current runtime behavior under accepted `V2` constraints
3. it does not by itself mean a pure `V2 runtime` already exists
4. `CP13-8A` and `CP13-9` must keep that interpretation explicit
### `CP13-8A`: Assignment-to-Publication Closure
Goal:
- close the control/runtime/publication contradiction exposed by `CP13-8` so the system no longer treats allocation or assignment presence as equivalent to replica publication readiness
Acceptance object:
1. `CP13-8A` accepts one bounded closure slice for assignment-to-publication truth on the accepted `RF=2 sync_all` path
2. it does not accept broad mode normalization, launch approval, or backend replacement by implication
3. it does not accept sleep-based or timing-based fixes that leave readiness semantics implicit
Execution steps:
1. Step 1: unify assignment lifecycle
- ensure assignment delivery flows through one authoritative path from role apply to receiver/shipper wiring to readiness bookkeeping
- remove semantic split between store-only role application and service-level replication/publication setup
2. Step 2: name readiness and publication truth
- define explicit readiness states for the chosen path
- rerun the bounded `CP13-8` workload package after closure lands
- determine whether the remaining contradiction is backend data visibility, adapter timing/publication, or a true core-rule gap
Required scope:
1. assignment-to-publication closure only
2. chosen path only: `RF=2 sync_all`
3. existing master / volume-server heartbeat path only
4. `blockvol` remains the execution backend
Must prove:
1. assignment delivered does not by itself imply receiver ready or publish healthy
2. replica publication requires explicit readiness closure rather than allocation completion or precomputed port presence
3. master lookup / REST / tester health checks consume the same bounded readiness truth
4. `CP13-8A` remains about closure, not mode normalization or backend redesign
Reuse discipline:
1. prefer `weed/server/*` and bridge-layer updates first because this is a surfaced control/runtime issue
2. update `weed/storage/blockvol/*` only if closure work exposes a concrete backend bug rather than a publication-path contradiction
3. keep `CP13-1..7` semantics fixed unless the closure work exposes a live contradiction
4. no checkpoint work may silently broaden into `CP13-9` mode policy or broad rollout claims
Verification mechanism:
1. one focused proof set around assignment lifecycle closure and readiness/publication gating
2. explicit tests that heartbeat / lookup / tester surfaces do not publish a replica before readiness closes
3. bounded `CP13-8` rerun or equivalent evidence showing the contradiction moves from mixed-state ambiguity to an attributable remaining cause
4. no-overclaim review so `CP13-8A` does not absorb `CP13-9`
Hard indicators:
1. one accepted lifecycle proof:
- assignment processing uses one authoritative path from role apply through runtime wiring
2. one accepted readiness proof:
- replica-ready is explicit and not inferred from mere existence/allocation
3. one accepted publication proof:
- lookup / heartbeat / tester gates do not publish a replica before readiness closure
4. one accepted boundedness proof:
- `CP13-8A` claims closure only and leaves broader mode policy untouched
Reject if:
1. assignment still reaches different semantic outcomes depending on whether it flows through heartbeat/store-only or service-level processing
2. a replica can still be surfaced as healthy/ready before receiver/session readiness closes
3. the slice relies on delays or ad hoc retries rather than explicit readiness semantics
4. delivery wording broadens into `CP13-9` mode normalization, launch approval, or generic backend replacement
Status:
- accepted
Carry-forward:
1. assignment/readiness/publication closure is now explicit enough for the bounded chosen path
2. the corrected path no longer treats replica allocation or assignment presence as equivalent to replica publication readiness
3. the remaining next step is mode-policy normalization on top of this closed assignment/publication path
### `CP13-9`: Mode Normalization Under `V2` Constraints
Goal:
- freeze one bounded mode-policy contract for the current chosen path so external health/publication meaning no longer drifts between implicit `V1` runtime behavior and `V2` constraint language
Acceptance object:
1. `CP13-9` accepts one bounded mode-normalization package for the accepted `RF=2 sync_all` path
2. it accepts mode/publication semantics for the current runtime only under explicit `V2` constraints
3. it does not accept pure `V2 core` extraction, launch approval, or broad transport/product expansion by implication
Execution steps:
1. Step 1: interpretation rule freeze
- make explicit that current integrated tests are evaluating `V1` runtime behavior under `V2` constraints
- define `CP13-9` as policy/meaning closure for the constrained current path, not proof that a completed `V2 runtime` already exists
2. Step 2: mode contract freeze
- define one bounded external mode set for the chosen path
- at minimum distinguish:
- allocated / assigned
- bootstrap-pending
- replica-ready
- publish-healthy
- degraded
- `NeedsRebuild`
- define what each surface is allowed to claim for each mode:
- heartbeat
- lookup / REST / tester surfaces
- operator/debug surfaces
3. Step 3: bootstrap-policy closure
- make the first-write / first-connect bootstrap behavior explicit
- ensure a freshly created `RF=2 sync_all` volume is not overclaimed as replicated-healthy before the first real replicated durability proof exists
4. Step 4: proof package
- prove all relevant surfaces agree on the bounded mode meanings
- prove no-overclaim around future pure-core extraction or broad launch claims
Required scope:
1. chosen path only: `RF=2 sync_all`
2. current master / volume-server heartbeat path only
3. `blockvol` remains the execution backend
4. current integrated runtime is interpreted as constrained `V1`, not yet as a completed `V2 runtime`
Must prove:
1. health/publication meaning is explicit and consistent across product/tester/operator surfaces
2. `bootstrap-pending` or equivalent first-write state is explicit rather than hidden inside ambiguous degraded/healthy output
3. publish/ready semantics remain fail-closed under the accepted replication contract
4. acceptance wording stays bounded to mode normalization for the constrained current path rather than `V2 core` extraction
Reuse discipline:
1. prefer surfaced policy/diagnostic/projection work first because this checkpoint is about external mode meaning
2. update `weed/storage/blockvol/*` only if mode normalization exposes a concrete backend leak rather than a surface-meaning gap
3. keep `CP13-1..8A` semantics fixed unless a live contradiction is exposed
4. no checkpoint work may silently broaden into `Phase 14` pure-core extraction or broad rollout claims
Verification mechanism:
1. one focused proof set around mode/publication semantics across heartbeat / lookup / tester / debug surfaces
2. explicit tests or bounded evidence that a fresh volume before first replicated write is not overpublished as replicated-healthy
3. explicit checks that degraded / rebuild-required surfaces remain distinguishable and bounded
4. no-overclaim review so `CP13-9` does not absorb `Phase 14`
Hard indicators:
1. one accepted interpretation proof:
- current integrated evidence is explicitly described as constrained `V1` under `V2` constraints
2. one accepted bootstrap proof:
- a fresh `RF=2 sync_all` volume before first replicated write is surfaced as bootstrap-pending or equivalent bounded non-healthy mode
3. one accepted surface-consistency proof:
- heartbeat / lookup / tester / debug surfaces agree on the same bounded mode meanings
4. one accepted boundedness proof:
- `CP13-9` claims mode normalization only and leaves pure-core extraction to later phases
Reject if:
1. the slice still uses one meaning of “healthy” for lookup and a different one for tester/debug/operator surfaces
2. a fresh volume can still appear fully replicated-healthy before first real replicated durability proof exists
3. the checkpoint quietly claims a completed `V2 runtime` already exists
4. delivery wording broadens into launch approval, broad productization, or `Phase 14` pure-core extraction
Status:
- accepted
Carry-forward:
1. one bounded mode set is now explicit for the current constrained chosen path:
- `allocated_only`
- `bootstrap_pending`
- `publish_healthy`
- `degraded`
- `needs_rebuild`
2. current integrated tests remain explicitly interpreted as constrained `V1` under `V2` constraints
3. `CP13-9` does not claim pure `V2 core` extraction, launch approval, or broad transport expansion
## Reuse Discipline
1. `weed/storage/blockvol/*` is the primary implementation surface and may be updated in place
2. focused unit/component/adversarial tests should carry the main proof burden
3. real-node / real-device validation belongs in testrunner or bounded component scenarios, not chat prose
4. `weed/server/*` may be updated only when replication correctness requires registry / assignment / heartbeat truth to change
5. no checkpoint may silently broaden into performance-optimization or broad rollout work
## Expected Outcome
`Phase 13` now succeeds with the following closure:
1. reconnect / catch-up / rebuild semantics become explicit and test-backed
2. `sync_all` correctness no longer depends on partial or implicit sender-state assumptions
3. later feature work can reuse a clearer replication contract instead of re-deriving durability semantics each time
4. one bounded real-workload package and one bounded mode-normalization package are both accepted on the current constrained path
1. make mode, readiness, and publication first-class core-owned meanings
Acceptance object:
1. `VolumeState`, normalized mode/readiness/publication state, and bounded
outward projection exist in `sw-block/engine/replication`
2. `publish_healthy` is derived from named semantic state rather than runtime
convenience
3. fail-closed mode distinctions stay explicit:
- `allocated_only`
- `bootstrap_pending`
- `replica_ready`
- `publish_healthy`
- `degraded`
- `needs_rebuild`
4. the structural acceptance tests prove:
- `replica_ready` and `publish_healthy` stay distinct
- no-replica path stays `allocated_only`
- `degraded` and `needs_rebuild` remain distinct fail-closed meanings
- the current integrated interpretation remains `constrained_v1`, not live
`v2_core` cutover
Ownership boundary:
1. `14A` owns semantic shell closure for:
- mode
- readiness
- publication
2. `14A` does not own:
- command-sequence closure
- durable-boundary closure
- recovery closure
Status:
1. delivered
### `14B`: Assignment / Command Semantics Closure
Goal:
1. make assignment transitions and command emission rules explicit from semantic
state rather than runtime convenience
Acceptance object:
1. assignment intent, role application, receiver start, shipper configuration,
and invalidation commands are emitted as bounded semantic decisions
2. one bounded event sequence produces one bounded command sequence
3. command emission does not depend on `weed/` internals
Ownership boundary:
1. `14B` owns command-sequence closure
2. `14B` does not redefine mode/publication ownership from `14A`
3. `14B` does not absorb durable-boundary or recovery closure from `14C`
Status:
1. delivered
### `14C`: Boundary / Recovery Semantic Closure
Goal:
1. make durable boundary and recovery semantics explicit in the same core owner
Acceptance object:
1. boundary truth distinguishes durable progress, checkpoint truth, and
diagnostic sender progress
2. recovery semantics preserve the accepted constraints around eligibility,
fail-closed degradation, and rebuild escalation
3. structural tests stay bounded and do not claim live path migration yet
Ownership boundary:
1. `14C` owns durable-boundary and recovery closure
2. `14C` may affect mode/publication only through explicit boundary/recovery
truth
3. `14C` does not reopen `14A` shell ownership or `14B` command-sequence
closure
Status:
1. delivered
## Manager Review Gate
Every `Phase 14` slice must survive one challenge review that asks:
1. which semantic constraint does this slice satisfy?
2. which overclaim does this slice prevent?
3. which accepted checkpoint proof does this slice preserve?
Reject the slice if any of those questions can only be answered by vague runtime
intuition.
## Immediate Next Step
Phase 14's first bounded core shell is now in place.
The best next step is `Phase 15A`:
1. connect one narrow adapter ingress into the explicit core
2. connect one bounded command path back out
3. prove the live path does not split semantic truth from the new core owner
Some files were not shown because too many files have changed in this diff
Show More
Reference in New Issue
Block a user
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.