pingqiu and Claude Opus 4.6
8ecc506452
V2 stabilization: 144/144 hardware actions PASS + design docs + SmartWAL prototype
...
Hardware scenarios (all PASS on m01/m02, 25Gbps RoCE):
- I-V3 auto-failover: 43/43 (create→write→kill→promote→verify IO)
- I-R8 rebuild-rejoin: 58/58 (failover→write→restart→1GB rebuild in 2s→verify data)
- Fast rejoin: 43/43 (kill replica→3s→restart→recovery→data verified)
Performance: V2 RF=1 = 46,666 IOPS vs V1.5 RF=1 = 47,233 IOPS (-1.2%, noise)
New test scenarios:
- v2-rebuild-rejoin.yaml: full failover→rebuild→second failover→data integrity
- v2-fast-rejoin-catchup.yaml: replica kill→fast restart→recovery
- v2-rebuild-failure-retry.yaml: kill during rebuild→restart→data verified
- rf1-perf-compare.yaml: RF=1 perf baseline for V1.5 vs V2 comparison
Design documents:
- protocol-anti-patterns.md: 7 anti-patterns with cases from SeaweedFS/Ceph/DRBD
- smartwal-design-memo.md: extent-first write algorithm research (BlueStore/ZFS/DRBD)
- smartwal-prototype-spec.md: prototype spec with 16/16 crash tests PASS
- v3-clean-recovery-draft.md: V3 semantic cleanup principles
- v2-integration-matrix.md: 25-row integration coverage map
- v2-acceptance-evidence.md: gap analysis for remaining work
SmartWAL prototype (16/16 tests PASS):
- smartwal.go, smartwal_record.go, smartwal_recovery.go: core implementation
- smartwal_test.go: 9 single-node crash tests
- smartwal_repl_test.go: 7 two-node replication crash tests
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-04-11 00:18:20 -07:00
pingqiu and Claude Opus 4.6
53246d2780
fix: recover TOCTOU + WAL pressure edge case tests
...
Fix recover path TOCTOU: re-Lookup after AddReplica so the primary
refresh assignment includes the freshly added replica addresses.
Previously, Lookup (copy) was called before AddReplica modified the
registry, so entry.Replicas was empty → primary got replicas=0 →
shipper never configured.
Add 2 WAL pressure edge case tests:
- ShipperCatchUpOrEscalate: 64KB WAL, 200 writes, aggressive flusher.
Proves no hang/deadlock/corruption. Shipper either keeps up or
correctly escalates to NeedsRebuild.
- RebuildWithPinWhilePrimaryWrites: rebuild session active while
primary writes 7600+ blocks in 2s. Proves primary never freezes
— rebuild pin is on replica only, primary WAL recycles freely.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-04-08 23:56:26 -07:00
pingqiu and Claude Opus 4.6
39f1232fe2
feat: validation matrix closure — Rebuild Ready 12/12, Restore Ready 10/10
...
Close all Rebuild Ready and Restore Ready matrix gaps. V2 Ready at 10/14
(2 partial, 2 missing — honest assessment).
New tests (tester-written):
- R1: syncAck-driven trigger via protocol engine decision
- R3: stale replica restart beyond WAL → rebuild converges
- R5: connection drop mid-base → cancel → fresh rebuild converges
- R10: failover-rejoin with forced WAL recycling, strict rebuild assert
- R11: divergent replica full overwrite convergence
- R12: crash mid-rebuild → fresh session converges (not resume)
- S2: corrupt WAL entry + corrupt base block both rejected
- S5: snapshot-tail rebuild (base + WAL tail replay)
- S7: crash between base install and tail replay
- S8: snapshot under concurrent writes
- V5: rebuild complete without DurableLSN blocks publish_healthy
- V9: mixed replica health aggregate projection
- V14: negative fail-closed matrix (epoch, kind, stale)
Bug fix: StartRebuildSession now clears stale dirty map + resets WAL +
updates checkpoint AFTER safety check but BEFORE session.Start(). Fixes
stale extent data shadowing rebuild base blocks on reopened replicas.
Cleanup: remove 14 obsolete design docs (migration batches, old WAL-v2
specs, simulator goals) — all superseded by current protocol docs.
34 component tests + 8 protocol engine tests + server tests all pass.
1GB CRC validation passes in 19s.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-04-08 16:31:55 -07:00
pingqiu and Claude Opus 4.6
59a36013d4
feat: rebuild hardening A1-A5 + session-controlled execution path
...
A1 Engine kind-routing fix:
SessionProgressObserved/Completed/Failed now respect active session
Kind. Rebuild progress no longer leaks into catch-up aggregate.
sessionKindMismatch guard + observeRebuildProgress helper.
2 regression tests lock kind isolation.
A2 Retention pin:
Rebuild session ack drives progress-based WAL retention floor.
Pin installed at base_lsn on accepted, advances with wal_applied_lsn,
released on completed/failed/cancelled. rebuildProgressPinFloor
returns min across all active replicas.
Retention pin test: 100 blocks fill WAL, 5 flusher cycles with
20 pinned rebuild entries — all verified correct.
A3 Progress ack emission:
Automatic sessionAck(running/base_complete/completed/failed) emitted
from rebuild session lifecycle transitions. sessionAckLocked builds
ack under session lock. emitRebuildSessionAck callback wired through
SetOnRebuildSessionAck on BlockVol.
ObserveReplicaRebuildSessionAck maps acks to core engine events.
WireLocalReplicaRebuildSessionAcks bridges local callback to server.
5 server tests proving ack→core, pin advance, pin cleanup.
A4 Deadline/timeout:
rebuildAckWatch watchdog: armed on accepted/running/base_complete,
refreshed on each ack, cleared on completed/failed. Timeout
cancels local session + clears pin + fail-closes.
2 tests: timeout→fail-close, progress→refresh.
A5 Session-controlled execution path:
v2bridge.Executor.TransferFullBase now uses session-controlled loop:
beginControlledFullBase → real sessionControl over TCP →
transferExtentToSession via RebuildTransportClient →
PrepareFullBaseRebuild → TryCompleteRebuildSession.
ReplicaReceiver control channel handles MsgSessionControl alongside
MsgBarrierReq. Session acks written back on same TCP connection.
RebuildSessionBase request type separates new per-block stream from
legacy raw extent stream. Full-base cleanup deferred until success.
Deadlock fix: ApplyBaseBlock releases session lock before ioMu.
Hydration skip for full-base sessions.
23 rebuild component tests (all pass):
11 kernel correctness, 8 transport/runtime, 3 scenario-scale,
including 1GB primary-initiated with CRC validation.
29 files changed, ~2500 insertions.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-04-08 14:39:11 -07:00
pingqiu and Claude Opus 4.6
d2d57851b0
feat: rebuild MVP — dual-lane session with bitmap protection
...
Rebuild session protocol implementation for v2-rebuild-mvp-session-protocol.md.
New files:
- rebuild_bitmap.go: RebuildBitmap — session-scoped dense bitset for
WAL-applied LBA tracking. MarkApplied on local WAL write (not receive).
ShouldApplyBase returns false for WAL-covered LBAs (WAL always wins).
- rebuild_session.go: RebuildSession — replica-side two-line rebuild.
WAL lane (ApplyWALEntry) + base lane (ApplyBaseBlock) with bitmap
conflict resolution. TryComplete requires BOTH base_complete AND
wal_applied_lsn >= target_lsn. Volume-level control surface:
StartRebuildSession, ApplyRebuildSessionWALEntry/BaseBlock,
MarkRebuildSessionBaseComplete, TryCompleteRebuildSession,
CancelRebuildSession, ActiveRebuildSession.
- rebuild_mvp_test.go: 4 correctness tests — base+WAL converge,
WAL-applied never overwritten by base, bitmap set on applied not
received, control surface start/supersede/complete.
- rebuild_transport_test.go: 2 transport-level tests — two-line with
real WAL shipping, live writes during base copy with bitmap conflict.
Design docs:
- v2-rebuild-mvp-session-protocol.md: MVP spec with message set, apply
rules, completion/failure/crash rules, test matrix
- v2-sync-recovery-protocol.md: full protocol context (keepup/catchup/
rebuild unified design, primary decision logic, two-line model)
- v2-session-protocol-shape.md: protocol shape overview
Protocol engine (reference, not production):
- sw-block/protocol/: 7-event engine with ~300 lines, 13 tests
6 rebuild tests pass, all existing component tests pass.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-04-07 14:30:34 -07:00
pingqiu and Claude Opus 4.6
44103a1bd7
feat: Phase 20 acceptance fixes + sw-test-runner suite mode
...
Acceptance rows closed:
- WriteLBA/SyncCache contract: code comments document write-back vs
durability fence semantics
- RF=2 stable identity: v2bridge always uses SetReplicaAddrs (preserves
ServerID); blockcmd dispatcher also fixed to use setupPrimaryReplicationMulti;
test asserts exact expected ReplicaID="vs-2" (not just non-empty)
- Tests treating WriteLBA as commit: replica_read_test rewritten with
SyncCache as durability fence
- publish_healthy contract: 3 gate tests with hard assertions including
gate 3 (PrimaryShipperConnected)
- SetReplicaAddr deprecation warning added
- WALShipper.ReplicaID() getter added for identity verification
Test runner enhancements:
- sw-test-runner suite command: build → deploy → run N scenarios in one
invocation with --skip-deploy support
- Suite YAML definitions for T6 Stage 0 and Stage 1
- deploy action: kill stale processes, clean dirs, cross-compile, upload
- run-phase20-t6.ps1 PowerShell script (deprecated by suite command)
Engine/runtime fixes:
- Recovery executor nil-safety improvements
- Recovery bundle BuildRecoveryBundle defensive checks
- ShipperGroup MinReplicaFlushedLSNAll surface
Docs: acceptance checklist refined, test matrix updated, T6 runbook,
engine maintainer tutorial, design README updated.
26 files changed, ~1600 insertions.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-04-06 11:30:54 -07:00
pingqiu and Claude Opus 4.6
d1a16fac03
feat: protocol-aware execution wave — phase gate for live WAL shipping
...
Add host-side protocol state seam that derives per-replica execution
state from V2 sender/session snapshots and blocks live-tail WAL
shipping while an active recovery session is in progress.
New file: weed/server/block_protocol_state.go
- replicaProtocolExecutionState derived from engine snapshots
- LiveEligible=false during active catch-up/rebuild sessions
- bindProtocolExecutionPolicy wires policy into BlockVol
- syncProtocolExecutionState called after assignments + core events
Data plane changes:
- WALShipper.Ship() checks liveShippingPolicy before dial/send
- BlockVol.SetLiveShippingPolicy persists across shipper group rebuilds
- ShipperGroup propagates policy to all shippers
Design contract: sw-block/design/v2-protocol-aware-execution.md
Scope: WAL-first rollout only. Prevents illegal live-tail delivery
during active recovery. Does not change snapshot/build behavior or
move backlog. Next wave: bounded WAL catch-up under same contract.
Tests: 4 unit/component tests for phase gate behavior, plus bootstrap
seam tests that confirmed the two pre-existing bugs locally.
13 files changed, 900 insertions, 69 deletions.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-04-05 23:47:07 -07:00
pingqiu and Claude Opus 4.6
cf16e53b04
feat: Phase 16M/17 + promote fixes + testrunner updates
...
Phase 16M: explicit replica readiness on heartbeat seam
- master.proto: optional bool replica_ready = 19 (proto regenerated on M01)
- block_heartbeat_proto.go: write/read ReplicaReady with presence semantics
- master_block_registry.go: replicaReadyObservedFromHeartbeat prefers
explicit proto field, falls back to address heuristic when absent
- volume_server_block.go: heartbeat emits ReplicaReady from core projection
Phase 17: host effects extraction + stop line
- phase-17-log.md: Batch 10/11 delivery notes
Promote fixes:
- master_block_failover.go: deterministic replica addrs from path hash
- qa_promote_replication_test.go: address-upgrade trigger test
- qa_promote_rejoin_live_test.go: new live rejoin test
Testrunner:
- devops.go: action improvements
- recovery-baseline-failover.yaml, suite-ha-failover.yaml: scenario updates
- cp11b3-manual-promote.yaml: promote scenario alignment
- fresh_volume_write_test.go: new component test
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-04-05 11:38:05 -07:00
pingqiu and Claude Opus 4.6
4c7fbefe25
feat: CP13-8 PASSES — real-workload validation on RF=2 sync_all
...
CP13-8 scenario results on m01/M02 (25Gbps RoCE):
fsck_ext4: CLEAN
file count: 200 (assert_equal PASS)
checksum match: MATCH (assert_contains PASS)
pgbench TPS: 565.69 (assert_greater PASS)
auto-failover: 10.0.0.1:18480 → 10.0.0.3:18480
Code changes (tester + scenario):
- volume_server_block.go: readiness state, assignment lifecycle cleanup
- block_heartbeat_loop.go: readiness-aware heartbeat reporting
- store_blockvol.go: readiness tracking
- master_server_handlers_block.go: block API handler updates
- cp13-8-real-workload-validation.yaml: redesigned scenario
(removed block_promote, use natural auto-failover flow,
bootstrap write before wait_volume_healthy)
- testrunner/actions/devops.go: scenario action improvements
- replica_read_test.go: component-level replica read test
Phase docs: CP13-7 accepted, CP13-8/8A technical packs updated,
design docs updated for protocol closure evidence.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-04-03 14:24:13 -07:00
pingqiu and Claude Opus 4.6
ebe95b6e2e
fix: flusher OOM on multi-block writes + testrunner enhancements
...
Bug: flusher.go:336 allocated make([]byte, entryLen) per dirty block
instead of per unique WAL entry. A 4MB WriteLBA creates 1024 dirty map
entries (one per 4KB block), all sharing the same WAL offset. The flusher
read the full 4MB WAL entry 1024 times into separate buffers:
1024 × 4MB = 4GB per 4MB write → OOM on mkfs.ext4.
Root cause: flusher assumed 1:1 dirty-block-to-WAL-entry mapping.
WriteLBA supports multi-block writes but the flusher never deduplicated
shared WAL offsets.
Fix: deduplicate WAL reads by WalOffset in flushOnceLocked(). Multiple
dirty blocks from the same WAL entry share one read buffer and one
DecodeWALEntry call. Memory: O(WAL_entries × size) not O(blocks × size).
For a 4MB write: 4GB → 4MB.
Verified on hardware (m01/M02 25Gbps RoCE):
- Before: mkfs.ext4 → VS RSS 100MB→25GB → OOM killed
- After: mkfs.ext4 → VS RSS 129MB stable, mkfs succeeds
- pgbench TPC-B c=4: 1,248 TPS (RF=1, previously blocked by OOM)
Tests added:
- flusher_test.go: flush_multiblock_shared_wal_read (16 blocks share
one WAL offset, flush dedup verified)
- flusher_test.go: flush_multiblock_data_correct (3 mixed multi-block
writes, all data correct after flush)
- test/component/large_write_test.go: 7 component tests (single 4MB,
sequential mkfs sim, concurrent, mixed sizes, production volume,
flusher throughput 30s sustained)
- iscsi/large_write_mem_test.go: 2 iSCSI session memory tests (4MB
R2T flow, slow device)
Testrunner enhancements (same commit — all tested on hardware):
- discover_primary action: maps primary IP → topology node name,
supports alt_ips for multi-NIC (RoCE + management)
- NodeSpec.AltIPs field for multi-NIC node identification
- 5 new YAML scenarios: ec3, ec5, degraded sync_all/best_effort, pgbench
- All 13 hardware-verified scenarios PASS
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-04-02 14:24:10 -07:00
pingqiu and Claude Opus 4.6
785a7d7efd
feat: wire real pinner into flusher retention + real WAL scan executor (Phase 07 P1)
...
Pinner wired to real retention:
- NewPinner calls vol.SetV2RetentionFloor(p.MinWALRetentionFloor)
- Flusher.RetentionFloorFn() / SetRetentionFloorFn() exposed
- SetV2RetentionFloor chains with existing shipper retention floor
- Holds actually prevent WAL reclaim (not just tracked state)
Executor uses real WAL scan:
- BlockVol.ScanWALEntries(fromLSN, callback) wraps wal.ScanFrom
with real fd, walOffset, checkpointLSN
- Executor.StreamWALEntries uses ScanWALEntries (not stub)
- Reads real WAL entries, tracks highest LSN scanned
CommittedLSN mapping:
- Explicitly documented as interim V1 model (committed = checkpointed)
- Will diverge when V2 distributed commit separates from local flush
Carry-forward:
- TransferSnapshot/TransferFullBase/TruncateWAL: stubs (need extent I/O)
- Control intent from confirmed failover: deferred
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-03-30 20:01:46 -07:00