Files
seaweedfs/weed/server/volume_grpc_erasure_coding_test.go
T
Chris LuandGitHub 79ac279fe1 fix(ec): don't mix EC shards from different encode runs (#9880)
* feat(ec): add encode_ts_ns to EC shard metadata and the shard read RPC

EcShardConfig and VolumeEcShardReadRequest gain an int64 encode_ts_ns
(encode time in unix nanos). It rides in .vif and the read request so a
read can be scoped to the encode run that produced the index.

* fix(ec): stamp each encode and reject cross-run shard reads

Generate stamps EncodeTsNs into the volume's .vif. Reads carry it to the
shard's owning volume (resolved together via FindEcVolumeWithShard, so a
multi-disk server validates the disk that actually serves the bytes) and
reject a shard from a different encode run, recovering from parity. A
zero on either side (pre-upgrade volume) skips the guard.

* fix(ec): stamp the encode identity on the worker-generated .vif

The worker-local encode path now writes EncodeTsNs (and the resolved EC
ratio) into the .vif, so the read guard is not silently off for volumes
encoded by the maintenance worker.

* fix(ec): wipe stale EC artifacts before re-encoding

VolumeEcShardsGenerate evicts any in-memory EcVolume for the volume and
removes its on-disk shard/index/sidecar files before writing fresh ones,
so a retried encode never builds on a partial prior run and the unlink
frees the inodes instead of leaving open fds serving old bytes.

* fix(ec): unmount EC shards across all disks

UnmountEcShards walked only the first disk holding the shard, leaving a
duplicate copy mounted on a sibling disk (split-disk reconciled volumes)
still serving and heartbeating. Traverse every disk and emit one
deletion delta per disk.

* fix(ec): delete orphan shards without a local .ecx

deleteEcShardIdsForEachLocation gated shard-file removal on a local .ecx,
so it could not clean an orphan .ecNN left by a failed copy on a disk
with no index. Delete the requested shard files unconditionally; the
index-file (.ecx/.ecj/.vif) routing stays gated as before.

* fix(ec): clear stale EC shards cluster-wide before re-encoding

ec.encode unmounts and deletes EC shards for the target volumes on every
node before regenerating: fatal for the shards the topology reports
(mounted leftovers), best-effort for the rest (a sweep that catches
unmounted failed-copy orphans). A down node is a no-op.

* fix(ec): don't nil EC fds on close so reads can't race eviction

A reader resolves an EcVolume/shard under the lock then reads after it is
released, so an eviction that nils ecxFile/ecdFile would race that read
and panic. Close the fds without nilling the fields: the field is now
write-once (no data race) and a concurrent read hits a closed fd, getting
a clean error that the caller recovers from parity.

* fix(ec): wipe stale EC artifacts on every disk and surface failures

The pre-encode wipe only deleted beside the source volume, so a stale
shard on a sibling disk survived and could be mounted against the new
index at reconcile. Sweep every disk. Removal also ignored os.Remove
errors, reporting a failed cleanup as success and letting a stale shard
join the next generation; surface the first real failure (treating
already-gone as success) from removeStaleEcArtifacts and the shard delete.

* fix(ec): log when a local shard is skipped for a different encode run

The cross-run guard returned errShardNotLocal, indistinguishable in logs
from a genuinely-absent shard. Add a V(1) line naming both EncodeTsNs so
operators can tell "wrong encode generation" from "shard not here".

* fix(ec): surface metadata removal failures in the shard delete path

deleteEcShardIdsForEachLocation still dropped os.Remove errors on the
.ecx/.ecj/.vif/sidecar cleanup. A surviving stale .ecx is the orphan-index
condition this path prevents, so route those through removeFileIfExists and
return the first real failure instead of reporting cleanup as success.

* fix(ec): fail orphan cleanup when a reachable node's delete fails

The pre-encode orphan sweep swallowed every error for unreported (node,
volume) pairs. That is only safe for an unreachable node, which cannot
receive this encode's new generation. A reachable node whose delete
genuinely failed (permission/IO) keeps an orphan shard that a later copy
re-stamps with the new run's volume-level .vif identity, so the read guard
would accept stale data. Surface those; stay best-effort only for
unreachable nodes (gRPC Unavailable / no status).

* fix(ec): guard ecjFile under its lock in the EC delete path

EcVolume.Close nils ecjFile under ecjFileAccessLock; a delete that resolved
its .ecx lookup before a concurrent eviction (the generate-time
UnloadEcVolume) could then reach the journal append with a nil fd. Bail
with a clear "volume closed" error under the lock instead.

* fix(ec): reject an unstamped shard when the caller has an encode identity

The read guard required both identities nonzero, so a current (stamped)
caller accepted a holder with identity 0 and could be served a stale
pre-upgrade shard. Reject when the caller is stamped and the holder
differs (including unstamped); stay lenient only when the caller itself
has no identity (pre-upgrade reader). A skipped shard recovers from parity.

* fix(ec): full-teardown delete so cluster cleanup wipes a whole generation

The pre-encode cluster sweep deleted only the listed canonical shards on
remote nodes, leaving index/sidecar (and, on builds with versioned
generations, those too) behind. Add a full_teardown flag to
VolumeEcShardsDelete that evicts the volume and wipes every EC artifact for
it on every disk via removeStaleEcArtifacts; the shell and worker pre-encode
cleanup paths set it. Other delete callers (balance/decode/repair) are
unchanged.

* fix(ec): take ecjFileAccessLock before the nil-check in Sync and Close

Sync and Close read ev.ecjFile before acquiring ecjFileAccessLock while
Close nils it under the lock, a data race on the field. Take the lock
first, then nil-check inside, in both.

* fix(ec): acknowledge full_teardown so a pre-upgrade server can't fake success

An old volume server silently ignores full_teardown and returns success
for an ordinary delete, so the caller wrongly believes the generation was
wiped and copies a fresh gen-0 onto an unwiped node. Echo full_teardown_done
in the response; the worker destination cleanup fails when it is absent, and
the shell cluster sweep fails for a reported (mounted) leftover while staying
best-effort for an unreported node. encode_ts_ns stays an accepted transient
(an old server just skips the new read guard, no regression).

* fix(ec): fail the pre-encode sweep for any reachable node that can't ack teardown

A reachable pre-upgrade server ignores full_teardown and returns success
without wiping an orphan, which a later copy then folds into the new
generation. Treat a missing full_teardown_done ack as fatal for every
reachable node (best-effort only for a gRPC-unreachable one), not just for
topology-reported pairs.

* fix(ec): return the served shard identity and validate it client-side

The encode identity was only enforced server-side, so a pre-upgrade server
ignored the request field and served bytes unchecked. Echo the served
shard's EncodeTsNs on every read response chunk and have the client reject a
mismatch (including 0 from an old server), so the guard holds regardless of
server version; a rejected read recovers from parity.

* fix(ec): reject a short/empty remote shard read instead of serving zeros

doReadRemoteEcShardInterval accepted an immediate EOF or a short stream and
returned success with a partly zero-filled, unvalidated buffer (the server
stamps the identity only on chunks that carry bytes). A non-deleted interval
must arrive whole: require n == len(buf), exempting the is_deleted
short-circuit (n=0), matching readLocalEcShardInterval's local check. A short
read now fails so the caller recovers from parity.

* test(ec): fake volume server echoes the full_teardown acknowledgement

The worker now fails a teardown delete that isn't acknowledged (so a
pre-upgrade server can't silently skip the wipe). The fake server's no-op
VolumeEcShardsDelete returned an empty response, which the worker read as a
skipped teardown and aborted the encode. Echo full_teardown_done.

* feat(ec): mirror the encode-run identity guard + full_teardown into the Rust volume server

The Go volume server stamps an encode-run identity (encode_ts_ns) into the .vif
and rejects a read served from a shard of a different run; full_teardown wipes a
whole generation and acknowledges it. The Rust volume server had none of it.
Mirror the shared logic: load encode_ts_ns from the .vif onto the EcVolume,
stamp it on every read response, and reject a request/response mismatch on both
the server and the distributed-read client (recovering from parity); handle
full_teardown by evicting the volume and wiping every EC artifact on each disk,
echoing full_teardown_done so the caller can detect a server that ignored it.

* fix(ec): remove a stale .vif on full teardown of a shard-only node

A shard copy installs shards + .ecx before .vif, so an interrupted copy after a
teardown could mount the new files under the previous run's identity / version /
shard ratio / dat_file_size carried by the surviving .vif. Remove .vif during
full teardown, gated on .idx absence so a source-volume holder keeps its live
.vif. In Rust this lives in a teardown-only helper so the reconcile / load-
fallback paths (which share the base removal) still preserve .vif.

* fix(ec): treat a missing teardown ack as fatal, not as an unreachable node

isNodeUnreachable returned true for any non-gRPC-status error, so a reachable
pre-upgrade server's missing full_teardown_done ack (a plain error) was
classified unreachable and the unreported pair was silently skipped. Classify
only a real codes.Unavailable as unreachable, and wrap the missing ack in a
sentinel the sweep treats as fatal regardless. A genuinely down node still
surfaces as Unavailable from the RPC and stays best-effort.

* fix(ec): reject a short shard read in the local EC needle reader

read_ec_shard_needle ignored the byte count from shard.read_at and appended the
whole pre-sized buffer, so a truncated shard's zero-filled tail passed the later
length check and parsed as garbage. Require n == buf.len() per interval, erroring
on a short read like the local interval reader already does.

* fix(ec): probe reachability before skipping a node that returns Unavailable

The pre-encode sweep skipped any node whose teardown delete returned
codes.Unavailable, but a reachable volume server in maintenance mode also
returns that code for the maintenance-gated delete, so its stale EC files were
left behind on a node that can still receive the new generation. Confirm with a
non-maintenance-gated empty-target Ping: skip only when the node fails the probe
too (genuinely unreachable).

* fix(ec): use try_exists for the teardown .vif .idx guard

The teardown-only .vif removal gated on Path::exists(), which returns false on a
permission/IO stat error, so a stat failure on a present .idx would read as a
shard-only node and delete the live source volume's .vif. Gate on
try_exists() == Ok(false) instead, preserving the sidecar on any stat error.

* fix(ec): only skip a sweep node when a Ping confirms it is transport-down

The pre-encode sweep skipped a node whenever its teardown delete and a liveness
Ping both failed, but it treated ANY Ping error as down — an application-level
Internal/ResourceExhausted, or Unimplemented from a pre-Ping server, left a
reachable node's stale generation in place. Classify the Ping tri-state and skip
only when it transport-fails with codes.Unavailable; a reachable or inconclusive
node stays fatal.

* fix(ec): exclude sweep-skipped nodes from the encode's rebalance

The pre-encode sweep skips a genuinely-down node best-effort, but the rebalance
then recollected the current topology — a node that recovered between the two
could become a copy target and receive the new generation while still holding
its stale, never-cleared shards. Have the sweep return the skipped set and
exclude those nodes from the rebalance for this encode, so a node we could not
clean cannot receive the new generation. Standalone ec.balance is unaffected.

* fix(ec): re-sweep recovered nodes before generation so they aren't stranded

A node skipped as down by the pre-encode sweep is excluded from the rebalance,
but it can recover and become the generation host — mounting all shards locally,
then being excluded from distribution. Union-only verification accepts all
shards on one node and deletes the originals: a single point of failure. Re-sweep
the skipped nodes just before generation; one whose teardown now succeeds leaves
the skipped set and rebalances normally, while a node still down stays skipped.

* fix(ec): abort the encode if a selected source is still skipped after re-sweep

The re-sweep un-skips a recovered node, but the source was selected before it and
a node can stay down through the re-sweep then recover just in time to be the
generation host — mounting all shards locally while still excluded from the
rebalance, which union-only verification accepts before deleting the originals.
Abort the encode when a selected source remains skipped after the re-sweep.

* fix(ec): batch delete returns retriable 503 when a volume became EC mid-batch

If a volume is not EC at the batch-delete classification but is encoded to EC and
its .dat deleted before the regular-volume mutation, the mutation returns an exact
"not found" that the filer chunk-GC treats as completed, dropping the delete.
Recheck EC presence under the mutation lock and return a retriable 503 with the
"try again" token so the filer requeues it onto the EC path.

* fix(ec): recheck EC state before the regular batch-delete mutation

ec.encode mounts EC shards (copied from the .dat) before deleting the originals,
so a volume can be EC while its .dat still exists. The batch delete only rechecked
EC after a NotFound, so a successful regular-volume delete in that window wrote a
tombstone to the soon-removed .dat — the delete was lost and the needle resurrected
from the pre-tombstone shards. Recheck has_ec_volume under the write lock before
delete_volume_needle and return a retriable 503 so the filer requeues onto the EC path.

* fix(volume): make the metrics push test independent of test order

test_push_metrics_once asserted the pushed body contains the request-counter
family without ever touching the counter — a CounterVec with no children emits
nothing, so the assertion only held when another test had already created a
labelset in the shared registry. Create one in the test itself.
2026-06-10 22:31:18 -07:00

295 lines
9.3 KiB
Go

package weed_server
import (
"context"
"os"
"path/filepath"
"sort"
"testing"
"github.com/seaweedfs/seaweedfs/weed/pb/volume_server_pb"
"github.com/seaweedfs/seaweedfs/weed/stats"
"github.com/seaweedfs/seaweedfs/weed/storage"
"github.com/seaweedfs/seaweedfs/weed/storage/erasure_coding"
"github.com/seaweedfs/seaweedfs/weed/storage/needle"
"github.com/seaweedfs/seaweedfs/weed/storage/types"
"github.com/seaweedfs/seaweedfs/weed/storage/volume_info"
"github.com/seaweedfs/seaweedfs/weed/util"
)
func TestCheckEcVolumeStatusCountOnlyDataShards(t *testing.T) {
tempDir := t.TempDir()
dataDir := filepath.Join(tempDir, "data")
idxDir := filepath.Join(tempDir, "idx")
if err := os.MkdirAll(dataDir, 0o755); err != nil {
t.Fatalf("mkdir data dir: %v", err)
}
if err := os.MkdirAll(idxDir, 0o755); err != nil {
t.Fatalf("mkdir idx dir: %v", err)
}
baseName := "7"
filesToCreate := []string{
filepath.Join(dataDir, baseName+".ec00"),
filepath.Join(dataDir, baseName+".ec09"),
filepath.Join(dataDir, baseName+".ec13"),
filepath.Join(idxDir, baseName+".ecx"),
filepath.Join(idxDir, baseName+".ecj"),
filepath.Join(idxDir, baseName+".idx"),
}
for _, fileName := range filesToCreate {
if err := os.WriteFile(fileName, []byte("x"), 0o644); err != nil {
t.Fatalf("create %s: %v", fileName, err)
}
}
location := &storage.DiskLocation{
Directory: dataDir,
IdxDirectory: idxDir,
}
hasEcxFile, hasIdxFile, shardCount, err := checkEcVolumeStatus(baseName, location)
if err != nil {
t.Fatalf("checkEcVolumeStatus: %v", err)
}
if !hasEcxFile {
t.Fatalf("expected hasEcxFile=true")
}
if !hasIdxFile {
t.Fatalf("expected hasIdxFile=true")
}
if shardCount != 3 {
t.Fatalf("expected shardCount=3, got %d", shardCount)
}
}
// TestRemoveStaleEcArtifacts: a fresh encode deletes every prior EC artifact
// (incl. shard ids beyond the default ratio and versioned bitrot sidecars)
// while leaving the source .dat/.idx/.vif untouched.
func TestRemoveStaleEcArtifacts(t *testing.T) {
tempDir := t.TempDir()
dataDir := filepath.Join(tempDir, "data")
idxDir := filepath.Join(tempDir, "idx")
for _, d := range []string{dataDir, idxDir} {
if err := os.MkdirAll(d, 0o755); err != nil {
t.Fatalf("mkdir %s: %v", d, err)
}
}
const baseName = "7"
dataBase := filepath.Join(dataDir, baseName)
idxBase := filepath.Join(idxDir, baseName)
// EC artifacts that must be removed, including a shard id past the default
// 10+4 ratio (proves the MaxShardCount scan) and a versioned sidecar.
var ecFiles []string
for _, id := range []int{0, 9, 13, 20, erasure_coding.MaxShardCount - 1} {
ecFiles = append(ecFiles, dataBase+erasure_coding.ToExt(id))
}
ecFiles = append(ecFiles,
dataBase+".ecx", dataBase+".ecj",
idxBase+".ecx", idxBase+".ecj",
dataBase+erasure_coding.BitrotSidecarExt, // .ecsum (generation 0)
dataBase+erasure_coding.BitrotSidecarExt+".v2", // versioned sidecar
)
// Source files that must survive — the authoritative input for the encode.
srcFiles := []string{dataBase + ".dat", idxBase + ".idx", dataBase + ".vif"}
for _, f := range append(append([]string{}, ecFiles...), srcFiles...) {
if err := os.WriteFile(f, []byte("x"), 0o644); err != nil {
t.Fatalf("create %s: %v", f, err)
}
}
removeStaleEcArtifacts(dataBase, idxBase, erasure_coding.MaxShardCount)
for _, f := range ecFiles {
if util.FileExists(f) {
t.Errorf("expected EC artifact removed: %s", f)
}
}
for _, f := range srcFiles {
if !util.FileExists(f) {
t.Errorf("expected source file preserved: %s", f)
}
}
}
// TestDeleteEcShardsWithoutLocalEcx: the delete handler removes the requested
// shard files even on a disk with no local .ecx, so a failed-copy orphan stays
// cleanable rather than being mounted later under a foreign index.
func TestDeleteEcShardsWithoutLocalEcx(t *testing.T) {
tempDir := t.TempDir()
dataDir := filepath.Join(tempDir, "data")
idxDir := filepath.Join(tempDir, "idx")
for _, d := range []string{dataDir, idxDir} {
if err := os.MkdirAll(d, 0o755); err != nil {
t.Fatalf("mkdir %s: %v", d, err)
}
}
const baseName = "7"
// Orphan shard files with NO .ecx/.idx anywhere — a failed-copy leftover.
orphans := []string{
filepath.Join(dataDir, baseName+".ec03"),
filepath.Join(dataDir, baseName+".ec11"),
}
for _, f := range orphans {
if err := os.WriteFile(f, []byte("x"), 0o644); err != nil {
t.Fatalf("create %s: %v", f, err)
}
}
location := &storage.DiskLocation{Directory: dataDir, IdxDirectory: idxDir}
if err := deleteEcShardIdsForEachLocation(baseName, location, []uint32{3, 11}); err != nil {
t.Fatalf("deleteEcShardIdsForEachLocation: %v", err)
}
for _, f := range orphans {
if util.FileExists(f) {
t.Errorf("expected orphan shard removed without a local .ecx: %s", f)
}
}
}
// TestVolumeEcShardsInfo_AggregatesAcrossDisks pins the multi-disk path:
// when a volume server mounts EC shards for the same volume on more than
// one disk (each disk holds its own EcVolume entry — Store.FindEcVolume
// returns only the first), VolumeEcShardsInfo used to report shards from
// a single disk. The ec.encode verification step (verifyEcShardsBeforeDelete)
// then refused to delete the source volume because the union across
// servers fell short of dataShards + parityShards. The handler must walk
// every DiskLocation so the response covers every shard the server holds.
func TestVolumeEcShardsInfo_AggregatesAcrossDisks(t *testing.T) {
tempDir := t.TempDir()
dir0 := filepath.Join(tempDir, "disk0")
dir1 := filepath.Join(tempDir, "disk1")
for _, d := range []string{dir0, dir1} {
if err := os.MkdirAll(d, 0o755); err != nil {
t.Fatalf("mkdir %s: %v", d, err)
}
}
const collection = "ec-multi-disk-info"
vid := needle.VolumeId(42)
const dataShards, parityShards = 10, 4
const datSize int64 = 10 * 1024 * 1024
// Two shards on disk0, two on disk1. The .ecx / .ecj / .vif live on
// disk0 so each disk's EcVolume can open the index files via the
// cross-disk fallback in NewEcVolume.
shardsOnDisk0 := []erasure_coding.ShardId{0, 5}
shardsOnDisk1 := []erasure_coding.ShardId{7, 12}
diskIOProbeConfig := stats.DefaultDiskIOProbeConfig()
store := storage.NewStore(nil, "localhost", 8080, 18080, "http://localhost:8080", "store-id",
[]string{dir0, dir1},
[]int32{100, 100},
[]util.MinFreeSpace{{}, {}},
"",
storage.NeedleMapInMemory,
[]types.DiskType{types.HardDriveType, types.HardDriveType},
nil,
3,
diskIOProbeConfig,
)
done := make(chan struct{})
go func() {
for {
select {
case <-store.NewEcShardsChan:
case <-store.NewVolumesChan:
case <-store.DeletedVolumesChan:
case <-store.DeletedEcShardsChan:
case <-store.StateUpdateChan:
case <-done:
return
}
}
}()
t.Cleanup(func() {
store.Close()
close(done)
})
base0 := erasure_coding.EcShardFileName(collection, dir0, int(vid))
base1 := erasure_coding.EcShardFileName(collection, dir1, int(vid))
// .ecx, .ecj, .vif live on disk0. NewEcVolume on disk1 falls back to
// disk0's idx dir. The .ecx needs >0 bytes so HasEcxFileOnDisk does
// not treat it as the corrupt-stub case; a single zero entry is the
// smallest valid index file (WalkIndex iterates one zero-sized needle).
if err := os.WriteFile(base0+".ecx", make([]byte, 16), 0o644); err != nil {
t.Fatalf("write .ecx: %v", err)
}
if err := os.WriteFile(base0+".ecj", nil, 0o644); err != nil {
t.Fatalf("write .ecj: %v", err)
}
if err := volume_info.SaveVolumeInfo(base0+".vif", &volume_server_pb.VolumeInfo{
Version: uint32(needle.Version3),
DatFileSize: datSize,
EcShardConfig: &volume_server_pb.EcShardConfig{
DataShards: dataShards,
ParityShards: parityShards,
},
}); err != nil {
t.Fatalf("save .vif: %v", err)
}
plant := func(base string, shardId erasure_coding.ShardId) {
t.Helper()
f, err := os.Create(base + erasure_coding.ToExt(int(shardId)))
if err != nil {
t.Fatalf("create shard %d: %v", shardId, err)
}
// MountEcShards does not validate shard size, so any non-empty
// truncate avoids the zero-byte ignore branch in loadAllEcShards.
if err := f.Truncate(1); err != nil {
f.Close()
t.Fatalf("truncate shard %d: %v", shardId, err)
}
f.Close()
}
for _, sid := range shardsOnDisk0 {
plant(base0, sid)
}
for _, sid := range shardsOnDisk1 {
plant(base1, sid)
}
for _, sid := range append([]erasure_coding.ShardId{}, append(shardsOnDisk0, shardsOnDisk1...)...) {
if err := store.MountEcShards(collection, vid, sid, ""); err != nil {
t.Fatalf("MountEcShards %d.%d: %v", vid, sid, err)
}
}
vs := &VolumeServer{store: store}
resp, err := vs.VolumeEcShardsInfo(context.Background(), &volume_server_pb.VolumeEcShardsInfoRequest{
VolumeId: uint32(vid),
})
if err != nil {
t.Fatalf("VolumeEcShardsInfo: %v", err)
}
gotShardIds := make([]int, 0, len(resp.GetEcShardInfos()))
for _, info := range resp.GetEcShardInfos() {
if info.GetVolumeId() != uint32(vid) {
t.Errorf("EcShardInfo VolumeId=%d, want %d", info.GetVolumeId(), vid)
}
gotShardIds = append(gotShardIds, int(info.GetShardId()))
}
sort.Ints(gotShardIds)
want := []int{0, 5, 7, 12}
if len(gotShardIds) != len(want) {
t.Fatalf("VolumeEcShardsInfo returned %d shards (ids=%v), want %d (ids=%v)",
len(gotShardIds), gotShardIds, len(want), want)
}
for i, sid := range want {
if gotShardIds[i] != sid {
t.Fatalf("VolumeEcShardsInfo shard ids=%v, want %v", gotShardIds, want)
}
}
}