mirror of
https://github.com/seaweedfs/seaweedfs.git
synced 2026-08-29 12:17:08 +00:00
* ec: uniform shard block layout An EC volume is striped as 1GiB blocks until less than one row remains, then 1MiB blocks, and consecutive blocks land on different shards. With ec.encode's -fullPercent 95 against the 30GiB default limit, ~30% of every volume sits in that 1MiB tail, so a 4MB filer chunk there is five stripes on five servers. New encodes now use one block per shard, sized ceil(datSize/dataShards) rounded up to 1MiB and recorded in the .vif (EcShardConfig.block_size, also carried by the .ecsum manifest). A needle now maps to one shard unless it is larger than the block or straddles a boundary. The chosen size equals the legacy layout's padded shard length for every input, so shard sizes, capacity math, and the shard-size credibility checks are unchanged; only the byte placement moved. Reads, decode, and scrub resolve the block sizes from the volume's .vif; absence keeps the legacy interpretation, so existing EC volumes read exactly as before. Rebuild is layout-agnostic. weed fix -ecx recovers the layout from the .vif, else the .ecsum sidecar, and with neither de-stripes under both candidate layouts and keeps the one that indexes more valid needles. Same change in the Rust volume server, which now also streams the encode in 256KB sub-batches like Go instead of allocating whole blocks, and computes the large-row count as shardSize/largeBlock to match Go on exact multiples. On a 26MB fixture both encoders produce byte-identical shards, and a Go-written .vif parses in Rust with the block size intact. * ec: resolve the rust ecx rebuild through the recorded layout The Rust rebuild path regenerated a lost .ecx by scanning the logical .dat through a hand-rolled pure-1MiB striping, which was already wrong for legacy volumes with large-block rows and is wrong for any uniform volume with a block past 1MiB. Route the scan through locate_data with the .vif-recorded block size, the same mapping the read path uses. Also seed the new tests' random data instead of the deprecated global math/rand.Read. * ec: fail the Rust ecx rebuild on any shard read error A read error mid-scan published the entries collected so far as a successful .ecx, and read_at's byte count was ignored so a legal short read passed as complete — a truncated or failing shard could produce a silently incomplete recovery index. Exact-read semantics in read_from_data_shards, error propagation in the needle walk, and a truncated-shard regression test. * ec: fail the mount on an unreadable or malformed vif Both servers silently fell back to the legacy layout when an existing .vif could not be read or parsed. Every new encode records a positive uniform block size there, so the fallback mounted the same shards with legacy offset math and could return wrong data. Absent stays legal (legacy volumes predate the sidecar), and a zero-byte stub still reads as absent (Go's MaybeLoadVolumeInfo convention, now mirrored in Rust); a present-but-unreadable or malformed .vif fails the mount instead. * ec: bound the reconstruct fan-out of one needle's intervals A degraded interval fans out a read to every reachable shard location, each with a buffer the size of the interval. Reading a needle's intervals in parallel multiplied that by the interval concurrency: a needle spanning 8 blocks could hold 8 x MaxShardCount remote reads and buffers at once, where the sequential version peaked at MaxShardCount. Give each needle a single reconstruct budget its intervals share, held for the buffer's lifetime, so separate reads stay independent but one read cannot multiply its own fan-out. * ec: drop the duplicated shard-size formula calculateExpectedShardSize reimplemented the padding rule that UniformBlockSize already owns — TestUniformBlockSizeMatchesLegacyShardSize asserts the two agree for every input — so a change to the rule would have had to be made in both. Defer to the helper, keeping the historic answer for an empty .dat. * ec: resolve the shard block layout from whatever records it Four places still answered the layout question by inference when a record of it was available, or accepted an answer that was not one: - A mount with no .vif defaulted to the legacy layout; the bitrot sidecar records the same config at encode time, so take it when present, as weed fix -ecx already does. The vif itself is now parsed once per mount rather than twice. - The Rust ecx rebuild derived its row count from the padded shard extent, which under the legacy layout reads a shard that is an exact large-block multiple as one row too many. Pass the encode-time .dat size from the .vif and keep the extent as the fallback. - weed fix -ecx read the block size outside the EC-config guard (collapsing the unknown sentinel into a definitive legacy), only wrote the recovered layout back when the .vif was absent rather than unusable, and broke a scan tie by candidate order instead of the documented reach. - The uniform layout tripped writeDatFile's large-block ambiguity guard, which cannot apply when the large and small blocks are the same size. * ec: give the index-recovery tests a parseable vif The fixtures wrote the literal bytes "volinfo" as the source .vif and the recovery copies it verbatim, so the receiving server then mounted the volume from a .vif it could not parse. That used to pass by silently defaulting to the legacy layout; a mount now refuses a vif it cannot read, which is what the tests were exercising all along without meaning to. * ec: validate the layout a vif records, not just its syntax Review follow-ups on the mount-strictness change: - A .vif can parse and still record a block size no encoder could have produced (negative, or not a whole number of small blocks). Both servers took it and mapped every read through it. ValidateBlockSize / the Rust mirror now refuse the mount, the same way an unparseable vif does; 0 stays valid as the legacy two-tier layout. - The bitrot-sidecar fallback accepted parity_shards == 0 and summed the counts in their own width, so values near the ceiling wrapped past the MaxShardCount bound. Require both counts and sum in a wider type. - weed fix -ecx treated a config with only DataShards > 0 as usable, so a half-written .vif suppressed the recovery paths AND survived the rewrite. Require a complete, in-range config before trusting it. - Returning the vif-load error left the .ecx and .ecj descriptors open; repeated mount attempts on malformed metadata could exhaust them. * ec: refuse to act on a layout the metadata does not establish - The worker encode only logged a failed .vif write and skipped it in the distribution set, and treated the .ecsum write as best-effort. A worker whose disk filled after the much larger shards landed could still distribute, mount, verify shard inventory, and delete the source replicas — leaving holders with shards whose geometry nothing records. Both writes and both inclusions are encode success conditions now. - A generation-matching .ecsum that disagreed with the .vif geometry only disabled checksums in Go, and in Rust was not compared at all, so protection stayed On while reads used the other layout. Both files record the layout their generation was encoded with, so a disagreement now fails the mount. * ec: reject an invalid recorded block size in weed fix -ecx A .vif with valid shard counts but a negative or unaligned block size was marked usable: a positive invalid value pinned the scan to a geometry that de-stripes to garbage, and a negative one ran the dual scan but left the invalid .vif in place afterwards. Validate it with the same rule the mount applies, and when it fails leave the layout unknown so the scan recovers it and the file is rewritten. * ec: validate the sidecar layout weed fix -ecx recovers from The .ecsum fallback was taken on DataShards > 0 alone, so a CRC-valid sidecar carrying the wrong generation, an incomplete ratio, or an unaligned block size would pin the reconstruction to one incorrect uniform-layout candidate instead of letting the dual scan decide. Require generation 0, a complete in-range ratio, and a valid block size; anything less leaves the layout unknown, which is the answer that still recovers by scanning. * ec: let only a genuinely absent sidecar choose the legacy layout With no .vif the bitrot sidecar is the only record of a volume's layout, and the mount fallback read a failed load, an unusable config, or a sidecar stamped for another generation as "assume legacy". A uniform generation-0 volume could therefore mount with legacy or another generation's geometry and answer reads with the wrong bytes. Present-but-unusable now fails the mount; only actual absence keeps the legacy defaults. Shared as EcShardConfigFromSidecar so every caller reads the sidecar the same way. * ec: treat a recorded-but-impossible layout as corruption, not as legacy - A .vif whose ecShardConfig is PRESENT but records an impossible ratio was answered with the default 10+4 and the legacy block layout, in both languages. That reads a uniform volume's shards at the wrong offsets and returns the wrong bytes. Only an entirely absent config still means "this predates the record"; a present one that cannot be true fails the mount. - The shard-count bound summed two uint32 counts as int, which wraps on a 32-bit build: 0x7fffffff + 0x7fffffff lands at -2 and slips under MaxShardCount. ValidEcShardCounts sums in uint64, and every EC call site that checked a recorded ratio now goes through it. * ec: rebuild on the geometry the sidecar records, and flag it when it disagrees The rebuild RPC passes BackgroundECContext, so RebuildEcFiles resolves the layout itself — and it resolved a missing or invalid .vif to the default 10+4 with the legacy block size. Two consequences: a 12+4 volume was reconstructed through a 10+4 matrix, which produces wrong bytes and never regenerates shards 14-15; and the chosen geometry then contradicted a valid uniform sidecar, which loadRebuildSidecar reported as BitrotOff — silently skipping the input and regenerated-shard checksum checks precisely when the volume had already lost its metadata. The layout now resolves from the bitrot sidecar (found across the server's disks, not just beside the base name) before falling back to the defaults, and a present-but-impossible ratio fails instead of being replaced. A sidecar that contradicts the chosen geometry is BitrotInvalid, which the existing unsafeIgnoreSidecar override still lets an operator push past. * ec: let the Rust rebuild read metadata off a sibling disk read_ec_shard_config searches only the location the rebuild writes into, so a volume whose .vif or generation-0 .ecsum sits on another of the server's disks resolved to the default 10+4 with the legacy block layout — the Rust half of the geometry-guessing the Go rebuild just stopped doing. It then reconstructs a custom-ratio or uniform volume through the wrong Reed-Solomon matrix and de-striping geometry. The rebuild now looks for the .vif in its own location and then each sibling, falls back to the generation-0 sidecar wherever that lives, and only defaults when neither exists anywhere. The encode-time .dat size the ecx rebuild needs is resolved the same way. * ec: resolve a rebuild's vif from every directory that may hold it RebuildEcFiles probed only <data-base>.vif. The caller knows the selected location's index directory and the sibling locations, but passed neither for metadata: additionalDirs carried shard directories only, and were searched for shards and the checksum sidecar. A split -dir/-dir.idx layout, or a disk holding only shards, therefore resolved a pre-sidecar custom-ratio volume to 10+4 and reconstructed through the wrong matrix — never regenerating shards 14-15. The caller now hands over the index and sibling directories, and the resolver probes the vif across all of them, matching what the Rust resolver already does for both the vif and the sidecar. * ec: make every rebuild consumer agree on the layout it resolved - The post-rebuild bitrot backfill re-derived the geometry from this directory's .vif alone and dropped the block size entirely, so a rebuild that resolved its layout from a sibling, the sidecar, or a uniform vif wrote a manifest describing a DIFFERENT layout — one later mounts reject, or that covers only the default shard count. The layout is resolved once now, through an exported ResolveRebuildECContext, and the rebuild and the backfill share that answer. - The Rust rebuild collected only each location's data directory, so a sibling's INDEX directory — where a split -dir/-dir.idx layout keeps .ecx/.ecj/.vif — was never probed, and a custom-ratio volume still resolved to 10+4 with the legacy layout. Both directories of every location are carried now, deduped against the rebuild's own. - A shard delivery can bring the checksum manifest with it, but the receive path only writes the file: a server that already had the volume mounted kept its resolved protection state (off) until a remount. The mount RPC re-resolves it once the shards it describes have been added. * ec: cover the rebuild's directory search with tests Reviewers flagged the sibling index directory twice, and the fix that closed it had no test of its own: the assembly sat inline in the rebuild handler, reachable only through a gRPC call against a populated store. Lifting it into rebuildSearchDirs / select_rebuild_location makes the rule assertable — a sibling contributes BOTH its data and its index directory, a shared index directory is listed once, and the rebuild's own data directory never repeats. Writing the Rust cases surfaced that the two implementations do not agree on where the rebuild's own index directory belongs, and both are right: Go's resolver takes a single directory list, so that directory has to be inside it, while Rust's takes the rebuild's data and index directories as their own arguments and would search them twice. The tests now state which contract each side is holding to, so neither drifts into the other's shape. Pure refactor otherwise; no behaviour change. * ec: search the index directory for the layout sidecar The Rust resolver looked for the generation-0 .ecsum in the rebuild's data directory and the sibling list, but not in the rebuild's own index directory — while the .vif lookup directly above it did, and Go's findBitrotSidecar has always checked both bases. On a split -dir/-dir.idx location that directory is where the metadata lives, and callers leave it out of the sibling list precisely because it is passed here separately, so nothing searched it. With no .vif anywhere the sidecar is the only surviving record of the layout. Missing it resolved a 12+4 uniform volume to 10+4 with the legacy striping — the test added here fails with (10, 4, 0) against the old code — and the rebuild then reconstructs through the wrong matrix and writes .ecx offsets that no reader can follow. * ec: let the rebuild see its own index directory The Rust rebuild takes a single flat directory list — the shape Go's RebuildEcFiles uses — so it cannot be handed the rebuild location's index directory separately the way the layout resolvers are, and the handler was passing the sibling list, which deliberately omits exactly that directory. On a split -dir/-dir.idx location that is where .ecx and .vif live, so the shard and index lookups could not see them. Go has always carried that directory in additionalDirs; this lines the two call sites up. * ec: let a config-free vif fall through to the layout sidecar A .vif that carries no ecShardConfig answers nothing about the layout, so it is no more informative than an absent one — but both trees treated its mere existence as the end of the search. Go went straight to the 10+4 legacy defaults without consulting the sidecar at all; Rust returned whatever ec_shard_config_from could make of a single directory. A 12+4 uniform volume with a legacy config-free vif therefore resolved as 10+4 legacy, and every read landed at the wrong shard offset. The sidecar lookup was also single-directory on both sides, while a split -dir/-dir.idx layout keeps .vif and .ecsum with the INDEX. Go's findBitrotSidecar has always taken both bases; the callers here passed only the data base, and the Rust bitrot resolver derived its path from the data base alone. Rust's layout resolver now takes a candidate directory list — data, index, then any siblings — and searches all of it, which also removes the early return that made the vif's presence decisive. load_vif_info_across_dirs reported `dir` even when load_vif_info had found the vif in `dir_idx`. Nothing reads that field today, so this changes no behaviour; it stops the next caller that resolves the rest of the volume's metadata against the answer from being sent to a disk holding none of it. Absence stays legal throughout: a volume with neither record is genuinely legacy. Present-but-unusable still fails the mount, now in the config-free-vif branch too. * ec: activate a delivered sidecar on every per-disk runtime A vid mounts as one EcVolume per disk, each with its own resolved protection state, but the post-delivery reload used the first-match lookup and so touched exactly one of them. The siblings kept reporting no protection until a remount — and since shard distribution deduplicates the metadata files onto the first target disk for a node, the runtime that got the .ecsum is not necessarily the one the lookup returns. Iterate every runtime instead, via a new FindAllEcVolumes and its Rust mut equivalent. Combined with each runtime now resolving its sidecar against its index directory as well as its data directory, a server sharing one -dir.idx across its disks activates all of them from the single delivered copy. The Rust volume server had no post-mount reload at all; it gets one here, matching Go. * ec: resolve the delivered sidecar across every EC metadata directory Reloading every per-disk runtime, added last round, did not by itself make the delivered manifest reachable. Startup mirroring copies .ecx/.ecj/.vif to every shard-bearing disk so each mounts self-contained, but deliberately not .ecsum, and a repair delivers exactly one copy. Each runtime was resolving against its own two directories, so every sibling of the disk that received the file kept reporting no protection however often it reloaded. Resolve one authoritative copy across every EC metadata directory instead of duplicating the file. Mirroring .ecsum would have to keep pace with a file that is rewritten as shards are repaired, and would not help the reported case at all: the delivery happens at runtime, and mirroring only runs at startup. The regression test pins both halves — a reload restricted to the volume's own directories still finds nothing, and the same reload given the server's metadata directories turns protection on. * ec: ask every directory before writing a TOFU baseline After a rebuild the opportunistic backfill asks whether this volume already has a checksum manifest, and answered from the data base alone. A split -dir/-dir.idx layout keeps the sidecar with the index, and a multi-disk server may keep it on a sibling, so an existing manifest read as absent. The consequence is worse than a missed read. On a false "no" the backfill writes a fresh sidecar at the data base from whatever the shards say right now — and the data base is the first candidate every resolver checks, so that TOFU baseline shadows the real manifest rather than sitting beside it. A shard that was silently corrupt gets blessed, and the record that would have caught it stops being consulted. FindBitrotSidecar exports the search the package already used internally, so the question is asked of the data base, the index base and the sibling disks — the same candidates the rebuild resolves its layout from. * ec: refuse a shard block size no encoder could have produced weed fix -ecx derived one from the raw shard extent, so a truncated or partially copied shard wrote a .vif that NewEcVolume then permanently refuses — the volume the tool was run to rescue could never mount again. An extent that is not a whole number of small blocks cannot have come from a uniform encode, so it is no longer offered as a candidate, and nothing unvalidated reaches the .vif. Claude-Session: https://claude.ai/code/session_011FRRoNKBiGbH58rs2AQyA7 * ec: derive the .vif's dat size and block size from one measurement VolumeEcShardsGenerate stat'ed the .dat before the encode while WriteEcFiles stat'ed it again to size the blocks. A write landing between the two produced a .vif whose own two fields describe different files. WriteEcFiles now leaves both on the context, and fills a placeholder context in place so the caller can read them back. Claude-Session: https://claude.ai/code/session_011FRRoNKBiGbH58rs2AQyA7 * ec: keep the source volume until every holder serves its shard layout The uniform layout rides in a .vif field older volume servers never knew: they discard it, mount the shards as legacy and return wrong bytes with nothing erroring, and the shard files are the same length either way so no other check notices. The upgrade order lived only in the release note. VolumeEcShardsInfo now reports the block size the holder actually serves, in both the Go and Rust servers, and the pre-delete verification refuses to drop the source unless every reachable holder echoes the one the shards were encoded with — while a rollback still exists. A server that predates the field answers 0, which is the negative answer. Claude-Session: https://claude.ai/code/session_011FRRoNKBiGbH58rs2AQyA7 * ec: drop the rebuild's dead block-size parameters generateMissingEcFiles never reads largeBlockSize/smallBlockSize — Reed-Solomon reconstruction is layout-agnostic — so passing the legacy constants only advertised a layout the rebuild does not use. Also move UniformBlockSize's doc off ValidateBlockSize. Claude-Session: https://claude.ai/code/session_011FRRoNKBiGbH58rs2AQyA7 * ec: warn about EC defaults only when the mount used them The "vif file not found, using defaults" warning fired even after the bitrot sidecar supplied a non-default layout, sending anyone triaging wrong bytes after the legacy layout the volume never mounted on. Claude-Session: https://claude.ai/code/session_011FRRoNKBiGbH58rs2AQyA7 * ec: stat the distributed bitrot sidecar once The strict check re-stat'ed the file immediately before the stat that already gates inclusion, and a failed sidecar write now fails the encode outright, so the first could only fire on a deletion between the two lines. Claude-Session: https://claude.ai/code/session_011FRRoNKBiGbH58rs2AQyA7 * ec: say what the reconstruct budget actually bounds A shard's buffer stays in bufs until its interval reconstructs, which is after the read that filled it released its permit, so the semaphore bounds round trips in flight and not retained bytes. Peak memory is the intervals reconstructing at once times the shards each reaches times the interval size. Claude-Session: https://claude.ai/code/session_011FRRoNKBiGbH58rs2AQyA7 * test: let the fake volume server report its delivered EC layout The pre-delete verification now asks each holder which shard block layout it serves, and a fake that always answered "unset" looked exactly like a volume server too old to know the field. Distribution ships the .vif to every holder alongside its shards, so read the layout back out of it as a real holder does. Claude-Session: https://claude.ai/code/session_011FRRoNKBiGbH58rs2AQyA7
1461 lines
51 KiB
Rust
1461 lines
51 KiB
Rust
//! EC encoding: convert a .dat file into 10 data + 4 parity shards.
|
|
//!
|
|
//! Uses Reed-Solomon erasure coding. The .dat file is split into blocks
|
|
//! (1GB large, 1MB small) and encoded across 14 shard files.
|
|
|
|
use std::fs::File;
|
|
use std::io;
|
|
#[cfg(not(unix))]
|
|
use std::io::{Read, Seek, SeekFrom};
|
|
|
|
use reed_solomon_erasure::galois_8::ReedSolomon;
|
|
|
|
use crate::pb::volume_server_pb::{
|
|
ChecksumAlgorithm, EcBitrotProtection, EcShardChecksums,
|
|
};
|
|
use crate::storage::erasure_coding::ec_bitrot::{
|
|
self, ShardChecksumBuilder, DEFAULT_BITROT_BLOCK_SIZE,
|
|
};
|
|
use crate::storage::erasure_coding::ec_shard::*;
|
|
use crate::storage::idx;
|
|
use crate::storage::types::*;
|
|
use crate::storage::volume::volume_file_name;
|
|
|
|
/// Encode a .dat file into EC shard files.
|
|
///
|
|
/// Creates .ec00-.ec13 files in the same directory.
|
|
/// Also creates a sorted .ecx index from the .idx file.
|
|
///
|
|
/// Always encodes with the uniform block layout, sized for this .dat, and
|
|
/// returns the block size so the caller can persist it to .vif. Mirrors Go's
|
|
/// WriteEcFiles.
|
|
pub fn write_ec_files(
|
|
dir: &str,
|
|
idx_dir: &str,
|
|
collection: &str,
|
|
volume_id: VolumeId,
|
|
data_shards: usize,
|
|
parity_shards: usize,
|
|
) -> io::Result<i64> {
|
|
let base = volume_file_name(dir, collection, volume_id);
|
|
let dat_path = format!("{}.dat", base);
|
|
let idx_base = volume_file_name(idx_dir, collection, volume_id);
|
|
let idx_path = format!("{}.idx", idx_base);
|
|
|
|
// Create sorted .ecx from .idx
|
|
write_sorted_ecx_from_idx(&idx_path, &format!("{}.ecx", base))?;
|
|
|
|
// Encode .dat into shards
|
|
let dat_file = File::open(&dat_path)?;
|
|
let dat_size = dat_file.metadata()?.len() as i64;
|
|
|
|
let rs = ReedSolomon::new(data_shards, parity_shards)
|
|
.map_err(|e| io::Error::new(io::ErrorKind::Other, format!("reed-solomon init: {:?}", e)))?;
|
|
|
|
// Create shard files
|
|
let total_shards = data_shards + parity_shards;
|
|
let mut shards: Vec<EcVolumeShard> = (0..total_shards as u8)
|
|
.map(|i| EcVolumeShard::new(dir, collection, volume_id, i))
|
|
.collect();
|
|
|
|
for shard in &mut shards {
|
|
shard.create()?;
|
|
}
|
|
|
|
// Per-shard bitrot checksum builders: accumulate a CRC32C for every
|
|
// DEFAULT_BITROT_BLOCK_SIZE block of each shard's byte stream as it is
|
|
// written, so the resulting `.ecsum` sidecar can later detect silent
|
|
// corruption in any shard (including cold parity).
|
|
let mut builders: Vec<ShardChecksumBuilder> = (0..total_shards)
|
|
.map(|_| ShardChecksumBuilder::new(DEFAULT_BITROT_BLOCK_SIZE as i64))
|
|
.collect();
|
|
|
|
let block_size = uniform_block_size(dat_size, data_shards);
|
|
encode_dat_file(
|
|
&dat_file,
|
|
dat_size,
|
|
&rs,
|
|
&mut shards,
|
|
&mut builders,
|
|
data_shards,
|
|
parity_shards,
|
|
ENCODE_BUFFER_SIZE,
|
|
block_size as usize,
|
|
block_size as usize,
|
|
)?;
|
|
|
|
// Close all shards
|
|
for shard in &mut shards {
|
|
shard.close();
|
|
}
|
|
|
|
// Write the generation-0 bitrot sidecar (`<base>.ecsum`). Finalizing each
|
|
// builder yields covered_size (== total bytes written to that shard) and
|
|
// the packed little-endian CRC32C array.
|
|
let mut shard_checksums: Vec<EcShardChecksums> = Vec::with_capacity(total_shards);
|
|
for (i, builder) in builders.into_iter().enumerate() {
|
|
let (covered_size, packed) = builder.finalize();
|
|
shard_checksums.push(EcShardChecksums {
|
|
shard_id: i as u32,
|
|
covered_size,
|
|
block_crc32c: packed,
|
|
});
|
|
}
|
|
let prot = EcBitrotProtection {
|
|
algorithm: ChecksumAlgorithm::ChecksumCrc32c as i32,
|
|
block_size: DEFAULT_BITROT_BLOCK_SIZE as u32,
|
|
generation: 0,
|
|
ec_shard_config: Some(ec_bitrot::ec_shard_config(
|
|
data_shards as u32,
|
|
parity_shards as u32,
|
|
block_size,
|
|
)),
|
|
shards: shard_checksums,
|
|
encode_uuid: ec_bitrot::new_encode_uuid(),
|
|
};
|
|
let sidecar_path = ec_bitrot::bitrot_sidecar_path(&base, 0);
|
|
if let Err(e) = ec_bitrot::save_bitrot_sidecar(&sidecar_path, &prot) {
|
|
// A failed sidecar must not fail the encode — the shards are already
|
|
// written and valid. The volume simply runs with bitrot protection
|
|
// off for this generation until the sidecar is regenerated.
|
|
tracing::warn!(
|
|
volume_id = volume_id.0,
|
|
path = %sidecar_path,
|
|
error = %e,
|
|
"ec encode: failed to write bitrot sidecar; protection off for this generation",
|
|
);
|
|
}
|
|
|
|
Ok(block_size)
|
|
}
|
|
|
|
/// uniform_block_size returns the per-shard block size of the uniform layout
|
|
/// for a .dat of the given size: ceil(dat_file_size/data_shards) rounded up to
|
|
/// a whole small block. For every input this equals the legacy layout's padded
|
|
/// shard size, so only the byte placement differs between the two layouts,
|
|
/// never the shard length. Mirrors Go's UniformBlockSize.
|
|
pub fn uniform_block_size(dat_file_size: i64, data_shards: usize) -> i64 {
|
|
let small = ERASURE_CODING_SMALL_BLOCK_SIZE as i64;
|
|
let per_shard = (dat_file_size + data_shards as i64 - 1) / data_shards as i64;
|
|
let blocks = ((per_shard + small - 1) / small).max(1);
|
|
blocks * small
|
|
}
|
|
|
|
/// Rebuild missing EC shard files from existing shards using Reed-Solomon reconstruct.
|
|
///
|
|
/// This does not require the `.dat` file, only the existing `.ecXX` shard files.
|
|
///
|
|
/// `additional_dirs` lists sibling disk locations to search for existing shards when
|
|
/// they are not found in `dir`. This is required on multi-disk volume servers where
|
|
/// shards for the same volume may be spread across disks.
|
|
pub fn rebuild_ec_files(
|
|
dir: &str,
|
|
collection: &str,
|
|
volume_id: VolumeId,
|
|
missing_shard_ids: &[u32],
|
|
data_shards: usize,
|
|
parity_shards: usize,
|
|
additional_dirs: &[&str],
|
|
) -> io::Result<()> {
|
|
if missing_shard_ids.is_empty() {
|
|
return Ok(());
|
|
}
|
|
|
|
let rs = ReedSolomon::new(data_shards, parity_shards)
|
|
.map_err(|e| io::Error::new(io::ErrorKind::Other, format!("reed-solomon init: {:?}", e)))?;
|
|
|
|
let total_shards = data_shards + parity_shards;
|
|
let mut shards: Vec<EcVolumeShard> = (0..total_shards as u8)
|
|
.map(|i| EcVolumeShard::new(dir, collection, volume_id, i))
|
|
.collect();
|
|
|
|
// Determine the exact shard size from the first available existing shard.
|
|
// When a shard is not found in `dir`, search `additional_dirs` (sibling disks
|
|
// on the same volume server) before failing — mirrors Go's findShardFile logic.
|
|
let mut shard_size = 0;
|
|
for (i, shard) in shards.iter_mut().enumerate() {
|
|
if !missing_shard_ids.contains(&(i as u32)) {
|
|
if let Ok(_) = shard.open() {
|
|
let size = shard.file_size();
|
|
if size > shard_size {
|
|
shard_size = size;
|
|
}
|
|
} else {
|
|
// Try sibling disk locations before giving up.
|
|
let mut found = false;
|
|
for &other_dir in additional_dirs {
|
|
let mut alt = EcVolumeShard::new(other_dir, collection, volume_id, i as u8);
|
|
if let Ok(_) = alt.open() {
|
|
let size = alt.file_size();
|
|
if size > shard_size {
|
|
shard_size = size;
|
|
}
|
|
*shard = alt;
|
|
found = true;
|
|
break;
|
|
}
|
|
}
|
|
if !found {
|
|
return Err(io::Error::new(
|
|
io::ErrorKind::NotFound,
|
|
format!("missing non-rebuild shard {}", i),
|
|
));
|
|
}
|
|
}
|
|
}
|
|
}
|
|
|
|
if shard_size == 0 {
|
|
return Err(io::Error::new(
|
|
io::ErrorKind::InvalidData,
|
|
"all existing shards are empty or cannot find an existing shard to determine size",
|
|
));
|
|
}
|
|
|
|
// Create the missing shards for writing
|
|
for i in missing_shard_ids {
|
|
if let Some(shard) = shards.get_mut(*i as usize) {
|
|
shard.create()?;
|
|
}
|
|
}
|
|
|
|
let block_size = ERASURE_CODING_SMALL_BLOCK_SIZE;
|
|
let mut remaining = shard_size;
|
|
let mut offset: u64 = 0;
|
|
|
|
// Process all data in blocks
|
|
while remaining > 0 {
|
|
let to_process = remaining.min(block_size as i64) as usize;
|
|
|
|
// Allocate buffers for all shards. Option<Vec<u8>> is required by rs.reconstruct()
|
|
let mut buffers: Vec<Option<Vec<u8>>> = vec![None; total_shards];
|
|
|
|
// Read available shards. A short read means a truncated/corrupt input
|
|
// shard; treat it as an error rather than reconstructing over a
|
|
// zero-padded tail and publishing the result as restored redundancy.
|
|
for (i, shard) in shards.iter().enumerate() {
|
|
if !missing_shard_ids.contains(&(i as u32)) {
|
|
let mut buf = vec![0u8; to_process];
|
|
let n = shard.read_at(&mut buf, offset)?;
|
|
if n != to_process {
|
|
return Err(io::Error::new(
|
|
io::ErrorKind::UnexpectedEof,
|
|
format!(
|
|
"ec rebuild short read shard {} at {}: got {} want {}",
|
|
i, offset, n, to_process
|
|
),
|
|
));
|
|
}
|
|
buffers[i] = Some(buf);
|
|
}
|
|
}
|
|
|
|
// Reconstruct missing shards
|
|
rs.reconstruct(&mut buffers).map_err(|e| {
|
|
io::Error::new(
|
|
io::ErrorKind::Other,
|
|
format!("reed-solomon reconstruct: {:?}", e),
|
|
)
|
|
})?;
|
|
|
|
// Write recovered data into the missing shards
|
|
for i in missing_shard_ids {
|
|
let idx = *i as usize;
|
|
if let Some(buf) = buffers[idx].take() {
|
|
shards[idx].write_all(&buf)?;
|
|
}
|
|
}
|
|
|
|
offset += to_process as u64;
|
|
remaining -= to_process as i64;
|
|
}
|
|
|
|
// Close all shards
|
|
for shard in &mut shards {
|
|
shard.close();
|
|
}
|
|
|
|
Ok(())
|
|
}
|
|
|
|
/// Reed-Solomon parity check over the locally-held shards: recompute parity from
|
|
/// the data shards and flag any shard whose bytes disagree. Wired into FULL
|
|
/// (scrub mode 2) as a TEMPORARY local parity/cold-region check — the per-needle
|
|
/// FULL walk only reads live data-shard intervals, so on its own it can't catch
|
|
/// bitrot in a parity shard or an unwalked region. Move to mode 4 (CHECKSUM) and
|
|
/// drop it from mode 2 once the `.ecsum` subsystem lands.
|
|
pub fn verify_ec_shards(
|
|
dir: &str,
|
|
collection: &str,
|
|
volume_id: VolumeId,
|
|
data_shards: usize,
|
|
parity_shards: usize,
|
|
) -> io::Result<(Vec<u32>, Vec<String>)> {
|
|
let rs = ReedSolomon::new(data_shards, parity_shards)
|
|
.map_err(|e| io::Error::new(io::ErrorKind::Other, format!("reed-solomon init: {:?}", e)))?;
|
|
|
|
let total_shards = data_shards + parity_shards;
|
|
let mut shards: Vec<EcVolumeShard> = (0..total_shards as u8)
|
|
.map(|i| EcVolumeShard::new(dir, collection, volume_id, i))
|
|
.collect();
|
|
|
|
let mut shard_size = 0;
|
|
let mut broken_shards = std::collections::HashSet::new();
|
|
let mut details = Vec::new();
|
|
|
|
for (i, shard) in shards.iter_mut().enumerate() {
|
|
if let Ok(_) = shard.open() {
|
|
let size = shard.file_size();
|
|
if size > shard_size {
|
|
shard_size = size;
|
|
}
|
|
} else {
|
|
broken_shards.insert(i as u32);
|
|
details.push(format!("failed to open or missing shard {}", i));
|
|
}
|
|
}
|
|
|
|
if shard_size == 0 || broken_shards.len() >= parity_shards {
|
|
// Can't do much if we don't know the size or have too many missing
|
|
return Ok((broken_shards.into_iter().collect(), details));
|
|
}
|
|
|
|
let block_size = ERASURE_CODING_SMALL_BLOCK_SIZE;
|
|
let mut remaining = shard_size;
|
|
let mut offset: u64 = 0;
|
|
|
|
while remaining > 0 {
|
|
let to_process = remaining.min(block_size as i64) as usize;
|
|
let mut buffers = vec![vec![0u8; to_process]; total_shards];
|
|
|
|
let mut read_failed = false;
|
|
for i in 0..total_shards {
|
|
if !broken_shards.contains(&(i as u32)) {
|
|
if let Err(e) = shards[i].read_at(&mut buffers[i], offset) {
|
|
broken_shards.insert(i as u32);
|
|
details.push(format!("read error shard {}: {}", i, e));
|
|
read_failed = true;
|
|
}
|
|
} else {
|
|
read_failed = true;
|
|
}
|
|
}
|
|
|
|
// Only do verification if all shards were readable
|
|
if !read_failed {
|
|
// Need to convert Vec<Vec<u8>> to &[&[u8]] for rs.verify
|
|
let slice_ptrs: Vec<&[u8]> = buffers.iter().map(|v| v.as_slice()).collect();
|
|
if let Ok(is_valid) = rs.verify(&slice_ptrs) {
|
|
if !is_valid {
|
|
// Reed-Solomon verification failed. We cannot easily pinpoint which shard
|
|
// is corrupted without recalculating parities or syndromes, so we just
|
|
// log that this batch has corruption. Wait, we can test each parity shard!
|
|
// Let's re-encode from the first `data_shards` and compare to the actual `parity_shards`.
|
|
|
|
let mut verify_buffers = buffers.clone();
|
|
// Clear the parity parts
|
|
for i in data_shards..total_shards {
|
|
verify_buffers[i].fill(0);
|
|
}
|
|
if rs.encode(&mut verify_buffers).is_ok() {
|
|
for i in 0..total_shards {
|
|
if buffers[i] != verify_buffers[i] {
|
|
broken_shards.insert(i as u32);
|
|
details.push(format!(
|
|
"parity mismatch on shard {} at offset {}",
|
|
i, offset
|
|
));
|
|
}
|
|
}
|
|
}
|
|
}
|
|
}
|
|
}
|
|
|
|
offset += to_process as u64;
|
|
remaining -= to_process as i64;
|
|
}
|
|
|
|
// Close all shards
|
|
for shard in &mut shards {
|
|
shard.close();
|
|
}
|
|
|
|
let mut broken_vec: Vec<u32> = broken_shards.into_iter().collect();
|
|
broken_vec.sort_unstable();
|
|
|
|
Ok((broken_vec, details))
|
|
}
|
|
|
|
/// Write sorted .ecx index from .idx file.
|
|
pub(crate) fn write_sorted_ecx_from_idx(idx_path: &str, ecx_path: &str) -> io::Result<()> {
|
|
if !std::path::Path::new(idx_path).exists() {
|
|
return Err(io::Error::new(
|
|
io::ErrorKind::NotFound,
|
|
"idx file not found",
|
|
));
|
|
}
|
|
|
|
// Read all idx entries
|
|
let mut idx_file = File::open(idx_path)?;
|
|
let mut entries: Vec<(NeedleId, Offset, Size)> = Vec::new();
|
|
|
|
idx::walk_index_file(&mut idx_file, 0, |key, offset, size| {
|
|
entries.push((key, offset, size));
|
|
Ok(())
|
|
})?;
|
|
|
|
// Sort by NeedleId, then by actual offset so later entries come last
|
|
entries.sort_by_key(|&(key, offset, _)| (key, offset.to_actual_offset()));
|
|
|
|
// Remove duplicates (keep last/latest entry for each key).
|
|
// dedup_by_key keeps the first in each run, so we reverse first,
|
|
// dedup, then reverse back.
|
|
entries.reverse();
|
|
entries.dedup_by_key(|entry| entry.0);
|
|
entries.reverse();
|
|
|
|
// Write sorted entries to .ecx
|
|
let mut ecx_file = File::create(ecx_path)?;
|
|
for &(key, offset, size) in &entries {
|
|
idx::write_index_entry(&mut ecx_file, key, offset, size)?;
|
|
}
|
|
|
|
Ok(())
|
|
}
|
|
|
|
/// Rebuild the .ecx index file by walking needles in the EC data shards.
|
|
///
|
|
/// This is the equivalent of Go's `RebuildEcxFile`. It reads the logical .dat
|
|
/// content from the EC data shards, walks through needle headers to extract
|
|
/// (needle_id, offset, size) entries, deduplicates them, and writes a sorted
|
|
/// .ecx index file.
|
|
///
|
|
/// `additional_dirs` lists sibling disk locations to search for data shards when
|
|
/// they are not found in `dir` — mirrors the same fallback used by rebuild_ec_files,
|
|
/// required on multi-disk volume servers where shards may be spread across disks.
|
|
pub fn rebuild_ecx_file(
|
|
dir: &str,
|
|
collection: &str,
|
|
volume_id: VolumeId,
|
|
data_shards: usize,
|
|
block_size: i64,
|
|
dat_file_size: i64,
|
|
additional_dirs: &[&str],
|
|
) -> io::Result<()> {
|
|
use crate::storage::needle::needle::get_actual_size;
|
|
use crate::storage::super_block::SUPER_BLOCK_SIZE;
|
|
|
|
let base = volume_file_name(dir, collection, volume_id);
|
|
let ecx_path = format!("{}.ecx", base);
|
|
|
|
// Open data shards to read logical .dat content. When a shard isn't found
|
|
// in `dir`, search `additional_dirs` (sibling disks on the same volume
|
|
// server) before giving up.
|
|
let mut shards: Vec<EcVolumeShard> = (0..data_shards as u8)
|
|
.map(|i| EcVolumeShard::new(dir, collection, volume_id, i))
|
|
.collect();
|
|
|
|
for (i, shard) in shards.iter_mut().enumerate() {
|
|
if let Err(_) = shard.open() {
|
|
let mut found = false;
|
|
for &other_dir in additional_dirs {
|
|
let mut alt = EcVolumeShard::new(other_dir, collection, volume_id, i as u8);
|
|
if alt.open().is_ok() {
|
|
*shard = alt;
|
|
found = true;
|
|
break;
|
|
}
|
|
}
|
|
if !found {
|
|
// If a data shard is missing, we can't rebuild ecx
|
|
for s in &mut shards {
|
|
s.close();
|
|
}
|
|
return Err(io::Error::new(
|
|
io::ErrorKind::NotFound,
|
|
format!("cannot open data shard for ecx rebuild"),
|
|
));
|
|
}
|
|
}
|
|
}
|
|
|
|
// Determine total logical data size from shard sizes
|
|
let shard_size = shards.iter().map(|s| s.file_size()).max().unwrap_or(0);
|
|
let total_data_size = shard_size as i64 * data_shards as i64;
|
|
// The volume's shard block layout: the .vif-recorded uniform block size,
|
|
// or the legacy two-tier sizes when 0. The row count comes from the shard
|
|
// length; -1 disambiguates a legacy shard that is an exact large-block
|
|
// multiple (mirrors the ecdFileSize-1 fallback in the read path).
|
|
let (large_block, small_block) = if block_size > 0 {
|
|
(block_size, block_size)
|
|
} else {
|
|
(
|
|
ERASURE_CODING_LARGE_BLOCK_SIZE as i64,
|
|
ERASURE_CODING_SMALL_BLOCK_SIZE as i64,
|
|
)
|
|
};
|
|
// The row count the de-stripe walks with. The encode-time .dat size is the
|
|
// authority — the same value the read path divides by data_shards — and the
|
|
// padded extent is only a fallback: under the legacy layout a shard that is
|
|
// an exact large-block multiple reads as one row too many, which
|
|
// re-interprets its last large row as small blocks and scrambles the
|
|
// recovered offsets. Subtracting one keeps that fallback on the safe side of
|
|
// the boundary, exactly as the read path's own fallback does.
|
|
let locate_shard_size = if dat_file_size > 0 {
|
|
dat_file_size / data_shards as i64
|
|
} else {
|
|
(shard_size as i64 - 1).max(0)
|
|
};
|
|
|
|
// Read version from superblock (first byte of logical data)
|
|
let mut sb_buf = [0u8; SUPER_BLOCK_SIZE];
|
|
read_from_data_shards(
|
|
&shards,
|
|
&mut sb_buf,
|
|
0,
|
|
data_shards,
|
|
locate_shard_size,
|
|
large_block,
|
|
small_block,
|
|
)?;
|
|
let version = Version(sb_buf[0]);
|
|
|
|
// Walk needles starting after superblock
|
|
let mut offset = SUPER_BLOCK_SIZE as i64;
|
|
let header_size = NEEDLE_HEADER_SIZE;
|
|
let mut entries: Vec<(NeedleId, Offset, Size)> = Vec::new();
|
|
|
|
while offset + header_size as i64 <= total_data_size {
|
|
// Read needle header (cookie + needle_id + size = 16 bytes).
|
|
// A read failure is NOT the end of the data — every offset in
|
|
// range maps into the shards, so an error means a truncated or
|
|
// unreadable shard. Publishing the entries collected so far as
|
|
// a successful .ecx would hand out a silently incomplete
|
|
// recovery index; propagate instead. (The scan still ends
|
|
// normally on the zero-cookie tail below.)
|
|
let mut header_buf = [0u8; NEEDLE_HEADER_SIZE];
|
|
if let Err(e) = read_from_data_shards(
|
|
&shards,
|
|
&mut header_buf,
|
|
offset as u64,
|
|
data_shards,
|
|
locate_shard_size,
|
|
large_block,
|
|
small_block,
|
|
) {
|
|
for s in &mut shards {
|
|
s.close();
|
|
}
|
|
return Err(io::Error::new(
|
|
e.kind(),
|
|
format!("scan needle header at offset {}: {}", offset, e),
|
|
));
|
|
}
|
|
|
|
let cookie = Cookie::from_bytes(&header_buf[..COOKIE_SIZE]);
|
|
let needle_id = NeedleId::from_bytes(&header_buf[COOKIE_SIZE..COOKIE_SIZE + NEEDLE_ID_SIZE]);
|
|
let size = Size::from_bytes(&header_buf[COOKIE_SIZE + NEEDLE_ID_SIZE..header_size]);
|
|
|
|
// Validate: stop if we hit zero cookie+id (end of data)
|
|
if cookie.0 == 0 && needle_id.0 == 0 {
|
|
break;
|
|
}
|
|
|
|
// Validate size is reasonable
|
|
if size.0 < 0 && !size.is_deleted() {
|
|
break;
|
|
}
|
|
|
|
let actual_size = get_actual_size(size, version);
|
|
if actual_size <= 0 || offset + actual_size > total_data_size {
|
|
break;
|
|
}
|
|
|
|
entries.push((needle_id, Offset::from_actual_offset(offset), size));
|
|
|
|
// Advance to next needle (aligned to NEEDLE_PADDING_SIZE)
|
|
offset += actual_size;
|
|
let padding_rem = offset % NEEDLE_PADDING_SIZE as i64;
|
|
if padding_rem != 0 {
|
|
offset += NEEDLE_PADDING_SIZE as i64 - padding_rem;
|
|
}
|
|
}
|
|
|
|
for shard in &mut shards {
|
|
shard.close();
|
|
}
|
|
|
|
// Sort by NeedleId, then by offset (later entries override earlier)
|
|
entries.sort_by_key(|&(key, offset, _)| (key, offset.to_actual_offset()));
|
|
|
|
// Deduplicate: keep latest entry per needle_id
|
|
entries.reverse();
|
|
entries.dedup_by_key(|entry| entry.0);
|
|
entries.reverse();
|
|
|
|
// Write sorted .ecx
|
|
let mut ecx_file = File::create(&ecx_path)?;
|
|
for &(key, offset, size) in &entries {
|
|
idx::write_index_entry(&mut ecx_file, key, offset, size)?;
|
|
}
|
|
ecx_file.sync_all()?;
|
|
|
|
Ok(())
|
|
}
|
|
|
|
/// Read bytes from EC data shards at a logical offset in the .dat file,
|
|
/// resolving the shard/offset through the volume's block layout via
|
|
/// locate_data — the same mapping the read path uses.
|
|
#[allow(clippy::too_many_arguments)]
|
|
fn read_from_data_shards(
|
|
shards: &[EcVolumeShard],
|
|
buf: &mut [u8],
|
|
logical_offset: u64,
|
|
data_shards: usize,
|
|
locate_shard_size: i64,
|
|
large_block_size: i64,
|
|
small_block_size: i64,
|
|
) -> io::Result<()> {
|
|
let intervals = crate::storage::erasure_coding::ec_locate::locate_data(
|
|
logical_offset as i64,
|
|
Size(buf.len() as i32),
|
|
locate_shard_size,
|
|
data_shards as u32,
|
|
large_block_size,
|
|
small_block_size,
|
|
);
|
|
let mut bytes_read = 0usize;
|
|
for interval in &intervals {
|
|
let (shard_id, shard_offset) =
|
|
interval.to_shard_id_and_offset(data_shards as u32, large_block_size, small_block_size);
|
|
if shard_id as usize >= data_shards {
|
|
return Err(io::Error::new(
|
|
io::ErrorKind::InvalidInput,
|
|
"shard index out of range",
|
|
));
|
|
}
|
|
let to_read = interval.size as usize;
|
|
let dest = &mut buf[bytes_read..bytes_read + to_read];
|
|
// Exact-read semantics: read_at may legally return fewer bytes
|
|
// than requested, and treating a short read as complete leaves
|
|
// the tail of `dest` as whatever the buffer held before. Loop
|
|
// until filled; zero bytes inside the mapped range means the
|
|
// shard is truncated — an error, not an end.
|
|
let mut filled = 0usize;
|
|
while filled < to_read {
|
|
let n = shards[shard_id as usize]
|
|
.read_at(&mut dest[filled..], shard_offset as u64 + filled as u64)?;
|
|
if n == 0 {
|
|
return Err(io::Error::new(
|
|
io::ErrorKind::UnexpectedEof,
|
|
format!(
|
|
"short read from data shard {}: {} of {} bytes at offset {}",
|
|
shard_id, filled, to_read, shard_offset
|
|
),
|
|
));
|
|
}
|
|
filled += n;
|
|
}
|
|
bytes_read += to_read;
|
|
}
|
|
if bytes_read != buf.len() {
|
|
return Err(io::Error::new(
|
|
io::ErrorKind::UnexpectedEof,
|
|
"short read from data shards",
|
|
));
|
|
}
|
|
Ok(())
|
|
}
|
|
|
|
/// Buffer size for one encode sub-batch per shard, mirroring Go's 256KB
|
|
/// bufferSize in WriteEcFiles. A block is processed in block_size/buffer_size
|
|
/// sub-batches, so memory stays at total_shards * 256KB no matter how large
|
|
/// the uniform block is.
|
|
const ENCODE_BUFFER_SIZE: usize = 256 * 1024;
|
|
|
|
/// Encode the .dat file data into shard files.
|
|
///
|
|
/// Uses a two-phase approach matching Go's ec_encoder.go:
|
|
/// 1. Process as many large blocks as possible
|
|
/// 2. Process remaining data with small blocks
|
|
///
|
|
/// `buffer_size` must divide both block sizes.
|
|
#[allow(clippy::too_many_arguments)]
|
|
pub(crate) fn encode_dat_file(
|
|
dat_file: &File,
|
|
dat_size: i64,
|
|
rs: &ReedSolomon,
|
|
shards: &mut [EcVolumeShard],
|
|
builders: &mut [ShardChecksumBuilder],
|
|
data_shards: usize,
|
|
parity_shards: usize,
|
|
buffer_size: usize,
|
|
large_block_size: usize,
|
|
small_block_size: usize,
|
|
) -> io::Result<()> {
|
|
let total_shards = data_shards + parity_shards;
|
|
let mut buffers: Vec<Vec<u8>> = (0..total_shards)
|
|
.map(|_| vec![0u8; buffer_size])
|
|
.collect();
|
|
|
|
let mut remaining = dat_size;
|
|
let mut offset: u64 = 0;
|
|
|
|
// Phase 1: process whole large-block rows while enough data remains
|
|
let large_row_size = large_block_size * data_shards;
|
|
|
|
while remaining >= large_row_size as i64 {
|
|
encode_data(
|
|
dat_file,
|
|
offset,
|
|
large_block_size,
|
|
rs,
|
|
&mut buffers,
|
|
shards,
|
|
builders,
|
|
data_shards,
|
|
)?;
|
|
offset += large_row_size as u64;
|
|
remaining -= large_row_size as i64;
|
|
}
|
|
|
|
// Phase 2: process remaining data with small blocks
|
|
let small_row_size = small_block_size * data_shards;
|
|
|
|
while remaining > 0 {
|
|
let to_process = remaining.min(small_row_size as i64);
|
|
encode_data(
|
|
dat_file,
|
|
offset,
|
|
small_block_size,
|
|
rs,
|
|
&mut buffers,
|
|
shards,
|
|
builders,
|
|
data_shards,
|
|
)?;
|
|
offset += to_process as u64;
|
|
remaining -= to_process;
|
|
}
|
|
|
|
Ok(())
|
|
}
|
|
|
|
/// Encode one row of blocks, streaming it in ENCODE_BUFFER_SIZE sub-batches so
|
|
/// arbitrarily large blocks never require block-sized allocations. Mirrors
|
|
/// Go's encodeData.
|
|
#[allow(clippy::too_many_arguments)]
|
|
fn encode_data(
|
|
dat_file: &File,
|
|
row_offset: u64,
|
|
block_size: usize,
|
|
rs: &ReedSolomon,
|
|
buffers: &mut [Vec<u8>],
|
|
shards: &mut [EcVolumeShard],
|
|
builders: &mut [ShardChecksumBuilder],
|
|
data_shards: usize,
|
|
) -> io::Result<()> {
|
|
let buffer_size = buffers[0].len();
|
|
if block_size % buffer_size != 0 {
|
|
return Err(io::Error::new(
|
|
io::ErrorKind::InvalidInput,
|
|
format!(
|
|
"unexpected block size {} buffer size {}",
|
|
block_size, buffer_size
|
|
),
|
|
));
|
|
}
|
|
let batch_count = block_size / buffer_size;
|
|
for b in 0..batch_count {
|
|
encode_one_batch(
|
|
dat_file,
|
|
row_offset + (b * buffer_size) as u64,
|
|
block_size,
|
|
rs,
|
|
buffers,
|
|
shards,
|
|
builders,
|
|
data_shards,
|
|
)?;
|
|
}
|
|
Ok(())
|
|
}
|
|
|
|
/// Encode one sub-batch: the same buffer-sized slice of every shard's block in
|
|
/// this row. Mirrors Go's encodeDataOneBatch.
|
|
#[allow(clippy::too_many_arguments)]
|
|
fn encode_one_batch(
|
|
dat_file: &File,
|
|
offset: u64,
|
|
block_size: usize,
|
|
rs: &ReedSolomon,
|
|
buffers: &mut [Vec<u8>],
|
|
shards: &mut [EcVolumeShard],
|
|
builders: &mut [ShardChecksumBuilder],
|
|
data_shards: usize,
|
|
) -> io::Result<()> {
|
|
// Read data shards from the .dat file, zero-filling past EOF — the buffers
|
|
// are reused across batches, so the tail must be cleared explicitly.
|
|
for i in 0..data_shards {
|
|
let read_offset = offset + (i * block_size) as u64;
|
|
let n = read_at_most(dat_file, &mut buffers[i], read_offset)?;
|
|
for b in buffers[i][n..].iter_mut() {
|
|
*b = 0;
|
|
}
|
|
}
|
|
|
|
// Encode parity shards
|
|
rs.encode(&mut *buffers).map_err(|e| {
|
|
io::Error::new(
|
|
io::ErrorKind::Other,
|
|
format!("reed-solomon encode: {:?}", e),
|
|
)
|
|
})?;
|
|
|
|
// Write all shard buffers to files and feed the same bytes to each
|
|
// shard's bitrot checksum builder, keeping covered_size == on-disk length.
|
|
for (i, buf) in buffers.iter().enumerate() {
|
|
shards[i].write_all(buf)?;
|
|
builders[i].write(buf);
|
|
}
|
|
|
|
Ok(())
|
|
}
|
|
|
|
/// Read into `buf` at `offset` until it is full or EOF; returns bytes read.
|
|
fn read_at_most(dat_file: &File, buf: &mut [u8], offset: u64) -> io::Result<usize> {
|
|
let mut n = 0;
|
|
while n < buf.len() {
|
|
#[cfg(unix)]
|
|
let r = {
|
|
use std::os::unix::fs::FileExt;
|
|
dat_file.read_at(&mut buf[n..], offset + n as u64)?
|
|
};
|
|
#[cfg(not(unix))]
|
|
let r = {
|
|
let mut f = dat_file.try_clone()?;
|
|
f.seek(SeekFrom::Start(offset + n as u64))?;
|
|
f.read(&mut buf[n..])?
|
|
};
|
|
if r == 0 {
|
|
break;
|
|
}
|
|
n += r;
|
|
}
|
|
Ok(n)
|
|
}
|
|
|
|
#[cfg(test)]
|
|
mod tests {
|
|
use super::*;
|
|
use crate::storage::needle::needle::Needle;
|
|
use crate::storage::needle_map::NeedleMapKind;
|
|
use crate::storage::volume::Volume;
|
|
use tempfile::TempDir;
|
|
|
|
#[test]
|
|
fn test_ec_encode_decode_round_trip() {
|
|
let tmp = TempDir::new().unwrap();
|
|
let dir = tmp.path().to_str().unwrap();
|
|
|
|
// Create a volume with some data
|
|
let mut v = Volume::new(
|
|
dir,
|
|
dir,
|
|
"",
|
|
VolumeId(1),
|
|
NeedleMapKind::InMemory,
|
|
None,
|
|
None,
|
|
0,
|
|
Version::current(),
|
|
)
|
|
.unwrap();
|
|
|
|
for i in 1..=5 {
|
|
let data = format!("test data for needle {}", i);
|
|
let mut n = Needle {
|
|
id: NeedleId(i),
|
|
cookie: Cookie(i as u32),
|
|
data: data.as_bytes().to_vec(),
|
|
data_size: data.len() as u32,
|
|
..Needle::default()
|
|
};
|
|
v.write_needle(&mut n, true, false).unwrap();
|
|
}
|
|
v.sync_to_disk().unwrap();
|
|
v.close();
|
|
|
|
// Encode to EC shards
|
|
let data_shards = 10;
|
|
let parity_shards = 4;
|
|
let total_shards = data_shards + parity_shards;
|
|
write_ec_files(dir, dir, "", VolumeId(1), data_shards, parity_shards).unwrap();
|
|
|
|
// Verify shard files exist
|
|
for i in 0..total_shards {
|
|
let path = format!("{}/{}.ec{:02}", dir, 1, i);
|
|
assert!(
|
|
std::path::Path::new(&path).exists(),
|
|
"shard file {} should exist",
|
|
path
|
|
);
|
|
}
|
|
|
|
// Verify .ecx exists
|
|
let ecx_path = format!("{}/1.ecx", dir);
|
|
assert!(std::path::Path::new(&ecx_path).exists());
|
|
}
|
|
|
|
fn make_volume_with_needles(n: u64) -> TempDir {
|
|
let tmp = TempDir::new().unwrap();
|
|
let dir = tmp.path().to_str().unwrap();
|
|
let mut v = Volume::new(
|
|
dir,
|
|
dir,
|
|
"",
|
|
VolumeId(1),
|
|
NeedleMapKind::InMemory,
|
|
None,
|
|
None,
|
|
0,
|
|
Version::current(),
|
|
)
|
|
.unwrap();
|
|
for i in 1..=n {
|
|
// Larger payloads so encoded shards span multiple bitrot blocks
|
|
// would require huge data; small payloads are fine for correctness.
|
|
let data = format!("test data for needle {} {}", i, "x".repeat(64));
|
|
let mut needle = Needle {
|
|
id: NeedleId(i),
|
|
cookie: Cookie(i as u32),
|
|
data: data.as_bytes().to_vec(),
|
|
data_size: data.len() as u32,
|
|
..Needle::default()
|
|
};
|
|
v.write_needle(&mut needle, true, false).unwrap();
|
|
}
|
|
v.sync_to_disk().unwrap();
|
|
v.close();
|
|
tmp
|
|
}
|
|
|
|
/// Encode-time capture writes a valid generation-0 `.ecsum` sidecar whose
|
|
/// recorded checksums match the actual on-disk shards.
|
|
#[test]
|
|
fn test_encode_writes_valid_bitrot_sidecar() {
|
|
use crate::storage::erasure_coding::ec_bitrot;
|
|
|
|
let tmp = make_volume_with_needles(5);
|
|
let dir = tmp.path().to_str().unwrap();
|
|
write_ec_files(dir, dir, "", VolumeId(1), 10, 4).unwrap();
|
|
|
|
let base = format!("{}/1", dir);
|
|
let sidecar_path = ec_bitrot::bitrot_sidecar_path(&base, 0);
|
|
assert!(
|
|
std::path::Path::new(&sidecar_path).exists(),
|
|
"generation-0 .ecsum sidecar should exist after encode"
|
|
);
|
|
|
|
let prot = ec_bitrot::load_bitrot_sidecar(&sidecar_path).unwrap();
|
|
ec_bitrot::validate_manifest(&prot, 10, 4).unwrap();
|
|
assert_eq!(prot.generation, 0);
|
|
assert_eq!(prot.shards.len(), 14);
|
|
assert_eq!(prot.encode_uuid.len(), 16);
|
|
assert_eq!(prot.block_size, ec_bitrot::DEFAULT_BITROT_BLOCK_SIZE as u32);
|
|
|
|
// Every shard's recorded covered_size must equal its on-disk length and
|
|
// its block CRCs must verify clean.
|
|
let bs = prot.block_size as i64;
|
|
for entry in &prot.shards {
|
|
let path = format!("{}.ec{:02}", base, entry.shard_id);
|
|
let on_disk = std::fs::metadata(&path).unwrap().len() as i64;
|
|
assert_eq!(
|
|
entry.covered_size, on_disk,
|
|
"covered_size must equal on-disk length for shard {}",
|
|
entry.shard_id
|
|
);
|
|
let mm = ec_bitrot::verify_shard_file_blocks(&path, entry, bs).unwrap();
|
|
assert!(
|
|
mm.is_empty(),
|
|
"shard {} should verify clean, got mismatches {:?}",
|
|
entry.shard_id,
|
|
mm
|
|
);
|
|
}
|
|
}
|
|
|
|
// encode_sample_volume writes a small volume and EC-encodes it, returning
|
|
// the dir path so a test can drop/truncate shards and rebuild.
|
|
fn encode_sample_volume(tmp: &TempDir) -> String {
|
|
let dir = tmp.path().to_str().unwrap().to_string();
|
|
let mut v = Volume::new(
|
|
&dir,
|
|
&dir,
|
|
"",
|
|
VolumeId(1),
|
|
NeedleMapKind::InMemory,
|
|
None,
|
|
None,
|
|
0,
|
|
Version::current(),
|
|
)
|
|
.unwrap();
|
|
for i in 1..=20 {
|
|
let data = format!("test data for needle {} padded with bytes", i).repeat(64);
|
|
let mut n = Needle {
|
|
id: NeedleId(i),
|
|
cookie: Cookie(i as u32),
|
|
data: data.as_bytes().to_vec(),
|
|
data_size: data.len() as u32,
|
|
..Needle::default()
|
|
};
|
|
v.write_needle(&mut n, true, false).unwrap();
|
|
}
|
|
v.sync_to_disk().unwrap();
|
|
v.close();
|
|
write_ec_files(&dir, &dir, "", VolumeId(1), 10, 4).unwrap();
|
|
dir
|
|
}
|
|
|
|
// A truncated/corrupt input shard must abort rebuild_ec_files with an error
|
|
// rather than reconstructing over a zero-padded tail and publishing a
|
|
// truncated shard as restored redundancy.
|
|
#[test]
|
|
fn test_rebuild_ec_files_short_read_input_errors() {
|
|
let tmp = TempDir::new().unwrap();
|
|
let dir = encode_sample_volume(&tmp);
|
|
|
|
// Truncate a present input shard to half its size.
|
|
let victim = format!("{}/1.ec03", dir);
|
|
let full = std::fs::metadata(&victim).unwrap().len();
|
|
assert!(full > 0, "encoded shard should be non-empty");
|
|
let f = std::fs::OpenOptions::new().write(true).open(&victim).unwrap();
|
|
f.set_len(full / 2).unwrap();
|
|
drop(f);
|
|
|
|
// Rebuild a different (genuinely missing) shard.
|
|
std::fs::remove_file(format!("{}/1.ec07", dir)).unwrap();
|
|
let res = rebuild_ec_files(&dir, "", VolumeId(1), &[7], 10, 4, &[]);
|
|
assert!(res.is_err(), "truncated input shard must abort the rebuild");
|
|
}
|
|
|
|
// Happy path: dropping a shard and rebuilding it from the rest must succeed
|
|
// and recreate the shard file (the short-read guard must not false-positive).
|
|
#[test]
|
|
fn test_rebuild_ec_files_happy_path() {
|
|
let tmp = TempDir::new().unwrap();
|
|
let dir = encode_sample_volume(&tmp);
|
|
|
|
let dropped = format!("{}/1.ec07", dir);
|
|
std::fs::remove_file(&dropped).unwrap();
|
|
rebuild_ec_files(&dir, "", VolumeId(1), &[7], 10, 4, &[]).unwrap();
|
|
assert!(
|
|
std::path::Path::new(&dropped).exists(),
|
|
"rebuilt shard .ec07 should exist"
|
|
);
|
|
assert!(
|
|
std::fs::metadata(&dropped).unwrap().len() > 0,
|
|
"rebuilt shard should be non-empty"
|
|
);
|
|
}
|
|
|
|
// Multi-disk rebuild: shards for the same volume are split across two
|
|
// directories (simulating a volume server with two disk locations). The
|
|
// shard to rebuild is in the primary dir; some of the input shards needed
|
|
// for RS reconstruction exist only in the secondary dir. Passing the
|
|
// secondary dir via `additional_dirs` must allow the rebuild to succeed
|
|
// where it would previously return "missing non-rebuild shard N".
|
|
#[test]
|
|
fn test_rebuild_ec_files_multi_disk() {
|
|
let tmp_primary = TempDir::new().unwrap();
|
|
let tmp_secondary = TempDir::new().unwrap();
|
|
let primary = tmp_primary.path().to_str().unwrap().to_string();
|
|
let secondary = tmp_secondary.path().to_str().unwrap().to_string();
|
|
|
|
// Encode into primary dir first (all 14 shards land there).
|
|
let dir = encode_sample_volume(&tmp_primary);
|
|
assert_eq!(dir, primary);
|
|
|
|
// Simulate a multi-disk layout: move shards 0, 4, 8 to the secondary
|
|
// dir, as if the master had placed them on a different disk.
|
|
let moved_shards: &[u8] = &[0, 4, 8];
|
|
for &shard_id in moved_shards {
|
|
let src = format!("{}/1.ec{:02}", primary, shard_id);
|
|
let dst = format!("{}/1.ec{:02}", secondary, shard_id);
|
|
std::fs::rename(&src, &dst).unwrap();
|
|
}
|
|
|
|
// Remove shard 7 from primary — this is the shard we want to rebuild.
|
|
let missing_path = format!("{}/1.ec07", primary);
|
|
std::fs::remove_file(&missing_path).unwrap();
|
|
|
|
// Without additional_dirs the rebuild must fail: shards 0, 4, 8 are
|
|
// not in primary and shard 7 cannot be reconstructed without them.
|
|
let res = rebuild_ec_files(&primary, "", VolumeId(1), &[7], 10, 4, &[]);
|
|
assert!(
|
|
res.is_err(),
|
|
"rebuild without additional_dirs must fail when input shards are on another disk"
|
|
);
|
|
assert!(
|
|
!std::path::Path::new(&missing_path).exists(),
|
|
"failed rebuild must not leave a partial shard file behind"
|
|
);
|
|
|
|
// With additional_dirs pointing at the secondary, the rebuild must succeed.
|
|
rebuild_ec_files(
|
|
&primary,
|
|
"",
|
|
VolumeId(1),
|
|
&[7],
|
|
10,
|
|
4,
|
|
&[secondary.as_str()],
|
|
)
|
|
.unwrap();
|
|
|
|
assert!(
|
|
std::path::Path::new(&missing_path).exists(),
|
|
"rebuilt shard .ec07 should exist in primary dir"
|
|
);
|
|
assert!(
|
|
std::fs::metadata(&missing_path).unwrap().len() > 0,
|
|
"rebuilt shard must be non-empty"
|
|
);
|
|
}
|
|
|
|
// Multi-disk .ecx rebuild: data shards (0-9) needed to reconstruct the
|
|
// logical .dat content are split across two directories, as on a
|
|
// multi-disk volume server. Passing the secondary dir via
|
|
// `additional_dirs` must allow rebuild_ecx_file to find them and
|
|
// succeed where it would previously fail with
|
|
// "cannot open data shard for ecx rebuild".
|
|
#[test]
|
|
fn test_rebuild_ecx_file_multi_disk() {
|
|
let tmp_primary = TempDir::new().unwrap();
|
|
let tmp_secondary = TempDir::new().unwrap();
|
|
let primary = tmp_primary.path().to_str().unwrap().to_string();
|
|
let secondary = tmp_secondary.path().to_str().unwrap().to_string();
|
|
|
|
let dir = encode_sample_volume(&tmp_primary);
|
|
assert_eq!(dir, primary);
|
|
|
|
// Move some data shards (0-9) to the secondary dir, simulating a
|
|
// multi-disk layout where not all data shards landed on the same disk.
|
|
let moved_shards: &[u8] = &[1, 3, 6];
|
|
for &shard_id in moved_shards {
|
|
let src = format!("{}/1.ec{:02}", primary, shard_id);
|
|
let dst = format!("{}/1.ec{:02}", secondary, shard_id);
|
|
std::fs::rename(&src, &dst).unwrap();
|
|
}
|
|
|
|
// Delete the existing .ecx to force a rebuild.
|
|
let ecx_path = format!("{}/1.ecx", primary);
|
|
std::fs::remove_file(&ecx_path).unwrap();
|
|
|
|
// Without additional_dirs, rebuild must fail: shards 1, 3, 6 are not
|
|
// in primary and the full logical .dat content can't be reconstructed.
|
|
let res = rebuild_ecx_file(&primary, "", VolumeId(1), 10, 0, 0, &[]);
|
|
assert!(
|
|
res.is_err(),
|
|
"ecx rebuild without additional_dirs must fail when data shards are on another disk"
|
|
);
|
|
assert!(
|
|
!std::path::Path::new(&ecx_path).exists(),
|
|
"failed ecx rebuild must not leave a partial .ecx file behind"
|
|
);
|
|
|
|
// With additional_dirs pointing at the secondary, the rebuild must succeed.
|
|
rebuild_ecx_file(&primary, "", VolumeId(1), 10, 0, 0, &[secondary.as_str()]).unwrap();
|
|
|
|
assert!(
|
|
std::path::Path::new(&ecx_path).exists(),
|
|
"rebuilt .ecx should exist in primary dir"
|
|
);
|
|
assert!(
|
|
std::fs::metadata(&ecx_path).unwrap().len() > 0,
|
|
"rebuilt .ecx must be non-empty"
|
|
);
|
|
}
|
|
|
|
// A uniform-layout volume (block size > 1MiB) must have its .ecx rebuilt
|
|
// through the recorded geometry; the legacy 1MiB mapping would scan
|
|
// garbage past the first block boundary.
|
|
#[test]
|
|
fn test_rebuild_ecx_file_uniform_layout() {
|
|
use crate::storage::needle_map::NeedleMapKind;
|
|
use crate::storage::volume::Volume;
|
|
let tmp = TempDir::new().unwrap();
|
|
let dir = tmp.path().to_str().unwrap().to_string();
|
|
let mut v = Volume::new(
|
|
&dir,
|
|
&dir,
|
|
"",
|
|
VolumeId(2),
|
|
NeedleMapKind::InMemory,
|
|
None,
|
|
None,
|
|
0,
|
|
Version::current(),
|
|
)
|
|
.unwrap();
|
|
for i in 1u64..=12 {
|
|
let data: Vec<u8> = (0..2 << 20)
|
|
.map(|b| ((b as u64).wrapping_mul(2654435761).wrapping_add(i) >> 8) as u8)
|
|
.collect();
|
|
let mut n = Needle {
|
|
id: NeedleId(i),
|
|
cookie: Cookie(i as u32),
|
|
data: data.clone(),
|
|
data_size: data.len() as u32,
|
|
..Needle::default()
|
|
};
|
|
v.write_needle(&mut n, true, false).unwrap();
|
|
}
|
|
v.sync_to_disk().unwrap();
|
|
v.close();
|
|
let block_size = write_ec_files(&dir, &dir, "", VolumeId(2), 10, 4).unwrap();
|
|
assert!(
|
|
block_size > ERASURE_CODING_SMALL_BLOCK_SIZE as i64,
|
|
"fixture must diverge from the legacy layout"
|
|
);
|
|
|
|
let ecx_path = format!("{}/2.ecx", dir);
|
|
let canonical = std::fs::read(&ecx_path).unwrap();
|
|
std::fs::remove_file(&ecx_path).unwrap();
|
|
|
|
rebuild_ecx_file(&dir, "", VolumeId(2), 10, block_size, 0, &[]).unwrap();
|
|
let rebuilt = std::fs::read(&ecx_path).unwrap();
|
|
assert_eq!(canonical, rebuilt, "rebuilt .ecx must match the encode-time .ecx");
|
|
}
|
|
|
|
// A truncated data shard must FAIL the .ecx rebuild, not publish the
|
|
// entries scanned so far as a successful (silently incomplete) index.
|
|
#[test]
|
|
fn test_rebuild_ecx_file_fails_on_truncated_shard() {
|
|
use crate::storage::needle_map::NeedleMapKind;
|
|
use crate::storage::volume::Volume;
|
|
let tmp = TempDir::new().unwrap();
|
|
let dir = tmp.path().to_str().unwrap().to_string();
|
|
let mut v = Volume::new(
|
|
&dir,
|
|
&dir,
|
|
"",
|
|
VolumeId(3),
|
|
NeedleMapKind::InMemory,
|
|
None,
|
|
None,
|
|
0,
|
|
Version::current(),
|
|
)
|
|
.unwrap();
|
|
for i in 1u64..=12 {
|
|
let data: Vec<u8> = (0..2 << 20)
|
|
.map(|b| ((b as u64).wrapping_mul(2654435761).wrapping_add(i) >> 8) as u8)
|
|
.collect();
|
|
let mut n = Needle {
|
|
id: NeedleId(i),
|
|
cookie: Cookie(i as u32),
|
|
data: data.clone(),
|
|
data_size: data.len() as u32,
|
|
..Needle::default()
|
|
};
|
|
v.write_needle(&mut n, true, false).unwrap();
|
|
}
|
|
v.sync_to_disk().unwrap();
|
|
v.close();
|
|
let block_size = write_ec_files(&dir, &dir, "", VolumeId(3), 10, 4).unwrap();
|
|
|
|
let ecx_path = format!("{}/3.ecx", dir);
|
|
std::fs::remove_file(&ecx_path).unwrap();
|
|
|
|
// Truncate shard 0 to just the superblock: the scan's very first
|
|
// needle-header read (offset SUPER_BLOCK_SIZE, shard 0 under the
|
|
// uniform layout) lands in the missing region. The pre-fix code
|
|
// broke the scan there and published an EMPTY .ecx as success.
|
|
let shard_path = format!("{}/3.ec00", dir);
|
|
let f = std::fs::OpenOptions::new()
|
|
.write(true)
|
|
.open(&shard_path)
|
|
.unwrap();
|
|
f.set_len(crate::storage::super_block::SUPER_BLOCK_SIZE as u64)
|
|
.unwrap();
|
|
drop(f);
|
|
|
|
let res = rebuild_ecx_file(&dir, "", VolumeId(3), 10, block_size, 0, &[]);
|
|
assert!(res.is_err(), "rebuild over a truncated shard must fail");
|
|
assert!(
|
|
!std::path::Path::new(&ecx_path).exists(),
|
|
"a failed rebuild must not leave a partial .ecx behind"
|
|
);
|
|
}
|
|
|
|
#[test]
|
|
fn test_reed_solomon_basic() {
|
|
let data_shards = 10;
|
|
let parity_shards = 4;
|
|
let total_shards = data_shards + parity_shards;
|
|
let rs = ReedSolomon::new(data_shards, parity_shards).unwrap();
|
|
let block_size = 1024;
|
|
let mut shards: Vec<Vec<u8>> = (0..total_shards)
|
|
.map(|i| {
|
|
if i < data_shards {
|
|
vec![(i as u8).wrapping_mul(7); block_size]
|
|
} else {
|
|
vec![0u8; block_size]
|
|
}
|
|
})
|
|
.collect();
|
|
|
|
// Encode
|
|
rs.encode(&mut shards).unwrap();
|
|
|
|
// Verify parity is non-zero (at least some)
|
|
let parity_nonzero: bool = shards[data_shards..]
|
|
.iter()
|
|
.any(|s| s.iter().any(|&b| b != 0));
|
|
assert!(parity_nonzero);
|
|
|
|
// Simulate losing 4 shards and reconstructing
|
|
let original_0 = shards[0].clone();
|
|
let original_1 = shards[1].clone();
|
|
|
|
let mut shard_opts: Vec<Option<Vec<u8>>> = shards.into_iter().map(Some).collect();
|
|
shard_opts[0] = None;
|
|
shard_opts[1] = None;
|
|
shard_opts[2] = None;
|
|
shard_opts[3] = None;
|
|
|
|
rs.reconstruct(&mut shard_opts).unwrap();
|
|
|
|
assert_eq!(shard_opts[0].as_ref().unwrap(), &original_0);
|
|
assert_eq!(shard_opts[1].as_ref().unwrap(), &original_1);
|
|
}
|
|
|
|
/// EC encode must read .idx from a separate index directory when configured.
|
|
#[test]
|
|
fn test_ec_encode_with_separate_idx_dir() {
|
|
let dat_tmp = TempDir::new().unwrap();
|
|
let idx_tmp = TempDir::new().unwrap();
|
|
let dat_dir = dat_tmp.path().to_str().unwrap();
|
|
let idx_dir = idx_tmp.path().to_str().unwrap();
|
|
|
|
// Create a volume with separate data and index directories
|
|
let mut v = Volume::new(
|
|
dat_dir,
|
|
idx_dir,
|
|
"",
|
|
VolumeId(1),
|
|
NeedleMapKind::InMemory,
|
|
None,
|
|
None,
|
|
0,
|
|
Version::current(),
|
|
)
|
|
.unwrap();
|
|
|
|
for i in 1..=5 {
|
|
let data = format!("needle {} payload", i);
|
|
let mut n = Needle {
|
|
id: NeedleId(i),
|
|
cookie: Cookie(i as u32),
|
|
data: data.as_bytes().to_vec(),
|
|
data_size: data.len() as u32,
|
|
..Needle::default()
|
|
};
|
|
v.write_needle(&mut n, true, false).unwrap();
|
|
}
|
|
v.sync_to_disk().unwrap();
|
|
v.close();
|
|
|
|
// Verify .dat is in data dir, .idx is in idx dir
|
|
assert!(std::path::Path::new(&format!("{}/1.dat", dat_dir)).exists());
|
|
assert!(!std::path::Path::new(&format!("{}/1.idx", dat_dir)).exists());
|
|
assert!(std::path::Path::new(&format!("{}/1.idx", idx_dir)).exists());
|
|
assert!(!std::path::Path::new(&format!("{}/1.dat", idx_dir)).exists());
|
|
|
|
// EC encode with separate idx dir
|
|
let data_shards = 10;
|
|
let parity_shards = 4;
|
|
let total_shards = data_shards + parity_shards;
|
|
write_ec_files(
|
|
dat_dir,
|
|
idx_dir,
|
|
"",
|
|
VolumeId(1),
|
|
data_shards,
|
|
parity_shards,
|
|
)
|
|
.unwrap();
|
|
|
|
// Verify all 14 shard files in data dir
|
|
for i in 0..total_shards {
|
|
let path = format!("{}/1.ec{:02}", dat_dir, i);
|
|
assert!(
|
|
std::path::Path::new(&path).exists(),
|
|
"shard {} should exist in data dir",
|
|
path
|
|
);
|
|
}
|
|
|
|
// Verify .ecx in data dir (not idx dir)
|
|
assert!(std::path::Path::new(&format!("{}/1.ecx", dat_dir)).exists());
|
|
assert!(!std::path::Path::new(&format!("{}/1.ecx", idx_dir)).exists());
|
|
|
|
// Verify no shard files leaked into idx dir
|
|
for i in 0..total_shards {
|
|
let path = format!("{}/1.ec{:02}", idx_dir, i);
|
|
assert!(
|
|
!std::path::Path::new(&path).exists(),
|
|
"shard {} should NOT exist in idx dir",
|
|
path
|
|
);
|
|
}
|
|
}
|
|
|
|
/// EC encode should fail gracefully when .idx is only in the data dir
|
|
/// but we pass a wrong idx_dir. This guards against regressions where
|
|
/// write_ec_files ignores the idx_dir parameter.
|
|
#[test]
|
|
fn test_ec_encode_fails_with_wrong_idx_dir() {
|
|
let dat_tmp = TempDir::new().unwrap();
|
|
let idx_tmp = TempDir::new().unwrap();
|
|
let wrong_tmp = TempDir::new().unwrap();
|
|
let dat_dir = dat_tmp.path().to_str().unwrap();
|
|
let idx_dir = idx_tmp.path().to_str().unwrap();
|
|
let wrong_dir = wrong_tmp.path().to_str().unwrap();
|
|
|
|
let mut v = Volume::new(
|
|
dat_dir,
|
|
idx_dir,
|
|
"",
|
|
VolumeId(1),
|
|
NeedleMapKind::InMemory,
|
|
None,
|
|
None,
|
|
0,
|
|
Version::current(),
|
|
)
|
|
.unwrap();
|
|
|
|
let mut n = Needle {
|
|
id: NeedleId(1),
|
|
cookie: Cookie(1),
|
|
data: b"hello".to_vec(),
|
|
data_size: 5,
|
|
..Needle::default()
|
|
};
|
|
v.write_needle(&mut n, true, false).unwrap();
|
|
v.sync_to_disk().unwrap();
|
|
v.close();
|
|
|
|
// Should fail: .idx is in idx_dir, not wrong_dir
|
|
let result = write_ec_files(dat_dir, wrong_dir, "", VolumeId(1), 10, 4);
|
|
assert!(
|
|
result.is_err(),
|
|
"should fail when idx_dir doesn't contain .idx"
|
|
);
|
|
}
|
|
}
|