Compare commits

...
80 Commits
Author SHA1 Message Date
Chris Lu 6b5942946f fix(s3api): cancel ListEntries stream in hasChildren
hasChildren opened a server-streaming ListEntries with an uncancelled
context and returned after one Recv, leaving gRPC's per-stream client
goroutine parked. Route through filer_pb.List, which cancels the stream
on return, so the goroutine is cleaned up.
2026-05-21 14:56:36 -07:00
Chris LuandGitHub 8fa769f29a feat(helm): add volume.rust toggle to run the Rust volume server (#9618)
feat(helm): add volume.rust to run the Rust volume server

When set, the volume statefulset execs /usr/bin/weed-volume instead of
'weed volume', dropping the Go-only -logtostderr/-logdir/-v flags and the
'volume' subcommand. All shared flags and extraArgs carry over unchanged.
2026-05-21 14:30:27 -07:00
Chris LuandGitHub 7c635c4508 fix(docker): restore executable bit on prebuilt weed-volume (#9616)
GitHub Actions artifacts drop the executable bit, so the pre-built Rust
volume server lands in the image as 0644 and 'weed-volume' fails to start
with 'exec: Permission denied'. chmod it 0755 after copying.
2026-05-21 14:11:08 -07:00
Chris LuandGitHub fbdcec1cba fix(s3): list empty directories as directory markers (#9615)
* fix(s3): list empty directories as directory markers

A real but empty directory created out of band (mount, mkdir, filer API)
carries no MIME, so it was hidden from S3 listings. hadoop-aws getFileStatus
probes LIST prefix=dir/ &delimiter=/ and reads an empty result as a missing
path, which breaks Spark's eventLog.dir when it points at an empty directory.

Surface such directories as directory markers, matching directories created
via PutObject with a trailing "/". Emptiness comes from the recursion result,
and the marker MIME is set only on the in-memory listing entry, so empty
directories stay eligible for empty-folder cleanup.

* fix(s3): only surface empty directory markers for explicit dir probes

Restrict the empty-directory marker to a trailing-slash prefix probe
(prefix=dir/), the pattern hadoop-aws getFileStatus uses. Plain listings
are left as before, so an empty directory left behind by deleted objects
(e.g. after lifecycle expiration) is no longer shown as a phantom key.
2026-05-21 14:05:16 -07:00
Chris Lu 0accff0e4a fix(ec): log EC destination planning failures at v=2
The maintenance scanner tries to plan EC destinations for every
eligible volume, so clusters that can't place EC logged a warning per
volume every cycle. The min-node gate already skips clusters with fewer
nodes than parity shards; demote the rest to V(2).
2026-05-21 10:35:34 -07:00
Chris LuandGitHub 9021225591 master: accept volume-server Ping targets on follower masters (#9614)
cluster.check asks every master to ping every volume server, but the
Ping gate validated volume-server targets only against the local
topology. Only the leader receives volume-server heartbeats, so a
follower's topology is empty and every probe through it failed with
"unknown ping target ... of type volumeServer".

Fall back to the volume-server set the master learns over its own
MasterClient subscription to the leader, the same source the filer gate
already trusts. The anti-SSRF intent is preserved: Ping still only dials
recognized cluster members.
2026-05-21 10:19:59 -07:00
Chris LuandGitHub 5b42287c22 fix(storage): surface stat error on zero-size idx scrub, mirror to rust (#9612)
fix(storage): harden zero-size idx scrub and mirror to rust

When a zero-size .idx is found, openIndex stats the backing .dat through
v.DataBackend: wrap that GetStat failure with %w, fix the indices typo, and
guard both openIndex and scrubVolumeData against a nil DataBackend (closed or
remote-only volumes) instead of panicking.

Add rust scrub tests for empty (superblock-only .dat, zero-size .idx) and
healthy volumes, keeping the volume server in parity with the go zero-size
scrub handling.
2026-05-21 10:17:23 -07:00
Chris LuandGitHub 3392493f0a test(volume): fix race in TestReplicatedUploadSucceedsImmediatelyAfterAllocate (#9613)
test(volume): wait for master to register both replicas before replicated upload

TestReplicatedUploadSucceedsImmediatelyAfterAllocate allocated the volume on
both nodes via direct AllocateVolume gRPC calls, then uploaded immediately. The
master only learns about replica locations through volume-server heartbeats,
which lag behind those direct gRPC calls, so the replicated write could look up
the master before the second replica was registered and fail with a 500
("replicating operations [1] is less than volume replication copy count [2]").

In production a client obtains its fid from the master assign flow, which
guarantees the master already knows every replica. The test crafts the fid by
hand, bypassing that guarantee, so wait until the master reports both replicas
before uploading.
2026-05-21 09:58:37 -07:00
Chris Lu d82b3a8d6a refactor(s3): drop unused source path in copy ETag check
ETagEntry derives the tag from chunks/Md5/remote-etag, never the entry path,
so the conditional-copy check no longer builds a bogus FullPath.
2026-05-21 09:51:50 -07:00
Lisandro PinandGitHub 39e9294907 Have volume scrubs account for zero-sized volumes. (#9609)
Fixes scrubbing for pre-allocated volumes with zero-size indeces by
reworking the validation code to allow zero-size indeces on zero-size
volumes.
2026-05-21 09:42:07 -07:00
Chris LuandGitHub 3825035f07 test(ec): deterministically populate disks before multi-disk EC balance check (#9611)
The disk-spread assertion raced volume growth and heartbeats. volume.grow
-count is a writable-target topup, not add-N, and swallows partial-failure
errors, so one grow could leave a node's data on a single disk; ec.encode
then piles all that node's shards there and ec.balance can't spread them.

Retry grow on under-spread nodes until the master topology shows every node
holding volumes on at least two physical disks, then encode.
2026-05-21 09:39:55 -07:00
Chris LuandGitHub 83b7ea5e7b fix(s3): keep server-side copy data in the bucket collection (#9607)
* fix(s3): keep server-side copy data in the bucket collection

UploadPartCopy and SSE-C CopyObject assigned destination volumes against
r.URL.Path, the S3 request URI. The filer derives a bucket's collection
only when the assign path sits under its buckets folder, so an S3 URI
routed copied bytes to the default collection instead of the destination
bucket's. Assign against the destination's real filer path.

* refactor(s3): centralize copy-part path and thread dstPath into SSE-C copy

Extract copyPartLocation so the fast path and writeEmptyCopyPart share one
definition of the .uploads/<id>/<n>_copy.part location. Pass the destination
filer path into copyChunksWithSSEC instead of re-deriving it from the request,
and thread it through key rotation so re-encrypt copies also assign in the
destination bucket's collection.
2026-05-21 09:35:42 -07:00
Chris LuandGitHub eae8f33db5 fix(filersink): return lock-free snapshot from ActiveTransfers (#9604)
ChunkTransferStatus embeds a sync.RWMutex, so returning a slice of it
made callers copy the lock when ranging. Split out a copyable
ChunkTransferSnapshot holding the data fields and return that instead.
2026-05-21 02:40:04 -07:00
Chris LuandGitHub 2c2b2d4d3e chore(skiplist): remove unused NameList/NameBatch implementation (#9603)
NameList, NameBatch and their serde were an earlier in-memory directory
batch implementation. The redis3 filer store uses its own ItemList
backed by Redis sorted sets, so these types had no production callers
(NameList only via its own test, LoadNameList none at all). Drop them
and the now-orphaned NameBatchData proto message, regenerating
skiplist.pb.go with the repo-standard protoc-gen-go v1.36.6.
2026-05-21 02:30:49 -07:00
Chris LuandGitHub cd15ae1395 fix(ec): bring ec.encode worker and EC/volume helpers to parity with shell (#9599)
* refactor(volume): extract replica sync/select into shared volume_replica package

Move the volume replica reconciliation helpers (status, union builder,
SyncAndSelectBestReplica, ReadNeedleMeta) out of the shell into a new
weed/storage/volume_replica package so both the shell (ec.encode, volume.tier.move,
volume.check.disk) and the EC encode worker can reuse them. No behavior change.

* fix(ec): bring ec.encode worker to parity with the shell

- Sync replicas and encode the most-complete one (via the shared
  volume_replica.SyncAndSelectBestReplica) instead of a possibly-stale replica,
  marking all replicas readonly first. Prevents silent data loss when a stale
  replica is encoded and the originals deleted.
- Skip remote/tiered volumes in detection (shell ec.encode excludes them).
- Min-node safety gate: refuse to encode when cluster nodes < parity shards.
- Align default thresholds with the shell (fullness 0.95, quiet 1h).

* fix(vacuum): plugin path honors min_volume_age_seconds override

deriveVacuumConfig hard-coded MinVolumeAgeSeconds=0, dropping any configured
value. Read it from worker config (default 0, matching the shell/master vacuum
which has no age gate) so an explicit override is honored.

* address review feedback

- config.go: align GetConfigSpec schema defaults (quiet_for_seconds=3600,
  fullness_ratio=0.95) with the runtime defaults so UI/bootstrap flows match the
  shell (coderabbitai).
- ec_task.go: roll back readonly when markReplicasReadonly fails partway, so
  already-marked replicas don't stay readonly (coderabbitai).
- volume_replica: pass the caller's replica statuses into buildUnionReplica instead
  of re-fetching them, and skip the per-needle ReadNeedleMeta RPC when the source
  replica is read-only (gemini-code-assist).

* test(plugin_workers/ec): make fixtures eligible under the new defaults

The default EC encode thresholds were raised to match the shell (fullness 0.95,
quiet 1h), but the plugin-worker integration fixtures still used 90%-full /
10-minute-old volumes, so detection found no eligible volumes and the tests failed
in CI. Bump the eligible fixtures to 96% full and 2h old.
2026-05-21 02:16:28 -07:00
Chris LuandGitHub 3f6410fdc3 fix(redis3): prevent filer crash from inconsistent skiplist ends (#9602)
* fix(redis3): prevent filer crash from inconsistent skiplist ends

DeleteByKey updated the two ends asymmetrically: the start side decided
whether to clear StartLevels[index] by comparing a cached reference key,
while the end side cleared EndLevels[index] structurally. The redis3
ItemList re-keys a node while keeping its id, so that cached key drifts.
When such a node was the only one at a level and got deleted, the stale
key comparison left StartLevels dangling while EndLevels was cleared to
nil. The next InsertByKey then dereferenced a nil EndLevels[0] and took
down the whole filer during a rename or delete.

Match the deleted node by its unique id so both ends stay consistent,
and guard each end in InsertByKey so an already-corrupted skiplist
persisted in Redis self-heals instead of crashing on load.

* fix(redis3): propagate errors from WriteName node split

The case 2.3 split path returned nil instead of the error in seven
branches. Because the split runs a multi-step sequence (DeleteByKey on
the skiplist, ItemAdd, redis range-delete, ItemAdd), a swallowed failure
let WriteName report success while the skiplist was half-updated, which
the caller then persisted - silently corrupting the directory listing
and setting up the very inconsistent-ends state that crashes the filer.
2026-05-21 02:12:00 -07:00
Chris LuandGitHub 87fdea5330 fix(admin): carry filer addresses as ServerAddress in plugin cluster context (#9600)
The plugin cluster context forwarded filers as gRPC-only addresses
(host:grpcPort). The admin-script worker stored that in
ShellOptions.FilerAddress, whose shell commands re-derive the gRPC port
via ToGrpcAddress() and re-add the +10000 offset, dialing a non-existent
host:28888.

Carry filers in pb.ServerAddress form (host:httpPort.grpcPort) and let
each consumer convert when it dials: the admin shell uses it verbatim,
while the s3_lifecycle and iceberg workers collapse it to a gRPC address.
Rename the proto field filer_grpc_addresses -> filer_addresses so the
name matches the content.
2026-05-21 02:10:27 -07:00
Chris LuandGitHub 303c2be38d feat(fix): rebuild lost EC index (.ecx) and .vif from local shards (#9596)
weed fix -ecx reconstructs the .dat from the local data shards, scans the
needles, and writes a fresh ascending-sorted .ecx containing only live
entries — the same on-disk index WriteSortedFileFromIdx emits at encode
time. When the .vif is also missing it is regenerated from the inferred
EC ratio (flags > .vif > shard-count inference / 10+4) and the .dat size
recovered from the scan.

When some data shards are missing but at least dataShards shards survive,
the missing shards are first reconstructed from the survivors via
Reed-Solomon, so a partial shard set is repaired too.

Also makes erasure_coding.WriteDatFile de-stripe using len(shardFileNames)
instead of the DataShardsCount constant, so the caller's actual data-shard
count is honored (behavior-preserving for the default 10, and fixing the
existing caller that already passes ECContext.DataShards).

This recovers an EC volume whose sealed index was lost from every node
while enough shards survive, a state neither ec.rebuild nor ec.decode can
repair because both require an existing .ecx.

Flags: -ecx, -ecDataShards, -ecParityShards. Run with the volume server
stopped.
2026-05-21 00:41:27 -07:00
9b9fdb5b76 fix(s3): sync IAM policies to advanced IAM Manager policy engine (#9577)
* fix(s3): sync IAM policies to advanced IAM Manager policy engine

* test(s3): add unit tests for PutPolicy/DeletePolicy IAM Manager sync

* fix(s3): flush loaded policies in SetIAMIntegration, drop extra reload

Sync the policies already loaded from the credential store into the IAM
Manager's engine from SetIAMIntegration itself, instead of re-running a
full LoadS3ApiConfigurationFromCredentialManager after setup. This covers
both startup orderings without a second filer round-trip or racing the
async loader goroutine: if the load won, the policies are in memory to
push; if SetIAMIntegration won, the load's own sync runs afterward.

Move the runtime PutPolicy/DeletePolicy sync out of the iam.m write lock
so the per-request auth RLock path isn't blocked by the policy recompile.

* fix(s3): serialize IAM manager policy resync to avoid stale snapshots

SyncRuntimePolicies replaces the manager's full policy set, so applying a
policy view captured before a later mutation can resurrect a deleted
policy or drop a new one. Funnel every path (PutPolicy, DeletePolicy,
SetIAMIntegration, and the credential-manager load) through a single
resyncIAMManagerPolicies that serializes on a dedicated mutex and reads
iam.policies fresh at apply time, so the live map always wins regardless
of interleaving. The load now installs the config into iam.policies
before resyncing, closing the window where the manager held policies the
map didn't yet have.

---------

Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-05-21 00:39:42 -07:00
Chris LuandGitHub 7e4691f2dc test(ec): make multi-disk EC balance disk-spread assertion deterministic (#9595)
test(ec): pre-populate disks so multi-disk EC balance spread is deterministic

The multidisk shard-loss regression asserts EC shards spread across more
than one disk per node, but that only holds for disks the balancer can see.
The master enumerates a physical disk only when it already holds a volume
or EC shard — an empty disk leaves no trace, since heartbeats aggregate
capacity per disk type, not per physical disk. So whether the post-encode
balance spread shards depended on how the master happened to place the
filler volumes across disks, which varies by environment: the test passed
locally (shards on 5 disks) but produced one disk per node in CI and failed
the "got 3 disks across 3 nodes" assertion.

Grow a few volumes on each server before encoding so every physical disk
holds a volume and is visible to the balancer. The volume server places
each new volume on its least-loaded disk, so a handful of grows touches
every disk, making the spread deterministic. The assertion still has teeth:
it counts disks holding shard files, so a balancer that failed to spread
would still collapse to one disk per node.
2026-05-21 00:17:14 -07:00
Chris LuandGitHub 391f543ff2 fix(ec): correct multi-disk disk counting and EC balance shard attribution (#9594)
* fix(shell): count physical disks in cluster.status on multi-disk nodes

The master keys DataNodeInfo.DiskInfos by disk type, so several same-type
physical disks on one node collapse into a single DiskInfo entry. cluster.status
(printClusterInfo) and CountTopologyResources counted len(DiskInfos), reporting
one disk per node instead of the real physical disk count, while volume.list and
the admin ActiveTopology already split per physical disk.

Route both counters through DiskInfo.SplitByPhysicalDisk so a node with N
same-type disks reports N. Cosmetic/diagnostic only; placement already uses the
per-disk activeDisk map.

* fix(ec): attribute EC balance source disk per shard and reject same-node moves

On multi-disk nodes the EC balance worker built a node-level view that kept only
the first physical disk id per (node, volume), so a move of a shard living on a
different disk reported the wrong source disk. That source disk drives the
per-disk capacity reservation, so the wrong disk drifts the capacity model the
EC placement planner relies on. Track shards per physical disk and resolve the
actual source disk for every emitted move (dedup, cross-rack, within-rack,
global), keeping the per-disk view consistent as simulated moves are applied.

Also close a data-loss trap: VolumeEcShardsDelete is node-wide (it removes the
shard from every disk on the node) and copyAndMountShard skips the copy when
source and target addresses match, so a same-node move would erase a shard it
never copied. isDedupPhase now requires the same node AND disk, and Validate /
Execute reject same-node cross-disk moves outright.

* fix(ec): spread EC balance moves across destination disks

Port the shell ec.balance pickBestDiskOnNode heuristic to the EC balance
worker so a moved shard is placed on a good physical disk instead of always
deferring to the volume server (target disk 0). The detection now builds a
per-physical-disk view of each node (free slots split from the node total, exact
EC shard count, disk type, discovered from both regular volumes and EC shards)
and, for each cross-rack, within-rack, and global move, chooses the destination
disk by ascending score:
  - fewer total EC shards on the disk,
  - far fewer shards of the same volume on the disk (spread a volume's shards
    across disks for fault tolerance), and
  - data/parity anti-affinity (a data shard avoids disks holding the volume's
    parity shards and vice versa).

Planned placements are reserved on the in-memory model during a run so multiple
shards moved to the same node spread across its disks rather than piling on one.

* fix(ec): bring EC balance worker to parity with shell ec.balance

The worker's cross-rack and within-rack balancing balanced shards by total
count; the shell balances data and parity shards separately with anti-affinity
and honors replica placement. Port that logic so the automatic balancer makes
the same fault-tolerance-aware decisions as the manual command:

- Cross-rack and within-rack now run a two-pass balance: data shards spread
  first, then parity shards spread while avoiding racks/nodes that already hold
  the volume's data shards (anti-affinity), mirroring doBalanceEcShardsAcrossRacks
  and doBalanceEcShardsWithinOneRack.
- Optional replica placement: a new replica_placement config (e.g. "020")
  constrains shards per rack (DiffRackCount) and per node (SameRackCount); empty
  keeps the previous even-spread behavior.
- The data/parity boundary is resolved from a per-collection EC ratio (standard
  10+4 here), replacing the previously hardcoded constant at the call sites.

Selection is deterministic (sorted keys) to keep behavior reproducible.

* refactor(ec): extract shared ecbalancer package for shell and worker

The EC shard balancing policy was duplicated between the shell ec.balance
command and the admin EC balance worker, and the two had drifted (multi-disk
handling, data/parity anti-affinity, replica placement). Extract the policy into
a new pure package, weed/storage/erasure_coding/ecbalancer, that both callers
share so it cannot drift again.

- ecbalancer.Plan(topology, options) runs the full policy (dedup, cross-rack and
  within-rack data/parity two-pass with anti-affinity, global per-rack balance,
  and diversity-aware disk selection) over a caller-built Topology snapshot and
  returns the shard Moves. It depends only on erasure_coding and super_block.
- The worker builds the Topology from the master topology and turns Moves into
  task proposals; the shell builds it from its EcNode model and executes Moves
  via the existing move/delete RPCs. Per-collection EC ratio resolution stays in
  each caller (passed as Options.Ratio).
- Options expose the two genuine policy differences: GlobalUtilizationBased
  (worker balances by fractional fullness; shell by raw count) and
  GlobalMaxMovesPerRack (worker moves incrementally across cycles; shell drains
  in one pass).

The shell keeps pickBestDiskOnNode for the evacuate command. Policy tests move to
the ecbalancer package; the shell and worker keep their adapter/execution tests.

* fix(ec): restore parallelism and per-type/full-range balancing after ecbalancer refactor

Address regressions and gaps from the ecbalancer extraction:

- Shell ec.balance honors -maxParallelization again: planned moves run phase by
  phase (preserving cross-phase dependencies) with bounded concurrency within a
  phase. Apply mode does only the RPCs concurrently; dry-run stays sequential and
  updates the in-memory model for inspection.
- Rack and node balancing gate on per-type spread (data and parity separately)
  instead of combined totals, so a data/parity skew is corrected even when the
  per-rack/node totals are even.
- Global rack balancing iterates the full shard-id space (MaxShardCount) so
  custom EC ratios with more than the standard total are candidates.
- Cross-rack planning decrements the destination node's free slots per planned
  move, so limited-capacity targets are no longer over-planned.

* fix(ec): make EC dedup keeper deterministic and capacity-aware

When a shard is duplicated across nodes, keep the copy on the node with the most
free slots and delete the duplicates from the more-constrained nodes, relieving
capacity pressure where it is tightest. Tie-break on node id so the choice is
deterministic. This unifies the shell and worker (the shell previously kept the
least-free node, an incidental default) on the more sensible behavior.

* fix(ec): restore global volume-diversity and per-volume move serialization

Two more behaviors lost in the ecbalancer refactor:

- Global rack balancing again prefers moving a shard of a volume the destination
  does not hold at all before adding another shard of an already-present volume
  (two-pass, mirroring the old balanceEcRack), keeping each volume's shards
  spread across nodes.
- Shell apply-mode execution serializes a single volume's moves within a phase
  while still running different volumes in parallel, so concurrent moves of the
  same volume cannot race on its shared .ecx/.ecj/.vif sidecar files.

* fix(ec): key EC balance shards by (collection, volume id)

A numeric volume id can be reused across collections, and EC identity is
(collection, vid) (see store_ec_attach_reservation.go). The ecbalancer keyed
Node.shards by vid alone, so volumes sharing an id across collections merged into
one entry — letting dedup delete a "duplicate" that is actually a different
collection's shard, and letting moves act across collections. Key shards by
(collection, vid) throughout so each volume stays distinct.

* fix(ec): credit freed capacity from dedup before later balance phases

Dedup deletions are simulated only by applyMovesToTopology, which cleared shard
bits but did not return the freed disk/node/rack slots. Later phases reject
destinations with no free slots, so a slot opened by dedup could not be reused in
the same Plan/ec.balance run. applyMovesToTopology now credits the freed
disk/node/rack capacity for dedup moves (non-dedup moves still rely on the inline
accounting their phase already did).

* test(ec): add multi-disk EC balance integration test

Cover issue 9593 end-to-end at the unit level the old tests missed: build the
master's actual multi-disk wire format (same-type disks collapsed into one
DiskInfo, real DiskId only in per-shard records), run it through a real
ActiveTopology and the Detection entry point, then replay the planned moves with
the volume server's true semantics (node-wide VolumeEcShardsDelete) and assert no
EC shard is ever lost. Covers a balanced spread, a one-node-concentrated volume,
and a multi-rack spread, and asserts moves are safe (no same-node cross-disk),
correctly attributed to the source disk, and redistribute concentrated volumes
across both other racks and multiple destination disks.

* fix(ec): aggregate per-disk EC shards when verifying multi-disk volumes

collectEcNodeShardsInfo overwrote its per-server entry for each EcShardInfo of a
volume. A multi-disk node reports one EcShardInfo per physical disk holding shards
of the volume, so only the last disk's shards survived — the node looked like it
was missing shards it actually had. This made ec.encode's pre-delete verification
(and ec.decode) under-count volumes whose shards are spread across disks on one
server, falsely aborting the encode on multi-disk clusters. Union the per-disk
shard sets per server instead.

Also make verifyEcShardsBeforeDelete poll briefly: shard relocations reach the
master via volume-server heartbeats, so a freshly distributed shard set may not be
fully visible the instant the balance returns. Retry before concluding the set is
incomplete; genuine loss still fails after the retries are exhausted.

* test(ec): end-to-end multi-disk EC balance shard-loss regression

Start a real cluster of multi-disk volume servers (3 servers x 4 disks),
EC-encode a volume, run ec.balance, and assert hard invariants the prior
integration tests only logged: after encode all 14 shards exist, ec.balance loses
no shard, shards span more than one disk per node, and cluster.status counts
physical disks (not one per node). This reproduces issue 9593 end to end and would
have caught the multi-disk shard-aggregation bug fixed alongside it.

* fix(ec): bring EC balance worker/plugin path to parity with shell

- Per-volume serialization and phase order: key the plugin proposal dedupe by
  (collection, volume) instead of (volume, shard, source), so the scheduler runs
  only one of a volume's moves at a time (within a run and against in-flight jobs).
  Concurrent same-volume moves raced on the volume's .ecx/.ecj/.vif sidecars; and
  because the planner emits a volume's moves in phase order, they now execute in
  order across detection cycles, matching the shell.
- disk_type "hdd": normalize via ToDiskType (hdd -> "" HardDriveType) while keeping
  a "filter requested" flag, so disk_type=hdd matches the empty-keyed HDD disks
  instead of nothing; apply the canonical type to planner options and move params.
- Replica placement: expose shard_replica_placement in the admin config form and
  read it into the worker config, mirroring ec.balance -shardReplicaPlacement.

* test(ec): rename worker in-process test (not a real integration test)

The worker-package multi-disk tests build a fake master topology and simulate
move execution; they are not real-cluster integration tests. Rename
integration_test.go -> multidisk_detection_test.go and drop the Integration
prefix so 'integration' refers only to the real-cluster E2Es in test/erasure_coding.

* ci(ec): remove redundant ec-integration workflow

ec-integration.yml duplicated EC Integration Tests under the same workflow name
but ran only 'go test ec_integration_test.go' (one file), so it never ran new
test files (e.g. multidisk_shardloss_test.go) and was a strict, path-filtered
subset of ec-integration-tests.yml, which already runs 'go test -v' over the whole
test/erasure_coding package on every push/PR.

* fix(ec): worker falls back to master default replication for EC balance

For strict parity with the shell, the EC balance worker now uses the master's
configured default replication as the replica-placement fallback when no explicit
shard_replica_placement is set, instead of always defaulting to even spread.

The maintenance scanner reads it via GetMasterConfiguration each cycle and passes
it through ClusterInfo.DefaultReplicaPlacement; detection resolves the constraint
(explicit config wins, else master default, else none) in resolveReplicaPlacement.
A zero-replication default (the common 000 case) still means even spread, so the
common configuration is unchanged.

* fix(ec): plugin path populates master default replication too

The plugin worker built ClusterInfo with only ActiveTopology, so the master
default replication fallback added for the maintenance path never reached
plugin-driven EC balance detection — empty shard_replica_placement still meant
even spread there. Fetch the master default via GetMasterConfiguration (new
pluginworker.FetchDefaultReplicaPlacement) and set ClusterInfo.DefaultReplicaPlacement
so both detection paths resolve replica placement identically to the shell.

* docs(ec): empty shard replica placement uses master default, not even spread

The EC balance config text (admin plugin form, legacy form help text, and
the struct/proto field comments) still said an empty shard_replica_placement
spreads evenly. The runtime resolves empty to the master default replication
(resolveReplicaPlacement), matching shell ec.balance, with even spread only
when that default is empty or zero. Update the text to match and regenerate
worker_pb for the proto comment change.
2026-05-20 23:31:21 -07:00
Chris LuandGitHub afcc491517 test: fix fd leak in the Samba DLM handoff test (promote xfail checks) (#9592)
test(mount): fix fd leak that deadlocked the DLM handoff check

The cross-mount handoff checks held a file open on mount 2 via fd 9 to
keep the distributed lock, then started the SMB writer in a background
subshell. The subshell inherited fd 9, so the SMB writer kept the file
open and waited on a lock held by its own descriptor; the put could
never complete, and the two checks were parked as expected-fail.

Close fd 9 in the subshell (9>&-) so the writer does not hold the file.
The waiter now acquires the freed lock within ~1s, so the two checks are
real assertions and the xfail machinery is gone.
2026-05-20 16:17:13 -07:00
Chris LuandGitHub a5d0e4a735 Samba-over-FUSE integration test and distributed-lock handoff fixes (#9590)
* test(mount): add Samba over FUSE integration test

Export a SeaweedFS FUSE mount over SMB with smbd and drive it with
smbclient: file round-trips, directories, rename, large-file chunking,
recursive upload, cross-protocol consistency, and deletes.

A second -dlm mount adds locking coverage: POSIX fcntl byte-range locks,
distributed-lock write coordination, and concurrent writers. The two
cross-mount handoff checks currently fail and pin a known limitation -
the distributed lock is released on FUSE Release, which the kernel can
delay under contention.

Runs locally via test/samba/run.sh or in Docker via the compose file;
wired into CI as samba-integration.yml.

* fix(cluster): release distributed lock without racing the renewal goroutine

Stop() closed the cancel channel, slept 10ms, then unlocked using
renewToken. A renewal in flight during that window rotates the token on
the server, so the unlock may be sent with a stale token, fail with a
mismatch, and leave the lock to linger until its TTL expires - stalling
other mounts waiting to write the same file.

Wait for the renewal goroutine to exit before unlocking. The channel
close also makes the renewToken read happen-after the last renewal.

* fix(cluster): poll for distributed lock acquisition without exponential backoff

A mount waiting to write a file held by another mount acquired through
util.RetryUntil, whose backoff grows to several seconds. Once the holder
released, the waiter could sleep that long before retrying, stretching
the cross-mount handoff past client timeouts.

Poll at the steady ~1s cadence AttemptToLock already enforces instead.

* test(mount): tighten Samba harness and mark the DLM handoff checks xfail

Run the workflow for weed/cluster changes, fail fast when the filer or
smbd port never opens, and fold the recursive mput result into its own
assertion so it cannot false-pass.

Mark the two cross-mount handoff checks expected-fail: they pin the
remaining DLM liveness bug (the lock is freed only on the delayed FUSE
Release) without failing CI, and turn the suite red if the handoff is
ever fixed.

* fix(cluster): keep a wedged renewal shutdown from sending a stale unlock

If the renewal goroutine is stuck in a slow RPC, Stop() fell through to
unlock anyway once it timed out waiting. A late renewal can rotate
renewToken, so that unlock races it, is rejected on a stale token, and
leaves the lock lingering until its TTL regardless. On the timeout path,
skip the unlock and let the TTL expire the lock instead.

* fix(cluster): wake the long-lived lock renewal loop promptly on Stop

StartLongLivedLock's renewal loop slept uninterruptibly between attempts,
up to 5*renewInterval (2.5*lockTTL) while unlocked. Stop() waits only
lockTTL+2s for the goroutine to exit, so a Stop() during that backoff
would time out before the goroutine woke and closed renewalDone,
breaking the shutdown synchronization. Sleep on a timer with a select on
cancelCh so the loop exits immediately.
2026-05-20 14:52:17 -07:00
Chris LuandGitHub a17dca7009 fix(filer): don't disable the SQL idle connection pool when unconfigured (#9591)
* fix(filer): don't disable the SQL idle connection pool when unconfigured

The mysql/mysql2/postgres stores called SetMaxIdleConns(maxIdle)
unconditionally, so an unset connection_max_idle (0) actively kept zero
idle connections - every query opened and closed a fresh connection
instead of reusing the pool.

Only apply the value when it's set; otherwise leave database/sql's
default idle pool of 2 in place.

* comments: shorten idle-pool note

* fix(filer): default the SQL idle pool via config, keep explicit 0 honored

Apply the idle-pool default at the config layer with SetDefault instead of
guarding the SetMaxIdleConns call. An absent connection_max_idle now reads
back as 2 (pool stays on), while an explicit 0 flows through to
SetMaxIdleConns(0) so operators can still disable idle pooling on purpose.
2026-05-20 14:04:23 -07:00
Chris LuandGitHub 024b59fb31 fix(ec): pack EC shards onto fewer disks instead of refusing the task (#9588)
The planner refused to create an EC task unless it found totalShards
distinct (server, disk_id) targets, so a cluster with fewer disks than
shards (e.g. 8 single-disk servers for a 10+4 scheme) could never encode.

A disk safely holds several distinct shards of one volume: each is its own
.ecNN file and ReceiveFile keys by that extension. Drop the strict check and
let createECTargets round-robin shards across the available disks, matching
ec.encode's "4,4,3,3" fallback. The minTotalDisks floor (ceil(total/parity))
already keeps any disk under parityShards shards, so the volume still
survives losing any one disk.

Reserve capacity for the actual per-disk shard count rather than assuming
one shard each, so packing doesn't over-commit disk slots.
2026-05-20 11:50:42 -07:00
Chris LuandGitHub 5af7d12f04 fix(filer.sync): keep sync_offset fresh while the source is read-only (#9589)
* fix(filer.sync): keep sync_offset fresh while the source is read-only

sync_offset holds the timestamp of the last replicated source event, so
monitoring derives lag from now-sync_offset. A read-only source emits no
metadata events, so the gauge froze at the last write and the derived lag
grew without bound, making thresholds unusable.

The source filer now sends an idle heartbeat carrying its current time
while a subscriber is caught up to the buffer head. filer.sync uses it to
advance the gauge, so now-sync_offset reflects real lag. Heartbeats are
opt-in (client_supports_idle_heartbeat), are never written to the metadata
log, and do not move the resume checkpoint, so a restart still resumes
from the last real event.

* fix(filer.sync): gate idle heartbeat on the read cursor, not SinceNs

In metadata-chunks mode persisted entries replay as log file refs and
never reach eachLogEntryFn, so lastSeenTsNs stays put and a caught-up
subscriber with an old SinceNs would never get a heartbeat. Use the
read cursor (lastReadTime), which advances in that mode too, max'd with
lastSeenTsNs so the in-memory backlog-then-idle case still works while
the cursor returned to the caller has not yet updated.
2026-05-20 11:26:37 -07:00
Chris LuandGitHub 4385b86bf1 fix(shell): volumeServer.evacuate no longer panics on a nil volume (#9587)
adjustAfterMove now removes the moved volume from the source disk's
VolumeInfos in place: it swaps the entry with the last one and nils the
tail. evacuateNormalVolumes ranges directly over that same slice, so the
niled tail slot is later read as a nil *VolumeInformationMessage and the
move attempt panics on vol.DiskType.

Iterate over a snapshot of the slice so in-place removals during a move
cannot leave nil holes in the loop.
2026-05-20 10:27:00 -07:00
Chris LuandGitHub c00aa90990 fix(s3/audit): populate requester for GET/HEAD/IAM operations (#9581)
Authentication records the identity with r.WithContext, which returns a
request copy. Handlers that log their own audit entry (PUT, DELETE,
tagging) see it, but GET/HEAD object and IAM operations rely on track()'s
fallback entry, which is built from the original request the auth copy
never reached - so requester came out empty.

Install a mutable identity holder on the request before authentication
and have SetIdentityNameInContext record into it. The holder is shared by
pointer across every request copy, so the fallback entry recovers the
authenticated requester. The per-request context value still takes
precedence, so nothing changes for handlers that see the auth copy.
2026-05-20 10:13:33 -07:00
Chris LuandGitHub e332b97d52 fix(shell): volume.balance no longer drains all volumes onto one server (#9579)
* fix(shell): volume.balance no longer drains all volumes onto one server

The density-based capacity function reads per-disk VolumeInfos sizes, but
adjustAfterMove only updated VolumeCount and the selectedVolumes map. The
planner re-read a stale topology after every move, so the source node's
density never dropped and it kept moving volumes until that node was empty.

Move the volume's size accounting between disks after each planned move so the
density recomputes and the loop converges to an even distribution.

* refactor(shell): O(1) volume removal and direct disk lookup in adjustAfterMove

removeVolumeInfo swaps with the last element instead of shifting, and the disk
is fetched by key rather than ranging the DiskInfos map.
2026-05-20 01:39:23 -07:00
Chris Lu 868849392c 4.27 2026-05-20 00:25:16 -07:00
Chris LuandGitHub a4415c39aa fix(mount): keep periodic metadata flush from dropping concurrent chunk uploads (#9574)
* fix(mount): keep periodic metadata flush from dropping concurrent chunk uploads

The periodic flush snapshotted entry.Chunks, then ran CompactFileChunks and
MaybeManifestize (the manifest upload is a network round trip) before
reassigning entry.Chunks. Async uploaders append freshly uploaded chunks
during that window, and the reassignment overwrote them: the data stayed on
the volumes but the file lost those chunk references, leaving zero-filled
holes on read. Large sequential writes such as cat of two 15 GiB files hit
several flush cycles and ended up corrupted.

Snapshot the chunk list under the entry lock with a length marker, do the
slow compaction and manifestization on the snapshot, then splice the
processed prefix back in front of whatever chunks arrived after the
snapshot.

* mount: drop redundant slice copies in the flush splice

processedPrefix is freshly built and the tail sub-slice is consumed
immediately under the entry lock, so append straight onto processedPrefix
instead of allocating two throwaway copies.
2026-05-19 20:47:52 -07:00
9914e6af30 chore(weed/command): prune unused functions (#9573)
* chore(weed/command): prune unused functions

* drop now-unused closed field and renderLocked guard

---------

Co-authored-by: Chris Lu <chris.lu@gmail.com>
2026-05-19 17:45:50 -07:00
Chris LuandGitHub cc5ef1b741 feat(s3): add TagUser, UntagUser, ListUserTags IAM actions (#9572)
* feat(s3): add TagUser, UntagUser, ListUserTags IAM actions

Adds AWS IAM-compatible user tag operations on the embedded IAM
endpoint. Tags persist in the Identity proto as a repeated UserTag
field; the existing 50-tag / 128-byte-key / 256-byte-value AWS limits
are enforced. Pagination is stubbed (IsTruncated=false) since the
50-tag cap means all tags fit in a single response.

* review: validate UntagUser TagKeys entries

parseTagKeysParams now rejects empty keys and keys past
MaxUserTagKeyLength; UntagUser additionally requires at least one
TagKeys.member.N entry to match AWS validation behavior.

* review: pre-allocate user-tag merge and filter slices

mergeUserTags now allocates the combined existing+incoming capacity
up front; UntagUser builds the filtered slice via make with the full
ident.Tags capacity instead of ident.Tags[:0:0], which forced a
reallocation on every append.

* review: cover duplicate-in-request and invalid TagKeys cases

Regression tests assert TagUser rejects two members with the same key
in one request, and UntagUser rejects missing/empty/oversized TagKeys
entries.
2026-05-19 17:35:44 -07:00
Chris LuandGitHub 37b6a14b0d feat(s3): add four bucket configuration handlers (#9570)
* feat(s3): add four bucket configuration handlers

- GetBucketPolicyStatus: computes IsPublic from the existing bucket policy
- PutBucketRequestPayment: companion writer to the existing GET; accepts
  only BucketOwner
- GetBucketAccelerateConfiguration: returns <Status>Suspended</Status>
- GetBucketLogging: returns an empty BucketLoggingStatus

Lets AWS SDK probes succeed instead of returning MethodNotAllowed.

* review: route GetBucketPolicyStatus through checkBucket

Mirrors the existence/auth gating used by other bucket handlers and
drops the bespoke filer_pb lookup so NoSuchBucket precedence is
consistent across the API surface.

* review: cap PutBucketRequestPayment body with MaxBytesReader

The body is unmarshalled as RequestPaymentConfiguration, which is a
handful of bytes; reject excessively large payloads up front and
defer Close immediately after wrapping.

* review: gate static getters on checkBucket

GetBucketAccelerateConfiguration and GetBucketLogging now run the
standard bucket existence check before returning the static
Suspended / empty-status response so a missing bucket cannot appear
to have valid configuration.

* review: share cache helper across misc tests; check io.ReadAll error

Accelerate and Logging tests now run through newMiscTestServer like
the others so the checkBucket guard sees a cached bucket; the
ReadAll error is explicitly checked.
2026-05-19 17:35:08 -07:00
Chris LuandGitHub cee2bf697c feat(s3): stub bucket configuration list endpoints (#9571)
* feat(s3): stub bucket configuration list endpoints

Adds Get and List handlers for Analytics, Inventory, IntelligentTiering,
and Metrics bucket configurations. List returns an empty result with
IsTruncated=false; single-get returns NoSuchConfiguration so SDK error
parsing remains predictable.

* review: gate stubs on bucket existence

All eight stub handlers now call checkBucket via stubBucketGuard so
NoSuchBucket takes precedence over NoSuchConfiguration / empty-list
responses, matching AWS S3 precedence. Tests provide a cached bucket
so the guard sees it as present.
2026-05-19 17:34:51 -07:00
Chris LuandGitHub 285025eb73 s3api: support group inline policies + Condition enforcement (#9569)
* test(s3api): cover IAM inline policy aws:SourceIp + group inline gap

Unit tests under weed/s3api/ drive PutUserPolicy / PutGroupPolicy → reload
→ VerifyActionPermission with a synthetic 127.0.0.1 request and assert that
the policy's IpAddress condition flips the outcome.

The user-policy cases pass on master (hydrateRuntimePolicies already routes
inline docs through the policy engine, so Condition blocks are honored end-
to-end). The group-policy case fails: PutGroupPolicy still returns
NotImplemented, so a group inline doc never lands in the engine.

Integration counterparts live under test/s3/iam/ and exercise the same
paths against a live SeaweedFS S3+IAM endpoint.

* s3api: support group inline policies + Condition enforcement

PutGroupPolicy/GetGroupPolicy/DeleteGroupPolicy/ListGroupPolicies used to
return NotImplemented in embedded IAM mode, so anything attached to a
group as an inline doc — including aws:SourceIp or any other Condition —
was simply unreachable.

Wire the four endpoints to the credential-store methods that were
already in place (memory, postgres, filer_etc all implement
GroupInlinePolicyStore). On every config reload, hydrateRuntimePolicies
now also walks LoadGroupInlinePolicies, registers each doc in the IAM
policy engine under __inline_group_policy__/<group>/<policy>, and
appends that key to Group.PolicyNames so evaluateIAMPolicies picks it up
through its existing group walk. PutGroupPolicy/DeleteGroupPolicy are
added to the ReloadConfiguration trigger list in DoActions.

Side fix: MemoryStore.LoadConfiguration now surfaces store.groups too.
Without it iam.groups never repopulated on a memory-store reload, so
group policy evaluation silently no-op'd whether the policy was inline
or attached. The existing tests didn't notice because no test reloaded
through cm after creating a group.

The NotImplemented unit test is inverted to drive the new round-trip.

* s3api: drop redundant refreshIAMConfiguration from Put/DeleteGroupPolicy

DoActions already triggers ReloadConfiguration for both actions via the
explicit reload list, so calling refreshIAMConfiguration inline runs the
load twice per request. Per PR review.

* s3api: scope group-policy resource names per test; tighten deny polling

- Integration test resource names get a per-test suffix so retried or
  parallel CI jobs don't trip EntityAlreadyExists / BucketAlreadyExists.
- Deny-path Eventually loops gate on AccessDenied via a typed helper
  rather than any non-nil error; transient setup errors no longer end
  the wait prematurely.
- ListGroupPolicies returns ServiceFailure when the credential manager
  is nil, matching Put/Get/DeleteGroupPolicy.

* test(s3 iam): cover both IPv4 and IPv6 loopback in allow CIDRs

CI runners with happy-eyeballs resolve `localhost` to ::1 first, in
which case a 127.0.0.0/8-only allow would silently never match and the
deny-driven enforcement test would hang for the allow case. Add ::1/128
to every loopback-matching policy so the allow path works regardless of
which loopback family the SDK lands on.
2026-05-19 16:03:45 -07:00
Chris LuandGitHub 77ac781bbd fix(ec): VolumeEcShardsInfo walks every disk on multi-disk servers (#9568)
* fix(ec): VolumeEcShardsInfo walks every disk on multi-disk servers

When a volume server holds EC shards for the same vid across more than
one disk, each DiskLocation registers its own EcVolume entry and
Store.FindEcVolume returns whichever one it hits first. The shard-info
RPC iterated only that single EcVolume's Shards, so the response missed
every shard mounted on a sibling disk.

The worker's verifyEcShardsBeforeDelete sums the per-server responses
into a union bitmap and refuses to delete the source volume when the
union falls short of dataShards+parityShards. On multi-disk
destinations, the union was systematically under-counted and source
deletion got blocked even though all shards were physically present and
mounted.

Walk every DiskLocation in the handler and emit the deduplicated union
of all shards. The .ecx-backed fields (file counts, volume size) still
come from a single EcVolume since every disk's entry opens the same
.ecx via NewEcVolume's cross-disk fallback.

Tests:
- TestVolumeEcShardsInfo_AggregatesAcrossDisks unit test in
  weed/server/.
- test/volume_server/grpc/ec_verify_multi_disk_test.go integration test
  drives the full generate -> mount -> redistribute -> restart ->
  reconcile path and asserts both VolumeEcShardsInfo and
  VerifyShardsAcrossServers + RequireFullShardSet (the production
  source-deletion gate) report all 14 shards.
- ec_multi_disk_lifecycle_test.go tightened: replaces the
  "VolumeEcShardsInfo only sees one disk's EcVolume" workaround with a
  full-shard-set assertion.

* review: use ShardBits bitmask + cap-pre-allocation for shard dedup
2026-05-19 14:58:56 -07:00
Chris LuandGitHub f72983c1fd fix(s3): stop S3 Tables routes from swallowing buckets named "buckets" or "get-table" (#9566)
* fix(s3): stop S3 Tables routes from swallowing buckets named "buckets" or "get-table"

The S3 Tables REST endpoints share top-level paths with the regular S3
API (/buckets for ListTableBuckets/CreateTableBucket, /get-table for
GetTable). They are registered first on the same router as the bucket
subrouter, so a path-style request such as GET /buckets?list-type=2 on
a bucket actually named "buckets" matched ListTableBuckets and returned
JSON. AWS SDK V2 (and Hadoop s3a / Spark) then failed XML parsing with
"Unexpected character '{' (code 123) in prolog".

Disambiguate by requiring the AWS V4 credential scope to name the
s3tables service on the colliding routes. Regular S3 SDKs sign with
service=s3, S3 Tables SDKs sign with service=s3tables, and the scope is
present in both the Authorization header and the X-Amz-Credential query
parameter for presigned URLs, so the matcher works for both flavors.

ARN-bearing S3 Tables routes (/buckets/<arn>, /namespaces/<arn>, etc.)
already cannot collide because colons are not valid in bucket names, so
they are left untouched.

* fix(s3): accept AWS JSON RPC content type as S3 Tables intent signal

The Iceberg catalog integration tests send unsigned PUT /buckets with
Content-Type: application/x-amz-json-1.1 to create table buckets. With
only the credential-scope check, those requests fell through to the
regular S3 CreateBucket handler and the suite went red on this branch.

Extend the matcher so a request is recognized as S3 Tables when either:

  - its AWS V4 credential scope names SERVICE=s3tables; or
  - it carries the canonical AWS JSON RPC 1.1 content type and is
    unsigned (a request explicitly signed for SERVICE=s3 still wins).

The regular S3 SDKs do not send application/x-amz-json-1.1, so the
signal is safe for the colliding paths (/buckets, /get-table).

Also add an AWS SDK V2 for Go integration test under
test/s3/sdk_v2_routing/ that drives the SDK's own XML deserializer
against a bucket literally named "buckets" and "get-table" — the SDK
errors before the test asserts if the server returns the wrong body
shape. Wired up via .github/workflows/s3-sdk-v2-routing-tests.yml,
mirroring the etag/acl workflow.

* s3api: extend service matcher to all S3 Tables routes; simplify scope check

- Apply serviceMatcher to every S3 Tables route, not just the bare-path
  ones. ARN-bearing paths could otherwise be hit by an S3 object key
  that starts with arn:aws:s3tables:..., inside a bucket named
  "buckets", "namespaces", "tables", or "tag". One matcher everywhere
  closes both collision classes.
- Replace strings.Split + index lookup with strings.Contains for the
  credential-scope check. The scope shape is fixed at
  AK/DATE/REGION/SERVICE/aws4_request, slashes only delimit components,
  and access keys are alphanumeric — so /s3tables/ matches iff SERVICE
  is exactly s3tables. Existing unit cases (including the
  access-key-substring case) still pass.
- Read the GetObject body in the SDK v2 routing test with io.ReadAll;
  the single Read could return short and make the equality check flaky.

* s3api: drop content-type fallback; sign s3 tables harness traffic instead

The content-type fallback in isS3TablesSignedRequest let an anonymous
regular-S3 request whose body type is application/x-amz-json-1.1 hit
an S3 Tables route when the path-style object key happened to be
shaped like an S3 Tables ARN (e.g. PutObject on bucket "buckets"
with key arn:aws:s3tables:.../bucket/foo/policy). Narrow the matcher
back to the AWS V4 credential scope so only requests signed for
SERVICE=s3tables match the S3 Tables routes.

Update the Iceberg catalog test harness — the only caller still
sending unsigned PUT /buckets — to sign with SERVICE=s3tables. The
mini instance runs in default-allow mode, so the signature itself is
not verified; only the credential scope matters for the route match.

Drop the stale unit cases for the JSON-RPC content-type signal and
the routing test that exercised unsigned harness traffic.
2026-05-19 14:24:25 -07:00
Chris LuandGitHub cfc08fbf6c fix(volume): tombstone integrity check no longer flips volumes read-only (fixes #9563) (#9565)
* fix(volume): pass on-disk tombstone size to ReadData in verifyDeletedNeedleIntegrity

verifyDeletedNeedleIntegrity was forwarding TombstoneFileSize (-1) into
Needle.ReadData. A deletion tombstone is appended to .dat with DataSize=0
so the on-disk needle header carries Size=0; TombstoneFileSize is only
the .idx sentinel for "this entry is deleted" and is never written into
a needle header.

ReadBytes' size check therefore mismatched on every tombstone
(-1 != 0), returned ErrorSizeMismatch, and triggered the
4-byte-offset wrap-around retry in ReadData (offset + 32 GB). On any
volume large enough that offset+32 GB exceeds dat fileSize the retry
read EOF, CheckVolumeDataIntegrity reported corruption, and the loader
set noWriteOrDelete = true. Every volume whose last 10 .idx entries
included a deletion went read-only on startup — i.e. any healthy
volume where the most recent operations included a delete.

Pass Size(0) so the size check matches the on-disk tombstone header.

Add a regression test that writes three needles, deletes one, and
asserts CheckVolumeDataIntegrity succeeds with a tombstone at the .idx
tail. Without this fix the test reproduces the exact log shape from
the bug report:

  read 0 dataSize 32 offset <orig+32GB> fileSize <much smaller>: EOF
  verifyDeletedNeedleIntegrity ...idx failed: read data [N,N+32) : EOF

The Rust port guards its integrity-check size comparison with
!size.is_deleted() (seaweed-volume/src/storage/volume.rs) and never
hits this path, so no Rust mirror change is needed.

* test(seaweed-volume): mirror Go regression for deletion-tombstone integrity

The Rust integrity check already guards its size-mismatch comparison
with !size.is_deleted() (volume.rs:1859) and reads tombstone AppendAtNs
with body_size=0, so the Go regression fixed in the previous commit
does not apply. Lock that guarantee in with a parallel reload test:
write three needles, delete one, sync, reopen via Volume::new, assert
the volume is not flipped read-only.

Catches any future change that removes the deleted-entry guard or
re-introduces a size-strict path in check_volume_data_integrity for
tombstones.

* fix(volume): propagate io.EOF and ErrorSizeMismatch from verifyDeletedNeedleIntegrity

CheckVolumeDataIntegrity relies on identity comparison against io.EOF
and ErrorSizeMismatch to walk back through the last ten .idx entries
and tolerate a partial truncation at the tail (the "fix and continue"
loop). The live-needle branch in doCheckAndFixVolumeData already
returns those sentinels unwrapped; the deletion branch wrapped them
in fmt.Errorf, so a genuine .dat truncation past a tombstone offset
broke the recovery and flipped the volume read-only.

Mirror the live-needle handling: both verifyDeletedNeedleIntegrity
and doCheckAndFixVolumeData now short-circuit on io.EOF /
ErrorSizeMismatch and pass them through unwrapped. Other errors keep
their existing context wrapping.

Also tighten the regression test to capture lastAppendAtNs and assert
it's non-zero, so a future regression that skips the tombstone body
(and therefore never populates AppendAtNs) is caught even when the
err check still passes.
2026-05-19 13:11:19 -07:00
Chris LuandGitHub d57de6dc20 fix(s3): keep anonymous access working with EnableIam default (fixes #9557) (#9567)
fix(s3): keep anonymous access working with EnableIam default

`docker run seaweedfs` (and `weed mini` with no config) start with
EnableIam=true but no IAM config file and no identities. The advanced-IAM
init path was failing in 4.25 because of the missing STS signing key,
which masked a latent bug: SetIAMIntegration unconditionally flipped
isAuthEnabled to true, and isEnabled() also treated a non-nil
iamIntegration as auth-on. Once the mini SSE-S3 KEK landed in 4.26 the
STS fallback started succeeding, the integration got installed end to
end, and every anonymous S3 request bounced as AccessDenied.

Separate the two concerns: SetIAMIntegration just plumbs in the OIDC /
embedded-IAM machinery, and a new EnableAuthEnforcement opts in to
enforcement. The startup path calls it only when -s3.iam.config is
actually provided, so operators with explicit IAM configs still get auth
(preserves #7726). isEnabled() now reads isAuthEnabled only.
2026-05-19 13:03:30 -07:00
4476cb282b feat(filer): add atime to FuseAttributes + TouchAccessTime RPC (#9556)
* feat(filer): add atime field and TouchAccessTime RPC to filer proto

Introduce POSIX-style access-time tracking on the filer:
- FuseAttributes gains atime (field 22) and atime_ns (field 23).
- New TouchAccessTime RPC (and Touch{Access,Time}{Request,Response})
  lets read paths bump atime without going through UpdateEntry's
  chunk-rewrite/EqualEntry short-circuit.

Additive proto changes only; zero atime is treated as unset and
existing clients are unaffected. Java client proto is kept in lock
step.

Co-authored-by: Cursor <cursoragent@cursor.com>

* feat(filer): wire Atime through Attr codec with mtime fallback

Add Attr.Atime and round-trip it through EntryAttributeToPb /
EntryAttributeToExistingPb / PbToEntryAttribute. A zero proto atime
decodes as Mtime, so legacy entries report a sensible value and
freshly-created/updated entries default Atime to Mtime when callers
do not set it explicitly.

CreateEntry and UpdateEntry stamp Atime = Mtime (or Crtime) when it
is zero. TouchAccessTime later bypasses this path to write atime
alone via Store.UpdateEntry.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(filer): preserve atime in first epoch second on decode

The Atime decode branch previously treated any attr.Atime == 0 as
unset and overwrote it with Mtime, which drops valid timestamps in
the first second of the unix epoch where attr.Atime is 0 but
attr.AtimeNs > 0. Check both fields so we only fall back to Mtime
when both are zero.

Co-authored-by: Cursor <cursoragent@cursor.com>

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-05-19 10:22:17 -07:00
Chris LuandGitHub b63610cf8f volume: accept legacy needle CRC encoding on read (#9564)
Volumes written by versions before 3.09 (commit 056c480eb) store the
needle checksum using the deprecated CRC.Value() transform. When the
read path moved into readNeedleTail, the fallback that accepts both
encodings was dropped, so .dat files copied from old installs now fail
verification with "invalid CRC ... data on disk corrupted" even though
the data is intact. Restore the dual check, matching the surviving
fallback in volume_read.go.
2026-05-19 09:58:47 -07:00
Chris LuandGitHub c61d227613 s3api: verify source permission on CopyObject and UploadPartCopy (#9555)
* s3api: verify source permission on CopyObject and UploadPartCopy

The Auth middleware only authorized the destination because routes key on
the request URL. The source from X-Amz-Copy-Source was never evaluated,
so an STS session token scoped to one prefix could copy from any other
prefix in the same bucket.

Add AuthorizeCopySource on IdentityAccessManagement to run the full
bucket-policy + IAM/identity flow against the source, using a synthetic
GetObject request so action resolution lands on s3:GetObject (or
s3:GetObjectVersion when a source versionId is supplied). Both
CopyObjectHandler and CopyObjectPartHandler now invoke it before reading
the source.

* s3api: preserve presigned-URL session token on copy-source check

Presigned CopyObject / UploadPartCopy requests carry the STS session
token in the query string (X-Amz-Security-Token), not in a header.
Rebuilding the synthetic source URL from scratch dropped that token, so
the source authorization would fall through to non-STS paths and miss
session policy enforcement. Forward X-Amz-Security-Token from the
original query (alongside versionId), still excluding unrelated params
like uploadId/partNumber that would steer ResolveS3Action away from
s3:GetObject.
2026-05-18 21:35:53 -07:00
Chris LuandGitHub 7c252e1f16 fix(volume): reopen .idx writable after MarkVolumeWritable (fixes #9515) (#9526)
* fix(volume): reopen .idx writable after MarkVolumeWritable

When .vif has ReadOnly=true, load() opens .idx as O_RDONLY and builds a
SortedFileNeedleMap whose Put returns os.ErrInvalid. MarkVolumeWritable
only flipped noWriteOrDelete back to false and rewrote .vif, so writes
still failed at v.nm.Put. Reopen .idx in O_RDWR and rebuild v.nm in its
writable form (in-memory or leveldb small/medium/large) before flipping
the flag.

Mirror the same fix in seaweed-volume: the Rust load path leaves
CompactNeedleMap/RedbNeedleMap with no idx_file writer when the volume
boots read-only, so post-MarkVolumeWritable puts silently succeeded
in-memory only and were lost on the next restart. set_writable now
reattaches an append-mode writer when one is missing.

* fix(volume): keep old needle map until replacement is built; defer writable flag

Go: build the writable needle map into a local before swapping. A
construction failure now leaves v.nm pointing at the original
SortedFileNeedleMap so MarkVolumeWritable can roll back, instead of
stranding the volume with v.nm == nil.

Rust: attach the .idx writer before flipping no_write_or_delete to
false. A transient open/metadata failure used to leave the volume
marked writable with no writer attached, and subsequent puts would
silently skip the on-disk append.
2026-05-18 20:51:04 -07:00
Chris LuandGitHub 7c5296dfb1 fix(admin): switch file browser upload/download to filer gRPC + volume HTTP (#9538)
* fix(admin): switch file browser upload/download to filer gRPC + volume HTTP

The admin file browser proxied uploads and downloads through the filer's
HTTP listener, so the whole feature 404'd against filers started with
-disableHttp=true even though S3 still worked on its own port. Re-route
through the filer gRPC service: LookupDirectoryEntry + StreamContent for
reads (chunks flow straight from the volume servers), AssignVolume +
volume HTTP POST + CreateEntry for writes. Volume read tokens come from
jwt.signing.read.key when configured; the old jwt.filer_signing tokens
no longer apply since the filer HTTP surface is bypassed.

* admin file browser: propagate request context + track response writes

Pass r.Context() into uploadFileToFiler so a client disconnect cancels
the in-flight chunked upload instead of letting it run to completion
against the volume servers. For DownloadFile, replace the Content-Type
probe with a small response-writer wrapper that records whether headers
or bytes have actually been sent, so the error path can't silently
convert a pre-stream failure into a partial response if future code
moves the header-setting around.
2026-05-18 20:33:16 -07:00
Chris LuandGitHub 58c3fa802c fix(s3): keep host-less bucket catch-all so reverse proxies work (#9540)
When s3.domainName is set, all bucket-prefix routes were gated on a
matching Host header. Requests that arrive via an IP, an unlisted
hostname, or a reverse proxy that rewrites Host hit no router and bounce
back as 405/404 (and 503 once a proxy maps the upstream error).

Register the path-style catch-all unconditionally, after the
host-specific routers, so it only fires when no Host matcher applies.
2026-05-18 19:44:19 -07:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
d3f80444df build(deps): bump github.com/cognusion/imaging from 1.0.2 to 1.0.3 (#9552)
Bumps [github.com/cognusion/imaging](https://github.com/cognusion/imaging) from 1.0.2 to 1.0.3.
- [Commits](https://github.com/cognusion/imaging/compare/v1.0.2...v1.0.3)

---
updated-dependencies:
- dependency-name: github.com/cognusion/imaging
  dependency-version: 1.0.3
  dependency-type: direct:production
  update-type: version-update:semver-patch
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-05-18 19:43:33 -07:00
Chris LuandGitHub 0dc65e7069 fix(admin.plugin): include disk_id in EC execution plan (#9547)
TaskSource and TaskTarget carry disk_id on the wire, but the execution
plan map built for the admin UI dropped the field entirely. On a
multi-disk node holding shards of the same volume, there was no way to
tell from the plan which disk would receive each shard. Include
disk_id on each endpoint and target_disk_id on each shard assignment,
and extend the existing execution-plan test to set and assert the
field.
2026-05-18 19:43:18 -07:00
ᎠᎡ. Ѕϵrgϵ ѴictorandGitHub 18c6c24e47 Revise MinIO comparison in README for accuracy (#9548)
Updated the README to reflect the current status of MinIO, noting its ceased development and security concerns, along with changes in the descriptions of its features compared to SeaweedFS.
2026-05-18 19:32:54 -07:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
120901c883 build(deps): bump github.com/parquet-go/parquet-go from 0.28.0 to 0.30.1 (#9549)
Bumps [github.com/parquet-go/parquet-go](https://github.com/parquet-go/parquet-go) from 0.28.0 to 0.30.1.
- [Release notes](https://github.com/parquet-go/parquet-go/releases)
- [Changelog](https://github.com/parquet-go/parquet-go/blob/main/CHANGELOG.md)
- [Commits](https://github.com/parquet-go/parquet-go/compare/v0.28.0...v0.30.1)

---
updated-dependencies:
- dependency-name: github.com/parquet-go/parquet-go
  dependency-version: 0.30.1
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-05-18 19:28:42 -07:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
a79880ed41 build(deps): bump github.com/redis/go-redis/v9 from 9.18.0 to 9.19.0 (#9550)
Bumps [github.com/redis/go-redis/v9](https://github.com/redis/go-redis) from 9.18.0 to 9.19.0.
- [Release notes](https://github.com/redis/go-redis/releases)
- [Changelog](https://github.com/redis/go-redis/blob/master/RELEASE-NOTES.md)
- [Commits](https://github.com/redis/go-redis/compare/v9.18.0...v9.19.0)

---
updated-dependencies:
- dependency-name: github.com/redis/go-redis/v9
  dependency-version: 9.19.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-05-18 19:28:32 -07:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
f5aa776742 build(deps): bump github.com/Azure/azure-sdk-for-go/sdk/storage/azblob from 1.6.4 to 1.7.0 (#9551)
build(deps): bump github.com/Azure/azure-sdk-for-go/sdk/storage/azblob

Bumps [github.com/Azure/azure-sdk-for-go/sdk/storage/azblob](https://github.com/Azure/azure-sdk-for-go) from 1.6.4 to 1.7.0.
- [Release notes](https://github.com/Azure/azure-sdk-for-go/releases)
- [Commits](https://github.com/Azure/azure-sdk-for-go/compare/sdk/storage/azblob/v1.6.4...sdk/azcore/v1.7.0)

---
updated-dependencies:
- dependency-name: github.com/Azure/azure-sdk-for-go/sdk/storage/azblob
  dependency-version: 1.7.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-05-18 19:28:24 -07:00
dependabot[bot]GitHubdependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
f3d6633aac build(deps): bump github.com/aws/aws-sdk-go-v2/service/s3 from 1.99.0 to 1.101.0 (#9553)
build(deps): bump github.com/aws/aws-sdk-go-v2/service/s3

Bumps [github.com/aws/aws-sdk-go-v2/service/s3](https://github.com/aws/aws-sdk-go-v2) from 1.99.0 to 1.101.0.
- [Release notes](https://github.com/aws/aws-sdk-go-v2/releases)
- [Commits](https://github.com/aws/aws-sdk-go-v2/compare/service/s3/v1.99.0...service/s3/v1.101.0)

---
updated-dependencies:
- dependency-name: github.com/aws/aws-sdk-go-v2/service/s3
  dependency-version: 1.101.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-05-18 19:28:06 -07:00
Chris LuandGitHub 68794fb94c fix(ec_distribute): remove partial files on copy stream error (#9543)
* fix(ec_distribute): remove partial files on copy stream error

writeToFile opens the destination with O_TRUNC and streams into it. On
a mid-stream receive / write / cancellation error it returned the
failure but left the destination behind in whatever state had been
written so far — typically 0 bytes when the source errored before
sending any FileContent. VolumeEcShardsCopy distributes .ecx by
calling doCopyFile, so this same stub-leaving behaviour produced the
0-byte .ecx files seen on EC encoding failures: the source claims a
non-zero ModifiedTsNs (so the existing "source not found" cleanup
doesn't fire), the stream then errors immediately, and the receiver
ends up with a 0-byte .ecx that downstream code mistook for a valid
empty index.

Clean up the partial file on every error path that returns from the
streaming loop (receive, write, and cancellation). Skip cleanup when
isAppend=true so resumable appends keep their existing content. As
defense in depth, VolumeEcShardsCopy also stats the .ecx after copy
and removes / errors on a 0-byte result so the orchestrator can pick
a different source.

The Rust volume server has only the source side of CopyFile (no
client-side stream-to-disk consumer) and no .ecx subsystem yet, so
this fix has no Rust mirror.

* fix(ec_distribute): close file before remove, fail fast on stat error

Address review feedback:

- writeToFile's mid-stream removeIncomplete called os.Remove while the
  destination file handle was still open. On Windows os.Remove fails
  while a handle is open, so the cleanup wouldn't run there. Wrap the
  handle close in a once-only helper, call it from removeIncomplete
  and from the existing "source not found" cleanup, and keep a deferred
  close as the safety net for the normal-return path.
- VolumeEcShardsCopy's post-copy .ecx check silently passed when
  os.Stat returned an error: doCopyFile had reported success but if
  the file was already gone, unreadable, or somehow a directory, the
  orchestrator only learned at mount time with no useful context.
  Treat any non-nil stat error and any directory result as a copy
  failure here and surface it immediately.
2026-05-18 15:19:51 -07:00
Chris LuandGitHub af8d4e00ee fix(ec_mount): reject 0-byte .ecx and aggregate cross-disk failures (#9542)
* fix(ec_mount): reject 0-byte .ecx and aggregate cross-disk failures

MountEcShards's per-disk loop bailed on the first disk returning a
non-ENOENT error, and NewEcVolume wrapped its ENOENT with %v so the
caller's `err == os.ErrNotExist` check never matched. On a multi-disk
volume server where ec.balance / ec.rebuild had distributed shards
across sibling disks while the matching .ecx never arrived, the mount
loop bailed after disk 0 with "cannot open ec volume index" and the
operator never saw that the rest of the disks were also empty. The
companion failure mode is a 0-byte .ecx stub left by EC distribute's
writeToFile after a mid-stream copy failure: Stat() succeeds, treating
the stub as a valid index, and downstream mount work proceeds against
an empty file.

Wrap the ec-volume open errors with %w, treat a 0-byte .ecx as
os.ErrNotExist (in NewEcVolume, findEcxIdxDirForVolume, and
HasEcxFileOnDisk), and have MountEcShards collect per-disk failures
before returning a single aggregated error. The "no .ecx anywhere"
case gets a distinct error so the orchestrator can re-copy the index
from a healthy replica rather than retry against the same broken
state.

* fix(ec_reconcile): indexEcxOwners also rejects 0-byte .ecx stubs

findEcxIdxDirForVolume already skipped 0-byte .ecx during MountEcShards,
but indexEcxOwners (used by reconcileEcShardsAcrossDisks at startup)
still recorded the first .ecx by name only. On a store where one disk
holds a 0-byte stub left by a failed EC distribute and a sibling disk
holds the real index, the stub would win the owner selection — and
NewEcVolume's new size check would then refuse to load against it,
leaving the orphan shards unloaded even though a valid index exists.

Mirror the size check from findEcxIdxDirForVolume: skip directory
entries whose .ecx Info() reports size 0 or whose Info() call fails.

* fix(ec_mount): accept 0-byte .ecx as valid empty index

The previous commit treated a 0-byte .ecx in NewEcVolume as
os.ErrNotExist, on the assumption that any empty .ecx was a stub left
by a failed copy stream. That broke the legitimate empty-volume case:
when an EC volume's source .idx has no live entries (e.g. all needles
deleted before WriteSortedFileFromIdx), the sorted .ecx is genuinely
0 bytes and must mount. The integration test
TestEcShardsToVolumeMissingShardAndNoLiveEntries fails with
"MountEcShards: no .ecx index found on any local disk" because the
mount path now refuses the legitimate empty index.

A 0-byte .ecx left by a failed copy stream is indistinguishable from
the legitimate empty case by file size alone. Preventing stub files
from being written is the receiver-side cleanup in writeToFile's job
(the companion EC distribute PR), not NewEcVolume's at mount time.

The cross-disk lookup helpers (findEcxIdxDirForVolume, HasEcxFileOnDisk,
indexEcxOwners) keep their size > 0 preference: when a real .ecx
exists on a sibling disk alongside a stub, we still want to route
mounts and reconcile at the real one. If no non-zero .ecx exists
anywhere, the per-disk fallback in MountEcShards can still open the
0-byte .ecx and the volume mounts.

Replace TestMountEcShards_ZeroByteEcxOnlyDisk with
TestMountEcShards_EmptyEcxMountsSuccessfully, which pins the
empty-volume invariant.
2026-05-18 15:00:33 -07:00
Chris LuandGitHub 41b6ad002b fix(volume.list): show one entry per physical disk on multi-disk nodes (#9541)
* fix(volume.list): show one entry per physical disk on multi-disk nodes

DataNodeInfo.DiskInfos is keyed by disk type, so several same-type
physical disks on one node collapse to a single map entry at the master.
volume.list iterated that map directly and reported one "Disk hdd ...
id:0" line per node, hiding the per-disk volume and shard layout. EC
operators on multi-disk volume servers had no way to verify which
physical disk a shard landed on.

Lift the per-physical-disk split into a DiskInfo.SplitByPhysicalDisk()
method on the proto type so consumers outside admin/topology can use
it. Apply it in writeDataNodeInfo so the verbose Disk block shows one
entry per physical disk, ordered by DiskId. Capacity counters are
split evenly across reconstructed disks since the wire format doesn't
carry per-disk capacity yet.

This is a display-only change. ActiveTopology already did the split on
its own and is now updated to call the shared helper.

* fix(volume.list): preserve totals, count active/remote exactly, dedupe header

Address review feedback on the per-physical-disk split:

- share() truncated remainders so reconstructed per-disk counters could
  sum to less than the original aggregate (10 / 3 = 3+3+3). Distribute
  the remainder to the lowest disk ids so MaxVolumeCount and
  FreeVolumeCount sum exactly back to the node totals.
- ActiveVolumeCount and RemoteVolumeCount are derivable per disk from
  the VolumeInfos already grouped by DiskId, so count them exactly
  (ReadOnly=false and RemoteStorageName!="" respectively) instead of
  approximating with an even split.
- writeDataNodeInfo's per-disk callback fired the DataNode header on
  every iteration after the split, so a node with 6 physical disks
  emitted 6 DataNode headers. Guard the callback with headerPrinted so
  the header still appears at most once per node.
- Sort split disks deterministically using explicit DiskId comparison
  to avoid int overflow risk on 32-bit systems.
- Tighten the volume.list test substring to "id:N\n" so unrelated
  tokens like "ec volume id:101" don't accidentally match the id:1
  needle, and assert the rack callback fires once.
2026-05-18 14:43:44 -07:00
Chris LuandGitHub a761441926 fix(test): reserve mini ports on all interfaces; bound risingwave cleanup shell (#9545)
The 127.0.0.1-only reservation in AllocateMiniPorts/AllocatePortSet let
another process hold the gRPC port on a different interface, so weed
mini's isPortAvailable check failed and it shifted master.grpc. weed
shell -master=<HTTP> still derives grpc as HTTP+10000 and dialed the
unused port, hanging until the 30s context deadline killed it. Bind the
reservation listeners on :port to match mini's check.

Also bound listFilerContents in catalog_risingwave with a 30s
exec.CommandContext so a hung weed shell during failure-cleanup can't
burn the 20-minute test budget.
2026-05-18 14:16:22 -07:00
Chris LuandGitHub 37e6263efe fix(shell): attach admin JWT for filer IAM gRPC calls (#9536)
When jwt.filer_signing.key is set, the filer's IamGrpcServer requires
a Bearer token on every IAM RPC. The shell's s3.* IAM commands dialed
without that header and failed with Unauthenticated. Route them through
a small helper that mints a token from the same key viper-loaded from
security.toml and appends it as outgoing metadata, matching the credential
grpc_store pattern.
2026-05-18 13:42:32 -07:00
Chris LuandGitHub 3d872a1416 fix(filer): load -s3.config static identities into the filer's CredentialManager (#9537)
When weed filer started its embedded S3 gateway with -s3 -s3.config, only
the S3 server loaded the s3.json static identities — the filer's own
CredentialManager stayed empty, so the IAM gRPC service backing the admin
UI and weed shell returned only dynamic users. Mirror the wiring weed
server already does and hand the same config path to the filer.
2026-05-18 13:41:30 -07:00
Chris LuandGitHub 4d04609bb8 fix(mount): don't release file handles from FUSE Forget (#9529)
fix(mount): don't release file handles from Forget

Forget(nodeid, nlookup) only decrements the kernel inode lookup count.
File handle lifecycle belongs to FUSE Open/Release. Driving the FH
refcount from Forget coupled two unrelated counters and could tear down
a still-live handle if Forget ever raced ahead of Release.

Drop the ReleaseByInode call (and the now-unused method).
2026-05-18 01:02:58 -07:00
Chris Lu 01b3e4a71c template 2026-05-17 23:12:04 -07:00
Chris LuandGitHub 6cab199400 fix(iceberg): dial filer gRPC address verbatim in plugin worker (#9527)
* fix(iceberg): dial filer gRPC address verbatim in plugin worker

dialFiler was running its address argument through pb.ServerAddress.ToGrpcAddress,
whose single-port fallback adds +10000 to any host:port — so when the admin
forwards ClusterContext.FilerGrpcAddresses (already host:grpcPort) to the worker,
the iceberg handler turns the real gRPC port (e.g. 18888) into a non-existent
28888 and dispatched jobs fail with connection refused.

Drop the conversion; the address is already dialable. Tests that produced fake
filer addresses in dual-port form now return host:grpcPort to match the new
contract.

* test(ec): use renamed detection_interval_minutes field

The admin_runtime.detection_interval_seconds field was renamed to
detection_interval_minutes back in May. This integration test was not
updated, so the unknown JSON field was silently ignored and the scheduler
fell back to the default detection interval (17 min for erasure_coding),
which exceeds the test's 5-minute wait and times out.

Switch to detection_interval_minutes: 1 — local run completes in ~120s.
2026-05-17 23:03:00 -07:00
Chris Lu 136eb1b7c8 4.26 2026-05-17 21:05:25 -07:00
Chris LuandGitHub c11ff6657b fix(ec): mirror EC sidecars onto every shard-bearing disk at startup (#9525)
* fix(ec): mirror EC sidecars onto every shard-bearing disk at startup

In a multi-disk volume server, ec.balance and ec.rebuild can land shards
on a disk that does not also hold the matching .ecx / .ecj / .vif index
files. The orphan-shard reconciler in reconcileEcShardsAcrossDisks
already loads those shards by pointing the EcVolume at the sibling
disk's index files; reads work, but any failure on the index-owning
disk silently disables every shard on the other disk, even though those
shards are physically fine.

This change adds mirrorEcMetadataToShardDisks, a startup pass that
physically replicates .ecx / .ecj / .vif onto each disk that holds
shards but is missing them. Each copy is atomic (tmp + fsync + rename)
and idempotent (a destination that already has the sidecar is
preserved). After mirroring, the cross-disk reconciler prefers the
local IdxDirectory so the EcVolume mounts self-contained; the
cross-disk virtual mount remains as a fallback for volumes whose mirror
failed (read-only target, out of space, partial copy on a previous
boot).

The same-disk invariant the EC lifecycle (encode / decode / balance /
vacuum / repair) was already documented as promising is now actually
restored at boot, so a future failure of one disk in a split-shards
layout no longer takes the other disk's shards with it.

Tests cover the orphan-layout mirror (dir0 receives the .ecx / .ecj /
.vif from dir1) and idempotency (an existing destination .ecx is not
overwritten with the owner's copy).

* fix(ec): handle legacy pre-dir.idx sidecar layout in mirror skip-check

hasAllEcSidecarsLocally checked only the modern destination path
(IdxDirectory for .ecx/.ecj, Directory for .vif). A destination disk
that still had a legacy .ecx in its data dir (written before -dir.idx
was set) would report "not present" and the mirror would write a
second copy to IdxDirectory, leaving two .ecx files on disk.

Matches HasEcxFileOnDisk's open-with-fallback contract: check the
modern path first, then the opposite directory. Factored the
exists-and-not-a-dir check into a small statRegular helper so the
fallback ladder stays readable.

* rust(seaweed-volume): mirror EC sidecars onto shard-bearing disks at startup

Port of the Go fix (commit 088e26ea6) to the Rust volume server.
Adds Store::mirror_ec_metadata_to_shard_disks, called from
add_location / load_new_volumes before the cross-disk orphan
reconciler. Physically copies .ecx / .ecj / .vif from the disk that
owns the index files onto every disk holding shards but missing
sidecars, so each shard-bearing disk ends up self-contained.

The reconciler now prefers the local idx_directory when the mirror
has installed a .ecx there; the cross-disk virtual mount remains as
the fallback for volumes whose mirror failed (read-only target, out
of space, partial copy on a previous boot). Adds ec_local_ecx_path
helper shared between reconcile and mirror to detect the post-mirror
fast path.

Mirrors the Go-side fallback in hasAllEcSidecarsLocally: when
-dir.idx is configured and the destination still has a legacy .ecx
in its data dir, that's recognized so the mirror does not write a
duplicate copy into idx_directory.

Tests cover the two key cases: orphan layout (dir0 receives the
sidecars from dir1) and idempotency (a pre-existing destination .ecx
is not overwritten).

* trim verbose comments on EC mirror code

Comments now lead with the WHY (non-obvious constraints, the
post-mirror fast path, why local copies are authoritative) and drop
restate-the-code blocks, headers, and section dividers. Behavior is
unchanged; all existing tests still pass on both the Go volume
server and the seaweed-volume Rust port.

* drop github issue refs from added comments

Two stray "#9212" references slipped into comments I added on the
cross-disk reconciler call site. The git log carries the issue
history; comments stand on their own.

* test(ec): accept rebuild on either disk after sidecar mirror

TestEcLifecycleAcrossMultipleDisks asserted the rebuilt shard 9 must
land at the disk-0 path. With the boot-time sidecar mirror, every
shard-bearing disk owns its own .ecx, so VolumeEcShardsRebuild now
picks whichever disk hosts the most shards — disk 1 in this layout
after the deletion. The shard can legitimately rebuild on either
disk; the test now accepts both and uses the chosen path for the
subsequent mount + read verification.
2026-05-17 19:55:15 -07:00
Chris LuandGitHub 6b94701213 mini: quieter startup with a docker-compose-style progress board (#9524)
* mini: quieter startup with a docker-compose-style progress board

Replaces noisy startup/shutdown logs with a single in-place progress
table on a TTY (or one line per state change off-TTY). Each component
renders as `pending -> starting -> ready` during startup and
`stopping -> stopped` during shutdown, with elapsed time on transition.

Also folds in a few cleanups uncovered while making this readable:

- route the admin.go startup prints through glog so quietMiniLogs()
  filters them under mini but standalone weed admin still shows them
- generate a dev SSE-S3 KEK + passphrase on first run via WEED_S3_SSE_KEK
  and WEED_S3_SSE_KEK_PASSPHRASE env vars (viper.Set has a nested-key
  conflict between s3.sse.kek and s3.sse.kek.passphrase); persisted under
  the data folder so restarts reuse the same key
- demote worker/master gRPC Recv 'context canceled' to V(1); those are
  the normal shutdown signal, not Errors/Warnings
- drop the 'Optimized Settings' block and the 'credentials loaded from
  environment variables' message from the welcome banner
- only show the credentials setup hints when no S3 identities exist
  (new s3api.HasAnyIdentity accessor backed by an atomic.Bool)
- use S3_BUCKET in the credentials hint so it pairs with
  AWS_ACCESS_KEY_ID / AWS_SECRET_ACCESS_KEY
- reorder running-services list to master / volume / filer / webdav /
  s3 / iceberg / admin

* mini: refuse in-memory-only SSE-S3 dev keys; surface admin serve errors

loadOrCreateMiniHexSecret returns "" when os.WriteFile fails, so SSE-S3
won't encrypt data under a KEK that the next restart can't reproduce
(which would orphan whatever was written this run). The caller already
treats "" as "skip setting WEED_S3_SSE_* env vars", so SSE-S3 and IAM
just stay disabled for this run.

startAdminServer's serve goroutine used to only log ListenAndServe
failures, so a bind error left the caller blocked on ctx.Done() with
no listener. Forward the error through a buffered channel and select
on it alongside ctx.Done().

* ci(s3-proxy-signature): match weed mini's new progress-board ready line

The readiness probe grepped for "S3 (gateway|service).*(started|ready)",
which matched weed mini's old "S3 service is ready at ..." line. Mini
now emits "  S3           ready (Xs)" from its progress board, so the
old pattern misses and the test timed out at the 30-second wait.

Widen the alternation to also accept "S3\s+ready". The curl HEAD
fallback already covers any remaining cases.
2026-05-17 19:13:09 -07:00
Chris LuandGitHub ff6f9fd90a iam: honor configured credential store for IAM API policies and propagate to S3 caches (fixes #9518) (#9522)
* iamapi: route managed policies through credential manager (fixes #9518)

CreatePolicy via the IAM API wrote straight to the filer
/etc/iam/policies.json, ignoring any non-filer credential store. When
credential.postgres was configured, policies created via the IAM API
landed only in the filer while the Admin UI wrote to postgres,
producing a split-brain where ListPolicies/GetPolicy never saw the
Admin UI's policies and vice versa.

GetPolicies/PutPolicies on IamS3ApiConfigure now load managed policies
from credentialManager and persist Create/Update/Delete as a delta
against the store. Inline user/group policies still live in the legacy
policies.json file (no credential-store API for them yet). Pre-existing
managed policies in the legacy file are merged on read so deployments
don't lose data, and re-persisted to the store on the next write so
the legacy file is drained over time.

* credential: route IAM API inline policies through credential manager

Extends the #9518 fix to user-inline and group-inline policies so the
IAM API never writes the legacy /etc/iam/policies.json bundle directly.
The previous patch only routed managed policies; this one finishes the
job for the other two policy types.

- Add GroupInlinePolicyStore + GroupInlinePoliciesLoader optional
  interfaces, mirroring the existing user-inline ones, and matching
  Put/Get/Delete/List/LoadAll wrappers on CredentialManager.
- Implement group-inline storage in memory (new map), filer_etc (new
  field on PoliciesCollection, reusing the legacy file under policyMu),
  and postgres (new group_inline_policies table with ON DELETE CASCADE
  off the groups FK).
- Wire the new methods through PropagatingCredentialStore so wrapped
  stores still delegate correctly.
- IamS3ApiConfigure.PutPolicies now applies managed + user-inline +
  group-inline as deltas through the credential manager; the legacy
  /etc/iam/policies.json file is never written when a credential
  manager is wired up. GetPolicies still reads the legacy bundle once
  as a fallback so unmigrated data is picked up and re-persisted into
  the store on the next write.

* credential: propagate SaveConfiguration writes to running S3 caches

Postgres (and any non-filer) credential stores never fired the S3 IAM
cache invalidation path on bulk identity / group updates. The
PropagatingCredentialStore had explicit Put/Remove handlers for
single-entity calls (CreateUser, PutPolicy, etc.) but inherited
SaveConfiguration unchanged from the embedded store, so the bulk path
the IAM API takes at the end of every handler was silent. Inline-policy
changes recompute identity.Actions and persist via SaveConfiguration,
so until restart the cached Actions on each S3 server stayed stale and
authorization decisions used the pre-change view.

Override SaveConfiguration to snapshot the prior user / group lists,
delegate the save, then fan out PutIdentity / PutGroup for what's in
the new config and RemoveIdentity / RemoveGroup for what got pruned.
Reuses the existing SeaweedS3IamCache RPCs, no protobuf changes.

* iamapi: drain legacy policies.json after authoritative credential-store writes

Review pointed out a resurrection bug: GetPolicies still reads
/etc/iam/policies.json as a one-way migration fallback, but PutPolicies
in the credential-manager path never wrote that file, so legacy-only
entries reappeared on the next read even after the IAM API "deleted"
them. PutPolicies now overwrites the bundle with an empty {} after a
successful credential-store write, unless the store is filer_etc
(which owns the bundle as its own inline-policy backing — clearing it
would wipe filer_etc's data). Also wraps the filer read, JSON
unmarshal, and marshal errors with context per the other review
comments.
2026-05-17 13:15:27 -07:00
Chris LuandGitHub b4289abb0a admin: convert filer address to gRPC form before dispatch (#9523)
The master returns each registered filer in pb.ServerAddress dual-port
form (host:httpPort.grpcPort, e.g. 10.0.0.1:8888.18888). The admin's
plugin context builder forwarded that string verbatim as
filer_grpc_address, so workers calling grpc.DialContext on it failed
every job in ~3ms with "dial tcp: lookup tcp/8888.18888: unknown port".

Run each entry through pb.ServerAddress.ToGrpcAddress before populating
ClusterContext.FilerGrpcAddresses.

The lifecycle integration test now pins filer.port.grpc to a value that
breaks the FILER_PORT+10000 assumption, and a new dispatch test drives
the admin's /api/plugin/job-types/s3_lifecycle/run path end-to-end and
asserts the dispatched job both reaches the filer and deletes the
backdated object.
2026-05-17 11:33:54 -07:00
Chris LuandGitHub 2a41e76101 fix(ec): blanket-clean every destination over the full shard range (#9512)
* fix(ec): blanket-clean every destination over the full shard range

The previous cleanup pass walked t.sources only, with the shard ids the
topology had reported at detection time. In the wild, a destination can
end up with EC shards mounted that the topology snapshot didn't list —
shards on a sibling disk that hadn't heartbeated, or shards left over
from a concurrent attempt's mount step. FindEcVolume still returns
true, so the next ReceiveFile trips the mounted-volume guard.

Cleanup now unions t.sources (with ShardIds) and t.targets and issues
unmount + delete over [0..totalShards-1] on each. Both RPCs are
idempotent on missing shards, so the wider sweep is free.

Two new tests cover the gap: shards mounted beyond what t.sources
lists, and a target-only destination with no source row.

* log(ec): include disk_id in EC unmount/delete/refusal log lines

The current logs identify the volume and shard but leave disk_id off,
which makes the cross-server cleanup story hard to follow when
multiple disks of one server hold pieces of the same volume:

  UnmountEcShards 4121.1                              -> add disk_id
  ec volume video-recordings_4121 shard delete [1 5]  -> add per-loc disk_id
  volume server X:Y deletes ec shards from 4121 [...] -> add disk_id
  ReceiveFile: ec volume 4121 is mounted; refusing... -> add disk_ids

ReceiveFile's refusal now names the disk_ids actually holding the
mount so operators can see whether the next cleanup pass needs to
target a sibling disk. Added Store.FindEcVolumeDiskIds /
Store::find_ec_volume_disk_ids as the supporting primitive.

Mirrored in seaweed-volume/src/ (unmount log in Store::unmount_ec_shard,
heartbeat delete log in diff_ec_shard_delta_messages, refusal in the
ReceiveFile handler).

* test(ec): stub VolumeEcShardsUnmount/Delete on the fake volume server

The plugin-worker EC tests boot a fake volume server that embeds
UnimplementedVolumeServerServer. After the worker started calling
VolumeEcShardsUnmount + VolumeEcShardsDelete pre-distribute, the
default Unimplemented response surfaced as fourteen "method not
implemented" errors and TestErasureCodingExecutionEncodesShards
failed. Both RPCs are no-ops here — nothing on the fake server has
mounted state or persisted shard files to remove.
2026-05-17 11:31:37 -07:00
Chris LuandGitHub bf9110ebd3 fix(ec): mount falls back to sibling-disk .ecx (fixes #9519) (#9521)
* fix(ec): mount falls back to sibling-disk .ecx (fixes #9519)

MountEcShards iterated DiskLocations and on each disk called LoadEcShard
with that disk's IdxDirectory as the .ecx home. When ec.balance lands the
.ec?? shard on disk A but the .ecx on sibling disk B of the same volume
server, NewEcVolume ENOENTs the .ecx and returns "cannot open ec volume
index ...". That error is not os.ErrNotExist, so the per-disk continue
branch did not engage and the mount loop bailed before trying any other
disk.

The startup reconciliation in reconcileEcShardsAcrossDisks already
handles this layout for orphan shards discovered on boot (issue #9212).
This change mirrors the same primitive on the mount path: look up the
.ecx owner across all DiskLocations once and route NewEcVolume at that
directory whenever the disk being mounted does not own its own copy of
the .ecx. Same-disk mounts are unaffected because HasEcxFileOnDisk keeps
LocalIdxDirectory in play.

Adds a regression test that plants the index files on a sibling disk
AFTER NewStore returns (so the startup reconcile is a no-op for that
vid) and verifies MountEcShards succeeds; also pins the same-disk
baseline against accidental re-routing.

* fix(ec): skip redundant stats in cross-disk .ecx lookup (review)

Two follow-ups from gemini-code-assist on #9521:

1. MountEcShards: when findEcxIdxDirForVolume already returned a path
   that lives on this disk's IdxDirectory or Directory, the disk owns
   the .ecx — skip the HasEcxFileOnDisk stat and use the local idx dir
   directly. Only re-check when the disk's directories are neither, so
   the duplicate-.ecx-on-multiple-disks edge case is still honored.

2. findEcxIdxDirForVolume: hoist the seen map across the location loop
   so a shared IdxDirectory (one -dir.idx paired with several -dir
   entries) is only stat'd once per call.

Both are I/O optimizations; behavior is unchanged. Existing
cross-disk and same-disk regression tests still pass.

* docs(ec): drop issue/PR references from cross-disk mount comments

Comments and test docstrings stand on their own; the issue number
adds nothing a reader can act on and goes stale across forks. Keep
the description of *what* the layout is and *why* the fallback
exists, just without the reference.
2026-05-17 11:10:37 -07:00
Chris LuandGitHub d51454adf4 rust(seaweed-volume): distributed EC read across peer servers (#9516)
* feat(seaweed-volume): distributed EC read across peer servers

EcVolume::read_ec_shard_needle previously errored with NotFound when
any interval's shard wasn't local. In an RS(10,4)-across-N deployment
each server holds one shard, so every read needed >=9 peer fetches and
post-EC GETs returned 404 on volumes whose shards lived on more than
one server.

Mirror of weed/storage/store_ec.go's readOneEcShardInterval ->
readRemoteEcShardInterval -> recoverOneRemoteEcShardInterval chain:

  * server/store_ec.rs (new): entry point
    read_ec_shard_needle_distributed. Snapshots locate-needle + local
    reads under the Store sync lock, drops the lock, then async-fetches
    missing intervals via the peer's VolumeEcShardRead RPC. Falls back
    to Reed-Solomon reconstruction (read every other shard at the same
    (shard_offset, size) and run rs.reconstruct) when the direct peer
    read fails. Refreshes the per-EcVolume shard_locations cache from
    the master's LookupEcVolume RPC using Go's freshness thresholds
    (11s / 7min / 37min).
  * erasure_coding/ec_volume.rs: shard_locations now sits behind a
    std::sync::RwLock so the read path can refresh the map without
    holding the Store write lock. Adds shard_locations_refresh_time
    (Mutex<Option<Instant>>) for the staleness heuristic. Mirrors Go's
    ShardLocationsLock / ShardLocationsRefreshTime fields. set/get
    helpers updated for interior mutability.
  * server/handlers.rs: GET handler now tries the local-only fast
    path first, then falls through to the distributed path on
    NotFound.

* review: address PR 9516 feedback on distributed EC read

Five of the six PR-review comments addressed; the sixth (JWT on
outgoing peer gRPC) is deferred with an explicit TODO because the
crate-wide outgoing-JWT signing surface doesn't exist yet — adding it
in this one call site would split the credential plumbing across
peer paths that already lack it (copy_file_from_source, batch_delete,
…). Revisit when an outgoing-JWT helper lands.

Fixed in this commit:

  * Handlers: drop the two-tier (local-first, then distributed)
    read in handlers.rs. read_ec_shard_needle_distributed already
    does the local-first pass under the same store read lock; the
    redundant outer attempt re-read local intervals twice for any
    needle that spanned mixed-locality shards.
  * Scanner snapshot: replace inline locate-needle math with
    `ecv.locate_needle(needle_id)`. Same routine the local-only
    read path uses, so byte-identical on shard-size + interval
    boundaries.
  * EcVolume::set_shard_locations also advances
    shard_locations_refresh_time so the staleness check honors
    callers that populate the cache directly without going through
    the master LookupEcVolume RPC.
  * parse_grpc_address moved from grpc_server.rs into
    grpc_client.rs as `pub` and is reused by both grpc_server.rs
    and the new store_ec module. Single source of truth for the
    HTTP↔gRPC port-offset convention.
  * Reconstruction (recover_one_remote_ec_shard_interval) now seeds
    bufs from locally-mounted survivor shards BEFORE the remote
    fan-out. Previously the fan-out was remote-only, so when the
    shard_locations cache was cold or the master lookup failed,
    reconstruction errored even though enough siblings were on
    local disk to recover the missing interval.

* review: tighten parse_grpc_address; atomic shard-locations cache swap

Two follow-up findings from the PR 9516 review round 2:

  * `parse_grpc_address` now validates BOTH port components in the
    dotted form (`host:port.grpcPort`) — previously a non-numeric
    HTTP port like `host:abc.18080` slipped through and tripped a
    less-useful downstream URI parse error. The implicit form
    (`host:port` → port + 10000) also gains an overflow check so
    inputs like `host:60000` (which silently wrap past u16) are
    rejected here instead of producing an opaque connection
    failure later. Six unit tests cover each rejection path.

  * `EcVolume::set_shard_locations` no longer bumps the per-volume
    refresh timestamp. The previous fix introduced a freshness
    race: a multi-shard population that inserts shard-by-shard
    would flip `needs_refresh == false` on the first write, letting
    a concurrent reader observe a half-populated map already
    marked "fresh" and return NotFound for the not-yet-inserted
    shards. Added `EcVolume::replace_shard_locations(map)` for the
    atomic bulk swap; `write_back_shard_locations` in the
    distributed-read path uses it so the cache transitions
    old → fresh in a single observable step.
2026-05-16 20:44:28 -07:00
Chris LuandGitHub f892b445b3 helm(admin): support secretExtraEnvironmentVars (refs #9511) (#9513)
* helm(admin): support secretExtraEnvironmentVars

The admin statefulset only honored extraEnvironmentVars, forcing the
OIDC client secret (and any other sensitive WEED_* value) to be inlined
as plain text in values.yaml — not GitOps-friendly. The filer chart has
had secretExtraEnvironmentVars for this exact case; mirror that pattern
on admin so secrets can be projected via valueFrom.secretKeyRef.

Surfaced by an enterprise OIDC deployment (issue #9511) where the only
workaround was hardcoding WEED_ADMIN_OIDC_CLIENT_SECRET in values.yaml.

* helm(admin): sort secretExtraEnvironmentVars keys for stable output

Helm/Go template map iteration is non-deterministic, so the env entries
could shuffle between renders and trigger spurious StatefulSet rollouts
in GitOps tooling (ArgoCD/Flux). Sort the keys with sortAlpha, mirroring
the extraEnvironmentVars block immediately above.

Flagged by gemini-code-assist and coderabbitai on PR #9513.
2026-05-15 13:19:05 -07:00
Chris LuandGitHub 62821964dd filer/iam-grpc: make admin Bearer auth opt-in (fixes #9509) (#9514)
PR #9442 made the filer refuse to register the IAM gRPC service unless
jwt.filer_signing.key was set in security.toml, which broke the admin
UI Users/Groups/Policies pages for every deployment that ships without
a security.toml — weed mini, plain Helm, vanilla weed filer. The Users
tab returns Unimplemented and the page is unusable. Issues #9504,
#9505 and #9509 all trace to this gap.

The rest of the filer's gRPC surface is unauthenticated by default;
treat IAM the same way. The service now always registers, and the
auth gate is a no-op when no signing key is configured. When the key
is set, every RPC still requires an admin-signed Bearer token, matching
the post-#9442 behaviour. Operators who expose the filer gRPC port
beyond a trusted network should set the key on both filer and admin.

The admin client (IamGrpcStore.withIamClient) already skips attaching
the authorization metadata when its key is empty, so no changes there.
2026-05-15 13:15:20 -07:00
Konstantin LebedevandGitHub 7d1b16fbcd fix: ListBucketsHandler for pathStyleDomains (#9510) 2026-05-15 13:12:55 -07:00
Chris LuandGitHub 2ed95d7ea9 helm: decouple JWT signing from cert-manager mTLS (fixes #9506) (#9508)
* helm(security): decouple JWT signing from cert-manager mTLS

The filer needs jwt.filer_signing.key to register the IAM gRPC service the
Admin UI Users tab calls (PR #9442). The chart only rendered security.toml
under enableSecurity, which also pulls in cert-manager for mTLS — much heavier
than the Admin UI needs. Operators on Helm without cert-manager have no way
to flip the JWT key on, so the Users tab fails with Unimplemented after
upgrading past 4.24.

Introduce seaweedfs.securityConfigEnabled, true when enableSecurity OR any
explicit jwtSigning toggle (volumeRead/filerWrite/filerRead) is set. The
configmap renders under that helper; the [grpc.*]/[https.*] sections inside
stay gated on enableSecurity. Each pod template splits the security-config
mount onto the helper and keeps the cert volume mounts on enableSecurity.

volumeWrite is intentionally excluded from the helper trigger because it
defaults to true; including it would silently start mounting security.toml on
every fresh install. With this change, enableSecurity=false + defaults
renders nothing (unchanged), enableSecurity=true renders the full toml
(unchanged), and enableSecurity=false + filerWrite=true renders just the
[jwt.*] sections so the Admin UI works without mTLS.

Fixes #9506.

* helm(security): trim verbose comments

* helm(security): handle null securityConfig in helper

Address review feedback: (.Values.global.seaweedfs.securityConfig).jwtSigning
errored if a user explicitly set securityConfig: null in their values. Drop
into intermediate $sec/$jwt with default dict at each step so a missing or
nulled-out parent is tolerated.

* helm(ci): cover IAM gRPC decoupling (issue #9506)

Five regression assertions exercised against the rendered chart so a
future change cannot silently re-couple jwt.filer_signing to mTLS:

1. defaults render no security-config ConfigMap (preserves baseline)
2. filerWrite=true alone renders [jwt.filer_signing] with no [grpc.*]
3. filerWrite=true mounts security-config on filer + admin without
   pulling in cert volumes — the actual fix for the Admin UI Users tab
4. enableSecurity=true still produces the full toml with [grpc.master]
5. securityConfig=null and securityConfig.jwtSigning=null both render
   cleanly (gemini-code-assist review nit, applied chart-wide)

Patch a pre-existing direct-access in filer-statefulset.yaml that
crashed on securityConfig=null, surfaced by the new null assertion.

* helm(ci): drop issue numbers from comments

* helm(ci): install pyyaml; assert [jwt.signing] in mTLS path

Address coderabbit review:

- The new IAM gRPC test block uses `import yaml` but ran before the
  later `pip install pyyaml -q` step that the security+S3 block
  performs. CI happens to pass because the runner image carries
  PyYAML, but make the dependency explicit so a future runner change
  cannot silently break the regression test.

- The enableSecurity=true assertion only checked for [grpc.master].
  Also assert [jwt.signing] so a refactor that drops the volume-side
  JWT stanza from the mTLS path fails the test instead of slipping
  through.
2026-05-14 23:43:24 -07:00
Chris LuandGitHub bfb2661fec fix(tests): make 32-bit GOARCH tests build and run (#9507)
fix(tests): make 32-bit GOARCH tests build and run (#9503)

verifyTestFilerClient had bare int64 atomic counters after a map header,
so atomic.AddInt64 panicked with "unaligned 64-bit atomic operation" on
linux/386. Switch to atomic.Int64, which the stdlib guarantees is
8-byte aligned on all platforms.

rpc_version_filter_test.go passed the untyped constant 0xdeadbeef to
t.Errorf, where it default-promoted to int and overflowed 32-bit int.
Bind it to a typed uint32 const used in both the comparison and the
error message.
2026-05-14 20:55:37 -07:00
Chris Lu 7acba59a5c 4.25 2026-05-14 12:16:26 -07:00
Chris LuandGitHub 2c1482f7a6 fix(ec): clear cross-server stale EC shards before re-distribute (#9478) (#9499)
* fix(ec): clear cross-server stale EC shards before re-distribute (#9478)

A previous failed encode leaves partial .ec?? shards mounted on
destination volume servers that are not the .dat owner. PR #9480 only
prunes when the .dat sits on a sibling disk of the SAME store, so the
cross-server case stays stuck: every retry trips
volume_grpc_copy.go:570's "ec volume %d is mounted; refusing overwrite"
guard and the scheduler loops.

Detection already lists existing EC shards as CleanupECShards sources;
plumb the shard ids through (ActiveTopology.GetECShardLocations,
TaskSourceSpec, TaskSource.shard_ids) and have the EC worker call
VolumeEcShardsUnmount + VolumeEcShardsDelete on each destination after
the local shard set is generated and before distributeEcShards. Skip
EC-shard sources in getReplicas so the post-encode VolumeDelete step
does not target destination-only nodes.

Integration test mounts a partial shard subset, asserts the
mounted-volume refusal, runs cleanupStaleEcShards, and asserts the
next ReceiveFile lands.

* chore(ec): tighten code comments in stale-shard cleanup

Drop issue-number refs from code comments and shorten the docstrings
on cleanupStaleEcShards / unmountAndDeleteEcShards / getReplicas plus
the new test file. Behavior unchanged.

* fix(ec): skip empty-ShardIds locations; dedupe getReplicas by node

GetECShardLocations dropped entries where ecShardMatchesCollection saw a
phantom info record with EcIndexBits=0 — without ShardIds, getReplicas
misread the resulting source as a regular replica and would have called
VolumeDelete on a destination-only node.

getReplicas now dedupes by Node since VolumeDelete is server-wide;
per-disk source rows on the same server collapse to one call.

* refactor(ec): use MaxShardCount and ShardBits in collectShardIdsForDisk

Drop the literal 32 bit-iteration bound for erasure_coding.MaxShardCount
and treat the EcIndexBits union as a ShardBits so Count() drives the
slice preallocation. Keeps the helper aligned with the rest of the EC
code and survives any future expansion of the shard-count ceiling.
2026-05-14 11:57:45 -07:00
Chris Lu e56a3ee4a2 ci(s3-lifecycle): split into per-test matrix jobs
Each test now runs against a fresh `weed mini`, so per-collection TTL
volume budget no longer leaks across tests and exhausts the pool.
2026-05-14 11:47:24 -07:00
Chris LuandGitHub db2d975b80 ci(docker): tag latest in unified release instead of rebuilding (#9500)
The separate container_latest.yml workflow rebuilt the latest image from
scratch on every tag push (full multi-arch build + QEMU + trivy gate),
which is slow and frequently fails — leaving `latest` stranded on the
prior release (e.g. 4.23 after 4.24 shipped, #9497).

Drop the rebuild. The unified release workflow already publishes the
exact same content as `<tag>` and `<tag>_large_disk`, so just re-tag
those manifests with `crane tag` on both GHCR and Docker Hub once
copy-to-dockerhub completes. Seconds, not hours, and no QEMU.

Move the trivy scan into the unified workflow as report-only: SARIF
still uploads to GitHub Security for visibility, but vuln findings
no longer block the release.

container_latest.yml stays as a workflow_dispatch-only manual fallback.

Refs #9497.
2026-05-14 11:26:28 -07:00
Chris LuandGitHub c47eab1a5d admin: attach admin-signed Bearer token on filer IAM gRPC calls (#9498)
* admin: attach admin-signed Bearer token on filer IAM gRPC calls

PR #9442 added Bearer-JWT enforcement on the filer's IAM gRPC service
but didn't update its only production client, IamGrpcStore. The admin
UI Users/Groups pages went through that client and started failing in
4.24 with either Unimplemented (filer refuses to register the service
when jwt.filer_signing.key is empty) or Unauthenticated (the client
sent no token). Issues #9495 and #9496 both trace to this gap.

Plumb jwt.filer_signing.key into IamGrpcStore via a new SetAdminSigning
hook called from the admin server, and append a freshly minted Bearer
token to outgoing metadata on every call. The mint helper
security.GenJwtForFilerAdmin existed since #9442 but had no production
caller; this wires it up.

Add an integration test alongside grpc_store.go that runs a real
IamGrpcServer over a real grpc.Server listener and exercises the store
end-to-end: matching key succeeds, wrong key returns Unauthenticated,
no key returns Unauthenticated. Without the client-side token attach
the success path fails, so the regression cannot land again.

* address review: include adminSigningExpiresAfterSec in mu comment
2026-05-14 10:51:04 -07:00
313 changed files with 17758 additions and 5799 deletions
+4 -3
View File
@@ -1,9 +1,10 @@
name: "docker: build latest container"
# Manual fallback only. On tag push, container_release_unified.yml already
# re-tags the released versioned image as `latest` / `latest_large_disk`,
# so a full rebuild here is unnecessary. Run this manually if you need to
# rebuild `latest` from an arbitrary ref.
on:
push:
tags:
- '*'
workflow_dispatch:
inputs:
source_ref:
+123 -1
View File
@@ -29,6 +29,7 @@ on:
permissions:
contents: read
security-events: write
env:
RELEASE_TAG: ${{ github.event_name == 'workflow_dispatch' && github.event.inputs.release_tag || github.ref_name }}
@@ -357,7 +358,128 @@ jobs:
ghcr.io/chrislusf/seaweedfs:${{ env.RELEASE_TAG }}${{ matrix.tag_suffix }} \
chrislusf/seaweedfs:${{ env.RELEASE_TAG }}${{ matrix.tag_suffix }}
echo "✓ Successfully copied ${{ matrix.variant }} to Docker Hub"
echo "Successfully copied ${{ matrix.variant }} to Docker Hub"
# Report-only trivy scan: uploads fixable HIGH/CRITICAL findings to GitHub
# Security for visibility, but never blocks the release. Releases (including
# `latest`) ship regardless — vulnerabilities are tracked, not gated, since
# we sometimes need to publish through known findings (e.g. unfixed upstream
# CVE, base-image lag).
trivy-scan:
runs-on: ubuntu-latest
needs: [build]
if: github.event_name == 'push'
continue-on-error: true
strategy:
fail-fast: false
matrix:
include:
- source_suffix: ""
variant: normal
- source_suffix: _large_disk
variant: large_disk
steps:
- name: Login to GHCR
uses: docker/login-action@v4.1.0
with:
registry: ghcr.io
username: ${{ secrets.GHCR_USERNAME }}
password: ${{ secrets.GHCR_TOKEN }}
- name: Trivy report (${{ matrix.variant }})
# Pin to SHA - mutable tags were compromised (GHSA-69fq-xp46-6x23)
uses: aquasecurity/trivy-action@ed142fd0673e97e23eac54620cfb913e5ce36c25 # v0.36.0
with:
scan-type: image
# Scan the multi-arch tag on GHCR (already pushed by the build job).
# Trivy scans the runner's native platform; OS packages are identical
# across architectures since they all share the same alpine base.
image-ref: ghcr.io/chrislusf/seaweedfs:${{ env.RELEASE_TAG }}${{ matrix.source_suffix }}
scanners: vuln
vuln-type: os,library
severity: HIGH,CRITICAL
ignore-unfixed: true
limit-severities-for-sarif: true
format: sarif
output: trivy-results.sarif
exit-code: '0'
- name: Upload Trivy scan results to GitHub Security
if: always()
uses: github/codeql-action/upload-sarif@v4
with:
sarif_file: trivy-results.sarif
category: trivy-${{ matrix.variant }}
# Point `latest` (and `latest_large_disk`) at the just-released versioned
# image. crane tag adds an extra tag to an existing manifest — no rebuild,
# no QEMU, no separate workflow. Replaces the old container_latest.yml
# rebuild that often failed or lagged behind the release. Independent of
# trivy-scan: vuln findings are reported but do not block `latest`.
tag-latest:
runs-on: ubuntu-latest
needs: [copy-to-dockerhub]
if: github.event_name == 'push'
strategy:
matrix:
include:
- source_suffix: ""
latest_tag: latest
- source_suffix: _large_disk
latest_tag: latest_large_disk
steps:
- name: Login to Docker Hub
uses: docker/login-action@v4.1.0
with:
username: ${{ secrets.DOCKER_USERNAME }}
password: ${{ secrets.DOCKER_PASSWORD }}
- name: Login to GHCR
uses: docker/login-action@v4.1.0
with:
registry: ghcr.io
username: ${{ secrets.GHCR_USERNAME }}
password: ${{ secrets.GHCR_TOKEN }}
- name: Install crane
run: |
cd $(mktemp -d)
curl -sL "https://github.com/google/go-containerregistry/releases/latest/download/go-containerregistry_Linux_x86_64.tar.gz" | tar xz
sudo mv crane /usr/local/bin/
crane version
- name: Re-tag ${{ env.RELEASE_TAG }}${{ matrix.source_suffix }} as ${{ matrix.latest_tag }}
run: |
retry_with_backoff() {
local max_attempts=5
local timeout=1
local attempt=1
local exit_code=0
while [ $attempt -le $max_attempts ]; do
if "$@"; then
return 0
else
exit_code=$?
fi
if [ $attempt -lt $max_attempts ]; then
echo "Attempt $attempt failed. Retrying in ${timeout}s..." >&2
sleep $timeout
timeout=$((timeout * 2))
fi
attempt=$((attempt + 1))
done
echo "Command failed after $max_attempts attempts" >&2
return $exit_code
}
SRC_TAG="${{ env.RELEASE_TAG }}${{ matrix.source_suffix }}"
DST_TAG="${{ matrix.latest_tag }}"
echo "Tagging ghcr.io/chrislusf/seaweedfs:${SRC_TAG} as ${DST_TAG}"
retry_with_backoff crane tag "ghcr.io/chrislusf/seaweedfs:${SRC_TAG}" "${DST_TAG}"
echo "Tagging chrislusf/seaweedfs:${SRC_TAG} as ${DST_TAG}"
retry_with_backoff crane tag "chrislusf/seaweedfs:${SRC_TAG}" "${DST_TAG}"
helm-release:
runs-on: ubuntu-latest
-49
View File
@@ -1,49 +0,0 @@
name: EC Integration Tests
on:
push:
branches: [ master ]
paths:
- 'weed/admin/**'
- 'weed/worker/**'
- 'test/erasure_coding/admin_dockertest/**'
- '.github/workflows/ec-integration.yml'
pull_request:
branches: [ master ]
paths:
- 'weed/admin/**'
- 'weed/worker/**'
- 'test/erasure_coding/admin_dockertest/**'
- '.github/workflows/ec-integration.yml'
jobs:
ec-integration-test:
runs-on: ubuntu-latest
timeout-minutes: 15
steps:
- name: Checkout code
uses: actions/checkout@v6
- name: Set up Go
uses: actions/setup-go@v6
with:
go-version-file: 'go.mod'
- name: Build weed binary
run: |
cd weed
go build -o ../weed_bin
- name: Run EC integration tests
run: |
cd test/erasure_coding/admin_dockertest
go test -v -timeout 15m ec_integration_test.go
- name: Upload test logs on failure
if: failure()
uses: actions/upload-artifact@v7
with:
name: ec-test-logs
path: test/erasure_coding/admin_dockertest/tmp/logs/
retention-days: 7
+129
View File
@@ -68,6 +68,135 @@ jobs:
grep -q "security-config" /tmp/security.yaml
echo "✓ Security configuration renders correctly"
echo ""
echo "=== Testing IAM gRPC opt-in path ==="
# Regression test: the filer registers the IAM gRPC service the
# Admin UI Users tab calls only when jwt.filer_signing.key is in
# security.toml. Operators must be able to enable that without
# the cert-manager mTLS bundle.
# Install PyYAML explicitly: this block runs before the later
# security+S3 block that does the same install, and we don't
# want to rely on the runner image shipping it.
pip install pyyaml -q
python3 - "$CHART_DIR" <<'PYEOF'
import subprocess, sys, yaml
chart = sys.argv[1]
def render(values):
args = ["helm", "template", "test", chart]
for k, v in values.items():
args += ["--set", f"{k}={v}"]
return subprocess.check_output(args, text=True)
def docs(manifest):
return [d for d in yaml.safe_load_all(manifest) if d]
def configmap(manifest, name):
for d in docs(manifest):
if d.get("kind") == "ConfigMap" and d["metadata"]["name"] == name:
return d
return None
def workload_mounts(manifest, name):
for d in docs(manifest):
if d.get("kind") not in ("Deployment", "StatefulSet"):
continue
if d["metadata"]["name"] != name:
continue
pod = d["spec"]["template"]["spec"]
vols = {v["name"] for v in pod.get("volumes", [])}
mounts = set()
for c in pod.get("containers", []):
for vm in c.get("volumeMounts", []):
mounts.add(vm["name"])
return vols, mounts
return None, None
failed = []
# Case 1: defaults. The chart historically rendered nothing
# security-related; preserve that so this PR is non-breaking on
# existing installs.
out = render({})
if configmap(out, "test-seaweedfs-security-config") is not None:
failed.append("defaults: security ConfigMap should not render")
else:
print("✓ defaults: no security-config ConfigMap (unchanged)")
# Case 2: filerWrite=true alone is the documented opt-in for
# the Admin UI Users tab. Configmap must render with
# [jwt.filer_signing] and NO [grpc.*] sections (cert paths
# only exist with mTLS).
out = render({
"global.seaweedfs.securityConfig.jwtSigning.filerWrite": "true",
"admin.enabled": "true",
})
cm = configmap(out, "test-seaweedfs-security-config")
if cm is None:
failed.append("filerWrite=true: security ConfigMap missing")
else:
toml = cm["data"]["security.toml"]
if "[jwt.filer_signing]" not in toml:
failed.append("filerWrite=true: security.toml missing [jwt.filer_signing]")
if "[grpc" in toml:
failed.append("filerWrite=true: security.toml unexpectedly has [grpc.*] (would need cert mounts)")
if "[jwt.filer_signing]" in toml and "[grpc" not in toml:
print("✓ filerWrite=true: security.toml has [jwt.filer_signing], no [grpc.*]")
# Case 3: filer + admin pods must MOUNT the security ConfigMap
# under filerWrite=true so the JWT key reaches both processes.
# Cert volumes must NOT be present (no mTLS).
for wl in ("test-seaweedfs-filer", "test-seaweedfs-admin"):
vols, mounts = workload_mounts(out, wl)
if vols is None:
failed.append(f"filerWrite=true: workload {wl} not found")
continue
if "security-config" not in vols or "security-config" not in mounts:
failed.append(f"filerWrite=true: {wl} does not mount security-config (IAM gRPC would still fail)")
else:
print(f"✓ filerWrite=true: {wl} mounts security-config")
cert_vols = {v for v in vols if v.endswith("-cert")}
if cert_vols:
failed.append(f"filerWrite=true: {wl} unexpectedly has cert volumes {sorted(cert_vols)}")
# Case 4: enableSecurity=true must still render the full toml
# with both [jwt.signing] and [grpc.*]. Guards against the
# decoupling change accidentally regressing the mTLS path.
out = render({"global.seaweedfs.enableSecurity": "true"})
cm = configmap(out, "test-seaweedfs-security-config")
if cm is None:
failed.append("enableSecurity=true: security ConfigMap missing")
else:
toml = cm["data"]["security.toml"]
missing = [s for s in ("[jwt.signing]", "[grpc.master]") if s not in toml]
if missing:
failed.append(f"enableSecurity=true: security.toml missing {missing}")
else:
print("✓ enableSecurity=true: security.toml has [jwt.signing] + [grpc.*] preserved")
# Case 5: helper must tolerate explicit nulls (gemini-code-assist
# PR review). securityConfig=null was the parens-pattern crash
# the helper review caught.
for null_path in ("global.seaweedfs.securityConfig",
"global.seaweedfs.securityConfig.jwtSigning"):
try:
out = render({null_path: "null"})
except subprocess.CalledProcessError as e:
failed.append(f"{null_path}=null: render failed: {e.output[:200] if e.output else e}")
continue
if configmap(out, "test-seaweedfs-security-config") is not None:
failed.append(f"{null_path}=null: should not render configmap")
else:
print(f"✓ {null_path}=null: render tolerates explicit null")
if failed:
print("\nFAIL:", file=sys.stderr)
for f in failed:
print(f" - {f}", file=sys.stderr)
sys.exit(1)
PYEOF
echo "✓ IAM gRPC decoupling tests passed"
echo "=== Testing with monitoring enabled ==="
helm template test $CHART_DIR \
--set global.seaweedfs.monitoring.enabled=true \
+31 -5
View File
@@ -248,7 +248,33 @@ jobs:
s3-lifecycle-tests:
name: S3 Lifecycle Tests
runs-on: ubuntu-22.04
timeout-minutes: 15
timeout-minutes: 10
strategy:
fail-fast: false
matrix:
# One job per test so each gets a fresh `weed mini` server, avoiding
# the cross-test volume-pool exhaustion that surfaced when several
# TTL-pinned bucket collections piled up in a single run.
test:
- TestLifecycleAbortIncompleteMultipartUpload
- TestLifecycleAdminDispatchSucceedsWithCustomFilerGrpcPort
- TestLifecycleBootstrapWalkOnExistingObjects
- TestLifecycleConfigUpdateBetweenSweeps
- TestLifecycleDeleteBucketLifecycleStopsDispatching
- TestLifecycleDisabledRuleSkipsObject
- TestLifecycleEmptyBucketSweepIsNoOp
- TestLifecycleExpirationDateInThePast
- TestLifecycleExpirationFiresOnBackdatedObject
- TestLifecycleExpiredDeleteMarkerCleanup
- TestLifecycleMultipleBucketsInOneSweep
- TestLifecycleMultipleRulesInOneBucket
- TestLifecycleNewerNoncurrentVersions
- TestLifecycleNoncurrentVersionExpiration
- TestLifecycleSizeFilterGreaterThan
- TestLifecycleSkipsObjectLockedObjects
- TestLifecycleSuspendedVersioningExpiration
- TestLifecycleTagFilter
- TestLifecycleVersionedBucketCreatesDeleteMarker
steps:
- name: Check out code
@@ -264,12 +290,12 @@ jobs:
run: |
go install -buildvcs=false
- name: Run S3 Lifecycle Tests
timeout-minutes: 12
- name: Run ${{ matrix.test }}
timeout-minutes: 8
working-directory: test/s3/lifecycle
run: |
set -x
make test-with-server
make test-with-server TEST_PATTERN='^${{ matrix.test }}$$'
- name: Show server logs on failure
if: failure()
@@ -286,7 +312,7 @@ jobs:
if: failure()
uses: actions/upload-artifact@v7
with:
name: s3-lifecycle-test-logs
name: s3-lifecycle-test-logs-${{ matrix.test }}
path: test/s3/lifecycle/weed-test*.log
retention-days: 3
@@ -71,8 +71,10 @@ jobs:
echo "Waiting for SeaweedFS S3 gateway to be ready via proxy..."
S3_READY=0
for i in $(seq 1 30); do
# Check logs first for startup message (weed mini says "S3 service is ready")
if docker compose logs seaweedfs 2>&1 | grep -qE "S3 (gateway|service).*(started|ready)"; then
# Check logs first for the readiness line. weed mini's progress
# board prints " S3 ready (Xs)"; older builds and the
# standalone S3 binary log "S3 (gateway|service) ... ready".
if docker compose logs seaweedfs 2>&1 | grep -qE "S3 (gateway|service).*(started|ready)|S3[[:space:]]+ready"; then
echo "SeaweedFS S3 gateway is ready"
S3_READY=1
break
@@ -0,0 +1,110 @@
name: "S3 SDK V2 Route Disambiguation Tests"
on:
push:
branches: [ master ]
paths:
- 'weed/s3api/**'
- 'test/s3/sdk_v2_routing/**'
- '.github/workflows/s3-sdk-v2-routing-tests.yml'
pull_request:
branches: [ master ]
paths:
- 'weed/s3api/**'
- 'test/s3/sdk_v2_routing/**'
- '.github/workflows/s3-sdk-v2-routing-tests.yml'
concurrency:
group: ${{ github.head_ref || github.ref }}/s3-sdk-v2-routing-tests
cancel-in-progress: true
permissions:
contents: read
jobs:
s3-sdk-v2-routing-tests:
name: S3 SDK V2 Routing Tests
runs-on: ubuntu-22.04
timeout-minutes: 10
steps:
- name: Check out code
uses: actions/checkout@v6
- name: Set up Go
uses: actions/setup-go@v6
with:
go-version-file: 'go.mod'
- name: Install SeaweedFS
run: |
cd weed && go install -buildvcs=false
- name: Start weed mini (S3 on :8333)
# Pins the regression for issue #9559: AWS SDK V2 / Hadoop s3a
# listing a bucket literally named "buckets" must get an XML
# ListObjectsV2 response, not the JSON ListTableBuckets body
# served by the S3 Tables REST endpoint on the same path.
run: |
mkdir -p /tmp/seaweedfs-sdk-v2-routing
cat > /tmp/seaweedfs-sdk-v2-routing-s3.json <<'JSON'
{
"identities": [
{
"name": "admin",
"credentials": [
{"accessKey": "some_access_key1", "secretKey": "some_secret_key1"}
],
"actions": ["Admin", "Read", "Write"]
}
]
}
JSON
AWS_ACCESS_KEY_ID=some_access_key1 \
AWS_SECRET_ACCESS_KEY=some_secret_key1 \
weed mini \
-dir=/tmp/seaweedfs-sdk-v2-routing \
-s3.port=8333 \
-s3.config=/tmp/seaweedfs-sdk-v2-routing-s3.json \
-ip=127.0.0.1 \
> /tmp/weed-mini.log 2>&1 &
echo $! > /tmp/weed-mini.pid
for i in $(seq 1 30); do
if curl -s -o /dev/null -w "%{http_code}" http://127.0.0.1:8333/ | grep -qE "^(200|403)$"; then
echo "weed mini is ready"
exit 0
fi
sleep 1
done
echo "weed mini failed to start within 30s"
tail -50 /tmp/weed-mini.log
exit 1
- name: Run SDK V2 routing tests
env:
S3_ENDPOINT: http://127.0.0.1:8333
AWS_ACCESS_KEY_ID: some_access_key1
AWS_SECRET_ACCESS_KEY: some_secret_key1
AWS_REGION: us-east-1
run: go test -v -timeout=5m ./test/s3/sdk_v2_routing/...
- name: Stop weed mini
if: always()
run: |
if [ -f /tmp/weed-mini.pid ]; then
kill "$(cat /tmp/weed-mini.pid)" 2>/dev/null || true
fi
- name: Show server log on failure
if: failure()
run: |
echo "=== weed mini log (last 200 lines) ==="
tail -n 200 /tmp/weed-mini.log 2>/dev/null || echo "no log available"
- name: Archive log
if: failure()
uses: actions/upload-artifact@v7
with:
name: s3-sdk-v2-routing-server-log
path: /tmp/weed-mini.log
retention-days: 3
+120
View File
@@ -0,0 +1,120 @@
name: "Samba on FUSE Integration"
on:
push:
branches: [ master, main ]
paths:
- 'weed/mount/**'
- 'weed/filer/**'
- 'weed/cluster/**'
- 'test/samba/**'
- '.github/workflows/samba-integration.yml'
pull_request:
branches: [ master, main ]
paths:
- 'weed/mount/**'
- 'weed/filer/**'
- 'weed/cluster/**'
- 'test/samba/**'
- '.github/workflows/samba-integration.yml'
workflow_dispatch:
concurrency:
group: samba-integration/${{ github.ref }}
cancel-in-progress: true
permissions:
contents: read
jobs:
samba-integration:
name: samba-integration
runs-on: ubuntu-22.04
timeout-minutes: 45
steps:
- name: Checkout code
uses: actions/checkout@v6
- name: Set up Go
uses: actions/setup-go@v6
with:
go-version-file: 'go.mod'
- name: Start local Docker registry
run: docker run -d --restart=always -p 5000:5000 --name registry registry:2
- name: Set up Docker Buildx
uses: docker/setup-buildx-action@v4
with:
driver-opts: network=host
- name: Build weed race binary
run: |
cd docker
make binary_race
- name: Build SeaweedFS e2e image
uses: docker/build-push-action@v7
with:
context: docker
file: docker/Dockerfile.e2e
tags: localhost:5000/chrislusf/seaweedfs:e2e
push: true
cache-from: type=gha,scope=samba-e2e
cache-to: type=gha,mode=max,scope=samba-e2e
- name: Tag e2e image for docker compose
run: |
docker pull localhost:5000/chrislusf/seaweedfs:e2e
docker tag localhost:5000/chrislusf/seaweedfs:e2e chrislusf/seaweedfs:e2e
- name: Build samba image
uses: docker/build-push-action@v7
with:
context: test/samba
build-contexts: |
chrislusf/seaweedfs:e2e=docker-image://localhost:5000/chrislusf/seaweedfs:e2e
tags: localhost:5000/chrislusf/seaweedfs:samba
push: true
cache-from: type=gha,scope=samba-harness
cache-to: type=gha,mode=max,scope=samba-harness
- name: Tag samba image for docker compose
run: |
docker pull localhost:5000/chrislusf/seaweedfs:samba
docker tag localhost:5000/chrislusf/seaweedfs:samba chrislusf/seaweedfs:samba
- name: Start SeaweedFS cluster and Samba
run: |
docker compose -f test/samba/docker-compose.yml up --wait
- name: Run Samba test battery
run: |
set -o pipefail
docker compose -f test/samba/docker-compose.yml exec -T samba \
/run_inside_container.sh 2>&1 | tee /tmp/samba-output.log
- name: Collect logs
if: always()
run: |
mkdir -p /tmp/samba-docker-logs
for svc in master volume filer samba; do
docker compose -f test/samba/docker-compose.yml logs "$svc" \
> "/tmp/samba-docker-logs/${svc}.log" 2>&1 || true
done
- name: Tear down
if: always()
run: |
docker compose -f test/samba/docker-compose.yml down -v
- name: Upload logs
if: always()
uses: actions/upload-artifact@v7
with:
name: samba-integration-results
path: |
/tmp/samba-output.log
/tmp/samba-docker-logs/
retention-days: 7
+9 -7
View File
@@ -509,20 +509,22 @@ SeaweedFS Filer uses off-the-shelf stores, such as MySql, Postgres, Sqlite, Mong
### Compared to MinIO ###
MinIO follows AWS S3 closely and is ideal for testing for S3 API. It has good UI, policies, versionings, etc. SeaweedFS is trying to catch up here. It is also possible to put MinIO as a gateway in front of SeaweedFS later.
Please note, as Apr 25, 2026 MinIO ceased developement. It's strongly discouraged to use that unmaintained software with multiple security bugs.
MinIO metadata are in simple files. Each file write will incur extra writes to corresponding meta file.
MinIO followed AWS S3 closely and was ideal for testing for S3 API. It had good UI, policies, versionings, etc. SeaweedFS is trying to catch up here.
MinIO does not have optimization for lots of small files. The files are simply stored as is to local disks.
MinIO metadata were in simple files. Each file write will incur extra writes to corresponding meta file.
MinIO did not have optimization for lots of small files. The files were simply stored as is to local disks.
Plus the extra meta file and shards for erasure coding, it only amplifies the LOSF problem.
MinIO has multiple disk IO to read one file. SeaweedFS has O(1) disk reads, even for erasure coded files.
MinIO had multiple disk IO to read one file. SeaweedFS has O(1) disk reads, even for erasure coded files.
MinIO has full-time erasure coding. SeaweedFS uses replication on hot data for faster speed and optionally applies erasure coding on warm data.
MinIO had full-time erasure coding. SeaweedFS uses replication on hot data for faster speed and optionally applies erasure coding on warm data.
MinIO does not have POSIX-like API support.
MinIO did not have POSIX-like API support.
MinIO has specific requirements on storage layout. It is not flexible to adjust capacity. In SeaweedFS, just start one volume server pointing to the master. That's all.
MinIO had specific requirements on storage layout. It is not flexible to adjust capacity. In SeaweedFS, just start one volume server pointing to the master. That's all.
## Dev Plan ##
+4
View File
@@ -42,6 +42,10 @@ RUN if [ -f "/prebuilt/weed-volume-${TARGETARCH}" ]; then \
echo "Skipping Rust build for $TARGETARCH (unsupported)" && \
touch /weed-volume; \
fi
# Pre-built binaries arrive via GitHub Actions artifacts, which drop the
# executable bit, so the copied file is 0644 and exec fails with "Permission
# denied". Restore it (no-op for the empty placeholder, which stays size 0).
RUN chmod 0755 /weed-volume
FROM alpine AS final
LABEL author="Chris Lu"
+22 -23
View File
@@ -15,7 +15,6 @@ require (
github.com/coreos/go-semver v0.3.1 // indirect
github.com/coreos/go-systemd/v22 v22.6.0 // indirect
github.com/davecgh/go-spew v1.1.2-0.20180830191138-d8f796af33cc // indirect
github.com/dgryski/go-rendezvous v0.0.0-20200823014737-9f7001d12a5f // indirect
github.com/dustin/go-humanize v1.0.1
github.com/eapache/go-resiliency v1.6.0 // indirect
github.com/eapache/go-xerial-snappy v0.0.0-20230731223053-c322873962e3 // indirect
@@ -92,13 +91,13 @@ require (
gocloud.dev v0.45.0
gocloud.dev/pubsub/natspubsub v0.45.0
gocloud.dev/pubsub/rabbitpubsub v0.45.0
golang.org/x/crypto v0.50.0
golang.org/x/crypto v0.51.0
golang.org/x/exp v0.0.0-20260410095643-746e56fc9e2f
golang.org/x/image v0.39.0
golang.org/x/net v0.53.0
golang.org/x/net v0.54.0
golang.org/x/oauth2 v0.36.0
golang.org/x/sys v0.43.0
golang.org/x/text v0.36.0 // indirect
golang.org/x/sys v0.44.0
golang.org/x/text v0.37.0 // indirect
golang.org/x/tools v0.44.0 // indirect
golang.org/x/xerrors v0.0.0-20240903120638-7835f813f4da // indirect
google.golang.org/api v0.278.0
@@ -123,11 +122,11 @@ require (
github.com/apple/foundationdb/bindings/go v0.0.0-20250911184653-27f7192f47c3
github.com/arangodb/go-driver v1.6.9
github.com/armon/go-metrics v0.4.1
github.com/aws/aws-sdk-go-v2 v1.41.6
github.com/aws/aws-sdk-go-v2 v1.41.7
github.com/aws/aws-sdk-go-v2/config v1.32.14
github.com/aws/aws-sdk-go-v2/credentials v1.19.14
github.com/aws/aws-sdk-go-v2/service/s3 v1.99.0
github.com/cognusion/imaging v1.0.2
github.com/aws/aws-sdk-go-v2/service/s3 v1.101.0
github.com/cognusion/imaging v1.0.3
github.com/fluent/fluent-logger-golang v1.10.1
github.com/getsentry/sentry-go v0.44.1
github.com/go-git/go-billy/v5 v5.9.0
@@ -141,12 +140,12 @@ require (
github.com/linkedin/goavro/v2 v2.15.0
github.com/minio/crc64nvme v1.1.1
github.com/orcaman/concurrent-map/v2 v2.0.1
github.com/parquet-go/parquet-go v0.28.0
github.com/parquet-go/parquet-go v0.30.1
github.com/pkg/sftp v1.13.10
github.com/rabbitmq/amqp091-go v1.11.0
github.com/rclone/rclone v1.74.1
github.com/rdleal/intervalst v1.5.0
github.com/redis/go-redis/v9 v9.18.0
github.com/redis/go-redis/v9 v9.19.0
github.com/schollz/progressbar/v3 v3.19.0
github.com/seaweedfs/go-fuse/v2 v2.9.3
github.com/shirou/gopsutil/v4 v4.26.3
@@ -298,13 +297,13 @@ require (
cloud.google.com/go/iam v1.7.0 // indirect
cloud.google.com/go/monitoring v1.24.3 // indirect
filippo.io/edwards25519 v1.1.1 // indirect
github.com/Azure/azure-sdk-for-go/sdk/azcore v1.21.0
github.com/Azure/azure-sdk-for-go/sdk/azcore v1.21.1
github.com/Azure/azure-sdk-for-go/sdk/azidentity v1.13.1
github.com/Azure/azure-sdk-for-go/sdk/internal v1.11.2 // indirect
github.com/Azure/azure-sdk-for-go/sdk/storage/azblob v1.6.4
github.com/Azure/azure-sdk-for-go/sdk/internal v1.12.0 // indirect
github.com/Azure/azure-sdk-for-go/sdk/storage/azblob v1.7.0
github.com/Azure/azure-sdk-for-go/sdk/storage/azfile v1.5.4 // indirect
github.com/Azure/go-ntlmssp v0.1.1 // indirect
github.com/AzureAD/microsoft-authentication-library-for-go v1.6.0 // indirect
github.com/AzureAD/microsoft-authentication-library-for-go v1.7.2 // indirect
github.com/Files-com/files-sdk-go/v3 v3.3.82 // indirect
github.com/GoogleCloudPlatform/opentelemetry-operations-go/detectors/gcp v1.31.0 // indirect
github.com/GoogleCloudPlatform/opentelemetry-operations-go/exporter/metric v0.55.0 // indirect
@@ -324,17 +323,17 @@ require (
github.com/andybalholm/cascadia v1.3.3 // indirect
github.com/appscode/go-querystring v0.0.0-20170504095604-0126cfb3f1dc // indirect
github.com/arangodb/go-velocypack v0.0.0-20200318135517-5af53c29c67e // indirect
github.com/aws/aws-sdk-go-v2/aws/protocol/eventstream v1.7.8 // indirect
github.com/aws/aws-sdk-go-v2/aws/protocol/eventstream v1.7.10 // indirect
github.com/aws/aws-sdk-go-v2/feature/ec2/imds v1.18.21 // indirect
github.com/aws/aws-sdk-go-v2/feature/s3/manager v1.22.13 // indirect
github.com/aws/aws-sdk-go-v2/internal/configsources v1.4.21 // indirect
github.com/aws/aws-sdk-go-v2/internal/endpoints/v2 v2.7.21 // indirect
github.com/aws/aws-sdk-go-v2/internal/configsources v1.4.23 // indirect
github.com/aws/aws-sdk-go-v2/internal/endpoints/v2 v2.7.23 // indirect
github.com/aws/aws-sdk-go-v2/internal/ini v1.8.6 // indirect
github.com/aws/aws-sdk-go-v2/internal/v4a v1.4.22 // indirect
github.com/aws/aws-sdk-go-v2/service/internal/accept-encoding v1.13.7 // indirect
github.com/aws/aws-sdk-go-v2/service/internal/checksum v1.9.13 // indirect
github.com/aws/aws-sdk-go-v2/service/internal/presigned-url v1.13.21 // indirect
github.com/aws/aws-sdk-go-v2/service/internal/s3shared v1.19.21 // indirect
github.com/aws/aws-sdk-go-v2/internal/v4a v1.4.24 // indirect
github.com/aws/aws-sdk-go-v2/service/internal/accept-encoding v1.13.9 // indirect
github.com/aws/aws-sdk-go-v2/service/internal/checksum v1.9.15 // indirect
github.com/aws/aws-sdk-go-v2/service/internal/presigned-url v1.13.23 // indirect
github.com/aws/aws-sdk-go-v2/service/internal/s3shared v1.19.23 // indirect
github.com/aws/aws-sdk-go-v2/service/sns v1.39.7 // indirect
github.com/aws/aws-sdk-go-v2/service/sqs v1.42.17 // indirect
github.com/aws/aws-sdk-go-v2/service/sso v1.30.15 // indirect
@@ -499,7 +498,7 @@ require (
go.opentelemetry.io/otel/trace v1.43.0 // indirect
go.uber.org/multierr v1.11.0 // indirect
go.uber.org/zap v1.27.1 // indirect
golang.org/x/term v0.42.0
golang.org/x/term v0.43.0
golang.org/x/time v0.15.0
google.golang.org/genproto/googleapis/api v0.0.0-20260401024825-9d38bb4040a9 // indirect
google.golang.org/genproto/googleapis/rpc v0.0.0-20260427160629-7cedc36a6bc4 // indirect
+44 -44
View File
@@ -553,22 +553,22 @@ gioui.org v0.0.0-20210308172011-57750fc8a0a6/go.mod h1:RSH6KIUZ0p2xy5zHDxgAM4zum
git.sr.ht/~sbinet/gg v0.3.1/go.mod h1:KGYtlADtqsqANL9ueOFkWymvzUvLMQllU5Ixo+8v3pc=
github.com/AdaLogics/go-fuzz-headers v0.0.0-20240806141605-e8a1dd7889d6 h1:He8afgbRMd7mFxO99hRNu+6tazq8nFF9lIwo9JFroBk=
github.com/AdaLogics/go-fuzz-headers v0.0.0-20240806141605-e8a1dd7889d6/go.mod h1:8o94RPi1/7XTJvwPpRSzSUedZrtlirdB3r9Z20bi2f8=
github.com/Azure/azure-sdk-for-go/sdk/azcore v1.21.0 h1:fou+2+WFTib47nS+nz/ozhEBnvU96bKHy6LjRsY4E28=
github.com/Azure/azure-sdk-for-go/sdk/azcore v1.21.0/go.mod h1:t76Ruy8AHvUAC8GfMWJMa0ElSbuIcO03NLpynfbgsPA=
github.com/Azure/azure-sdk-for-go/sdk/azcore v1.21.1 h1:jHb/wfvRikGdxMXYV3QG/SzUOPYN9KEUUuC0Yd0/vC0=
github.com/Azure/azure-sdk-for-go/sdk/azcore v1.21.1/go.mod h1:pzBXCYn05zvYIrwLgtK8Ap8QcjRg+0i76tMQdWN6wOk=
github.com/Azure/azure-sdk-for-go/sdk/azidentity v1.13.1 h1:Hk5QBxZQC1jb2Fwj6mpzme37xbCDdNTxU7O9eb5+LB4=
github.com/Azure/azure-sdk-for-go/sdk/azidentity v1.13.1/go.mod h1:IYus9qsFobWIc2YVwe/WPjcnyCkPKtnHAqUYeebc8z0=
github.com/Azure/azure-sdk-for-go/sdk/azidentity/cache v0.3.2 h1:yz1bePFlP5Vws5+8ez6T3HWXPmwOK7Yvq8QxDBD3SKY=
github.com/Azure/azure-sdk-for-go/sdk/azidentity/cache v0.3.2/go.mod h1:Pa9ZNPuoNu/GztvBSKk9J1cDJW6vk/n0zLtV4mgd8N8=
github.com/Azure/azure-sdk-for-go/sdk/internal v1.11.2 h1:9iefClla7iYpfYWdzPCRDozdmndjTm8DXdpCzPajMgA=
github.com/Azure/azure-sdk-for-go/sdk/internal v1.11.2/go.mod h1:XtLgD3ZD34DAaVIIAyG3objl5DynM3CQ/vMcbBNJZGI=
github.com/Azure/azure-sdk-for-go/sdk/internal v1.12.0 h1:fhqpLE3UEXi9lPaBRpQ6XuRW0nU7hgg4zlmZZa+a9q4=
github.com/Azure/azure-sdk-for-go/sdk/internal v1.12.0/go.mod h1:7dCRMLwisfRH3dBupKeNCioWYUZ4SS09Z14H+7i8ZoY=
github.com/Azure/azure-sdk-for-go/sdk/keyvault/azkeys v0.10.0 h1:m/sWOGCREuSBqg2htVQTBY8nOZpyajYztF0vUvSZTuM=
github.com/Azure/azure-sdk-for-go/sdk/keyvault/azkeys v0.10.0/go.mod h1:Pu5Zksi2KrU7LPbZbNINx6fuVrUp/ffvpxdDj+i8LeE=
github.com/Azure/azure-sdk-for-go/sdk/keyvault/internal v0.7.1 h1:FbH3BbSb4bvGluTesZZ+ttN/MDsnMmQP36OSnDuSXqw=
github.com/Azure/azure-sdk-for-go/sdk/keyvault/internal v0.7.1/go.mod h1:9V2j0jn9jDEkCkv8w/bKTNppX/d0FVA1ud77xCIP4KA=
github.com/Azure/azure-sdk-for-go/sdk/resourcemanager/storage/armstorage v1.8.1 h1:/Zt+cDPnpC3OVDm/JKLOs7M2DKmLRIIp3XIx9pHHiig=
github.com/Azure/azure-sdk-for-go/sdk/resourcemanager/storage/armstorage v1.8.1/go.mod h1:Ng3urmn6dYe8gnbCMoHHVl5APYz2txho3koEkV2o2HA=
github.com/Azure/azure-sdk-for-go/sdk/storage/azblob v1.6.4 h1:jWQK1GI+LeGGUKBADtcH2rRqPxYB1Ljwms5gFA2LqrM=
github.com/Azure/azure-sdk-for-go/sdk/storage/azblob v1.6.4/go.mod h1:8mwH4klAm9DUgR2EEHyEEAQlRDvLPyg5fQry3y+cDew=
github.com/Azure/azure-sdk-for-go/sdk/storage/azblob v1.7.0 h1:BM85pSYlVYQHdq00nxyPoOkyLF5NArJG3bOsrmbwr4k=
github.com/Azure/azure-sdk-for-go/sdk/storage/azblob v1.7.0/go.mod h1:QYjP2cB7ZYtS/8jAbE0VSBZde/tjExqGjp+8JY6/+ts=
github.com/Azure/azure-sdk-for-go/sdk/storage/azfile v1.5.4 h1:tZh20RjgfMxKBxJiIS75iTVAKIUxrST5X2dVHMTptL4=
github.com/Azure/azure-sdk-for-go/sdk/storage/azfile v1.5.4/go.mod h1:vGYAk36rhMVCfTP7v+RVruCR0zmPe6S+36KRpDCLySw=
github.com/Azure/go-ansiterm v0.0.0-20250102033503-faa5f7b0171c h1:udKWzYgxTojEKWjV8V+WSxDXJ4NFATAsZjh8iIbsQIg=
@@ -581,8 +581,8 @@ github.com/Azure/go-ntlmssp v0.1.1 h1:l+FM/EEMb0U9QZE7mKNEDw5Mu3mFiaa2GKOoTSsNDP
github.com/Azure/go-ntlmssp v0.1.1/go.mod h1:NYqdhxd/8aAct/s4qSYZEerdPuH1liG2/X9DiVTbhpk=
github.com/AzureAD/microsoft-authentication-extensions-for-go/cache v0.1.1 h1:WJTmL004Abzc5wDB5VtZG2PJk5ndYDgVacGqfirKxjM=
github.com/AzureAD/microsoft-authentication-extensions-for-go/cache v0.1.1/go.mod h1:tCcJZ0uHAmvjsVYzEFivsRTN00oz5BEsRgQHu5JZ9WE=
github.com/AzureAD/microsoft-authentication-library-for-go v1.6.0 h1:XRzhVemXdgvJqCH0sFfrBUTnUJSBrBf7++ypk+twtRs=
github.com/AzureAD/microsoft-authentication-library-for-go v1.6.0/go.mod h1:HKpQxkWaGLJ+D/5H8QRpyQXA1eKjxkFlOMwck5+33Jk=
github.com/AzureAD/microsoft-authentication-library-for-go v1.7.2 h1:RHK7bS+HQMslb1sZpAokUt+zTVmue0hKSs2C791hhzU=
github.com/AzureAD/microsoft-authentication-library-for-go v1.7.2/go.mod h1:HKpQxkWaGLJ+D/5H8QRpyQXA1eKjxkFlOMwck5+33Jk=
github.com/BurntSushi/toml v0.3.1/go.mod h1:xHWCNGjB5oqiDr8zfno3MHue2Ht5sIBksp03qcyfWMU=
github.com/BurntSushi/xgb v0.0.0-20160522181843-27f122750802/go.mod h1:IVnqGOEym/WlBOVXweHU+Q+/VP0lqqI8lqeDx9IjBqo=
github.com/Codefor/geohash v0.0.0-20140723084247-1b41c28e3a9d h1:iG9B49Q218F/XxXNRM7k/vWf7MKmLIS8AcJV9cGN4nA=
@@ -715,10 +715,10 @@ github.com/armon/go-metrics v0.4.1/go.mod h1:E6amYzXo6aW1tqzoZGT755KkbgrJsSdpwZ+
github.com/atomicgo/cursor v0.0.1/go.mod h1:cBON2QmmrysudxNBFthvMtN32r3jxVRIvzkUiF/RuIk=
github.com/aws/aws-sdk-go v1.55.8 h1:JRmEUbU52aJQZ2AjX4q4Wu7t4uZjOu71uyNmaWlUkJQ=
github.com/aws/aws-sdk-go v1.55.8/go.mod h1:ZkViS9AqA6otK+JBBNH2++sx1sgxrPKcSzPPvQkUtXk=
github.com/aws/aws-sdk-go-v2 v1.41.6 h1:1AX0AthnBQzMx1vbmir3Y4WsnJgiydmnJjiLu+LvXOg=
github.com/aws/aws-sdk-go-v2 v1.41.6/go.mod h1:dy0UzBIfwSeot4grGvY1AqFWN5zgziMmWGzysDnHFcQ=
github.com/aws/aws-sdk-go-v2/aws/protocol/eventstream v1.7.8 h1:eBMB84YGghSocM7PsjmmPffTa+1FBUeNvGvFou6V/4o=
github.com/aws/aws-sdk-go-v2/aws/protocol/eventstream v1.7.8/go.mod h1:lyw7GFp3qENLh7kwzf7iMzAxDn+NzjXEAGjKS2UOKqI=
github.com/aws/aws-sdk-go-v2 v1.41.7 h1:DWpAJt66FmnnaRIOT/8ASTucrvuDPZASqhhLey6tLY8=
github.com/aws/aws-sdk-go-v2 v1.41.7/go.mod h1:4LAfZOPHNVNQEckOACQx60Y8pSRjIkNZQz1w92xpMJc=
github.com/aws/aws-sdk-go-v2/aws/protocol/eventstream v1.7.10 h1:gx1AwW1Iyk9Z9dD9F4akX5gnN3QZwUB20GGKH/I+Rho=
github.com/aws/aws-sdk-go-v2/aws/protocol/eventstream v1.7.10/go.mod h1:qqY157uZoqm5OXq/amuaBJyC9hgBCBQnsaWnPe905GY=
github.com/aws/aws-sdk-go-v2/config v1.32.14 h1:opVIRo/ZbbI8OIqSOKmpFaY7IwfFUOCCXBsUpJOwDdI=
github.com/aws/aws-sdk-go-v2/config v1.32.14/go.mod h1:U4/V0uKxh0Tl5sxmCBZ3AecYny4UNlVmObYjKuuaiOo=
github.com/aws/aws-sdk-go-v2/credentials v1.19.14 h1:n+UcGWAIZHkXzYt87uMFBv/l8THYELoX6gVcUvgl6fI=
@@ -727,24 +727,24 @@ github.com/aws/aws-sdk-go-v2/feature/ec2/imds v1.18.21 h1:NUS3K4BTDArQqNu2ih7yeD
github.com/aws/aws-sdk-go-v2/feature/ec2/imds v1.18.21/go.mod h1:YWNWJQNjKigKY1RHVJCuupeWDrrHjRqHm0N9rdrWzYI=
github.com/aws/aws-sdk-go-v2/feature/s3/manager v1.22.13 h1:uMC4oL6G3MNhodo358QEqSDjrgvzV3TUQ58nyQSGq2E=
github.com/aws/aws-sdk-go-v2/feature/s3/manager v1.22.13/go.mod h1:Cer86AE2686DvVUe57LPve3jUBmbujuaonSX8pNzGgw=
github.com/aws/aws-sdk-go-v2/internal/configsources v1.4.21 h1:Rgg6wvjjtX8bNHcvi9OnXWwcE0a2vGpbwmtICOsvcf4=
github.com/aws/aws-sdk-go-v2/internal/configsources v1.4.21/go.mod h1:A/kJFst/nm//cyqonihbdpQZwiUhhzpqTsdbhDdRF9c=
github.com/aws/aws-sdk-go-v2/internal/endpoints/v2 v2.7.21 h1:PEgGVtPoB6NTpPrBgqSE5hE/o47Ij9qk/SEZFbUOe9A=
github.com/aws/aws-sdk-go-v2/internal/endpoints/v2 v2.7.21/go.mod h1:p+hz+PRAYlY3zcpJhPwXlLC4C+kqn70WIHwnzAfs6ps=
github.com/aws/aws-sdk-go-v2/internal/configsources v1.4.23 h1:GpT/TrnBYuE5gan2cZbTtvP+JlHsutdmlV2YfEyNde0=
github.com/aws/aws-sdk-go-v2/internal/configsources v1.4.23/go.mod h1:xYWD6BS9ywC5bS3sz9Xh04whO/hzK2plt2Zkyrp4JuA=
github.com/aws/aws-sdk-go-v2/internal/endpoints/v2 v2.7.23 h1:bpd8vxhlQi2r1hiueOw02f/duEPTMK59Q4QMAoTTtTo=
github.com/aws/aws-sdk-go-v2/internal/endpoints/v2 v2.7.23/go.mod h1:15DfR2nw+CRHIk0tqNyifu3G1YdAOy68RftkhMDDwYk=
github.com/aws/aws-sdk-go-v2/internal/ini v1.8.6 h1:qYQ4pzQ2Oz6WpQ8T3HvGHnZydA72MnLuFK9tJwmrbHw=
github.com/aws/aws-sdk-go-v2/internal/ini v1.8.6/go.mod h1:O3h0IK87yXci+kg6flUKzJnWeziQUKciKrLjcatSNcY=
github.com/aws/aws-sdk-go-v2/internal/v4a v1.4.22 h1:rWyie/PxDRIdhNf4DzRk0lvjVOqFJuNnO8WwaIRVxzQ=
github.com/aws/aws-sdk-go-v2/internal/v4a v1.4.22/go.mod h1:zd/JsJ4P7oGfUhXn1VyLqaRZwPmZwg44Jf2dS84Dm3Y=
github.com/aws/aws-sdk-go-v2/service/internal/accept-encoding v1.13.7 h1:5EniKhLZe4xzL7a+fU3C2tfUN4nWIqlLesfrjkuPFTY=
github.com/aws/aws-sdk-go-v2/service/internal/accept-encoding v1.13.7/go.mod h1:x0nZssQ3qZSnIcePWLvcoFisRXJzcTVvYpAAdYX8+GI=
github.com/aws/aws-sdk-go-v2/service/internal/checksum v1.9.13 h1:JRaIgADQS/U6uXDqlPiefP32yXTda7Kqfx+LgspooZM=
github.com/aws/aws-sdk-go-v2/service/internal/checksum v1.9.13/go.mod h1:CEuVn5WqOMilYl+tbccq8+N2ieCy0gVn3OtRb0vBNNM=
github.com/aws/aws-sdk-go-v2/service/internal/presigned-url v1.13.21 h1:c31//R3xgIJMSC8S6hEVq+38DcvUlgFY0FM6mSI5oto=
github.com/aws/aws-sdk-go-v2/service/internal/presigned-url v1.13.21/go.mod h1:r6+pf23ouCB718FUxaqzZdbpYFyDtehyZcmP5KL9FkA=
github.com/aws/aws-sdk-go-v2/service/internal/s3shared v1.19.21 h1:ZlvrNcHSFFWURB8avufQq9gFsheUgjVD9536obIknfM=
github.com/aws/aws-sdk-go-v2/service/internal/s3shared v1.19.21/go.mod h1:cv3TNhVrssKR0O/xxLJVRfd2oazSnZnkUeTf6ctUwfQ=
github.com/aws/aws-sdk-go-v2/service/s3 v1.99.0 h1:hlSuz394kV0vhv9drL5lhuEFbEOEP1VyQpy15qWh1Pk=
github.com/aws/aws-sdk-go-v2/service/s3 v1.99.0/go.mod h1:uoA43SdFwacedBfSgfFSjjCvYe8aYBS7EnU5GZ/YKMM=
github.com/aws/aws-sdk-go-v2/internal/v4a v1.4.24 h1:OQqn11BtaYv1WLUowvcA30MpzIu8Ti4pcLPIIyoKZrA=
github.com/aws/aws-sdk-go-v2/internal/v4a v1.4.24/go.mod h1:X5ZJyfwVrWA96GzPmUCWFQaEARPR7gCrpq2E92PJwAE=
github.com/aws/aws-sdk-go-v2/service/internal/accept-encoding v1.13.9 h1:FLudkZLt5ci0ozzgkVo8BJGwvqNaZbTWb3UcucAateA=
github.com/aws/aws-sdk-go-v2/service/internal/accept-encoding v1.13.9/go.mod h1:w7wZ/s9qK7c8g4al+UyoF1Sp/Z45UwMGcqIzLWVQHWk=
github.com/aws/aws-sdk-go-v2/service/internal/checksum v1.9.15 h1:ieLCO1JxUWuxTZ1cRd0GAaeX7O6cIxnwk7tc1LsQhC4=
github.com/aws/aws-sdk-go-v2/service/internal/checksum v1.9.15/go.mod h1:e3IzZvQ3kAWNykvE0Tr0RDZCMFInMvhku3qNpcIQXhM=
github.com/aws/aws-sdk-go-v2/service/internal/presigned-url v1.13.23 h1:pbrxO/kuIwgEsOPLkaHu0O+m4fNgLU8B3vxQ+72jTPw=
github.com/aws/aws-sdk-go-v2/service/internal/presigned-url v1.13.23/go.mod h1:/CMNUqoj46HpS3MNRDEDIwcgEnrtZlKRaHNaHxIFpNA=
github.com/aws/aws-sdk-go-v2/service/internal/s3shared v1.19.23 h1:03xatSQO4+AM1lTAbnRg5OK528EUg744nW7F73U8DKw=
github.com/aws/aws-sdk-go-v2/service/internal/s3shared v1.19.23/go.mod h1:M8l3mwgx5ToK7wot2sBBce/ojzgnPzZXUV445gTSyE8=
github.com/aws/aws-sdk-go-v2/service/s3 v1.101.0 h1:etqBTKY581iwLL/H/S2sVgk3C9lAsTJFeXWFDsDcWOU=
github.com/aws/aws-sdk-go-v2/service/s3 v1.101.0/go.mod h1:L2dcoOgS2VSgbPLvpak2NyUPsO1TBN7M45Z4H7DlRc4=
github.com/aws/aws-sdk-go-v2/service/signin v1.0.9 h1:QKZH0S178gCmFEgst8hN0mCX1KxLgHBKKY/CLqwP8lg=
github.com/aws/aws-sdk-go-v2/service/signin v1.0.9/go.mod h1:7yuQJoT+OoH8aqIxw9vwF+8KpvLZ8AWmvmUWHsGQZvI=
github.com/aws/aws-sdk-go-v2/service/sns v1.39.7 h1:fovS7qGMT+BBSuifkySdVaMWxXTyaYT6qaBx/1y6Ij4=
@@ -866,8 +866,8 @@ github.com/cockroachdb/redact v1.1.5 h1:u1PMllDkdFfPWaNGMyLD1+so+aq3uUItthCFqzwP
github.com/cockroachdb/redact v1.1.5/go.mod h1:BVNblN9mBWFyMyqK1k3AAiSxhvhfK2oOZZ2lK+dpvRg=
github.com/cockroachdb/version v0.0.0-20250314144055-3860cd14adf2 h1:8Vfw2iNEpYIV6aLtMwT5UOGuPmp9MKlEKWKFTuB+MPU=
github.com/cockroachdb/version v0.0.0-20250314144055-3860cd14adf2/go.mod h1:P9WiZOdQ1R/ZZDL0WzF5wlyRvrjtfhNOwMZymFpBwjE=
github.com/cognusion/imaging v1.0.2 h1:BQwBV8V8eF3+dwffp8Udl9xF1JKh5Z0z5JkJwAi98Mc=
github.com/cognusion/imaging v1.0.2/go.mod h1:mj7FvH7cT2dlFogQOSUQRtotBxJ4gFQ2ySMSmBm5dSk=
github.com/cognusion/imaging v1.0.3 h1:nHyIeEVDV8JkBbuhgx8iSBW72W8rHbDEyruA0Jh7Lnk=
github.com/cognusion/imaging v1.0.3/go.mod h1:38tFLFhGK81ORThZG8dVXPtp1uhd+3xM1WXuMn+unOA=
github.com/colinmarc/hdfs/v2 v2.4.0 h1:v6R8oBx/Wu9fHpdPoJJjpGSUxo8NhHIwrwsfhFvU9W0=
github.com/colinmarc/hdfs/v2 v2.4.0/go.mod h1:0NAO+/3knbMx6+5pCv+Hcbaz4xn/Zzbn9+WIib2rKVI=
github.com/compose-spec/compose-go/v2 v2.9.0 h1:UHSv/QHlo6QJtrT4igF1rdORgIUhDo1gWuyJUoiNNIM=
@@ -1667,8 +1667,8 @@ github.com/parquet-go/bitpack v1.0.0 h1:AUqzlKzPPXf2bCdjfj4sTeacrUwsT7NlcYDMUQxP
github.com/parquet-go/bitpack v1.0.0/go.mod h1:XnVk9TH+O40eOOmvpAVZ7K2ocQFrQwysLMnc6M/8lgs=
github.com/parquet-go/jsonlite v1.0.0 h1:87QNdi56wOfsE5bdgas0vRzHPxfJgzrXGml1zZdd7VU=
github.com/parquet-go/jsonlite v1.0.0/go.mod h1:nDjpkpL4EOtqs6NQugUsi0Rleq9sW/OtC1NnZEnxzF0=
github.com/parquet-go/parquet-go v0.28.0 h1:ECyksyv8T2pOrlLsN7aWJIoQakyk/HtxQ2lchgS4els=
github.com/parquet-go/parquet-go v0.28.0/go.mod h1:navtkAYr2LGoJVp141oXPlO/sxLvaOe3la2JEoD8+rg=
github.com/parquet-go/parquet-go v0.30.1 h1:Oy6ganNrAdFiVwy7wNmWagfPTWA2X9Z3tVHBc7JtuX8=
github.com/parquet-go/parquet-go v0.30.1/go.mod h1:navtkAYr2LGoJVp141oXPlO/sxLvaOe3la2JEoD8+rg=
github.com/pascaldekloe/goe v0.1.0 h1:cBOtyMzM9HTpWjXfbbunk26uA6nG3a8n06Wieeh0MwY=
github.com/pascaldekloe/goe v0.1.0/go.mod h1:lzWF7FIEvWOWxwDKqyGYQf6ZUaNfKdP144TG7ZOy1lc=
github.com/patrickmn/go-cache v2.1.0+incompatible h1:HRMgzkcYKYpi3C8ajMPV8OFXaaRUnok+kx1WdO15EQc=
@@ -1795,8 +1795,8 @@ github.com/rcrowley/go-metrics v0.0.0-20201227073835-cf1acfcdf475 h1:N/ElC8H3+5X
github.com/rcrowley/go-metrics v0.0.0-20201227073835-cf1acfcdf475/go.mod h1:bCqnVzQkZxMG4s8nGwiZ5l3QUCyqpo9Y+/ZMZ9VjZe4=
github.com/rdleal/intervalst v1.5.0 h1:SEB9bCFz5IqD1yhfH1Wv8IBnY/JQxDplwkxHjT6hamU=
github.com/rdleal/intervalst v1.5.0/go.mod h1:xO89Z6BC+LQDH+IPQQw/OESt5UADgFD41tYMUINGpxQ=
github.com/redis/go-redis/v9 v9.18.0 h1:pMkxYPkEbMPwRdenAzUNyFNrDgHx9U+DrBabWNfSRQs=
github.com/redis/go-redis/v9 v9.18.0/go.mod h1:k3ufPphLU5YXwNTUcCRXGxUoF1fqxnhFQmscfkCoDA0=
github.com/redis/go-redis/v9 v9.19.0 h1:XPVaaPSnG6RhYf7p+rmSa9zZfeVAnWsH5h3lxthOm/k=
github.com/redis/go-redis/v9 v9.19.0/go.mod h1:v/M13XI1PVCDcm01VtPFOADfZtHf8YW3baQf57KlIkA=
github.com/redis/rueidis v1.0.71 h1:pODtnAR5GAB7j4ekhldZ29HKOxe4Hph0GTDGk1ayEQY=
github.com/redis/rueidis v1.0.71/go.mod h1:lfdcZzJ1oKGKL37vh9fO3ymwt+0TdjkkUCJxbgpmcgQ=
github.com/redis/rueidis/rueidiscompat v1.0.71 h1:wNZ//kEjMZgBM0KCk7ncOX8KmAgROU2kDdDNpwheG4w=
@@ -2216,8 +2216,8 @@ golang.org/x/crypto v0.14.0/go.mod h1:MVFd36DqK4CsrnJYDkBA3VC4m2GkXAM0PvzMCn4JQf
golang.org/x/crypto v0.19.0/go.mod h1:Iy9bg/ha4yyC70EfRS8jz+B6ybOBKMaSxLj6P6oBDfU=
golang.org/x/crypto v0.23.0/go.mod h1:CKFgDieR+mRhux2Lsu27y0fO304Db0wZe70UKqHu0v8=
golang.org/x/crypto v0.31.0/go.mod h1:kDsLvtWBEx7MV9tJOj9bnXsPbxwJQ6csT/x4KIN4Ssk=
golang.org/x/crypto v0.50.0 h1:zO47/JPrL6vsNkINmLoo/PH1gcxpls50DNogFvB5ZGI=
golang.org/x/crypto v0.50.0/go.mod h1:3muZ7vA7PBCE6xgPX7nkzzjiUq87kRItoJQM1Yo8S+Q=
golang.org/x/crypto v0.51.0 h1:IBPXwPfKxY7cWQZ38ZCIRPI50YLeevDLlLnyC5wRGTI=
golang.org/x/crypto v0.51.0/go.mod h1:8AdwkbraGNABw2kOX6YFPs3WM22XqI4EXEd8g+x7Oc8=
golang.org/x/exp v0.0.0-20180321215751-8460e604b9de/go.mod h1:CJ0aWSM057203Lf6IL+f9T1iT9GByDxfZKAQTCR3kQA=
golang.org/x/exp v0.0.0-20180807140117-3d87b88a115f/go.mod h1:CJ0aWSM057203Lf6IL+f9T1iT9GByDxfZKAQTCR3kQA=
golang.org/x/exp v0.0.0-20190121172915-509febef88a4/go.mod h1:CJ0aWSM057203Lf6IL+f9T1iT9GByDxfZKAQTCR3kQA=
@@ -2352,8 +2352,8 @@ golang.org/x/net v0.16.0/go.mod h1:NxSsAGuq816PNPmqtQdLE42eU2Fs7NoRIZrHJAlaCOE=
golang.org/x/net v0.21.0/go.mod h1:bIjVDfnllIU7BJ2DNgfnXvpSvtn8VRwhlsaeUTyUS44=
golang.org/x/net v0.25.0/go.mod h1:JkAGAh7GEvH74S6FOH42FLoXpXbE/aqXSrIQjXgsiwM=
golang.org/x/net v0.33.0/go.mod h1:HXLR5J+9DxmrqMwG9qjGCxZ+zKXxBru04zlTvWlWuN4=
golang.org/x/net v0.53.0 h1:d+qAbo5L0orcWAr0a9JweQpjXF19LMXJE8Ey7hwOdUA=
golang.org/x/net v0.53.0/go.mod h1:JvMuJH7rrdiCfbeHoo3fCQU24Lf5JJwT9W3sJFulfgs=
golang.org/x/net v0.54.0 h1:2zJIZAxAHV/OHCDTCOHAYehQzLfSXuf/5SoL/Dv6w/w=
golang.org/x/net v0.54.0/go.mod h1:Sj4oj8jK6XmHpBZU/zWHw3BV3abl4Kvi+Ut7cQcY+cQ=
golang.org/x/oauth2 v0.0.0-20180821212333-d2e6202438be/go.mod h1:N/0e6XlmueqKjAGxoOufVs8QHGRruUQn6yWY3a++T0U=
golang.org/x/oauth2 v0.0.0-20190226205417-e64efc72b421/go.mod h1:gOpvHmFTYa4IltrdGE7lF6nIHvwfUNPOp7c8zoXwtLw=
golang.org/x/oauth2 v0.0.0-20190604053449-0f29369cfe45/go.mod h1:gOpvHmFTYa4IltrdGE7lF6nIHvwfUNPOp7c8zoXwtLw=
@@ -2511,8 +2511,8 @@ golang.org/x/sys v0.13.0/go.mod h1:oPkhp1MJrh7nUepCBck5+mAzfO9JrbApNNgaTdGDITg=
golang.org/x/sys v0.17.0/go.mod h1:/VUhepiaJMQUp4+oa/7Zr1D23ma6VTLIYjOOTFZPUcA=
golang.org/x/sys v0.20.0/go.mod h1:/VUhepiaJMQUp4+oa/7Zr1D23ma6VTLIYjOOTFZPUcA=
golang.org/x/sys v0.28.0/go.mod h1:/VUhepiaJMQUp4+oa/7Zr1D23ma6VTLIYjOOTFZPUcA=
golang.org/x/sys v0.43.0 h1:Rlag2XtaFTxp19wS8MXlJwTvoh8ArU6ezoyFsMyCTNI=
golang.org/x/sys v0.43.0/go.mod h1:4GL1E5IUh+htKOUEOaiffhrAeqysfVGipDYzABqnCmw=
golang.org/x/sys v0.44.0 h1:ildZl3J4uzeKP07r2F++Op7E9B29JRUy+a27EibtBTQ=
golang.org/x/sys v0.44.0/go.mod h1:4GL1E5IUh+htKOUEOaiffhrAeqysfVGipDYzABqnCmw=
golang.org/x/telemetry v0.0.0-20240228155512-f48c80bd79b2/go.mod h1:TeRTkGYfJXctD9OcfyVLyj2J3IxLnKwHJR8f4D8a3YE=
golang.org/x/telemetry v0.0.0-20260409153401-be6f6cb8b1fa h1:efT73AJZfAAUV7SOip6pWGkwJDzIGiKBZGVzHYa+ve4=
golang.org/x/telemetry v0.0.0-20260409153401-be6f6cb8b1fa/go.mod h1:kHjTxDEnAu6/Nl9lDkzjWpR+bmKfxeiRuSDlsMb70gE=
@@ -2531,8 +2531,8 @@ golang.org/x/term v0.13.0/go.mod h1:LTmsnFJwVN6bCy1rVCoS+qHT1HhALEFxKncY3WNNh4U=
golang.org/x/term v0.17.0/go.mod h1:lLRBjIVuehSbZlaOtGMbcMncT+aqLLLmKrsjNrUguwk=
golang.org/x/term v0.20.0/go.mod h1:8UkIAJTvZgivsXaD6/pH6U9ecQzZ45awqEOzuCvwpFY=
golang.org/x/term v0.27.0/go.mod h1:iMsnZpn0cago0GOrHO2+Y7u7JPn5AylBrcoWkElMTSM=
golang.org/x/term v0.42.0 h1:UiKe+zDFmJobeJ5ggPwOshJIVt6/Ft0rcfrXZDLWAWY=
golang.org/x/term v0.42.0/go.mod h1:Dq/D+snpsbazcBG5+F9Q1n2rXV8Ma+71xEjTRufARgY=
golang.org/x/term v0.43.0 h1:S4RLU2sB31O/NCl+zFN9Aru9A/Cq2aqKpTZJ6B+DwT4=
golang.org/x/term v0.43.0/go.mod h1:lrhlHNdQJHO+1qVYiHfFKVuVioJIheAc3fBSMFYEIsk=
golang.org/x/text v0.0.0-20170915032832-14c0d48ead0c/go.mod h1:NqM8EUOU14njkJ3fqMW+pc6Ldnwhi/IjpwHt7yyuwOQ=
golang.org/x/text v0.3.0/go.mod h1:NqM8EUOU14njkJ3fqMW+pc6Ldnwhi/IjpwHt7yyuwOQ=
golang.org/x/text v0.3.1-0.20180807135948-17ff2d5776d2/go.mod h1:NqM8EUOU14njkJ3fqMW+pc6Ldnwhi/IjpwHt7yyuwOQ=
@@ -2553,8 +2553,8 @@ golang.org/x/text v0.13.0/go.mod h1:TvPlkZtksWOMsz7fbANvkp4WM8x/WCo/om8BMLbz+aE=
golang.org/x/text v0.14.0/go.mod h1:18ZOQIKpY8NJVqYksKHtTdi31H5itFRjB5/qKTNYzSU=
golang.org/x/text v0.15.0/go.mod h1:18ZOQIKpY8NJVqYksKHtTdi31H5itFRjB5/qKTNYzSU=
golang.org/x/text v0.21.0/go.mod h1:4IBbMaMmOPCJ8SecivzSH54+73PCFmPWxNTLm+vZkEQ=
golang.org/x/text v0.36.0 h1:JfKh3XmcRPqZPKevfXVpI1wXPTqbkE5f7JA92a55Yxg=
golang.org/x/text v0.36.0/go.mod h1:NIdBknypM8iqVmPiuco0Dh6P5Jcdk8lJL0CUebqK164=
golang.org/x/text v0.37.0 h1:Cqjiwd9eSg8e0QAkyCaQTNHFIIzWtidPahFWR83rTrc=
golang.org/x/text v0.37.0/go.mod h1:a5sjxXGs9hsn/AJVwuElvCAo9v8QYLzvavO5z2PiM38=
golang.org/x/time v0.0.0-20181108054448-85acf8d2951c/go.mod h1:tRJNPiyCQ0inRvYxbN9jk5I+vvW/OXSQhTDSoE431IQ=
golang.org/x/time v0.0.0-20190308202827-9d24e82272b4/go.mod h1:tRJNPiyCQ0inRvYxbN9jk5I+vvW/OXSQhTDSoE431IQ=
golang.org/x/time v0.0.0-20191024005414-555d28b269f0/go.mod h1:tRJNPiyCQ0inRvYxbN9jk5I+vvW/OXSQhTDSoE431IQ=
+2 -2
View File
@@ -1,6 +1,6 @@
apiVersion: v1
description: SeaweedFS
name: seaweedfs
appVersion: "4.24"
appVersion: "4.27"
# Dev note: Trigger a helm chart release by `git tag -a helm-<version>`
version: 4.24.0
version: 4.27.0
@@ -135,6 +135,14 @@ spec:
{{ toYaml $value | nindent 16 | trim }}
{{- end -}}
{{- end }}
{{- $secretExtraEnvironmentVars := .Values.admin.secretExtraEnvironmentVars }}
{{- if $secretExtraEnvironmentVars }}
{{- range $key := keys $secretExtraEnvironmentVars | sortAlpha }}
{{- $value := index $secretExtraEnvironmentVars $key }}
- name: {{ $key }}
valueFrom: {{ toYaml $value | nindent 16 }}
{{- end }}
{{- end }}
command:
- "/bin/sh"
- "-ec"
@@ -180,11 +188,13 @@ spec:
- name: admin-logs
mountPath: /logs
{{- end }}
{{- if .Values.global.seaweedfs.enableSecurity }}
{{- if include "seaweedfs.securityConfigEnabled" . }}
- name: security-config
readOnly: true
mountPath: /etc/seaweedfs/security.toml
subPath: security.toml
{{- end }}
{{- if .Values.global.seaweedfs.enableSecurity }}
- name: ca-cert
readOnly: true
mountPath: /usr/local/share/ca-certificates/ca/
@@ -275,10 +285,12 @@ spec:
persistentVolumeClaim:
claimName: {{ .Values.admin.logs.claimName }}
{{- end }}
{{- if .Values.global.seaweedfs.enableSecurity }}
{{- if include "seaweedfs.securityConfigEnabled" . }}
- name: security-config
configMap:
name: {{ include "seaweedfs.fullname" . }}-security-config
{{- end }}
{{- if .Values.global.seaweedfs.enableSecurity }}
- name: ca-cert
secret:
secretName: {{ include "seaweedfs.fullname" . }}-ca-cert
@@ -330,11 +330,13 @@ spec:
mountPath: /etc/seaweedfs/master.toml
subPath: master.toml
readOnly: true
{{- if .Values.global.seaweedfs.enableSecurity }}
{{- if include "seaweedfs.securityConfigEnabled" . }}
- name: security-config
mountPath: /etc/seaweedfs/security.toml
subPath: security.toml
readOnly: true
{{- end }}
{{- if .Values.global.seaweedfs.enableSecurity }}
- name: ca-cert
mountPath: /usr/local/share/ca-certificates/ca/
readOnly: true
@@ -461,10 +463,12 @@ spec:
- name: master-config
configMap:
name: {{ include "seaweedfs.fullname" . }}-master-config
{{- if .Values.global.seaweedfs.enableSecurity }}
{{- if include "seaweedfs.securityConfigEnabled" . }}
- name: security-config
configMap:
name: {{ include "seaweedfs.fullname" . }}-security-config
{{- end }}
{{- if .Values.global.seaweedfs.enableSecurity }}
- name: ca-cert
secret:
secretName: {{ include "seaweedfs.fullname" . }}-ca-cert
@@ -117,11 +117,13 @@ spec:
name: config-users
readOnly: true
{{- end }}
{{- if .Values.global.seaweedfs.enableSecurity }}
{{- if include "seaweedfs.securityConfigEnabled" . }}
- name: security-config
readOnly: true
mountPath: /etc/seaweedfs/security.toml
subPath: security.toml
{{- end }}
{{- if .Values.global.seaweedfs.enableSecurity }}
- name: ca-cert
readOnly: true
mountPath: /usr/local/share/ca-certificates/ca/
@@ -179,10 +181,12 @@ spec:
secretName: {{ include "seaweedfs.fullname" . }}-s3-secret
{{- end }}
{{- end }}
{{- if .Values.global.seaweedfs.enableSecurity }}
{{- if include "seaweedfs.securityConfigEnabled" . }}
- name: security-config
configMap:
name: {{ include "seaweedfs.fullname" . }}-security-config
{{- end }}
{{- if .Values.global.seaweedfs.enableSecurity }}
- name: ca-cert
secret:
secretName: {{ include "seaweedfs.fullname" . }}-ca-cert
@@ -234,11 +234,13 @@ spec:
mountPath: /etc/seaweedfs/notification.toml
subPath: notification.toml
{{- end }}
{{- if .Values.global.seaweedfs.enableSecurity }}
{{- if include "seaweedfs.securityConfigEnabled" . }}
- name: security-config
readOnly: true
mountPath: /etc/seaweedfs/security.toml
subPath: security.toml
{{- end }}
{{- if .Values.global.seaweedfs.enableSecurity }}
- name: ca-cert
readOnly: true
mountPath: /usr/local/share/ca-certificates/ca/
@@ -274,7 +276,8 @@ spec:
name: swfs-s3-tls
{{- end }}
{{- end }}
{{- $isJwtEnabled := or .Values.global.seaweedfs.securityConfig.jwtSigning.filerWrite .Values.global.seaweedfs.securityConfig.jwtSigning.filerRead }}
{{- $jwt := (.Values.global.seaweedfs.securityConfig).jwtSigning | default dict }}
{{- $isJwtEnabled := or $jwt.filerWrite $jwt.filerRead }}
{{- if .Values.filer.readinessProbe.enabled }}
readinessProbe:
{{- if or $isJwtEnabled .Values.filer.readinessProbe.tcpSocket }}
@@ -368,10 +371,12 @@ spec:
configMap:
name: {{ include "seaweedfs.fullname" . }}-notification-config
{{- end }}
{{- if .Values.global.seaweedfs.enableSecurity }}
{{- if include "seaweedfs.securityConfigEnabled" . }}
- name: security-config
configMap:
name: {{ include "seaweedfs.fullname" . }}-security-config
{{- end }}
{{- if .Values.global.seaweedfs.enableSecurity }}
- name: ca-cert
secret:
secretName: {{ include "seaweedfs.fullname" . }}-ca-cert
@@ -188,11 +188,13 @@ spec:
readOnly: true
mountPath: /etc/seaweedfs/master.toml
subPath: master.toml
{{- if .Values.global.seaweedfs.enableSecurity }}
{{- if include "seaweedfs.securityConfigEnabled" . }}
- name: security-config
readOnly: true
mountPath: /etc/seaweedfs/security.toml
subPath: security.toml
{{- end }}
{{- if .Values.global.seaweedfs.enableSecurity }}
- name: ca-cert
readOnly: true
mountPath: /usr/local/share/ca-certificates/ca/
@@ -287,10 +289,12 @@ spec:
- name: master-config
configMap:
name: {{ include "seaweedfs.fullname" . }}-master-config
{{- if .Values.global.seaweedfs.enableSecurity }}
{{- if include "seaweedfs.securityConfigEnabled" . }}
- name: security-config
configMap:
name: {{ include "seaweedfs.fullname" . }}-security-config
{{- end }}
{{- if .Values.global.seaweedfs.enableSecurity }}
- name: ca-cert
secret:
secretName: {{ include "seaweedfs.fullname" . }}-ca-cert
@@ -156,11 +156,13 @@ spec:
name: config-users
readOnly: true
{{- end }}
{{- if .Values.global.seaweedfs.enableSecurity }}
{{- if include "seaweedfs.securityConfigEnabled" . }}
- name: security-config
readOnly: true
mountPath: /etc/seaweedfs/security.toml
subPath: security.toml
{{- end }}
{{- if .Values.global.seaweedfs.enableSecurity }}
- name: ca-cert
readOnly: true
mountPath: /usr/local/share/ca-certificates/ca/
@@ -249,10 +251,12 @@ spec:
- name: logs
emptyDir: {}
{{- end }}
{{- if .Values.global.seaweedfs.enableSecurity }}
{{- if include "seaweedfs.securityConfigEnabled" . }}
- name: security-config
configMap:
name: {{ include "seaweedfs.fullname" . }}-security-config
{{- end }}
{{- if .Values.global.seaweedfs.enableSecurity }}
- name: ca-cert
secret:
secretName: {{ include "seaweedfs.fullname" . }}-ca-cert
@@ -176,11 +176,13 @@ spec:
- mountPath: /etc/sw/ssh
name: config-ssh
readOnly: true
{{- if .Values.global.seaweedfs.enableSecurity }}
{{- if include "seaweedfs.securityConfigEnabled" . }}
- name: security-config
readOnly: true
mountPath: /etc/seaweedfs/security.toml
subPath: security.toml
{{- end }}
{{- if .Values.global.seaweedfs.enableSecurity }}
- name: ca-cert
readOnly: true
mountPath: /usr/local/share/ca-certificates/ca/
@@ -264,10 +266,12 @@ spec:
- name: logs
emptyDir: {}
{{- end }}
{{- if .Values.global.seaweedfs.enableSecurity }}
{{- if include "seaweedfs.securityConfigEnabled" . }}
- name: security-config
configMap:
name: {{ include "seaweedfs.fullname" . }}-security-config
{{- end }}
{{- if .Values.global.seaweedfs.enableSecurity }}
- name: ca-cert
secret:
secretName: {{ include "seaweedfs.fullname" . }}-ca-cert
@@ -332,6 +332,16 @@ Create the name of the service account to use
{{- .Values.global.seaweedfs.serviceAccountName | default "seaweedfs" -}}
{{- end -}}
{{/* True when security.toml should be rendered and mounted. volumeWrite is
excluded since it defaults to true. */}}
{{- define "seaweedfs.securityConfigEnabled" -}}
{{- $sec := (.Values.global.seaweedfs).securityConfig | default dict -}}
{{- $jwt := $sec.jwtSigning | default dict -}}
{{- if or .Values.global.seaweedfs.enableSecurity $jwt.volumeRead $jwt.filerWrite $jwt.filerRead -}}
true
{{- end -}}
{{- end -}}
{{/* S3 TLS cert/key arguments, using custom secret if s3.tlsSecret is set */}}
{{- define "seaweedfs.s3.tlsArgs" -}}
{{- $prefix := .prefix -}}
@@ -1,5 +1,5 @@
{{- include "seaweedfs.compat" . -}}
{{- if .Values.global.seaweedfs.enableSecurity }}
{{- if include "seaweedfs.securityConfigEnabled" . }}
apiVersion: v1
kind: ConfigMap
metadata:
@@ -55,6 +55,7 @@ data:
key = "{{ dig "jwt" "filer_signing" "read" "key" (randAlphaNum 10 | b64enc) $securityConfig }}"
{{- end }}
{{- if .Values.global.seaweedfs.enableSecurity }}
# all grpc tls authentications are mutual
# the values for the following ca, cert, and key are paths to the PERM files.
[grpc]
@@ -94,4 +95,5 @@ data:
[https.volume]
cert = ""
key = ""
{{- end }}
{{- end }}
@@ -137,6 +137,9 @@ spec:
- "/bin/sh"
- "-ec"
- |
{{- if $volume.rust }}
exec /usr/bin/weed-volume \
{{- else }}
exec /usr/bin/weed \
{{- if $volume.logs }}
-logdir=/logs \
@@ -149,6 +152,7 @@ spec:
-v={{ $.Values.global.seaweedfs.loggingLevel }} \
{{- end }}
volume \
{{- end }}
-port={{ $volume.port }} \
{{- if $volume.metricsPort }}
-metricsPort={{ $volume.metricsPort }} \
@@ -180,7 +184,7 @@ spec:
{{- if $volume.imagesFixOrientation }}
-images.fix.orientation \
{{- end }}
{{- if $volume.pulseSeconds }}
{{- if and $volume.pulseSeconds (not $volume.rust) }}
-pulseSeconds={{ $volume.pulseSeconds }} \
{{- end }}
{{- if $volume.index }}
@@ -211,11 +215,13 @@ spec:
- name: idx
mountPath: "/idx/"
{{- end }}
{{- if $.Values.global.seaweedfs.enableSecurity }}
{{- if include "seaweedfs.securityConfigEnabled" $ }}
- name: security-config
readOnly: true
mountPath: /etc/seaweedfs/security.toml
subPath: security.toml
{{- end }}
{{- if $.Values.global.seaweedfs.enableSecurity }}
- name: ca-cert
readOnly: true
mountPath: /usr/local/share/ca-certificates/ca/
@@ -333,10 +339,12 @@ spec:
emptyDir: {}
{{- end }}
{{- end }}
{{- if $.Values.global.seaweedfs.enableSecurity }}
{{- if include "seaweedfs.securityConfigEnabled" $ }}
- name: security-config
configMap:
name: {{ include "seaweedfs.fullname" $ }}-security-config
{{- end }}
{{- if $.Values.global.seaweedfs.enableSecurity }}
- name: ca-cert
secret:
secretName: {{ include "seaweedfs.fullname" $ }}-ca-cert
@@ -149,11 +149,13 @@ spec:
- name: worker-logs
mountPath: /logs
{{- end }}
{{- if .Values.global.seaweedfs.enableSecurity }}
{{- if include "seaweedfs.securityConfigEnabled" . }}
- name: security-config
readOnly: true
mountPath: /etc/seaweedfs/security.toml
subPath: security.toml
{{- end }}
{{- if .Values.global.seaweedfs.enableSecurity }}
- name: ca-cert
readOnly: true
mountPath: /usr/local/share/ca-certificates/ca/
@@ -252,10 +254,12 @@ spec:
persistentVolumeClaim:
claimName: {{ .Values.worker.logs.claimName }}
{{- end }}
{{- if .Values.global.seaweedfs.enableSecurity }}
{{- if include "seaweedfs.securityConfigEnabled" . }}
- name: security-config
configMap:
name: {{ include "seaweedfs.fullname" . }}-security-config
{{- end }}
{{- if .Values.global.seaweedfs.enableSecurity }}
- name: ca-cert
secret:
secretName: {{ include "seaweedfs.fullname" . }}-ca-cert
+14
View File
@@ -18,6 +18,8 @@ global:
loggingLevel: 1
enableSecurity: false
masterServer: null
# filerWrite: true mounts security.toml on filer + admin without needing
# enableSecurity (mTLS); required for the Admin UI Users tab.
securityConfig:
jwtSigning:
volumeWrite: true
@@ -303,6 +305,11 @@ volume:
enabled: true
imageOverride: null
restartPolicy: null
# Run the Rust volume server (/usr/bin/weed-volume) instead of the Go one.
# Requires an image that ships the Rust binary (amd64/arm64). The Go-only
# log flags (-logtostderr/-logdir/-v) and -pulseSeconds are dropped; set log
# level via the RUST_LOG env var in extraEnvironmentVars if needed.
rust: false
port: 8080
grpcPort: 18080
metricsPort: 9327
@@ -1286,6 +1293,13 @@ admin:
extraEnvironmentVars: {}
# secret env variables (e.g. for injecting OIDC client secret from a Kubernetes Secret)
secretExtraEnvironmentVars: {}
# WEED_ADMIN_OIDC_CLIENT_SECRET:
# secretKeyRef:
# name: seaweedfs-admin-oidc
# key: client_secret
# Health checks
livenessProbe:
enabled: true
@@ -22,6 +22,9 @@ service SeaweedFiler {
rpc UpdateEntry (UpdateEntryRequest) returns (UpdateEntryResponse) {
}
rpc TouchAccessTime (TouchAccessTimeRequest) returns (TouchAccessTimeResponse) {
}
rpc AppendToEntry (AppendToEntryRequest) returns (AppendToEntryResponse) {
}
@@ -208,6 +211,8 @@ message FuseAttributes {
int32 mtime_ns = 19; // nanosecond component of mtime (0-999999999)
int32 ctime_ns = 20; // nanosecond component of ctime (0-999999999)
int32 crtime_ns = 21; // nanosecond component of crtime (0-999999999)
int64 atime = 22; // unix time in seconds, last access time
int32 atime_ns = 23; // nanosecond component of atime (0-999999999)
}
message CreateEntryRequest {
@@ -247,6 +252,16 @@ message UpdateEntryResponse {
SubscribeMetadataResponse metadata_event = 1;
}
message TouchAccessTimeRequest {
string directory = 1;
string name = 2;
int64 client_atime_ns = 3; // nanoseconds since epoch; filer may override with relatime
}
message TouchAccessTimeResponse {
int64 persisted_atime_ns = 1; // nanoseconds since epoch; 0 if no update was performed
bool updated = 2;
}
message AppendToEntryRequest {
string directory = 1;
string entry_name = 2;
+97
View File
@@ -107,6 +107,57 @@ pub fn build_grpc_endpoint(
Ok(endpoint)
}
/// Parse a SeaweedFS server address (`"ip:port.grpcPort"` or
/// `"ip:port"`) into the `host:grpcPort` form `build_grpc_endpoint`
/// expects. With the trailing `.grpcPort` segment, that segment IS
/// the gRPC port; without it, the gRPC port is `port + 10000`
/// (SeaweedFS's HTTP↔gRPC port-offset convention).
///
/// Shared between `grpc_server.rs` and the distributed-EC-read path
/// in `store_ec.rs` — keep this as the single source of truth so the
/// HTTP↔gRPC port translation can't drift between callers.
pub fn parse_grpc_address(source: &str) -> Result<String, String> {
let colon_idx = source
.rfind(':')
.ok_or_else(|| format!("cannot parse address: {}", source))?;
let host = &source[..colon_idx];
let port_part = &source[colon_idx + 1..];
if let Some(dot_idx) = port_part.rfind('.') {
// Format: "ip:port.grpcPort". Validate BOTH ports as u16
// so a malformed HTTP port (e.g. `host:abc.18080`) is
// rejected here rather than tripping a downstream
// `build_grpc_endpoint` URI parse failure with a less
// useful error.
let http_port = &port_part[..dot_idx];
let grpc_port = &port_part[dot_idx + 1..];
http_port
.parse::<u16>()
.map_err(|e| format!("invalid http port {:?}: {}", http_port, e))?;
grpc_port
.parse::<u16>()
.map_err(|e| format!("invalid grpc port {:?}: {}", grpc_port, e))?;
return Ok(format!("{}:{}", host, grpc_port));
}
// Format: "ip:port" → grpc = port + 10000. Reject inputs whose
// implicit grpc port would overflow the TCP port range (e.g.
// `host:60000` produces 70000 — invalid). Without this check
// the cast silently wraps and the endpoint call later fails
// with an opaque connection error.
let port: u16 = port_part
.parse()
.map_err(|e| format!("invalid port {:?}: {}", port_part, e))?;
let grpc_port = port as u32 + 10000;
if grpc_port > u16::MAX as u32 {
return Err(format!(
"implicit grpc port out of range: {} + 10000 = {}",
port, grpc_port
));
}
Ok(format!("{}:{}", host, grpc_port))
}
#[cfg(test)]
mod tests {
use super::{build_grpc_endpoint, grpc_endpoint_uri, load_outgoing_grpc_tls};
@@ -203,4 +254,50 @@ mod tests {
let endpoint = build_grpc_endpoint("127.0.0.1:19333", None).unwrap();
assert_eq!(endpoint.uri().scheme_str(), Some("http"));
}
#[test]
fn test_parse_grpc_address_dotted_form() {
use super::parse_grpc_address;
assert_eq!(
parse_grpc_address("127.0.0.1:8080.18080").unwrap(),
"127.0.0.1:18080"
);
}
#[test]
fn test_parse_grpc_address_implicit_form_adds_10000() {
use super::parse_grpc_address;
assert_eq!(
parse_grpc_address("127.0.0.1:8080").unwrap(),
"127.0.0.1:18080"
);
}
#[test]
fn test_parse_grpc_address_rejects_non_numeric_http_port_in_dotted_form() {
use super::parse_grpc_address;
let err = parse_grpc_address("host:abc.18080").unwrap_err();
assert!(err.contains("invalid http port"), "{}", err);
}
#[test]
fn test_parse_grpc_address_rejects_non_numeric_grpc_port_in_dotted_form() {
use super::parse_grpc_address;
let err = parse_grpc_address("host:8080.xyz").unwrap_err();
assert!(err.contains("invalid grpc port"), "{}", err);
}
#[test]
fn test_parse_grpc_address_rejects_implicit_port_that_overflows() {
use super::parse_grpc_address;
let err = parse_grpc_address("127.0.0.1:60000").unwrap_err();
assert!(err.contains("out of range"), "{}", err);
}
#[test]
fn test_parse_grpc_address_rejects_input_without_colon() {
use super::parse_grpc_address;
let err = parse_grpc_address("hostname").unwrap_err();
assert!(err.contains("cannot parse"), "{}", err);
}
}
+10 -25
View File
@@ -1511,9 +1511,11 @@ impl VolumeServer for VolumeGrpcService {
// EcVolume holds fds on the same inodes, so overwriting
// corrupts live readers.
if store.has_ec_volume(VolumeId(info.volume_id)) {
let mounted_disks =
store.find_ec_volume_disk_ids(VolumeId(info.volume_id));
resp_error = Some(format!(
"ec volume {} is mounted; unmount before ReceiveFile",
info.volume_id
"ec volume {} is mounted on disk_ids:{:?}; unmount before ReceiveFile",
info.volume_id, mounted_disks
));
break;
}
@@ -4051,29 +4053,12 @@ async fn ping_filer_target(
Ok(resp.into_inner().start_time_ns)
}
/// Parse a SeaweedFS server address ("ip:port.grpcPort" or "ip:port") into a gRPC address.
fn parse_grpc_address(source: &str) -> Result<String, String> {
if let Some(colon_idx) = source.rfind(':') {
let port_part = &source[colon_idx + 1..];
if let Some(dot_idx) = port_part.rfind('.') {
// Format: "ip:port.grpcPort"
let host = &source[..colon_idx];
let grpc_port = &port_part[dot_idx + 1..];
grpc_port
.parse::<u16>()
.map_err(|e| format!("invalid grpc port: {}", e))?;
return Ok(format!("{}:{}", host, grpc_port));
}
// Format: "ip:port" → grpc = port + 10000
let port: u16 = port_part
.parse()
.map_err(|e| format!("invalid port: {}", e))?;
let grpc_port = port as u32 + 10000;
let host = &source[..colon_idx];
return Ok(format!("{}:{}", host, grpc_port));
}
Err(format!("cannot parse address: {}", source))
}
// parse_grpc_address moved to super::grpc_client::parse_grpc_address
// for sharing with the distributed-EC-read path in server/store_ec.rs.
// In-file callers below still write `parse_grpc_address(...)`; this
// `use` makes them resolve to the new home without churning every
// call site.
use super::grpc_client::parse_grpc_address;
/// Set the modification time of a file from nanoseconds since Unix epoch.
fn set_file_mtime(path: &str, modified_ts_ns: i64) {
+32 -31
View File
@@ -1019,43 +1019,44 @@ async fn get_or_head_handler_inner(
if has_ec_volume && !has_volume {
// ---- EC volume read path (always full read, no streaming) ----
let store = state.store.read().unwrap();
match store.find_ec_volume(vid) {
Some(ecv) => match ecv.read_ec_shard_needle(needle_id) {
Ok(Some(ec_needle)) => {
n = ec_needle;
}
Ok(None) => {
metrics::HANDLER_COUNTER
.with_label_values(&[metrics::ERROR_GET_NOT_FOUND])
.inc();
return StatusCode::NOT_FOUND.into_response();
}
Err(e) => {
if e.kind() == std::io::ErrorKind::NotFound {
metrics::HANDLER_COUNTER
.with_label_values(&[metrics::ERROR_GET_NOT_FOUND])
.inc();
return StatusCode::NOT_FOUND.into_response();
}
metrics::HANDLER_COUNTER
.with_label_values(&[metrics::ERROR_GET_INTERNAL])
.inc();
return (
StatusCode::INTERNAL_SERVER_ERROR,
format!("read ec error: {}", e),
)
.into_response();
}
},
None => {
//
// The distributed read path already does a local-first pass
// in its Snapshot phase under the same store read lock the
// legacy code would have taken — so calling it directly
// serves both the "all shards local" fast case and the
// "some intervals need peer fetch + reconstruct" general
// case without paying for the local interval reads twice.
match crate::server::store_ec::read_ec_shard_needle_distributed(
&state, vid, needle_id,
)
.await
{
Ok(Some(ec_needle)) => {
n = ec_needle;
}
Ok(None) => {
metrics::HANDLER_COUNTER
.with_label_values(&[metrics::ERROR_GET_NOT_FOUND])
.inc();
return StatusCode::NOT_FOUND.into_response();
}
Err(e) => {
let kind = if e.kind() == std::io::ErrorKind::NotFound {
metrics::ERROR_GET_NOT_FOUND
} else {
metrics::ERROR_GET_INTERNAL
};
metrics::HANDLER_COUNTER.with_label_values(&[kind]).inc();
if e.kind() == std::io::ErrorKind::NotFound {
return StatusCode::NOT_FOUND.into_response();
}
return (
StatusCode::INTERNAL_SERVER_ERROR,
format!("ec read: {}", e),
)
.into_response();
}
}
drop(store);
// Validate cookie (matches Go behavior after ReadEcShardNeedle)
if n.cookie != cookie {
+6
View File
@@ -344,6 +344,12 @@ fn diff_ec_shard_delta_messages(
if !current.contains_key(key) {
let mut deleted = message.clone();
deleted.shard_sizes = vec![0];
tracing::info!(
volume_id = deleted.id,
disk_id = deleted.disk_id,
ec_index_bits = deleted.ec_index_bits,
"deletes ec shards"
);
deleted_ec_shards.push(deleted);
}
}
+1
View File
@@ -9,6 +9,7 @@ pub mod memory_status;
pub mod profiling;
pub mod request_id;
pub mod server_stats;
pub mod store_ec;
pub mod ui;
pub mod volume_server;
pub mod write_queue;
+678
View File
@@ -0,0 +1,678 @@
//! Distributed EC read path. Mirror of `weed/storage/store_ec.go`'s
//! `readEcShardIntervals` → `readOneEcShardInterval` →
//! `readRemoteEcShardInterval` → `recoverOneRemoteEcShardInterval`
//! chain.
//!
//! The existing `EcVolume::read_ec_shard_needle` reads only locally-
//! mounted shards and returns `NotFound` if any interval requires a
//! shard held on a peer server. In a standard RS(10,4)-across-14
//! deployment each server holds one shard, so every read needs >=9
//! peer fetches. This module fills the gap by:
//!
//! 1. Locating the needle in `.ecx` (under the Store read lock) and
//! computing the per-interval (shard_id, shard_offset, size).
//! 2. Reading the local-resident intervals while still holding the
//! lock — same path the local-only helper uses.
//! 3. Dropping the lock and, for any remaining intervals, fetching
//! from peer volume servers via `VolumeEcShardRead`. If the
//! direct peer read fails, fan-out reads to other shards at the
//! same (shard_offset, size) and rebuild the missing shard via
//! Reed-Solomon — exactly Go's flow.
//! 4. Refreshing the per-EcVolume `shard_locations` cache from the
//! master's `LookupEcVolume` RPC when the cached map is stale.
//!
//! All gRPC IO is async; the file IO portion runs under the sync
//! Store read lock, matching Go's `readLocalEcShardInterval`. The
//! cache write-back briefly reacquires the EcVolume's internal
//! `RwLock` so we do not contend with the Store-level lock at all.
use std::collections::HashMap;
use std::io;
use std::sync::Arc;
use std::time::{Duration, Instant};
use futures::future::join_all;
use reed_solomon_erasure::galois_8::ReedSolomon;
use tonic::Request;
use crate::pb::master_pb::{self, seaweed_client::SeaweedClient, LookupEcVolumeRequest};
use crate::pb::volume_server_pb::{
volume_server_client::VolumeServerClient, VolumeEcShardReadRequest,
};
use crate::server::grpc_client::{build_grpc_endpoint, parse_grpc_address, GRPC_MAX_MESSAGE_SIZE};
use crate::server::request_id::outgoing_request_id_interceptor;
use crate::server::volume_server::VolumeServerState;
use crate::storage::erasure_coding::ec_shard::ShardId;
use crate::storage::needle::needle::{get_actual_size, Needle};
use crate::storage::types::*;
/// One interval's data after Phase A.
enum IntervalResult {
/// Already read from a locally-mounted shard.
Local(Vec<u8>),
/// Shard not local (or local read failed). Must be fetched.
NeedRemote {
shard_id: ShardId,
shard_offset: i64,
size: usize,
},
}
/// Snapshot extracted under the Store read lock so Phases B/C can run
/// without holding any sync lock across `.await`.
struct Snapshot {
data_shards: u32,
parity_shards: u32,
version: Version,
actual_size: usize,
offset: Offset,
size_for_parse: Size,
intervals: Vec<IntervalResult>,
cached_locations: HashMap<ShardId, Vec<String>>,
cache_refreshed_at: Option<Instant>,
}
/// Top-level entry point. Returns `Ok(None)` for "not found" (matches
/// Go's `ReadEcShardNeedle`); errors propagate as `io::Error`.
pub async fn read_ec_shard_needle_distributed(
state: &Arc<VolumeServerState>,
vid: VolumeId,
needle_id: NeedleId,
) -> io::Result<Option<Needle>> {
// Phase A — under the Store read lock, locate the needle, compute
// intervals, and read any locally-mounted shard intervals. We must
// not `.await` while holding this guard (std::sync::RwLockReadGuard
// is !Send).
let snapshot = match snapshot_under_lock(state, vid, needle_id)? {
Some(s) => s,
None => return Ok(None),
};
// Phase B — refresh the shard_locations cache from the master if
// it is stale. Do this lazily: if every needed interval was read
// locally we can skip the master RPC entirely.
let any_remote = snapshot
.intervals
.iter()
.any(|r| matches!(r, IntervalResult::NeedRemote { .. }));
let total_shards = (snapshot.data_shards + snapshot.parity_shards) as usize;
let mut shard_locations = snapshot.cached_locations.clone();
if any_remote
&& needs_refresh(
&shard_locations,
snapshot.cache_refreshed_at,
snapshot.data_shards as usize,
total_shards,
)
{
match cached_lookup_ec_shard_locations(state, vid).await {
Ok(fresh) => {
shard_locations = fresh.clone();
write_back_shard_locations(state, vid, fresh);
}
Err(e) => {
// Lookup failed — proceed with cached values. If cache
// is empty, the remote fetch below will fail and we
// surface a NotFound (matching Go's behavior when no
// locations are known).
tracing::warn!(
"ec lookup failed for volume {}: {} — using cached locations ({} entries)",
vid.0,
e,
shard_locations.len(),
);
}
}
}
// Phase C — fetch missing intervals, reconstructing when the
// direct peer read fails.
let mut assembled: Vec<Vec<u8>> = Vec::with_capacity(snapshot.intervals.len());
for res in snapshot.intervals {
match res {
IntervalResult::Local(buf) => assembled.push(buf),
IntervalResult::NeedRemote {
shard_id,
shard_offset,
size,
} => {
let buf = fetch_one_interval(
state,
vid,
needle_id,
shard_id,
shard_offset,
size,
&shard_locations,
snapshot.data_shards as usize,
snapshot.parity_shards as usize,
)
.await?;
assembled.push(buf);
}
}
}
// Phase D — assemble and parse the Needle. Mirrors the tail of
// `EcVolume::read_ec_shard_needle`.
let mut bytes = Vec::with_capacity(snapshot.actual_size);
for chunk in assembled {
bytes.extend_from_slice(&chunk);
}
bytes.truncate(snapshot.actual_size);
if bytes.len() < snapshot.actual_size {
return Err(io::Error::new(
io::ErrorKind::UnexpectedEof,
format!(
"read {} bytes but need {} for needle {}",
bytes.len(),
snapshot.actual_size,
needle_id
),
));
}
let mut n = Needle::default();
n.id = needle_id;
n.read_bytes(
&bytes,
snapshot.offset.to_actual_offset(),
snapshot.size_for_parse,
snapshot.version,
)
.map_err(|e| io::Error::new(io::ErrorKind::InvalidData, format!("{}", e)))?;
Ok(Some(n))
}
fn snapshot_under_lock(
state: &Arc<VolumeServerState>,
vid: VolumeId,
needle_id: NeedleId,
) -> io::Result<Option<Snapshot>> {
let store = state.store.read().unwrap();
let ecv = match store.find_ec_volume(vid) {
Some(v) => v,
None => return Ok(None),
};
// Reuse EcVolume::locate_needle for offset/size resolution AND
// the per-needle shard-interval math — it's the same routine the
// local-only read path uses, so we stay byte-identical on the
// shard-size + interval boundaries.
let (offset, size, intervals) = match ecv.locate_needle(needle_id)? {
Some(v) => v,
None => return Ok(None),
};
if intervals.is_empty() {
return Err(io::Error::new(
io::ErrorKind::InvalidData,
"no intervals for needle",
));
}
let actual = get_actual_size(size, ecv.version);
// Phase A.local: for each interval read the local shard if we
// hold it; otherwise fall through to remote. We accumulate the
// results in interval order so the assembly step is just a
// concat.
let mut interval_results = Vec::with_capacity(intervals.len());
for interval in &intervals {
let (shard_id, shard_offset) = interval.to_shard_id_and_offset(ecv.data_shards);
let buf_size = interval.size as usize;
let local = ecv
.shards
.get(shard_id as usize)
.and_then(|s| s.as_ref());
match local {
Some(shard) => {
let mut buf = vec![0u8; buf_size];
match shard.read_at(&mut buf, shard_offset as u64) {
Ok(n) if n == buf_size => {
interval_results.push(IntervalResult::Local(buf));
}
_ => interval_results.push(IntervalResult::NeedRemote {
shard_id,
shard_offset,
size: buf_size,
}),
}
}
None => interval_results.push(IntervalResult::NeedRemote {
shard_id,
shard_offset,
size: buf_size,
}),
}
}
let cached_locations = ecv.shard_locations.read().unwrap().clone();
let cache_refreshed_at = *ecv.shard_locations_refresh_time.lock().unwrap();
Ok(Some(Snapshot {
data_shards: ecv.data_shards,
parity_shards: ecv.parity_shards,
version: ecv.version,
actual_size: actual as usize,
offset,
size_for_parse: size,
intervals: interval_results,
cached_locations,
cache_refreshed_at,
}))
}
/// Master `LookupEcVolume` freshness rules — match Go's
/// `cachedLookupEcShardLocations` thresholds in store_ec.go.
fn needs_refresh(
locations: &HashMap<ShardId, Vec<String>>,
refreshed_at: Option<Instant>,
data_shards: usize,
total_shards: usize,
) -> bool {
let now = Instant::now();
let age = match refreshed_at {
Some(t) => now.saturating_duration_since(t),
None => return true,
};
let shard_count = locations.len();
if shard_count < data_shards && age < Duration::from_secs(11) {
return false;
}
if shard_count == total_shards && age < Duration::from_secs(37 * 60) {
return false;
}
if shard_count >= data_shards && age < Duration::from_secs(7 * 60) {
return false;
}
true
}
async fn cached_lookup_ec_shard_locations(
state: &Arc<VolumeServerState>,
vid: VolumeId,
) -> io::Result<HashMap<ShardId, Vec<String>>> {
let master = {
let live = state.current_master_url.read().await.clone();
if !live.is_empty() {
live
} else {
state.master_url.clone()
}
};
if master.is_empty() {
return Err(io::Error::new(
io::ErrorKind::Other,
"no master configured for ec shard lookup",
));
}
let grpc_addr = parse_grpc_address(&master)
.map_err(|e| io::Error::new(io::ErrorKind::InvalidInput, e))?;
let endpoint = build_grpc_endpoint(&grpc_addr, state.outgoing_grpc_tls.as_ref())
.map_err(|e| io::Error::new(io::ErrorKind::Other, e.to_string()))?;
let channel = endpoint
.connect_timeout(Duration::from_secs(5))
.timeout(Duration::from_secs(10))
.connect()
.await
.map_err(|e| io::Error::new(io::ErrorKind::Other, format!("master connect: {}", e)))?;
let mut client = SeaweedClient::with_interceptor(channel, outgoing_request_id_interceptor)
.max_decoding_message_size(GRPC_MAX_MESSAGE_SIZE)
.max_encoding_message_size(GRPC_MAX_MESSAGE_SIZE);
let resp = client
.lookup_ec_volume(Request::new(LookupEcVolumeRequest { volume_id: vid.0 }))
.await
.map_err(|e| io::Error::new(io::ErrorKind::Other, format!("lookup_ec_volume: {}", e)))?;
let resp = resp.into_inner();
let mut out = HashMap::new();
for entry in resp.shard_id_locations {
let addrs: Vec<String> = entry
.locations
.iter()
.map(format_location_as_server_address)
.collect();
out.insert(entry.shard_id as ShardId, addrs);
}
Ok(out)
}
fn write_back_shard_locations(
state: &Arc<VolumeServerState>,
vid: VolumeId,
locations: HashMap<ShardId, Vec<String>>,
) {
let store = state.store.read().unwrap();
if let Some(ecv) = store.find_ec_volume(vid) {
// Atomic swap + freshness stamp so a concurrent reader sees
// either the prior cache or the fresh one — never an
// intermediate half-replaced map with the freshness flag
// already flipped.
ecv.replace_shard_locations(locations);
}
}
/// Build a SeaweedFS-style `host:httpPort.grpcPort` address from a
/// master `Location` so the result is what `parse_grpc_address` (and
/// the heartbeat path) already understand.
fn format_location_as_server_address(loc: &master_pb::Location) -> String {
let raw = loc
.url
.trim_start_matches("http://")
.trim_start_matches("https://");
if loc.grpc_port > 0 {
if let Some((host, http_port)) = raw.rsplit_once(':') {
return format!("{}:{}.{}", host, http_port, loc.grpc_port);
}
}
raw.to_string()
}
/// Try direct peer read; on failure, reconstruct via Reed-Solomon
/// from the other shards. Mirrors `readOneEcShardInterval`'s tail.
async fn fetch_one_interval(
state: &Arc<VolumeServerState>,
vid: VolumeId,
needle_id: NeedleId,
shard_id: ShardId,
shard_offset: i64,
size: usize,
shard_locations: &HashMap<ShardId, Vec<String>>,
data_shards: usize,
parity_shards: usize,
) -> io::Result<Vec<u8>> {
// Direct peer read against the cached locations for this shard.
if let Some(sources) = shard_locations.get(&shard_id) {
if !sources.is_empty() {
match read_remote_ec_shard_interval(
state,
sources,
vid,
needle_id,
shard_id,
shard_offset,
size,
)
.await
{
Ok(buf) => return Ok(buf),
Err(e) => {
tracing::debug!(
"direct read ec shard {}.{} from {:?} failed: {} — will reconstruct",
vid.0,
shard_id,
sources,
e
);
}
}
}
}
// Reconstruct: fan-out reads to every other shard at the same
// (shard_offset, size). Mirrors `recoverOneRemoteEcShardInterval`.
recover_one_remote_ec_shard_interval(
state,
vid,
needle_id,
shard_id,
shard_offset,
size,
shard_locations,
data_shards,
parity_shards,
)
.await
}
async fn read_remote_ec_shard_interval(
state: &Arc<VolumeServerState>,
sources: &[String],
vid: VolumeId,
needle_id: NeedleId,
shard_id: ShardId,
shard_offset: i64,
size: usize,
) -> io::Result<Vec<u8>> {
let mut last_err: Option<io::Error> = None;
for src in sources {
match do_read_remote_ec_shard_interval(
state,
src,
vid,
needle_id,
shard_id,
shard_offset,
size,
)
.await
{
Ok(buf) => return Ok(buf),
Err(e) => last_err = Some(e),
}
}
Err(last_err.unwrap_or_else(|| {
io::Error::new(
io::ErrorKind::NotFound,
format!("no source for ec shard {}.{}", vid.0, shard_id),
)
}))
}
async fn do_read_remote_ec_shard_interval(
state: &Arc<VolumeServerState>,
source: &str,
vid: VolumeId,
needle_id: NeedleId,
shard_id: ShardId,
shard_offset: i64,
size: usize,
) -> io::Result<Vec<u8>> {
let grpc_addr =
parse_grpc_address(source).map_err(|e| io::Error::new(io::ErrorKind::InvalidInput, e))?;
let endpoint = build_grpc_endpoint(&grpc_addr, state.outgoing_grpc_tls.as_ref())
.map_err(|e| io::Error::new(io::ErrorKind::Other, e.to_string()))?;
let channel = endpoint
.connect_timeout(Duration::from_secs(5))
.timeout(Duration::from_secs(30))
.connect()
.await
.map_err(|e| {
io::Error::new(
io::ErrorKind::Other,
format!("connect to {}: {}", source, e),
)
})?;
// TODO(grpc-jwt): clusters with `jwt.signing.key` configured will
// reject peer-to-peer VolumeEcShardRead calls until the Rust
// crate grows an outgoing-JWT interceptor. The gap is shared
// with every other peer gRPC call from this binary
// (`copy_file_from_source`, `batch_delete`, …) — handling it
// here in isolation would split the credential plumbing across
// call sites. Re-visit when outgoing JWT signing lands as a
// server-wide helper.
let mut client = VolumeServerClient::with_interceptor(channel, outgoing_request_id_interceptor)
.max_decoding_message_size(GRPC_MAX_MESSAGE_SIZE)
.max_encoding_message_size(GRPC_MAX_MESSAGE_SIZE);
let req = VolumeEcShardReadRequest {
volume_id: vid.0,
shard_id: shard_id as u32,
offset: shard_offset,
size: size as i64,
file_key: needle_id.0,
};
let resp = client
.volume_ec_shard_read(Request::new(req))
.await
.map_err(|e| {
io::Error::new(
io::ErrorKind::Other,
format!("volume_ec_shard_read {}.{} from {}: {}", vid.0, shard_id, source, e),
)
})?;
let mut stream = resp.into_inner();
let mut out = Vec::with_capacity(size);
while let Some(msg) = stream
.message()
.await
.map_err(|e| io::Error::new(io::ErrorKind::Other, format!("recv: {}", e)))?
{
if !msg.data.is_empty() {
out.extend_from_slice(&msg.data);
}
}
if out.len() < size {
return Err(io::Error::new(
io::ErrorKind::UnexpectedEof,
format!(
"short read from {} for ec shard {}.{}: got {} want {}",
source,
vid.0,
shard_id,
out.len(),
size
),
));
}
out.truncate(size);
Ok(out)
}
async fn recover_one_remote_ec_shard_interval(
state: &Arc<VolumeServerState>,
vid: VolumeId,
needle_id: NeedleId,
shard_id_to_recover: ShardId,
shard_offset: i64,
size: usize,
shard_locations: &HashMap<ShardId, Vec<String>>,
data_shards: usize,
parity_shards: usize,
) -> io::Result<Vec<u8>> {
let total_shards = data_shards + parity_shards;
let rs = ReedSolomon::new(data_shards, parity_shards).map_err(|e| {
io::Error::new(
io::ErrorKind::Other,
format!("reed-solomon init: {:?}", e),
)
})?;
let mut bufs: Vec<Option<Vec<u8>>> = vec![None; total_shards];
// Phase 0: seed bufs from LOCALLY mounted shards. If this node
// already holds enough sibling shards, reconstruction completes
// without any peer fan-out — and even with a cold/incomplete
// shard_locations cache or a failed master lookup, local
// survivors still contribute. Mirrors Go's
// recoverOneRemoteEcShardInterval behaviour, which is implicitly
// local-aware because the Store fan-out targets ALL known
// locations (including the caller's own server address); the
// Rust port had been remote-only, so reconstructing with a cold
// cache failed even when enough siblings were on disk.
{
let store = state.store.read().unwrap();
if let Some(ecv) = store.find_ec_volume(vid) {
for sid in 0..total_shards {
if sid as ShardId == shard_id_to_recover {
continue;
}
if let Some(Some(shard)) = ecv.shards.get(sid) {
let mut buf = vec![0u8; size];
if shard.read_at(&mut buf, shard_offset as u64).map(|n| n == size).unwrap_or(false) {
bufs[sid] = Some(buf);
}
}
}
}
}
// Phase 1: remote fan-out — one task per known shard location
// we DON'T already have locally and DON'T need to recover.
let mut tasks = Vec::new();
for (sid, locs) in shard_locations {
if *sid == shard_id_to_recover || locs.is_empty() {
continue;
}
if bufs[*sid as usize].is_some() {
continue;
}
let sid = *sid;
let locs = locs.clone();
let state = state.clone();
tasks.push(async move {
let res = read_remote_ec_shard_interval(
&state,
&locs,
vid,
needle_id,
sid,
shard_offset,
size,
)
.await;
(sid, res)
});
}
let results = join_all(tasks).await;
for (sid, res) in results {
match res {
Ok(buf) => {
if (sid as usize) < total_shards {
bufs[sid as usize] = Some(buf);
}
}
Err(e) => {
tracing::debug!(
"recover: read {}.{} for needle {} failed: {}",
vid.0,
sid,
needle_id,
e
);
}
}
}
let available = bufs.iter().filter(|b| b.is_some()).count();
if available < data_shards {
return Err(io::Error::new(
io::ErrorKind::Other,
format!(
"cannot recover ec shard {}.{}: only {} shards available, need at least {}",
vid.0, shard_id_to_recover, available, data_shards
),
));
}
rs.reconstruct(&mut bufs).map_err(|e| {
io::Error::new(
io::ErrorKind::Other,
format!(
"reed-solomon reconstruct ec shard {}.{}: {:?}",
vid.0, shard_id_to_recover, e
),
)
})?;
match bufs.into_iter().nth(shard_id_to_recover as usize).flatten() {
Some(buf) => Ok(buf),
None => Err(io::Error::new(
io::ErrorKind::Other,
format!(
"reconstructed buffer for shard {}.{} missing after RS reconstruct",
vid.0, shard_id_to_recover
),
)),
}
}
// parse_grpc_address lives in `grpc_client.rs` and is re-exported
// here via the use above so this module shares a single
// HTTP↔gRPC port-translation routine with grpc_server.rs.
@@ -44,8 +44,17 @@ pub struct EcVolume {
/// Directory where .ecx/.ecj were actually found (may differ from dir_idx after fallback).
ecx_actual_dir: String,
/// Maps shard ID -> list of server addresses where that shard exists.
/// Used for distributed EC reads across the cluster.
pub shard_locations: HashMap<ShardId, Vec<String>>,
/// Used for distributed EC reads across the cluster. Wrapped in
/// `RwLock` so the read path can refresh the map (under master
/// lookup) without holding the Store write lock — mirrors Go's
/// `ShardLocationsLock sync.RWMutex` in `weed/storage/erasure_coding/ec_volume.go`.
pub shard_locations: std::sync::RwLock<HashMap<ShardId, Vec<String>>>,
/// Wall-clock timestamp of the most recent successful
/// `LookupEcVolume` refresh of `shard_locations`. `None` until the
/// first refresh. Drives the staleness heuristic in
/// `cached_lookup_ec_shard_locations` (mirrors Go's
/// `ShardLocationsRefreshTime`).
pub shard_locations_refresh_time: std::sync::Mutex<Option<std::time::Instant>>,
/// EC volume expiration time (unix epoch seconds), set during EC encode from TTL.
pub expire_at_sec: u64,
}
@@ -165,7 +174,8 @@ impl EcVolume {
deleted_needles: RwLock::new(HashSet::new()),
disk_type: DiskType::default(),
ecx_actual_dir: dir_idx.to_string(),
shard_locations: HashMap::new(),
shard_locations: std::sync::RwLock::new(HashMap::new()),
shard_locations_refresh_time: std::sync::Mutex::new(None),
expire_at_sec,
};
@@ -430,17 +440,45 @@ impl EcVolume {
// ---- Shard locations (distributed tracking) ----
/// Set the list of server addresses for a given shard ID.
pub fn set_shard_locations(&mut self, shard_id: ShardId, locations: Vec<String>) {
self.shard_locations.insert(shard_id, locations);
/// Set the list of server addresses for a single shard ID. Does
/// NOT touch `shard_locations_refresh_time` — a per-shard write
/// from inside a multi-shard population (e.g. iterating the
/// `LookupEcVolume` response shard-by-shard) would otherwise
/// flip the staleness flag while the map is still incomplete,
/// letting a concurrent reader observe `needs_refresh == false`
/// against a half-populated cache and return NotFound for the
/// not-yet-inserted shards.
///
/// Callers populating the whole cache atomically should use
/// [`Self::replace_shard_locations`] instead — it swaps the
/// entire map under the write lock and advances the refresh
/// timestamp in one step.
pub fn set_shard_locations(&self, shard_id: ShardId, locations: Vec<String>) {
self.shard_locations
.write()
.unwrap()
.insert(shard_id, locations);
}
/// Get the list of server addresses for a given shard ID.
pub fn get_shard_locations(&self, shard_id: ShardId) -> &[String] {
/// Atomically replace the entire shard-locations map and stamp
/// the refresh time. Used by the distributed-read path's
/// post-`LookupEcVolume` write-back so the cache transitions
/// from old → fresh in a single observable step — concurrent
/// readers either see the full prior map or the full new map,
/// never an intermediate state with the freshness flag flipped.
pub fn replace_shard_locations(&self, locations: HashMap<ShardId, Vec<String>>) {
*self.shard_locations.write().unwrap() = locations;
*self.shard_locations_refresh_time.lock().unwrap() = Some(std::time::Instant::now());
}
/// Get a cloned list of server addresses for a given shard ID.
pub fn get_shard_locations(&self, shard_id: ShardId) -> Vec<String> {
self.shard_locations
.read()
.unwrap()
.get(&shard_id)
.map(|v| v.as_slice())
.unwrap_or(&[])
.cloned()
.unwrap_or_default()
}
// ---- Index operations ----
+1
View File
@@ -4,6 +4,7 @@ pub mod idx;
pub mod needle;
pub mod needle_map;
pub mod store;
pub mod store_ec_mirror;
pub mod store_ec_reconcile;
pub mod super_block;
pub mod types;
+20
View File
@@ -212,6 +212,13 @@ impl CompactNeedleMap {
self.idx_file_offset = offset;
}
/// True when an .idx file writer is attached. A read-only load leaves
/// this `false` — set_writable() must reattach a writer or subsequent
/// puts silently skip the disk append.
pub fn has_idx_writer(&self) -> bool {
self.idx_file.is_some()
}
// ---- Map operations ----
/// Insert or update an entry. Appends to .idx file if present.
@@ -705,6 +712,11 @@ impl RedbNeedleMap {
self.idx_file_offset = offset;
}
/// True when an .idx file writer is attached. See CompactNeedleMap.
pub fn has_idx_writer(&self) -> bool {
self.idx_file.is_some()
}
// ---- Map operations ----
/// Insert or update an entry. Writes to idx file first, then redb.
@@ -1000,6 +1012,14 @@ impl NeedleMap {
}
}
/// True when an .idx file writer is attached.
pub fn has_idx_writer(&self) -> bool {
match self {
NeedleMap::InMemory(nm) => nm.has_idx_writer(),
NeedleMap::Redb(nm) => nm.has_idx_writer(),
}
}
/// Content byte count.
pub fn content_size(&self) -> u64 {
match self {
+39 -11
View File
@@ -96,14 +96,16 @@ impl Store {
// re-loading shards we just cleaned up.
self.prune_incomplete_ec_with_sibling_dat();
// After every disk has finished its per-disk EC scan, sweep
// the store for shards that live on a disk without local index
// files and load them by reaching across to a sibling disk's
// .ecx / .ecj / .vif (seaweedfs/seaweedfs#9212 / #9244).
// ec.balance / ec.rebuild can move shards onto a destination
// node's second disk while leaving the index on the disk that
// already held the volume; without this pass those orphan
// shards stay invisible to the master.
// Physically mirror EC sidecars onto every shard-bearing disk
// so each disk mounts self-contained. Must run before the
// cross-disk reconciler so the orphan pass can prefer the
// local idx_directory.
self.mirror_ec_metadata_to_shard_disks();
// Cross-disk fallback for orphan shards — ec.balance can land
// shards on one disk while leaving the index on another. Still
// needed after the mirror pass for volumes whose mirror failed
// (read-only target, partial copy).
self.reconcile_ec_shards_across_disks();
Ok(())
@@ -118,6 +120,7 @@ impl Store {
}
}
self.prune_incomplete_ec_with_sibling_dat();
self.mirror_ec_metadata_to_shard_disks();
self.reconcile_ec_shards_across_disks();
}
@@ -699,10 +702,20 @@ impl Store {
// Walk all locations rather than stopping at the first with the
// vid — split-disk reconciled volumes can have the same vid on
// multiple disks, with the target shard on any of them.
for loc in &mut self.locations {
if loc.has_ec_volume(vid) {
loc.unmount_ec_shards(vid, &[shard_id]);
for disk_id in 0..self.locations.len() {
let has_shard = self.locations[disk_id]
.find_ec_volume(vid)
.is_some_and(|ec_vol| ec_vol.has_shard(shard_id as u8));
if !has_shard {
continue;
}
tracing::info!(
volume_id = vid.0,
shard_id,
disk_id,
"UnmountEcShards"
);
self.locations[disk_id].unmount_ec_shards(vid, &[shard_id]);
}
// Go returns nil if shard not found (no error)
Ok(())
@@ -733,6 +746,21 @@ impl Store {
self.locations.iter().any(|loc| loc.has_ec_volume(vid))
}
/// Returns every disk_id on this store that has an EcVolume entry
/// for `vid`. Useful for diagnostic logging when a single
/// `has_ec_volume` hit hides which disk is actually holding the
/// mount (e.g., the ReceiveFile mounted-volume guard).
/// Mirrors Go's `Store.FindEcVolumeDiskIds`.
pub fn find_ec_volume_disk_ids(&self, vid: VolumeId) -> Vec<u32> {
let mut ids = Vec::new();
for (idx, loc) in self.locations.iter().enumerate() {
if loc.has_ec_volume(vid) {
ids.push(idx as u32);
}
}
ids
}
/// Returns the index of the disk location that has `(vid, shard_id)`
/// mounted, if any. Mirrors Go's `Store.findEcShard` and is the
/// right primitive for read/unmount/delete operations on a single
@@ -0,0 +1,421 @@
//! Physical EC sidecar mirroring across disks of the same volume
//! server. Mirrors `weed/storage/store_ec_mirror.go`.
use std::collections::HashMap;
use std::fs;
use std::io::{self, Read, Write};
use std::path::Path;
use tracing::{info, warn};
use crate::storage::disk_location::{parse_collection_volume_id_pub, DiskLocation};
use crate::storage::store::Store;
use crate::storage::types::VolumeId;
// Listed in `EcVolume::new`'s open order.
const EC_MIRRORED_SIDECARS: &[&str] = &[".ecx", ".ecj", ".vif"];
#[derive(Clone, Debug, PartialEq, Eq, Hash)]
struct EcKey {
collection: String,
vid: VolumeId,
}
#[derive(Clone, Debug)]
struct EcxOwner {
location: usize,
idx_dir: String,
data_dir: String,
}
impl Store {
/// Mirror EC sidecars onto every shard-bearing disk that lacks
/// them, so each disk mounts self-contained. Runs before
/// `reconcile_ec_shards_across_disks` so the orphan pass can
/// prefer the local idx_directory.
pub fn mirror_ec_metadata_to_shard_disks(&mut self) {
if self.locations.len() < 2 {
return;
}
let owners = self.index_ecx_owners_for_mirror();
if owners.is_empty() {
return;
}
// Two-pass: gather work under an immutable borrow, then apply
// copies under independent mutable borrows.
struct Mirror<'a> {
target_idx: usize,
owner: &'a EcxOwner,
collection: String,
vid: VolumeId,
}
let mut mirrors: Vec<Mirror> = Vec::new();
for (loc_idx, loc) in self.locations.iter().enumerate() {
let orphans = collect_shard_disk_volumes(loc);
for (key, _shards) in orphans {
let Some(owner) = owners.get(&key) else {
continue;
};
if owner.location == loc_idx {
continue;
}
if disk_has_all_sidecars(loc, &key.collection, key.vid) {
continue;
}
mirrors.push(Mirror {
target_idx: loc_idx,
owner,
collection: key.collection,
vid: key.vid,
});
}
}
for m in mirrors {
let loc_dir = self.locations[m.target_idx].directory.clone();
let loc_idx_dir = self.locations[m.target_idx].idx_directory.clone();
match mirror_sidecars_for_volume(
&m.owner.idx_dir,
&m.owner.data_dir,
&loc_dir,
&loc_idx_dir,
&m.collection,
m.vid,
) {
Ok(0) => {}
Ok(copied) => {
info!(
volume_id = m.vid.0,
collection = %m.collection,
from = %m.owner.data_dir,
to = %loc_dir,
copied,
"mirrored EC sidecar(s) for same-disk invariant",
);
}
Err(e) => {
warn!(
volume_id = m.vid.0,
collection = %m.collection,
from = %m.owner.data_dir,
to = %loc_dir,
error = %e,
"mirror EC sidecars failed; cross-disk fallback will handle this volume",
);
}
}
}
}
// Records both idx_dir (where .ecx was found) and data_dir, so
// the mirror can resolve .vif from data_dir even when .ecx lives
// in idx_directory.
fn index_ecx_owners_for_mirror(&self) -> HashMap<EcKey, EcxOwner> {
let mut owners: HashMap<EcKey, EcxOwner> = HashMap::new();
for (loc_idx, loc) in self.locations.iter().enumerate() {
let mut seen: Vec<&str> = Vec::with_capacity(2);
for scan in [loc.idx_directory.as_str(), loc.directory.as_str()] {
if scan.is_empty() || seen.contains(&scan) {
continue;
}
seen.push(scan);
let Ok(read) = fs::read_dir(scan) else {
continue;
};
for ent in read.flatten() {
if ent.file_type().map(|ft| ft.is_dir()).unwrap_or(false) {
continue;
}
let name = ent.file_name().to_string_lossy().into_owned();
let Some(base) = name.strip_suffix(".ecx") else {
continue;
};
let Some((collection, vid)) = parse_collection_volume_id_pub(base) else {
continue;
};
owners
.entry(EcKey { collection, vid })
.or_insert_with(|| EcxOwner {
location: loc_idx,
idx_dir: scan.to_string(),
data_dir: loc.directory.clone(),
});
}
}
}
owners
}
}
// Checks the modern routing and the opposite directory — without
// that fallback, a destination with a legacy pre-`-dir.idx` .ecx in
// its data dir would be re-mirrored into idx_directory.
fn disk_has_all_sidecars(loc: &DiskLocation, collection: &str, vid: VolumeId) -> bool {
for ext in EC_MIRRORED_SIDECARS {
let primary = sidecar_dest_path(&loc.directory, &loc.idx_directory, collection, vid, ext);
if path_is_regular_file(&primary) {
continue;
}
if loc.idx_directory != loc.directory {
let (fallback_data, fallback_idx) = if *ext == ".vif" {
(loc.idx_directory.as_str(), loc.directory.as_str())
} else {
(loc.directory.as_str(), loc.directory.as_str())
};
let fallback = sidecar_dest_path(fallback_data, fallback_idx, collection, vid, ext);
if path_is_regular_file(&fallback) {
continue;
}
}
return false;
}
true
}
fn path_is_regular_file(path: &str) -> bool {
fs::metadata(path).map(|m| !m.is_dir()).unwrap_or(false)
}
// `.ecx`/`.ecj` route to idx_directory, `.vif` to directory.
fn sidecar_dest_path(
data_dir: &str,
idx_dir: &str,
collection: &str,
vid: VolumeId,
ext: &str,
) -> String {
let dir = if ext == ".vif" { data_dir } else { idx_dir };
if collection.is_empty() {
format!("{}/{}{}", dir, vid.0, ext)
} else {
format!("{}/{}_{}{}", dir, collection, vid.0, ext)
}
}
fn mirror_sidecars_for_volume(
src_idx_dir: &str,
src_data_dir: &str,
dst_data_dir: &str,
dst_idx_dir: &str,
collection: &str,
vid: VolumeId,
) -> io::Result<usize> {
let mut copied = 0usize;
for ext in EC_MIRRORED_SIDECARS {
let dst = sidecar_dest_path(dst_data_dir, dst_idx_dir, collection, vid, ext);
// An existing local copy is authoritative — it may be newer
// than the owner's after a delete journal append.
if fs::metadata(&dst).is_ok() {
continue;
}
let candidates = [
sidecar_dest_path(src_data_dir, src_idx_dir, collection, vid, ext),
sidecar_dest_path(src_idx_dir, src_data_dir, collection, vid, ext),
];
let mut src_path: Option<String> = None;
for c in candidates.iter() {
if fs::metadata(c).map(|m| !m.is_dir()).unwrap_or(false) {
src_path = Some(c.clone());
break;
}
}
let Some(src) = src_path else {
continue;
};
copy_sidecar_atomic(Path::new(&src), Path::new(&dst))?;
copied += 1;
}
Ok(copied)
}
fn copy_sidecar_atomic(src: &Path, dst: &Path) -> io::Result<()> {
if let Some(parent) = dst.parent() {
fs::create_dir_all(parent)?;
}
let mut src_file = fs::File::open(src)?;
let tmp = {
let mut s = dst.as_os_str().to_owned();
s.push(".mirror.tmp");
std::path::PathBuf::from(s)
};
let _ = fs::remove_file(&tmp);
let mut dst_file = fs::OpenOptions::new()
.write(true)
.create_new(true)
.open(&tmp)?;
let mut buf = [0u8; 64 * 1024];
loop {
let n = src_file.read(&mut buf)?;
if n == 0 {
break;
}
dst_file.write_all(&buf[..n])?;
}
dst_file.sync_all()?;
drop(dst_file);
if let Err(e) = fs::rename(&tmp, dst) {
let _ = fs::remove_file(&tmp);
return Err(e);
}
Ok(())
}
fn collect_shard_disk_volumes(loc: &DiskLocation) -> HashMap<EcKey, Vec<String>> {
let mut out: HashMap<EcKey, Vec<String>> = HashMap::new();
let Ok(read) = fs::read_dir(&loc.directory) else {
return out;
};
for ent in read.flatten() {
if ent.file_type().map(|ft| ft.is_dir()).unwrap_or(false) {
continue;
}
let name = ent.file_name().to_string_lossy().into_owned();
let Some(dot) = name.rfind('.') else {
continue;
};
let (base, ext) = name.split_at(dot);
if crate::storage::disk_location::is_ec_shard_extension(ext).is_none() {
continue;
}
match ent.metadata() {
Ok(meta) if meta.len() > 0 => {}
_ => continue,
}
let Some((collection, vid)) = parse_collection_volume_id_pub(base) else {
continue;
};
out.entry(EcKey { collection, vid })
.or_default()
.push(name);
}
out
}
#[cfg(test)]
mod tests {
use super::*;
use crate::config::MinFreeSpace;
use crate::storage::needle_map::NeedleMapKind;
use crate::storage::types::DiskType;
use crate::storage::volume::{VifEcShardConfig, VifVolumeInfo};
use tempfile::TempDir;
fn plant_shard(dir: &Path, collection: &str, vid: u32, shard_id: u8) {
let path = if collection.is_empty() {
dir.join(format!("{}.ec{:02}", vid, shard_id))
} else {
dir.join(format!("{}_{}.ec{:02}", collection, vid, shard_id))
};
fs::write(&path, b"shard data nonempty").unwrap();
}
fn plant_ecx(dir: &Path, collection: &str, vid: u32, bytes: &[u8]) {
let path = dir.join(format!("{}_{}.ecx", collection, vid));
fs::write(&path, bytes).unwrap();
}
fn plant_ecj(dir: &Path, collection: &str, vid: u32, bytes: &[u8]) {
let path = dir.join(format!("{}_{}.ecj", collection, vid));
fs::write(&path, bytes).unwrap();
}
fn plant_vif(dir: &Path, collection: &str, vid: u32, data_shards: u32, parity_shards: u32) {
let vif = VifVolumeInfo {
version: 3,
ec_shard_config: Some(VifEcShardConfig {
data_shards,
parity_shards,
}),
..Default::default()
};
let path = dir.join(format!("{}_{}.vif", collection, vid));
fs::write(&path, serde_json::to_string(&vif).unwrap()).unwrap();
}
fn add_loc(store: &mut Store, dir: &Path) {
store
.add_location(
dir.to_str().unwrap(),
dir.to_str().unwrap(),
100,
DiskType::HardDrive,
MinFreeSpace::Percent(0.0),
Vec::new(),
)
.unwrap();
}
#[test]
fn mirror_copies_sidecars_to_shard_only_disk() {
let tmp = TempDir::new().unwrap();
let dir0 = tmp.path().join("data0");
let dir1 = tmp.path().join("data1");
fs::create_dir_all(&dir0).unwrap();
fs::create_dir_all(&dir1).unwrap();
let collection = "video-recordings";
let vid = 4121u32;
plant_shard(&dir0, collection, vid, 0);
plant_shard(&dir0, collection, vid, 12);
plant_shard(&dir1, collection, vid, 1);
let ecx = vec![0xA1u8; 20];
let ecj = vec![0xB2u8; 16];
plant_ecx(&dir1, collection, vid, &ecx);
plant_ecj(&dir1, collection, vid, &ecj);
plant_vif(&dir1, collection, vid, 10, 4);
let mut store = Store::new(NeedleMapKind::InMemory);
add_loc(&mut store, &dir0);
add_loc(&mut store, &dir1);
for ext in [".ecx", ".ecj", ".vif"] {
let dst = dir0.join(format!("{}_{}{}", collection, vid, ext));
assert!(
dst.exists(),
"mirror did not install sidecar {} on dir0",
ext
);
}
let ecx_dst = fs::read(dir0.join(format!("{}_{}.ecx", collection, vid))).unwrap();
assert_eq!(ecx_dst, ecx, ".ecx mirrored bytes differ from source");
let ecj_dst = fs::read(dir0.join(format!("{}_{}.ecj", collection, vid))).unwrap();
assert_eq!(ecj_dst, ecj, ".ecj mirrored bytes differ from source");
}
#[test]
fn mirror_preserves_existing_destination_sidecars() {
let tmp = TempDir::new().unwrap();
let dir0 = tmp.path().join("data0");
let dir1 = tmp.path().join("data1");
fs::create_dir_all(&dir0).unwrap();
fs::create_dir_all(&dir1).unwrap();
let collection = "video-recordings";
let vid = 7777u32;
plant_shard(&dir0, collection, vid, 0);
plant_shard(&dir0, collection, vid, 12);
plant_shard(&dir1, collection, vid, 1);
let ecx_owner = vec![0xC3u8; 20];
let ecx_local = vec![0x5Au8; 20];
let ecj_bytes = vec![0xD4u8; 16];
plant_ecx(&dir1, collection, vid, &ecx_owner);
plant_ecj(&dir1, collection, vid, &ecj_bytes);
plant_vif(&dir1, collection, vid, 10, 4);
plant_ecx(&dir0, collection, vid, &ecx_local);
plant_ecj(&dir0, collection, vid, &ecj_bytes);
plant_vif(&dir0, collection, vid, 10, 4);
let mut store = Store::new(NeedleMapKind::InMemory);
add_loc(&mut store, &dir0);
add_loc(&mut store, &dir1);
let post = fs::read(dir0.join(format!("{}_{}.ecx", collection, vid))).unwrap();
assert_eq!(post, ecx_local, "mirror overwrote dir0's existing .ecx");
}
}
@@ -26,6 +26,14 @@ use crate::storage::store::Store;
use crate::storage::super_block::SUPER_BLOCK_SIZE;
use crate::storage::types::VolumeId;
pub(crate) fn ec_local_ecx_path(dir: &str, collection: &str, vid: VolumeId) -> String {
if collection.is_empty() {
format!("{}/{}.ecx", dir, vid.0)
} else {
format!("{}/{}_{}.ecx", dir, collection, vid.0)
}
}
/// Sibling-disk `.dat` candidate for `prune_incomplete_ec_with_sibling_dat`.
/// We record both the disk index and the file's size: the size is
/// consulted before deleting any EC artefacts. A zero-byte or truncated
@@ -76,10 +84,10 @@ impl Store {
return;
}
// Snapshot of orphan shards, keyed by (loc_idx, ec_key) so we
// can release the immutable borrow on self.locations before
// calling mount_ec_shards_with_idx_dir (which needs &mut).
let mut to_load: Vec<(usize, EcKey, Vec<(String, u32)>, EcxOwnerInfo)> = Vec::new();
// `use_local_idx` is the post-mirror fast path: when the
// mirror already installed sidecars locally, mount against
// loc.idx_directory instead of the owner disk.
let mut to_load: Vec<(usize, EcKey, Vec<(String, u32)>, EcxOwnerInfo, bool)> = Vec::new();
for (loc_idx, loc) in self.locations.iter().enumerate() {
let orphans = collect_orphan_ec_shards(loc, loc_idx);
for (key, shards) in orphans {
@@ -93,34 +101,57 @@ impl Store {
);
continue;
};
if owner.location == loc_idx && owner.idx_dir == loc.idx_directory {
// Normal same-disk case: load_all_ec_shards already
// attempted the mount via `loc.idx_directory` and
// logged the underlying failure. No point retrying
// the same call.
let local_ecx = ec_local_ecx_path(&loc.idx_directory, &key.collection, key.vid);
let local_ecx_in_data = ec_local_ecx_path(&loc.directory, &key.collection, key.vid);
let use_local_idx = std::path::Path::new(&local_ecx).exists()
|| std::path::Path::new(&local_ecx_in_data).exists();
if !use_local_idx
&& owner.location == loc_idx
&& owner.idx_dir == loc.idx_directory
{
// Same-disk no-op: load_all_ec_shards already
// tried and logged the failure.
continue;
}
// Either a cross-disk owner OR a same-disk owner whose
// `.ecx` actually lives in `loc.directory` (the legacy
// pre-`-dir.idx` layout). The latter wasn't tried by
// load_all_ec_shards, which only looked in
// `self.idx_directory`, so we still need to retry it
// here with the owner's discovered idx_dir.
to_load.push((loc_idx, key, shards, owner.clone()));
to_load.push((loc_idx, key, shards, owner.clone(), use_local_idx));
}
}
for (loc_idx, key, shards, owner) in to_load {
for (loc_idx, key, shards, owner, use_local_idx) in to_load {
let shard_names: Vec<&str> = shards.iter().map(|(n, _)| n.as_str()).collect();
let loc_dir = self.locations[loc_idx].directory.clone();
let shard_ids: Vec<u32> = shards.iter().map(|(_, sid)| *sid).collect();
if use_local_idx {
info!(
volume_id = key.vid.0,
collection = %key.collection,
directory = %loc_dir,
"loading orphan EC shards against locally-mirrored sidecars: {:?}",
shard_names,
);
let loc = &mut self.locations[loc_idx];
if let Err(e) = loc.mount_ec_shards(key.vid, &key.collection, &shard_ids, "") {
loc.unmount_ec_shards(key.vid, &shard_ids);
warn!(
volume_id = key.vid.0,
directory = %loc_dir,
"local-mirror shard load failed: {}",
e,
);
}
continue;
}
info!(
volume_id = key.vid.0,
collection = %key.collection,
from = %self.locations[owner.location].directory,
to = %self.locations[loc_idx].directory,
to = %loc_dir,
"loading orphan EC shards using index files from sibling disk (issue #9212): {:?}",
shard_names,
);
let shard_ids: Vec<u32> = shards.iter().map(|(_, sid)| *sid).collect();
let owner_idx_dir = owner.idx_dir.clone();
let loc = &mut self.locations[loc_idx];
if let Err(e) = loc.mount_ec_shards_with_idx_dir(
@@ -141,7 +172,7 @@ impl Store {
loc.unmount_ec_shards(key.vid, &shard_ids);
warn!(
volume_id = key.vid.0,
directory = %loc.directory,
directory = %loc_dir,
"cross-disk shard load failed: {}",
e,
);
+224
View File
@@ -2122,7 +2122,34 @@ impl Volume {
}
/// Mark this volume as writable (allow writes and deletes).
///
/// If the volume booted with .vif ReadOnly=true, `load_index` built the
/// needle map without an .idx writer attached, so subsequent puts would
/// silently skip the on-disk append and only mutate in-memory state —
/// surviving until the next restart, then vanishing. Re-attach a writer
/// here so writes persist again.
pub fn set_writable(&mut self) -> Result<(), VolumeError> {
// Attach the writer (if missing) before flipping the flag — otherwise
// a transient open/metadata failure would leave the volume marked
// writable with no .idx writer, and subsequent puts would silently
// skip the on-disk append and vanish on the next restart.
let needs_idx_writer = self
.nm
.as_ref()
.map(|nm| !nm.has_idx_writer())
.unwrap_or(false);
if needs_idx_writer {
let idx_path = self.file_name(".idx");
let write_file = OpenOptions::new()
.write(true)
.append(true)
.create(true)
.open(&idx_path)?;
let idx_size = write_file.metadata()?.len();
if let Some(ref mut nm) = self.nm {
nm.set_idx_file(Box::new(write_file), idx_size);
}
}
self.no_write_or_delete = false;
self.save_vif()
}
@@ -3565,6 +3592,118 @@ mod tests {
assert!(matches!(err, VolumeError::Deleted));
}
// Guard the Rust integrity-check tombstone path against the Go regression
// where verifyDeletedNeedleIntegrity forwarded TombstoneFileSize into the
// needle-size check, mismatched against the on-disk Size=0 header, and
// sent every volume with a trailing deletion read-only on load. The Rust
// check guards its size comparison with !size.is_deleted(); this test
// keeps that guarantee from silently regressing.
#[test]
fn test_check_volume_data_integrity_with_deletion_tombstone() {
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
{
let mut v = make_test_volume(dir);
for i in 1..=3 {
let data = format!("data {}", i);
let mut n = Needle {
id: NeedleId(i),
cookie: Cookie(i as u32),
data: data.as_bytes().to_vec(),
data_size: data.len() as u32,
..Needle::default()
};
v.write_needle(&mut n, true).unwrap();
}
v.delete_needle(&mut Needle {
id: NeedleId(2),
cookie: Cookie(2),
..Needle::default()
})
.unwrap();
v.sync_to_disk().unwrap();
}
let v = Volume::new(
dir,
dir,
"",
VolumeId(1),
NeedleMapKind::InMemory,
None,
None,
0,
Version::current(),
)
.unwrap();
assert!(
!v.is_no_write_or_delete(),
"volume should not be read-only after reload with trailing deletion tombstone"
);
}
#[test]
fn test_scrub_empty_volume() {
// Mirror of Go's TestScrubVolumeData "zero-size volume without index"
// case (weed/storage/volume_checking_test.go): a freshly created /
// pre-allocated volume has a superblock-only .dat and a zero-size .idx,
// and must scrub clean instead of being flagged as corrupt.
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
let v = make_test_volume(dir);
// .dat holds only the superblock; .idx is empty.
assert_eq!(v.dat_file_size().unwrap(), SUPER_BLOCK_SIZE as u64);
let (files_checked, broken) = v.scrub().unwrap();
assert_eq!(files_checked, 0);
assert!(
broken.is_empty(),
"empty volume should scrub clean, got {:?}",
broken
);
// The index-only mode must agree.
let (idx_checked, idx_broken) = v.scrub_index().unwrap();
assert_eq!(idx_checked, 0);
assert!(
idx_broken.is_empty(),
"empty volume should scrub_index clean, got {:?}",
idx_broken
);
}
#[test]
fn test_scrub_healthy_volume() {
// Mirror of Go's TestScrubVolumeData "healthy volume" case: a volume
// with live needles scrubs clean and the .dat size accounting matches.
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
let mut v = make_test_volume(dir);
for i in 1..=5 {
let data = format!("needle data {}", i);
let mut n = Needle {
id: NeedleId(i),
cookie: Cookie(i as u32),
data: data.as_bytes().to_vec(),
data_size: data.len() as u32,
..Needle::default()
};
v.write_needle(&mut n, true).unwrap();
}
v.sync_to_disk().unwrap();
let (files_checked, broken) = v.scrub().unwrap();
assert_eq!(files_checked, 5);
assert!(
broken.is_empty(),
"healthy volume should scrub clean, got {:?}",
broken
);
}
#[test]
fn test_volume_multiple_needles() {
let tmp = TempDir::new().unwrap();
@@ -4191,6 +4330,91 @@ mod tests {
assert!(v.no_write_can_delete);
}
// A volume booted with .vif ReadOnly=true used to come back stuck —
// load_index_inmemory built the CompactNeedleMap without an .idx writer
// attached, and set_writable only flipped the flag and rewrote .vif.
// The next put silently skipped the .idx append, so the write landed in
// memory only and was lost on the next restart.
#[test]
fn test_set_writable_reattaches_idx_writer_after_persisted_readonly() {
let tmp = TempDir::new().unwrap();
let dir = tmp.path().to_str().unwrap();
{
let mut v = make_test_volume(dir);
let mut n = Needle {
id: NeedleId(1),
cookie: Cookie(1),
data: b"initial".to_vec(),
data_size: 7,
..Needle::default()
};
v.write_needle(&mut n, true).unwrap();
v.set_read_only_persist(true).unwrap();
v.sync_to_disk().unwrap();
}
let mut v = Volume::new(
dir,
dir,
"",
VolumeId(1),
NeedleMapKind::InMemory,
None,
None,
0,
Version::current(),
)
.unwrap();
assert!(
v.no_write_or_delete,
"reloaded volume should be read-only from .vif"
);
assert!(
!v.nm.as_ref().unwrap().has_idx_writer(),
"read-only load should not attach an .idx writer"
);
v.set_writable().unwrap();
assert!(!v.is_read_only());
assert!(
v.nm.as_ref().unwrap().has_idx_writer(),
"set_writable must reattach the .idx writer or post-restart writes vanish"
);
let mut n = Needle {
id: NeedleId(2),
cookie: Cookie(2),
data: b"after-mark-writable".to_vec(),
data_size: 19,
..Needle::default()
};
v.write_needle(&mut n, true).unwrap();
v.sync_to_disk().unwrap();
// Reload one more time — the .idx must contain the post-mark-writable
// entry, not just have it in memory.
drop(v);
let v = Volume::new(
dir,
dir,
"",
VolumeId(1),
NeedleMapKind::InMemory,
None,
None,
0,
Version::current(),
)
.unwrap();
let mut probe = Needle {
id: NeedleId(2),
..Needle::default()
};
v.read_needle(&mut probe).unwrap();
assert_eq!(std::str::from_utf8(&probe.data).unwrap(), "after-mark-writable");
}
#[test]
fn test_load_vif_defaults_local_version_and_bytes_offset() {
let tmp = TempDir::new().unwrap();
@@ -247,7 +247,7 @@ func TestEcEndToEnd(t *testing.T) {
"job_type": "erasure_coding",
"admin_runtime": map[string]interface{}{
"enabled": true,
"detection_interval_seconds": 1,
"detection_interval_minutes": 1,
"global_execution_concurrency": 4,
"per_worker_execution_concurrency": 4,
"max_jobs_per_detection": 100,
@@ -0,0 +1,265 @@
package erasure_coding
import (
"context"
"fmt"
"path/filepath"
"regexp"
"strconv"
"strings"
"testing"
"time"
"github.com/seaweedfs/seaweedfs/weed/pb/master_pb"
"github.com/seaweedfs/seaweedfs/weed/shell"
"github.com/seaweedfs/seaweedfs/weed/storage/needle"
"github.com/stretchr/testify/assert"
"github.com/stretchr/testify/require"
"google.golang.org/grpc"
)
// TestMultiDiskECBalanceNoShardLoss is the end-to-end regression for issue 9593.
// It runs a real cluster of multi-disk volume servers (3 servers x 4 disks),
// EC-encodes a volume, then runs ec.balance, asserting hard invariants the older
// integration tests only logged:
//
// - after encode the full set of 14 EC shards exists,
// - ec.balance never loses a shard (still 14 distinct shards afterwards),
// - shards end up spread across more than one disk per node, and
// - cluster.status counts physical disks (not one per node) and matches the
// real on-disk distribution.
func TestMultiDiskECBalanceNoShardLoss(t *testing.T) {
if testing.Short() {
t.Skip("Skipping multi-disk EC integration test in short mode")
}
testDir := t.TempDir()
ctx, cancel := context.WithTimeout(context.Background(), 240*time.Second)
defer cancel()
cluster, err := startMultiDiskCluster(ctx, testDir)
require.NoError(t, err)
defer cluster.Stop()
require.NoError(t, waitForServer("127.0.0.1:9334", 30*time.Second))
for i := 0; i < 3; i++ {
require.NoError(t, waitForServer(fmt.Sprintf("127.0.0.1:809%d", i), 30*time.Second))
}
t.Log("waiting for multi-disk volume servers to register...")
time.Sleep(10 * time.Second)
commandEnv := shell.NewCommandEnv(&shell.ShellOptions{
Masters: stringPtr("127.0.0.1:9334"),
GrpcDialOption: grpc.WithInsecure(),
FilerGroup: stringPtr("default"),
})
connectToMasterAndSync(ctx, t, commandEnv)
// Upload enough small files that the volume holds real data to encode.
var volumeId needle.VolumeId
for retry := 0; retry < 5; retry++ {
volumeId, err = uploadTestDataToMaster([]byte(strings.Repeat("multidisk-ec-9593 ", 64)), "127.0.0.1:9334")
if err == nil {
break
}
time.Sleep(3 * time.Second)
}
require.NoError(t, err, "failed to upload test data")
for i := 0; i < 40; i++ {
if _, e := uploadTestDataToMaster([]byte(strings.Repeat("filler ", 128)), "127.0.0.1:9334"); e != nil {
break
}
}
t.Logf("using volume %d", volumeId)
time.Sleep(3 * time.Second)
// Populate every server's disks with volumes so the encode can see and target
// each physical disk. The master only enumerates disks that already hold a
// volume or EC shard — an empty disk leaves no trace in the topology (heartbeats
// aggregate capacity per disk type, not per physical disk). ec.encode therefore
// spreads a volume's shards only across the disks the master already knows hold
// data on each node; if a node's data sits on a single disk, all its shards land
// there and ec.balance cannot redistribute them (it has no within-node
// cross-disk move). So spreading must be set up before encoding.
//
// volume.grow only tops up toward a writable target and stops on the first
// allocation error, so a single -count grow can create far fewer volumes than
// asked and leave a node on one disk. Grow repeatedly on the nodes that have not
// spread yet (the volume server places each new volume on its least-loaded disk)
// until the master's topology shows every node holding volumes on at least two
// physical disks. This makes the multi-disk layout — and thus the post-encode
// disk spread — deterministic instead of racing volume-growth and heartbeat.
require.Eventually(t, func() bool {
spread := nodeVolumeDiskCounts(t, commandEnv)
if len(spread) == 3 && allAtLeast(spread, 2) {
return true
}
for i := 0; i < 3; i++ {
server := fmt.Sprintf("127.0.0.1:809%d", i)
if spread[server] < 2 {
captureCommandOutput(t, shell.Commands[findCommandIndex("volume.grow")],
[]string{"-collection", "test", "-dataNode", server, "-count", "4"}, commandEnv)
}
}
return false
}, 60*time.Second, 2*time.Second,
"volumes never spread across >=2 disks on all 3 nodes")
locked, unlock := tryLockWithTimeout(t, commandEnv, 15*time.Second)
require.True(t, locked, "could not acquire shell lock")
defer unlock()
// EC-encode the volume.
out, err := captureCommandOutput(t, shell.Commands[findCommandIndex("ec.encode")],
[]string{"-volumeId", fmt.Sprintf("%d", volumeId), "-collection", "test", "-force"}, commandEnv)
t.Logf("ec.encode output:\n%s", out)
require.NoError(t, err, "ec.encode failed")
// All 14 shards must exist after encoding.
require.Eventually(t, func() bool {
return len(collectDistinctShardIDs(testDir, uint32(volumeId))) == erasureShardCount
}, 30*time.Second, time.Second, "expected all %d EC shards after encode, got %v",
erasureShardCount, collectDistinctShardIDs(testDir, uint32(volumeId)))
beforeBalance := collectDistinctShardIDs(testDir, uint32(volumeId))
t.Logf("after encode: %d distinct shards on %d disks", len(beforeBalance), disksWithShards(testDir, uint32(volumeId)))
// Run ec.balance.
out, err = captureCommandOutput(t, shell.Commands[findCommandIndex("ec.balance")],
[]string{"-collection", "test", "-force"}, commandEnv)
t.Logf("ec.balance output:\n%s", out)
require.NoError(t, err, "ec.balance failed")
time.Sleep(3 * time.Second)
// The core regression: ec.balance must not lose any shard.
afterBalance := collectDistinctShardIDs(testDir, uint32(volumeId))
require.Equal(t, erasureShardCount, len(afterBalance),
"ec.balance lost shards on multi-disk nodes: had %v, now %v", sortedKeysOf(beforeBalance), sortedKeysOf(afterBalance))
// Shards must be spread across more than one physical disk per node overall.
usedDisks := disksWithShards(testDir, uint32(volumeId))
assert.Greater(t, usedDisks, 3, "EC shards should span more than one disk per node (got %d disks across 3 nodes)", usedDisks)
// cluster.status must count physical disks, not collapse to one per node: it
// must report at least the disks actually holding this volume's shards (which
// is already >3 across the 3 nodes). Before the fix it reported 3 (node count).
require.Eventually(t, func() bool {
n, ok := clusterStatusDiskCount(t, commandEnv)
return ok && n >= usedDisks
}, 30*time.Second, 2*time.Second, "cluster.status never reported the >=%d physical disks holding shards (multi-disk count)", usedDisks)
n, _ := clusterStatusDiskCount(t, commandEnv)
t.Logf("cluster.status reports %d physical disks (>= %d holding this volume's shards)", n, usedDisks)
}
const erasureShardCount = 14 // 10 data + 4 parity
// collectDistinctShardIDs returns the set of EC shard ids present for a volume
// across every disk of every server in the multi-disk test layout.
func collectDistinctShardIDs(testDir string, volumeId uint32) map[int]bool {
ids := map[int]bool{}
for server := 0; server < 3; server++ {
for disk := 0; disk < 4; disk++ {
diskDir := filepath.Join(testDir, fmt.Sprintf("server%d_disk%d", server, disk))
files, err := listECShardFiles(diskDir, volumeId)
if err != nil {
continue
}
for _, f := range files {
i := strings.LastIndex(f, ".ec")
if i < 0 {
continue
}
if n, err := strconv.Atoi(f[i+3:]); err == nil && n >= 0 && n < erasureShardCount {
ids[n] = true
}
}
}
}
return ids
}
// disksWithShards counts how many physical disks hold at least one shard.
func disksWithShards(testDir string, volumeId uint32) int {
n := 0
for _, disks := range countShardsPerDisk(testDir, volumeId) {
for _, c := range disks {
if c > 0 {
n++
}
}
}
return n
}
// nodeVolumeDiskCounts returns, per volume server id, how many distinct physical
// disks hold at least one volume according to the master's topology. The master
// only enumerates disks that already hold a volume or EC shard (heartbeats
// aggregate capacity per disk type, not per physical disk), so this reports the
// disks ec.encode can actually spread a volume's shards across on each node.
func nodeVolumeDiskCounts(t *testing.T, commandEnv *shell.CommandEnv) map[string]int {
t.Helper()
var resp *master_pb.VolumeListResponse
err := commandEnv.MasterClient.WithClient(false, func(client master_pb.SeaweedClient) error {
var e error
resp, e = client.VolumeList(context.Background(), &master_pb.VolumeListRequest{})
return e
})
counts := map[string]int{}
if err != nil || resp.GetTopologyInfo() == nil {
return counts
}
for _, dc := range resp.GetTopologyInfo().GetDataCenterInfos() {
for _, r := range dc.GetRackInfos() {
for _, dn := range r.GetDataNodeInfos() {
disks := map[uint32]bool{}
for _, di := range dn.GetDiskInfos() {
for _, vi := range di.GetVolumeInfos() {
disks[vi.GetDiskId()] = true
}
}
counts[dn.Id] = len(disks)
}
}
}
return counts
}
func allAtLeast(counts map[string]int, min int) bool {
for _, c := range counts {
if c < min {
return false
}
}
return true
}
var diskCountRe = regexp.MustCompile(`(\d+)\s+disks?`)
// clusterStatusDiskCount runs cluster.status and parses the reported disk count.
func clusterStatusDiskCount(t *testing.T, commandEnv *shell.CommandEnv) (int, bool) {
t.Helper()
out, err := captureCommandOutput(t, shell.Commands[findCommandIndex("cluster.status")], []string{}, commandEnv)
if err != nil {
return 0, false
}
m := diskCountRe.FindStringSubmatch(out)
if m == nil {
return 0, false
}
n, err := strconv.Atoi(m[1])
return n, err == nil
}
func sortedKeysOf(m map[int]bool) []int {
out := make([]int, 0, len(m))
for k := range m {
out = append(out, k)
}
for i := 1; i < len(out); i++ {
for j := i; j > 0 && out[j-1] > out[j]; j-- {
out[j-1], out[j] = out[j], out[j-1]
}
}
return out
}
@@ -207,8 +207,10 @@ func buildVolumeListResponse(t *testing.T, spec topologySpec, volumeID uint32) *
t.Helper()
volumeSizeLimitMB := uint64(100)
volumeSize := uint64(90) * 1024 * 1024
volumeModifiedAt := time.Now().Add(-10 * time.Minute).Unix()
// Exceed the default fullness (0.95) and quiet (1h) thresholds so volumes are
// EC-eligible.
volumeSize := uint64(96) * 1024 * 1024
volumeModifiedAt := time.Now().Add(-2 * time.Hour).Unix()
diskTypes := spec.diskTypes
if len(diskTypes) == 0 {
@@ -29,9 +29,11 @@ func TestErasureCodingDetectionLargeTopology(t *testing.T) {
}
nodesPerRack := serverCount / rackCount
eligibleSize := uint64(90) * 1024 * 1024
// Eligible volumes must exceed the default fullness (0.95) and quiet (1h)
// thresholds; ineligible ones fall below the fullness threshold.
eligibleSize := uint64(96) * 1024 * 1024
ineligibleSize := uint64(10) * 1024 * 1024
modifiedAt := time.Now().Add(-10 * time.Minute).Unix()
modifiedAt := time.Now().Add(-2 * time.Hour).Unix()
volumeID := uint32(1)
dataCenters := make([]*master_pb.DataCenterInfo, 0, 1)
+14
View File
@@ -292,6 +292,20 @@ func (v *VolumeServer) VolumeEcShardsMount(ctx context.Context, req *volume_serv
return &volume_server_pb.VolumeEcShardsMountResponse{}, nil
}
// VolumeEcShardsUnmount is a no-op stub: the worker's pre-distribute
// cleanup calls it against every destination, and the fake server has no
// mounted state to clear.
func (v *VolumeServer) VolumeEcShardsUnmount(ctx context.Context, req *volume_server_pb.VolumeEcShardsUnmountRequest) (*volume_server_pb.VolumeEcShardsUnmountResponse, error) {
return &volume_server_pb.VolumeEcShardsUnmountResponse{}, nil
}
// VolumeEcShardsDelete is a no-op stub paired with VolumeEcShardsUnmount
// above; the fake server doesn't persist shard files beyond what
// ReceiveFile wrote, so there's nothing to remove.
func (v *VolumeServer) VolumeEcShardsDelete(ctx context.Context, req *volume_server_pb.VolumeEcShardsDeleteRequest) (*volume_server_pb.VolumeEcShardsDeleteResponse, error) {
return &volume_server_pb.VolumeEcShardsDeleteResponse{}, nil
}
func (v *VolumeServer) VolumeEcShardsInfo(ctx context.Context, req *volume_server_pb.VolumeEcShardsInfoRequest) (*volume_server_pb.VolumeEcShardsInfoResponse, error) {
if req == nil {
return nil, fmt.Errorf("VolumeEcShardsInfo request is nil")
@@ -0,0 +1,309 @@
package iam
import (
"fmt"
mathrand "math/rand"
"strings"
"testing"
"time"
"github.com/aws/aws-sdk-go/aws"
"github.com/aws/aws-sdk-go/aws/awserr"
"github.com/aws/aws-sdk-go/service/iam"
"github.com/aws/aws-sdk-go/service/s3"
"github.com/stretchr/testify/assert"
"github.com/stretchr/testify/require"
)
// uniqueResourceSuffix returns a lowercased per-test, per-invocation suffix
// safe for use in IAM resource and S3 bucket names. Avoids EntityAlreadyExists
// / BucketAlreadyExists collisions when integration jobs retry or run in
// parallel against a shared stack.
func uniqueResourceSuffix(t *testing.T) string {
name := strings.ToLower(t.Name())
name = strings.ReplaceAll(name, "/", "-")
name = strings.ReplaceAll(name, "_", "-")
return fmt.Sprintf("%s-%d", name, mathrand.Intn(10000))
}
// isAccessDenied returns true when err is an AWS error with code "AccessDenied".
// Used to gate deny-path polling so transient setup errors don't end the
// Eventually loop prematurely.
func isAccessDenied(err error) bool {
if err == nil {
return false
}
awsErr, ok := err.(awserr.Error)
return ok && awsErr.Code() == "AccessDenied"
}
// TestIAMUserInlinePolicySourceIpCondition verifies that an aws:SourceIp condition
// on a user inline policy is honored. Tests run from localhost (127.0.0.1), so a
// policy that only allows access from a non-loopback CIDR must deny the request,
// and a policy that allows access from 127.0.0.0/8 must allow it.
func TestIAMUserInlinePolicySourceIpCondition(t *testing.T) {
framework := NewS3IAMTestFramework(t)
defer framework.Cleanup()
iamClient, err := framework.CreateIAMClientWithJWT("admin-user", "TestAdminRole")
require.NoError(t, err)
suffix := uniqueResourceSuffix(t)
userName := "user-" + suffix
policyName := "policy-" + suffix
bucketName := "bucket-" + suffix
_, err = iamClient.CreateUser(&iam.CreateUserInput{UserName: aws.String(userName)})
require.NoError(t, err)
keyResp, err := iamClient.CreateAccessKey(&iam.CreateAccessKeyInput{
UserName: aws.String(userName),
})
require.NoError(t, err)
accessKeyId := *keyResp.AccessKey.AccessKeyId
secretKey := *keyResp.AccessKey.SecretAccessKey
userS3 := createS3Client(t, accessKeyId, secretKey)
adminS3, err := framework.CreateS3ClientWithJWT("admin-user", "TestAdminRole")
require.NoError(t, err)
require.NoError(t, framework.CreateBucketWithCleanup(adminS3, bucketName))
t.Cleanup(func() {
if _, err := iamClient.DeleteUserPolicy(&iam.DeleteUserPolicyInput{
UserName: aws.String(userName),
PolicyName: aws.String(policyName),
}); err != nil {
t.Logf("cleanup: failed to delete user policy: %v", err)
}
if _, err := iamClient.DeleteAccessKey(&iam.DeleteAccessKeyInput{
UserName: aws.String(userName),
AccessKeyId: keyResp.AccessKey.AccessKeyId,
}); err != nil {
t.Logf("cleanup: failed to delete access key: %v", err)
}
if _, err := iamClient.DeleteUser(&iam.DeleteUserInput{UserName: aws.String(userName)}); err != nil {
t.Logf("cleanup: failed to delete user: %v", err)
}
})
policyDoc := func(cidrs ...string) string {
quoted := make([]string, len(cidrs))
for i, c := range cidrs {
quoted[i] = `"` + c + `"`
}
return `{
"Version":"2012-10-17",
"Statement":[{
"Effect":"Allow",
"Action":"s3:*",
"Resource":["arn:aws:s3:::` + bucketName + `","arn:aws:s3:::` + bucketName + `/*"],
"Condition":{"IpAddress":{"aws:SourceIp":[` + strings.Join(quoted, ",") + `]}}
}]
}`
}
t.Run("denies_when_source_ip_does_not_match", func(t *testing.T) {
// SourceIp 198.51.100.0/24 is RFC5737 TEST-NET-2; the test client is on
// loopback (127.0.0.1 or ::1 depending on resolver), so the condition
// must fail and the action must be denied.
_, err = iamClient.PutUserPolicy(&iam.PutUserPolicyInput{
UserName: aws.String(userName),
PolicyName: aws.String(policyName),
PolicyDocument: aws.String(policyDoc("198.51.100.0/24")),
})
require.NoError(t, err)
var lastErr error
require.Eventually(t, func() bool {
_, lastErr = userS3.PutObject(&s3.PutObjectInput{
Bucket: aws.String(bucketName),
Key: aws.String("denied.txt"),
Body: aws.ReadSeekCloser(strings.NewReader("nope")),
})
return isAccessDenied(lastErr)
}, 10*time.Second, 500*time.Millisecond,
"PutObject must be denied with AccessDenied when aws:SourceIp condition does not match (last error: %v)", lastErr)
})
t.Run("allows_when_source_ip_matches", func(t *testing.T) {
// Cover both IPv4 and IPv6 loopback: on CI runners `localhost` may
// resolve to ::1 first, in which case a 127.0.0.0/8-only allow would
// silently never match and the test would hang.
_, err = iamClient.PutUserPolicy(&iam.PutUserPolicyInput{
UserName: aws.String(userName),
PolicyName: aws.String(policyName),
PolicyDocument: aws.String(policyDoc("127.0.0.0/8", "::1/128")),
})
require.NoError(t, err)
require.Eventually(t, func() bool {
_, err := userS3.PutObject(&s3.PutObjectInput{
Bucket: aws.String(bucketName),
Key: aws.String("allowed.txt"),
Body: aws.ReadSeekCloser(strings.NewReader("ok")),
})
return err == nil
}, 10*time.Second, 500*time.Millisecond,
"PutObject must succeed when aws:SourceIp condition matches the loopback range")
})
}
// TestIAMGroupInlinePolicyEnforcement verifies that PutGroupPolicy is supported
// and that the resulting inline policy is enforced for members of the group,
// including its Condition block.
func TestIAMGroupInlinePolicyEnforcement(t *testing.T) {
framework := NewS3IAMTestFramework(t)
defer framework.Cleanup()
iamClient, err := framework.CreateIAMClientWithJWT("admin-user", "TestAdminRole")
require.NoError(t, err)
suffix := uniqueResourceSuffix(t)
groupName := "group-" + suffix
userName := "user-" + suffix
policyName := "policy-" + suffix
bucketName := "bucket-" + suffix
_, err = iamClient.CreateUser(&iam.CreateUserInput{UserName: aws.String(userName)})
require.NoError(t, err)
keyResp, err := iamClient.CreateAccessKey(&iam.CreateAccessKeyInput{
UserName: aws.String(userName),
})
require.NoError(t, err)
_, err = iamClient.CreateGroup(&iam.CreateGroupInput{GroupName: aws.String(groupName)})
require.NoError(t, err)
_, err = iamClient.AddUserToGroup(&iam.AddUserToGroupInput{
GroupName: aws.String(groupName),
UserName: aws.String(userName),
})
require.NoError(t, err)
userS3 := createS3Client(t, *keyResp.AccessKey.AccessKeyId, *keyResp.AccessKey.SecretAccessKey)
adminS3, err := framework.CreateS3ClientWithJWT("admin-user", "TestAdminRole")
require.NoError(t, err)
require.NoError(t, framework.CreateBucketWithCleanup(adminS3, bucketName))
t.Cleanup(func() {
if _, err := iamClient.DeleteGroupPolicy(&iam.DeleteGroupPolicyInput{
GroupName: aws.String(groupName),
PolicyName: aws.String(policyName),
}); err != nil {
t.Logf("cleanup: failed to delete group policy: %v", err)
}
if _, err := iamClient.RemoveUserFromGroup(&iam.RemoveUserFromGroupInput{
GroupName: aws.String(groupName),
UserName: aws.String(userName),
}); err != nil {
t.Logf("cleanup: failed to remove user from group: %v", err)
}
if _, err := iamClient.DeleteAccessKey(&iam.DeleteAccessKeyInput{
UserName: aws.String(userName),
AccessKeyId: keyResp.AccessKey.AccessKeyId,
}); err != nil {
t.Logf("cleanup: failed to delete access key: %v", err)
}
if _, err := iamClient.DeleteUser(&iam.DeleteUserInput{UserName: aws.String(userName)}); err != nil {
t.Logf("cleanup: failed to delete user: %v", err)
}
if _, err := iamClient.DeleteGroup(&iam.DeleteGroupInput{GroupName: aws.String(groupName)}); err != nil {
t.Logf("cleanup: failed to delete group: %v", err)
}
})
// Cover both IPv4 and IPv6 loopback in the allow CIDR list: on CI runners
// `localhost` may resolve to ::1 first, in which case a 127.0.0.0/8-only
// allow would silently never match and the test would hang.
allowDoc := `{
"Version":"2012-10-17",
"Statement":[{
"Effect":"Allow",
"Action":"s3:*",
"Resource":["arn:aws:s3:::` + bucketName + `","arn:aws:s3:::` + bucketName + `/*"],
"Condition":{"IpAddress":{"aws:SourceIp":["127.0.0.0/8","::1/128"]}}
}]
}`
denyDoc := `{
"Version":"2012-10-17",
"Statement":[{
"Effect":"Allow",
"Action":"s3:*",
"Resource":["arn:aws:s3:::` + bucketName + `","arn:aws:s3:::` + bucketName + `/*"],
"Condition":{"IpAddress":{"aws:SourceIp":"198.51.100.0/24"}}
}]
}`
t.Run("crud_round_trip", func(t *testing.T) {
_, err := iamClient.PutGroupPolicy(&iam.PutGroupPolicyInput{
GroupName: aws.String(groupName),
PolicyName: aws.String(policyName),
PolicyDocument: aws.String(allowDoc),
})
require.NoError(t, err, "PutGroupPolicy must succeed (no longer NotImplemented)")
listResp, err := iamClient.ListGroupPolicies(&iam.ListGroupPoliciesInput{
GroupName: aws.String(groupName),
})
require.NoError(t, err)
found := false
for _, name := range listResp.PolicyNames {
if name != nil && *name == policyName {
found = true
break
}
}
assert.True(t, found, "ListGroupPolicies must return the freshly added policy")
getResp, err := iamClient.GetGroupPolicy(&iam.GetGroupPolicyInput{
GroupName: aws.String(groupName),
PolicyName: aws.String(policyName),
})
require.NoError(t, err)
require.NotNil(t, getResp.PolicyDocument)
assert.Contains(t, *getResp.PolicyDocument, "aws:SourceIp",
"GetGroupPolicy must round-trip the Condition block")
})
t.Run("enforces_allow_when_condition_matches", func(t *testing.T) {
_, err := iamClient.PutGroupPolicy(&iam.PutGroupPolicyInput{
GroupName: aws.String(groupName),
PolicyName: aws.String(policyName),
PolicyDocument: aws.String(allowDoc),
})
require.NoError(t, err)
require.Eventually(t, func() bool {
_, err := userS3.PutObject(&s3.PutObjectInput{
Bucket: aws.String(bucketName),
Key: aws.String("group-allowed.txt"),
Body: aws.ReadSeekCloser(strings.NewReader("ok")),
})
return err == nil
}, 10*time.Second, 500*time.Millisecond,
"group member must be allowed when the group policy condition matches")
})
t.Run("enforces_deny_when_condition_does_not_match", func(t *testing.T) {
_, err := iamClient.PutGroupPolicy(&iam.PutGroupPolicyInput{
GroupName: aws.String(groupName),
PolicyName: aws.String(policyName),
PolicyDocument: aws.String(denyDoc),
})
require.NoError(t, err)
var lastErr error
require.Eventually(t, func() bool {
_, lastErr = userS3.PutObject(&s3.PutObjectInput{
Bucket: aws.String(bucketName),
Key: aws.String("group-denied.txt"),
Body: aws.ReadSeekCloser(strings.NewReader("nope")),
})
return isAccessDenied(lastErr)
}, 10*time.Second, 500*time.Millisecond,
"group member must be denied with AccessDenied when the group policy condition does not match (last error: %v)", lastErr)
})
}
+7 -1
View File
@@ -9,6 +9,10 @@ S3_PORT := 8333
MASTER_PORT := 9333
VOLUME_PORT := 8080
FILER_PORT := 8888
# Pin the filer gRPC port off the FILER_PORT+10000 convention so any
# code path that assumes the offset breaks here, not in production.
FILER_GRPC_PORT := 18890
ADMIN_PORT := 23646
ACCESS_KEY ?= some_access_key1
SECRET_KEY ?= some_secret_key1
TEST_TIMEOUT := 10m
@@ -39,6 +43,7 @@ start-server: build-weed
@AWS_ACCESS_KEY_ID=$(ACCESS_KEY) AWS_SECRET_ACCESS_KEY=$(SECRET_KEY) $(WEED_BINARY) mini \
-dir=$(SERVER_DIR) \
-s3.port=$(S3_PORT) \
-filer.port.grpc=$(FILER_GRPC_PORT) \
> weed-test.log 2>&1 & \
echo $$! > weed-server.pid
@for i in $$(seq 1 90); do \
@@ -68,7 +73,8 @@ test:
S3_ENDPOINT=http://localhost:$(S3_PORT) \
S3_GRPC_ENDPOINT=localhost:$$(($(S3_PORT) + 10000)) \
MASTER_ENDPOINT=http://localhost:$(MASTER_PORT) \
FILER_GRPC_ADDRESS=localhost:$$(($(FILER_PORT) + 10000)) \
FILER_GRPC_ADDRESS=localhost:$(FILER_GRPC_PORT) \
ADMIN_ENDPOINT=http://localhost:$(ADMIN_PORT) \
go test -v -timeout $(TEST_TIMEOUT) -run $(TEST_PATTERN)
test-with-server: start-server
@@ -0,0 +1,113 @@
package lifecycle
import (
"bytes"
"context"
"encoding/json"
"fmt"
"net/http"
"strings"
"testing"
"time"
"github.com/aws/aws-sdk-go-v2/aws"
"github.com/aws/aws-sdk-go-v2/service/s3"
"github.com/stretchr/testify/require"
)
// Drives the real admin->worker dispatch path via the run endpoint and
// asserts the dispatch reaches the filer (no dial error) and deletes
// the backdated object. The Makefile pins filer.port.grpc off the
// FILER_PORT+10000 convention so a raw-address forwarding regression
// resurfaces here.
func TestLifecycleAdminDispatchSucceedsWithCustomFilerGrpcPort(t *testing.T) {
adminEndpoint := envOr("ADMIN_ENDPOINT", defaultAdminEndpoint)
c := s3Client(t)
fc, fcClose := filerClient(t)
defer fcClose()
bucket := uniqueBucket("admin-dispatch")
mustCreateBucket(t, c, bucket)
putExpirationLifecycle(t, c, bucket, "expire/", 1)
const oldKey = "expire/old.txt"
putObject(t, c, bucket, oldKey, "old")
backdateMtime(t, fc, bucket, oldKey, 30)
waitForLifecycleWorkerReady(t, adminEndpoint)
// Lifecycle is a long-running batch; the run endpoint cancels it at
// this timeout, which converts a healthy run into canceled_count=1.
const runTimeoutSeconds = 30
body, err := json.Marshal(map[string]any{
"timeout_seconds": runTimeoutSeconds,
})
require.NoError(t, err)
req, err := http.NewRequestWithContext(
context.Background(), http.MethodPost,
adminEndpoint+"/api/plugin/job-types/s3_lifecycle/run",
bytes.NewReader(body),
)
require.NoError(t, err)
req.Header.Set("Content-Type", "application/json")
resp, err := http.DefaultClient.Do(req)
require.NoError(t, err)
defer resp.Body.Close()
var payload map[string]any
require.NoError(t, json.NewDecoder(resp.Body).Decode(&payload))
t.Logf("admin /api/plugin/job-types/s3_lifecycle/run response: %v", payload)
require.Equal(t, http.StatusOK, resp.StatusCode, "admin run endpoint failed: %v", payload)
require.GreaterOrEqual(t, jsonNumber(t, payload, "detected_count"), 1)
require.Equal(t, 0, jsonNumber(t, payload, "error_count"),
"dispatched job errored — likely filer_grpc_address was raw host:httpPort.grpcPort")
require.Eventuallyf(t, func() bool {
_, err := c.HeadObject(context.Background(), &s3.HeadObjectInput{
Bucket: aws.String(bucket), Key: aws.String(oldKey),
})
return err != nil
}, 30*time.Second, 500*time.Millisecond,
"expected %s/%s to be deleted after admin-dispatched lifecycle run", bucket, oldKey)
}
func jsonNumber(t *testing.T, payload map[string]any, key string) int {
t.Helper()
raw, ok := payload[key]
require.Truef(t, ok, "response missing key %q: %v", key, payload)
switch v := raw.(type) {
case float64:
return int(v)
case json.Number:
n, err := v.Int64()
require.NoErrorf(t, err, "key %q is not an int: %v", key, raw)
return int(n)
default:
t.Fatalf("key %q has unexpected type %T (%v)", key, raw, raw)
return 0
}
}
func waitForLifecycleWorkerReady(t *testing.T, adminEndpoint string) {
t.Helper()
deadline := time.Now().Add(30 * time.Second)
for time.Now().Before(deadline) {
req, err := http.NewRequest(http.MethodGet, adminEndpoint+"/api/plugin/scheduler-states", nil)
require.NoError(t, err)
resp, err := http.DefaultClient.Do(req)
if err != nil {
time.Sleep(250 * time.Millisecond)
continue
}
var payload any
_ = json.NewDecoder(resp.Body).Decode(&payload)
resp.Body.Close()
if resp.StatusCode == http.StatusOK && strings.Contains(fmt.Sprintf("%v", payload), "s3_lifecycle") {
return
}
time.Sleep(250 * time.Millisecond)
}
t.Fatalf("admin never reported an s3_lifecycle-capable worker")
}
+5 -3
View File
@@ -40,12 +40,14 @@ const (
defaultS3Endpoint = "http://localhost:8333"
defaultS3GrpcEndpoint = "localhost:18333"
defaultMasterEndpt = "http://localhost:9333"
defaultFilerGRPC = "localhost:18888"
defaultAdminEndpoint = "http://localhost:23646"
// Pinned off the FILER_PORT+10000 convention; see Makefile.
defaultFilerGRPC = "localhost:18890"
bucketLifecycleXMLKey = "s3-bucket-lifecycle-configuration-xml"
bucketsPath = "/buckets"
accessKey = "some_access_key1"
secretKey = "some_secret_key1"
region = "us-east-1"
secretKey = "some_secret_key1"
region = "us-east-1"
)
func envOr(key, def string) string {
+41
View File
@@ -0,0 +1,41 @@
# AWS SDK V2 Route Disambiguation Integration Tests
#
# Pins the regression for the route collision between the regular S3 API
# and the S3 Tables REST API on shared top-level paths (/buckets,
# /get-table). The tests use the real AWS SDK V2 for Go so the SDK's own
# XML deserializer is the assertion — a JSON body produces an SDK error
# before any test code runs.
#
# Prerequisites:
# - SeaweedFS running with S3 API enabled on port 8333
# - Go 1.21+
#
# Usage:
# make test - Run the SDK V2 routing tests
# make test-verbose - Run with verbose output
# make clean - Clean test cache
.PHONY: all test test-verbose clean help
S3_ENDPOINT ?= http://127.0.0.1:8333
all: test
test:
@echo "Running SDK V2 routing tests against $(S3_ENDPOINT)..."
S3_ENDPOINT=$(S3_ENDPOINT) go test -v -timeout 5m ./...
test-verbose:
S3_ENDPOINT=$(S3_ENDPOINT) go test -v -timeout 5m -count=1 ./...
clean:
go clean -testcache
help:
@echo "AWS SDK V2 Route Disambiguation Tests"
@echo "Targets:"
@echo " test Run the routing tests"
@echo " test-verbose Run with verbose output"
@echo " clean Clean test cache"
@echo "Environment Variables:"
@echo " S3_ENDPOINT S3 endpoint URL (default: http://127.0.0.1:8333)"
@@ -0,0 +1,210 @@
// Package sdkv2routing_test exercises route disambiguation between the
// regular S3 API and the S3 Tables REST API on top-level paths the two
// share (/buckets, /get-table). The bug it pins:
//
// When a user has an S3 bucket named "buckets" (or "get-table"), a
// path-style ListObjectsV2 request sent by AWS SDK V2 / Hadoop s3a /
// Spark would be routed to the S3 Tables ListTableBuckets handler and
// receive a JSON body. AWS SDK V2 then fails XML parsing with
// "Unexpected character '{' (code 123) in prolog".
//
// These tests use the real AWS SDK V2 for Go, so the SDK's own
// deserializer is the assertion: if the server returns the wrong
// content type, the SDK errors out before any test assertion runs.
package sdkv2routing_test
import (
"context"
"io"
"os"
"strings"
"testing"
"time"
"github.com/aws/aws-sdk-go-v2/aws"
"github.com/aws/aws-sdk-go-v2/config"
"github.com/aws/aws-sdk-go-v2/credentials"
"github.com/aws/aws-sdk-go-v2/service/s3"
"github.com/stretchr/testify/require"
)
const (
defaultEndpoint = "http://127.0.0.1:8333"
defaultAccessKey = "some_access_key1"
defaultSecretKey = "some_secret_key1"
defaultRegion = "us-east-1"
)
func getS3Client(t *testing.T) *s3.Client {
t.Helper()
endpoint := os.Getenv("S3_ENDPOINT")
if endpoint == "" {
endpoint = defaultEndpoint
}
accessKey := os.Getenv("AWS_ACCESS_KEY_ID")
if accessKey == "" {
accessKey = defaultAccessKey
}
secretKey := os.Getenv("AWS_SECRET_ACCESS_KEY")
if secretKey == "" {
secretKey = defaultSecretKey
}
region := os.Getenv("AWS_REGION")
if region == "" {
region = defaultRegion
}
cfg, err := config.LoadDefaultConfig(context.TODO(),
config.WithRegion(region),
config.WithCredentialsProvider(credentials.NewStaticCredentialsProvider(accessKey, secretKey, "")),
config.WithEndpointResolverWithOptions(aws.EndpointResolverWithOptionsFunc(
func(service, region string, options ...interface{}) (aws.Endpoint, error) {
return aws.Endpoint{
URL: endpoint,
SigningRegion: defaultRegion,
HostnameImmutable: true,
}, nil
})),
)
require.NoError(t, err)
return s3.NewFromConfig(cfg, func(o *s3.Options) {
o.UsePathStyle = true
})
}
// ensureBucket creates bucket if it doesn't already exist. It tolerates
// BucketAlreadyOwnedByYou / BucketAlreadyExists so the tests are
// idempotent across local re-runs.
func ensureBucket(t *testing.T, ctx context.Context, client *s3.Client, bucket string) {
t.Helper()
_, err := client.CreateBucket(ctx, &s3.CreateBucketInput{Bucket: aws.String(bucket)})
if err == nil {
return
}
msg := err.Error()
if strings.Contains(msg, "BucketAlreadyOwnedByYou") || strings.Contains(msg, "BucketAlreadyExists") {
return
}
t.Fatalf("CreateBucket(%q) failed: %v", bucket, err)
}
// deleteBucket best-effort cleans up a bucket and any objects in it.
// Test does not fail if cleanup fails — the next run is idempotent.
func deleteBucket(ctx context.Context, client *s3.Client, bucket string) {
paginator := s3.NewListObjectsV2Paginator(client, &s3.ListObjectsV2Input{Bucket: aws.String(bucket)})
for paginator.HasMorePages() {
page, err := paginator.NextPage(ctx)
if err != nil {
break
}
for _, obj := range page.Contents {
client.DeleteObject(ctx, &s3.DeleteObjectInput{Bucket: aws.String(bucket), Key: obj.Key})
}
}
client.DeleteBucket(ctx, &s3.DeleteBucketInput{Bucket: aws.String(bucket)})
}
// TestListObjectsV2_OnBucketNamedBuckets is the direct reproducer for
// issue #9559: Spark / Hadoop s3a does a ListObjectsV2 against bucket
// "buckets" via AWS SDK V2, which fails with
// "Could not parse XML response. ... Unexpected character '{' (code 123)
// in prolog" when SeaweedFS routes the request to the JSON-returning
// ListTableBuckets handler. The SDK's response deserializer is the
// real assertion here — a JSON body produces an SDK error before we
// reach require.NoError.
func TestListObjectsV2_OnBucketNamedBuckets(t *testing.T) {
ctx, cancel := context.WithTimeout(context.Background(), 30*time.Second)
defer cancel()
client := getS3Client(t)
const bucket = "buckets"
ensureBucket(t, ctx, client, bucket)
t.Cleanup(func() { deleteBucket(ctx, client, bucket) })
out, err := client.ListObjectsV2(ctx, &s3.ListObjectsV2Input{
Bucket: aws.String(bucket),
Prefix: aws.String("logs/"),
})
require.NoError(t, err, "AWS SDK V2 must parse the response as XML")
require.NotNil(t, out)
}
// TestPutGetObject_OnBucketNamedBuckets exercises the full read/write
// round-trip on the colliding bucket name. PutObject and GetObject go
// through different routes than ListObjectsV2, and verifying them
// guards against future regressions that re-route only some verbs.
func TestPutGetObject_OnBucketNamedBuckets(t *testing.T) {
ctx, cancel := context.WithTimeout(context.Background(), 30*time.Second)
defer cancel()
client := getS3Client(t)
const bucket = "buckets"
const key = "logs/hello.txt"
const body = "hello from issue 9559"
ensureBucket(t, ctx, client, bucket)
t.Cleanup(func() { deleteBucket(ctx, client, bucket) })
_, err := client.PutObject(ctx, &s3.PutObjectInput{
Bucket: aws.String(bucket),
Key: aws.String(key),
Body: strings.NewReader(body),
})
require.NoError(t, err)
out, err := client.GetObject(ctx, &s3.GetObjectInput{Bucket: aws.String(bucket), Key: aws.String(key)})
require.NoError(t, err)
defer out.Body.Close()
got, err := io.ReadAll(out.Body)
require.NoError(t, err)
require.Equal(t, body, string(got))
}
// TestListObjectsV2_OnBucketNamedGetTable covers the second colliding
// path. The S3 Tables GET /get-table endpoint shares its path with a
// bucket literally named "get-table", which is a legal S3 bucket name.
func TestListObjectsV2_OnBucketNamedGetTable(t *testing.T) {
ctx, cancel := context.WithTimeout(context.Background(), 30*time.Second)
defer cancel()
client := getS3Client(t)
const bucket = "get-table"
ensureBucket(t, ctx, client, bucket)
t.Cleanup(func() { deleteBucket(ctx, client, bucket) })
out, err := client.ListObjectsV2(ctx, &s3.ListObjectsV2Input{Bucket: aws.String(bucket)})
require.NoError(t, err)
require.NotNil(t, out)
}
// TestCreateAndListBuckets_ServiceLevel verifies the SDK's service-level
// ListBuckets still parses as XML when a bucket named "buckets" exists.
// ListBuckets goes through GET / (root) which is unaffected by the
// /buckets route collision — this is a guard against the matcher
// accidentally widening to top-level paths.
func TestCreateAndListBuckets_ServiceLevel(t *testing.T) {
ctx, cancel := context.WithTimeout(context.Background(), 30*time.Second)
defer cancel()
client := getS3Client(t)
const bucket = "buckets"
ensureBucket(t, ctx, client, bucket)
t.Cleanup(func() { deleteBucket(ctx, client, bucket) })
out, err := client.ListBuckets(ctx, &s3.ListBucketsInput{})
require.NoError(t, err)
require.NotNil(t, out)
found := false
for _, b := range out.Buckets {
if aws.ToString(b.Name) == bucket {
found = true
break
}
}
require.True(t, found, "bucket %q must appear in service-level ListBuckets", bucket)
}
+23 -3
View File
@@ -17,6 +17,9 @@ import (
"testing"
"time"
"github.com/aws/aws-sdk-go/aws/credentials"
v4 "github.com/aws/aws-sdk-go/aws/signer/v4"
"github.com/seaweedfs/seaweedfs/test/testutil"
)
@@ -451,19 +454,22 @@ func icebergPath(prefix, path string) string {
return withPrefix
}
// createTableBucket creates a table bucket via the S3Tables REST API
// createTableBucket creates a table bucket via the S3Tables REST API.
// The request is AWS V4 signed for SERVICE=s3tables so the S3 Tables
// route matcher accepts it; signing with regular SERVICE=s3 would let
// the request fall through to the S3 CreateBucket handler.
func createTableBucket(t *testing.T, env *TestEnvironment, bucketName string) {
t.Helper()
// Use S3Tables REST API to create the bucket
endpoint := fmt.Sprintf("http://localhost:%d/buckets", env.s3Port)
reqBody := fmt.Sprintf(`{"name":"%s"}`, bucketName)
req, err := http.NewRequest(http.MethodPut, endpoint, strings.NewReader(reqBody))
if err != nil {
t.Fatalf("Failed to create request: %v", err)
}
req.Header.Set("Content-Type", "application/x-amz-json-1.1")
signS3TablesRequest(t, req, reqBody)
resp, err := http.DefaultClient.Do(req)
if err != nil {
@@ -480,6 +486,20 @@ func createTableBucket(t *testing.T, env *TestEnvironment, bucketName string) {
t.Logf("Created table bucket %s", bucketName)
}
// signS3TablesRequest signs req with AWS V4 for SERVICE=s3tables. The
// underlying weed mini instance runs in default-allow mode so the
// signature itself is not verified; only the credential scope matters,
// because the S3 Tables route matcher requires SERVICE=s3tables to
// distinguish S3 Tables traffic from regular S3 calls on the same paths.
func signS3TablesRequest(t *testing.T, req *http.Request, body string) {
t.Helper()
creds := credentials.NewStaticCredentials("test-ak", "test-sk", "")
signer := v4.NewSigner(creds)
if _, err := signer.Sign(req, strings.NewReader(body), "s3tables", "us-east-1", time.Now()); err != nil {
t.Fatalf("Failed to sign S3 Tables request: %v", err)
}
}
// randomSuffix returns a short random hex suffix for unique resource naming.
func randomSuffix() string {
return fmt.Sprintf("%x", time.Now().UnixNano()&0xffffffff)
@@ -411,7 +411,12 @@ func createIcebergTable(t *testing.T, env *TestEnvironment, bucketName, namespac
func listFilerContents(t *testing.T, env *TestEnvironment, path string) {
t.Helper()
cmd := exec.Command("weed", "shell",
// Bound diagnostic listing so a hung weed shell during cleanup can't
// burn the whole 20-minute test timeout.
ctx, cancel := context.WithTimeout(context.Background(), 30*time.Second)
defer cancel()
cmd := exec.CommandContext(ctx, "weed", "shell",
fmt.Sprintf("-master=%s", env.hostMasterAddress()),
)
cmd.Stdin = strings.NewReader(fmt.Sprintf("fs.ls -R %s\nexit\n", path))
+20
View File
@@ -0,0 +1,20 @@
FROM chrislusf/seaweedfs:e2e
RUN apt-get -o Acquire::Retries=5 -o Acquire::http::Timeout=30 update && \
DEBIAN_FRONTEND=noninteractive apt-get -o Acquire::Retries=5 -o Acquire::http::Timeout=30 install -y \
--no-install-recommends \
--no-install-suggests \
samba \
smbclient \
python3-minimal \
&& apt-get clean \
&& rm -rf /var/lib/apt/lists/*
COPY smb.conf.template /smb.conf.template
COPY smb_tests.sh /smb_tests.sh
COPY lock_tests.sh /lock_tests.sh
COPY entrypoint.sh /entrypoint.sh
COPY run_inside_container.sh /run_inside_container.sh
RUN chmod +x /smb_tests.sh /lock_tests.sh /entrypoint.sh /run_inside_container.sh
ENTRYPOINT ["/entrypoint.sh"]
+96
View File
@@ -0,0 +1,96 @@
# Samba on FUSE integration test
Exports a SeaweedFS FUSE mount over SMB with Samba's `smbd` and drives it with
`smbclient`, verifying that SMB file operations work correctly on top of the
mount and that data stays consistent across both protocols.
## What it checks
The functional battery in `smb_tests.sh` covers:
- connecting to the share and listing the root
- 1 MiB upload/download round-trip with content verification
- subdirectory creation and writes into it
- file rename
- 64 MiB upload/download (exercises SeaweedFS chunk splitting)
- recursive upload of a directory tree
- cross-protocol consistency: files written over SMB appear on the FUSE mount
with identical content, and files written directly on the FUSE mount are
readable over SMB
- deleting files and directory trees
The locking / concurrency battery in `lock_tests.sh` covers the harder cases a
network-filesystem backend has to get right:
- **POSIX `fcntl` byte-range locking** on the FUSE mount: a held exclusive lock
denies a conflicting lock, allows a non-overlapping range, and is reacquirable
after release (exercises the mount's `SetLk`/`GetLk`)
- **Distributed locking** (`-dlm`): a file held open for writing on one mount
blocks a writer on a second mount until it is released
- **Distributed-lock integrity**: concurrent writers to the same file from two
mounts leave exactly one intact payload, never a torn mix
- **Concurrency**: parallel writers to distinct files all succeed
Both FUSE mounts are started with `-dlm` (distributed lock manager). The second
mount (`/mnt/seaweedfs2`) exists only to contend with the smbd-backed mount in
the distributed-locking tests; both see the same filer path, so `.../share` is
the same data on each.
> Note on DLM semantics: `-dlm` coordinates *write access* (one mount writes a
> file at a time) and guarantees writes are not torn. It does not guarantee
> which concurrent writer wins or instant cross-mount read convergence — the
> holder's buffered data is flushed on close, asynchronously to lock release.
> When a holder closes the file, a writer on another mount acquires the freed
> lock within ~1s and completes.
## Layout
| File | Purpose |
| --- | --- |
| `smb_tests.sh` | SMB functional battery. Shared by both runners. |
| `lock_tests.sh` | SMB locking / concurrency battery. Shared by both runners. |
| `smb.conf.template` | Samba config; placeholders are filled in at run time. |
| `run.sh` | Local runner: `weed mini` + two `-dlm` mounts + `smbd` + both batteries, all as the current user on unprivileged ports. |
| `entrypoint.sh` | Container entrypoint: starts two `-dlm` FUSE mounts and runs `smbd`. |
| `run_inside_container.sh` | Runs both batteries inside the container against the local `smbd`. |
| `Dockerfile` | Adds Samba to the `chrislusf/seaweedfs:e2e` image. |
| `docker-compose.yml` | master + volume + filer + samba services. |
## Running locally
Requirements: `weed` on `$PATH`, `fusermount3`, and Samba's `smbd` /
`smbclient` / `smbpasswd` (Debian/Ubuntu: `apt-get install samba smbclient`).
```sh
test/samba/run.sh
```
No `sudo` is needed: `smbd` runs as the current user on port 4450 and all state
lives under a temp work dir that is cleaned up on exit.
## Running with Docker
Mirrors the CI job. Requires `/dev/fuse` and `SYS_ADMIN` (provided in the
compose file).
```sh
# build the base e2e image first (from the repo's docker/ dir)
docker compose -f test/samba/docker-compose.yml up --wait
docker compose -f test/samba/docker-compose.yml exec -T samba /run_inside_container.sh
docker compose -f test/samba/docker-compose.yml down -v
```
## CI
`.github/workflows/samba-integration.yml` runs on changes to `weed/mount/**`,
`weed/filer/**`, or `test/samba/**`. It builds the e2e image, builds the Samba
harness image on top, brings up the cluster, runs the battery, and uploads
server logs as artifacts.
## Notes
- The share disables Samba's DOS-attribute / xattr mapping and oplocks. The
SeaweedFS FUSE mount does not implement that surface, and leaving it on
produces `NT_STATUS_NOT_SUPPORTED` errors unrelated to data integrity.
- The share path is a subdirectory of the mount (`.../share`) so the runner can
verify SMB-side operations directly on the FUSE side.
+60
View File
@@ -0,0 +1,60 @@
services:
master:
image: chrislusf/seaweedfs:e2e
command: "-v=4 master -ip=master -ip.bind=0.0.0.0 -raftBootstrap"
healthcheck:
test: ["CMD", "curl", "--fail", "-I", "http://localhost:9333/cluster/healthz"]
interval: 2s
timeout: 10s
retries: 30
start_period: 10s
volume:
image: chrislusf/seaweedfs:e2e
command: "-v=4 volume -master=master:9333 -ip=volume -ip.bind=0.0.0.0 -preStopSeconds=1"
healthcheck:
test: ["CMD", "curl", "--fail", "-I", "http://localhost:8080/healthz"]
interval: 2s
timeout: 10s
retries: 15
start_period: 5s
depends_on:
master:
condition: service_healthy
filer:
image: chrislusf/seaweedfs:e2e
command: "-v=4 filer -master=master:9333 -ip=filer -ip.bind=0.0.0.0"
healthcheck:
test: ["CMD", "curl", "--fail", "-I", "http://localhost:8888/healthz"]
interval: 2s
timeout: 10s
retries: 15
start_period: 5s
depends_on:
volume:
condition: service_healthy
samba:
image: chrislusf/seaweedfs:samba
build:
context: .
environment:
FILER: filer:8888
cap_add:
- SYS_ADMIN
devices:
- /dev/fuse
security_opt:
- apparmor:unconfined
healthcheck:
test:
- "CMD-SHELL"
- "mountpoint -q /mnt/seaweedfs && mountpoint -q /mnt/seaweedfs2 && smbclient -L 127.0.0.1 -p 445 -U smbtest%smbtest -m SMB3 >/dev/null 2>&1"
interval: 3s
timeout: 10s
retries: 20
start_period: 15s
depends_on:
filer:
condition: service_healthy
+76
View File
@@ -0,0 +1,76 @@
#!/usr/bin/env bash
#
# Entrypoint for the samba test container.
#
# Mounts SeaweedFS over FUSE twice, both with distributed locking (-dlm) so the
# locking tests can exercise cross-mount write coordination:
# - MOUNT_DIR (/mnt/seaweedfs) is exported over SMB by smbd
# - MOUNT2_DIR (/mnt/seaweedfs2) is a second, independent mount of the same
# filer used to contend with the SMB writer
#
# Both mounts see the same filer path, so .../share is the same data on each.
# smbd runs in the foreground (as root, which owns the mounts), so the share
# uses "force user = root".
set -euo pipefail
FILER="${FILER:-filer:8888}"
MOUNT_DIR="${MOUNT_DIR:-/mnt/seaweedfs}"
MOUNT2_DIR="${MOUNT2_DIR:-/mnt/seaweedfs2}"
SHARE_DIR="${MOUNT_DIR}/share"
STATE_DIR="${STATE_DIR:-/var/lib/samba-test}"
SMB_PORT="${SMB_PORT:-445}"
SMB_USER="${SMB_USER:-smbtest}"
SMB_PASS="${SMB_PASS:-smbtest}"
mkdir -p "${MOUNT_DIR}" "${MOUNT2_DIR}" \
"${STATE_DIR}/private" "${STATE_DIR}/state" "${STATE_DIR}/cache" \
"${STATE_DIR}/lock" "${STATE_DIR}/pid" "${STATE_DIR}/ncalrpc"
# mount_seaweedfs <mountpoint> <logfile> — mount with -dlm and wait for it.
mount_seaweedfs() {
local dir="$1" log="$2"
echo "==> Mounting SeaweedFS (${FILER}) at ${dir} with -dlm"
weed -v=1 mount \
-filer="${FILER}" \
-dir="${dir}" \
-filer.path=/ \
-dirAutoCreate \
-allowOthers \
-dlm \
>"${log}" 2>&1 &
local pid=$!
for _ in $(seq 1 120); do
if mountpoint -q "${dir}"; then
return 0
fi
if ! kill -0 "${pid}" 2>/dev/null; then
echo "weed mount (${dir}) exited early; log tail:" >&2
tail -n 100 "${log}" >&2 || true
exit 1
fi
sleep 0.5
done
echo "FUSE mount ${dir} did not come up" >&2
tail -n 100 "${log}" >&2 || true
exit 1
}
mount_seaweedfs "${MOUNT_DIR}" /var/log/weed-mount.log
mount_seaweedfs "${MOUNT2_DIR}" /var/log/weed-mount2.log
mkdir -p "${SHARE_DIR}"
chmod 0777 "${SHARE_DIR}"
# --- configure and start smbd ----------------------------------------------
echo "==> Configuring Samba share on port ${SMB_PORT}"
sed -e "s#@SHARE_PATH@#${SHARE_DIR}#g" \
-e "s#@STATE_DIR@#${STATE_DIR}#g" \
-e "s#@SMB_PORT@#${SMB_PORT}#g" \
-e "s#@FORCE_USER@#root#g" \
/smb.conf.template >/etc/samba/smb.conf
id -u "${SMB_USER}" >/dev/null 2>&1 || useradd -M -s /usr/sbin/nologin "${SMB_USER}"
printf '%s\n%s\n' "${SMB_PASS}" "${SMB_PASS}" | smbpasswd -a -s "${SMB_USER}"
echo "==> Starting smbd"
exec smbd -F --no-process-group -s /etc/samba/smb.conf
+218
View File
@@ -0,0 +1,218 @@
#!/usr/bin/env bash
#
# Locking / concurrency test battery for Samba on a SeaweedFS FUSE mount.
#
# Covers the challenges a network-filesystem backend has to get right:
# 1. POSIX fcntl byte-range locking on the FUSE mount (SetLk/GetLk)
# 2. Distributed locking (-dlm): a write held open on one mount blocks a
# writer on another mount until it is released
# 3. Distributed locking integrity: concurrent writers to the same file from
# two mounts produce intact (non-torn) data
# 4. Concurrent writers to distinct files all succeed
#
# Required env:
# SMB_USER, SMB_PASS samba credentials
# MOUNT_SHARE dir on the smbd-backed FUSE mount (mount 1)
# MOUNT2_SHARE dir on the second FUSE mount (mount 2)
# Optional env:
# SMB_HOST (127.0.0.1), SMB_SHARE (seaweedfs), SMB_PORT (445)
set -uo pipefail
SMB_HOST="${SMB_HOST:-127.0.0.1}"
SMB_SHARE="${SMB_SHARE:-seaweedfs}"
SMB_PORT="${SMB_PORT:-445}"
SMB_USER="${SMB_USER:?SMB_USER is required}"
SMB_PASS="${SMB_PASS:?SMB_PASS is required}"
MOUNT_SHARE="${MOUNT_SHARE:?MOUNT_SHARE is required}"
MOUNT2_SHARE="${MOUNT2_SHARE:?MOUNT2_SHARE is required}"
WORK="$(mktemp -d /tmp/samba-locktest.XXXXXX)"
trap 'rm -rf "${WORK}"' EXIT
PASS=0
FAIL=0
pass() { printf ' [PASS] %s\n' "$1"; PASS=$((PASS + 1)); }
fail() { printf ' [FAIL] %s\n' "$1"; FAIL=$((FAIL + 1)); }
smb() {
smbclient "//${SMB_HOST}/${SMB_SHARE}" -p "${SMB_PORT}" \
-U "${SMB_USER}%${SMB_PASS}" -m SMB3 -c "$1"
}
md5() { md5sum "$1" | awk '{print $1}'; }
# 1. POSIX fcntl byte-range locking on the FUSE mount ------------------------
# Exercises the mount's SetLk/GetLk via two processes contending over fcntl
# (F_SETLK) byte-range locks. python3's fcntl.lockf issues real POSIX locks.
echo "==> 1. POSIX fcntl byte-range locking (FUSE mount SetLk/GetLk)"
lockfile="${MOUNT_SHARE}/fcntl_lock.dat"
: >"${lockfile}"
fcntl_out="$(python3 - "${lockfile}" <<'PY'
import fcntl, os, sys
path = sys.argv[1]
parent_to_child_r, parent_to_child_w = os.pipe() # release signal
child_to_parent_r, child_to_parent_w = os.pipe() # locked signal
pid = os.fork()
if pid == 0: # child: hold an exclusive lock on [0,100)
fd = os.open(path, os.O_RDWR | os.O_CREAT, 0o644)
fcntl.lockf(fd, fcntl.LOCK_EX, 100, 0, 0)
os.write(child_to_parent_w, b"L")
os.read(parent_to_child_r, 1) # wait until parent says release
fcntl.lockf(fd, fcntl.LOCK_UN, 100, 0, 0)
os.close(fd)
os._exit(0)
# parent
os.read(child_to_parent_r, 1) # wait until child holds the lock
fd = os.open(path, os.O_RDWR | os.O_CREAT, 0o644)
results = []
# a. a conflicting exclusive lock must be denied while the child holds it
try:
fcntl.lockf(fd, fcntl.LOCK_EX | fcntl.LOCK_NB, 100, 0, 0)
fcntl.lockf(fd, fcntl.LOCK_UN, 100, 0, 0)
results.append(("conflicting exclusive lock denied while held", False))
except OSError:
results.append(("conflicting exclusive lock denied while held", True))
# b. a non-overlapping range must be grantable
try:
fcntl.lockf(fd, fcntl.LOCK_EX | fcntl.LOCK_NB, 100, 200, 0)
fcntl.lockf(fd, fcntl.LOCK_UN, 100, 200, 0)
results.append(("non-overlapping range lock granted", True))
except OSError:
results.append(("non-overlapping range lock granted", False))
# c. after the holder releases, the lock must be acquirable
os.write(parent_to_child_w, b"R")
os.waitpid(pid, 0)
try:
fcntl.lockf(fd, fcntl.LOCK_EX | fcntl.LOCK_NB, 100, 0, 0)
fcntl.lockf(fd, fcntl.LOCK_UN, 100, 0, 0)
results.append(("lock acquirable after holder releases", True))
except OSError:
results.append(("lock acquirable after holder releases", False))
for name, ok in results:
print((" [PASS] " if ok else " [FAIL] ") + name)
sys.exit(0 if all(ok for _, ok in results) else 1)
PY
)"
echo "${fcntl_out}"
PASS=$((PASS + $(grep -c '\[PASS\]' <<<"${fcntl_out}")))
FAIL=$((FAIL + $(grep -c '\[FAIL\]' <<<"${fcntl_out}")))
# 2. Distributed lock blocks a cross-mount writer, then hands it off ----------
# mount 2 holds a file open for writing (holding the DLM lock on its path).
# An SMB put of the same file goes through mount 1 and must (a) block while
# mount 2 holds it and (b) succeed once mount 2 releases, leaving the SMB
# writer's payload on disk. smbclient gets a long client timeout (-t) so we are
# testing the lock handoff itself, not smbclient's own ~20s default timeout.
echo "==> 2. distributed lock: cross-mount write coordination"
dlmfile="dlm_coord.bin"
newdata="${WORK}/dlm_new.bin"
head -c 4096 /dev/urandom >"${newdata}"
# Hold the file open for writing on mount 2 via fd 9 -> holds the DLM lock.
exec 9>"${MOUNT2_SHARE}/${dlmfile}"
printf 'held-by-mount2' >&9
# Start the SMB write; record its real exit code when it returns. The subshell
# must NOT inherit fd 9 (9>&-): otherwise the SMB writer keeps the file open and
# waits on a DLM lock held by its own inherited descriptor, deadlocking the
# handoff this test is meant to exercise.
rm -f "${WORK}/dlm_put.rc"
(
smbclient "//${SMB_HOST}/${SMB_SHARE}" -p "${SMB_PORT}" \
-U "${SMB_USER}%${SMB_PASS}" -m SMB3 -t 120 \
-c "put ${newdata} ${dlmfile}" >/dev/null 2>&1
echo "$?" >"${WORK}/dlm_put.rc"
) 9>&- &
smb_bg=$!
sleep 4
if [[ ! -f "${WORK}/dlm_put.rc" ]]; then
pass "SMB write blocks while another mount holds the file open"
else
fail "SMB write returned early instead of blocking (rc=$(cat "${WORK}/dlm_put.rc"))"
fi
# Release mount 2's DLM lock; the blocked SMB write must now complete.
exec 9>&-
# Wait (bounded) for the SMB put to finish so a stuck handoff fails the test
# instead of hanging the suite.
put_rc="timeout"
for _ in $(seq 1 20); do
if [[ -f "${WORK}/dlm_put.rc" ]]; then
put_rc="$(cat "${WORK}/dlm_put.rc")"
break
fi
sleep 1
done
kill "${smb_bg}" 2>/dev/null
wait "${smb_bg}" 2>/dev/null
if [[ "${put_rc}" == "0" ]]; then
pass "blocked SMB write succeeds after the other mount releases"
else
fail "blocked SMB write succeeds after the other mount releases (rc=${put_rc})"
fi
# A correct handoff leaves the SMB writer's payload on disk: mount 1 acquired
# the lock and wrote after mount 2 released.
got="${WORK}/dlm_got.bin"
if smb "get ${dlmfile} ${got}" >/dev/null 2>&1 && [[ "$(md5 "${got}")" == "$(md5 "${newdata}")" ]]; then
pass "post-release content is the SMB writer's payload (correct handoff)"
else
fail "post-release content is the SMB writer's payload (correct handoff)"
fi
# 3. Distributed lock integrity: concurrent writers, same file ---------------
# An SMB writer (mount 1) and a direct writer (mount 2) race on one file. DLM
# serializes them, so the result must be exactly one of the two payloads.
echo "==> 3. distributed lock: concurrent writers produce intact data"
racefile="dlm_race.bin"
payloadA="${WORK}/dlm_raceA.bin"
head -c 1048576 /dev/urandom >"${payloadA}"
payloadB="direct-write-from-mount2-payload"
(smb "put ${payloadA} ${racefile}" >/dev/null 2>&1) &
(printf '%s' "${payloadB}" >"${MOUNT2_SHARE}/${racefile}") &
wait
racegot="${WORK}/dlm_race_got.bin"
if smb "get ${racefile} ${racegot}" >/dev/null 2>&1 &&
{ [[ "$(md5 "${racegot}")" == "$(md5 "${payloadA}")" ]] || [[ "$(cat "${racegot}")" == "${payloadB}" ]]; }; then
pass "concurrent same-file writers leave one intact payload"
else
fail "concurrent same-file writers leave one intact payload"
fi
# 4. Concurrent writers to distinct files ------------------------------------
echo "==> 4. concurrent writers to distinct files"
n=6
declare -a srcs=()
for i in $(seq 1 "${n}"); do
s="${WORK}/cc_${i}.bin"
head -c 1048576 /dev/urandom >"${s}"
srcs+=("${s}")
(smb "put ${s} concurrent_${i}.bin" >/dev/null 2>&1) &
done
wait
all_ok=true
for i in $(seq 1 "${n}"); do
g="${WORK}/cc_got_${i}.bin"
if ! smb "get concurrent_${i}.bin ${g}" >/dev/null 2>&1 ||
[[ "$(md5 "${srcs[$((i - 1))]}")" != "$(md5 "${g}")" ]]; then
all_ok=false
fi
done
if ${all_ok}; then
pass "${n} concurrent distinct-file writes all intact"
else
fail "${n} concurrent distinct-file writes all intact"
fi
echo
echo "==> Summary: ${PASS} passed, ${FAIL} failed"
[[ "${FAIL}" -eq 0 ]]
+203
View File
@@ -0,0 +1,203 @@
#!/usr/bin/env bash
#
# Run the SMB (Samba) integration test against a SeaweedFS FUSE mount.
#
# Pipeline:
# 1. start a self-contained "weed mini" (master + volume + filer in one)
# 2. mount the filesystem with "weed mount"
# 3. export a subdirectory of the mount over SMB with smbd
# 4. drive the share with smbclient (test/samba/smb_tests.sh)
#
# Everything runs as the current user on unprivileged ports, so no sudo is
# required. State lives under a temp work dir and is removed on exit.
#
# Requirements: weed in $PATH, fusermount3, and Samba's smbd / smbclient /
# smbpasswd (Debian/Ubuntu: apt-get install samba smbclient).
#
# Usage:
# test/samba/run.sh
set -euo pipefail
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
WEED_BIN="${WEED_BIN:-weed}"
WORK_DIR="${WORK_DIR:-$(mktemp -d /tmp/seaweedfs-samba.XXXXXX)}"
MOUNT_DIR="${MOUNT_DIR:-${WORK_DIR}/mnt}"
MOUNT2_DIR="${MOUNT2_DIR:-${WORK_DIR}/mnt2}"
DATA_DIR="${DATA_DIR:-${WORK_DIR}/data}"
LOG_DIR="${LOG_DIR:-${WORK_DIR}/logs}"
STATE_DIR="${WORK_DIR}/samba"
SHARE_DIR="${MOUNT_DIR}/share"
SHARE_DIR2="${MOUNT2_DIR}/share"
FILER_PORT="${FILER_PORT:-28888}"
FILER_ADDR="127.0.0.1:${FILER_PORT}"
SMB_PORT="${SMB_PORT:-4450}"
SMB_SHARE="seaweedfs"
SMB_USER="${SMB_USER:-$(id -un)}"
SMB_PASS="${SMB_PASS:-seaweedfs}"
SMBD_BIN="$(command -v smbd || echo /usr/sbin/smbd)"
SMBPASSWD_BIN="$(command -v smbpasswd || echo /usr/bin/smbpasswd)"
CI_LOG_DIR="/tmp/seaweedfs-samba-logs"
mini_pid=""
mount_pid=""
mount2_pid=""
smbd_pid=""
unmount_dir() {
local dir="$1"
if mountpoint -q "${dir}" 2>/dev/null; then
fusermount3 -u "${dir}" 2>/dev/null ||
fusermount -u "${dir}" 2>/dev/null || true
fi
}
cleanup() {
set +e
if [[ -n "${smbd_pid}" ]] && kill -0 "${smbd_pid}" 2>/dev/null; then
kill -TERM "${smbd_pid}" 2>/dev/null || true
wait "${smbd_pid}" 2>/dev/null || true
fi
for p in "${mount_pid}" "${mount2_pid}"; do
if [[ -n "${p}" ]] && kill -0 "${p}" 2>/dev/null; then
kill -TERM "${p}" 2>/dev/null || true
wait "${p}" 2>/dev/null || true
fi
done
unmount_dir "${MOUNT_DIR}"
unmount_dir "${MOUNT2_DIR}"
if [[ -n "${mini_pid}" ]] && kill -0 "${mini_pid}" 2>/dev/null; then
kill -TERM "${mini_pid}" 2>/dev/null || true
wait "${mini_pid}" 2>/dev/null || true
fi
# Copy logs to a fixed path for CI artifact upload.
mkdir -p "${CI_LOG_DIR}"
cp "${LOG_DIR}"/*.log "${LOG_DIR}"/*.out "${STATE_DIR}/smbd.log" "${CI_LOG_DIR}/" 2>/dev/null || true
}
trap cleanup EXIT INT TERM
mkdir -p "${MOUNT_DIR}" "${MOUNT2_DIR}" "${DATA_DIR}" "${LOG_DIR}" \
"${STATE_DIR}/private" "${STATE_DIR}/state" "${STATE_DIR}/cache" \
"${STATE_DIR}/lock" "${STATE_DIR}/pid" "${STATE_DIR}/ncalrpc"
# --- 1. weed mini -----------------------------------------------------------
echo "==> Starting weed mini on ${FILER_ADDR}"
"${WEED_BIN}" mini \
-dir="${DATA_DIR}" \
-ip=127.0.0.1 \
-filer.port="${FILER_PORT}" \
-s3=false \
-webdav=false \
-admin.ui=false \
>"${LOG_DIR}/mini.log" 2>&1 &
mini_pid=$!
for i in $(seq 1 60); do
if (echo >"/dev/tcp/127.0.0.1/${FILER_PORT}") 2>/dev/null; then
break
fi
if ! kill -0 "${mini_pid}" 2>/dev/null; then
echo "weed mini exited early; log tail:" >&2
tail -n 100 "${LOG_DIR}/mini.log" >&2 || true
exit 1
fi
sleep 0.5
done
if ! (echo >"/dev/tcp/127.0.0.1/${FILER_PORT}") 2>/dev/null; then
echo "weed mini filer did not become reachable within 30s; log tail:" >&2
tail -n 100 "${LOG_DIR}/mini.log" >&2 || true
exit 1
fi
# --- 2. weed mount (two mounts, both with -dlm) -----------------------------
# mount_with_dlm <mountpoint> <logfile> <pid-var-name>
mount_with_dlm() {
local dir="$1" log="$2" pidvar="$3" pid
echo "==> Mounting SeaweedFS at ${dir} with -dlm"
"${WEED_BIN}" mount \
-filer="${FILER_ADDR}" \
-dir="${dir}" \
-filer.path=/ \
-dirAutoCreate \
-dlm \
>"${log}" 2>&1 &
pid=$!
printf -v "${pidvar}" '%s' "${pid}"
for _ in $(seq 1 60); do
if mountpoint -q "${dir}"; then
return 0
fi
if ! kill -0 "${pid}" 2>/dev/null; then
echo "weed mount (${dir}) exited early; log tail:" >&2
tail -n 100 "${log}" >&2 || true
exit 1
fi
sleep 0.5
done
echo "FUSE mount ${dir} did not come up within 30s" >&2
tail -n 100 "${log}" >&2 || true
exit 1
}
mount_with_dlm "${MOUNT_DIR}" "${LOG_DIR}/mount.log" mount_pid
mount_with_dlm "${MOUNT2_DIR}" "${LOG_DIR}/mount2.log" mount2_pid
mkdir -p "${SHARE_DIR}"
# --- 3. smbd ----------------------------------------------------------------
echo "==> Generating smb.conf and starting smbd on port ${SMB_PORT}"
SMB_CONF="${STATE_DIR}/smb.conf"
sed -e "s#@SHARE_PATH@#${SHARE_DIR}#g" \
-e "s#@STATE_DIR@#${STATE_DIR}#g" \
-e "s#@SMB_PORT@#${SMB_PORT}#g" \
-e "s#@FORCE_USER@#${SMB_USER}#g" \
"${SCRIPT_DIR}/smb.conf.template" >"${SMB_CONF}"
printf '%s\n%s\n' "${SMB_PASS}" "${SMB_PASS}" |
"${SMBPASSWD_BIN}" -c "${SMB_CONF}" -a -s "${SMB_USER}"
"${SMBD_BIN}" -F --no-process-group -s "${SMB_CONF}" >"${LOG_DIR}/smbd.out" 2>&1 &
smbd_pid=$!
for i in $(seq 1 60); do
if (echo >"/dev/tcp/127.0.0.1/${SMB_PORT}") 2>/dev/null; then
break
fi
if ! kill -0 "${smbd_pid}" 2>/dev/null; then
echo "smbd exited early; log tail:" >&2
tail -n 100 "${LOG_DIR}/smbd.out" "${STATE_DIR}/smbd.log" 2>/dev/null >&2 || true
exit 1
fi
sleep 0.5
done
if ! (echo >"/dev/tcp/127.0.0.1/${SMB_PORT}") 2>/dev/null; then
echo "smbd did not become reachable within 30s; log tail:" >&2
tail -n 100 "${LOG_DIR}/smbd.out" "${STATE_DIR}/smbd.log" 2>/dev/null >&2 || true
exit 1
fi
# --- 4. run the test batteries ---------------------------------------------
rc=0
echo "==> Running SMB functional test battery"
SMB_HOST=127.0.0.1 \
SMB_SHARE="${SMB_SHARE}" \
SMB_PORT="${SMB_PORT}" \
SMB_USER="${SMB_USER}" \
SMB_PASS="${SMB_PASS}" \
SHARE_FS_PATH="${SHARE_DIR}" \
"${SCRIPT_DIR}/smb_tests.sh" || rc=1
echo "==> Running SMB locking / concurrency test battery"
SMB_HOST=127.0.0.1 \
SMB_SHARE="${SMB_SHARE}" \
SMB_PORT="${SMB_PORT}" \
SMB_USER="${SMB_USER}" \
SMB_PASS="${SMB_PASS}" \
MOUNT_SHARE="${SHARE_DIR}" \
MOUNT2_SHARE="${SHARE_DIR2}" \
"${SCRIPT_DIR}/lock_tests.sh" || rc=1
exit "${rc}"
+24
View File
@@ -0,0 +1,24 @@
#!/usr/bin/env bash
#
# Runs the SMB test batteries inside the samba container against the local smbd,
# which serves /mnt/seaweedfs/share over a SeaweedFS FUSE mount. A second FUSE
# mount (/mnt/seaweedfs2) backs the distributed-locking tests.
# Invoked via: docker compose exec samba /run_inside_container.sh
set -euo pipefail
export SMB_HOST=127.0.0.1
export SMB_SHARE=seaweedfs
export SMB_PORT="${SMB_PORT:-445}"
export SMB_USER="${SMB_USER:-smbtest}"
export SMB_PASS="${SMB_PASS:-smbtest}"
export SHARE_FS_PATH="${SHARE_FS_PATH:-/mnt/seaweedfs/share}"
export MOUNT_SHARE="${MOUNT_SHARE:-/mnt/seaweedfs/share}"
export MOUNT2_SHARE="${MOUNT2_SHARE:-/mnt/seaweedfs2/share}"
rc=0
echo "############ SMB functional tests ############"
/smb_tests.sh || rc=1
echo
echo "############ SMB locking / concurrency tests ############"
/lock_tests.sh || rc=1
exit "${rc}"
+52
View File
@@ -0,0 +1,52 @@
[global]
server role = standalone server
workgroup = WORKGROUP
server string = SeaweedFS FUSE Samba test
security = user
server min protocol = SMB2
smb ports = @SMB_PORT@
bind interfaces only = yes
interfaces = lo 127.0.0.1
# Self-contained state so smbd can run rootless and leaves nothing behind
# outside the test work directory.
private dir = @STATE_DIR@/private
state directory = @STATE_DIR@/state
cache directory = @STATE_DIR@/cache
lock directory = @STATE_DIR@/lock
pid directory = @STATE_DIR@/pid
ncalrpc dir = @STATE_DIR@/ncalrpc
log file = @STATE_DIR@/smbd.log
log level = 1
usershare max shares = 0
# No printing subsystem in a file-server test.
load printers = no
printing = bsd
printcap name = /dev/null
disable spoolss = yes
# The SeaweedFS FUSE mount does not implement the full xattr / DOS-attribute
# surface Samba uses by default. Disabling these avoids spurious
# NT_STATUS_NOT_SUPPORTED / EOPNOTSUPP errors unrelated to data integrity.
ea support = no
store dos attributes = no
map archive = no
map hidden = no
map system = no
map readonly = no
# A network-filesystem backend should not advertise local oplocks/leases.
oplocks = no
level2 oplocks = no
kernel oplocks = no
posix locking = no
[seaweedfs]
path = @SHARE_PATH@
comment = SeaweedFS share backed by a FUSE mount
browseable = yes
read only = no
create mask = 0644
directory mask = 0755
force user = @FORCE_USER@
+172
View File
@@ -0,0 +1,172 @@
#!/usr/bin/env bash
#
# SMB protocol test battery against a Samba share backed by a SeaweedFS FUSE
# mount. Driven both by the local runner (test/samba/run.sh) and by the Docker
# harness (run_inside_container.sh).
#
# Required env:
# SMB_USER samba username
# SMB_PASS samba password
# Optional env:
# SMB_HOST samba host (default 127.0.0.1)
# SMB_SHARE share name (default seaweedfs)
# SMB_PORT smbd port (default 445)
# SHARE_FS_PATH directory on the FUSE mount that backs the share. When set,
# the suite also checks cross-protocol consistency: data written
# over SMB is visible on the FUSE mount, and vice versa.
set -uo pipefail
SMB_HOST="${SMB_HOST:-127.0.0.1}"
SMB_SHARE="${SMB_SHARE:-seaweedfs}"
SMB_PORT="${SMB_PORT:-445}"
SMB_USER="${SMB_USER:?SMB_USER is required}"
SMB_PASS="${SMB_PASS:?SMB_PASS is required}"
SHARE_FS_PATH="${SHARE_FS_PATH:-}"
WORK="$(mktemp -d /tmp/samba-smbtest.XXXXXX)"
trap 'rm -rf "${WORK}"' EXIT
PASS=0
FAIL=0
pass() { printf ' [PASS] %s\n' "$1"; PASS=$((PASS + 1)); }
fail() { printf ' [FAIL] %s\n' "$1"; FAIL=$((FAIL + 1)); }
# Run one or more smbclient commands (separated by ';') against the share.
smb() {
smbclient "//${SMB_HOST}/${SMB_SHARE}" -p "${SMB_PORT}" \
-U "${SMB_USER}%${SMB_PASS}" -m SMB3 -c "$1"
}
md5() { md5sum "$1" | awk '{print $1}'; }
echo "==> Target //${SMB_HOST}/${SMB_SHARE} (port ${SMB_PORT}) as ${SMB_USER}"
[[ -n "${SHARE_FS_PATH}" ]] && echo "==> Cross-protocol checks against ${SHARE_FS_PATH}"
# 1. Connectivity ------------------------------------------------------------
echo "==> 1. connectivity"
if smb "ls" >/dev/null 2>&1; then
pass "connect and list share root"
else
fail "connect and list share root"
fi
# 2. Upload / download round-trip -------------------------------------------
echo "==> 2. upload / download round-trip"
src="${WORK}/src.bin"
head -c 1048576 /dev/urandom >"${src}" # 1 MiB
if smb "put ${src} roundtrip.bin" >/dev/null 2>&1; then
pass "put 1 MiB file"
else
fail "put 1 MiB file"
fi
got="${WORK}/got.bin"
if smb "get roundtrip.bin ${got}" >/dev/null 2>&1 && [[ "$(md5 "${src}")" == "$(md5 "${got}")" ]]; then
pass "get returns identical content"
else
fail "get returns identical content"
fi
if [[ -n "${SHARE_FS_PATH}" ]]; then
if [[ -f "${SHARE_FS_PATH}/roundtrip.bin" ]] && [[ "$(md5 "${SHARE_FS_PATH}/roundtrip.bin")" == "$(md5 "${src}")" ]]; then
pass "SMB-written file visible on FUSE mount with identical content"
else
fail "SMB-written file visible on FUSE mount with identical content"
fi
fi
# 3. Directory operations ----------------------------------------------------
echo "==> 3. directory operations"
if smb "mkdir docs; cd docs; put ${src} nested.bin; ls" >/dev/null 2>&1; then
pass "mkdir + put into subdirectory"
else
fail "mkdir + put into subdirectory"
fi
if [[ -z "${SHARE_FS_PATH}" || -f "${SHARE_FS_PATH}/docs/nested.bin" ]]; then
pass "nested file present"
else
fail "nested file present"
fi
# 4. Rename ------------------------------------------------------------------
echo "==> 4. rename"
if smb "rename roundtrip.bin renamed.bin" >/dev/null 2>&1; then
pass "rename file"
else
fail "rename file"
fi
renback="${WORK}/renamed.bin"
if smb "get renamed.bin ${renback}" >/dev/null 2>&1 && [[ "$(md5 "${renback}")" == "$(md5 "${src}")" ]]; then
pass "renamed file readable with original content"
else
fail "renamed file readable with original content"
fi
if [[ -n "${SHARE_FS_PATH}" ]]; then
if [[ -f "${SHARE_FS_PATH}/renamed.bin" && ! -e "${SHARE_FS_PATH}/roundtrip.bin" ]]; then
pass "rename reflected on FUSE mount"
else
fail "rename reflected on FUSE mount"
fi
fi
# 5. Large file (exercises SeaweedFS chunking) -------------------------------
echo "==> 5. large file (SeaweedFS chunking)"
big="${WORK}/big.bin"
head -c 67108864 /dev/urandom >"${big}" # 64 MiB
bigback="${WORK}/big.back"
if smb "put ${big} big.bin" >/dev/null 2>&1 &&
smb "get big.bin ${bigback}" >/dev/null 2>&1 &&
[[ "$(md5 "${big}")" == "$(md5 "${bigback}")" ]]; then
pass "64 MiB put/get round-trip"
else
fail "64 MiB put/get round-trip"
fi
# 6. Recursive upload --------------------------------------------------------
echo "==> 6. recursive upload"
tree="${WORK}/tree"
mkdir -p "${tree}/a/b"
echo one >"${tree}/f1.txt"
echo two >"${tree}/a/f2.txt"
echo three >"${tree}/a/b/f3.txt"
if (cd "${WORK}" && smb "recurse ON; prompt OFF; mput tree" >/dev/null 2>&1) &&
{ [[ -z "${SHARE_FS_PATH}" ]] || [[ -f "${SHARE_FS_PATH}/tree/a/b/f3.txt" ]]; }; then
pass "recursive mput"
else
fail "recursive mput"
fi
# 7. Cross-protocol read (FUSE writes, SMB reads) ----------------------------
if [[ -n "${SHARE_FS_PATH}" ]]; then
echo "==> 7. cross-protocol read (FUSE write -> SMB read)"
echo "written-via-fuse" >"${SHARE_FS_PATH}/from_fuse.txt"
cpb="${WORK}/from_fuse.back"
if smb "get from_fuse.txt ${cpb}" >/dev/null 2>&1 && grep -q written-via-fuse "${cpb}"; then
pass "FUSE-written file readable over SMB"
else
fail "FUSE-written file readable over SMB"
fi
fi
# 8. Delete ------------------------------------------------------------------
echo "==> 8. delete"
smb "del renamed.bin" >/dev/null 2>&1
smb "del big.bin" >/dev/null 2>&1
smb "deltree docs" >/dev/null 2>&1
smb "deltree tree" >/dev/null 2>&1
if [[ -n "${SHARE_FS_PATH}" ]]; then
if [[ ! -e "${SHARE_FS_PATH}/renamed.bin" && ! -e "${SHARE_FS_PATH}/big.bin" &&
! -e "${SHARE_FS_PATH}/docs" && ! -e "${SHARE_FS_PATH}/tree" ]]; then
pass "delete files and directory trees"
else
fail "delete files and directory trees"
fi
else
if ! smb "get renamed.bin /dev/null" >/dev/null 2>&1; then
pass "deleted file no longer retrievable"
else
fail "deleted file no longer retrievable"
fi
fi
echo
echo "==> Summary: ${PASS} passed, ${FAIL} failed"
[[ "${FAIL}" -eq 0 ]]
+12 -5
View File
@@ -51,6 +51,13 @@ func MustAllocatePorts(t *testing.T, count int) []int {
// from recycling ports between allocations. Use this when ports will be
// passed to weed mini without explicit gRPC port flags, so mini will
// derive gRPC ports as HTTP + 10000.
//
// Listeners are bound on all interfaces (":port") rather than 127.0.0.1
// to match weed mini's availability check (isPortAvailable). A port can
// be free on loopback but held by another process on a different
// interface; reserving only on loopback lets mini's check fail and
// trigger gRPC port shifting, which then causes weed shell to dial the
// wrong port and hang.
func AllocateMiniPorts(count int) ([]int, error) {
const (
minPort = 10000
@@ -75,12 +82,12 @@ func AllocateMiniPorts(count int) ([]int, error) {
continue
}
l1, err := net.Listen("tcp", fmt.Sprintf("127.0.0.1:%d", port))
l1, err := net.Listen("tcp", fmt.Sprintf(":%d", port))
if err != nil {
continue
}
l2, err := net.Listen("tcp", fmt.Sprintf("127.0.0.1:%d", grpcPort))
l2, err := net.Listen("tcp", fmt.Sprintf(":%d", grpcPort))
if err != nil {
l1.Close()
continue
@@ -145,11 +152,11 @@ func AllocatePortSet(miniCount, regularCount int) (mini []int, regular []int, er
if reserved[port] || reserved[grpcPort] {
continue
}
l1, lErr := net.Listen("tcp", fmt.Sprintf("127.0.0.1:%d", port))
l1, lErr := net.Listen("tcp", fmt.Sprintf(":%d", port))
if lErr != nil {
continue
}
l2, lErr := net.Listen("tcp", fmt.Sprintf("127.0.0.1:%d", grpcPort))
l2, lErr := net.Listen("tcp", fmt.Sprintf(":%d", grpcPort))
if lErr != nil {
l1.Close()
continue
@@ -168,7 +175,7 @@ func AllocatePortSet(miniCount, regularCount int) (mini []int, regular []int, er
regular = make([]int, 0, regularCount)
for i := 0; i < regularCount; i++ {
l, lErr := net.Listen("tcp", "127.0.0.1:0")
l, lErr := net.Listen("tcp", ":0")
if lErr != nil {
return nil, nil, lErr
}
@@ -136,8 +136,6 @@ func TestEcLifecycleAcrossMultipleDisks(t *testing.T) {
conn2, grpcClient2 := framework.DialVolumeServer(t, clusterHarness.VolumeGRPCAddress())
defer conn2.Close()
// VolumeEcShardsInfo only sees one disk's EcVolume; filesystem layout is
// the ground truth for the whole-store shard count.
postReconcileLayout := scanShardLayout(t, dataDirs, collection, volumeID)
if got, want := totalShardsInLayout(postReconcileLayout), erasure_coding.TotalShardsCount; got != want {
t.Fatalf("post-reconcile: total shards on disk mismatch: got %d, want %d (layout=%v)", got, want, postReconcileLayout)
@@ -148,11 +146,28 @@ func TestEcLifecycleAcrossMultipleDisks(t *testing.T) {
if got, want := len(postReconcileLayout[1]), splitAt; got != want {
t.Fatalf("post-reconcile: disk 1 shard count drift: got %d, want %d (layout=%v)", got, want, postReconcileLayout)
}
if _, err := grpcClient2.VolumeEcShardsInfo(ctx, &volume_server_pb.VolumeEcShardsInfoRequest{
// VolumeEcShardsInfo must walk every DiskLocation and report the full
// shard set — the verification step in ec_task.go gates source-volume
// deletion on this RPC returning a complete shard inventory.
infoResp, err := grpcClient2.VolumeEcShardsInfo(ctx, &volume_server_pb.VolumeEcShardsInfoRequest{
VolumeId: volumeID,
}); err != nil {
})
if err != nil {
t.Fatalf("VolumeEcShardsInfo after redistribute restart: %v", err)
}
if got, want := len(infoResp.GetEcShardInfos()), erasure_coding.TotalShardsCount; got != want {
t.Fatalf("VolumeEcShardsInfo after redistribute restart: got %d shards, want %d (per-disk layout=%v)",
got, want, postReconcileLayout)
}
gotShardIds := make(map[uint32]struct{}, len(infoResp.GetEcShardInfos()))
for _, info := range infoResp.GetEcShardInfos() {
gotShardIds[info.GetShardId()] = struct{}{}
}
for shardId := uint32(0); shardId < uint32(erasure_coding.TotalShardsCount); shardId++ {
if _, ok := gotShardIds[shardId]; !ok {
t.Fatalf("VolumeEcShardsInfo missing shard %d (per-disk layout=%v)", shardId, postReconcileLayout)
}
}
for _, n := range needles {
verifyHTTPRead(t, httpClient, clusterHarness.VolumeAdminURL(), n.fid, n.payload, "after-cross-disk-reconcile")
}
@@ -205,8 +220,21 @@ func TestEcLifecycleAcrossMultipleDisks(t *testing.T) {
t.Fatalf("VolumeEcShardsRebuild expected to rebuild shard %d, got %v",
repairTargetShard, rebuildResp.GetRebuiltShardIds())
}
if _, statErr := os.Stat(shardPath); statErr != nil {
t.Fatalf("rebuild did not restore shard %d on disk 0 (%s): %v", repairTargetShard, shardPath, statErr)
// The rebuilder picks whichever disk hosts the most shards plus a
// matching .ecx. After the boot-time mirror runs, every shard-
// bearing disk owns its own sidecars, so rebuild can legitimately
// restore the shard on either disk — accept both.
rebuiltOn := -1
for i, dir := range dataDirs {
candidate := filepath.Join(dir, shardFileName(collection, volumeID, repairTargetShard))
if _, statErr := os.Stat(candidate); statErr == nil {
rebuiltOn = i
shardPath = candidate
break
}
}
if rebuiltOn < 0 {
t.Fatalf("rebuild did not restore shard %d on any disk (checked %v)", repairTargetShard, dataDirs)
}
if _, err := grpcClient3.VolumeEcShardsMount(ctx, &volume_server_pb.VolumeEcShardsMountRequest{
@@ -0,0 +1,175 @@
package volume_server_grpc_test
import (
"context"
"net/http"
"sort"
"testing"
"time"
"github.com/seaweedfs/seaweedfs/test/volume_server/framework"
"github.com/seaweedfs/seaweedfs/test/volume_server/matrix"
"github.com/seaweedfs/seaweedfs/weed/pb/volume_server_pb"
"github.com/seaweedfs/seaweedfs/weed/storage/erasure_coding"
"google.golang.org/grpc"
"google.golang.org/grpc/credentials/insecure"
)
// TestVolumeEcShardsInfoReturnsAllShardsAcrossDisks drives the full path
// behind the ec.encode source-deletion gate. A multi-disk volume server
// ends up with EC shards split across disks (each registers its own
// EcVolume entry in DiskLocation.ecVolumes), and the volume server's
// VolumeEcShardsInfo RPC must walk every DiskLocation rather than
// reporting whichever disk Store.FindEcVolume picks first.
//
// Pre-fix, verifyEcShardsBeforeDelete refused to delete source volumes —
// the shard-bitmap union across destinations fell short of dataShards +
// parityShards because each destination only reported shards on one of
// its disks. With the handler fix, the same VerifyShardsAcrossServers
// call returns a complete bitmap and the gate opens.
func TestVolumeEcShardsInfoReturnsAllShardsAcrossDisks(t *testing.T) {
if testing.Short() {
t.Skip("skipping integration test in short mode")
}
const (
dataDirCount = 2
volumeID = uint32(9558)
collection = "ec-multi-disk-verify"
)
clusterHarness := framework.StartSingleVolumeClusterWithDataDirs(t, matrix.P1(), dataDirCount)
dataDirs := clusterHarness.VolumeDataDirs()
if len(dataDirs) != dataDirCount {
t.Fatalf("expected %d data dirs, got %d: %v", dataDirCount, len(dataDirs), dataDirs)
}
conn, grpcClient := framework.DialVolumeServer(t, clusterHarness.VolumeGRPCAddress())
defer conn.Close()
ctx, cancel := context.WithTimeout(context.Background(), 60*time.Second)
defer cancel()
framework.AllocateVolume(t, grpcClient, volumeID, collection)
httpClient := framework.NewHTTPClient()
needles := []struct {
fid string
payload []byte
}{
{framework.NewFileID(volumeID, 9559, 0xC0FFEE01), bytesOfLen(64, 0xB1)},
{framework.NewFileID(volumeID, 9560, 0xC0FFEE02), bytesOfLen(8192, 0xB2)},
{framework.NewFileID(volumeID, 9561, 0xC0FFEE03), bytesOfLen(131072, 0xB3)},
}
for _, n := range needles {
resp := framework.UploadBytes(t, httpClient, clusterHarness.VolumeAdminURL(), n.fid, n.payload)
_ = framework.ReadAllAndClose(t, resp)
if resp.StatusCode != http.StatusCreated {
t.Fatalf("upload %s expected 201, got %d", n.fid, resp.StatusCode)
}
}
if _, err := grpcClient.VolumeEcShardsGenerate(ctx, &volume_server_pb.VolumeEcShardsGenerateRequest{
VolumeId: volumeID,
Collection: collection,
}); err != nil {
t.Fatalf("VolumeEcShardsGenerate: %v", err)
}
// Generate places every shard plus the .ecx/.ecj/.vif on the .dat's
// disk (disk 0). Mount all 14 there first so the next step's restart
// has a steady starting state.
allShards := make([]uint32, erasure_coding.TotalShardsCount)
for i := range allShards {
allShards[i] = uint32(i)
}
if _, err := grpcClient.VolumeEcShardsMount(ctx, &volume_server_pb.VolumeEcShardsMountRequest{
VolumeId: volumeID,
Collection: collection,
ShardIds: allShards,
}); err != nil {
t.Fatalf("VolumeEcShardsMount all shards: %v", err)
}
// Drop the .dat so the EC shards are the only data path — mirrors the
// real ec.encode flow before verifyEcShardsBeforeDelete fires.
if _, err := grpcClient.VolumeDelete(ctx, &volume_server_pb.VolumeDeleteRequest{
VolumeId: volumeID,
}); err != nil {
t.Fatalf("VolumeDelete (drop .dat): %v", err)
}
// Move half the shards onto disk 1, leaving .ecx on disk 0. After
// restart, the cross-disk reconcile path attaches each disk's shards
// against its own EcVolume entry — the exact in-memory shape the bug
// reporter saw on a multi-disk destination.
clusterHarness.StopVolumeServer()
const splitAt = 7
for shard := 0; shard < splitAt; shard++ {
movedFile(t, dataDirs[0], dataDirs[1], collection, volumeID, erasure_coding.ToExt(shard))
}
if fileExistsIn(dataDirs[1], collection, volumeID, ".ecx") {
t.Fatalf("setup: .ecx must stay on disk 0 to exercise the multi-disk path")
}
clusterHarness.RestartVolumeServer()
conn2, grpcClient2 := framework.DialVolumeServer(t, clusterHarness.VolumeGRPCAddress())
defer conn2.Close()
postReconcileLayout := scanShardLayout(t, dataDirs, collection, volumeID)
if got, want := totalShardsInLayout(postReconcileLayout), erasure_coding.TotalShardsCount; got != want {
t.Fatalf("post-reconcile: total shards on disk mismatch: got %d, want %d (layout=%v)", got, want, postReconcileLayout)
}
if len(postReconcileLayout[0]) == 0 || len(postReconcileLayout[1]) == 0 {
t.Fatalf("post-reconcile: expected shards on BOTH disks, got per-disk layout %v", postReconcileLayout)
}
// Direct RPC assertion: VolumeEcShardsInfo must report every shard
// the server holds, not just the ones registered against the first
// matching DiskLocation.
infoResp, err := grpcClient2.VolumeEcShardsInfo(ctx, &volume_server_pb.VolumeEcShardsInfoRequest{
VolumeId: volumeID,
})
if err != nil {
t.Fatalf("VolumeEcShardsInfo: %v", err)
}
gotShardIds := make([]int, 0, len(infoResp.GetEcShardInfos()))
for _, info := range infoResp.GetEcShardInfos() {
if info.GetVolumeId() != volumeID {
t.Errorf("EcShardInfo VolumeId=%d, want %d", info.GetVolumeId(), volumeID)
}
gotShardIds = append(gotShardIds, int(info.GetShardId()))
}
sort.Ints(gotShardIds)
wantShardIds := make([]int, erasure_coding.TotalShardsCount)
for i := range wantShardIds {
wantShardIds[i] = i
}
if len(gotShardIds) != len(wantShardIds) {
t.Fatalf("VolumeEcShardsInfo returned %d shards (ids=%v), want %d (ids=%v) — per-disk layout=%v",
len(gotShardIds), gotShardIds, len(wantShardIds), wantShardIds, postReconcileLayout)
}
for i, sid := range wantShardIds {
if gotShardIds[i] != sid {
t.Fatalf("VolumeEcShardsInfo shard ids=%v, want %v (per-disk layout=%v)",
gotShardIds, wantShardIds, postReconcileLayout)
}
}
// End-to-end assertion via the same helper the worker uses to gate
// source-volume deletion (weed/worker/tasks/erasure_coding/ec_task.go
// verifyEcShardsBeforeDelete). The union across destinations is what
// RequireFullShardSet measures; with one destination that holds every
// shard, the union must cover dataShards + parityShards.
dialOption := grpc.WithTransportCredentials(insecure.NewCredentials())
servers := []string{clusterHarness.VolumeServerAddress()}
union, perServer := erasure_coding.VerifyShardsAcrossServers(ctx, volumeID, servers, dialOption)
if err := erasure_coding.RequireFullShardSet(volumeID, union, erasure_coding.TotalShardsCount); err != nil {
t.Fatalf("verifyEcShardsBeforeDelete-equivalent gate failed: %v\nper-server inventory: %s\nper-disk layout: %v",
err, erasure_coding.SummarizeShardInventory(perServer), postReconcileLayout)
}
if got, want := union.Count(), erasure_coding.TotalShardsCount; got != want {
t.Fatalf("VerifyShardsAcrossServers union covered %d/%d shards (per-server=%s, layout=%v)",
got, want, erasure_coding.SummarizeShardInventory(perServer), postReconcileLayout)
}
}
@@ -2,6 +2,8 @@ package volume_server_http_test
import (
"context"
"encoding/json"
"fmt"
"net/http"
"testing"
"time"
@@ -44,6 +46,17 @@ func TestReplicatedUploadSucceedsImmediatelyAfterAllocate(t *testing.T) {
fid := framework.NewFileID(volumeID, 881001, 0x0B0C0D0E)
payload := []byte("replicated-upload-after-allocate")
// The master only learns about replica locations through volume-server
// heartbeats, which lag behind the direct AllocateVolume gRPC calls above.
// In production a client obtains its fid from the master assign flow, which
// guarantees the master already knows every replica; this test crafts the
// fid by hand, so the replicated write would otherwise look up the master
// before the second replica is registered and fail with a 500. Wait until
// the master reports both replicas before uploading.
if !waitForMasterReplicaCount(t, client, clusterHarness.MasterURL(), volumeID, 2, 10*time.Second) {
t.Fatalf("master did not report 2 replica locations for volume %d within deadline", volumeID)
}
uploadResp := framework.UploadBytes(t, client, clusterHarness.VolumeAdminURL(0), fid, payload)
_ = framework.ReadAllAndClose(t, uploadResp)
if uploadResp.StatusCode != http.StatusCreated {
@@ -61,3 +74,29 @@ func TestReplicatedUploadSucceedsImmediatelyAfterAllocate(t *testing.T) {
t.Fatalf("replica body mismatch: got %q want %q", string(replicaBody), string(payload))
}
}
// waitForMasterReplicaCount polls the master volume lookup until it reports at
// least want locations for volumeID, or the timeout elapses.
func waitForMasterReplicaCount(t testing.TB, client *http.Client, masterURL string, volumeID uint32, want int, timeout time.Duration) bool {
t.Helper()
lookupURL := fmt.Sprintf("%s/dir/lookup?volumeId=%d", masterURL, volumeID)
deadline := time.Now().Add(timeout)
for time.Now().Before(deadline) {
resp := framework.DoRequest(t, client, mustNewRequest(t, http.MethodGet, lookupURL))
body := framework.ReadAllAndClose(t, resp)
if resp.StatusCode == http.StatusOK {
var result struct {
Locations []struct {
Url string `json:"url"`
} `json:"locations"`
}
if err := json.Unmarshal(body, &result); err == nil && len(result.Locations) >= want {
return true
}
}
time.Sleep(200 * time.Millisecond)
}
return false
}
+16
View File
@@ -193,6 +193,22 @@ func NewAdminServer(masters string, templateFS http.FileSystem, dataDir string,
} else {
glog.V(0).Infof("Credential store %s does not support filer address function", store.GetName())
}
// Mirror the filer's jwt.filer_signing.key so the admin UI's
// Users/Groups pages can present a valid Bearer token when the
// filer enforces IAM gRPC auth. When the key is empty, both
// sides run unauthenticated and no token is sent.
if signer, ok := store.(interface {
SetAdminSigning(security.SigningKey, int)
}); ok {
viper := util.GetViper()
key := security.SigningKey(viper.GetString("jwt.filer_signing.key"))
expires := viper.GetInt("jwt.filer_signing.expires_after_seconds")
signer.SetAdminSigning(key, expires)
if len(key) > 0 {
glog.V(0).Infof("Credential store configured with admin Bearer token signing")
}
}
}
}
+17
View File
@@ -12,6 +12,8 @@ import (
"github.com/seaweedfs/seaweedfs/weed/pb/filer_pb"
"github.com/seaweedfs/seaweedfs/weed/pb/master_pb"
"github.com/seaweedfs/seaweedfs/weed/pb/volume_server_pb"
"github.com/seaweedfs/seaweedfs/weed/wdclient"
"google.golang.org/grpc"
)
// WithMasterClient executes a function with a master client connection
@@ -38,6 +40,21 @@ func (s *AdminServer) WithVolumeServerClient(address pb.ServerAddress, f func(cl
})
}
// GetMasterClient returns the admin server's wdclient.MasterClient. It is used
// by file browser download paths that stream chunks straight from the volume
// servers via filer.PrepareStreamContent so they keep working when the filer
// has -disableHttp=true.
func (s *AdminServer) GetMasterClient() *wdclient.MasterClient {
return s.masterClient
}
// GetGrpcDialOption returns the dial option used for all admin-originated
// gRPC connections (TLS or insecure). File browser uploads need this when
// they perform the assign + volume HTTP POST + create-entry flow.
func (s *AdminServer) GetGrpcDialOption() grpc.DialOption {
return s.grpcDialOption
}
// GetFilerAddress returns a filer address, discovering from masters if needed
func (s *AdminServer) GetFilerAddress() string {
// Discover filers from masters
+8 -5
View File
@@ -734,8 +734,8 @@ func (s *AdminServer) parseOrBuildClusterContext(raw json.RawMessage) (*plugin_p
if len(contextMessage.MasterGrpcAddresses) == 0 {
contextMessage.MasterGrpcAddresses = append(contextMessage.MasterGrpcAddresses, fallback.MasterGrpcAddresses...)
}
if len(contextMessage.FilerGrpcAddresses) == 0 {
contextMessage.FilerGrpcAddresses = append(contextMessage.FilerGrpcAddresses, fallback.FilerGrpcAddresses...)
if len(contextMessage.FilerAddresses) == 0 {
contextMessage.FilerAddresses = append(contextMessage.FilerAddresses, fallback.FilerAddresses...)
}
if len(contextMessage.VolumeGrpcAddresses) == 0 {
contextMessage.VolumeGrpcAddresses = append(contextMessage.VolumeGrpcAddresses, fallback.VolumeGrpcAddresses...)
@@ -754,7 +754,7 @@ func (s *AdminServer) parseOrBuildClusterContext(raw json.RawMessage) (*plugin_p
func (s *AdminServer) buildDefaultPluginClusterContext() *plugin_pb.ClusterContext {
clusterContext := &plugin_pb.ClusterContext{
MasterGrpcAddresses: make([]string, 0),
FilerGrpcAddresses: make([]string, 0),
FilerAddresses: make([]string, 0),
VolumeGrpcAddresses: make([]string, 0),
S3GrpcAddresses: make([]string, 0),
Metadata: map[string]string{
@@ -767,6 +767,9 @@ func (s *AdminServer) buildDefaultPluginClusterContext() *plugin_pb.ClusterConte
clusterContext.MasterGrpcAddresses = append(clusterContext.MasterGrpcAddresses, masterAddress)
}
// Master returns filers in pb.ServerAddress form (host:httpPort.grpcPort).
// Forward that verbatim; each worker converts to a gRPC or HTTP address as
// it needs (dialing wants gRPC, the admin shell wants the ServerAddress).
filerSeen := map[string]struct{}{}
for _, filer := range s.GetAllFilers() {
filer = strings.TrimSpace(filer)
@@ -777,7 +780,7 @@ func (s *AdminServer) buildDefaultPluginClusterContext() *plugin_pb.ClusterConte
continue
}
filerSeen[filer] = struct{}{}
clusterContext.FilerGrpcAddresses = append(clusterContext.FilerGrpcAddresses, filer)
clusterContext.FilerAddresses = append(clusterContext.FilerAddresses, filer)
}
volumeSeen := map[string]struct{}{}
@@ -825,7 +828,7 @@ func (s *AdminServer) buildDefaultPluginClusterContext() *plugin_pb.ClusterConte
}
sort.Strings(clusterContext.MasterGrpcAddresses)
sort.Strings(clusterContext.FilerGrpcAddresses)
sort.Strings(clusterContext.FilerAddresses)
sort.Strings(clusterContext.VolumeGrpcAddresses)
sort.Strings(clusterContext.S3GrpcAddresses)
+9 -2
View File
@@ -2,6 +2,7 @@ package dash
import (
"context"
"errors"
"fmt"
"io"
"net"
@@ -17,7 +18,9 @@ import (
"github.com/seaweedfs/seaweedfs/weed/security"
"github.com/seaweedfs/seaweedfs/weed/util"
"google.golang.org/grpc"
"google.golang.org/grpc/codes"
"google.golang.org/grpc/peer"
"google.golang.org/grpc/status"
)
const (
@@ -273,9 +276,13 @@ func (s *WorkerGrpcServer) WorkerStream(stream worker_pb.WorkerService_WorkerStr
msg, err := stream.Recv()
if err != nil {
if err == io.EOF {
switch {
case err == io.EOF:
glog.Infof("Worker %s disconnected", workerID)
} else {
case errors.Is(err, context.Canceled), status.Code(err) == codes.Canceled:
// Graceful shutdown on either side cancels the stream.
glog.V(1).Infof("Worker %s stream canceled: %v", workerID, err)
default:
glog.Errorf("Error receiving from worker %s: %v", workerID, err)
}
s.unregisterWorker(conn)
+291
View File
@@ -0,0 +1,291 @@
package handlers
import (
"bytes"
"context"
"fmt"
"io"
"mime"
"net/http"
"path"
"strconv"
"time"
"github.com/seaweedfs/seaweedfs/weed/filer"
"github.com/seaweedfs/seaweedfs/weed/glog"
"github.com/seaweedfs/seaweedfs/weed/operation"
"github.com/seaweedfs/seaweedfs/weed/pb/filer_pb"
"github.com/seaweedfs/seaweedfs/weed/security"
"github.com/seaweedfs/seaweedfs/weed/util"
)
// Admin file browser upload chunk sizing — kept in sync with the values
// s3api uses so files end up split into the same fid-sized pieces the rest of
// the cluster expects.
const (
adminUploadChunkSize = 8 * 1024 * 1024
adminUploadSmallFileLimit = 256 * 1024
)
// File browser handlers backed by the filer gRPC service. They bypass the
// filer's HTTP listener so the UI keeps working when the filer is started
// with -disableHttp=true; chunk bytes still flow through the volume server
// HTTP endpoints (which run on their own ports).
// fetchFileContentGrpc reads file content via the filer gRPC service, looking
// the entry up and then streaming the chunks straight from the volume servers.
// When maxBytes > 0 the stream is truncated to that many bytes — used by the
// "is this text?" sniff so unknown-MIME files don't get fully downloaded.
func (h *FileBrowserHandlers) fetchFileContentGrpc(ctx context.Context, filePath string, maxBytes int) (string, error) {
cleanFilePath, err := h.validateAndCleanFilePath(filePath)
if err != nil {
return "", err
}
entry, err := h.lookupEntry(ctx, cleanFilePath)
if err != nil {
return "", err
}
if entry.IsDirectory {
return "", fmt.Errorf("path is a directory")
}
size := int64(filer.FileSize(entry))
streamSize := size
if maxBytes > 0 && streamSize > int64(maxBytes) {
streamSize = int64(maxBytes)
}
var buf bytes.Buffer
if err := h.streamEntryContent(ctx, entry, streamSize, &buf); err != nil {
return "", err
}
return buf.String(), nil
}
// downloadFileGrpc streams a file via gRPC + volume server HTTP. The
// response writer receives the canonical attachment headers and the raw
// bytes; this replaces the HTTP-to-filer proxy that used to run in
// DownloadFile.
func (h *FileBrowserHandlers) downloadFileGrpc(ctx context.Context, filePath string, w http.ResponseWriter) error {
cleanFilePath, err := h.validateAndCleanFilePath(filePath)
if err != nil {
return err
}
entry, err := h.lookupEntry(ctx, cleanFilePath)
if err != nil {
return err
}
if entry.IsDirectory {
return fmt.Errorf("path is a directory")
}
size := int64(filer.FileSize(entry))
fileName := path.Base(cleanFilePath)
w.Header().Set("Content-Disposition", mime.FormatMediaType("attachment", map[string]string{"filename": fileName}))
contentType := ""
if entry.Attributes != nil {
contentType = entry.Attributes.Mime
}
if contentType == "" {
contentType = "application/octet-stream"
}
w.Header().Set("Content-Type", contentType)
w.Header().Set("Content-Length", strconv.FormatInt(size, 10))
w.WriteHeader(http.StatusOK)
return h.streamEntryContent(ctx, entry, size, w)
}
// uploadFileGrpc streams the upload to volume servers in 8 MiB chunks via the
// shared chunked-upload helper, then registers the assembled entry through the
// filer gRPC service. Bytes never enter the admin process's heap as a whole —
// each chunk is sized to adminUploadChunkSize. Small files (< 256 KiB) are
// stored inline on the entry, matching the S3 server's behaviour.
func (h *FileBrowserHandlers) uploadFileGrpc(ctx context.Context, filePath string, fileName string, mimeType string, reader io.Reader) error {
cleanFilePath, err := h.validateAndCleanFilePath(filePath)
if err != nil {
return err
}
dir := path.Dir(cleanFilePath)
if dir == "." {
dir = "/"
}
entryName := path.Base(cleanFilePath)
if mimeType == "" {
mimeType = "application/octet-stream"
}
assignFunc := func(ctx context.Context, count int, expectedDataSize uint64) (*operation.VolumeAssignRequest, *operation.AssignResult, error) {
var assignResp *filer_pb.AssignVolumeResponse
err := h.adminServer.WithFilerClient(func(client filer_pb.SeaweedFilerClient) error {
resp, assignErr := client.AssignVolume(ctx, &filer_pb.AssignVolumeRequest{
Count: int32(count),
Path: cleanFilePath,
ExpectedDataSize: expectedDataSize,
})
if assignErr != nil {
return assignErr
}
if resp.Error != "" {
return fmt.Errorf("%s", resp.Error)
}
assignResp = resp
return nil
})
if err != nil {
return nil, nil, err
}
if assignResp.Location == nil || assignResp.FileId == "" {
return nil, nil, fmt.Errorf("assign volume returned empty location")
}
return nil, &operation.AssignResult{
Fid: assignResp.FileId,
Url: assignResp.Location.Url,
PublicUrl: assignResp.Location.PublicUrl,
Count: uint64(count),
Auth: security.EncodedJwt(assignResp.Auth),
}, nil
}
chunkResult, err := operation.UploadReaderInChunks(ctx, reader, &operation.ChunkedUploadOption{
ChunkSize: adminUploadChunkSize,
SmallFileLimit: adminUploadSmallFileLimit,
SaveSmallInline: true,
MimeType: mimeType,
AssignFunc: assignFunc,
})
if err != nil {
// Partial chunks come back even on error so we can clean them up rather
// than leaving orphaned data on volume servers.
if chunkResult != nil && len(chunkResult.FileChunks) > 0 {
h.deleteOrphanedChunks(chunkResult.FileChunks)
}
return fmt.Errorf("upload: %w", err)
}
now := time.Now()
entry := &filer_pb.Entry{
Name: entryName,
Attributes: &filer_pb.FuseAttributes{
FileSize: uint64(chunkResult.TotalSize),
Mtime: now.Unix(),
Crtime: now.Unix(),
FileMode: 0644,
Mime: mimeType,
},
}
if len(chunkResult.SmallContent) > 0 {
entry.Content = chunkResult.SmallContent
} else {
entry.Chunks = chunkResult.FileChunks
}
err = h.adminServer.WithFilerClient(func(client filer_pb.SeaweedFilerClient) error {
_, createErr := client.CreateEntry(ctx, &filer_pb.CreateEntryRequest{
Directory: dir,
Entry: entry,
})
return createErr
})
if err != nil {
if len(chunkResult.FileChunks) > 0 {
h.deleteOrphanedChunks(chunkResult.FileChunks)
}
return fmt.Errorf("create entry: %w", err)
}
return nil
}
// deleteOrphanedChunks best-effort removes the chunk fids when an upload
// fails partway through. Errors are logged; we can't surface them past the
// caller's primary failure.
func (h *FileBrowserHandlers) deleteOrphanedChunks(chunks []*filer_pb.FileChunk) {
fileIds := make([]string, 0, len(chunks))
for _, c := range chunks {
if fid := c.GetFileIdString(); fid != "" {
fileIds = append(fileIds, fid)
}
}
if len(fileIds) == 0 {
return
}
master := h.adminServer.GetMasterClient()
results := operation.DeleteFileIds(master.GetMaster, false, h.adminServer.GetGrpcDialOption(), fileIds)
for _, r := range results {
if r.Error != "" {
glog.Warningf("admin file browser: orphan chunk %s cleanup: %s", r.FileId, r.Error)
}
}
}
func (h *FileBrowserHandlers) lookupEntry(ctx context.Context, cleanFilePath string) (*filer_pb.Entry, error) {
dir := path.Dir(cleanFilePath)
if dir == "." {
dir = "/"
}
name := path.Base(cleanFilePath)
var entry *filer_pb.Entry
err := h.adminServer.WithFilerClient(func(client filer_pb.SeaweedFilerClient) error {
resp, lookupErr := client.LookupDirectoryEntry(ctx, &filer_pb.LookupDirectoryEntryRequest{
Directory: dir,
Name: name,
})
if lookupErr != nil {
return lookupErr
}
if resp.Entry == nil {
return fmt.Errorf("not found")
}
entry = resp.Entry
return nil
})
return entry, err
}
func (h *FileBrowserHandlers) streamEntryContent(ctx context.Context, entry *filer_pb.Entry, size int64, w io.Writer) error {
if size == 0 {
// Inline content (small files stored directly on the entry) skip the
// chunk pipeline entirely.
if len(entry.Content) > 0 {
_, err := w.Write(entry.Content)
return err
}
return nil
}
if len(entry.Content) > 0 && len(entry.GetChunks()) == 0 {
_, err := w.Write(entry.Content)
return err
}
streamFn, err := filer.PrepareStreamContentWithThrottler(
ctx,
h.adminServer.GetMasterClient(),
volumeServerReadJwt,
entry.GetChunks(),
0,
size,
0,
)
if err != nil {
return fmt.Errorf("prepare stream: %w", err)
}
return streamFn(w)
}
// volumeServerReadJwt mints a per-fileId Bearer token for reads against a
// volume server when jwt.signing.read.key is configured. The volume servers
// are unaware of jwt.filer_signing.read.key — that one only gates the filer
// HTTP surface, which this code path doesn't touch.
func volumeServerReadJwt(fileId string) string {
v := util.GetViper()
signingKey := security.SigningKey(v.GetString("jwt.signing.read.key"))
if len(signingKey) == 0 {
return ""
}
expiresAfterSec := v.GetInt("jwt.signing.read.expires_after_seconds")
return string(security.GenJwtForVolumeServer(signingKey, expiresAfterSec, fileId))
}
+51 -370
View File
@@ -1,15 +1,10 @@
package handlers
import (
"bytes"
"context"
"fmt"
"io"
"mime"
"mime/multipart"
"net"
"net/http"
"net/url"
"os"
"path"
"path/filepath"
@@ -21,9 +16,7 @@ import (
"github.com/seaweedfs/seaweedfs/weed/admin/view/app"
"github.com/seaweedfs/seaweedfs/weed/admin/view/layout"
"github.com/seaweedfs/seaweedfs/weed/glog"
"github.com/seaweedfs/seaweedfs/weed/pb"
"github.com/seaweedfs/seaweedfs/weed/pb/filer_pb"
"github.com/seaweedfs/seaweedfs/weed/security"
"github.com/seaweedfs/seaweedfs/weed/util"
"github.com/seaweedfs/seaweedfs/weed/util/http/client"
)
@@ -49,15 +42,6 @@ func NewFileBrowserHandlers(adminServer *dash.AdminServer) *FileBrowserHandlers
}
}
// newClientWithTimeout creates a temporary http.Client with the specified timeout,
// reusing the TLS transport from the shared httpClient.
func (h *FileBrowserHandlers) newClientWithTimeout(timeout time.Duration) http.Client {
return http.Client{
Transport: h.httpClient.Client.Transport,
Timeout: timeout,
}
}
// ShowFileBrowser renders the file browser page
func (h *FileBrowserHandlers) ShowFileBrowser(w http.ResponseWriter, r *http.Request) {
// Get path from query parameter, default to root
@@ -309,7 +293,7 @@ func (h *FileBrowserHandlers) UploadFile(w http.ResponseWriter, r *http.Request)
}
// Upload file to filer
err = h.uploadFileToFiler(fullPath, fileHeader)
err = h.uploadFileToFiler(r.Context(), fullPath, fileHeader)
if err != nil {
failedUploads = append(failedUploads, fmt.Sprintf("%s: %v", fileName, err))
@@ -346,139 +330,21 @@ func (h *FileBrowserHandlers) UploadFile(w http.ResponseWriter, r *http.Request)
}
}
// uploadFileToFiler uploads a file directly to the filer using multipart form data
func (h *FileBrowserHandlers) uploadFileToFiler(filePath string, fileHeader *multipart.FileHeader) error {
// Get filer address from admin server
filerAddress := h.adminServer.GetFilerAddress()
if filerAddress == "" {
return fmt.Errorf("filer address not configured")
}
// Validate and sanitize the filer address
if err := h.validateFilerAddress(filerAddress); err != nil {
return fmt.Errorf("invalid filer address: %w", err)
}
filerHttpAddress := pb.ServerAddress(filerAddress).ToHttpAddress()
// Validate and sanitize the file path
cleanFilePath, err := h.validateAndCleanFilePath(filePath)
if err != nil {
return fmt.Errorf("invalid file path: %w", err)
}
// Open the file
// uploadFileToFiler uploads a file to the cluster via filer gRPC + volume
// HTTP. This works whether or not the filer is running with -disableHttp=true,
// since the bytes never traverse the filer's HTTP listener. The multipart
// file is streamed through the chunked uploader, so the admin process never
// buffers the entire payload in memory. The caller passes the request
// context so a client disconnect cancels the in-flight chunk uploads instead
// of letting them run to completion against the volume servers.
func (h *FileBrowserHandlers) uploadFileToFiler(ctx context.Context, filePath string, fileHeader *multipart.FileHeader) error {
file, err := fileHeader.Open()
if err != nil {
return fmt.Errorf("failed to open file: %w", err)
}
defer file.Close()
// Create multipart form data
var body bytes.Buffer
writer := multipart.NewWriter(&body)
// Create form file field with normalized base filename
// Use path.Base (not filepath.Base) since cleanFilePath uses URL path semantics
baseFileName := path.Base(cleanFilePath)
part, err := writer.CreateFormFile("file", baseFileName)
if err != nil {
return fmt.Errorf("failed to create form file: %w", err)
}
// Copy file content to form
_, err = io.Copy(part, file)
if err != nil {
return fmt.Errorf("failed to copy file content: %w", err)
}
// Close the writer to finalize the form
err = writer.Close()
if err != nil {
return fmt.Errorf("failed to close multipart writer: %w", err)
}
// Create the upload URL - the httpClient will normalize to the correct scheme (http/https)
// based on the https.client configuration in security.toml
uploadURL := filerFileURL(filerHttpAddress, cleanFilePath)
// Normalize the URL scheme based on TLS configuration
uploadURL, err = h.httpClient.NormalizeHttpScheme(uploadURL)
if err != nil {
return fmt.Errorf("failed to normalize URL scheme: %w", err)
}
// Create HTTP request
req, err := http.NewRequest("POST", uploadURL, &body)
if err != nil {
return fmt.Errorf("failed to create request: %w", err)
}
// Set content type with boundary
req.Header.Set("Content-Type", writer.FormDataContentType())
// Add JWT Token to Authorization Header
h.setupFilerJwtAuth(req, "jwt.filer_signing.key", "jwt.filer_signing.expires_after_seconds", "filer upload")
// Send request using TLS-aware HTTP client with 60s timeout for large file uploads
// lgtm[go/ssrf]
// Safe: filerAddress validated by validateFilerAddress() to match configured filer
// Safe: cleanFilePath validated and cleaned by validateAndCleanFilePath() to prevent path traversal
client := h.newClientWithTimeout(60 * time.Second)
resp, err := client.Do(req)
if err != nil {
return fmt.Errorf("failed to upload file: %w", err)
}
defer resp.Body.Close()
// Check response
if resp.StatusCode != http.StatusOK && resp.StatusCode != http.StatusCreated {
responseBody, _ := io.ReadAll(resp.Body)
return fmt.Errorf("upload failed with status %d: %s", resp.StatusCode, string(responseBody))
}
return nil
}
// validateFilerAddress validates that the filer address is safe to use
func (h *FileBrowserHandlers) validateFilerAddress(address string) error {
if address == "" {
return fmt.Errorf("filer address cannot be empty")
}
// CRITICAL: Only allow the configured filer address to prevent SSRF
configuredFiler := h.adminServer.GetFilerAddress()
normalizedAddress := pb.ServerAddress(address).ToHttpAddress()
normalizedConfigured := pb.ServerAddress(configuredFiler).ToHttpAddress()
if normalizedAddress != normalizedConfigured {
return fmt.Errorf("address does not match configured filer: got %s, expected %s", address, configuredFiler)
}
// Parse the normalized HTTP address to validate it's a proper host:port format.
host, port, err := net.SplitHostPort(normalizedAddress)
if err != nil {
return fmt.Errorf("invalid address format: %w", err)
}
// Validate host is not empty
if host == "" {
return fmt.Errorf("host cannot be empty")
}
// Validate port is numeric and in valid range
if port == "" {
return fmt.Errorf("port cannot be empty")
}
portNum, err := strconv.Atoi(port)
if err != nil {
return fmt.Errorf("invalid port number: %w", err)
}
if portNum < 1 || portNum > 65535 {
return fmt.Errorf("port number must be between 1 and 65535")
}
return nil
return h.uploadFileGrpc(ctx, filePath, fileHeader.Filename, fileHeader.Header.Get("Content-Type"), file)
}
// validateAndCleanFilePath validates and cleans the file path to prevent path traversal
@@ -507,161 +373,56 @@ func (h *FileBrowserHandlers) validateAndCleanFilePath(filePath string) (string,
return cleanPath, nil
}
// filerFileURL joins the filer HTTP address with a validated file path, URL-escaping
// the path so that control characters and other bytes that are legal in S3 object keys
// cannot inject into the HTTP request target.
func filerFileURL(filerHttpAddress, cleanFilePath string) string {
return filerHttpAddress + (&url.URL{Path: cleanFilePath}).EscapedPath()
}
// fetchFileContent fetches file content from the filer and returns the content or an error.
// fetchFileContent fetches file content via the filer gRPC service. It is
// used for the "view as text" path, so the maxBytes cap matches the 1 MB
// limit the caller already applies before invoking us.
func (h *FileBrowserHandlers) fetchFileContent(filePath string, timeout time.Duration) (string, error) {
filerAddress := h.adminServer.GetFilerAddress()
if filerAddress == "" {
return "", fmt.Errorf("filer address not configured")
}
if err := h.validateFilerAddress(filerAddress); err != nil {
return "", fmt.Errorf("invalid filer address configuration: %w", err)
}
filerHttpAddress := pb.ServerAddress(filerAddress).ToHttpAddress()
cleanFilePath, err := h.validateAndCleanFilePath(filePath)
if err != nil {
return "", err
}
// Create the file URL with proper scheme based on TLS configuration
fileURL := filerFileURL(filerHttpAddress, cleanFilePath)
fileURL, err = h.httpClient.NormalizeHttpScheme(fileURL)
if err != nil {
return "", fmt.Errorf("failed to construct file URL: %w", err)
}
// lgtm[go/ssrf]
// Safe: filerAddress validated by validateFilerAddress() to match configured filer
// Safe: cleanFilePath validated and cleaned by validateAndCleanFilePath() to prevent path traversal
client := h.newClientWithTimeout(timeout)
req, err := http.NewRequest("GET", fileURL, nil)
if err != nil {
return "", fmt.Errorf("failed to create request: %w", err)
}
h.addFilerJwtAuthHeader(req)
resp, err := client.Do(req)
if err != nil {
return "", fmt.Errorf("failed to fetch file from filer: %w", err)
}
defer resp.Body.Close()
if resp.StatusCode != http.StatusOK {
body, err := io.ReadAll(resp.Body)
if err != nil {
return "", fmt.Errorf("filer returned status %d but failed to read response body: %w", resp.StatusCode, err)
}
return "", fmt.Errorf("filer returned status %d: %s", resp.StatusCode, string(body))
}
contentBytes, err := io.ReadAll(resp.Body)
if err != nil {
return "", fmt.Errorf("failed to read file content: %w", err)
}
return string(contentBytes), nil
ctx, cancel := context.WithTimeout(context.Background(), timeout)
defer cancel()
return h.fetchFileContentGrpc(ctx, filePath, 0)
}
// DownloadFile handles file download requests by proxying through the Admin UI server
// This ensures mTLS works correctly since the Admin UI server has the client certificates
// DownloadFile streams a file straight from the volume servers via the filer
// gRPC service, so the admin file browser keeps working even when the filer
// is started with -disableHttp=true.
func (h *FileBrowserHandlers) DownloadFile(w http.ResponseWriter, r *http.Request) {
filePath := r.URL.Query().Get("path")
if filePath == "" {
writeJSONError(w, http.StatusBadRequest, "File path is required")
return
}
// Get filer address
filerAddress := h.adminServer.GetFilerAddress()
if filerAddress == "" {
writeJSONError(w, http.StatusInternalServerError, "Filer address not configured")
return
}
// Validate filer address to prevent SSRF
if err := h.validateFilerAddress(filerAddress); err != nil {
writeJSONError(w, http.StatusInternalServerError, "Invalid filer address configuration")
return
}
filerHttpAddress := pb.ServerAddress(filerAddress).ToHttpAddress()
// Validate and sanitize the file path
cleanFilePath, err := h.validateAndCleanFilePath(filePath)
if err != nil {
writeJSONError(w, http.StatusBadRequest, "Invalid file path: "+err.Error())
return
}
// Create the download URL with proper scheme based on TLS configuration
downloadURL := filerFileURL(filerHttpAddress, cleanFilePath)
downloadURL, err = h.httpClient.NormalizeHttpScheme(downloadURL)
if err != nil {
writeJSONError(w, http.StatusInternalServerError, "Failed to construct download URL: "+err.Error())
return
}
// Proxy the download through the Admin UI server to support mTLS
// lgtm[go/ssrf]
// Safe: filerAddress validated by validateFilerAddress() to match configured filer
// Safe: cleanFilePath validated and cleaned by validateAndCleanFilePath() to prevent path traversal
// Use request context so download is cancelled when client disconnects
req, err := http.NewRequestWithContext(r.Context(), "GET", downloadURL, nil)
if err != nil {
writeJSONError(w, http.StatusInternalServerError, "Failed to create request: "+err.Error())
return
}
client := h.newClientWithTimeout(5 * time.Minute) // Longer timeout for large file downloads
h.addFilerJwtAuthHeader(req)
resp, err := client.Do(req)
if err != nil {
writeJSONError(w, http.StatusBadGateway, "Failed to fetch file from filer: "+err.Error())
return
}
defer resp.Body.Close()
if resp.StatusCode != http.StatusOK {
body, err := io.ReadAll(resp.Body)
if err != nil {
writeJSONError(w, resp.StatusCode, fmt.Sprintf("Filer returned status %d but failed to read response body: %v", resp.StatusCode, err))
tracker := &responseWriteTracker{ResponseWriter: w}
if err := h.downloadFileGrpc(r.Context(), filePath, tracker); err != nil {
// Once bytes have been written we can't switch to a JSON error body
// without corrupting the partial response — log and stop. Before any
// write the response is still uncommitted, so a 502 with details is
// safe.
if tracker.committed {
glog.Errorf("Error streaming file download: %v", err)
return
}
writeJSONError(w, resp.StatusCode, fmt.Sprintf("Filer returned status %d: %s", resp.StatusCode, string(body)))
return
writeJSONError(w, http.StatusBadGateway, "Failed to fetch file: "+err.Error())
}
}
// Set headers for file download
fileName := filepath.Base(cleanFilePath)
// Use mime.FormatMediaType for RFC 6266 compliant Content-Disposition,
// properly handling non-ASCII characters and special characters
w.Header().Set("Content-Disposition", mime.FormatMediaType("attachment", map[string]string{"filename": fileName}))
// responseWriteTracker wraps http.ResponseWriter to record whether the
// response has been committed (status line + headers sent). DownloadFile
// uses this instead of probing Header() so future header-setting code
// reorganization can't silently break the "did we already send bytes?"
// detection.
type responseWriteTracker struct {
http.ResponseWriter
committed bool
}
// Use content type from filer response, or default to octet-stream
contentType := resp.Header.Get("Content-Type")
if contentType == "" {
contentType = "application/octet-stream"
}
w.Header().Set("Content-Type", contentType)
func (t *responseWriteTracker) WriteHeader(code int) {
t.committed = true
t.ResponseWriter.WriteHeader(code)
}
// Set content length if available
if resp.ContentLength > 0 {
w.Header().Set("Content-Length", fmt.Sprintf("%d", resp.ContentLength))
}
// Stream the response body to the client
w.WriteHeader(http.StatusOK)
_, err = io.Copy(w, resp.Body)
if err != nil {
glog.Errorf("Error streaming file download: %v", err)
}
func (t *responseWriteTracker) Write(p []byte) (int, error) {
t.committed = true
return t.ResponseWriter.Write(p)
}
// ViewFile handles file viewing requests (for text files, images, etc.)
@@ -883,64 +644,16 @@ func (h *FileBrowserHandlers) formatBytes(bytes int64) string {
// Helper function to check if a file is likely a text file by checking content
func (h *FileBrowserHandlers) isLikelyTextFile(filePath string, maxCheckSize int64) bool {
filerAddress := h.adminServer.GetFilerAddress()
if filerAddress == "" {
return false
}
// Validate filer address to prevent SSRF
if err := h.validateFilerAddress(filerAddress); err != nil {
glog.Errorf("Invalid filer address: %v", err)
return false
}
filerHttpAddress := pb.ServerAddress(filerAddress).ToHttpAddress()
cleanFilePath, err := h.validateAndCleanFilePath(filePath)
ctx, cancel := context.WithTimeout(context.Background(), 10*time.Second)
defer cancel()
content, err := h.fetchFileContentGrpc(ctx, filePath, int(maxCheckSize))
if err != nil {
return false
}
// Create the file URL with proper scheme based on TLS configuration
fileURL := filerFileURL(filerHttpAddress, cleanFilePath)
fileURL, err = h.httpClient.NormalizeHttpScheme(fileURL)
if err != nil {
glog.Errorf("Failed to normalize URL scheme: %v", err)
return false
if len(content) == 0 {
return true
}
// lgtm[go/ssrf]
// Safe: filerAddress validated by validateFilerAddress() to match configured filer
// Safe: cleanFilePath validated and cleaned by validateAndCleanFilePath() to prevent path traversal
client := h.newClientWithTimeout(10 * time.Second)
req, err := http.NewRequest("GET", fileURL, nil)
if err != nil {
glog.Errorf("Failed to create request: %v", err)
return false
}
h.addFilerJwtAuthHeader(req)
resp, err := client.Do(req)
if err != nil {
return false
}
defer resp.Body.Close()
if resp.StatusCode != http.StatusOK {
return false
}
// Read first few bytes to check if it's text
buffer := make([]byte, min(maxCheckSize, 512))
n, err := resp.Body.Read(buffer)
if err != nil && err != io.EOF {
return false
}
if n == 0 {
return true // Empty file can be considered text
}
// Check if content is printable text
return h.isPrintableText(buffer[:n])
return h.isPrintableText([]byte(content))
}
// Helper function to check if content is printable text
@@ -973,35 +686,3 @@ func min(a, b int64) int64 {
return b
}
// setupFilerJwtAuth generates a JWT token and adds it to the request Authorization header if configured.
func (h *FileBrowserHandlers) setupFilerJwtAuth(req *http.Request, keyPath, expiresPath, operation string) {
// Load security configuration
v := util.GetViper()
// Read Filer JWT token from security.toml
signingKey := security.SigningKey(v.GetString(keyPath))
expiresAfterSec := v.GetInt(expiresPath)
// Generate JWT token to authenticate with Filer
var jwtToken security.EncodedJwt
if len(signingKey) > 0 {
jwtToken = security.GenJwtForFilerServer(signingKey, expiresAfterSec)
glog.V(4).Infof("Generated JWT token for %s (expires in %d sec)", operation, expiresAfterSec)
} else {
if v.GetString("jwt.signing.key") != "" {
glog.Warningf("JWT %s key not configured, but general JWT security is enabled. %s without authentication.", keyPath, operation)
} else {
glog.V(1).Infof("No JWT signing key configured, %s without authentication", operation)
}
}
// Add JWT Token to Authorization Header
if jwtToken != "" {
req.Header.Set("Authorization", fmt.Sprintf("Bearer %s", string(jwtToken)))
glog.V(4).Infof("Added JWT authorization header for %s", operation)
}
}
func (h *FileBrowserHandlers) addFilerJwtAuthHeader(req *http.Request) {
h.setupFilerJwtAuth(req, "jwt.filer_signing.read.key", "jwt.filer_signing.read.expires_after_seconds", "filer request")
}
@@ -42,21 +42,3 @@ func TestValidateAndCleanFilePath_RejectsEmpty(t *testing.T) {
}
}
func TestFilerFileURL_EscapesControlChars(t *testing.T) {
cases := []struct {
addr string
path string
want string
}{
{"http://127.0.0.1:8888", "/buckets/profilebuilder/3testGB.zip\n ", "http://127.0.0.1:8888/buckets/profilebuilder/3testGB.zip%0A%20"},
{"http://127.0.0.1:8888", "/buckets/profilebuilder/file\rname", "http://127.0.0.1:8888/buckets/profilebuilder/file%0Dname"},
{"http://127.0.0.1:8888", "/buckets/profilebuilder/file\x00name", "http://127.0.0.1:8888/buckets/profilebuilder/file%00name"},
// Plain path round-trips unchanged.
{"http://h:1", "/a/b.txt", "http://h:1/a/b.txt"},
}
for _, tc := range cases {
if got := filerFileURL(tc.addr, tc.path); got != tc.want {
t.Errorf("filerFileURL(%q, %q) = %q, want %q", tc.addr, tc.path, got, tc.want)
}
}
}
@@ -26,6 +26,10 @@ type MaintenanceIntegration struct {
// Active topology for task detection and target selection
activeTopology *topology.ActiveTopology
// Master's default replication, refreshed by the scanner each cycle and
// passed to detectors as the replica-placement fallback (matches the shell).
defaultReplicaPlacement string
// Type conversion maps
taskTypeMap map[types.TaskType]MaintenanceTaskType
revTaskTypeMap map[MaintenanceTaskType]types.TaskType
@@ -219,9 +223,10 @@ func (s *MaintenanceIntegration) ScanWithTaskDetectors(volumeMetrics []*types.Vo
// Create cluster info
clusterInfo := &types.ClusterInfo{
TotalVolumes: len(filteredMetrics),
LastUpdated: time.Now(),
ActiveTopology: s.activeTopology, // Provide ActiveTopology for destination planning
TotalVolumes: len(filteredMetrics),
LastUpdated: time.Now(),
ActiveTopology: s.activeTopology, // Provide ActiveTopology for destination planning
DefaultReplicaPlacement: s.defaultReplicaPlacement,
}
// Run detection for each registered task type
@@ -271,6 +276,12 @@ func (s *MaintenanceIntegration) ScanWithTaskDetectors(volumeMetrics []*types.Vo
return allResults, nil
}
// SetDefaultReplicaPlacement records the master's default replication so detectors
// can use it as the replica-placement fallback (matching the shell).
func (s *MaintenanceIntegration) SetDefaultReplicaPlacement(replicaPlacement string) {
s.defaultReplicaPlacement = replicaPlacement
}
// UpdateTopologyInfo updates the volume shard tracker with topology information for empty servers
func (s *MaintenanceIntegration) UpdateTopologyInfo(topologyInfo *master_pb.TopologyInfo) error {
// Log topology details before update for diagnostics
@@ -51,6 +51,10 @@ func (ms *MaintenanceScanner) ScanForMaintenanceTasks() ([]*TaskDetectionResult,
}
}
// Refresh the master's default replication so detectors can use it as the
// replica-placement fallback (matches the shell ec.balance default).
ms.integration.SetDefaultReplicaPlacement(ms.getDefaultReplicaPlacement())
// Use task detection system with complete cluster information
results, err := ms.integration.ScanWithTaskDetectors(taskMetrics)
if err != nil {
@@ -67,6 +71,26 @@ func (ms *MaintenanceScanner) ScanForMaintenanceTasks() ([]*TaskDetectionResult,
return []*TaskDetectionResult{}, nil
}
// getDefaultReplicaPlacement reads the master's configured default replication,
// used by detectors as the replica-placement fallback. Returns "" on error so
// detectors fall back to even spread rather than failing the scan.
func (ms *MaintenanceScanner) getDefaultReplicaPlacement() string {
var replicaPlacement string
err := ms.adminClient.WithMasterClient(func(client master_pb.SeaweedClient) error {
resp, err := client.GetMasterConfiguration(context.Background(), &master_pb.GetMasterConfigurationRequest{})
if err != nil {
return err
}
replicaPlacement = resp.DefaultReplication
return nil
})
if err != nil {
glog.V(1).Infof("could not fetch master default replication: %v", err)
return ""
}
return replicaPlacement
}
// getVolumeHealthMetrics collects health information for all volumes.
// Returns metrics in task-system format directly (no intermediate copy) and
// the topology info for updating the active topology.
+1 -1
View File
@@ -113,7 +113,7 @@ func cloneClusterContext(in *plugin_pb.ClusterContext) *plugin_pb.ClusterContext
}
out := &plugin_pb.ClusterContext{
MasterGrpcAddresses: in.MasterGrpcAddresses,
FilerGrpcAddresses: in.FilerGrpcAddresses,
FilerAddresses: in.FilerAddresses,
VolumeGrpcAddresses: in.VolumeGrpcAddresses,
S3GrpcAddresses: in.S3GrpcAddresses,
}
+5
View File
@@ -115,6 +115,7 @@ func buildErasureCodingExecutionPlan(params *worker_pb.TaskParams) map[string]in
source.DataCenter,
source.Rack,
source.VolumeId,
source.DiskId,
source.ShardIds,
dataShards,
))
@@ -132,6 +133,7 @@ func buildErasureCodingExecutionPlan(params *worker_pb.TaskParams) map[string]in
target.DataCenter,
target.Rack,
target.VolumeId,
target.DiskId,
target.ShardIds,
dataShards,
))
@@ -147,6 +149,7 @@ func buildErasureCodingExecutionPlan(params *worker_pb.TaskParams) map[string]in
"target_data_center": strings.TrimSpace(target.DataCenter),
"target_rack": strings.TrimSpace(target.Rack),
"target_volume_id": int(target.VolumeId),
"target_disk_id": int(target.DiskId),
})
}
}
@@ -182,6 +185,7 @@ func buildExecutionEndpoint(
dataCenter string,
rack string,
volumeID uint32,
diskID uint32,
shardIDs []uint32,
dataShardCount int,
) map[string]interface{} {
@@ -201,6 +205,7 @@ func buildExecutionEndpoint(
"data_center": strings.TrimSpace(dataCenter),
"rack": strings.TrimSpace(rack),
"volume_id": int(volumeID),
"disk_id": int(diskID),
"shard_ids": allShards,
"data_shard_ids": dataShards,
"parity_shard_ids": parityShards,
+31
View File
@@ -423,6 +423,7 @@ func TestTrackExecutionStartStoresErasureCodingExecutionPlan(t *testing.T) {
DataCenter: "dc1",
Rack: "rack1",
VolumeId: 29,
DiskId: 5,
},
},
Targets: []*worker_pb.TaskTarget{
@@ -431,6 +432,7 @@ func TestTrackExecutionStartStoresErasureCodingExecutionPlan(t *testing.T) {
DataCenter: "dc1",
Rack: "rack2",
VolumeId: 29,
DiskId: 2,
ShardIds: []uint32{0, 10},
},
{
@@ -438,6 +440,7 @@ func TestTrackExecutionStartStoresErasureCodingExecutionPlan(t *testing.T) {
DataCenter: "dc2",
Rack: "rack3",
VolumeId: 29,
DiskId: 3,
ShardIds: []uint32{1, 11},
},
},
@@ -486,10 +489,28 @@ func TestTrackExecutionStartStoresErasureCodingExecutionPlan(t *testing.T) {
if plan["volume_id"] != float64(29) {
t.Fatalf("unexpected execution plan volume id: %+v", plan["volume_id"])
}
sourcesRaw, ok := plan["sources"].([]interface{})
if !ok || len(sourcesRaw) != 1 {
t.Fatalf("unexpected sources in execution plan: %+v", plan["sources"])
}
firstSource, ok := sourcesRaw[0].(map[string]interface{})
if !ok {
t.Fatalf("unexpected source payload: %+v", sourcesRaw[0])
}
if firstSource["disk_id"] != float64(5) {
t.Fatalf("unexpected source disk_id: %+v", firstSource["disk_id"])
}
targets, ok := plan["targets"].([]interface{})
if !ok || len(targets) != 2 {
t.Fatalf("unexpected targets in execution plan: %+v", plan["targets"])
}
firstTarget, ok := targets[0].(map[string]interface{})
if !ok {
t.Fatalf("unexpected target payload: %+v", targets[0])
}
if firstTarget["disk_id"] != float64(2) {
t.Fatalf("unexpected target disk_id: %+v", firstTarget["disk_id"])
}
assignments, ok := plan["shard_assignments"].([]interface{})
if !ok || len(assignments) != 4 {
t.Fatalf("unexpected shard assignments in execution plan: %+v", plan["shard_assignments"])
@@ -501,6 +522,16 @@ func TestTrackExecutionStartStoresErasureCodingExecutionPlan(t *testing.T) {
if firstAssignment["shard_id"] != float64(0) || firstAssignment["kind"] != "data" {
t.Fatalf("unexpected first assignment: %+v", firstAssignment)
}
if firstAssignment["target_disk_id"] != float64(2) {
t.Fatalf("unexpected first assignment target_disk_id: %+v", firstAssignment["target_disk_id"])
}
secondAssignment, ok := assignments[1].(map[string]interface{})
if !ok {
t.Fatalf("unexpected second assignment payload: %+v", assignments[1])
}
if secondAssignment["shard_id"] != float64(1) || secondAssignment["target_disk_id"] != float64(3) {
t.Fatalf("unexpected second assignment: %+v", secondAssignment)
}
}
func TestBuildJobDetailIncludesActivitiesAndRunRecord(t *testing.T) {
@@ -0,0 +1,54 @@
package topology
import (
"testing"
"github.com/seaweedfs/seaweedfs/weed/pb/master_pb"
)
// TestCountTopologyResources_multiDiskPerNode covers the case where the master
// keys DiskInfos by disk type, so several same-type physical disks on a node
// collapse into a single DiskInfo entry. Counting len(DiskInfos) under-reports
// the physical disk count and disagrees with the per-disk activeDisk map that
// the rest of the admin topology builds via SplitByPhysicalDisk.
func TestCountTopologyResources_multiDiskPerNode(t *testing.T) {
makeNode := func(id string) *master_pb.DataNodeInfo {
var ecShardInfos []*master_pb.VolumeEcShardInformationMessage
for diskId := uint32(0); diskId < 6; diskId++ {
ecShardInfos = append(ecShardInfos, &master_pb.VolumeEcShardInformationMessage{
Id: diskId + 1,
DiskId: diskId,
EcIndexBits: 1,
})
}
return &master_pb.DataNodeInfo{
Id: id,
DiskInfos: map[string]*master_pb.DiskInfo{
"": {Type: "", MaxVolumeCount: 60, EcShardInfos: ecShardInfos},
},
}
}
topo := &master_pb.TopologyInfo{
Id: "multi_disk_topo",
DataCenterInfos: []*master_pb.DataCenterInfo{{
Id: "dc1",
RackInfos: []*master_pb.RackInfo{{
Id: "rack1",
DataNodeInfos: []*master_pb.DataNodeInfo{
makeNode("node1"), makeNode("node2"), makeNode("node3"),
},
}},
}},
}
dcCount, nodeCount, diskCount := CountTopologyResources(topo)
if dcCount != 1 {
t.Errorf("dcCount = %d, want 1", dcCount)
}
if nodeCount != 3 {
t.Errorf("nodeCount = %d, want 3", nodeCount)
}
if diskCount != 18 {
t.Errorf("diskCount = %d, want 18 (6 physical disks x 3 nodes)", diskCount)
}
}
+8 -5
View File
@@ -113,10 +113,13 @@ type MultiDestinationPlan struct {
SuccessfulDCs int `json:"successful_dcs"`
}
// VolumeReplica represents a replica location with server and disk information
// VolumeReplica represents a replica location with server and disk information.
// ShardIds is populated only by GetECShardLocations — it lists the EC shards
// the disk holds for the volume.
type VolumeReplica struct {
ServerID string `json:"server_id"`
DiskID uint32 `json:"disk_id"`
DataCenter string `json:"data_center"`
Rack string `json:"rack"`
ServerID string `json:"server_id"`
DiskID uint32 `json:"disk_id"`
DataCenter string `json:"data_center"`
Rack string `json:"rack"`
ShardIds []uint32 `json:"shard_ids,omitempty"`
}
+1
View File
@@ -341,6 +341,7 @@ type TaskSourceSpec struct {
DataCenter string // Data center of the source server
Rack string // Rack of the source server
CleanupType SourceCleanupType // For EC: volume replica vs existing shards
ShardIds []uint32 // For CleanupECShards: shard ids on the source disk to clear before re-distributing
StorageImpact *StorageSlotChange // Optional: manual override
EstimatedSize *int64 // Optional: manual override
}
+58 -82
View File
@@ -6,76 +6,9 @@ import (
"github.com/seaweedfs/seaweedfs/weed/glog"
"github.com/seaweedfs/seaweedfs/weed/pb/master_pb"
"github.com/seaweedfs/seaweedfs/weed/storage/erasure_coding"
)
// splitDiskInfoByPhysicalDisk returns one master_pb.DiskInfo per physical
// disk_id observed in VolumeInfos / EcShardInfos. Multiple same-type physical
// disks collapse to one DiskInfo at the master; per-volume/per-shard records
// keep the original disk_id and are the authoritative signal here. Capacity
// is split evenly — the wire format doesn't carry per-disk capacity yet.
func splitDiskInfoByPhysicalDisk(diskInfo *master_pb.DiskInfo) []*master_pb.DiskInfo {
if diskInfo == nil {
return nil
}
// Records with DiskId=0 and a non-zero outer DiskId belong to the outer
// disk — handles older payloads / fixtures that omit the per-record id.
normalize := func(id uint32) uint32 {
if id == 0 && diskInfo.DiskId != 0 {
return diskInfo.DiskId
}
return id
}
diskIDs := make(map[uint32]struct{})
for _, vi := range diskInfo.VolumeInfos {
diskIDs[normalize(vi.DiskId)] = struct{}{}
}
for _, eci := range diskInfo.EcShardInfos {
diskIDs[normalize(eci.DiskId)] = struct{}{}
}
if len(diskIDs) == 0 {
diskIDs[diskInfo.DiskId] = struct{}{}
}
if len(diskIDs) == 1 {
for diskID := range diskIDs {
if diskID == diskInfo.DiskId {
return []*master_pb.DiskInfo{diskInfo}
}
}
}
perDiskVolumes := make(map[uint32][]*master_pb.VolumeInformationMessage)
for _, vi := range diskInfo.VolumeInfos {
perDiskVolumes[normalize(vi.DiskId)] = append(perDiskVolumes[normalize(vi.DiskId)], vi)
}
perDiskShards := make(map[uint32][]*master_pb.VolumeEcShardInformationMessage)
for _, eci := range diskInfo.EcShardInfos {
perDiskShards[normalize(eci.DiskId)] = append(perDiskShards[normalize(eci.DiskId)], eci)
}
count := int64(len(diskIDs))
share := func(total int64) int64 { return total / count }
result := make([]*master_pb.DiskInfo, 0, len(diskIDs))
for diskID := range diskIDs {
result = append(result, &master_pb.DiskInfo{
Type: diskInfo.Type,
MaxVolumeCount: share(diskInfo.MaxVolumeCount),
VolumeCount: int64(len(perDiskVolumes[diskID])),
FreeVolumeCount: share(diskInfo.FreeVolumeCount),
ActiveVolumeCount: share(diskInfo.ActiveVolumeCount),
RemoteVolumeCount: share(diskInfo.RemoteVolumeCount),
VolumeInfos: perDiskVolumes[diskID],
EcShardInfos: perDiskShards[diskID],
DiskId: diskID,
Tags: append([]string(nil), diskInfo.Tags...),
})
}
return result
}
// CountTopologyResources counts datacenters, nodes, and disks in topology info
func CountTopologyResources(topologyInfo *master_pb.TopologyInfo) (dcCount, nodeCount, diskCount int) {
if topologyInfo == nil {
@@ -86,7 +19,12 @@ func CountTopologyResources(topologyInfo *master_pb.TopologyInfo) (dcCount, node
for _, rack := range dc.RackInfos {
nodeCount += len(rack.DataNodeInfos)
for _, node := range rack.DataNodeInfos {
diskCount += len(node.DiskInfos)
// DiskInfos is keyed by disk type, so same-type physical disks
// collapse into one entry. Count physical disks so the number
// matches the per-disk activeDisk map.
for _, diskInfo := range node.DiskInfos {
diskCount += len(diskInfo.SplitByPhysicalDisk())
}
}
}
}
@@ -141,10 +79,10 @@ func (at *ActiveTopology) UpdateTopology(topologyInfo *master_pb.TopologyInfo) e
disks: make(map[uint32]*activeDisk),
}
// One activeDisk per physical disk_id (#9369): the master keys
// One activeDisk per physical disk_id: the master keys
// DiskInfos by disk type, so same-type disks must be split out.
for diskType, diskInfo := range nodeInfo.DiskInfos {
perDiskInfos := splitDiskInfoByPhysicalDisk(diskInfo)
perDiskInfos := diskInfo.SplitByPhysicalDisk()
for _, perDisk := range perDiskInfos {
disk := &activeDisk{
DiskInfo: &DiskInfo{
@@ -333,7 +271,8 @@ func (at *ActiveTopology) GetVolumeLocations(volumeID uint32, collection string)
return replicas
}
// GetECShardLocations returns the disk locations for EC shards using O(1) lookup
// GetECShardLocations returns the disk locations for EC shards using O(1) lookup.
// Each VolumeReplica.ShardIds lists the shard ids on that disk.
func (at *ActiveTopology) GetECShardLocations(volumeID uint32, collection string) []VolumeReplica {
at.mutex.RLock()
defer at.mutex.RUnlock()
@@ -345,22 +284,59 @@ func (at *ActiveTopology) GetECShardLocations(volumeID uint32, collection string
var ecShards []VolumeReplica
for _, diskKey := range diskKeys {
if disk, diskExists := at.disks[diskKey]; diskExists {
// Verify collection matches (since index doesn't include collection)
if at.ecShardMatchesCollection(disk, volumeID, collection) {
ecShards = append(ecShards, VolumeReplica{
ServerID: disk.NodeID,
DiskID: disk.DiskID,
DataCenter: disk.DataCenter,
Rack: disk.Rack,
})
}
disk, diskExists := at.disks[diskKey]
if !diskExists {
continue
}
if !at.ecShardMatchesCollection(disk, volumeID, collection) {
continue
}
shardIds := collectShardIdsForDisk(disk, volumeID, collection)
if len(shardIds) == 0 {
// ecShardMatchesCollection saw an info entry but every
// EcIndexBits is zero — phantom shard record; emitting it
// would feed an EC-cleanup source with no shard ids and
// confuse the len(ShardIds) discriminator downstream.
continue
}
ecShards = append(ecShards, VolumeReplica{
ServerID: disk.NodeID,
DiskID: disk.DiskID,
DataCenter: disk.DataCenter,
Rack: disk.Rack,
ShardIds: shardIds,
})
}
return ecShards
}
// collectShardIdsForDisk unions every matching EcIndexBits on the disk and
// expands the bitmap into shard ids, so multiple info entries for the same
// volume don't produce duplicates.
func collectShardIdsForDisk(disk *activeDisk, volumeID uint32, collection string) []uint32 {
if disk == nil || disk.DiskInfo == nil || disk.DiskInfo.DiskInfo == nil {
return nil
}
var bits erasure_coding.ShardBits
for _, ecShardInfo := range disk.DiskInfo.DiskInfo.EcShardInfos {
if ecShardInfo.Id != volumeID || ecShardInfo.Collection != collection {
continue
}
bits |= erasure_coding.ShardBits(ecShardInfo.EcIndexBits)
}
if bits == 0 {
return nil
}
ids := make([]uint32, 0, bits.Count())
for id := uint32(0); id < erasure_coding.MaxShardCount; id++ {
if uint32(bits)&(1<<id) != 0 {
ids = append(ids, id)
}
}
return ids
}
// volumeMatchesCollection checks if a volume on a disk matches the given collection
func (at *ActiveTopology) volumeMatchesCollection(disk *activeDisk, volumeID uint32, collection string) bool {
if disk.DiskInfo == nil || disk.DiskInfo.DiskInfo == nil {
+1 -1
View File
@@ -1,6 +1,6 @@
// Code generated by templ - DO NOT EDIT.
// templ: version: v0.3.977
// templ: version: v0.3.1001
package app
//lint:file-ignore SA4006 This context is only used if a nested component is present.
+1 -1
View File
@@ -1,6 +1,6 @@
// Code generated by templ - DO NOT EDIT.
// templ: version: v0.3.977
// templ: version: v0.3.1001
package app
//lint:file-ignore SA4006 This context is only used if a nested component is present.
@@ -1,6 +1,6 @@
// Code generated by templ - DO NOT EDIT.
// templ: version: v0.3.977
// templ: version: v0.3.1001
package app
//lint:file-ignore SA4006 This context is only used if a nested component is present.
@@ -1,6 +1,6 @@
// Code generated by templ - DO NOT EDIT.
// templ: version: v0.3.977
// templ: version: v0.3.1001
package app
//lint:file-ignore SA4006 This context is only used if a nested component is present.
@@ -1,6 +1,6 @@
// Code generated by templ - DO NOT EDIT.
// templ: version: v0.3.977
// templ: version: v0.3.1001
package app
//lint:file-ignore SA4006 This context is only used if a nested component is present.
+1 -1
View File
@@ -1,6 +1,6 @@
// Code generated by templ - DO NOT EDIT.
// templ: version: v0.3.977
// templ: version: v0.3.1001
package app
//lint:file-ignore SA4006 This context is only used if a nested component is present.
+1 -1
View File
@@ -1,6 +1,6 @@
// Code generated by templ - DO NOT EDIT.
// templ: version: v0.3.977
// templ: version: v0.3.1001
package app
//lint:file-ignore SA4006 This context is only used if a nested component is present.
@@ -1,6 +1,6 @@
// Code generated by templ - DO NOT EDIT.
// templ: version: v0.3.977
// templ: version: v0.3.1001
package app
//lint:file-ignore SA4006 This context is only used if a nested component is present.
+1 -1
View File
@@ -1,6 +1,6 @@
// Code generated by templ - DO NOT EDIT.
// templ: version: v0.3.977
// templ: version: v0.3.1001
package app
//lint:file-ignore SA4006 This context is only used if a nested component is present.
@@ -1,6 +1,6 @@
// Code generated by templ - DO NOT EDIT.
// templ: version: v0.3.977
// templ: version: v0.3.1001
package app
//lint:file-ignore SA4006 This context is only used if a nested component is present.
@@ -1,6 +1,6 @@
// Code generated by templ - DO NOT EDIT.
// templ: version: v0.3.977
// templ: version: v0.3.1001
package app
//lint:file-ignore SA4006 This context is only used if a nested component is present.
+11 -11
View File
@@ -1,6 +1,6 @@
// Code generated by templ - DO NOT EDIT.
// templ: version: v0.3.977
// templ: version: v0.3.1001
package app
//lint:file-ignore SA4006 This context is only used if a nested component is present.
@@ -142,7 +142,7 @@ func FileBrowser(data dash.FileBrowserData) templ.Component {
var templ_7745c5c3_Var6 templ.SafeURL
templ_7745c5c3_Var6, templ_7745c5c3_Err = templ.JoinURLErrs(dash.PUrl(ctx, fileBrowserPathURL(crumb.Path)))
if templ_7745c5c3_Err != nil {
return templ.Error{Err: templ_7745c5c3_Err, FileName: `view/app/file_browser.templ`, Line: 62, Col: 73}
return templ.Error{Err: templ_7745c5c3_Err, FileName: `view/app/file_browser.templ`, Line: 62, Col: 62}
}
_, templ_7745c5c3_Err = templ_7745c5c3_Buffer.WriteString(templ.EscapeString(templ_7745c5c3_Var6))
if templ_7745c5c3_Err != nil {
@@ -173,7 +173,7 @@ func FileBrowser(data dash.FileBrowserData) templ.Component {
var templ_7745c5c3_Var8 templ.SafeURL
templ_7745c5c3_Var8, templ_7745c5c3_Err = templ.JoinURLErrs(dash.PUrl(ctx, fileBrowserPathURL(crumb.Path)))
if templ_7745c5c3_Err != nil {
return templ.Error{Err: templ_7745c5c3_Err, FileName: `view/app/file_browser.templ`, Line: 68, Col: 73}
return templ.Error{Err: templ_7745c5c3_Err, FileName: `view/app/file_browser.templ`, Line: 68, Col: 62}
}
_, templ_7745c5c3_Err = templ_7745c5c3_Buffer.WriteString(templ.EscapeString(templ_7745c5c3_Var8))
if templ_7745c5c3_Err != nil {
@@ -216,7 +216,7 @@ func FileBrowser(data dash.FileBrowserData) templ.Component {
var templ_7745c5c3_Var10 templ.SafeURL
templ_7745c5c3_Var10, templ_7745c5c3_Err = templ.JoinURLErrs(dash.PUrl(ctx, fileBrowserPathURL("/")))
if templ_7745c5c3_Err != nil {
return templ.Error{Err: templ_7745c5c3_Err, FileName: `view/app/file_browser.templ`, Line: 86, Col: 45}
return templ.Error{Err: templ_7745c5c3_Err, FileName: `view/app/file_browser.templ`, Line: 86, Col: 53}
}
_, templ_7745c5c3_Err = templ_7745c5c3_Buffer.WriteString(templ.EscapeString(templ_7745c5c3_Var10))
if templ_7745c5c3_Err != nil {
@@ -234,7 +234,7 @@ func FileBrowser(data dash.FileBrowserData) templ.Component {
var templ_7745c5c3_Var11 templ.SafeURL
templ_7745c5c3_Var11, templ_7745c5c3_Err = templ.JoinURLErrs(dash.PUrl(ctx, fileBrowserPathURL("/buckets")))
if templ_7745c5c3_Err != nil {
return templ.Error{Err: templ_7745c5c3_Err, FileName: `view/app/file_browser.templ`, Line: 88, Col: 58}
return templ.Error{Err: templ_7745c5c3_Err, FileName: `view/app/file_browser.templ`, Line: 88, Col: 66}
}
_, templ_7745c5c3_Err = templ_7745c5c3_Buffer.WriteString(templ.EscapeString(templ_7745c5c3_Var11))
if templ_7745c5c3_Err != nil {
@@ -265,7 +265,7 @@ func FileBrowser(data dash.FileBrowserData) templ.Component {
var templ_7745c5c3_Var13 templ.SafeURL
templ_7745c5c3_Var13, templ_7745c5c3_Err = templ.JoinURLErrs(dash.PUrl(ctx, fileBrowserPathURL(data.CurrentPath)))
if templ_7745c5c3_Err != nil {
return templ.Error{Err: templ_7745c5c3_Err, FileName: `view/app/file_browser.templ`, Line: 93, Col: 77}
return templ.Error{Err: templ_7745c5c3_Err, FileName: `view/app/file_browser.templ`, Line: 93, Col: 66}
}
_, templ_7745c5c3_Err = templ_7745c5c3_Buffer.WriteString(templ.EscapeString(templ_7745c5c3_Var13))
if templ_7745c5c3_Err != nil {
@@ -278,7 +278,7 @@ func FileBrowser(data dash.FileBrowserData) templ.Component {
var templ_7745c5c3_Var14 string
templ_7745c5c3_Var14, templ_7745c5c3_Err = templ.JoinStringErrs(filepath.Base(data.CurrentPath))
if templ_7745c5c3_Err != nil {
return templ.Error{Err: templ_7745c5c3_Err, FileName: `view/app/file_browser.templ`, Line: 93, Col: 155}
return templ.Error{Err: templ_7745c5c3_Err, FileName: `view/app/file_browser.templ`, Line: 93, Col: 144}
}
_, templ_7745c5c3_Err = templ_7745c5c3_Buffer.WriteString(templ.EscapeString(templ_7745c5c3_Var14))
if templ_7745c5c3_Err != nil {
@@ -358,7 +358,7 @@ func FileBrowser(data dash.FileBrowserData) templ.Component {
var templ_7745c5c3_Var16 templ.SafeURL
templ_7745c5c3_Var16, templ_7745c5c3_Err = templ.JoinURLErrs(dash.PUrl(ctx, fileBrowserPageURL(data.CurrentPath, data.LastFileName, data.PageSize)))
if templ_7745c5c3_Err != nil {
return templ.Error{Err: templ_7745c5c3_Err, FileName: `view/app/file_browser.templ`, Line: 110, Col: 138}
return templ.Error{Err: templ_7745c5c3_Err, FileName: `view/app/file_browser.templ`, Line: 110, Col: 102}
}
_, templ_7745c5c3_Err = templ_7745c5c3_Buffer.WriteString(templ.EscapeString(templ_7745c5c3_Var16))
if templ_7745c5c3_Err != nil {
@@ -382,7 +382,7 @@ func FileBrowser(data dash.FileBrowserData) templ.Component {
var templ_7745c5c3_Var17 templ.SafeURL
templ_7745c5c3_Var17, templ_7745c5c3_Err = templ.JoinURLErrs(dash.PUrl(ctx, fileBrowserPathURL(data.ParentPath)))
if templ_7745c5c3_Err != nil {
return templ.Error{Err: templ_7745c5c3_Err, FileName: `view/app/file_browser.templ`, Line: 119, Col: 78}
return templ.Error{Err: templ_7745c5c3_Err, FileName: `view/app/file_browser.templ`, Line: 119, Col: 67}
}
_, templ_7745c5c3_Err = templ_7745c5c3_Buffer.WriteString(templ.EscapeString(templ_7745c5c3_Var17))
if templ_7745c5c3_Err != nil {
@@ -428,7 +428,7 @@ func FileBrowser(data dash.FileBrowserData) templ.Component {
var templ_7745c5c3_Var19 templ.SafeURL
templ_7745c5c3_Var19, templ_7745c5c3_Err = templ.JoinURLErrs(dash.PUrl(ctx, fileBrowserPathURL(entry.FullPath)))
if templ_7745c5c3_Err != nil {
return templ.Error{Err: templ_7745c5c3_Err, FileName: `view/app/file_browser.templ`, Line: 153, Col: 83}
return templ.Error{Err: templ_7745c5c3_Err, FileName: `view/app/file_browser.templ`, Line: 153, Col: 72}
}
_, templ_7745c5c3_Err = templ_7745c5c3_Buffer.WriteString(templ.EscapeString(templ_7745c5c3_Var19))
if templ_7745c5c3_Err != nil {
@@ -733,7 +733,7 @@ func FileBrowser(data dash.FileBrowserData) templ.Component {
var templ_7745c5c3_Var35 templ.SafeURL
templ_7745c5c3_Var35, templ_7745c5c3_Err = templ.JoinURLErrs(dash.PUrl(ctx, fileBrowserPageURL(data.CurrentPath, data.LastFileName, data.PageSize)))
if templ_7745c5c3_Err != nil {
return templ.Error{Err: templ_7745c5c3_Err, FileName: `view/app/file_browser.templ`, Line: 241, Col: 138}
return templ.Error{Err: templ_7745c5c3_Err, FileName: `view/app/file_browser.templ`, Line: 241, Col: 102}
}
_, templ_7745c5c3_Err = templ_7745c5c3_Buffer.WriteString(templ.EscapeString(templ_7745c5c3_Var35))
if templ_7745c5c3_Err != nil {
+1 -1
View File
@@ -1,6 +1,6 @@
// Code generated by templ - DO NOT EDIT.
// templ: version: v0.3.977
// templ: version: v0.3.1001
package app
//lint:file-ignore SA4006 This context is only used if a nested component is present.
+1 -1
View File
@@ -1,6 +1,6 @@
// Code generated by templ - DO NOT EDIT.
// templ: version: v0.3.977
// templ: version: v0.3.1001
package app
//lint:file-ignore SA4006 This context is only used if a nested component is present.
@@ -1,6 +1,6 @@
// Code generated by templ - DO NOT EDIT.
// templ: version: v0.3.977
// templ: version: v0.3.1001
package app
//lint:file-ignore SA4006 This context is only used if a nested component is present.
@@ -1,6 +1,6 @@
// Code generated by templ - DO NOT EDIT.
// templ: version: v0.3.977
// templ: version: v0.3.1001
package app
//lint:file-ignore SA4006 This context is only used if a nested component is present.
+1 -1
View File
@@ -1,6 +1,6 @@
// Code generated by templ - DO NOT EDIT.
// templ: version: v0.3.977
// templ: version: v0.3.1001
package app
//lint:file-ignore SA4006 This context is only used if a nested component is present.
@@ -1,6 +1,6 @@
// Code generated by templ - DO NOT EDIT.
// templ: version: v0.3.977
// templ: version: v0.3.1001
package app
//lint:file-ignore SA4006 This context is only used if a nested component is present.
+1 -1
View File
@@ -1,6 +1,6 @@
// Code generated by templ - DO NOT EDIT.
// templ: version: v0.3.977
// templ: version: v0.3.1001
package app
//lint:file-ignore SA4006 This context is only used if a nested component is present.

Some files were not shown because too many files have changed in this diff Show More